AI Answer Safety

Every AI answer passes through an AWS Bedrock Guardrail, in Australia, before the visitor sees it. You choose one of three levels for your organisation in Settings. The level applies to every site and search group in the organisation.

What every level does

  • Prompt attacks are blocked. Questions such as "ignore all previous instructions" get a refusal, not an answer.
  • Harmful content is blocked. Hate, insults, sexual content, violence and misconduct, in questions and answers.
  • Personal details in questions are masked. Tax file numbers, Medicare numbers, card numbers, email addresses and phone numbers are replaced with placeholders such as {AU Tax File Number} before the question reaches the AI, your analytics or the logs. The visitor still gets an answer.
  • Identity numbers are masked in answers. Your links, email addresses, phone numbers and staff names are never removed: they come from your own pages.
  • Safety questions are never refused. A question about suicide, self-harm, family violence or sexual assault is never blocked or replaced with "no reliable answer". Your own answer is shown with a line in front telling the person to contact local emergency services if they are in danger. If there is no usable answer, a short general support message is shown. We do not add phone numbers or service names, because the right ones depend on where your visitor is: put your recommended services in your own content and the answer will include them.

The three levels

CheckOFFICIALOFFICIAL:SensitivePROTECTED
Prompt attacks, harmful contentYesYesYes
TFN, Medicare, card, email, phone masked in questionsYesYesYes
Names and IP addresses masked in questionsNoYesYes
Licence, passport and vehicle identifiers masked in questionsNoNoYes
Classification markings (e.g. TOP SECRET) blocked in questionsNoYesYes
Questions on intelligence sources and cabinet documents refusedNoNoYes
Grounding checkNoYesYes

The grounding check

On OFFICIAL:Sensitive and PROTECTED, each answer is scored against the pages it was drawn from. An answer your content does not support is replaced with "I couldn't give a reliable answer from this website's content", and the visitor still sees the search results.

We set the threshold by measurement, not by guess. On real questions against live sites, invented answers scored 0.00 to 0.05 and correct answers scored 0.02 to 0.99 (median 0.86). The threshold is 0.06: in our test it caught every invented answer and replaced 1 correct answer in 48, a long step-by-step technical explanation. A higher threshold catches no more invented answers and replaces more correct ones, so PROTECTED uses the same threshold and is stricter on questions instead.

Without the grounding check, answers are still limited to your indexed content: the AI is given only the passages found in your site, and when nothing is found it says so.

Choosing a level

  • OFFICIAL suits most public websites. It is the default.
  • OFFICIAL:Sensitive suits sites where a wrong answer has a real cost, such as health, legal or benefits information.
  • PROTECTED adds stricter screening of questions for sites in security-related portfolios.

If your organisation also uses Quant AI through the Quant Cloud dashboard, search follows the matching level from your AI governance settings until you choose one here.