Guardrail coverage
The eight categories, every sensitivity and action setting on each side, and the readiness checks that watch them.
Reference for what the guardrails cover: the categories, the settings each side offers, and what runs no matter how the page is configured.
The eight categories
Picking an action does nothing until at least one category is switched on.
| Category | What it can do |
|---|---|
| Harassment | Stop a reply |
| Hate | Stop a reply |
| Sexual | Stop a reply |
| Violence | Stop a reply |
| Illicit or criminal | Stop a reply |
| Self-harm | Steer the next reply |
| Child safety | Steer the next reply |
| Regulated professional advice | Steer the next reply |
Why the split
The five scored categories map onto content filters that return a confidence, so a match can stop the reply at the sensitivity you chose. The other three have no score behind them, so they are on-and-off switches and they steer the agent's next reply instead. The editor groups those three under Topics the filter cannot score and marks them Steering only.
Sensitivity
Each scored category carries one of four settings, and so does the caller side's Screen for prompt attacks. The scale describes how sure the check has to be before it fires.
| Setting | Fires when | Effect |
|---|---|---|
| Off | Never | The category is not checked |
| Only when certain | The match is confident | Catches the least, and rarely stops an ordinary turn |
| Balanced | The match is reasonably likely | The middle setting |
| Catch the most | Even a weak match | Catches the most, and will occasionally stop a turn that was fine |
Setting all five at once
The Sensitivity control on the agent's card writes all five scored categories together.
| Sensitivity | Each of the five becomes |
|---|---|
| Off | Off |
| Let most through | Only when certain |
| Balanced | Balanced |
| Catch the most | Catch the most |
Change one category on its own and the control reads Custom. Custom is a readout of what the five categories say, not a setting you can pick.
Actions
One action covers each side, and each side offers its own list. Watch only is the recommended starting point on both: it records what would have been caught and changes nothing the caller hears.
When a reply is caught
| Action | What happens |
|---|---|
| Watch only | The reply is spoken unchanged and the hit is recorded |
| Say a safe line instead | The reply is cut and the agent speaks the safe line |
| Put the caller through to a person | The agent speaks the safe line, then transfers the caller. Needs a transfer destination |
| End the call | The agent speaks the safe line, then the call ends once the line finishes |
The safe line is the one you write under What the agent says instead. Left empty, the agent speaks the platform line:
Sorry, I can't help with that. Is there something else I can do for you?
When a caller turn is caught
Two actions here rather than four. The caller's own words are not something the agent can replace or hand to a person, so the only lever on this side is the attempt count the built-in screen already keeps.
| Action | What happens |
|---|---|
| Watch only | The hit is recorded and nothing about the call changes |
| Count it as an attempt | The turn counts one attempt toward the three-attempt limit, and the third attempt ends the call |
Counting an attempt gives the graded check a share of the built-in screen's count, so three attempts spread across the two still ends the call. The built-in screen counts its own matches whichever action you pick.
What runs regardless of your settings
| Piece | When it runs |
|---|---|
| The injection screen | Every call, whether or not guardrails are configured. It cannot be switched off |
| The three-attempt sign-off and hang-up | The same calls |
Both cover real inbound calls and in-app test calls.
Where the checks run
Both sides use the same moderation service, running in Sydney, and there is nothing to configure. A check is capped at two seconds, after which the turn goes ahead unchecked, and any other failure resolves the same way. See Data residency.
What the validator checks
Three readiness checks watch this page. All three are risks rather than errors, so none of them blocks publishing. They catch settings that read as protective in the editor and check nothing.
| Fires when | Message |
|---|---|
| An action is set for the agent's replies and no category or topic is switched on | "The guardrails act on what the agent says, but no category or topic is switched on, so nothing is checked. Switch one on under Guardrails, or set When a reply is caught to Watch only." |
| An action is set for the caller's turns and the prompt-attack check is off | "The guardrails act on what the caller says, but Screen for prompt attacks is off, so nothing is checked. Switch it on under Guardrails, or set When a caller turn is caught to Watch only." |
| The agent's replies hand the caller to a person and no transfer action has a number to dial | "The guardrails hand the caller to a person, but no transfer_to_human action has a number to dial, so the agent reads the safe line instead and no one is put through. Set a transfer destination under Tools, or pick a different guardrail action." |
All three read the action you chose yourself rather than the one the page starts on, and Watch only never triggers the first two, because it is not being asked to act on anything.