Quality and safety

Guardrail coverage

The eight categories, every sensitivity and action setting on each side, and the readiness checks that watch them.

Reference for what the guardrails cover: the categories, the settings each side offers, and what runs no matter how the page is configured.

The eight categories

Picking an action does nothing until at least one category is switched on.

CategoryWhat it can do
HarassmentStop a reply
HateStop a reply
SexualStop a reply
ViolenceStop a reply
Illicit or criminalStop a reply
Self-harmSteer the next reply
Child safetySteer the next reply
Regulated professional adviceSteer the next reply

Why the split

The five scored categories map onto content filters that return a confidence, so a match can stop the reply at the sensitivity you chose. The other three have no score behind them, so they are on-and-off switches and they steer the agent's next reply instead. The editor groups those three under Topics the filter cannot score and marks them Steering only.

Sensitivity

Each scored category carries one of four settings, and so does the caller side's Screen for prompt attacks. The scale describes how sure the check has to be before it fires.

SettingFires whenEffect
OffNeverThe category is not checked
Only when certainThe match is confidentCatches the least, and rarely stops an ordinary turn
BalancedThe match is reasonably likelyThe middle setting
Catch the mostEven a weak matchCatches the most, and will occasionally stop a turn that was fine

Setting all five at once

The Sensitivity control on the agent's card writes all five scored categories together.

SensitivityEach of the five becomes
OffOff
Let most throughOnly when certain
BalancedBalanced
Catch the mostCatch the most

Change one category on its own and the control reads Custom. Custom is a readout of what the five categories say, not a setting you can pick.

Actions

One action covers each side, and each side offers its own list. Watch only is the recommended starting point on both: it records what would have been caught and changes nothing the caller hears.

When a reply is caught

ActionWhat happens
Watch onlyThe reply is spoken unchanged and the hit is recorded
Say a safe line insteadThe reply is cut and the agent speaks the safe line
Put the caller through to a personThe agent speaks the safe line, then transfers the caller. Needs a transfer destination
End the callThe agent speaks the safe line, then the call ends once the line finishes

The safe line is the one you write under What the agent says instead. Left empty, the agent speaks the platform line:

Sorry, I can't help with that. Is there something else I can do for you?

When a caller turn is caught

Two actions here rather than four. The caller's own words are not something the agent can replace or hand to a person, so the only lever on this side is the attempt count the built-in screen already keeps.

ActionWhat happens
Watch onlyThe hit is recorded and nothing about the call changes
Count it as an attemptThe turn counts one attempt toward the three-attempt limit, and the third attempt ends the call

Counting an attempt gives the graded check a share of the built-in screen's count, so three attempts spread across the two still ends the call. The built-in screen counts its own matches whichever action you pick.

What runs regardless of your settings

PieceWhen it runs
The injection screenEvery call, whether or not guardrails are configured. It cannot be switched off
The three-attempt sign-off and hang-upThe same calls

Both cover real inbound calls and in-app test calls.

Where the checks run

Both sides use the same moderation service, running in Sydney, and there is nothing to configure. A check is capped at two seconds, after which the turn goes ahead unchecked, and any other failure resolves the same way. See Data residency.

What the validator checks

Three readiness checks watch this page. All three are risks rather than errors, so none of them blocks publishing. They catch settings that read as protective in the editor and check nothing.

Fires whenMessage
An action is set for the agent's replies and no category or topic is switched on"The guardrails act on what the agent says, but no category or topic is switched on, so nothing is checked. Switch one on under Guardrails, or set When a reply is caught to Watch only."
An action is set for the caller's turns and the prompt-attack check is off"The guardrails act on what the caller says, but Screen for prompt attacks is off, so nothing is checked. Switch it on under Guardrails, or set When a caller turn is caught to Watch only."
The agent's replies hand the caller to a person and no transfer action has a number to dial"The guardrails hand the caller to a person, but no transfer_to_human action has a number to dial, so the agent reads the safe line instead and no one is put through. Set a transfer destination under Tools, or pick a different guardrail action."

All three read the action you chose yourself rather than the one the page starts on, and Watch only never triggers the first two, because it is not being asked to act on anything.

On this page