Quality and safety

Scoring checks

The two kinds of check, the catalogue the platform runs on a call, which checks you can switch off, and the shapes a metric result takes.

Metrics are defined once for the workspace and reused across agents. A metric with no explicit setting on an agent is active on that agent. See Scoring for how checks run and where results land.

Two kinds of check

DeterministicJudge
How it decidesMeasured from the transcript, timings or audioA language model answers a question about the transcript
CostFree, with one exception belowPriced per metric per call
DeterminismSame call, same answer, every timeSame call, near-identical answers, not guaranteed identical
Good forLatency, dead air, talk ratio, whether a tool succeededWhether the caller got what they rang for, whether the agent invented something

The one that is not free

Transcription errors is the one deterministic check that costs money: it runs the recording through a second speech-to-text pass to compare against. The other audio-analysis checks are free local signal processing.

Checks you can switch off

Each is on by default and sits on the Checks card of the agent's scoring page. Switch one off when the agent's design makes that check misleading.

CheckSwitch it off when
VerbosityThe agent must read out compliance wording, instructions or detailed quotes
Talk ratioThe agent reads a script or runs a survey, so the caller is not expected to talk much
Natural turn-takingThe agent is deliberately brisk or menu-driven
CSATCallers arrive unhappy by nature of the line, such as collections or complaints
SentimentSame as CSAT: it marks the agent down for delivering unwelcome news
Eager tool before clarifyThe agent deliberately looks the caller up on their number before asking
Appropriate terminationA screening agent takes a message and hangs up on purpose
Appropriate termination by userYou do not want a check that grades the caller rather than the agent
RelevancyThe agent deliberately refuses to be drawn off script
Unnecessary repetition scoreThe agent reads details back to confirm them, such as an address or card number
Response consistencyAn answer legitimately depends on stock, price or availability that moved mid-call
Task successThe line is not meant to resolve anything, for example one that only takes a message
HallucinationThe agent answers from live tool data the check cannot see

The last two are critical checks, so switching one off changes how a call is judged successful.

The rest of the catalogue

Judge metrics

  • Search relevance
  • Topic of call
  • Drop-off stage
  • Voicemail detection
  • Three human-conversation checks that only ever score discovery calls where a person answered

Deterministic checks

  • Tool call success
  • Tool retry stall
  • Dead air
  • Latency at the median
  • Talk ratio
  • Silence detection
  • Thinking leak
  • Unnecessary repetition count
  • Promised but not sent
  • Call disposition
  • Cost per call
  • Disclosure spoken
  • A heuristic search-relevance check

Audio checks

Transcription errors (word error rate), average pitch, gibberish, and agent audio dropouts. These need a recording, so they only run on voice calls with recording on.

What a result looks like

  • A pass or fail verdict.
  • A number with a unit and a direction that counts as better.
  • One of a fixed set of values, such as call disposition or topic of call.

Numeric metrics carry bands that decide whether a value reads as good, borderline or failing.

Checks that sit a call out

Discovery-only checks never score an ordinary call. Checks that assume an agent answered (task success, tool call success) are skipped on a discovery call rather than failed.

On this page