Scoring checks
The two kinds of check, the catalogue the platform runs on a call, which checks you can switch off, and the shapes a metric result takes.
Metrics are defined once for the workspace and reused across agents. A metric with no explicit setting on an agent is active on that agent. See Scoring for how checks run and where results land.
Two kinds of check
| Deterministic | Judge | |
|---|---|---|
| How it decides | Measured from the transcript, timings or audio | A language model answers a question about the transcript |
| Cost | Free, with one exception below | Priced per metric per call |
| Determinism | Same call, same answer, every time | Same call, near-identical answers, not guaranteed identical |
| Good for | Latency, dead air, talk ratio, whether a tool succeeded | Whether the caller got what they rang for, whether the agent invented something |
The one that is not free
Transcription errors is the one deterministic check that costs money: it runs the recording through a second speech-to-text pass to compare against. The other audio-analysis checks are free local signal processing.
Checks you can switch off
Each is on by default and sits on the Checks card of the agent's scoring page. Switch one off when the agent's design makes that check misleading.
| Check | Switch it off when |
|---|---|
| Verbosity | The agent must read out compliance wording, instructions or detailed quotes |
| Talk ratio | The agent reads a script or runs a survey, so the caller is not expected to talk much |
| Natural turn-taking | The agent is deliberately brisk or menu-driven |
| CSAT | Callers arrive unhappy by nature of the line, such as collections or complaints |
| Sentiment | Same as CSAT: it marks the agent down for delivering unwelcome news |
| Eager tool before clarify | The agent deliberately looks the caller up on their number before asking |
| Appropriate termination | A screening agent takes a message and hangs up on purpose |
| Appropriate termination by user | You do not want a check that grades the caller rather than the agent |
| Relevancy | The agent deliberately refuses to be drawn off script |
| Unnecessary repetition score | The agent reads details back to confirm them, such as an address or card number |
| Response consistency | An answer legitimately depends on stock, price or availability that moved mid-call |
| Task success | The line is not meant to resolve anything, for example one that only takes a message |
| Hallucination | The agent answers from live tool data the check cannot see |
The last two are critical checks, so switching one off changes how a call is judged successful.
The rest of the catalogue
Judge metrics
- Search relevance
- Topic of call
- Drop-off stage
- Voicemail detection
- Three human-conversation checks that only ever score discovery calls where a person answered
Deterministic checks
- Tool call success
- Tool retry stall
- Dead air
- Latency at the median
- Talk ratio
- Silence detection
- Thinking leak
- Unnecessary repetition count
- Promised but not sent
- Call disposition
- Cost per call
- Disclosure spoken
- A heuristic search-relevance check
Audio checks
Transcription errors (word error rate), average pitch, gibberish, and agent audio dropouts. These need a recording, so they only run on voice calls with recording on.
What a result looks like
- A pass or fail verdict.
- A number with a unit and a direction that counts as better.
- One of a fixed set of values, such as call disposition or topic of call.
Numeric metrics carry bands that decide whether a value reads as good, borderline or failing.
Checks that sit a call out
Discovery-only checks never score an ordinary call. Checks that assume an agent answered (task success, tool call success) are skipped on a discovery call rather than failed.