Quality and safety

Scoring

How every call is scored, what it costs, and where the scores appear.

Scoring decides what counts as a good call. The same checks that grade an evaluation also grade real customer calls, so a failing check on Monday's calls is directly comparable to a failing check in a test.

Every check is on unless you switch it off. A new agent is fully scored from its first call, with no setup.

What runs after a call

When a call ends, three passes run in order.

  1. Deterministic checks. Measured from the transcript, the timings and the recording. No language model is involved, so they cost nothing and always run.
  2. Judge metrics. A language model reads the transcript and answers one question per metric. Only the metrics active on this agent are judged, so an agent with no active judge metrics makes no model calls and is charged nothing for this pass.
  3. The call summary and verdict. One cheap model pass that writes the summary, sentiment, outcome and extracted facts, and rolls the call into caller memory when memory is on. This pass is the one the Grade every call switch controls.

A text or bridge evaluation scores the conversation but fires nothing downstream. A phone evaluation is an ordinary inbound call, so it runs the whole chain.

The scoring page

The page has three cards:

  • Call summary and verdict. One switch, Grade every call. When off, calls are still recorded, transcribed and scored by the checks below, but no summary, sentiment or outcome is written.
  • Checks. The checks you can switch off, each with its price per call and a note explaining when switching it off makes sense.
  • Always on. Paid checks with no switch, because a failure means something went wrong rather than that this agent works differently. The card heading shows their combined price per graded call.

Free checks are not listed, which is why switching everything visible off does not switch scoring off. A check that does not apply to a call sits it out rather than failing it.

How a call is judged successful

There is no separate success rule to author. Success is computed from the critical checks alone:

  1. A check is critical if its definition says it affects call success. Four do today: task success, hallucination, tool call success and disclosure spoken.
  2. A critical check that fails makes the call a failure.
  3. A critical check that was not computed for this call is skipped, and counts as a pass rather than blocking.
  4. A call with no critical checks, or with all of them passing, is a success.

A worked example

task success        pass
hallucination       switched off, so it cannot fail the call
tool call success   not computed, so skipped and counted as a pass
disclosure spoken   pass
                    => success

Switching a critical check off

Rule 3 catches people out. Switching a critical check off does not blank the verdict: it removes that check's ability to fail the call, so the call reports success without it.

In an evaluation, the same rule acts as a veto: a critical failure turns a would-be success into a failure, never the other way around.

What scoring costs

Scoring is billed per active judge metric on the call, never as a flat rate.

  • Deterministic checks are free.
  • Each judge metric has its own price, shown on the scoring page next to its switch, and priced at cost.
  • Only a check that actually produced a result is billed. Errored, untriggered or skipped checks are not.

See Scoring checks for the full catalogue and which single deterministic check is not free.

A scenario scoring three judge metrics costs materially less than one scoring twenty. The run estimate before launching an evaluation adds up the metrics active on that scenario rather than applying a flat price. Treat any estimate as an estimate: billing follows the metered minutes and the metrics that actually ran.

Switching off a judge metric that does not suit the line removes its price from every call. That is what the Checks card is for.

Where scores appear

Scores land on the call, not in a separate report.

The call page opens with a headline strip: task success, hallucination, CSAT and latency. Metrics that errored or were skipped are left out. Below it, a one-line list of failing checks links to their cards. A metrics panel below groups results into Judge metrics, Quality checks and Audio quality, with the judge's reasoning and a thumbs-down to flag a verdict you disagree with. Evaluate metrics and Re-evaluate metrics score a call on demand.

On the calls list, each row lifts task success, hallucination and latency, with a metrics pill whose tooltip names up to five failing checks.

On this page