Skip to main content
Evals go deeper than metrics — each one runs a purpose-built AI judge against the transcript to score a specific behavior. Tuner ships with 21 of them across four categories, optimized using patterns from real-world voice AI deployments. Enable any combination with a single click, or write your own for behavior specific to your use case. Predefined evals library in Tuner

Accuracy

Checks that the agent captured and confirmed information correctly.

Actions & tools

Validates that tools and workflows were invoked correctly.

Conversation quality

Measures how well the agent communicates throughout the call.

Safety & compliance

Ensures the agent stays within required guardrails.

Adding them to your agent

Pick from the library

Open the evals panel, browse by category, and add any predefined eval with one click. Each ships with a prompt optimized on real voice AI use cases.

Optimized AI judges

Every predefined eval is backed by a judge tuned on patterns from actual deployments — not generic LLM prompts. Fewer false positives, scores you can act on.

Write your own

Have a behavior specific to your use case? Describe it in plain language and Tuner runs it as a first-class judge alongside these.
Predefined evals run on every production call and every simulated call — the same criteria score both.

Next steps

Agents with Dynamic Instructions

How to use call to evaluate dynamic agents.

Red flags and labels

Turn eval results into tags you can filter and alert on.