> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usetuner.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Capture failures and prove the fix

> Save the call that failed, reproduce it with the exact same caller audio, prove your fix on it, and keep it as a test that runs on every change.

Voice agents rarely fail on the calls you tested. They fail on the ones you didn't see coming. Tuner already flags the failure and shows you where the call went wrong. The hard part is proving your fix works on that exact call.

Calling your agent from a quiet office doesn't prove it. A generated simulation gets close, but it can't reproduce the exact voice, noise, and timing that broke it. The only real test is the call that failed, so save it to your Dataset and replay it.

***

## The calls you can't make twice

| What broke | What it sounds like |
| - | - |
| **Speech-to-text** | A caller on speakerphone in a moving car, and the agent hears "Tuesday" instead of "Thursday." Or an accent your speech-to-text struggles with, so the agent keeps asking the caller to repeat themselves. |
| **Turn detection** | A caller who talks over the agent or pauses mid-sentence, and the agent cuts them off. |
| **Tool calls** | A caller who reads out a date in a way your agent didn't expect, and the booking tool call fails. Or a tool that errors, and the agent carries on as if it worked. |
| **Logic** | The agent makes up a policy, skips a step in the workflow, or lands the wrong Call Outcome. |
| **Slow responses** | The caller waits in silence after "Let me check that for you" while a tool runs. |

You'll find these calls in production:

* Calls with a **failed eval**, like a made-up policy or a missed confirmation
* Calls tagged with a [red flag](/docs/evaluation/red-flags)
* Calls whose **Call Outcome** was a failure or an escalation you didn't want
* Calls where the agent misread what the caller wanted, their **Intent**
* Calls a customer complained about

***

## From failure to verified fix

<Steps>
  <Step title="Spot the failure">
    Open the call and find where it went wrong: the eval that failed, the tool call that errored, or the turn the agent misheard.

    <Frame>
      <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/failure-call.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=ae373982a097fc9ff4343ade699d32a4" alt="A production call where the agent heard “Friday the fourth” instead of “the fourteenth”, the reschedule tool call failed, and the eval failed" width="1360" height="1466" data-path="images/datasets/failure-call.png" />
    </Frame>
  </Step>

  <Step title="Save the call">
    On **Simulations > Dataset**, click **Add call** and pick it from your production calls, or upload the recording. Background noise and all. Name it after the problem, for example **\[tool] Reschedule, road noise**.

    <Frame>
      <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/add-call-production.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=4535db08dbfd88dd892a6c96527a2641" alt="The Add call dialog with the failed production call selected" width="1480" height="1926" data-path="images/datasets/add-call-production.png" />
    </Frame>
  </Step>

  <Step title="Reproduce it">
    Replay it 3 times against your current agent before you change anything. If it fails every time, it's a real problem, not a one-off glitch, and you have a test that fails for the right reason. If it fails once in three, it's intermittent, and a single passing replay after your fix won't prove much.

    <Frame>
      <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/fix-reproduce.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=cc44234b954d4f78b1f5c98835472267" alt="The failed call replayed three times before any change, failing all three times" width="1400" height="676" data-path="images/datasets/fix-reproduce.png" />
    </Frame>
  </Step>

  <Step title="Fix your agent">
    Change the prompt, try a different speech-to-text model, rewrite a tool's description, or clean up the date before the tool call. Whatever you think fixes it.
  </Step>

  <Step title="Replay it again">
    Same caller, same noise, same words, the same number of times. If the call now passes your evals, you watched the fix work on the call that broke.

    <Frame>
      <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/fix-verified.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=b7985d0a8cfc17d9058e33ccb73053e7" alt="The same call replayed three times after the fix, passing all three times" width="1400" height="674" data-path="images/datasets/fix-verified.png" />
    </Frame>
  </Step>
</Steps>

***

## How a replay is graded

* **Evals:** each replay is graded against **all** of your agent's **Pass/Fail** evals. Score (1 to 5) evals aren't graded on replays. See [Custom evals](/docs/evaluation/custom-evals).
* **Intents and Call Outcomes:** classified on each replayed call like on any other call. See [Call classification](/docs/evaluation/call-classification).

On the **Replays** tab, each call shows a green check when none of its evals failed, or a red X when at least one did. Open a call to see which eval failed and why.

<Frame>
  <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/replay-evals.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=13b58265619f0271e188f2e1c35c28e4" alt="The Evals of a replayed call, with one eval failing and the evidence for it" width="1240" height="966" data-path="images/datasets/replay-evals.png" />
</Frame>

<Note>
  If your agent has no Pass/Fail evals, a replay only measures latency, and each call shows as **ran** instead of passed or failed. Add a Pass/Fail eval that catches the failure before you replay it.
</Note>

### Write evals that catch it

* **Make each eval about one behaviour**, so a failure tells you exactly what to fix.
* **Anchor evals on the agent's words and actions**, like a confirmation, a refusal, or a transfer, not on what the caller said after.
* **For tools, test how the tool was used.** For example: "The agent called `book_appointment` before confirming a time to the caller," or "If a tool returned an error, the agent didn't tell the caller the action succeeded."

See [Custom evals](/docs/evaluation/custom-evals) and the [business evals framework](/docs/guides/business-evals-framework) for how to write them.

***

## Check a tool call on a replay

Open a replayed call from the **Replays** tab. Tool calls appear in the transcript, in order, between the turns. Each one shows its **name** and how long it took, and a tool call that returned an **error** shows in red with a warning icon. Click a tool call to expand its **Input** and **Output** (or **Error**).

<Frame>
  <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/datasets/replay-tool-call.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=27c4a71f1502ffdccb0c6376a94c476c" alt="A replayed call where reschedule_appointment() returned an error, with its input and error expanded" width="1400" height="1550" data-path="images/datasets/replay-tool-call.png" />
</Frame>

If you send OpenTelemetry spans, the **Traces** tab shows what ran inside each tool call. See [Traces](/docs/observability/traces).

<Warning>
  **Replays call your real agent, so your agent calls your real tools.** A replayed refund request can issue a real refund. Before replaying calls that create bookings, move money, or change customer records, point your agent at test or sandbox endpoints, or use test accounts.
</Warning>

Tools also return live data, and a replay can happen weeks after the original call, when the slot the caller asked for is already taken. Point tools at **test data** that doesn't change between runs, or write evals about **how** the agent used the tool rather than the exact result.

***

## Your Dataset becomes your test suite

That call stays in your Dataset, and so does every hard call you add: the noisy ones, the strong accents, the interruptions, the edge cases. Replay the whole set with every prompt update, model swap, or deployment. If an old failure comes back, you'll see it before your callers hear it.

<Frame>
  <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/datasets/suite-run.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=71778c7ea4fc911c8c322e2cc019e51b" alt="The whole Dataset replayed as a test suite: six calls passed and an old failure came back" width="1400" height="1068" data-path="images/datasets/suite-run.png" />
</Frame>

Add a few **golden calls** too: ordinary calls your agent handles well today. They're how you notice when a fix for one problem quietly breaks something that used to work.

***

## The caller doesn't adapt

A replay plays the caller's recorded turns in order, whatever your agent says. If your fixed agent now answers differently, the caller still says what they said in the original call. For example, they may still say "Great, thanks" after your agent has refused a request.

This is what keeps replays comparable, but it affects which calls make good tests:

* **Choose calls where the failure happens early**, before the conversation depends on what the agent said.
* **Keep calls short and focused** on one behaviour.
* **Judge the agent's turns, not the caller's.** Write evals about what the agent did, like "didn't promise a refund", rather than about how the conversation ended.

For conversations that should branch on the agent's answers, use a [generated simulation](/docs/simulation/overview) instead. Its caller reacts to your agent in real time.

### Next steps

<CardGroup cols={2}>
  <Card title="Benchmarking" icon="scale-balanced" href="/docs/datasets/use-cases/benchmarking">
    Pick your stack on your own calls.
  </Card>

  <Card title="Dataset best practices" icon="lightbulb" href="/docs/datasets/best-practices">
    Naming, repeats, and keeping your Dataset useful over time.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.