> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usetuner.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Replays

> Run a call through your agent again to see what changed. Replay a real call from your Dataset with the exact same caller audio, or replay a saved simulation scenario.

## What is a replay?

When you change your agent, you need to know what got better and what broke. A replay runs a call through your agent again, and Tuner scores it with your evals, so you can compare the result with the run before.

Use replays to:

* Reproduce a call that failed.
* Prove that a fix worked.
* Compare two setups, like a new model or provider, on the same calls.

<Frame>
  <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/fix-before-after.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=4c6754303d4679a0993f0d7884486843" alt="The same Dataset call replayed three times before a fix, failing every time, and three times after it, passing every time" width="1400" height="1444" data-path="images/datasets/fix-before-after.png" />
</Frame>

***

## How it works

Every replay follows the same three steps:

1. Pick what to replay: a call from your Dataset, or a saved simulation scenario.
2. Tuner calls your agent and plays the call, like a real customer would.
3. Tuner scores each call with your evals, so you can see what changed.

### Two kinds of replay

What differs is the caller.

<CardGroup cols={2}>
  <Card title="Replay a Dataset call" icon="database">
    The caller is a recording of a real call, played back exactly as it was said. Best for reproducing a failure.
  </Card>

  <Card title="Replay a simulation" icon="flask">
    The caller is a simulated customer following a saved scenario. Best for re-checking a scenario after every change.
  </Card>
</CardGroup>

### How they compare

| | Replay a Dataset call | Replay a simulation |
| - | - | - |
| **Starts from** | A real call in your [Dataset](/docs/datasets/overview) | A scenario saved from a simulation run, or one you wrote |
| **The caller is** | The real caller's recorded audio | A simulated caller following the scenario |
| **Same conversation every run?** | Yes, the caller audio is identical | Similar, not identical: the caller improvises |
| **Language, noise, test profile** | Whatever the real call had | Chosen on every run |
| **Agents** | Inbound only | Inbound and outbound |
| **Where** | **Replays** tab | **Scenarios** tab, or **From library** in Run Simulation |

### Step by step

Pick a kind in any step, and the other steps switch with it.

<Steps>
  <Step title="Pick what to replay">
    <Tabs>
      <Tab title="Replay a Dataset call">
        Choose which calls from your Dataset to play to your agent.

        * Go to **Simulations > Replays** and click **New Replay**.
        * Click **Change** next to **Calls** and tick up to 10 Dataset calls. Calls that are still extracting, or that failed, can't be picked.
        * Dataset replays run against **inbound** agents on Vapi, Retell, Dograh, LiveKit, Pipecat, or a Custom API integration. Set up [inbound simulation](/docs/simulation/inbound/setup) first, and add at least one **Ready** call to your [Dataset](/docs/datasets/add-calls).

        <Frame>
          <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/simulation/new-replay.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=c37c891883c371ec6c98af1b2bd9dfc6" alt="The New Replay dialog with three Dataset calls selected, each replayed three times" width="1040" height="1936" data-path="images/simulation/new-replay.png" />
        </Frame>
      </Tab>

      <Tab title="Replay a simulation">
        Save the scenario behind a call you want to run again.

        * On the **Runs** tab, expand a generated run that is **Complete** or **Partial**, and tick the calls you want to keep. A bookmark marks calls that are already in your library.
        * Click **Save to library**, give each scenario a name, and optionally add **Tags for all**. Then click **Save**.
        * Tuner saves each call's **Situation**, **Goal**, **Stop conditions**, type (Routine or Pressure), **Intent**, and **Target eval**. It doesn't save the language, background noise, or test profile. You can also write a scenario yourself with **New scenario** on the **Scenarios** tab.

        <Frame>
          <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/simulation/save-to-library.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=e3ea4ef42549fe262a5176d2eab07c23" alt="A finished simulation run with two calls selected and the Save to library button" width="1480" height="976" data-path="images/simulation/save-to-library.png" />
        </Frame>
      </Tab>
    </Tabs>
  </Step>

  <Step title="Run it">
    <Tabs>
      <Tab title="Replay a Dataset call">
        Tuner calls your agent and plays each caller turn from the recorded call.

        * Set **Times to replay each call** (1 to 20, at most 20 calls in total) and **Max call duration** (1 to 12 minutes). **Longest possible call** warns you if a call won't fit.
        * Click **Run**. The run appears on the **Replays** tab and fills in as calls finish.
        * Tuner waits for your agent's greeting, plays each caller turn in full, even if your agent talks over it, and waits for your agent's reply before the next turn. After the last turn it waits for one more reply, then hangs up. A call also ends if your agent hangs up or it reaches the max call duration.
        * The caller never reacts to what your agent says. They say what they said in the original call. See [The caller doesn't adapt](/docs/datasets/use-cases/capturing-failures#the-caller-doesnt-adapt).
        * To change how long Tuner waits for your agent, see [Turn-taking settings](#turn-taking-settings).
      </Tab>

      <Tab title="Replay a simulation">
        Tuner places a new simulated call to your agent for each scenario.

        * On the **Scenarios** tab or a scenario's page, click **Run this scenario**.
        * Or, in **Run Simulation**, click **Change** next to **Scenarios**, choose **From library**, and pick up to 20 scenarios.
        * Language, accent, background noise, test profile, and max call duration work as in any [simulation run](/docs/simulation/overview).
        * The simulated caller follows the scenario but improvises, so the wording and length of each call can differ. To replay the exact same caller, replay a Dataset call.
      </Tab>
    </Tabs>
  </Step>

  <Step title="Read the results">
    <Tabs>
      <Tab title="Replay a Dataset call">
        Every replay is graded, so you can see what changed.

        <Frame>
          <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/simulation/replays-tab.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=ede1ba6f055a2469f1ba2b436ac9b542" alt="A replay run with nine calls: six passed and three failed" width="1400" height="1850" data-path="images/simulation/replays-tab.png" />
        </Frame>

        * Each run card shows how many calls **passed** and **failed**, and the calls, repeats, and turn-taking settings it used.
        * Every call is graded against all of your agent's **Pass/Fail** evals. A green check means none failed. A red X means at least one failed, or the call didn't run. If your agent has no Pass/Fail evals, calls count as **ran** and the run measures latency only.
        * Click a row to open the call. A **Replay call** banner names the Dataset call it came from, and **Replay details** shows the source call and the exact settings used. The recording is stereo: caller on the left, agent on the right.
        * To compare two setups, keep the same calls, repeats, max call duration, and turn-taking settings, then compare the run cards and open matching calls. There's no side-by-side view yet. See [Benchmarking](/docs/datasets/use-cases/benchmarking#compare-two-setups).

        <Frame>
          <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/simulation/replay-call-banner.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=26a6bb9200eeafed127a2f29dd8f84b5" alt="The Replay call banner on a call's details page" width="1360" height="196" data-path="images/simulation/replay-call-banner.png" />
        </Frame>
      </Tab>

      <Tab title="Replay a simulation">
        Every simulated call is scored, and each scenario keeps its history.

        <Frame>
          <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/simulation/scenario-history.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=4a1bf79fbf29a086d3dc3396258aa1e5" alt="A saved scenario with its definition and a run history of three passes out of the last five runs" width="1560" height="1354" data-path="images/simulation/scenario-history.png" />
        </Frame>

        * The run appears on the **Runs** tab and is scored against your Evals and Intents. A Pressure scenario with a **Target eval** passes or fails on that eval.
        * Each scenario's page keeps a **Run history**: how it did on its last runs, and a link to every call. Run it again after every change to see whether it still passes.
      </Tab>
    </Tabs>
  </Step>
</Steps>

***

## Turn-taking settings

This is optional, and it applies to Dataset replays. Expand **Advanced turn-taking** in **New Replay** to change how long Tuner waits for your agent.

<Frame>
  <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/simulation/replay-turn-taking.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=1b9439f5ab45f469a34168f2effee594" alt="Advanced turn-taking settings with the default values" width="896" height="978" data-path="images/simulation/replay-turn-taking.png" />
</Frame>

| Setting | Default | What it does |
| - | - | - |
| **Initial agent wait (ms)** | 12,000 | How long to wait for the agent's greeting before the first caller turn |
| **Silence threshold (ms)** | 700 | How much silence counts as the agent having finished |
| **Minimum speech (ms)** | 250 | Agent speech shorter than this doesn't count as a reply |
| **Inter-turn delay (ms)** | 300 | The pause before the next caller turn plays |
| **Max wait (ms)** | 15,000 | The longest Tuner waits for the agent before moving on |

***

## Limits

| Limit | Replay a Dataset call | Replay a simulation |
| - | - | - |
| Per run | 1 to 10 calls, each 1 to 20 times, at most 20 calls | 1 to 20 scenarios |
| Max call duration | 1 to 12 minutes | 1 to 12 minutes |
| Runs in progress per agent | 1, shared by all simulation runs | 1, shared by all simulation runs |

Both are also available through the public API, in the **Simulation** group of the API Reference.

### Next steps

<CardGroup cols={2}>
  <Card title="Introduction to Datasets" icon="database" href="/docs/datasets/overview">
    Save the calls you want to replay: failures, hard callers, and edge cases.
  </Card>

  <Card title="Introduction to Call Simulation" icon="flask" href="/docs/simulation/overview">
    Generate new scenarios to find failures you haven't seen yet.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.