> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usetuner.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Introduction to Datasets

> Save real calls and replay them against your agent with the exact same caller audio. Reproduce the call that failed, prove your fix, and pick your stack on your own calls.

<Info>
  Build Datasets on the **Dataset** tab and replay them from the **Replays** tab, both under **Simulations**. You need a plan with Call Simulation and an agent set up for [inbound simulation](/docs/simulation/inbound/setup).
</Info>

## What is a Dataset?

A Dataset is a set of real calls you save to test your agent. Each saved call, a Dataset call, keeps only what the caller said, cut into turns. When you replay it, Tuner calls your agent and plays those turns back, and your agent answers as it would on a real call. The caller never changes, so when a result changes, your agent changed.

The loop is simple: save a call, replay it, change your agent, and replay it again.

<Frame>
  <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/dataset-tab.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=4feee0bc6e87c8d81dfc86380c8925fc" alt="The Dataset tab with calls that are Ready, Extracting, and Failed" width="1400" height="1778" data-path="images/datasets/dataset-tab.png" />
</Frame>

***

## Why use a Dataset

Most teams use a Dataset for two things.

### Reproduce the call that failed

Prove your fix on the exact call that broke.

* Save the call that failed, like a noisy line where your agent misheard a date.
* Replay it to see the failure, fix your agent, then replay it again.
* Keep it in your Dataset as a test, so the failure can't quietly come back.

<Frame>
  <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/fix-before-after.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=4c6754303d4679a0993f0d7884486843" alt="The same Dataset call replayed three times before the fix, failing every time, and three times after it, passing every time" width="1400" height="1444" data-path="images/datasets/fix-before-after.png" />
</Frame>

<Card title="Capturing failures" icon="bug" href="/docs/datasets/use-cases/capturing-failures">
  From failure to verified fix, step by step.
</Card>

### Pick your stack on your own calls

Build your own benchmark to test a new model or provider on your callers before you switch.

* Public benchmarks use generic audio, not your callers or your stack.
* Save your typical and hardest calls to your Dataset. That's your benchmark.
* Replay it on your current setup and on the new one: same calls, same scoring, latency measured on every run.

<Frame>
  <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/compare-setups.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=dde9686b274fa56558aea506a91baf9c" alt="Two replay runs of the same three calls: all three pass on the new setup, two fail on the current one" width="1400" height="1444" data-path="images/datasets/compare-setups.png" />
</Frame>

<Card title="Benchmarking" icon="scale-balanced" href="/docs/datasets/use-cases/benchmarking">
  How to build your benchmark and compare two setups.
</Card>

***

## How a Dataset works

Either way, you start with real calls in your Dataset.

<Steps>
  <Step title="Add a Dataset call">
    **Start on day one with your own calls**, or use real calls Tuner recorded from your agent. Open **Simulations > Dataset**, click **Add call**, name it, and pick a source.

    <CardGroup cols={2}>
      <Card title="Upload your own calls" icon="upload" href="/docs/datasets/add-calls#from-an-upload">
        Audio recordings you already have, even before your agent takes any calls in Tuner. Stereo, 1.5 credits per minute.
      </Card>

      <Card title="Pick a production call" icon="phone-volume" href="/docs/datasets/add-calls#from-production">
        A real call your agent took that Tuner recorded, like one that just failed. Free.
      </Card>
    </CardGroup>
  </Step>

  <Step title="Tuner extracts the caller">
    It keeps only the caller's side, cut into **caller turns**. The call shows as **Extracting**, then **Ready**, usually within a minute.
  </Step>

  <Step title="Check the turns">
    Click the headphones icon to listen to each turn before you rely on the call.

    <Frame>
      <img src="https://mintcdn.com/tuner/SUvdaDyWHmpRoVGZ/images/datasets/caller-turns.png?fit=max&auto=format&n=SUvdaDyWHmpRoVGZ&q=85&s=aa20967892e713fd0492b255ea197f90" alt="The caller turns of a Dataset call, each with a play button, time range, and transcript" width="1400" height="1282" data-path="images/datasets/caller-turns.png" />
    </Frame>
  </Step>
</Steps>

<Card title="Add a Dataset call" icon="circle-plus" href="/docs/datasets/add-calls" horizontal>
  Upload your own recording or pick a production call.
</Card>

***

## How a Dataset is used

Your Dataset is built. Now use it in a **Replay**, a separate feature under Simulations. Tuner calls your agent and plays back the caller turns, like a live call, then grades each call against your Pass/Fail evals and measures latency.

<Frame>
  <img src="https://mintcdn.com/tuner/bVl6uahM1YhEolIF/images/datasets/step-replay.png?fit=max&auto=format&n=bVl6uahM1YhEolIF&q=85&s=a7bdf1206c42bd4a4c5bbd0f8a97ebf3" alt="Picking Dataset calls in New Replay and choosing how many times to replay each one" width="1000" height="856" data-path="images/datasets/step-replay.png" />
</Frame>

<Card title="Replays" icon="rotate-right" href="/docs/simulation/replays#replay-a-dataset-call">
  Run a replay, read the results, and compare runs.
</Card>

***

## FAQ

<AccordionGroup>
  <Accordion title="What's inside a Dataset call?">
    | Part | What it is |
    | - | - |
    | **Caller turns** | Each stretch of the caller speaking, played in order during a replay |
    | **Caller audio** | A caller-only track, the audio your agent hears |
    | **Reference transcript** | What the caller said in each turn, for your review |
    | **Channel mode** | **Multi-channel**, or **Mono source** when the agent's voice can bleed in |
    | **Status** | **Extracting**, **Ready**, or **Failed** with a reason |
  </Accordion>

  <Accordion title="How is a Dataset different from generated simulations?">
    | | Generated simulations | Replays of Dataset calls |
    | - | - | - |
    | **Caller** | A simulation agent that improvises | The recorded caller from a real call |
    | **Same caller every run?** | No | Yes |
    | **Best for** | Finding new failure modes | Reproducing a call, proving a fix, comparing setups |

    Found a failure in a generated run? [Save its scenario to your library](/docs/simulation/replays#replay-a-simulation) and run it again.
  </Accordion>

  <Accordion title="How should I choose and name calls?">
    Start from a question each call answers, and add a prefix like **\[stt]**, **\[tool]**, or **\[eval]** so related calls group together. See [Dataset best practices](/docs/datasets/best-practices).
  </Accordion>

  <Accordion title="Does adding a call cost credits?">
    Production calls are free. Uploads cost **1.5 credits per minute** of audio, charged when the call is **Ready**. Replays use the standard simulation rate.
  </Accordion>

  <Accordion title="Can I edit a Dataset call?">
    No, so every replay stays comparable. To cut a call differently, re-extract it, or upload the file again if it was an upload. The original stays as it is.
  </Accordion>

  <Accordion title="Which agents can replay a Dataset?">
    Inbound agents on Vapi, Retell, Dograh, LiveKit, Pipecat, or a Custom API integration. Outbound agents can't replay calls yet.
  </Accordion>
</AccordionGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.