Skip to main content

Start from a question

Every Dataset call should answer something specific. Is the new STT better with accents? Did the prompt fix stop the made-up policy? Does booking still work after the tool change? If you can’t name the question a call answers, it probably doesn’t belong in the set.

Name calls so they group

Your agent has one Dataset, and you pick which calls to replay together. Put the purpose at the front of each name so the set you need is easy to find in New Replay: Use the Description to note why the call is in the set, like “Agent promised a refund the policy doesn’t allow”. Names are unique per agent.

Keep calls short and focused

Shorter calls replay faster, cost less, and fit within the Max call duration. Each caller turn can add up to the Max wait of waiting for your agent, so with the default settings a 3-minute cap fits roughly 8 caller turns. The Longest possible call line in New Replay tells you before you run whether a call fits. If only part of a long call matters, upload a trimmed recording of that part instead.

Prefer multi-channel recordings

A Multi-channel call keeps the caller on their own channel, so the replay is clean. A Mono source mixes caller and agent, so fragments of the original agent can leak into the caller’s turns, and your agent will hear them. Use mono calls only when there’s no better recording, and listen to them first.

Listen before you trust

Open every new call with the headphones icon and play a few turns. Check for clipped words, agent speech inside a caller turn, and turns split in the wrong place. A bad cut makes every replay of that call misleading. See Review the caller turns.

Keep your failures, and some golden calls

Every call your agent got wrong is a test you get for free. Add it, fix the agent, and keep replaying it so the failure can’t come back unnoticed. Balance failures with a few golden calls your agent handles well today, so fixes that break something else show up too.

Replay more than once

Your agent’s responses vary from run to run even when the caller doesn’t. Replay each call 3 to 5 times:
  • Failing every time means the problem is consistent.
  • Failing now and then means it’s intermittent, and one passing replay after a fix doesn’t prove much.
The limit is 20 replayed calls per run, counting repeats, so 4 calls replayed 5 times fits in one run.

Change one thing at a time

Between two runs you compare, change only one thing, and keep the same calls, repeats, Max call duration, and Advanced turn-taking settings. Then any difference has one cause. See Compare two setups.

Remember the caller doesn’t adapt

The caller says what they said in the original call, whatever your agent answers now. Pick calls where the behaviour you’re testing happens early, and write evals about the agent’s turns, not about how the conversation ends. See The caller doesn’t adapt.

Protect your tools and data

Replays reach your real agent, and your agent calls its real tools. Point it at test endpoints or test accounts before replaying calls that book, charge, refund, or change records. See Check a tool call on a replay.

Refresh the set

Add calls when you launch new Intents, enter a new market, or see a new kind of failure in production. Delete calls that no longer match how your product works. Past runs keep their results after you delete a call.

Mind privacy

Dataset calls keep your callers’ real voices and words, and a Dataset call keeps its own copy of the caller’s audio even if the source call is later removed. Add only calls you’re allowed to reuse for testing, and delete the ones you no longer need.

Keep an eye on cost

Next steps

Add Dataset calls

Pick production calls or upload recordings.

Replays

Replay your Dataset calls against your agent.