Simulations let you test how your AI Voice Agent will behave in real-world situations before it speaks to a single customer. You describe a caller scenario and what a successful conversation looks like, and Aircall runs simulated conversations between an AI-played caller and your agent. Each conversation is then scored against your success criteria, so you can see exactly where your agent performs well and where it needs adjusting.
This article explains what simulations are, why they matter, and how to create scenarios, run simulations, and read the results.
What is a simulation?
A simulation is a text-based reproduction of a call. It's not a real phone call, and no audio is involved. Instead, an AI plays the role of a caller and has a conversation with your AI Voice Agent, turn by turn, just as a real caller would.
Your agent responds using its actual configuration, including its call script, knowledge sources, intake questions, call transfer rules, and AI actions. This means the simulation reflects how your agent would really behave on a live call.
Every simulation follows three stages:
- Define the scenario: describe who is calling, why they're calling, and what the agent needs to achieve for the conversation to be a success.
- Run the conversation: a simulated caller talks to your AI Voice Agent and stays in character until the conversation ends.
- Evaluate the result: an AI evaluator reviews the conversation against each success criterion and explains its verdict using evidence from the transcript and any AI actions the agent performed.
Note: Because simulations are text-based and not real calls, they don't count towards your call minutes.
Why run simulations?
AI Voice Agents are powered by large language models (LLMs), which means they can respond slightly differently from one conversation to the next. A single manual test call only shows you one possible outcome. Simulations help you:
- Catch issues before your customers do: see how your agent handles a situation before it goes live.
- Validate changes safely: test edits to your call script, knowledge sources, or AI actions and confirm they work as expected.
- Prevent regressions: re-run your saved scenarios after every change to make sure something that used to work still does.
- Test difficult situations: check how your agent handles frustrated callers, unclear requests, or questions it shouldn't answer.
- Measure consistency: run the same scenario multiple times to see how reliably your agent reaches the right outcome.
Cost and usage limits
- Cost: running simulations is free.
- Call minutes: simulations are text-based and don't count towards your call minutes.
- Daily limit: your company can run up to 100 simulated conversations per day, shared across all users and all AI Voice Agents.
- Scenarios per run: you can run one scenario at a time, with multiple iterations of that scenario.
Important: Each iteration counts as one simulated conversation. For example, running one scenario with 5 iterations uses 5 of your 100 daily conversations. If you reach the limit, you'll see a warning and won't be able to start new simulations until the following day; your saved scenarios and past results remain available. If you need a higher limit, contact Aircall Support (https://support.aircall.io/en-gb/articles/10375397243293).
Accessing Simulations
- In your Aircall Dashboard, go to AI Voice Agent and select the agent you want to test.
- In the left-hand menu, under Deploy, select Simulations.
- The Scenarios tab lists all scenarios saved for this agent, including the number of success criteria and when each was last saved.
Creating scenarios
A scenario is a reusable test case for your agent. It describes what the caller will do and what a successful conversation should accomplish. Scenarios are saved to the agent, so you can edit, reuse, and re-run them whenever you make changes.
Each scenario includes:
- Success criteria: what the agent must do for the scenario to pass, for example, "Caller gets transferred to the billing department."
- Reason for the call: what the caller wants, for example, "Caller needs to find out how to refund their order."
- Simulated caller (optional): how the caller behaves, including their speaking style and emotional state, and the details they'll provide when asked, such as an account number.
You can create scenarios in two ways: by writing them yourself, or by generating them with AI.
Option 1: Create a scenario manually
Use this option when you already know exactly what you want to test.
- On the Scenarios tab, select Create scenario.
- Enter a Scenario name, for example, "Refund request."
- Add your Success criteria. Select + to add more than one criterion.
- Describe the Reason for the call.
- Optionally, define the Simulated caller:
- Speaking style: for example, "Speaks quickly, interrupts often."
- Emotional state: for example, "Confused."
- Details for caller to provide: information the caller will give when the agent asks for it, such as an account number or order number. Select + to add more details.
- Optionally, expand Additional context to add details such as the Caller phone number or Aircall phone number.
- Select Save scenario.
Option 2: Generate scenarios with AI
Use this option to quickly build a set of test cases. AI generates scenarios based on your agent's configuration, such as its call script, transfer rules, intake questions, AI actions, and knowledge sources. It also generates the success criteria for each scenario.
- On the Scenarios tab, select Create with AI.
- Choose the Number of scenarios you want to generate.
- Optionally, add Instructions to focus the scenarios on what you want to test, for example, "Focus on frustrated callers disputing a charge" or "Ask for an order number."
- Select Draft scenarios.
- Review the generated scenarios. Each one shows its type and its success criteria. Expand a scenario to see all of its criteria.
- Select Save scenarios to add them to your scenario list.
Generated scenarios cover different types of situations so you get a well-rounded view of your agent:
- Ideal: straightforward conversations that establish a baseline.
- Standard: conversations covering your agent's configured workflows and decision points.
- Edge cases: requests outside what your agent is configured to handle, to check that it responds appropriately.
Tip: Once saved, AI-generated scenarios work exactly like manual ones. You can open and edit them to fine-tune the success criteria before running a simulation.
Writing good success criteria
The AI evaluator checks each success criterion independently, so clear criteria lead to clear results.
- One outcome per criterion: instead of "Agent collects the order number and transfers to billing," split this into two criteria.
- Describe observable behavior: "The agent confirms the new delivery date and time with the caller" is easier to evaluate than "The agent is helpful."
- Include actions when relevant: "The agent looks up the contact before creating a new one" checks what the agent actually did, not just what it said.
Running a simulation
- On the Scenarios tab, find the scenario you want to test and select Run simulation.
- In the Run simulation panel, enter a Simulation name, for example, "First test batch."
- Select the AI agent version to test: the last published version, or your current draft.
- Choose whether to Use real API actions. This setting is off by default, which means API actions are simulated. See Simulated vs. real API actions below.
- Set the number of Iterations, the number of conversations to run for this scenario.
- Select Run simulation.
Your simulation appears at the top of the Run history tab, and results appear as each conversation finishes.
Tip: Testing your current draft lets you see the effect of an unpublished change before any customer experiences it.
Simulated vs. real API actions
If your agent uses AI actions that call external systems, such as looking up or creating a contact in your CRM, you can choose how those actions behave during the simulation.
| Simulated API actions (default) | Real API actions | |
|---|---|---|
| What happens | Aircall simulates the API response. No request is sent to your connected apps. | The agent calls the relevant apps set up in your integrations library, just like on a live call. |
| Best for | Quickly testing how your agent handles different outcomes, including both successful and failed actions. | End-to-end testing of your integrations. |
| Advantages | Fast, and leaves your live data untouched. | You can check your CRM or other connected app to see whether the action succeeded, what was sent, and how it was recorded. |
| Things to consider | The response is simulated, so it doesn't confirm that your integration is set up correctly. | Actions are performed on your live systems, so records may be created or updated. |
Important: When using real API actions, make sure your scenario provides every piece of information the action requires, for example by adding it under Details for caller to provide. If a required field is missing, the API call may fail, and the simulation result will reflect the missing data rather than how your agent truly performs.
Why run multiple iterations?
Because AI Voice Agents are powered by an LLM, the same scenario can play out a little differently each time. The agent may phrase a question differently, ask things in a different order, or take a different path through the conversation.
Running several iterations of the same scenario shows you the range of conversations your agent might have and how consistently it achieves the right outcome. A scenario that passes once might not pass every time, and multiple iterations help you spot that before your customers do.
Tip: We recommend running each scenario several times rather than just once. If some iterations fail, compare the transcripts to see what differed and adjust your agent's configuration.
Viewing simulation results
Go to the Run history tab to see all past simulations for this agent. Each simulation shows the number of iterations, the agent version tested, and how many iterations passed. Select a simulation to open its Simulation report.
The simulation report
The Overview section summarizes how the agent performed across all iterations:
- Scenario pass rate: the percentage of iterations that met all success criteria.
- Average conversation length: the average number of turns and duration per conversation.
- Lowest scoring criterion: the success criterion your agent struggled with most, which is usually the best place to start improving.
- Success criteria: a breakdown showing how many iterations met each criterion.
Below the overview, the Scenario iterations section lists each individual conversation. Each iteration is scored against your success criteria and receives one of the following results:
- Passed: the agent met all success criteria.
- Partially passed: the agent met some, but not all, success criteria.
- Failed: the agent didn't meet the success criteria.
Use the dropdown to filter iterations by result.
Reviewing a conversation transcript
Select View transcript on any iteration to see exactly what happened. The transcript panel shows:
- The verdict for each success criterion, with an explanation of why it was met or not met.
- The full conversation, turn by turn, between the simulated caller and your agent.
- The AI actions the agent performed, which you can expand to see the details sent and returned.
Best practices
- Start with AI-generated scenarios: generate a first set of scenarios to quickly cover your agent's main workflows, then add manual scenarios for situations specific to your business.
- Run multiple iterations: run each scenario several times to see how consistently your agent behaves.
- Test your draft before publishing: run your key scenarios against your draft after every change to catch regressions before they reach customers.
- Start with simulated API actions: use simulated actions to iterate quickly, then switch to real API actions for a final end-to-end check of your integrations.
- Provide complete caller details: when using real API actions, include every field the action needs so results reflect your agent's true performance.
- Focus on the lowest scoring criterion: use the report to find where your agent struggles most, review the related transcripts, and adjust your configuration.