Why Simulation Test Runs Matter
Simulation test runs provide deep insights into conversational agent performance:Test Runs Dashboard
Navigate to Evaluation → Test Runs from the left navigation panel to see simulation test runs.
Filtering and Search
- Date Range: Filter runs by time period to compare performance over time
- Search: Find specific test runs by agent or dataset name
- Sort: Order by date, status, or dataset
Viewing Test Run Details
Click on any simulation test run to access detailed results.
Summary Metrics
The top of the detail view shows aggregated performance data:Viewing Scenario Details
Click on any test run item to view the detailed scenario results. This opens a modal with three tabs.Tab 1: Conversation
The Conversation tab shows the full multi-turn dialogue between the simulated user and your agent.
- Turn-by-Turn Display: Each conversation turn is clearly separated
- User Messages: Shows what the simulated user said
- Agent Responses: Shows what your agent replied
- Trace Links: Click View Trace on any turn to see detailed execution traces
- Turn Index: Track which turn number you’re viewing (Turn 1, Turn 2, etc.)
- Exit Reason: Shows why the conversation ended
Tab 2: Evaluation Results
The Evaluation Results tab shows scores from all configured evaluators.
Tab 3: Scenario Details
The Scenario Details tab shows the complete configuration used for this simulation.
User Data Section:
Shows all context data provided to the simulated user:
The Scenario Details tab is crucial for understanding the context of each
simulation. It shows exactly what data the simulated user had access to and
what facts the agent was expected to communicate.
Analyzing Simulation Results
Identifying Patterns
When reviewing simulation test runs, look for:- Goal achievement rates: What percentage of simulations achieved their goals?
- Persona differences: Does your agent perform better with certain personas? Run the same scenarios with all persona types and compare results.
- Turn efficiency: Are conversations longer than necessary? Compare turn counts for successful vs failed scenarios.
- Common failure points: Which turns typically cause issues?
- Fact accuracy: Are specific facts consistently missed?
- Cost trends: Monitor total cost across test runs and identify scenarios that consume excessive turns.
Debugging Failed Simulations
For each failed scenario:- Review the Conversation tab: Identify where the conversation went wrong
- Check the Evaluation Results tab: See which evaluators failed and why
- Examine the Scenario Details tab: Verify the user data and facts were correct
- Click View Trace: Inspect the full execution flow for problematic turns — check LLM inputs, tool calls, and latency breakdowns
Comparing Across Runs
To track improvement or regression:- Run simulations after each agent update
- Compare goal achievement rates and evaluator scores across runs
- Investigate scenarios that changed from pass to fail
- Track turn efficiency and cost trends over time
Best Practices
- Test after every agent change: Run simulations when updating your agent to catch regressions early
- Create baseline runs: Establish performance benchmarks before making changes
- Always check traces for failures: Don’t just read the conversation — inspect the execution flow, LLM context, and tool calls
- Review latency: Identify slow turns that might frustrate real users
Related
- Simulation Overview - Understand the full simulation framework
- Datasets - Create scenarios that generate test runs
- Evaluators - Configure scoring logic for simulations
- Traces - Debug simulation turns with execution traces
