Simulation evaluators assess entire multi-turn conversations at the session level. After a simulated conversation completes, evaluators determine whether your agent achieved its goal, communicated facts correctly, and maintained quality throughout the interaction.
Why Simulation Evaluators Matter
Multi-turn conversations require different evaluation approaches than single-turn responses:
Evaluators Dashboard
Navigate to Evaluation → Evaluators from the left navigation panel. Switch to the Library tab and filter by Multi turn to see the simulation evaluators.
Netra organizes simulation evaluators into two categories: Quality and Agentic.
Library Evaluators
Netra provides 8 preconfigured library evaluators across two categories. All evaluators run at the session level, assessing the entire conversation after it completes.
Quality Evaluators
Quality evaluators assess how well your agent maintains conversation standards.
Agentic Evaluators
Agentic evaluators assess goal-directed and information-gathering behavior.
Evaluator Configuration
All 8 library evaluators share the same configuration:
You can adjust the pass criteria threshold for any evaluator based on your requirements. A higher threshold enforces stricter quality standards.
Using Evaluators in Simulations
When configuring a multi-turn dataset, you select and configure evaluators in Step 4 of the dataset creation flow. You can choose any combination of Quality and Agentic evaluators based on what you want to measure.
Best Practices
Choosing Evaluators by Scenario Type
Getting Started with Evaluators
- Start with Goal Fulfillment and Factual Accuracy — these cover the most critical aspects of any simulation
- Add Quality evaluators based on your use case — Conversation Completeness and Guideline Adherence are strong defaults
- Adjust pass criteria if the default threshold of 0.6 is too lenient or strict for your needs
- Monitor results across the first few test runs to ensure evaluators align with your expectations