Why Simulation Datasets Matter
Simulation datasets transform simple Q&A testing into realistic conversation testing:Dataset Dashboard
Navigate to Evaluation → Datasets from the left navigation panel. Filter by Multi turn type to see simulation datasets.
Creating a Multi-Turn Dataset
Click the Create Dataset button in the top right corner of the Datasets page.1
Configure Basics

Import from traces and CSV import for multi-turn datasets are coming soon.
2
Configure Scenario

- Lower (1-3): Quick interactions like single-question support
- Medium (4-6): Standard support conversations
- Higher (7-10): Complex, multi-step problem resolution
Provider and Model — Choose the LLM provider and model that will generate simulated user responses (e.g., OpenAI / GPT-4.1).
3
Add User Data and Facts

Example (JSON):
Example (JSON):
4
Select Evaluators

- Agentic: Goal Fulfillment, Information Elicitation
- Quality: Factual Accuracy, Conversation Completeness, Guideline Adherence
- Scenario fields: Goal, persona, user data
- Agent response: What the agent said in each turn
- Conversation metadata: Turn index, conversation history
- Execution data: Latency, tokens, model
5
Configure Evaluators

- Rename (optional) — Rename any evaluator to match your use case (e.g., “Refund Goal Fulfillment” instead of “Goal Fulfillment”)
- Select Provider and Model — For each evaluator, choose the provider and model that will run the LLM-as-Judge evaluation (e.g., OpenAI / GPT-4.1)
Running a Simulation
Once your dataset is configured, you can run simulations:1
Get Dataset ID
Open your dataset and copy the Dataset ID displayed at the top of the page.

2
Trigger Simulation
Use the Dataset ID in your simulation code. The simulation runs automatically
through the Netra SDK.
3
View Results
Monitor progress and results in Test Runs.
Best Practices
Crafting Effective Scenarios
- Be specific: “Get a refund for a damaged product” is better than “Ask about returns”
- Include context: Provide enough detail for realistic simulation (order details, timeline, issue description)
- Include edge cases: Create scenarios that challenge your agent’s boundaries
Choosing User Personas
- Neutral: Best for baseline performance testing
- Friendly: Tests whether your agent maintains professionalism even when not challenged
- Frustrated: Critical for customer support agents—tests patience and de-escalation
- Confused: Tests clarity and explanation quality
- Custom: Use for industry-specific personas (technical users, non-native speakers, etc.)
Defining User Data
- Provide realistic data: Use representative order numbers, dates, and values
- Include edge cases: Test with missing fields, unusual values, or conflicting data
- Keep it relevant: Only include data that matters for the scenario
- Use consistent formats: Standardize date formats, currency, and naming
Setting Fact Checkers
- Focus on critical facts: What MUST the agent communicate correctly?
- Be precise: “5-7 business days” is better than “about a week”
- Test compliance: Include regulatory or policy-critical information
- Verify, don’t duplicate: Don’t repeat information already in user data
Related
- Simulation Overview - Understand the full simulation framework
- Evaluators - Configure scoring logic for simulations
- Test Runs - View simulation results and conversation transcripts
- Traces - Debug simulation turns with execution traces
