Skip to main content
Evaluators are the scoring logic that determines whether your AI system meets quality standards. They transform subjective assessments into measurable metrics — from semantic correctness and tool execution accuracy to audio naturalness and visual composition. Use them with evaluations to build automated quality pipelines.

Quick Start: Evaluation

New to evaluations? Get your first evaluation running in minutes.

Evaluator Categories

Netra organizes evaluators by the modality they assess:

Text Evaluators

LLM-as-Judge, code-based, and rule-based evaluators for text outputs. Assess answer correctness, relevance, safety, tool execution, and custom business logic.

Voice Evaluators

Audio-aware evaluators for TTS, STT, and conversational delivery. Score voice naturalness, transcription accuracy, pronunciation, and speaking behavior.

Image Evaluators

Multimodal and rule-based evaluators for image generation and editing. Assess composition, style, format compliance, and image-text alignment.

Evaluator Types

Across all categories, Netra provides several evaluation approaches:

Cross-Modal Evaluation

Evaluators are not limited to their primary modality. Text evaluators can evaluate transcript content from voice interactions, making them useful for scoring conversation quality even when the input is audio-derived.
When a voice simulation completes, Netra produces both audio and transcript data. You can attach text evaluators to the same evaluation to evaluate the conversation content, while voice evaluators assess the audio delivery — giving you full coverage of both what was said and how it sounded.

Evaluator Library

Netra provides a library of 49 preconfigured evaluators across 12 categories, ready to use or customize:

Creating Evaluators

You can create evaluators in two ways:
  1. From the Library — Browse prebuilt evaluators, click Add, customize the prompt or parameters, and save to My Evaluators
  2. Custom — Click Add Custom, choose the evaluator type (LLM as Judge, Code, or rule-based), and configure from scratch
Every evaluator supports Playground testing so you can validate it against sample data before attaching it to evaluations.

Getting Started

1

Choose Your Modality

Decide whether you’re evaluating text, voice, or image outputs — or a combination.
2

Browse the Library

Start with prebuilt evaluators from the Library. Customize thresholds and prompts to match your quality standards.
3

Test in Playground

Validate evaluators against sample data before deploying to production evaluations.
4

Attach to Evaluations

Connect evaluators to evaluations and map variables to your data fields.
5

Run and Analyze

Execute evaluations and analyze results in Test Runs. Iterate on prompts and thresholds.
Last modified on August 28, 2026