Technical
5. Evaluator design
For: engineers and SREsPrerequisites: Have read chapter 4 of the guide, especially step 5.
Evaluators are the most domain-specific component of an AI observability platform. Off-the-shelf scorers (Ragas, DeepEval, Phoenix evals) handle the generic dimensions (faithfulness, toxicity, format). Domain-specific quality requires custom evaluators, almost always implemented as LLM-as-judge calls with a carefully written rubric. This chapter covers the practical mechanics.
5.1 Evaluator catalog
Section titled “5.1 Evaluator catalog”A practical starting set for most deployments.
| Evaluator | Applies to | What it measures |
|---|---|---|
| Faithfulness | RAG | Whether the answer is supported by the retrieved context |
| Answer relevancy | LLM, RAG | Whether the answer addresses the question asked |
| Context precision | RAG | Proportion of retrieved chunks that are relevant |
| Context recall | RAG | Proportion of relevant chunks that were retrieved |
| Toxicity | LLM | Presence of harmful, harassing or hateful content |
| PII leakage | LLM | Presence of personal identifying data in output |
| Format compliance | LLM | Conformance of output to a required schema |
| Hallucination rate | LLM, RAG | Fabricated entities or claims not present in source |
| Refusal correctness | LLM | Whether the model refused (or did not refuse) appropriately |
| Tool selection | Agent | Whether the correct tool was chosen at each step |
| Argument validity | Agent | Whether tool arguments matched the schema |
| Loop bound | Agent | Whether the agent exceeded an acceptable iteration count |
| Domain rubric | Any | Custom criteria specific to the application |
5.2 Evaluator types
Section titled “5.2 Evaluator types”Three implementation classes, in order of cost and flexibility.
- Rule-based: regex, schema validation, keyword presence. Fast, cheap, deterministic. Limited to surface properties.
- Classifier-based: small models trained for specific tasks (toxicity, PII detection, sentiment). Fast at inference time, deterministic, require labeled training data.
- LLM-as-judge: a model scores another model’s output against a written rubric. Flexible, expensive, non-deterministic. Required for any semantic criterion.
A common error is to use LLM-as-judge for everything. Simple pattern detection (a card number, a banned word) is a regex job. Schema validation and format compliance of structured output are a JSON Schema check. Reserve LLM-as-judge for criteria that genuinely require semantic understanding.
5.3 LLM-as-judge prompt patterns
Section titled “5.3 LLM-as-judge prompt patterns”Three patterns cover most use cases.
Pattern 1: binary classification
Section titled “Pattern 1: binary classification”The judge returns 0 or 1. Used for hallucination detection or refusal correctness. Format compliance of structured output is better checked with JSON Schema; keep the judge for format criteria that require a semantic reading (tone, register).
SYSTEM:You are an evaluator. Given a question, a context and an answer, returnexactly one JSON object with the schema: {"grounded": 0 or 1, "reason": "one short sentence"}A grounded answer is one where every factual claim is supported by thecontext. Inferences that are not stated in the context are not grounded.
USER:Question: {question}Context: {context}Answer: {answer}Pattern 2: rubric-based scoring
Section titled “Pattern 2: rubric-based scoring”The judge returns a numeric score against a written rubric. Used for relevance, completeness, tone.
SYSTEM:Score the answer on a scale of 1 to 5 against the following rubric: 1 = irrelevant or off-topic 2 = partially addresses the question, missing key information 3 = addresses the question but with errors or omissions 4 = correct and complete, lacks polish 5 = correct, complete, well-structuredReturn JSON: {"score": int, "reason": "short justification"}
USER:Question: {question}Answer: {answer}Pattern 3: pairwise comparison
Section titled “Pattern 3: pairwise comparison”The judge picks the better of two answers. Used for A/B testing of prompts or models. Less sensitive to absolute calibration than single-answer scoring.
SYSTEM:Given a question and two candidate answers, return JSON: {"winner": "A" or "B" or "tie", "reason": "short justification"}Judge on factual correctness first, then completeness, then clarity.
USER:Question: {question}Answer A: {answer_a}Answer B: {answer_b}Pairwise judges have a position bias: they tend to favor the answer shown first (or second, depending on the model). Run each pair twice with A and B swapped, and only declare a winner when both runs agree; otherwise, count a tie.
5.4 Calibration
Section titled “5.4 Calibration”LLM-as-judge scores are not absolute. They depend on the judge model, the prompt and the temperature. Calibration means establishing how the judge’s scores correlate with human judgment on a curated set.
- Build a calibration set with human-assigned ground truth. Order of magnitude: a few dozen to a few hundred traces (for example 50 to 200), with enough positive and negative cases; the rarer or subtler the criterion, the more you need.
- Run the judge against the calibration set. Measure agreement with a chance-corrected statistic, such as Cohen’s kappa for binary or categorical verdicts. Raw agreement is misleading when one class dominates: if 90 percent of answers are grounded, a judge that always answers “grounded” gets 90 percent agreement. For numeric scores, use a rank correlation or a weighted kappa.
- Set the acceptance threshold before measuring, according to the stakes. As a reference point, the Landis and Koch (1977) scale calls a kappa between 0.61 and 0.80 “substantial” agreement; a threshold of 0.7 is an example, not a standard. Below your threshold, the judge prompt or the judge model needs work.
- Recalibrate when the judge model is upgraded. Provider model upgrades shift judge behavior.
5.5 Score aggregation
Section titled “5.5 Score aggregation”Per-trace scores are useful for forensic queries. Aggregate scores are useful for monitoring. Three aggregation patterns.
- Mean over a window: simple, hides distribution shape. Default for stable evaluators.
- Percentile (often p10 or p25): surfaces the worst tail rather than the average. Better for catching regressions.
- Rate below threshold: fraction of traces with score under a fixed value. This is the basis of a quality SLO, which is stated as a proportion: for example “at least 95 percent of evaluated answers score above 0.8 on faithfulness over a rolling 7-day window” (illustrative values).
This proportion is estimated on a sample of evaluated answers: you therefore compare not its point value with the target but the lower bound of its confidence interval (Wilson bound). Defining the quality SLO, sizing the sample, the error budget and the judge’s meta-observability are covered in the GenAI method, Part IV, which is the reference on this point.
5.6 Self-consistency and judge ensembles
Section titled “5.6 Self-consistency and judge ensembles”A single judge call is noisy. Two strategies reduce variance.
- Self-consistency: run the same judge prompt three or five times and take the majority. Multiplies judge cost by three (or five) and markedly reduces variance.
- Judge ensembles: use two different judge models and require agreement. Higher confidence on positive cases, more disagreement to investigate.
Use self-consistency on high-stakes evaluators (refusal, safety). Use single-call judges on cheap continuous scorers (relevance).
5.7 Sampling for evaluation
Section titled “5.7 Sampling for evaluation”Running every evaluator on every trace is rarely affordable at scale. Tier the evaluation.
A principle shared across the site: what sets the precision of a quality indicator is the number of judgments per segment (feature, tenant, language, request type) and per period, not the sampled percentage. The same 1 percent rate yields 100,000 judgments a day on 10 million requests, but only 20 on 2,000 requests, too few to track a proportion. First set the number of judgments needed per segment and per window (the GenAI method, Part IV gives the sizing formula), derive the rate from it, then oversample rare or risky segments. The tiers below illustrate a cost-based organization; their percentages are examples.
- Tier 1 (cheap, always-on): rule-based format checks, classifier-based PII and toxicity. For example on 100 percent of traces.
- Tier 2 (medium cost): small LLM-as-judge calls for faithfulness and relevance. For example on 10 percent of traces.
- Tier 3 (high cost): large judge models or self-consistency for domain rubrics. For example on 1 percent of traces, plus 100 percent of traces flagged by Tier 1 or Tier 2.
These rates are examples, to be adjusted to your volume and budget: check that they yield enough judgments per segment and per period for the precision you need.
Next: 6. The tools landscape.
Revised on 4 October 2026: quality SLO stated as a proportion with a link to the GenAI method (Wilson bound), judge sampling principle (number of judgments per segment and per period rather than percentage), tiers presented as examples.