Skip to content

TechnicalPractitioner

5. Evaluator design

For: engineers and SREsPrerequisites: Have read chapter 4 of the guide, especially step 5.

Evaluators are the most domain-specific component of an AI observability platform. Off-the-shelf scorers (Ragas, DeepEval, Phoenix evals) handle the generic dimensions (faithfulness, toxicity, format). Domain-specific quality requires custom evaluators, almost always implemented as LLM-as-judge calls with a carefully written rubric. This chapter covers the practical mechanics.

A practical starting set for most deployments.

EvaluatorApplies toWhat it measures
FaithfulnessRAGWhether the answer is supported by the retrieved context
Answer relevancyLLM, RAGWhether the answer addresses the question asked
Context precisionRAGProportion of retrieved chunks that are relevant
Context recallRAGProportion of relevant chunks that were retrieved
ToxicityLLMPresence of harmful, harassing or hateful content
PII leakageLLMPresence of personal identifying data in output
Format complianceLLMConformance of output to a required schema
Hallucination rateLLM, RAGFabricated entities or claims not present in source
Refusal correctnessLLMWhether the model refused (or did not refuse) appropriately
Tool selectionAgentWhether the correct tool was chosen at each step
Argument validityAgentWhether tool arguments matched the schema
Loop boundAgentWhether the agent exceeded an acceptable iteration count
Domain rubricAnyCustom criteria specific to the application

Three implementation classes, in order of cost and flexibility.

  • Rule-based: regex, schema validation, keyword presence. Fast, cheap, deterministic. Limited to surface properties.
  • Classifier-based: small models trained for specific tasks (toxicity, PII detection, sentiment). Fast at inference time, deterministic, require labeled training data.
  • LLM-as-judge: a model scores another model’s output against a written rubric. Flexible, expensive, non-deterministic. Required for any semantic criterion.

A common error is to use LLM-as-judge for everything. Simple pattern detection (a card number, a banned word) is a regex job. Schema validation and format compliance of structured output are a JSON Schema check. Reserve LLM-as-judge for criteria that genuinely require semantic understanding.

Three patterns cover most use cases.

The judge returns 0 or 1. Used for hallucination detection or refusal correctness. Format compliance of structured output is better checked with JSON Schema; keep the judge for format criteria that require a semantic reading (tone, register).

SYSTEM:
You are an evaluator. Given a question, a context and an answer, return
exactly one JSON object with the schema:
{"grounded": 0 or 1, "reason": "one short sentence"}
A grounded answer is one where every factual claim is supported by the
context. Inferences that are not stated in the context are not grounded.
USER:
Question: {question}
Context: {context}
Answer: {answer}

The judge returns a numeric score against a written rubric. Used for relevance, completeness, tone.

SYSTEM:
Score the answer on a scale of 1 to 5 against the following rubric:
1 = irrelevant or off-topic
2 = partially addresses the question, missing key information
3 = addresses the question but with errors or omissions
4 = correct and complete, lacks polish
5 = correct, complete, well-structured
Return JSON: {"score": int, "reason": "short justification"}
USER:
Question: {question}
Answer: {answer}

The judge picks the better of two answers. Used for A/B testing of prompts or models. Less sensitive to absolute calibration than single-answer scoring.

SYSTEM:
Given a question and two candidate answers, return JSON:
{"winner": "A" or "B" or "tie", "reason": "short justification"}
Judge on factual correctness first, then completeness, then clarity.
USER:
Question: {question}
Answer A: {answer_a}
Answer B: {answer_b}

Pairwise judges have a position bias: they tend to favor the answer shown first (or second, depending on the model). Run each pair twice with A and B swapped, and only declare a winner when both runs agree; otherwise, count a tie.

LLM-as-judge scores are not absolute. They depend on the judge model, the prompt and the temperature. Calibration means establishing how the judge’s scores correlate with human judgment on a curated set.

  1. Build a calibration set with human-assigned ground truth. Order of magnitude: a few dozen to a few hundred traces (for example 50 to 200), with enough positive and negative cases; the rarer or subtler the criterion, the more you need.
  2. Run the judge against the calibration set. Measure agreement with a chance-corrected statistic, such as Cohen’s kappa for binary or categorical verdicts. Raw agreement is misleading when one class dominates: if 90 percent of answers are grounded, a judge that always answers “grounded” gets 90 percent agreement. For numeric scores, use a rank correlation or a weighted kappa.
  3. Set the acceptance threshold before measuring, according to the stakes. As a reference point, the Landis and Koch (1977) scale calls a kappa between 0.61 and 0.80 “substantial” agreement; a threshold of 0.7 is an example, not a standard. Below your threshold, the judge prompt or the judge model needs work.
  4. Recalibrate when the judge model is upgraded. Provider model upgrades shift judge behavior.

Per-trace scores are useful for forensic queries. Aggregate scores are useful for monitoring. Three aggregation patterns.

  • Mean over a window: simple, hides distribution shape. Default for stable evaluators.
  • Percentile (often p10 or p25): surfaces the worst tail rather than the average. Better for catching regressions.
  • Rate below threshold: fraction of traces with score under a fixed value. This is the basis of a quality SLO, which is stated as a proportion: for example “at least 95 percent of evaluated answers score above 0.8 on faithfulness over a rolling 7-day window” (illustrative values).

This proportion is estimated on a sample of evaluated answers: you therefore compare not its point value with the target but the lower bound of its confidence interval (Wilson bound). Defining the quality SLO, sizing the sample, the error budget and the judge’s meta-observability are covered in the GenAI method, Part IV, which is the reference on this point.

A single judge call is noisy. Two strategies reduce variance.

  • Self-consistency: run the same judge prompt three or five times and take the majority. Multiplies judge cost by three (or five) and markedly reduces variance.
  • Judge ensembles: use two different judge models and require agreement. Higher confidence on positive cases, more disagreement to investigate.

Use self-consistency on high-stakes evaluators (refusal, safety). Use single-call judges on cheap continuous scorers (relevance).

Running every evaluator on every trace is rarely affordable at scale. Tier the evaluation.

A principle shared across the site: what sets the precision of a quality indicator is the number of judgments per segment (feature, tenant, language, request type) and per period, not the sampled percentage. The same 1 percent rate yields 100,000 judgments a day on 10 million requests, but only 20 on 2,000 requests, too few to track a proportion. First set the number of judgments needed per segment and per window (the GenAI method, Part IV gives the sizing formula), derive the rate from it, then oversample rare or risky segments. The tiers below illustrate a cost-based organization; their percentages are examples.

  • Tier 1 (cheap, always-on): rule-based format checks, classifier-based PII and toxicity. For example on 100 percent of traces.
  • Tier 2 (medium cost): small LLM-as-judge calls for faithfulness and relevance. For example on 10 percent of traces.
  • Tier 3 (high cost): large judge models or self-consistency for domain rubrics. For example on 1 percent of traces, plus 100 percent of traces flagged by Tier 1 or Tier 2.

These rates are examples, to be adjusted to your volume and budget: check that they yield enough judgments per segment and per period for the precision you need.

Next: 6. The tools landscape.

Revised on 4 October 2026: quality SLO stated as a proportion with a link to the GenAI method (Wilson bound), judge sampling principle (number of judgments per segment and per period rather than percentage), tiers presented as examples.