Skip to content

TechnicalPractitioner

Module 4: continuous evaluation in production

For: engineers and SREsPrerequisites: Have read module 3 of the course on the three drifts.

The two complement each other and answer different questions.

OnlineOffline
On whatreal traffic, sampled (for example 1 to 5%, see below)a reference set (order of magnitude: a few hundred to a thousand question and answer pairs)
Whencontinuouslyat each version, before deployment
Costkept under control by sampling, since each judgment consumes tokenshigher, with a human in the loop
Strengthfast detection, near real-time alertingrigorous comparison between versions
Limitno oracle, only another modeldoes not see drifts after release to production

In practice, you evaluate offline before each release to production and online continuously.

How many answers should be judged online? The principle used across the site: what makes a score precise is the number of judgments per segment and per period (per feature or per customer, per hour or per day), not the percentage of traffic. The 1 to 5% quoted here are examples: on heavy traffic, 1% already yields many judgments; on a rarely used segment, even 100% may yield only a few per day. The design of sampling (evaluator tiers, traces that are always evaluated) has its reference in section 5.7 of the guide; this module keeps what is specific to the labs: judge biases and RAG diagnosis.

The two evaluations form a single loop:

flowchart TB
  V["New version"] --> OFF{"Offline<br/>reference set"}
  OFF -->|"regression"| FX["Fix"]
  FX --> V
  OFF -->|"compliant"| PR["Production"]
  PR --> ON["Online<br/>1 to 5% sample,<br/>different judge"]
  PR --> FB["User feedback<br/>linked to the trace_id"]
  ON --> AL{"Score<br/>dropping"}
  FB --> AL
  AL -->|"yes"| EN["Investigation<br/>on the trace"]
  EN -.->|"new cases"| OFF

Having a model grade the answers is fast and yields a score you can use as a metric, but it is expensive and non-deterministic: pinning the judge model version and setting temperature to zero improve reproducibility without guaranteeing it. Agreement with human judgment can be high: on comparisons of chatbot answers, Zheng et al. measure over 80% agreement between GPT-4 and human evaluators, as much as between humans (Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, 2023). This result does not carry over as is to your judge or your criteria: measure agreement on a sample you have had annotated.

Its biases are known and must be offset:

BiasEffect
verbosityprefers long answers
self-preferenceprefers its own style
positionfavors the first option in an A/B comparison
confidencetolerates a well-written hallucination
hidden costthe judge’s tokens sometimes exceed those of the evaluated model

To offset them, use a structured evaluation prompt, a rigid output format, a judge different from the evaluated model, and cross-sampling with human annotations.

MetricQuestionTools
Faithfulnessare the claims in the answer supported by the retrieved context?RAGAS, Phoenix
Context recalldoes the context contain the needed information?RAGAS, ground truth
Context precisionis the context dense or noisy?RAGAS
Answer relevancedoes the answer address the question asked?RAGAS, LLM judge

Low faithfulness with low context recall points to retrieval. Low faithfulness with high context recall points to generation.

flowchart LR
  F["Low faithfulness"] --> R{"Context recall"}
  R -->|"low"| RET["Look at retrieval"]
  R -->|"high"| GEN["Look at generation"]

Three levels, to combine according to risk:

  1. Rules: a figure, a date or a proper name in the answer that does not appear in the context.
  2. Specialized models: entailment classifiers between context and answer.
  3. Weak signals: negative feedback, immediate rewordings of the same question, abandonment.

Thumbs up or down is the cheapest and most direct signal. Link it to the trace through its trace_id: negative feedback then becomes an entry point to the exact context that produced the bad answer.

  • Offline evaluation gates each release to production; online evaluation runs continuously.
  • An LLM judge is only useful if it differs from the evaluated model and its biases are offset.
  • Read together, faithfulness and context recall tell whether the problem comes from retrieval or generation.
  • User feedback linked to its trace gives the exact context of the error, which no aggregate score provides.

Revised on 2 October 2026: the unsourced correlation figure between judge and humans is replaced by the study by Zheng et al. (2023), sampling rates and reference set size are presented as orders of magnitude.

Revised on 4 October 2026: judge sampling principle stated (the number of judgments per segment and per period matters more than the percentage), rates presented as examples and pointer to section 5.7 of the guide.