Technical
Module 4: continuous evaluation in production
For: engineers and SREsPrerequisites: Have read module 3 of the course on the three drifts.
Online or offline
Section titled “Online or offline”The two complement each other and answer different questions.
| Online | Offline | |
|---|---|---|
| On what | real traffic, sampled (for example 1 to 5%, see below) | a reference set (order of magnitude: a few hundred to a thousand question and answer pairs) |
| When | continuously | at each version, before deployment |
| Cost | kept under control by sampling, since each judgment consumes tokens | higher, with a human in the loop |
| Strength | fast detection, near real-time alerting | rigorous comparison between versions |
| Limit | no oracle, only another model | does not see drifts after release to production |
In practice, you evaluate offline before each release to production and online continuously.
How many answers should be judged online? The principle used across the site: what makes a score precise is the number of judgments per segment and per period (per feature or per customer, per hour or per day), not the percentage of traffic. The 1 to 5% quoted here are examples: on heavy traffic, 1% already yields many judgments; on a rarely used segment, even 100% may yield only a few per day. The design of sampling (evaluator tiers, traces that are always evaluated) has its reference in section 5.7 of the guide; this module keeps what is specific to the labs: judge biases and RAG diagnosis.
The two evaluations form a single loop:
flowchart TB
V["New version"] --> OFF{"Offline<br/>reference set"}
OFF -->|"regression"| FX["Fix"]
FX --> V
OFF -->|"compliant"| PR["Production"]
PR --> ON["Online<br/>1 to 5% sample,<br/>different judge"]
PR --> FB["User feedback<br/>linked to the trace_id"]
ON --> AL{"Score<br/>dropping"}
FB --> AL
AL -->|"yes"| EN["Investigation<br/>on the trace"]
EN -.->|"new cases"| OFF
The LLM as a judge
Section titled “The LLM as a judge”Having a model grade the answers is fast and yields a score you can use as a metric, but it is expensive and non-deterministic: pinning the judge model version and setting temperature to zero improve reproducibility without guaranteeing it. Agreement with human judgment can be high: on comparisons of chatbot answers, Zheng et al. measure over 80% agreement between GPT-4 and human evaluators, as much as between humans (Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, 2023). This result does not carry over as is to your judge or your criteria: measure agreement on a sample you have had annotated.
Its biases are known and must be offset:
| Bias | Effect |
|---|---|
| verbosity | prefers long answers |
| self-preference | prefers its own style |
| position | favors the first option in an A/B comparison |
| confidence | tolerates a well-written hallucination |
| hidden cost | the judge’s tokens sometimes exceed those of the evaluated model |
To offset them, use a structured evaluation prompt, a rigid output format, a judge different from the evaluated model, and cross-sampling with human annotations.
Four RAG-specific metrics
Section titled “Four RAG-specific metrics”| Metric | Question | Tools |
|---|---|---|
| Faithfulness | are the claims in the answer supported by the retrieved context? | RAGAS, Phoenix |
| Context recall | does the context contain the needed information? | RAGAS, ground truth |
| Context precision | is the context dense or noisy? | RAGAS |
| Answer relevance | does the answer address the question asked? | RAGAS, LLM judge |
Low faithfulness with low context recall points to retrieval. Low faithfulness with high context recall points to generation.
flowchart LR
F["Low faithfulness"] --> R{"Context recall"}
R -->|"low"| RET["Look at retrieval"]
R -->|"high"| GEN["Look at generation"]
Detecting hallucinations
Section titled “Detecting hallucinations”Three levels, to combine according to risk:
- Rules: a figure, a date or a proper name in the answer that does not appear in the context.
- Specialized models: entailment classifiers between context and answer.
- Weak signals: negative feedback, immediate rewordings of the same question, abandonment.
User feedback
Section titled “User feedback”Thumbs up or down is the cheapest and most direct signal. Link it to the trace through its trace_id: negative feedback then becomes an entry point to the exact context that produced the bad answer.
In summary
Section titled “In summary”- Offline evaluation gates each release to production; online evaluation runs continuously.
- An LLM judge is only useful if it differs from the evaluated model and its biases are offset.
- Read together, faithfulness and context recall tell whether the problem comes from retrieval or generation.
- User feedback linked to its trace gives the exact context of the error, which no aggregate score provides.
Revised on 2 October 2026: the unsourced correlation figure between judge and humans is replaced by the study by Zheng et al. (2023), sampling rates and reference set size are presented as orders of magnitude.
Revised on 4 October 2026: judge sampling principle stated (the number of judgments per segment and per period matters more than the percentage), rates presented as examples and pointer to section 5.7 of the guide.