Technical
Part IV. The SRE core: quality SLOs and error budgets
For: engineers and SREs · architects · team managersPrerequisites: Have read part I; notions of SLOs, error budgets and burn rate.
This is the most specific contribution of the method. The SRE discipline of SLOs and error budgets applies, but quality is probabilistic and measured by sampling, which changes everything.
12. Deterministic versus probabilistic SLIs
Section titled “12. Deterministic versus probabilistic SLIs”Three SLI families:
- Service (deterministic): availability, latency, error ratio. Measured on 100 percent of traffic, as in classical SRE.
- Quality (probabilistic): faithfulness, hallucination, refusal, relevance. Measured on a sample, by an evaluator that is itself non-deterministic.
- Cost: average cost per request, session, task and drift.
The fundamental difference: a quality SLI is an estimate carrying sampling uncertainty. Treating it as an exact measure is a frequent methodological fault, covered as an anti-pattern (Part VIII, pitfall 3).
13. Defining a quality SLO
Section titled “13. Defining a quality SLO”This part is the site’s reference for quality SLOs. The simplified thresholds shown elsewhere, for example in the labs or the guide, are pedagogical SLOs that point here.
Why a threshold on a mean is not a quality SLO. A mean of scores (“average faithfulness above 0.8”) does not tell you how many answers are bad: a few very bad answers hide behind many good ones, and the mean barely moves when the tail degrades. Nor does it provide a budget: you cannot derive an error budget or a burn rate from it. An SLO counts good events among valid events. For quality, the good event is an answer whose score reaches a threshold, so the SLI is a proportion and, since that proportion is estimated on a sample, you compare it to the target through its lower confidence bound.
Define a quality threshold tau (for example a minimum faithfulness score), then:
SLI_quality = (number of sampled responses with score >= tau) / (sample size)Since this is a proportion estimated over n observations, you compare to the target not the point value but the lower bound of the confidence interval at 95 percent (Wilson bound, robust on small samples):
SLO met <=> lower_bound_95(SLI_quality, n) >= targetSample size: for a margin of error e at 95 percent around a proportion p, size by n ~= (1.96 / e)^2 * p * (1 - p). Example: to estimate a faithfulness rate near 0.95 to within plus or minus 2 points, you need on the order of 450 to 500 evaluated responses per window. Stratified sampling is recommended: oversample the at-risk segments (sensitive requests, new flows) rather than a uniform draw.
What sets the precision is the number of judgments per segment and per window, not the percentage of traffic. Illustrative example: 10 percent of a segment of 200 requests a day gives 20 judgments, far too few for the margin targeted above; 1 percent of a flow of 100,000 requests gives 1,000. The per-tier percentages in the guide, 5.7 are examples to convert this way: start from the margin you want, derive n per segment and per window, then each segment’s sampling rate.
Typical wording of a quality SLO: “at least 95 percent of responses, estimated on a stratified sample of about 500 per day, score above the faithfulness threshold, with the lower bound of the 95 percent confidence interval above target, over a 7-day rolling window”.
14. Error budget and burn rate on quality
Section titled “14. Error budget and burn rate on quality”The error budget applies to quality, not only to availability:
quality_error_budget = 1 - targetburn_rate = (1 - SLI_quality) / (1 - target)A burn_rate of 1 consumes the budget exactly at the planned pace, above that it depletes too fast. Multi-window alerting, adapted to the dynamics of quality:
- short window (for example 1 h, high burn rate): catches a sudden collapse, for example after a prompt deployment or a model change.
- long window (for example 24 h to 7 d, moderate burn rate): catches slow drift, which is the proper pathology of quality.
Latency is not the right analogue: quality moves slowly, its windows are longer than availability windows.
15. The evaluator problem (meta-observability)
Section titled “15. The evaluator problem (meta-observability)”The blind spot of most setups: the LLM-as-judge evaluator is itself a model, non-deterministic, and it drifts. If the judge drifts, the quality SLI is corrupted with nothing to indicate it. So you must observe the judge:
- golden set: a human-annotated reference set, stable.
- calibration: agreement between the judge and the golden set, preferably with a chance-corrected measure such as Cohen’s kappa (see the guide, 5.4), measured and tracked as a metric in its own right. A dropping agreement invalidates the quality measures.
- judge versioning: pin the judge model version and the evaluation prompt. Any judge change is an event to trace, on the same footing as a production model change.
- periodic recalibration: revalidate the judge against humans at a regular interval, and after any version change.
- ground truth economics: the bottleneck is human annotation. Strategy: build the golden set once, use it to calibrate the judge, let the judge scale and re-engage humans only for periodic revalidation and low-confidence cases.
Operationalizing the Wilson bound. It is not computed in MetricsQL. It lives in a quality SLO sidecar that reads the sampled scores (from the eval platform or the trace store), computes the bound over a rolling window and exposes it as an OTel metric. The SLO alert then reads that metric. End-to-end chain:
# quality SLO sidecar: Wilson lower bound exposed as an OTel metricfrom math import sqrt
def wilson_lower(k, n, z=1.96): if n == 0: return 0.0 p = k / n d = 1 + z*z/n centre = p + z*z/(2*n) margin = z * sqrt(p*(1-p)/n + z*z/(4*n*n)) return (centre - margin) / d
# k = sampled responses with score >= tau, n = sample size# emitted as quality_slo.faithfulness.lower_bound (OTel gauge), outside# the gen_ai.* namespace, reserved for the semantic conventionsslo_met = wilson_lower(k, n) >= targetThe alert fires on the lower bound dropping below target, not on the point value. The full lab of the method, in preparation (Part IX), will ship this sidecar: its hallucination scenario is meant to drive the bound below target where latency stays green.
16. Composite SLOs
Section titled “16. Composite SLOs”A RAG has a retrieval SLO and a generation SLO. The user-facing SLO is the conjunction: the answer is only good if retrieval brought the right context and generation respected it. Decomposing allows correct attribution of a degradation to the right link. An agent has a task success SLO, the conjunction of the tool and model SLOs over the trajectory.
17. Quality gating in continuous integration
Section titled “17. Quality gating in continuous integration”The offline counterpart of SLOs. Any change of prompt, model or tool version goes through an eval suite on a reference dataset before deployment. The rule: block the deployment if the suite regresses beyond budget. Tools: DeepEval, RAGAS, Promptfoo. Principle: a light eval that runs on every pull request beats a full quarterly eval you forget to run.
Revised on 4 October 2026: this part becomes the site’s reference for quality SLOs; added why a threshold on a mean is not an SLO, the judge sampling principle (number of judgments per segment and per window, with a pointer to guide 5.7) and chance-corrected agreement for calibration.