Skip to content

TechnicalExpert

Simulator: quality SLO and Wilson bound

For: engineers and SREs · business and productPrerequisites: Notions of SLOs and basic statistics (proportion, confidence interval).

Reading mode

A quality SLO, for instance on the faithfulness of a generative assistant’s answers, is measured on an evaluated sample: k answers above the threshold out of n. The observed proportion k / n is only an estimate. Comparing that point value to the target leads to declaring the SLO met by luck on a small sample. The right comparison uses the lower bound of the confidence interval, computed here with the Wilson interval, which is robust on small samples and near 0 or 1.

Sample
SLO and precision

Result

Observed proportion
Wilson interval
Verdict
Size for margin e
n ≈ (z / e)² p (1 - p), observed p
Size to demonstrate the target
if the observed proportion holds
Confidence interval and target
How the interval narrows as n grows, at constant proportion
Wilson intervalObserved proportionTargetcurrent n

Wilson score interval, without continuity correction. Assumption: independent responses and a representative sample; stratified sampling requires per-stratum weighting. The sizing formula n ≈ (z / e)² p (1 - p) is the classic normal approximation: it underestimates the need when p is close to 0 or 1.

  1. The defaults: 475 answers out of 500, exactly 95%, for a 95% target. The interval runs from about 92.7% to 96.6%: the target lies inside, the SLO is “not demonstrated”. A proportion equal to the target never proves the target.
  2. Set k = 49 and n = 50: 98% observed, but the lower bound drops to about 89.5%. An excellent score on a small sample demonstrates nothing.
  3. Go back to n = 500 and set k = 460: 92% observed, upper bound around 94.1%, below target. The SLO is violated, and this time the sample is large enough to say so.
  4. Type 96 in the proportion field (n = 500): the SLO is not demonstrated yet. The “Size to demonstrate the target” tile shows you would need about 1,800 evaluated answers if that proportion holds. Switch confidence to 99%: you need more than 3,000.
  • Met if the lower bound is above the target;
  • violated if the upper bound is below it;
  • not demonstrated otherwise: you need more data, not a conclusion.

The Wilson interval, with z = 1.645, 1.96 or 2.576 for 90, 95 or 99% confidence:

center = (p + z² / 2n) / (1 + z² / n)
half-width = z / (1 + z² / n) × sqrt( p (1 - p) / n + z² / 4n² )

To size the sample around a proportion p with a margin e, the approximation n ≈ (z / e)² p (1 - p) gives the order of magnitude: about 460 answers for p = 0.95 and e = 2 points at 95%.

The full method, including how to run the bound in a dedicated component and quality error budgets, is described in the SRE core: quality SLOs. See also the error budget simulator, the method calculator and the LLM, RAG and MCP page.