Technical
Lab 4: cost, performance, SLOs and alerts
For: engineers and SREs · architectsPrerequisites: Have completed labs 1 to 3 of the course.
Duration: 1 h 35, plus the time spent waiting for alerts (up to 30 minutes for the quality rule). Prerequisites: labs 1 to 3.
Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list.
Step 1: simulated traffic (15 min)
Section titled “Step 1: simulated traffic (15 min)”cd app && python lab4_cost_quality.pyThe script simulates 80 requests mixing 4 customers, 6 models and 4 features. No model is called: tokens, latencies and errors are drawn at random. Options: --n (number of requests), --interval (seconds between two requests), --tenant (all traffic on one customer), --token-factor (multiplies the drawn tokens, like a growing context), --latency-factor, --error-rate, --prices (price list, see below).
Input tokens are drawn uniformly between 80 and 1,200 (640 on average), output tokens between 40 and 600 (320 on average): 960 tokens per successful request. A failed request (1.5% by default) consumes no tokens, hence about 946 tokens per request on average, errors included.
Measure in tokens, convert to euros with your price list
Section titled “Measure in tokens, convert to euros with your price list”The lab budget and SLOs are expressed in tokens: it is the unit in which most API providers bill, it also measures the load of a self-hosted model and it does not depend on any price. Without a price list, the lab works fully and exports no cost metric.
To get a cost in euros, copy prices.example.json (at the root of the kit, with no values) to prices.json, fill in the version, the effective date, the source, the number of tokens the prices apply to and an input and output price for each of the six simulated models, then:
cd app && python lab4_cost_quality.py --prices ../prices.jsonThe script then computes, for each successful request, cost = t_in x p_in + t_out x p_out, where t_in and t_out are the input and output tokens, and p_in and p_out your prices brought down to one token. It exports mttl_client_cost_total with the mttl_pricing_version label (version and date of the price list): you thus know which price list produced a past cost. An incomplete or undated price list is rejected, rather than producing a partial cost. A self-hosted model has no price per token but an infrastructure cost: enter your internal cost brought down to one token, or 0 if you allocate it separately.
Step 2: the Cost & Quality dashboard (25 min)
Section titled “Step 2: the Cost & Quality dashboard (25 min)”Identify the customer that consumes the most, the feature that consumes the most tokens and the average number of tokens per request: about 950 with the default settings. The euro panels, at the bottom of the dashboard, stay empty until you supply a price list. Question for discussion: how would you relate this consumption to the quality score of labs 2 and 3? (the kit does not link the two yet)
Step 3: define the SLOs (20 min)
Section titled “Step 3: define the SLOs (20 min)”The lab SLOs, for a customer support assistant. They are the same values as the SLO dictionary in the script and as the alert thresholds:
| Indicator | Objective | Associated alert or check |
|---|---|---|
p95 latency of chat operations | 3 s at most | over 3 s for 10 min |
| Quality score (simplified teaching SLO) | 0.6 average faithfulness over 1 h, at least | below 0.6 for 30 min |
| Error rate | 2% over 15 min, at most | over 2% for 15 min |
| Average tokens per request (input and output) | 2,000 at most | dashboard panel, red above 2,000 |
| Token budget per customer | 1,000,000 per rolling hour | over 1,000,000 tokens over the last hour, for 5 min |
Simulated latency follows a log-normal distribution with parameters 0.2 and 0.4: median of about 1.2 s, p95 of about 2.4 s, below the objective. A local model on a workstation without a GPU (labs 1 to 3) can exceed 3 s: the latency alert can then fire on those labs, and that is a real finding. In a contract, add a customer commitment column, looser than the internal objective. If you also track the euro cost, set that budget in your currency from your price list and keep the token budget, which does not move when prices change.
To see the tokens-per-request SLO fail, run python lab4_cost_quality.py --token-factor 2.5: the expected average rises to 946 x 2.5, about 2,360 tokens per request, and the script’s final table shows “HORS SLO” (out of SLO). Over twenty draws of 80 requests, the observed average stayed between 2,180 and 2,630.
The row that sets this table apart from a classic set of SLOs is the quality one. Without it, the SLOs see none of the failures specific to an LLM: a service that is available, fast and within budget can very well give wrong answers.
Step 4: five reference alerts (15 min, plus the wait)
Section titled “Step 4: five reference alerts (15 min, plus the wait)”The rules in grafana/provisioning/alerting/alerts.yaml are loaded when Grafana starts (MTTL folder). Grafana evaluates them every minute; a rule fires when its condition stays true for the whole stated duration. The script must therefore run long enough: that is what --interval is for.
| Alert | Condition | Action | How to trigger it |
|---|---|---|---|
| p95 latency | over 3 s for 10 min | check GPU load and prompt size | python lab4_cost_quality.py --n 960 --interval 1 --latency-factor 2 |
| Token budget per customer | over 1,000,000 tokens over the last hour, for 5 min | contact the customer, check context size and rate limiting | python lab4_cost_quality.py --n 2000 --interval 0.2 --tenant acme-corp |
| Degraded quality | average score below 0.6 over 1 h, for 30 min | suspect drift, run a judge on a sample | python lab3_drift_detection.py --phase data --repeat 2 |
| Error rate | over 2% over 15 min, for 15 min | incident analysis, check the model | python lab4_cost_quality.py --n 1200 --interval 1 --tenant acme-corp --error-rate 0.1 |
| Drift | at least 3 events of the same type in 30 min, for 5 min | review of the pipeline and the prompt, controlled A/B | python lab3_drift_detection.py --phase data --repeat 4 |
The calculation behind each command:
- Latency: with
--latency-factor 2, each latency is doubled; median around 2.4 s, p95 around 4.7 s. Estimated from the kit’s histogram bounds (2.56 and 5.12 s), the p95 is about 5.0 s, above 3 s. The query covers a rolling 5 minutes: 960 requests at one per second make 16 minutes of traffic, enough to hold the condition for 10 minutes. - Token budget: about 946 tokens per request, errors included; 2,000 requests on a single customer give about 2,000 x 946, or 1,890,000 tokens, nearly twice the 1,000,000 threshold. The alert reads
gen_ai_client_token_usage_sum, the sum of the token histogram, which theprometheusremotewrite0.108.0 exporter names without a unit suffix (the{token}unit, in braces, is ignored); this customer has 48 series there (6 models, 4 features, 2 token types). At one request every 0.2 s, the 2,000 requests spread over about 7 minutes, hence over many metric exports (one every 5 s). The sum exceeds 1,000,000 after about 1,060 requests (1,000,000 / 946), around 3 min 30; even if the first value of each of the 48 series were not counted, a few tens of thousands of tokens would be missing, a few seconds of traffic. Grafana sees the overrun at the next evaluation, then waits 5 minutes: the alert fires between 9 and 10 minutes after launch, and the sum stays above the threshold for the following hour. - Errors: 10% of errors drawn, on a single customer. The rate read by VictoriaMetrics is a little lower, because rate() ignores the first value of each new series, but it stays several times above 2% during the 20 minutes of the run.
- Quality: the data drift phase of lab 3 asks out-of-domain questions, whose faithfulness usually falls below 0.6 (the value depends on the model). The average covers the past hour: run the command on a stack with no other quality traffic in the hour, then wait about 30 minutes.
- Drift: each pass of the drifting data phase emits one event; four passes in the same process give four, above the threshold of three even if the first exported value does not count in the computed increase.
The token budget also counts the tokens of labs 1 to 3, which carry the same customer label: on a shared stack, run the command on a customer those labs have not used in the hour, or take their consumption into account.
Stopping Ollama triggers nothing in lab 4: its traffic is simulated. To see real errors, stop Ollama during lab 1 (step 5).
Check that each alert is routed to the right recipient.
Step 5: your 30-, 60- and 90-day plan (20 min)
Section titled “Step 5: your 30-, 60- and 90-day plan (20 min)”Write your plan using the template in module 5.
Lab recap
Section titled “Lab recap”| Notion | What the lab shows |
|---|---|
| Budget | in tokens, per request and per customer: it does not depend on any price |
| Euro cost | computed in the application only with your price list, versioned and dated |
| SLOs | always a quality SLO on top of the classic ones; the lab’s is a simplified teaching SLO |
| Alerts | every alert has an owner and a procedure |
| Order | instrument, then evaluate, then alert |
Revised on 2 October 2026, prices removed: box on the kit status, a single definition of the SLOs (script, lesson, alerts), budget and SLOs in tokens, per-customer token budget alert instead of the cost alert, euro cost only with a dated price list supplied by the user, simulator options (--n, --interval, --tenant, --token-factor, --latency-factor, --error-rate, --prices) and calculation of each trigger, error alert limited to one customer because of rate(), cost and quality question presented as open.
Revised on 4 October 2026: the lab’s quality SLO is labeled a simplified teaching SLO (threshold on a mean), box on the production quality SLO (proportion, Wilson bound) with a pointer to part IV of the method.