Skip to content

TechnicalPractitioner

Lab 4: cost, performance, SLOs and alerts

For: engineers and SREs · architectsPrerequisites: Have completed labs 1 to 3 of the course.

Duration: 1 h 35, plus the time spent waiting for alerts (up to 30 minutes for the quality rule). Prerequisites: labs 1 to 3.

Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list.

Fenêtre de terminal
cd app && python lab4_cost_quality.py

The script simulates 80 requests mixing 4 customers, 6 models and 4 features. No model is called: tokens, latencies and errors are drawn at random. Options: --n (number of requests), --interval (seconds between two requests), --tenant (all traffic on one customer), --token-factor (multiplies the drawn tokens, like a growing context), --latency-factor, --error-rate, --prices (price list, see below).

Input tokens are drawn uniformly between 80 and 1,200 (640 on average), output tokens between 40 and 600 (320 on average): 960 tokens per successful request. A failed request (1.5% by default) consumes no tokens, hence about 946 tokens per request on average, errors included.

Measure in tokens, convert to euros with your price list

Section titled “Measure in tokens, convert to euros with your price list”

The lab budget and SLOs are expressed in tokens: it is the unit in which most API providers bill, it also measures the load of a self-hosted model and it does not depend on any price. Without a price list, the lab works fully and exports no cost metric.

To get a cost in euros, copy prices.example.json (at the root of the kit, with no values) to prices.json, fill in the version, the effective date, the source, the number of tokens the prices apply to and an input and output price for each of the six simulated models, then:

Fenêtre de terminal
cd app && python lab4_cost_quality.py --prices ../prices.json

The script then computes, for each successful request, cost = t_in x p_in + t_out x p_out, where t_in and t_out are the input and output tokens, and p_in and p_out your prices brought down to one token. It exports mttl_client_cost_total with the mttl_pricing_version label (version and date of the price list): you thus know which price list produced a past cost. An incomplete or undated price list is rejected, rather than producing a partial cost. A self-hosted model has no price per token but an infrastructure cost: enter your internal cost brought down to one token, or 0 if you allocate it separately.

Step 2: the Cost & Quality dashboard (25 min)

Section titled “Step 2: the Cost & Quality dashboard (25 min)”

Identify the customer that consumes the most, the feature that consumes the most tokens and the average number of tokens per request: about 950 with the default settings. The euro panels, at the bottom of the dashboard, stay empty until you supply a price list. Question for discussion: how would you relate this consumption to the quality score of labs 2 and 3? (the kit does not link the two yet)

The lab SLOs, for a customer support assistant. They are the same values as the SLO dictionary in the script and as the alert thresholds:

IndicatorObjectiveAssociated alert or check
p95 latency of chat operations3 s at mostover 3 s for 10 min
Quality score (simplified teaching SLO)0.6 average faithfulness over 1 h, at leastbelow 0.6 for 30 min
Error rate2% over 15 min, at mostover 2% for 15 min
Average tokens per request (input and output)2,000 at mostdashboard panel, red above 2,000
Token budget per customer1,000,000 per rolling hourover 1,000,000 tokens over the last hour, for 5 min

Simulated latency follows a log-normal distribution with parameters 0.2 and 0.4: median of about 1.2 s, p95 of about 2.4 s, below the objective. A local model on a workstation without a GPU (labs 1 to 3) can exceed 3 s: the latency alert can then fire on those labs, and that is a real finding. In a contract, add a customer commitment column, looser than the internal objective. If you also track the euro cost, set that budget in your currency from your price list and keep the token budget, which does not move when prices change.

To see the tokens-per-request SLO fail, run python lab4_cost_quality.py --token-factor 2.5: the expected average rises to 946 x 2.5, about 2,360 tokens per request, and the script’s final table shows “HORS SLO” (out of SLO). Over twenty draws of 80 requests, the observed average stayed between 2,180 and 2,630.

The row that sets this table apart from a classic set of SLOs is the quality one. Without it, the SLOs see none of the failures specific to an LLM: a service that is available, fast and within budget can very well give wrong answers.

Step 4: five reference alerts (15 min, plus the wait)

Section titled “Step 4: five reference alerts (15 min, plus the wait)”

The rules in grafana/provisioning/alerting/alerts.yaml are loaded when Grafana starts (MTTL folder). Grafana evaluates them every minute; a rule fires when its condition stays true for the whole stated duration. The script must therefore run long enough: that is what --interval is for.

AlertConditionActionHow to trigger it
p95 latencyover 3 s for 10 mincheck GPU load and prompt sizepython lab4_cost_quality.py --n 960 --interval 1 --latency-factor 2
Token budget per customerover 1,000,000 tokens over the last hour, for 5 mincontact the customer, check context size and rate limitingpython lab4_cost_quality.py --n 2000 --interval 0.2 --tenant acme-corp
Degraded qualityaverage score below 0.6 over 1 h, for 30 minsuspect drift, run a judge on a samplepython lab3_drift_detection.py --phase data --repeat 2
Error rateover 2% over 15 min, for 15 minincident analysis, check the modelpython lab4_cost_quality.py --n 1200 --interval 1 --tenant acme-corp --error-rate 0.1
Driftat least 3 events of the same type in 30 min, for 5 minreview of the pipeline and the prompt, controlled A/Bpython lab3_drift_detection.py --phase data --repeat 4

The calculation behind each command:

  • Latency: with --latency-factor 2, each latency is doubled; median around 2.4 s, p95 around 4.7 s. Estimated from the kit’s histogram bounds (2.56 and 5.12 s), the p95 is about 5.0 s, above 3 s. The query covers a rolling 5 minutes: 960 requests at one per second make 16 minutes of traffic, enough to hold the condition for 10 minutes.
  • Token budget: about 946 tokens per request, errors included; 2,000 requests on a single customer give about 2,000 x 946, or 1,890,000 tokens, nearly twice the 1,000,000 threshold. The alert reads gen_ai_client_token_usage_sum, the sum of the token histogram, which the prometheusremotewrite 0.108.0 exporter names without a unit suffix (the {token} unit, in braces, is ignored); this customer has 48 series there (6 models, 4 features, 2 token types). At one request every 0.2 s, the 2,000 requests spread over about 7 minutes, hence over many metric exports (one every 5 s). The sum exceeds 1,000,000 after about 1,060 requests (1,000,000 / 946), around 3 min 30; even if the first value of each of the 48 series were not counted, a few tens of thousands of tokens would be missing, a few seconds of traffic. Grafana sees the overrun at the next evaluation, then waits 5 minutes: the alert fires between 9 and 10 minutes after launch, and the sum stays above the threshold for the following hour.
  • Errors: 10% of errors drawn, on a single customer. The rate read by VictoriaMetrics is a little lower, because rate() ignores the first value of each new series, but it stays several times above 2% during the 20 minutes of the run.
  • Quality: the data drift phase of lab 3 asks out-of-domain questions, whose faithfulness usually falls below 0.6 (the value depends on the model). The average covers the past hour: run the command on a stack with no other quality traffic in the hour, then wait about 30 minutes.
  • Drift: each pass of the drifting data phase emits one event; four passes in the same process give four, above the threshold of three even if the first exported value does not count in the computed increase.

The token budget also counts the tokens of labs 1 to 3, which carry the same customer label: on a shared stack, run the command on a customer those labs have not used in the hour, or take their consumption into account.

Stopping Ollama triggers nothing in lab 4: its traffic is simulated. To see real errors, stop Ollama during lab 1 (step 5).

Check that each alert is routed to the right recipient.

Step 5: your 30-, 60- and 90-day plan (20 min)

Section titled “Step 5: your 30-, 60- and 90-day plan (20 min)”

Write your plan using the template in module 5.

NotionWhat the lab shows
Budgetin tokens, per request and per customer: it does not depend on any price
Euro costcomputed in the application only with your price list, versioned and dated
SLOsalways a quality SLO on top of the classic ones; the lab’s is a simplified teaching SLO
Alertsevery alert has an owner and a procedure
Orderinstrument, then evaluate, then alert

Revised on 2 October 2026, prices removed: box on the kit status, a single definition of the SLOs (script, lesson, alerts), budget and SLOs in tokens, per-customer token budget alert instead of the cost alert, euro cost only with a dated price list supplied by the user, simulator options (--n, --interval, --tenant, --token-factor, --latency-factor, --error-rate, --prices) and calculation of each trigger, error alert limited to one customer because of rate(), cost and quality question presented as open.

Revised on 4 October 2026: the lab’s quality SLO is labeled a simplified teaching SLO (threshold on a mean), box on the production quality SLO (proportion, Wilson bound) with a pointer to part IV of the method.