Cross-cutting
Annexes
For: engineers and SREs · architectsPrerequisites: Have gone through the parts of the method.
Annex A. Metrics summary per component and axis
Section titled “Annex A. Metrics summary per component and axis”| Component | Reliability | Availability | Cost |
|---|---|---|---|
| LLM | finish_reasons, refusal rate, hallucination, drift | TTFT, E2E latency, error ratio, GPU and KV cache saturation | tokens in/out, cost per request, cache hit |
| RAG | context precision and recall, faithfulness, relevance, coverage | retrieval latency, index health, base freshness | embeddings cost, vector store cost, reranking |
| Agent | task success, tool selection, planning, guardrails | latency per step and per task, tool failure rate | steps per task, cost per task, loop detection |
| MCP | server mutations, tool poisoning, auth failures, injections | server availability, session health, tool latency | invocation rate-limiting |
Annex B. Tools
Section titled “Annex B. Tools”The tools landscape (self-hosted or managed, reference stack, selection criteria, licenses) is in chapter 6 of the guide. Tools cited by the method and absent from that chapter: MCP security (mcp-scan, Cisco mcp-scanner, MCP gateway with audit), GPU metrics (DCGM exporter), inference metrics (Prometheus metrics from vLLM and TGI), the Laminar and Confident AI evaluation platforms, multi-provider cost tracking (Helicone).
Annex C. Method-specific terms
Section titled “Annex C. Method-specific terms”General terms (TTFT, faithfulness, LLM-as-judge, golden set, burn rate, Wilson bound, tool poisoning, silent failure, OTLP) are defined in the site glossary. Only the ones it does not cover remain here:
- TPOT: time per output token, average time per output token after the first, computed per request: (total latency minus TTFT) divided by (output tokens minus one).
- ITL: inter-token latency, the delay between two consecutive tokens, measured for each pair; track its distribution (p95, p99) to see the stutters that the TPOT average smooths out.
- KV cache: the inference server key-value cache, decisive for latency.
- Context precision and recall: retrieval quality in a RAG.
- Rug pull: a malicious update of a trusted MCP tool.
- Tool shadowing: a fake tool duplicating a legitimate one.
- Spanmetrics: derivation of metrics from spans, at the OTel collector.
Annex D. Building your dated price table
Section titled “Annex D. Building your dated price table”Model prices change often and depend on the provider, the region, the call mode (synchronous, batch) and caching. Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. Build your own table, versioned like code, with these fields:
| Field | Content |
|---|---|
| Model | exact identifier, as it appears in gen_ai.request.model |
| Provider | value of gen_ai.provider.name |
| Input price | per million tokens |
| Output price | per million tokens |
| Cached input price | if the provider bills prompt cache reads separately |
| Currency | currency of the price list; if converted, rate and rate date |
| Source | URL of the provider’s official price list or contract reference |
| Checked on | day the price was verified |
| Effective from | when the price applies, to recompute the past with the right price |
Template to fill in from your provider’s dated price list:
| Model | Provider | Input (€ / M tokens) | Output (€ / M tokens) | Currency | Source | Checked on |
|---|---|---|---|---|---|---|
| demo-llm (large model) | provider-a | to fill in | to fill in | EUR | to fill in | to fill in |
| demo-llm-small (small model) | provider-a | to fill in | to fill in | EUR | to fill in | to fill in |
Conversion for the cost rules: a price of P € per million tokens is P x 1e-6 € per token. A job reads the table and publishes the price_in_eur_per_token and price_out_eur_per_token series used by the Part II recording rules; the same values serve as inputs to the exposure calculator.
Revised on 2 October 2026: prices removed from annex D, replaced with a template to fill in.
Revised on 4 October 2026: annex B points to chapter 6 of the guide and keeps only the tools it does not cite; annex C points to the site glossary and keeps only method-specific terms; annex D, which contains no prices, is kept.