Skip to content

Cross-cuttingExpert

Annexes

For: engineers and SREs · architectsPrerequisites: Have gone through the parts of the method.

Annex A. Metrics summary per component and axis

Section titled “Annex A. Metrics summary per component and axis”
ComponentReliabilityAvailabilityCost
LLMfinish_reasons, refusal rate, hallucination, driftTTFT, E2E latency, error ratio, GPU and KV cache saturationtokens in/out, cost per request, cache hit
RAGcontext precision and recall, faithfulness, relevance, coverageretrieval latency, index health, base freshnessembeddings cost, vector store cost, reranking
Agenttask success, tool selection, planning, guardrailslatency per step and per task, tool failure ratesteps per task, cost per task, loop detection
MCPserver mutations, tool poisoning, auth failures, injectionsserver availability, session health, tool latencyinvocation rate-limiting

The tools landscape (self-hosted or managed, reference stack, selection criteria, licenses) is in chapter 6 of the guide. Tools cited by the method and absent from that chapter: MCP security (mcp-scan, Cisco mcp-scanner, MCP gateway with audit), GPU metrics (DCGM exporter), inference metrics (Prometheus metrics from vLLM and TGI), the Laminar and Confident AI evaluation platforms, multi-provider cost tracking (Helicone).

General terms (TTFT, faithfulness, LLM-as-judge, golden set, burn rate, Wilson bound, tool poisoning, silent failure, OTLP) are defined in the site glossary. Only the ones it does not cover remain here:

  • TPOT: time per output token, average time per output token after the first, computed per request: (total latency minus TTFT) divided by (output tokens minus one).
  • ITL: inter-token latency, the delay between two consecutive tokens, measured for each pair; track its distribution (p95, p99) to see the stutters that the TPOT average smooths out.
  • KV cache: the inference server key-value cache, decisive for latency.
  • Context precision and recall: retrieval quality in a RAG.
  • Rug pull: a malicious update of a trusted MCP tool.
  • Tool shadowing: a fake tool duplicating a legitimate one.
  • Spanmetrics: derivation of metrics from spans, at the OTel collector.

Model prices change often and depend on the provider, the region, the call mode (synchronous, batch) and caching. Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. Build your own table, versioned like code, with these fields:

FieldContent
Modelexact identifier, as it appears in gen_ai.request.model
Providervalue of gen_ai.provider.name
Input priceper million tokens
Output priceper million tokens
Cached input priceif the provider bills prompt cache reads separately
Currencycurrency of the price list; if converted, rate and rate date
SourceURL of the provider’s official price list or contract reference
Checked onday the price was verified
Effective fromwhen the price applies, to recompute the past with the right price

Template to fill in from your provider’s dated price list:

ModelProviderInput (€ / M tokens)Output (€ / M tokens)CurrencySourceChecked on
demo-llm (large model)provider-ato fill into fill inEURto fill into fill in
demo-llm-small (small model)provider-ato fill into fill inEURto fill into fill in

Conversion for the cost rules: a price of P € per million tokens is P x 1e-6 € per token. A job reads the table and publishes the price_in_eur_per_token and price_out_eur_per_token series used by the Part II recording rules; the same values serve as inputs to the exposure calculator.


Revised on 2 October 2026: prices removed from annex D, replaced with a template to fill in.

Revised on 4 October 2026: annex B points to chapter 6 of the guide and keeps only the tools it does not cite; annex C points to the site glossary and keeps only method-specific terms; annex D, which contains no prices, is kept.