Skip to content

TechnicalPractitioner

LLM metrics taxonomy and Python instrumentation

For: engineers and SREsPrerequisites: Have read lesson 2 of the course and be able to read Python.

The taxonomy covers the four dimensions of LLM observability: response quality, performance and latency, costs and tokens, and GPU/CPU infrastructure.

  • Prefix llm_ for all LLM-specific metrics.
  • Prefix rag_ for retrieval-specific metrics.
  • Use the standard DCGM names for GPU metrics (DCGM_FI_*).
  • Standard suffixes: _total (counter), _seconds or _ms (latency histogram), _ratio (a unitless value between 0 and 1, whatever the type: gauge or histogram, such as rag_context_pressure_ratio), _bytes (gauge).
MetricPrometheus typeCardinalityEssential labels
llm_requests_totalCounterMediummodel, provider, pipeline_id, status
llm_request_duration_msHistogramMediummodel, provider, pipeline_id
llm_tokens_total{type}CounterLowmodel, provider, type=input or output
llm_cost_usd_totalCounterLowmodel, provider, tenant_id, use_case
llm_errors_totalCounterMediummodel, provider, error_type
llm_ttft_msHistogramMediummodel, provider
rag_retrieval_score_avgGaugeLowindex_id, pipeline_id, strategy
rag_context_pressure_ratioHistogramLowpipeline_id, model
rag_chunks_dropped_totalCounterLowpipeline_id, reason
DCGM_FI_DEV_GPU_UTILGaugeLowgpu, UUID, pod, namespace, node
vllm:num_requests_runningGaugeLowmodel_name, instance

Module 6: Python instrumentation, pushing metrics

Section titled “Module 6: Python instrumentation, pushing metrics”
  • Prometheus client library: a /metrics endpoint is exposed and scraped by vmagent. Simple and widely used.
  • OTLP via the OpenTelemetry SDK: metrics and traces in the same pipeline. Recommended for new applications.
Fenêtre de terminal
# Option 1: prometheus_client
pip install prometheus-client
# Option 2: OTel SDK (unified metrics and traces)
pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc
from prometheus_client import Counter, Histogram, Gauge, start_http_server
import time
LLM_REQUESTS = Counter(
'llm_requests_total', 'Total LLM requests',
['model', 'provider', 'pipeline_id', 'status']
)
LLM_ERRORS = Counter(
'llm_errors_total', 'Total failed LLM requests',
['model', 'provider', 'error_type']
)
LLM_DURATION = Histogram(
'llm_request_duration_ms', 'LLM request duration in ms',
['model', 'provider', 'pipeline_id'],
buckets=[50, 100, 200, 500, 1000, 2000, 5000, 10000, 30000]
)
LLM_TOKENS = Counter(
'llm_tokens_total', 'Total tokens processed',
['model', 'provider', 'type'] # type: input or output
)
LLM_COST = Counter(
'llm_cost_usd_total', 'Total cost in USD',
['model', 'provider', 'tenant_id', 'use_case']
)
LLM_TTFT = Histogram(
'llm_ttft_ms', 'Time to first token in ms',
['model', 'provider'],
buckets=[50, 100, 200, 500, 1000, 2000, 5000]
)
RAG_RETRIEVAL_SCORE = Gauge(
'rag_retrieval_score_avg', 'Average retrieval similarity score',
['index_id', 'pipeline_id', 'strategy']
)
RAG_CONTEXT_PRESSURE = Histogram(
'rag_context_pressure_ratio', 'Context window pressure ratio (0-1)',
['pipeline_id', 'model'],
buckets=[0.1, 0.3, 0.5, 0.7, 0.8, 0.85, 0.9, 0.95, 1.0]
)
RAG_CHUNKS_DROPPED = Counter(
'rag_chunks_dropped_total', 'Chunks dropped from the context',
['pipeline_id', 'reason']
)
# Expose /metrics on port 8000 (9090 is Prometheus' usual port)
start_http_server(8000)

6.3 Instrumentation decorator: automatic cost tracking

Section titled “6.3 Instrumentation decorator: automatic cost tracking”
import functools
import json
import os
# Price grid in USD per token, read from a versioned JSON file that you supply;
# no values here. To fill in from your provider's dated price list:
# {"effective_date": "YYYY-MM-DD", "currency": "USD",
# "models": {"large-model": {"input": null, "output": null}}}
# Without a file, TOKEN_PRICES stays empty: the decorator counts tokens and does
# not emit llm_cost_usd_total.
PRICE_GRID_FILE = os.getenv('LLM_PRICE_GRID_FILE')
PRICE_GRID_EFFECTIVE_DATE = None
TOKEN_PRICES = {}
if PRICE_GRID_FILE:
with open(PRICE_GRID_FILE, encoding='utf-8') as fh:
_grid = json.load(fh)
PRICE_GRID_EFFECTIVE_DATE = _grid['effective_date']
TOKEN_PRICES = {m: p for m, p in _grid['models'].items()
if p.get('input') is not None and p.get('output') is not None}
def track_llm_call(model, provider, pipeline_id, tenant_id='default', use_case='chat'):
def decorator(func):
@functools.wraps(func)
async def wrapper(*args, **kwargs):
labels = {'model': model, 'provider': provider, 'pipeline_id': pipeline_id}
t0 = time.monotonic()
status = 'success'
try:
result = await func(*args, **kwargs)
if hasattr(result, 'usage'):
in_tok = result.usage.prompt_tokens
out_tok = result.usage.completion_tokens
LLM_TOKENS.labels(model=model, provider=provider, type='input').inc(in_tok)
LLM_TOKENS.labels(model=model, provider=provider, type='output').inc(out_tok)
if model in TOKEN_PRICES:
p = TOKEN_PRICES[model]
cost = in_tok * p['input'] + out_tok * p['output']
LLM_COST.labels(model=model, provider=provider,
tenant_id=tenant_id, use_case=use_case).inc(cost)
return result
except Exception as e:
status = type(e).__name__
LLM_ERRORS.labels(model=model, provider=provider, error_type=status).inc()
raise
finally:
LLM_DURATION.labels(**labels).observe((time.monotonic() - t0) * 1000)
LLM_REQUESTS.labels(**labels, status=status).inc()
return wrapper
return decorator
@track_llm_call(model='large-model', provider='provider-a', pipeline_id='support-rag', tenant_id='acme')
async def call_llm(messages: list) -> dict:
# client: your provider's SDK client; use the exact model identifier
return await client.chat.completions.create(model='model-identifier', messages=messages)

The status label takes the exception class name: keep the number of distinct exception types small, or map them to a short list of categories, to bound cardinality.

Revised on 2 October 2026: prices removed; the decorator’s price grid is read from a file you supply, with no values in the guide, and cost is emitted only for models whose price is filled in.