Technical
LLM metrics taxonomy and Python instrumentation
For: engineers and SREsPrerequisites: Have read lesson 2 of the course and be able to read Python.
Module 5: LLM metrics taxonomy
Section titled “Module 5: LLM metrics taxonomy”The taxonomy covers the four dimensions of LLM observability: response quality, performance and latency, costs and tokens, and GPU/CPU infrastructure.
5.1 Recommended naming conventions
Section titled “5.1 Recommended naming conventions”- Prefix
llm_for all LLM-specific metrics. - Prefix
rag_for retrieval-specific metrics. - Use the standard DCGM names for GPU metrics (
DCGM_FI_*). - Standard suffixes:
_total(counter),_secondsor_ms(latency histogram),_ratio(a unitless value between 0 and 1, whatever the type: gauge or histogram, such asrag_context_pressure_ratio),_bytes(gauge).
| Metric | Prometheus type | Cardinality | Essential labels |
|---|---|---|---|
llm_requests_total | Counter | Medium | model, provider, pipeline_id, status |
llm_request_duration_ms | Histogram | Medium | model, provider, pipeline_id |
llm_tokens_total{type} | Counter | Low | model, provider, type=input or output |
llm_cost_usd_total | Counter | Low | model, provider, tenant_id, use_case |
llm_errors_total | Counter | Medium | model, provider, error_type |
llm_ttft_ms | Histogram | Medium | model, provider |
rag_retrieval_score_avg | Gauge | Low | index_id, pipeline_id, strategy |
rag_context_pressure_ratio | Histogram | Low | pipeline_id, model |
rag_chunks_dropped_total | Counter | Low | pipeline_id, reason |
DCGM_FI_DEV_GPU_UTIL | Gauge | Low | gpu, UUID, pod, namespace, node |
vllm:num_requests_running | Gauge | Low | model_name, instance |
Module 6: Python instrumentation, pushing metrics
Section titled “Module 6: Python instrumentation, pushing metrics”6.1 Two complementary approaches
Section titled “6.1 Two complementary approaches”- Prometheus client library: a
/metricsendpoint is exposed and scraped by vmagent. Simple and widely used. - OTLP via the OpenTelemetry SDK: metrics and traces in the same pipeline. Recommended for new applications.
# Option 1: prometheus_clientpip install prometheus-client
# Option 2: OTel SDK (unified metrics and traces)pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc6.2 Prometheus client: core LLM metrics
Section titled “6.2 Prometheus client: core LLM metrics”from prometheus_client import Counter, Histogram, Gauge, start_http_serverimport time
LLM_REQUESTS = Counter( 'llm_requests_total', 'Total LLM requests', ['model', 'provider', 'pipeline_id', 'status'])LLM_ERRORS = Counter( 'llm_errors_total', 'Total failed LLM requests', ['model', 'provider', 'error_type'])LLM_DURATION = Histogram( 'llm_request_duration_ms', 'LLM request duration in ms', ['model', 'provider', 'pipeline_id'], buckets=[50, 100, 200, 500, 1000, 2000, 5000, 10000, 30000])LLM_TOKENS = Counter( 'llm_tokens_total', 'Total tokens processed', ['model', 'provider', 'type'] # type: input or output)LLM_COST = Counter( 'llm_cost_usd_total', 'Total cost in USD', ['model', 'provider', 'tenant_id', 'use_case'])LLM_TTFT = Histogram( 'llm_ttft_ms', 'Time to first token in ms', ['model', 'provider'], buckets=[50, 100, 200, 500, 1000, 2000, 5000])RAG_RETRIEVAL_SCORE = Gauge( 'rag_retrieval_score_avg', 'Average retrieval similarity score', ['index_id', 'pipeline_id', 'strategy'])RAG_CONTEXT_PRESSURE = Histogram( 'rag_context_pressure_ratio', 'Context window pressure ratio (0-1)', ['pipeline_id', 'model'], buckets=[0.1, 0.3, 0.5, 0.7, 0.8, 0.85, 0.9, 0.95, 1.0])RAG_CHUNKS_DROPPED = Counter( 'rag_chunks_dropped_total', 'Chunks dropped from the context', ['pipeline_id', 'reason'])
# Expose /metrics on port 8000 (9090 is Prometheus' usual port)start_http_server(8000)6.3 Instrumentation decorator: automatic cost tracking
Section titled “6.3 Instrumentation decorator: automatic cost tracking”import functoolsimport jsonimport os
# Price grid in USD per token, read from a versioned JSON file that you supply;# no values here. To fill in from your provider's dated price list:# {"effective_date": "YYYY-MM-DD", "currency": "USD",# "models": {"large-model": {"input": null, "output": null}}}# Without a file, TOKEN_PRICES stays empty: the decorator counts tokens and does# not emit llm_cost_usd_total.PRICE_GRID_FILE = os.getenv('LLM_PRICE_GRID_FILE')PRICE_GRID_EFFECTIVE_DATE = NoneTOKEN_PRICES = {}if PRICE_GRID_FILE: with open(PRICE_GRID_FILE, encoding='utf-8') as fh: _grid = json.load(fh) PRICE_GRID_EFFECTIVE_DATE = _grid['effective_date'] TOKEN_PRICES = {m: p for m, p in _grid['models'].items() if p.get('input') is not None and p.get('output') is not None}
def track_llm_call(model, provider, pipeline_id, tenant_id='default', use_case='chat'): def decorator(func): @functools.wraps(func) async def wrapper(*args, **kwargs): labels = {'model': model, 'provider': provider, 'pipeline_id': pipeline_id} t0 = time.monotonic() status = 'success' try: result = await func(*args, **kwargs) if hasattr(result, 'usage'): in_tok = result.usage.prompt_tokens out_tok = result.usage.completion_tokens LLM_TOKENS.labels(model=model, provider=provider, type='input').inc(in_tok) LLM_TOKENS.labels(model=model, provider=provider, type='output').inc(out_tok) if model in TOKEN_PRICES: p = TOKEN_PRICES[model] cost = in_tok * p['input'] + out_tok * p['output'] LLM_COST.labels(model=model, provider=provider, tenant_id=tenant_id, use_case=use_case).inc(cost) return result except Exception as e: status = type(e).__name__ LLM_ERRORS.labels(model=model, provider=provider, error_type=status).inc() raise finally: LLM_DURATION.labels(**labels).observe((time.monotonic() - t0) * 1000) LLM_REQUESTS.labels(**labels, status=status).inc() return wrapper return decorator
@track_llm_call(model='large-model', provider='provider-a', pipeline_id='support-rag', tenant_id='acme')async def call_llm(messages: list) -> dict: # client: your provider's SDK client; use the exact model identifier return await client.chat.completions.create(model='model-identifier', messages=messages)The status label takes the exception class name: keep the number of distinct exception types small, or map them to a short list of categories, to bound cardinality.
Revised on 2 October 2026: prices removed; the decorator’s price grid is read from a file you supply, with no values in the guide, and cost is emitted only for models whose price is filled in.