Skip to content

TechnicalPractitioner

Lab 1: instrument an LLM application

For: engineers and SREsPrerequisites: Have read modules 1 and 2 of the course, know how to use Docker and read Python.

Duration: 1 h 45. Prerequisites: modules 1 and 2, Docker, Python 3.10 or later, 16 GB of memory recommended.

Fenêtre de terminal
cd labs/llm-observability
docker compose up -d
# Local models (5 to 10 minutes on first launch)
docker exec -it mttl-ollama ollama pull mistral:7b
docker exec -it mttl-ollama ollama pull nomic-embed-text
cd app && pip install -r requirements.txt
ServiceAddressNature
Grafanahttp://localhost:3000interface (admin / admin)
VictoriaMetricshttp://localhost:8428/vmuiquery interface
Tempohttp://localhost:3200HTTP API, no interface: traces are explored from Grafana
Phoenixhttp://localhost:6006interface

Image versions in the kit: Collector contrib 0.108.0, VictoriaMetrics v1.110.0, Tempo 2.6.0, Grafana 11.2.0, Phoenix version-5.9.1, Qdrant v1.12.0, Ollama 0.3.12. They are pinned in docker-compose.yml; Python dependencies are pinned in app/requirements.txt.

Open app/lab1_llm_app.py and locate:

  • init_telemetry(), which creates the trace and metric providers and the OTLP exporters;
  • the manual span around the model call;
  • the gen_ai.* attributes set on that span;
  • the instruments created by create_genai_metrics(): tokens and duration (standard), cost and errors (mttl.* extensions). Lab 1 emits no cost: the euro cost is computed only in lab 4, from a price list you supply.

The core of the instrumentation fits in a few lines:

labels = {
GEN_AI_PROVIDER_NAME: Provider.OLLAMA, # gen_ai.provider.name
GEN_AI_OPERATION_NAME: Operation.CHAT,
GEN_AI_REQUEST_MODEL: MODEL,
MTTL_TENANT_ID: tenant, # in-house extension: mttl.tenant.id
MTTL_FEATURE: feature, # in-house extension: mttl.feature
}
with tracer.start_as_current_span(f"chat {MODEL}", attributes={**labels, ...}) as span:
try:
response = get_ollama().chat(model=MODEL, messages=[...], options={...})
input_tokens = response.get("prompt_eval_count", 0)
output_tokens = response.get("eval_count", 0)
span.set_attribute(GEN_AI_USAGE_INPUT_TOKENS, input_tokens)
span.set_attribute(GEN_AI_USAGE_OUTPUT_TOKENS, output_tokens)
metrics["duration"].record(time.time() - start, labels)
metrics["tokens"].record(input_tokens, {**labels, GEN_AI_TOKEN_TYPE: "input"})
metrics["tokens"].record(output_tokens, {**labels, GEN_AI_TOKEN_TYPE: "output"})
except Exception as e:
err = {**labels, ERROR_TYPE: type(e).__name__}
span.record_exception(e)
span.set_status(trace.Status(trace.StatusCode.ERROR, str(e)))
metrics["duration"].record(time.time() - start, err)
metrics["errors"].add(1, err)
raise

The constants (GEN_AI_PROVIDER_NAME, MTTL_TENANT_ID…) are defined in app/shared/conventions.py: a single file to change when the convention evolves. Standard names stay in gen_ai.*, the kit extensions go in mttl.*.

In VictoriaMetrics, the Collector exporter adds unit and type suffixes: the duration becomes gen_ai_client_operation_duration_seconds_bucket, the error counter mttl_client_errors_total, and the mttl.tenant.id attribute becomes the mttl_tenant_id label. These are the names that dashboards and alerts query.

Fenêtre de terminal
python lab1_llm_app.py

In Grafana:

  1. Explore, Tempo source: look for the lab1-llm-app service and open a trace.
  2. Check that the provider, model and usage attributes are present.
  3. LLM Overview dashboard: tokens and latency should appear.
  1. Add the gen_ai.request.top_p attribute to the span.
  2. Vary the temperature from 0.0 to 1.0 and watch the spread of latency.
  3. Add a span event for the prompt: span.add_event("prompt", {...}). This capture is acceptable in this local lab; in production it stays off by default, is enabled only by an explicit choice (opt-in) and comes with masking of sensitive data (see module 2 and §2 of the article Observing an LLM system).

Step 4: several customers and features (25 min)

Section titled “Step 4: several customers and features (25 min)”

The script sends five requests spread across two customers (acme-corp, client-test) and three features (support, search, summary). Exercise to do yourself: add to the PROMPTS list a third customer (globex-energy) and a fourth feature (translation), run again, then check on the LLM Overview dashboard that metrics break down by customer and by feature (panels “Jetons par client” and “Jetons par fonctionnalité”, that is tokens per customer and per feature; model and tenant variables at the top of the dashboard).

Fenêtre de terminal
docker stop mttl-ollama
python lab1_llm_app.py

Watch the spans in error in Tempo (with the error.type attribute), the increment of the mttl_client_errors_total counter and the absence of consumed tokens. Then run docker start mttl-ollama again.

NotionWhat the lab shows
Tracesin this lab, an LLM call produces one span; child spans per step arrive in lab 2
Metricsalways labeled by model, customer and feature
Errorsspan.set_status(StatusCode.ERROR) plus a dedicated error counter
Costtracked here in tokens; in euros only in lab 4, with your versioned and dated price list
  • putting full prompts in attributes;
  • leaving a free-form customer identifier if you have thousands of them: cardinality explodes;
  • forgetting that auto-instrumentation exists for standard cases: the official OpenTelemetry project package for the OpenAI client is opentelemetry-instrumentation-openai-v2; opentelemetry-instrumentation-openai, without suffix, is the one from Traceloop’s OpenLLMetry project.

Revised on 2 October 2026: box on the kit status, gen_ai.system replaced by gen_ai.provider.name, kit extensions moved from gen_ai.* to mttl.*, metric names with the suffixes added by the exporter, Tempo presented as an API, exercise aligned with the code (two customers, three features), official opentelemetry-instrumentation-openai-v2 package specified; no more zero cost emitted for the local model, since the euro cost is computed only with a supplied price list.