Technical
Lab 1: instrument an LLM application
For: engineers and SREsPrerequisites: Have read modules 1 and 2 of the course, know how to use Docker and read Python.
Duration: 1 h 45. Prerequisites: modules 1 and 2, Docker, Python 3.10 or later, 16 GB of memory recommended.
Prepare the environment
Section titled “Prepare the environment”cd labs/llm-observabilitydocker compose up -d
# Local models (5 to 10 minutes on first launch)docker exec -it mttl-ollama ollama pull mistral:7bdocker exec -it mttl-ollama ollama pull nomic-embed-text
cd app && pip install -r requirements.txt| Service | Address | Nature |
|---|---|---|
| Grafana | http://localhost:3000 | interface (admin / admin) |
| VictoriaMetrics | http://localhost:8428/vmui | query interface |
| Tempo | http://localhost:3200 | HTTP API, no interface: traces are explored from Grafana |
| Phoenix | http://localhost:6006 | interface |
Image versions in the kit: Collector contrib 0.108.0, VictoriaMetrics v1.110.0, Tempo 2.6.0, Grafana 11.2.0, Phoenix version-5.9.1, Qdrant v1.12.0, Ollama 0.3.12. They are pinned in docker-compose.yml; Python dependencies are pinned in app/requirements.txt.
Step 1: read the code (15 min)
Section titled “Step 1: read the code (15 min)”Open app/lab1_llm_app.py and locate:
init_telemetry(), which creates the trace and metric providers and the OTLP exporters;- the manual span around the model call;
- the
gen_ai.*attributes set on that span; - the instruments created by
create_genai_metrics(): tokens and duration (standard), cost and errors (mttl.*extensions). Lab 1 emits no cost: the euro cost is computed only in lab 4, from a price list you supply.
The core of the instrumentation fits in a few lines:
labels = { GEN_AI_PROVIDER_NAME: Provider.OLLAMA, # gen_ai.provider.name GEN_AI_OPERATION_NAME: Operation.CHAT, GEN_AI_REQUEST_MODEL: MODEL, MTTL_TENANT_ID: tenant, # in-house extension: mttl.tenant.id MTTL_FEATURE: feature, # in-house extension: mttl.feature}with tracer.start_as_current_span(f"chat {MODEL}", attributes={**labels, ...}) as span: try: response = get_ollama().chat(model=MODEL, messages=[...], options={...}) input_tokens = response.get("prompt_eval_count", 0) output_tokens = response.get("eval_count", 0)
span.set_attribute(GEN_AI_USAGE_INPUT_TOKENS, input_tokens) span.set_attribute(GEN_AI_USAGE_OUTPUT_TOKENS, output_tokens) metrics["duration"].record(time.time() - start, labels) metrics["tokens"].record(input_tokens, {**labels, GEN_AI_TOKEN_TYPE: "input"}) metrics["tokens"].record(output_tokens, {**labels, GEN_AI_TOKEN_TYPE: "output"}) except Exception as e: err = {**labels, ERROR_TYPE: type(e).__name__} span.record_exception(e) span.set_status(trace.Status(trace.StatusCode.ERROR, str(e))) metrics["duration"].record(time.time() - start, err) metrics["errors"].add(1, err) raiseThe constants (GEN_AI_PROVIDER_NAME, MTTL_TENANT_ID…) are defined in app/shared/conventions.py: a single file to change when the convention evolves. Standard names stay in gen_ai.*, the kit extensions go in mttl.*.
In VictoriaMetrics, the Collector exporter adds unit and type suffixes: the duration becomes gen_ai_client_operation_duration_seconds_bucket, the error counter mttl_client_errors_total, and the mttl.tenant.id attribute becomes the mttl_tenant_id label. These are the names that dashboards and alerts query.
Step 2: first run (20 min)
Section titled “Step 2: first run (20 min)”python lab1_llm_app.pyIn Grafana:
- Explore, Tempo source: look for the
lab1-llm-appservice and open a trace. - Check that the provider, model and usage attributes are present.
- LLM Overview dashboard: tokens and latency should appear.
Step 3: guided changes (25 min)
Section titled “Step 3: guided changes (25 min)”- Add the
gen_ai.request.top_pattribute to the span. - Vary the temperature from 0.0 to 1.0 and watch the spread of latency.
- Add a span event for the prompt:
span.add_event("prompt", {...}). This capture is acceptable in this local lab; in production it stays off by default, is enabled only by an explicit choice (opt-in) and comes with masking of sensitive data (see module 2 and §2 of the article Observing an LLM system).
Step 4: several customers and features (25 min)
Section titled “Step 4: several customers and features (25 min)”The script sends five requests spread across two customers (acme-corp, client-test) and three features (support, search, summary). Exercise to do yourself: add to the PROMPTS list a third customer (globex-energy) and a fourth feature (translation), run again, then check on the LLM Overview dashboard that metrics break down by customer and by feature (panels “Jetons par client” and “Jetons par fonctionnalité”, that is tokens per customer and per feature; model and tenant variables at the top of the dashboard).
Step 5: deliberate failure (20 min)
Section titled “Step 5: deliberate failure (20 min)”docker stop mttl-ollamapython lab1_llm_app.pyWatch the spans in error in Tempo (with the error.type attribute), the increment of the mttl_client_errors_total counter and the absence of consumed tokens. Then run docker start mttl-ollama again.
Lab recap
Section titled “Lab recap”| Notion | What the lab shows |
|---|---|
| Traces | in this lab, an LLM call produces one span; child spans per step arrive in lab 2 |
| Metrics | always labeled by model, customer and feature |
| Errors | span.set_status(StatusCode.ERROR) plus a dedicated error counter |
| Cost | tracked here in tokens; in euros only in lab 4, with your versioned and dated price list |
Pitfalls to avoid
Section titled “Pitfalls to avoid”- putting full prompts in attributes;
- leaving a free-form customer identifier if you have thousands of them: cardinality explodes;
- forgetting that auto-instrumentation exists for standard cases: the official OpenTelemetry project package for the OpenAI client is
opentelemetry-instrumentation-openai-v2;opentelemetry-instrumentation-openai, without suffix, is the one from Traceloop’s OpenLLMetry project.
Revised on 2 October 2026: box on the kit status, gen_ai.system replaced by gen_ai.provider.name, kit extensions moved from gen_ai.* to mttl.*, metric names with the suffixes added by the exporter, Tempo presented as an API, exercise aligned with the code (two customers, three features), official opentelemetry-instrumentation-openai-v2 package specified; no more zero cost emitted for the local model, since the euro cost is computed only with a supplied price list.