Technical
Module 2: OpenTelemetry GenAI conventions
For: engineers and SREsPrerequisites: Have read module 1 of the course.
OpenTelemetry in four notions
Section titled “OpenTelemetry in four notions”| Notion | Role |
|---|---|
| Traces | break an operation down into hierarchical spans, each carrying attributes and events |
| Metrics | time series (counter, gauge, histogram), exported over OTLP, compatible with Prometheus remote write |
| Logs | structured logs, correlated with traces through trace_id and span_id |
| Semantic conventions | standardized vocabulary: http.*, db.*, messaging.*, gen_ai.*; this is what makes instrumentation portable |
The essential GenAI attributes
Section titled “The essential GenAI attributes”| Attribute | Description | Example |
|---|---|---|
gen_ai.provider.name | model provider | well-known values: openai, anthropic, mistral_ai, azure.ai.openai, etc.; a custom value is allowed when none applies (the kit uses ollama) |
gen_ai.operation.name | type of operation | chat, embeddings, text_completion |
gen_ai.request.model | requested model | mistral:7b |
gen_ai.request.temperature | requested temperature | 0.7 |
gen_ai.usage.input_tokens | prompt tokens | 245 |
gen_ai.usage.output_tokens | generated tokens | 128 |
gen_ai.response.finish_reasons | finish reason | stop, length, tool_call |
Anatomy of a GenAI span
Section titled “Anatomy of a GenAI span”Here is what a production span should contain:
span chat mistral:7b trace_id 5b8efff798038103d269b633813fc60c duration 1.847s attributes gen_ai.provider.name = ollama gen_ai.request.model = mistral:7b gen_ai.operation.name = chat gen_ai.usage.input_tokens = 245 gen_ai.usage.output_tokens = 128 app.tenant.id = acme-corp app.feature = support-chat app.quality.score = 0.78 events [t+0.000s] request.start [t+0.182s] first_token_receivedThe value ollama is not in the list of well-known values of gen_ai.provider.name: the convention requires a well-known value when one applies and otherwise allows a custom value, which is the case for a model served locally by Ollama.
The last three attributes are not standard: they are in-house extensions. Put them in your own namespace (app.*, or the name of your organization) rather than in gen_ai.*, so they do not collide with a future version of the convention. The lab kit puts its own under mttl.* (mttl.tenant.id, mttl.feature, mttl.quality.score).
The first_token_received event gives the time to first token, the indicator of perceived latency in streaming; the convention also measures it as a metric (below).
Metrics to emit by default
Section titled “Metrics to emit by default”| Metric | Type | Content |
|---|---|---|
gen_ai.client.token.usage | histogram (standard) | tokens consumed per operation, by input or output type; being replaced, see the box above |
gen_ai.client.operation.duration | histogram (standard) | duration of an LLM operation, including failures with the error.type attribute |
gen_ai.client.operation.time_to_first_chunk | histogram (standard) | time until the first chunk is received in streaming |
estimated cost (mttl.client.cost in the kit) | counter (extension) | euros, computed from tokens and a dated price list you supply; without a price list, the metric is not emitted |
errors (mttl.client.errors in the kit) | counter (extension) | timeouts, refusals, provider errors |
quality score (mttl.quality.score in the kit) | histogram (extension) | score from 0 to 1 coming from evaluation |
For duration histograms, set bucket boundaries suited to seconds: the convention recommends some (0.01; 0.02; 0.04 … 81.92 s). Without them, the Python SDK applies default boundaries designed for milliseconds, and the computed quantiles no longer make sense.
Always label by model, customer and feature. These are the three breakdown axes that finance, product and operations will ask for.
Traces, metrics or events?
Section titled “Traces, metrics or events?”- The trace is for diagnosing one request: which document, which prompt, which tool call.
- The metric is for aggregation and alerting: cost per hour, p95, average score.
- The event (or log) carries bulky or sensitive content, such as the prompt and the response, when you have explicitly enabled its capture: shorter retention and masking.
In summary
Section titled “In summary”The GenAI conventions give providers and frameworks a common vocabulary. They are still moving, so translation between versions happens in the Collector and never in dashboards. Your own attributes go in your namespace. Finally, every LLM metric is broken down at least by model, customer and feature.
Revised on 2 October 2026: move of the GenAI conventions to the semantic-conventions-genai repository (version 1.42) and ongoing rework of token metrics noted, standard time to first chunk metric, known values of gen_ai.provider.name, names of the kit extensions, kit cost emitted only with a price list supplied by the user.
Revised on 4 October 2026: history of the conventions replaced by a pointer to §2 of the article Observing an LLM system; content capture rule aligned with the GenAI conventions (off by default, opt-in, outside indexed attributes in production), checked against the semantic-conventions-genai repository; ollama presented as an allowed custom value of gen_ai.provider.name.