Technical
Trace explorer
For: engineers and SREsPrerequisites: Basic notions of distributed traces.
A dashboard says a request is slow or failing; a trace says where and why. It splits a request into nested spans, each with its duration, status and attributes. The three traces below are hand-built following the OpenTelemetry semantic conventions as of 2 October 2026 (GenAI conventions from the semantic-conventions-genai repository, still at Development status, and MCP conventions for protocol version 2025-06-18). MCP version 2026-07-28 removes sessions; these traces keep 2025-06-18, the version used in the convention’s example. One deliberate exception: the evaluate spans are an application choice outside the conventions; the scores themselves are carried by the gen_ai.evaluation.result event the conventions define. Two traces involve a RAG agent, the third a classic checkout, to show that the reading method is the same with or without AI.
Red outline: span in error. Amber outline: evaluation below its threshold. "key" marker in the details: attribute worth reading.
What this trace tells us
Hand-built example traces in a simplified format (times in milliseconds relative to the start of the trace). Attributes follow the OpenTelemetry semantic conventions; the generative AI, MCP and evaluation conventions are still in development and may change. Attributes prefixed with app. are application-specific. The "evaluate" spans are an application choice outside the conventions: the conventions define no evaluation span, they carry scores in the gen_ai.evaluation.result event, to be attached preferably to the span of the evaluated operation, otherwise through gen_ai.response.id; here it is placed in the application span.
Try this
Section titled “Try this”- Read the baseline: on the nominal trace, find the two
chatcalls to the model. They account for most of the duration; compare theirgen_ai.usage.input_tokenswith the embeddings call. - Find the silent failure: switch to the “silent failure” trace. No error, HTTP 200, a faster request. Only the faithfulness evaluation is flagged “low score”: open it, its
gen_ai.evaluation.resultevent carries the 0.34 score. Then clickquery kb_support_frandinvoke_agent:app.retrieval.top_kis 1 and theapp.config.loadedevent names the faulty configuration version. - Compare tokens: the “LLM tokens” tile drops from about 9,000 to 2,600. On a cost dashboard, the incident looks like a successful optimization.
- Step outside AI: on the e-commerce trace, use the keyboard (up and down arrows) to reach the error span. The 3 s timeout, then the retry (
http.request.resend_count), explain the 4.3 s the customer waited.
What to remember
Section titled “What to remember”- Time reads top to bottom, the cause sits in the attributes. The waterfall shows where time goes; attributes and events say why. A trace without business attributes (configuration version,
top_k, idempotency key) cannot support a conclusion. - An OK status does not mean a good answer. For a generative system, evaluation must be attached to the trace: otherwise the silent failure stays invisible. The conventions provide the
gen_ai.evaluation.resultevent for this, attached to the evaluated operation or, failing that, to its answer throughgen_ai.response.id; here, each event sits in an application evaluation span, which makes the evaluation time visible in the waterfall. - Beware of homonyms.
gen_ai.request.top_krefers to model sampling, not to the number of retrieved passages: that is why the trace uses an application attribute,app.retrieval.top_k.
Go further: GenAI conventions in the LLM observability training, the three drifts of a RAG system in the dedicated module, quality SLOs in the GenAI method, and the stack that carries these traces in the clickable architecture.
Revised on 2 October 2026: evaluation scores move to the conventions’ gen_ai.evaluation.result event, evaluation spans are flagged as an application choice, and each trace states its build date and the convention versions it follows.