Skip to content

TechnicalPractitioner

Lab 2: an end-to-end observed RAG pipeline

For: engineers and SREsPrerequisites: Have completed lab 1 of the course.

Duration: 1 h 30. Prerequisites: lab 1 completed, Qdrant reachable on localhost:6333.

StepRoleSpanAttribute
Embeddingvectorize the questionembeddings nomic-embed-textmttl.rag.stage = embedding
Retrievalsearch in Qdrant (top_k)rag retrievalmttl.rag.stage = retrieval
Assemblyconcatenate context and questionrag prompt assemblymttl.rag.stage = prompt_assembly
Generationcall the LLMchat mistral:7bmttl.rag.stage = generation
Evaluationfaithfulness, hallucinationrag evaluationmttl.rag.stage = evaluation

The two model calls follow the GenAI convention naming (operation then model) and feed gen_ai.client.operation.duration. Each step also feeds the mttl.rag.stage.duration histogram, broken down by the step attribute: this is what the RAG dashboard reads.

Fenêtre de terminal
cd app && python lab2_rag_pipeline.py

On first launch, ensure_collection() creates the Qdrant collection and indexes seven documents. Later launches do not reindex, except with the --reindex option: lab 3 relies on this behavior.

In Tempo, open a rag query trace:

rag query
├── embeddings nomic-embed-text ~150 ms
├── rag retrieval ~50 ms
├── rag prompt assembly ~5 ms
├── chat mistral:7b ~1500 ms
└── rag evaluation ~10 ms

These durations are an example: on a workstation without a GPU, generation can take several seconds. The same spans placed on a timeline make the proportion visible:

gantt
  title rag query trace, about 1.7 s
  dateFormat x
  axisFormat %-S.%L s
  section Root
  rag query          :q, 0, 1715ms
  section Steps
  embedding          :e1, 0, 150ms
  retrieval          :e2, after e1, 50ms
  prompt assembly    :e3, after e2, 5ms
  generation         :e4, after e3, 1500ms
  evaluation         :e5, after e4, 10ms

In this example, generation accounts for 1,500 ms out of 1,715, about 87% of the total time. This is expected in this lab: generating 200 tokens costs far more than a search over seven documents. Yet that is not where quality problems hide: retrieval decides what the model has in front of it.

Step 3: three measured optimizations (25 min)

Section titled “Step 3: three measured optimizations (25 min)”
  1. --top-k 1 versus --top-k 5: impact on retrieval latency and on faithfulness.
  2. Embedding cache: add an in-memory dictionary in rag_query yourself, then measure the gain.
  3. --stream: generation switches to streaming and the time to first chunk shows up in the standard gen_ai.client.operation.time_to_first_chunk metric.

Write down each result: the first optimization shows that a latency gain can cost a lot in quality.

PanelWhat you should see
p95 duration per stepgeneration dominates, retrieval stays fast
Throughput per stepconstant across steps, with no drop
Average faithfulnessabove 0.6, the threshold shared by the quality alert and the lab 4 SLO
Hallucinations detected over 1 hideally zero; investigate from the first one

The hallucination counter (mttl.hallucinations) relies on the kit heuristic: a response containing more than two numbers absent from the context. It is a coarse signal, to be complemented by a judge.

Increase the load (python lab2_rag_pipeline.py --repeat 4) and find the step that drops first.

The trace-to-metric link is preconfigured in the Tempo source (grafana/provisioning/datasources/datasources.yaml, tracesToMetrics block): from a slow span, open the p95 latency or the per-step duration of the same service, with the filter already applied. Add a query to it, for example the average faithfulness of the service.

NotionWhat the lab shows
Hierarchyone RAG root span, one child span per step
Step attributeallows breakdown by step without double instrumentation
Online evaluationthe lab’s faithfulness is a ratio of words shared by context and response, to be complemented by a judge
Sizeunder 1 KB per attribute; the identifier of chunks, not their content
  • add a reranking step (bge-reranker) and measure the faithfulness gain;
  • export question and answer pairs to Phoenix for batch evaluation.

Revised on 2 October 2026: box on the kit status, Qdrant search migrated to query_points, mttl.hallucinations counter actually emitted, faithfulness threshold aligned on 0.6, share of generation recalculated (87%).