Technical
Lab 2: an end-to-end observed RAG pipeline
For: engineers and SREsPrerequisites: Have completed lab 1 of the course.
Duration: 1 h 30. Prerequisites: lab 1 completed, Qdrant reachable on localhost:6333.
The five steps and their spans
Section titled “The five steps and their spans”| Step | Role | Span | Attribute |
|---|---|---|---|
| Embedding | vectorize the question | embeddings nomic-embed-text | mttl.rag.stage = embedding |
| Retrieval | search in Qdrant (top_k) | rag retrieval | mttl.rag.stage = retrieval |
| Assembly | concatenate context and question | rag prompt assembly | mttl.rag.stage = prompt_assembly |
| Generation | call the LLM | chat mistral:7b | mttl.rag.stage = generation |
| Evaluation | faithfulness, hallucination | rag evaluation | mttl.rag.stage = evaluation |
The two model calls follow the GenAI convention naming (operation then model) and feed gen_ai.client.operation.duration. Each step also feeds the mttl.rag.stage.duration histogram, broken down by the step attribute: this is what the RAG dashboard reads.
Step 1: indexing (10 min)
Section titled “Step 1: indexing (10 min)”cd app && python lab2_rag_pipeline.pyOn first launch, ensure_collection() creates the Qdrant collection and indexes seven documents. Later launches do not reindex, except with the --reindex option: lab 3 relies on this behavior.
Step 2: read the span tree (20 min)
Section titled “Step 2: read the span tree (20 min)”In Tempo, open a rag query trace:
rag query├── embeddings nomic-embed-text ~150 ms├── rag retrieval ~50 ms├── rag prompt assembly ~5 ms├── chat mistral:7b ~1500 ms└── rag evaluation ~10 msThese durations are an example: on a workstation without a GPU, generation can take several seconds. The same spans placed on a timeline make the proportion visible:
gantt title rag query trace, about 1.7 s dateFormat x axisFormat %-S.%L s section Root rag query :q, 0, 1715ms section Steps embedding :e1, 0, 150ms retrieval :e2, after e1, 50ms prompt assembly :e3, after e2, 5ms generation :e4, after e3, 1500ms evaluation :e5, after e4, 10ms
In this example, generation accounts for 1,500 ms out of 1,715, about 87% of the total time. This is expected in this lab: generating 200 tokens costs far more than a search over seven documents. Yet that is not where quality problems hide: retrieval decides what the model has in front of it.
Step 3: three measured optimizations (25 min)
Section titled “Step 3: three measured optimizations (25 min)”--top-k 1versus--top-k 5: impact on retrieval latency and on faithfulness.- Embedding cache: add an in-memory dictionary in
rag_queryyourself, then measure the gain. --stream: generation switches to streaming and the time to first chunk shows up in the standardgen_ai.client.operation.time_to_first_chunkmetric.
Write down each result: the first optimization shows that a latency gain can cost a lot in quality.
Step 4: the RAG dashboard (20 min)
Section titled “Step 4: the RAG dashboard (20 min)”| Panel | What you should see |
|---|---|
| p95 duration per step | generation dominates, retrieval stays fast |
| Throughput per step | constant across steps, with no drop |
| Average faithfulness | above 0.6, the threshold shared by the quality alert and the lab 4 SLO |
| Hallucinations detected over 1 h | ideally zero; investigate from the first one |
The hallucination counter (mttl.hallucinations) relies on the kit heuristic: a response containing more than two numbers absent from the context. It is a coarse signal, to be complemented by a judge.
Increase the load (python lab2_rag_pipeline.py --repeat 4) and find the step that drops first.
Step 5: from trace to metric (15 min)
Section titled “Step 5: from trace to metric (15 min)”The trace-to-metric link is preconfigured in the Tempo source (grafana/provisioning/datasources/datasources.yaml, tracesToMetrics block): from a slow span, open the p95 latency or the per-step duration of the same service, with the filter already applied. Add a query to it, for example the average faithfulness of the service.
Lab recap
Section titled “Lab recap”| Notion | What the lab shows |
|---|---|
| Hierarchy | one RAG root span, one child span per step |
| Step attribute | allows breakdown by step without double instrumentation |
| Online evaluation | the lab’s faithfulness is a ratio of words shared by context and response, to be complemented by a judge |
| Size | under 1 KB per attribute; the identifier of chunks, not their content |
Going further
Section titled “Going further”- add a reranking step (
bge-reranker) and measure the faithfulness gain; - export question and answer pairs to Phoenix for batch evaluation.
Revised on 2 October 2026: box on the kit status, Qdrant search migrated to query_points, mttl.hallucinations counter actually emitted, faithfulness threshold aligned on 0.6, share of generation recalculated (87%).