Technical
8. Versioning, replay and experimentation
For: engineers and SREs · architectsPrerequisites: Have read chapters 3 to 5 of the guide.
Once the platform reaches Level 3, three advanced capabilities become tractable: versioning, replay and experimentation. Each multiplies the value of the observability data already collected.
8.1 Versioning every artifact
Section titled “8.1 Versioning every artifact”Every artifact that influences output should be versioned and that version captured as a span attribute. Without this, you cannot answer the question: did the change of prompt version cause the regression?
Artifacts to version
Section titled “Artifacts to version”- Prompt templates (system and user) by content hash and a semantic version.
- Model identity, including the actually-served version returned by the provider, not just the requested name.
- Retriever index by snapshot or build identifier.
- Reranker model and weights.
- Tool schemas exposed to the agent.
- Evaluator prompts, when LLM-as-judge is used.
- The entire feature configuration as a single config hash.
Captured as attributes
Section titled “Captured as attributes”acme.prompt.version = "rag-answer-v17"acme.prompt.hash = "sha256:..."acme.retriever.index = "kb-prod-2026-05-12"acme.retriever.version = "hybrid-v3"acme.reranker.model = "bge-rerank-v2"acme.config.hash = "sha256:..."gen_ai.response.model = "provider-model-2026-03-15"If a regression appears, the first query is: group by acme.prompt.version and compute mean faithfulness. The answer is one query away or it is unanswerable.
8.2 Replay
Section titled “8.2 Replay”Replay means running an old trace through a different version of the system to compare outputs. It is the foundation of safe rollouts.
What replay enables
Section titled “What replay enables”- Test a new prompt version against last week’s traffic before deploying.
- Compare a new model against the current production model on identical inputs.
- Re-evaluate historical traces with an updated evaluator to see the trend with the new rubric.
- Reproduce a customer complaint to validate that a fix works.
What replay requires
Section titled “What replay requires”- Full prompt and tool argument capture at the time of the original trace.
- A way to bypass non-deterministic external calls (frozen MCP responses, frozen retrieval results).
- A separate execution path that does not contaminate production traces.
Common pitfalls
Section titled “Common pitfalls”- Forgetting that retrieval results have changed since the original trace. Either replay against a frozen index or accept that retrieval differences are part of the comparison.
- Replaying with the original temperature: sampling noise dominates small score differences. Use temperature zero or self-consistency. Temperature zero reduces variability without removing it: it is not deterministic with every provider, and some reasoning models do not accept this parameter. In that case, self-consistency or several runs per trace are the only option.
- Letting replay traces enter production dashboards. Tag them explicitly with a replay marker and filter them out of operational alerts.
8.3 Experimentation: shadow mode and A/B testing
Section titled “8.3 Experimentation: shadow mode and A/B testing”Three deployment patterns for testing changes against live traffic.
Shadow mode
Section titled “Shadow mode”- The new system runs in parallel with production but its output is not returned to the user.
- Both outputs are logged with a shared trace ID.
- Evaluators score both. Differences are surfaced in a comparison dashboard.
- Used for high-confidence rollouts where the new system is expected to match or exceed production.
A/B testing
Section titled “A/B testing”- A fraction of traffic is routed to the new system. The user sees the experimental output.
- Trace attribute
acme.experiment.variantcaptures which arm the request was assigned to. - Metrics and evaluator scores are aggregated by variant.
- Statistical significance is required before promoting a variant. A 1 percent change in faithfulness on 1,000 traces is not significant.
Canary
Section titled “Canary”- Same as A/B but with a small initial fraction (1 to 5 percent), expanded gradually.
- Automatic rollback on quality regression detected by online evaluators.
- Most appropriate for high-traffic features where statistical signal arrives quickly.
8.4 Regression gating in CI
Section titled “8.4 Regression gating in CI”The final maturity move: block deployments on offline evaluation results.
- Maintain a regression dataset of inputs covering the failure modes seen in production.
- Run the candidate prompt or model against the regression dataset on every CI build.
- Score each output with the relevant evaluators.
- Fail the build if any evaluator score regresses below the previous best.
- Maintain the regression dataset as a living artifact: every triaged production failure becomes a new entry.