Skip to content

TechnicalPractitioner

8. Versioning, replay and experimentation

For: engineers and SREs · architectsPrerequisites: Have read chapters 3 to 5 of the guide.

Once the platform reaches Level 3, three advanced capabilities become tractable: versioning, replay and experimentation. Each multiplies the value of the observability data already collected.

Every artifact that influences output should be versioned and that version captured as a span attribute. Without this, you cannot answer the question: did the change of prompt version cause the regression?

  • Prompt templates (system and user) by content hash and a semantic version.
  • Model identity, including the actually-served version returned by the provider, not just the requested name.
  • Retriever index by snapshot or build identifier.
  • Reranker model and weights.
  • Tool schemas exposed to the agent.
  • Evaluator prompts, when LLM-as-judge is used.
  • The entire feature configuration as a single config hash.
acme.prompt.version = "rag-answer-v17"
acme.prompt.hash = "sha256:..."
acme.retriever.index = "kb-prod-2026-05-12"
acme.retriever.version = "hybrid-v3"
acme.reranker.model = "bge-rerank-v2"
acme.config.hash = "sha256:..."
gen_ai.response.model = "provider-model-2026-03-15"

If a regression appears, the first query is: group by acme.prompt.version and compute mean faithfulness. The answer is one query away or it is unanswerable.

Replay means running an old trace through a different version of the system to compare outputs. It is the foundation of safe rollouts.

  • Test a new prompt version against last week’s traffic before deploying.
  • Compare a new model against the current production model on identical inputs.
  • Re-evaluate historical traces with an updated evaluator to see the trend with the new rubric.
  • Reproduce a customer complaint to validate that a fix works.
  • Full prompt and tool argument capture at the time of the original trace.
  • A way to bypass non-deterministic external calls (frozen MCP responses, frozen retrieval results).
  • A separate execution path that does not contaminate production traces.
  • Forgetting that retrieval results have changed since the original trace. Either replay against a frozen index or accept that retrieval differences are part of the comparison.
  • Replaying with the original temperature: sampling noise dominates small score differences. Use temperature zero or self-consistency. Temperature zero reduces variability without removing it: it is not deterministic with every provider, and some reasoning models do not accept this parameter. In that case, self-consistency or several runs per trace are the only option.
  • Letting replay traces enter production dashboards. Tag them explicitly with a replay marker and filter them out of operational alerts.

8.3 Experimentation: shadow mode and A/B testing

Section titled “8.3 Experimentation: shadow mode and A/B testing”

Three deployment patterns for testing changes against live traffic.

  • The new system runs in parallel with production but its output is not returned to the user.
  • Both outputs are logged with a shared trace ID.
  • Evaluators score both. Differences are surfaced in a comparison dashboard.
  • Used for high-confidence rollouts where the new system is expected to match or exceed production.
  • A fraction of traffic is routed to the new system. The user sees the experimental output.
  • Trace attribute acme.experiment.variant captures which arm the request was assigned to.
  • Metrics and evaluator scores are aggregated by variant.
  • Statistical significance is required before promoting a variant. A 1 percent change in faithfulness on 1,000 traces is not significant.
  • Same as A/B but with a small initial fraction (1 to 5 percent), expanded gradually.
  • Automatic rollback on quality regression detected by online evaluators.
  • Most appropriate for high-traffic features where statistical signal arrives quickly.

The final maturity move: block deployments on offline evaluation results.

  1. Maintain a regression dataset of inputs covering the failure modes seen in production.
  2. Run the candidate prompt or model against the regression dataset on every CI build.
  3. Score each output with the relevant evaluators.
  4. Fail the build if any evaluator score regresses below the previous best.
  5. Maintain the regression dataset as a living artifact: every triaged production failure becomes a new entry.

Next: 9. Anti-patterns and common pitfalls.