Technical
9. Anti-patterns and common pitfalls
For: engineers and SREs · architects · team managersPrerequisites: Have read chapters 1 to 4 of the guide.
Twelve common pitfalls, described here as typical scenarios, that produce data nobody uses, dashboards nobody reads or platforms nobody trusts. Avoiding them is cheaper than fixing them.
9.1 Instrumenting before defining questions
Section titled “9.1 Instrumenting before defining questions”Symptom: a comprehensive trace tree with hundreds of attributes per span, but nobody can answer “what is the faithfulness score by tenant?”. The relevant attributes were never indexed.
Fix: Step 1 of Chapter 4. Write the questions first.
9.2 Treating logs as observability
Section titled “9.2 Treating logs as observability”Symptom: large structured log volume in Loki or Splunk, no spans, no evaluations. Cost is high, value extraction is manual.
Fix: move to Level 2. Logs are a fallback channel, not the primary signal.
9.3 Storing prompts and responses without redaction
Section titled “9.3 Storing prompts and responses without redaction”Symptom: a production trace contains an API key a user pasted into the chat. The trace is now retained in a queryable backend with broad access.
Fix: redaction in the collector, before any exporter. Never trust the application to scrub.
9.4 LLM-as-judge on every trace
Section titled “9.4 LLM-as-judge on every trace”Symptom: the evaluation token cost exceeds the production LLM cost. The eval data is so noisy that nobody acts on it.
Fix: tier evaluators. For example: cheap rule-based scorers on 100 percent; medium LLM judges on 10 percent; expensive judges on 1 percent plus flagged traces. Use self-consistency only where it matters.
9.5 Dashboards without alerts
Section titled “9.5 Dashboards without alerts”Symptom: beautiful Grafana boards, nobody looks at them, regressions are discovered by users.
Fix: every dashboard panel should have an alert if it represents a real SLO. Dashboards are diagnostic, alerts are operational.
9.6 High-cardinality attributes on metrics
Section titled “9.6 High-cardinality attributes on metrics”Symptom: metrics backend memory pressure, slow queries, sudden cost explosion when a new tenant onboards.
Fix: attribute hygiene. High-cardinality keys belong on traces. Metrics dimensions must be bounded.
9.7 No versioning of prompts and configs
Section titled “9.7 No versioning of prompts and configs”Symptom: quality regression detected, but nobody can identify which change caused it. Multiple things shipped that week.
Fix: every artifact versioned, version captured on every span, deployments correlated with version changes.
9.8 Treating drift as a problem to be fixed
Section titled “9.8 Treating drift as a problem to be fixed”Symptom: a drift alarm fires, the team rebuilds the index or swaps the model, drift returns the next week.
Fix: drift is a signal, not a defect. It indicates the world has changed. The right response is investigation and often the right action is no action. Only some drift correlates with quality loss.
9.9 Buying a tool before defining the stack
Section titled “9.9 Buying a tool before defining the stack”Symptom: the team subscribes to a commercial LLM observability product, integrates it, then realizes it cannot deploy in the required jurisdiction or cannot handle the cardinality.
Fix: the choice between open source and commercial is a strategic decision made against sovereignty, cost and integration constraints. Make that decision before evaluating products.
9.10 Evaluating without calibrating
Section titled “9.10 Evaluating without calibrating”Symptom (illustrative scenario): faithfulness scores look stable around 0.85, but spot checks reveal the judge often contradicts human judgment. Nobody trusts the metric.
Fix: build a calibration set with human ground truth. Measure agreement with a chance-corrected statistic (Cohen’s kappa). Iterate on the judge prompt until it reaches the threshold set in advance. Recalibrate when the judge model is upgraded.
9.11 Skipping the closed loop
Section titled “9.11 Skipping the closed loop”Symptom: evaluation runs, scores are stored, nothing changes. The platform produces data but no improvement.
Fix: formalize the weekly review. Failed traces become regression set entries. Without this ritual, observability collapses to expensive read-only storage.
9.12 Conflating MLOps and AI observability
Section titled “9.12 Conflating MLOps and AI observability”Symptom: an MLOps team owns the AI observability stack and tries to instrument the LLM inference layer the same way they instrument model training. The semantics do not match.
Fix: AI observability is a production observability concern, owned by the platform team in collaboration with ML and security. MLOps owns model lifecycle, not request lifecycle.
Next: Appendix.