Technical
7. Operating the platform
For: engineers and SREs · architects · team managersPrerequisites: Have read chapters 4 and 5 of the guide.
Implementation is only the start. Running an AI observability platform sustainably requires deliberate choices around what data to keep, how much it costs and how to keep it safe.
7.1 Sampling strategy
Section titled “7.1 Sampling strategy”Full-fidelity tracing of every LLM request is expensive in storage and processing. Sampling becomes necessary above a few hundred requests per second.

The figure simplifies: only scores from synchronous evaluations are available when the sampling decision is made (see below).
Figure 10. Head sampling decides at trace start, randomly. Tail sampling decides at trace end, by policy. Tail sampling preserves errors, slow traces and traces flagged by a synchronous evaluation deterministically.
Recommended policy
Section titled “Recommended policy”- Always keep error traces.
- Always keep abnormally slow traces.
- Always keep traces where a synchronous evaluation (toxicity filter, refusal classifier, run during the request) is below threshold.
- Probabilistic 10 percent on the remainder, for baseline visibility.
Implementation: OpenTelemetry Collector tail_sampling processor. It can only decide on what the trace contains when its wait time expires (decision_wait, 30 seconds by default): errors, latencies, attributes set during the request. Asynchronous evaluator scores and user feedback arrive much later; they cannot drive this decision. To keep traces that are scored after the fact, keep all raw traces on a short retention (a few days), then promote to long-term storage those that an evaluator scored poorly or that received negative feedback.
The processor’s memory cost depends on the number of traces held pending (num_traces) and on the decision_wait duration, multiplied by the average trace size: size these two parameters from your throughput.
7.2 Cost modeling
Section titled “7.2 Cost modeling”This section is the site’s reference on the cost of observability itself. The cost of non-observability (the economic exposure of a poorly observed system, with its calculator) is covered in the GenAI method, Part V.
At scale, the cost of AI observability can rival the cost of the LLM calls themselves. A simple model:
total_obs_cost = storage_cost + processing_cost + judge_cost
stored_GB = RPS * avg_payload_KB * 86400 * retention_days / 1024 / 1024storage_cost = stored_GB * cost_per_GB_month (per month)processing_cost = collector_cost + ingestion_costjudge_cost = RPS * 86400 * sample_rate * tokens_per_judge * cost_per_token (per day)Example at 100 RPS, 4 KB average payload, 30-day hot retention: stored_GB = 100 * 4 * 86400 * 30 / 1024 / 1024 = 988.8 GB, about 1 TB kept in hot storage storage_cost = 988.8 * cost_per_GB_month (per month)
Example judge at 100 RPS, 10 percent sample, 1500 tokens per judge: judgments = 100 * 86400 * 0.10 = 864,000 judgments per day judge tokens = 864,000 * 1500 = 1,296 million tokens per day for one evaluator judge_cost = 1,296,000,000 * cost_per_token (per day)Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list for cost_per_GB_month and cost_per_token. Compare both items over the same period: storage is billed per month, the judge consumes every day (about 39 billion tokens over 30 days in this example). A single per-token rate is also a simplification, since output tokens are usually billed higher than input tokens.
Cost drivers, in order of importance
Section titled “Cost drivers, in order of importance”- Payload size: prompts, completions and retrieved contexts can be tens of kilobytes each. Storage volume grows linearly with traffic.
- LLM-as-judge evaluations: consume tokens at evaluation time. A trivial evaluator running on every trace can double total token usage.
- High-cardinality attributes:
user_id,request_id,session_idinflate metrics storage if attached to metrics rather than traces.
Mitigation strategies
Section titled “Mitigation strategies”- Compress and tier old trace storage (hot for seven days, warm for thirty, cold archived beyond).
- Use small models for online evaluation, large models only for periodic audit batches.
- Restrict high-cardinality attributes to traces, keep metrics low-cardinality through aggregation.
- Sample LLM-as-judge runs rather than evaluating every trace.
- Tier evaluators by cost and run expensive ones on a small sample plus flagged traces.
7.3 Cardinality management
Section titled “7.3 Cardinality management”Metrics cardinality is the single largest scaling risk in an AI observability platform. AI workloads have natural high-cardinality dimensions: prompt versions, model versions, retriever index versions, tenant IDs, user IDs.
Rules of thumb
Section titled “Rules of thumb”- Attributes with thousands of distinct values per day belong on traces, not on metrics.
- Metrics dimensions should be bounded:
model_name(tens),feature(tens),tenant(hundreds at most). - Per-user metrics are an anti-pattern: query traces by
user_idinstead. - Changing backends does not solve the problem. Some backends (VictoriaMetrics, Grafana Mimir, Thanos, or managed offerings) claim better handling of high cardinality than Prometheus alone; these are vendor statements, to be measured on your workload. The real answer remains to fix the metrics design.
7.4 PII and data sovereignty
Section titled “7.4 PII and data sovereignty”Prompts and completions frequently contain personal data. Three controls are recommended, all the more so in regulated contexts.
Redaction before storage
Section titled “Redaction before storage”- Regex-based for known patterns: email, phone, IBAN, credit card, social security numbers. The Collector’s
redactionprocessor does this (blocked values masked or hashed, non-allowed keys deleted). - NER-based for named entities: person names, organizations, locations. The Collector does not do this natively: it takes an external service (for example a personal data detection engine called by the application or by an intermediate component) or a custom processor.
- LLM-based for context-sensitive scrubbing where regex and NER miss patterns. Same constraint: an external component, with its cost and latency.
- Run regex redaction in the collector, before export, and place NER or LLM processing upstream of storage; never in the storage backend, once the data has been written.
Access control
Section titled “Access control”- Role-based access control on trace queries is the minimum.
- Field-level access where the backend supports it: attribute keys are readable by all, but prompt and completion bodies require elevated access.
- Audit log on all trace query access, retained per regulatory requirement.
Retention policy
Section titled “Retention policy”- Aligned with the storage limitation principle (GDPR Article 5(1)(e)) and with sectoral regulations.
- Distinct retention for trace metadata (longer) and trace bodies (shorter).
- Automated purge processes verified by audit.
7.5 Security
Section titled “7.5 Security”LLM observability introduces a new attack surface that classical APM does not face.
Threats
Section titled “Threats”- Stored prompts may contain credentials accidentally pasted by users (API keys, tokens, secrets).
- Stored completions may contain output that leaked training data through memorization.
- Evaluation pipelines calling external LLM judges may exfiltrate sensitive data to third-party providers.
- Trace query interfaces expose the full content of past interactions to anyone with access.
- Prompt-injection attacks may target the judge model itself, manipulating evaluator scores.
Mitigations
Section titled “Mitigations”- Pre-storage scrubbing for credential patterns alongside PII.
- Default-deny on outbound LLM judge calls, with an explicit allowlist of permitted providers.
- Network egress controls from the collector to judge endpoints.
- Audit log on all trace query access, with periodic review.
- Encryption at rest and in transit for trace storage.
- Judge prompts hardened against injection: sanitize the input fields, use delimiters, never let the rated content occupy the system prompt.
7.6 Organizational ownership
Section titled “7.6 Organizational ownership”- Platform owns the stack (collectors, backends, dashboards, alert routing), ML owns the evaluators, quality thresholds and the regression set, security owns redaction rules, access policies and audit review.
- All three jointly run the weekly review of failed and low-score traces: in my view, the most useful recurring forum and the only one that closes the loop.
- The full RACI matrix (including product, compliance and quality on-call) is in the GenAI method, Part VII, which is the reference; the AI Act and regulatory evidence are in Part VI.
Next: 8. Versioning, replay and experimentation.
Revised on 2 October 2026: prices removed, tail sampling limits (asynchronous evaluations and user feedback), memory sizing, NER and LLM redaction outside the Collector, GDPR Article 5, NIS2 and DORA location requirements corrected.
Revised on 4 October 2026: section 7.2 designated as the reference on the cost of observability, with a link to the method for the cost of non-observability; section 7.6 summarized in three lines with a link to the method’s RACI.