Skip to content

TechnicalPractitioner

7. Operating the platform

For: engineers and SREs · architects · team managersPrerequisites: Have read chapters 4 and 5 of the guide.

Implementation is only the start. Running an AI observability platform sustainably requires deliberate choices around what data to keep, how much it costs and how to keep it safe.

Full-fidelity tracing of every LLM request is expensive in storage and processing. Sampling becomes necessary above a few hundred requests per second.

Head sampling, decided randomly at trace start, versus tail sampling, decided by policy at trace end

The figure simplifies: only scores from synchronous evaluations are available when the sampling decision is made (see below).

Figure 10. Head sampling decides at trace start, randomly. Tail sampling decides at trace end, by policy. Tail sampling preserves errors, slow traces and traces flagged by a synchronous evaluation deterministically.

  • Always keep error traces.
  • Always keep abnormally slow traces.
  • Always keep traces where a synchronous evaluation (toxicity filter, refusal classifier, run during the request) is below threshold.
  • Probabilistic 10 percent on the remainder, for baseline visibility.

Implementation: OpenTelemetry Collector tail_sampling processor. It can only decide on what the trace contains when its wait time expires (decision_wait, 30 seconds by default): errors, latencies, attributes set during the request. Asynchronous evaluator scores and user feedback arrive much later; they cannot drive this decision. To keep traces that are scored after the fact, keep all raw traces on a short retention (a few days), then promote to long-term storage those that an evaluator scored poorly or that received negative feedback.

The processor’s memory cost depends on the number of traces held pending (num_traces) and on the decision_wait duration, multiplied by the average trace size: size these two parameters from your throughput.

This section is the site’s reference on the cost of observability itself. The cost of non-observability (the economic exposure of a poorly observed system, with its calculator) is covered in the GenAI method, Part V.

At scale, the cost of AI observability can rival the cost of the LLM calls themselves. A simple model:

total_obs_cost = storage_cost + processing_cost + judge_cost
stored_GB = RPS * avg_payload_KB * 86400 * retention_days / 1024 / 1024
storage_cost = stored_GB * cost_per_GB_month (per month)
processing_cost = collector_cost + ingestion_cost
judge_cost = RPS * 86400 * sample_rate * tokens_per_judge * cost_per_token (per day)
Example at 100 RPS, 4 KB average payload, 30-day hot retention:
stored_GB = 100 * 4 * 86400 * 30 / 1024 / 1024
= 988.8 GB, about 1 TB kept in hot storage
storage_cost = 988.8 * cost_per_GB_month (per month)
Example judge at 100 RPS, 10 percent sample, 1500 tokens per judge:
judgments = 100 * 86400 * 0.10
= 864,000 judgments per day
judge tokens = 864,000 * 1500
= 1,296 million tokens per day for one evaluator
judge_cost = 1,296,000,000 * cost_per_token (per day)

Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list for cost_per_GB_month and cost_per_token. Compare both items over the same period: storage is billed per month, the judge consumes every day (about 39 billion tokens over 30 days in this example). A single per-token rate is also a simplification, since output tokens are usually billed higher than input tokens.

  1. Payload size: prompts, completions and retrieved contexts can be tens of kilobytes each. Storage volume grows linearly with traffic.
  2. LLM-as-judge evaluations: consume tokens at evaluation time. A trivial evaluator running on every trace can double total token usage.
  3. High-cardinality attributes: user_id, request_id, session_id inflate metrics storage if attached to metrics rather than traces.
  • Compress and tier old trace storage (hot for seven days, warm for thirty, cold archived beyond).
  • Use small models for online evaluation, large models only for periodic audit batches.
  • Restrict high-cardinality attributes to traces, keep metrics low-cardinality through aggregation.
  • Sample LLM-as-judge runs rather than evaluating every trace.
  • Tier evaluators by cost and run expensive ones on a small sample plus flagged traces.

Metrics cardinality is the single largest scaling risk in an AI observability platform. AI workloads have natural high-cardinality dimensions: prompt versions, model versions, retriever index versions, tenant IDs, user IDs.

  • Attributes with thousands of distinct values per day belong on traces, not on metrics.
  • Metrics dimensions should be bounded: model_name (tens), feature (tens), tenant (hundreds at most).
  • Per-user metrics are an anti-pattern: query traces by user_id instead.
  • Changing backends does not solve the problem. Some backends (VictoriaMetrics, Grafana Mimir, Thanos, or managed offerings) claim better handling of high cardinality than Prometheus alone; these are vendor statements, to be measured on your workload. The real answer remains to fix the metrics design.

Prompts and completions frequently contain personal data. Three controls are recommended, all the more so in regulated contexts.

  • Regex-based for known patterns: email, phone, IBAN, credit card, social security numbers. The Collector’s redaction processor does this (blocked values masked or hashed, non-allowed keys deleted).
  • NER-based for named entities: person names, organizations, locations. The Collector does not do this natively: it takes an external service (for example a personal data detection engine called by the application or by an intermediate component) or a custom processor.
  • LLM-based for context-sensitive scrubbing where regex and NER miss patterns. Same constraint: an external component, with its cost and latency.
  • Run regex redaction in the collector, before export, and place NER or LLM processing upstream of storage; never in the storage backend, once the data has been written.
  • Role-based access control on trace queries is the minimum.
  • Field-level access where the backend supports it: attribute keys are readable by all, but prompt and completion bodies require elevated access.
  • Audit log on all trace query access, retained per regulatory requirement.
  • Aligned with the storage limitation principle (GDPR Article 5(1)(e)) and with sectoral regulations.
  • Distinct retention for trace metadata (longer) and trace bodies (shorter).
  • Automated purge processes verified by audit.

LLM observability introduces a new attack surface that classical APM does not face.

  • Stored prompts may contain credentials accidentally pasted by users (API keys, tokens, secrets).
  • Stored completions may contain output that leaked training data through memorization.
  • Evaluation pipelines calling external LLM judges may exfiltrate sensitive data to third-party providers.
  • Trace query interfaces expose the full content of past interactions to anyone with access.
  • Prompt-injection attacks may target the judge model itself, manipulating evaluator scores.
  • Pre-storage scrubbing for credential patterns alongside PII.
  • Default-deny on outbound LLM judge calls, with an explicit allowlist of permitted providers.
  • Network egress controls from the collector to judge endpoints.
  • Audit log on all trace query access, with periodic review.
  • Encryption at rest and in transit for trace storage.
  • Judge prompts hardened against injection: sanitize the input fields, use delimiters, never let the rated content occupy the system prompt.
  • Platform owns the stack (collectors, backends, dashboards, alert routing), ML owns the evaluators, quality thresholds and the regression set, security owns redaction rules, access policies and audit review.
  • All three jointly run the weekly review of failed and low-score traces: in my view, the most useful recurring forum and the only one that closes the loop.
  • The full RACI matrix (including product, compliance and quality on-call) is in the GenAI method, Part VII, which is the reference; the AI Act and regulatory evidence are in Part VI.

Next: 8. Versioning, replay and experimentation.

Revised on 2 October 2026: prices removed, tail sampling limits (asynchronous evaluations and user feedback), memory sizing, NER and LLM redaction outside the Collector, GDPR Article 5, NIS2 and DORA location requirements corrected.

Revised on 4 October 2026: section 7.2 designated as the reference on the cost of observability, with a link to the method for the cost of non-observability; section 7.6 summarized in three lines with a link to the method’s RACI.