Technical
Simulator: tail sampling
For: engineers and SREs · architectsPrerequisites: Basic notions of distributed traces and the OpenTelemetry Collector.
Keeping 100% of traces is expensive; keeping 10% at random means throwing away 90% of incidents. Head sampling decides when the request starts, before knowing whether it will fail. Tail sampling waits until the trace is complete, then decides: every error, every slow trace, and a small percentage of the rest. The price of that intelligence is Collector memory, which must hold every trace until the decision is made.
Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. Without a price, the storage tile shows the volume to be billed in GB-months; enter your price per GB-month to get an amount.
Result
Breakdown by category
| Category | Incoming | Kept (tail) | Kept (head) |
|---|
Assumptions: error and slowness are treated as independent (a trace can be both; it is then counted as an error). KB and GB in powers of 1,000. Real Collector memory is often several times the wire size (in-memory structures, traffic peaks): plan headroom. The storage price is empty by default: enter the one from your provider’s dated price list to get an amount.
Things to try
Section titled “Things to try”- Look at the “Errors kept” tile: with the defaults, tail sampling keeps 100% of errors; head sampling, for the same stored volume, keeps only about 13%. The chart says the same thing: two bars of equal length, two very different compositions. Both policies keep about 269 traces per second (13% of traffic), that is 466 GB per day and 6,985 GB-months to be billed over 15 days of retention.
- Untick “Keep 100% of error traces”: errors fall to the probabilistic rate like everything else. You are back to roughly head sampling behavior, with the memory cost on top.
- Raise the share of slow traces to 30%: the kept volume explodes, from 269 to 725 traces per second (36% of traffic) and from 6,985 to 18,789 GB-months. A badly tuned latency threshold cancels the saving: it should target the tail of the distribution, not a third of the traffic.
- Raise the wait before decision from 30 to 120 seconds: Collector memory quadruples, from 1.2 to 4.8 GB. Long traces (asynchronous processing) force a longer wait, hence bigger Collectors or traces spread across instances by trace ID.
How it works
Section titled “How it works”With head sampling every trace has the same probability of being kept: the sample mirrors the traffic, and rare errors stay rare. With tail sampling the decision is based on the trace content, so what matters for diagnosis can be over-represented. Two practical constraints: all spans of a trace must reach the same Collector (load balancing by trace ID), and memory scales roughly as traces/s × wait × size, with headroom for peaks.
The simulator assumes error and slowness are independent; in reality they are often correlated (a dependency timing out yields a trace that is both slow and failed), which slightly reduces the kept volume.
Further reading: the business dimension on telemetry cost, the Collector cost simulator, the cardinality simulator and the glossary.
Revised on 2 October 2026: prices removed; storage shows in GB-months, and in euros only with your price.