Technical
Animated diagram: the OpenTelemetry Collector pipeline
For: engineers and SREs · architectsPrerequisites: Know what an OpenTelemetry Collector is for.
An OpenTelemetry Collector receives signals from applications and ships them to storage. In between, a chain of processors decides what is kept, transformed or dropped. It is the last place where you control the data before it becomes a line on the bill, exposed personal data or audit evidence. Below, each dot is a signal: its color gives the type, its letter what it carries. Click a processor to turn it off.
Click a processor (or Tab then Enter) to disable it.
Collector configuration
Recommended order: memory_limiter first; filtering and redaction early, so nothing sensitive or useless travels further; tail_sampling after redaction; batch last, right before export, or, instead, batching inside the exporter (sending_queue.batch).
What reaches the storage backends
| Destination | Received | Noise | Sensitive | Errors |
|---|
Illustrative simulation: signal mix, 10% sampling and memory threshold are assumptions. Signals refused by memory_limiter are in reality retried by the sender, within the limits of its own queue.
Try this
Section titled “Try this”- Turn off
filter: dots marked H (health checks) and D (debug logs) reach Tempo and VictoriaLogs. The “Noise” column climbs and stored volume grows with no diagnostic gain. - Turn off
transform: email addresses (@) arrive in clear and metrics keep theiruser_idlabel (U). The sensitive data counter turns red. Note thatuser_idon a metric also blows up cardinality: see the cardinality simulator. - Turn off
memory_limiter, then click “Traffic spike”: the Collector’s memory exceeds its capacity, it gets killed (OOM) and every in-flight signal is lost. Turn it back on and spike again: the Collector refuses part of the sends, which senders retry, but it stays up. - Turn off
tail_sampling, thenrouting: without sampling, every nominal trace is stored although all errors were already kept; without routing, audit logs (A) stay in short-retention hot storage and the “Audit logs in cold storage” rate drops to zero.
Order matters
Section titled “Order matters”Processors run in the order they are listed in service.pipelines, not in the order of the processors section. Four rules to remember:
memory_limiterfirst. It must be able to refuse data before the Collector spends memory processing it.- Filter and redact early. What is dropped early costs nothing downstream, and no sensitive data should travel through sampling, batching or an exporter.
tail_samplingafter redaction. The decision is made on the complete trace (every error, a fraction of nominal traffic). It assumes all spans of a trace reach the same Collector, hence aloadbalancingexporter upstream as soon as you run several instances.batchlast, right before export: batching data that will be dropped afterwards is wasted work. The alternative is to batch inside the exporter itself (sending_queuewithbatch: {}, off by default): this is the direction the project is taking to replace the processor (deprecation proposal), but as of 2 October 2026 thebatchprocessor is still beta and not formally deprecated.
Routing to cold storage goes through the routing connector: the former routing processor was removed in favor of the connector. The exact syntax of its routing table has changed across Collector contrib versions: the configuration shown follows v0.162 (prefixed OTTL condition, log.attributes[...], no context field); check the documentation for your version before reusing it. The awss3 exporter targets S3-compatible storage outside AWS here, hence endpoint and s3_force_path_style.
What to measure for real
Section titled “What to measure for real”The Collector observes itself: it exposes its own metrics (Prometheus format on port 8888 by default, configurable under service.telemetry). Things to watch:
| Question | Collector internal metric (prefix otelcol_) |
|---|---|
| How many signals come in? | receiver_accepted_spans, receiver_accepted_log_records, receiver_accepted_metric_points |
| How many are refused at the door? | receiver_refused_*: refusals by memory_limiter or a saturated exporter |
| How many go out, how many fail? | exporter_sent_*, exporter_send_failed_* |
| Is the export queue saturating? | exporter_queue_size relative to exporter_queue_capacity |
| Is the Collector short of memory? | process_memory_rss, process_runtime_heap_alloc_bytes, and container restarts on the Kubernetes side |
Exact names vary slightly across versions (_total suffixes, newer processor metric names): start from your Collector’s /metrics page. For sensitive data, no metric replaces a check: a scheduled query in VictoriaLogs or Tempo looking for an email pattern, with an alert if it finds anything. For audit, compare the number of audit logs emitted with the number of objects written to cold storage.
Going further: the glossary, the technical view, and on the cost side the business page.
Revised on 2 October 2026: configuration aligned with Collector contrib v0.162 (trace_conditions and log_conditions filter syntax, prefixed transform statements, otlp_grpc and otlp_http exporters, S3 storage endpoint), and exporter-side batching mentioned.