Technical
OpenTelemetry end to end
For: engineers and SREs · architects · team managersPrerequisites: Having read Observability in 10 minutes, or knowing what a metric, a log and a trace are.
Observability in 10 minutes introduced the signals on one example: a slow payment. This page shows how those signals are produced and transported with OpenTelemetry. We follow a single span, the bank call, from the application to storage.
What OpenTelemetry is
Section titled “What OpenTelemetry is”OpenTelemetry (often shortened to OTel) is an open source project of the CNCF, the foundation that also hosts Kubernetes and Prometheus. It was born from the merger of two earlier projects, OpenTracing and OpenCensus (official overview).
OpenTelemetry is not a storage or visualization tool. It produces and transports telemetry. Storage is your choice: Prometheus, Thanos, Mimir or VictoriaMetrics for metrics, Tempo or Jaeger for traces, Loki or Elasticsearch for logs, or a commercial solution.
The project brings together six building blocks:
| Building block | Role, in one sentence |
|---|---|
| Specification | the text that says how everything must work, language by language |
| API | the functions code calls to create a span or a metric |
| SDK | the library implementing the API: it samples, batches and exports |
| OTLP | the transport protocol shared by all signals |
| Collector | a standalone program that receives, transforms and forwards telemetry |
| Semantic conventions | the list of standard attribute names and their meaning |
The journey of a span
Section titled “The journey of a span”Here is the complete path. Each step is detailed below.
flowchart LR
A["Application<br/>payment service"] -->|"API"| S["SDK<br/>sampling, batching"]
S -->|"OTLP"| R
subgraph COL["Collector"]
R["Receiver<br/>otlp"] --> P["Processors<br/>memory_limiter, batch"] --> E["Exporters"]
end
E -->|"OTLP"| B[("Trace<br/>storage")]
E --> D["debug<br/>console"]
1. Instrumentation creates the span
Section titled “1. Instrumentation creates the span”Instrumenting means adding to a program the code that produces telemetry. There are two ways to do it.
Auto-instrumentation (zero-code) requires no code change. An agent or launcher intercepts well-known libraries: web server, HTTP client, database driver. It creates spans on its own. It is the best way to start.
Manual instrumentation means calling the API in your own code. It adds what auto-instrumentation cannot guess: a business step, an attribute such as the card type.
In practice, you combine both. Auto-instrumentation gives the skeleton of the trace. Manual instrumentation adds meaning.
2. The SDK prepares the export
Section titled “2. The SDK prepares the export”The SDK receives finished spans. It decides whether to keep them, which is sampling. It groups them into batches to limit the number of sends. It adds the resource attributes. Then it hands them to an exporter.
3. OTLP transports
Section titled “3. OTLP transports”OTLP (OpenTelemetry Protocol) is the transport format. It runs over gRPC (port 4317 by default) or HTTP (port 4318) (Collector configuration). The same protocol carries metrics, logs and traces.
4. The Collector controls
Section titled “4. The Collector controls”The Collector is a program that sits between applications and storage. It has three kinds of components.
- Receivers take data in. The
otlpreceiver listens on ports 4317 and 4318. - Processors transform data, in the order they are listed.
memory_limiterrefuses data when memory gets close to its limit, so the Collector does not crash.batchgroups data before sending. - Exporters send data to one or more storage backends.
A pipeline connects receivers, processors and exporters for one signal type. You write one pipeline for traces, another for metrics.
The Collector is not mandatory: an SDK can export straight to a backend. I still recommend using it from day one. It is the single place where you protect, filter and route data without touching the applications.
5. Storage keeps and displays
Section titled “5. Storage keeps and displays”The backend indexes spans by trace id. The interface rebuilds the waterfall seen on the previous page. The trace explorer shows what that result looks like.
A minimal example
Section titled “A minimal example”The following example instruments a Python service and runs a Collector that prints received data to its console.
Instrument a Python service without touching code
Section titled “Instrument a Python service without touching code”The commands come from the OpenTelemetry Python zero-code documentation. app.py is your application.
pip install opentelemetry-distro opentelemetry-exporter-otlpopentelemetry-bootstrap -a install
export OTEL_SERVICE_NAME=paymentexport OTEL_RESOURCE_ATTRIBUTES=service.version=1.4.2,deployment.environment.name=devexport OTEL_TRACES_EXPORTER=otlpexport OTEL_METRICS_EXPORTER=otlpexport OTEL_EXPORTER_OTLP_PROTOCOL=grpcexport OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
opentelemetry-instrument python app.pyopentelemetry-bootstrap detects installed libraries and adds the matching instrumentations. opentelemetry-instrument starts the application with those instrumentations active.
Add a business span by hand
Section titled “Add a business span by hand”To see the “authorize” step in the trace, with the card type, a few lines are enough:
from opentelemetry import trace
tracer = trace.get_tracer("payment")
def authorize(order): with tracer.start_as_current_span("authorize_payment") as span: span.set_attribute("app.payment.card_type", order.card_type) return call_bank(order)This span becomes the child of the HTTP span created automatically. The app. prefix flags an application-specific attribute, outside the standard conventions.
Configure the Collector
Section titled “Configure the Collector”receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318
processors: memory_limiter: check_interval: 1s limit_percentage: 80 spike_limit_percentage: 20 batch: {}
exporters: debug: verbosity: detailed # To a trace backend that accepts OTLP (Tempo, Jaeger, etc.). otlp_grpc/traces: endpoint: traces-backend:4317 tls: insecure: true # lab only, never in production
service: pipelines: traces: receivers: [otlp] processors: [memory_limiter, batch] exporters: [debug, otlp_grpc/traces] metrics: receivers: [otlp] processors: [memory_limiter, batch] exporters: [debug]Three notes to read this file.
memory_limitercomes first: it must refuse data before any other processing.batchcomes last, after processors that filter or transform.- The
debugexporter prints each span to the Collector console. It is the simplest way to check that data arrives. - Recent Collector versions name the OTLP gRPC exporter
otlp_grpcand the OTLP HTTP exporterotlp_http. The old namesotlpandotlphttpare still accepted as deprecated aliases: some configurations on this site still use them. Check the name your version expects.
The 0.0.0.0 address listens on all interfaces. The Collector documentation uses it for convenience and recommends localhost when all clients are local.
Resources and semantic conventions
Section titled “Resources and semantic conventions”Two notions come up everywhere. They are simple, but they make all the difference when searching.
The resource describes the emitter. It is attached once and for all to everything the service produces. Examples: service.name="payment", service.version="1.4.2", deployment.environment.name="dev". In the example, it comes from the OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES variables. Without service.name, there is no way to know which service emitted a span.
Semantic conventions set the name and meaning of common attributes (semantic conventions). An HTTP request carries http.request.method and http.response.status_code. A route carries http.route. Everyone uses the same names, whatever the language.
The benefit is concrete. If a Java service and a Python service follow the conventions, a single query finds all HTTP 500 errors from both. If each invents its own names (status, httpCode, return_code), you need one query per team.
The conventions are still evolving. Some names have changed in recent years. I recommend pinning the version you follow and reviewing it at every SDK upgrade.
Going further
Section titled “Going further”- Animated Collector pipeline: switch processors on and off to see their effect.
- Tail sampling: keep every error trace without storing everything.
- Collector cost: estimate the volume each processing step saves.
- Clickable architecture: a complete stack, block by block.
- SLOs and alerting: turn these signals into objectives and useful alerts.
- The glossary defines Collector, span, trace and semantic convention.
Sources
Section titled “Sources”- OpenTelemetry, What is OpenTelemetry?: origin of the project, building blocks and backend independence.
- OpenTelemetry, Python zero-code instrumentation: install and launch commands.
- OpenTelemetry, Collector configuration: receivers, processors, exporters, pipelines, ports 4317 and 4318.
- OpenTelemetry, Semantic conventions: standard attribute names.