Skip to content

TechnicalBeginner

OpenTelemetry end to end

For: engineers and SREs · architects · team managersPrerequisites: Having read Observability in 10 minutes, or knowing what a metric, a log and a trace are.

Reading mode

Observability in 10 minutes introduced the signals on one example: a slow payment. This page shows how those signals are produced and transported with OpenTelemetry. We follow a single span, the bank call, from the application to storage.

OpenTelemetry (often shortened to OTel) is an open source project of the CNCF, the foundation that also hosts Kubernetes and Prometheus. It was born from the merger of two earlier projects, OpenTracing and OpenCensus (official overview).

OpenTelemetry is not a storage or visualization tool. It produces and transports telemetry. Storage is your choice: Prometheus, Thanos, Mimir or VictoriaMetrics for metrics, Tempo or Jaeger for traces, Loki or Elasticsearch for logs, or a commercial solution.

The project brings together six building blocks:

Building blockRole, in one sentence
Specificationthe text that says how everything must work, language by language
APIthe functions code calls to create a span or a metric
SDKthe library implementing the API: it samples, batches and exports
OTLPthe transport protocol shared by all signals
Collectora standalone program that receives, transforms and forwards telemetry
Semantic conventionsthe list of standard attribute names and their meaning

Here is the complete path. Each step is detailed below.

flowchart LR
  A["Application<br/>payment service"] -->|"API"| S["SDK<br/>sampling, batching"]
  S -->|"OTLP"| R
  subgraph COL["Collector"]
    R["Receiver<br/>otlp"] --> P["Processors<br/>memory_limiter, batch"] --> E["Exporters"]
  end
  E -->|"OTLP"| B[("Trace<br/>storage")]
  E --> D["debug<br/>console"]

Instrumenting means adding to a program the code that produces telemetry. There are two ways to do it.

Auto-instrumentation (zero-code) requires no code change. An agent or launcher intercepts well-known libraries: web server, HTTP client, database driver. It creates spans on its own. It is the best way to start.

Manual instrumentation means calling the API in your own code. It adds what auto-instrumentation cannot guess: a business step, an attribute such as the card type.

In practice, you combine both. Auto-instrumentation gives the skeleton of the trace. Manual instrumentation adds meaning.

The SDK receives finished spans. It decides whether to keep them, which is sampling. It groups them into batches to limit the number of sends. It adds the resource attributes. Then it hands them to an exporter.

OTLP (OpenTelemetry Protocol) is the transport format. It runs over gRPC (port 4317 by default) or HTTP (port 4318) (Collector configuration). The same protocol carries metrics, logs and traces.

The Collector is a program that sits between applications and storage. It has three kinds of components.

  • Receivers take data in. The otlp receiver listens on ports 4317 and 4318.
  • Processors transform data, in the order they are listed. memory_limiter refuses data when memory gets close to its limit, so the Collector does not crash. batch groups data before sending.
  • Exporters send data to one or more storage backends.

A pipeline connects receivers, processors and exporters for one signal type. You write one pipeline for traces, another for metrics.

The Collector is not mandatory: an SDK can export straight to a backend. I still recommend using it from day one. It is the single place where you protect, filter and route data without touching the applications.

The backend indexes spans by trace id. The interface rebuilds the waterfall seen on the previous page. The trace explorer shows what that result looks like.

The following example instruments a Python service and runs a Collector that prints received data to its console.

Instrument a Python service without touching code

Section titled “Instrument a Python service without touching code”

The commands come from the OpenTelemetry Python zero-code documentation. app.py is your application.

Fenêtre de terminal
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install
export OTEL_SERVICE_NAME=payment
export OTEL_RESOURCE_ATTRIBUTES=service.version=1.4.2,deployment.environment.name=dev
export OTEL_TRACES_EXPORTER=otlp
export OTEL_METRICS_EXPORTER=otlp
export OTEL_EXPORTER_OTLP_PROTOCOL=grpc
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
opentelemetry-instrument python app.py

opentelemetry-bootstrap detects installed libraries and adds the matching instrumentations. opentelemetry-instrument starts the application with those instrumentations active.

To see the “authorize” step in the trace, with the card type, a few lines are enough:

from opentelemetry import trace
tracer = trace.get_tracer("payment")
def authorize(order):
with tracer.start_as_current_span("authorize_payment") as span:
span.set_attribute("app.payment.card_type", order.card_type)
return call_bank(order)

This span becomes the child of the HTTP span created automatically. The app. prefix flags an application-specific attribute, outside the standard conventions.

receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
batch: {}
exporters:
debug:
verbosity: detailed
# To a trace backend that accepts OTLP (Tempo, Jaeger, etc.).
otlp_grpc/traces:
endpoint: traces-backend:4317
tls:
insecure: true # lab only, never in production
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug, otlp_grpc/traces]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [debug]

Three notes to read this file.

  • memory_limiter comes first: it must refuse data before any other processing. batch comes last, after processors that filter or transform.
  • The debug exporter prints each span to the Collector console. It is the simplest way to check that data arrives.
  • Recent Collector versions name the OTLP gRPC exporter otlp_grpc and the OTLP HTTP exporter otlp_http. The old names otlp and otlphttp are still accepted as deprecated aliases: some configurations on this site still use them. Check the name your version expects.

The 0.0.0.0 address listens on all interfaces. The Collector documentation uses it for convenience and recommends localhost when all clients are local.

Two notions come up everywhere. They are simple, but they make all the difference when searching.

The resource describes the emitter. It is attached once and for all to everything the service produces. Examples: service.name="payment", service.version="1.4.2", deployment.environment.name="dev". In the example, it comes from the OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES variables. Without service.name, there is no way to know which service emitted a span.

Semantic conventions set the name and meaning of common attributes (semantic conventions). An HTTP request carries http.request.method and http.response.status_code. A route carries http.route. Everyone uses the same names, whatever the language.

The benefit is concrete. If a Java service and a Python service follow the conventions, a single query finds all HTTP 500 errors from both. If each invents its own names (status, httpCode, return_code), you need one query per team.

The conventions are still evolving. Some names have changed in recent years. I recommend pinning the version you follow and reviewing it at every SDK upgrade.