Skip to content

TechnicalBeginner

Observability in 10 minutes

For: engineers and SREs · architects · team managers · business and productPrerequisites: None. Knowing what a web application is will do.

Reading mode

This page is for someone discovering the topic. It sets out the vocabulary used everywhere else on the site. A single example runs through it: a slow payment on an online shop.

Picture an online shop. The customer clicks “Pay”. Usually, the confirmation page shows up in under a second. This morning, some customers wait several seconds. A few give up.

This scenario is illustrative. It is deliberately ordinary, because this is the kind of problem observability should help solve.

Behind the “Pay” button, several services are at work. A service is a program that performs one specific function. Here, a checkout service receives the order. It calls a payment service, which in turn calls the bank. The checkout service also reads the cart from a database.

flowchart LR
  C["Customer"] --> K["checkout"]
  K --> DB[("Cart database")]
  K --> P["payment"]
  P --> B["Bank API"]

The question is simple: why are some payments slow?

Monitoring watches indicators chosen in advance. For example, you decide to track the error rate and the response time. You set a threshold. An alert fires when the threshold is crossed.

Monitoring answers planned questions well. “Is the service up?” “Is the response time above 2 seconds?”

Our problem does not fit those boxes. Only some payments are slow. Maybe those of a specific bank, a country or a card type. Nobody had planned that question.

Observability is the ability to answer a question you had not planned. You ask it of the data already collected, without changing code or redeploying. It requires rich data that is linked together.

The two are not opposed. Monitoring tells you there is a problem. Observability helps you understand which one.

A signal is a type of data an application emits about its own behavior. This is also called telemetry. Three signals are classic. Two more complete the picture.

A metric is a numeric measurement taken at regular intervals. For example, the number of payments per minute or their duration.

Each metric has a name and labels. A label is a key and value pair that specifies what is measured. For example service="payment" or status="error".

Metrics are compact and cheap to keep for a long time. They feed dashboards and alerts. In our example, a chart shows that payment duration has gone up since 9 am. It does not say why.

A trace tells the journey of one request through all the services. It is made of spans. A span is a step with a start, an end and attributes. For example “call to the bank, 3.8 seconds”.

Spans nest inside each other. The checkout span contains the payment span, which contains the bank call span. You can see at a glance where the time goes.

POST /checkout (checkout) |=========================================| 4,200 ms
read cart (database) |=| 70 ms
authorize (payment) |=======================================| 3,950 ms
bank API call (bank) |=====================================| 3,850 ms

The durations above are illustrative. The trace shows that the payment service is waiting for the bank. The problem is neither in the database nor in checkout.

A log is a timestamped message describing an event. For example “retrying connection to the bank, timeout exceeded”.

A log is all the more useful when it is structured, meaning written as fields (bank="bank-b", retry=2) rather than as free text. You can then filter and count it.

In our example, the payment service logs show retries towards a single bank. We have the likely cause.

A profile measures where a program spends CPU or memory, function by function. It answers “which line of code is expensive?”. In OpenTelemetry, profiles are still under development (signals documentation).

An event marks a one-off fact that matters to operations. A deployment, a configuration change, a failover. Overlaid on a chart, it answers “what changed at 9 am?”.

SignalQuestion it answersIn the example
MetricIs there a problem, since when, how big?payment duration rising since 9 am
TraceWhere, in which service, at which step?time spent in the bank call
LogWhat exactly happened?retries towards a single bank
ProfileWhich code consumes resources?useful if the service itself were slow
EventWhat changed?a payment deployment at 8:55 am

Taken separately, each signal gives one piece of the story. Observability starts when you can jump from one to another in a click. Two mechanisms make it possible.

The trace id. Every trace gets a unique identifier, the trace id. It travels from service to service with the request, in a standard HTTP header defined by the W3C (Trace Context). If logs also carry this trace id, you can jump from a slow span to the exact logs of that request.

A metric can also point to a trace. This link is called an exemplar: a point on the chart leads to a representative request.

Resource attributes. A resource describes who emits the signal: the service name, its version, the environment, the host. If metrics, logs and traces carry the same resource attributes, you filter all three signals the same way. For example service.name="payment".

flowchart LR
  M["Metric<br/>payment duration"] -->|"exemplar"| T["Trace<br/>slow bank call"]
  T -->|"trace id"| L["Logs<br/>retries"]
  M -.->|"service.name"| L

This is what OpenTelemetry does, the open standard presented in OpenTelemetry end to end. It gives the same identifiers and the same attribute names to every signal.

The cardinality of a metric is the number of distinct combinations of its labels. Each combination creates a time series, a sequence of values to store (Prometheus documentation). A bank label with 20 values and a status label with 3 values give at most 60 series. Adding a customer_id label with 100,000 customers multiplies that number by 100,000 (illustrative calculation). Cost and slowness follow. The rule I recommend: a unique identifier (customer, order, request) goes into traces and logs, never into the labels of a metric. The cardinality simulator lets you check this on your own numbers.

  • OpenTelemetry, Signals: traces, metrics, logs, baggage and the status of profiles and events.
  • W3C, Trace Context: format of the header that carries the trace id.
  • Prometheus, Metric and label naming: every label combination creates a new time series.