Cross-cutting
The layers of observability
For: engineers and SREs · architectsPrerequisites: None.
Observability is spoken of as one topic, yet the automation engineer, the network engineer, the platform team, the developer, the data engineer and the ML team observe neither the same things, nor with the same signals, nor with the same tools. This diagram stacks them in the manner of the OSI model, from the industrial field up to AI models, with two bands that cross them all: security, and cost weighed against business value. The tools named are examples, open source and commercial, listed alphabetically with no ranking.
Click a layer or a band, or walk through them with the keyboard (Tab, or up and down arrows). The example follows a fictitious shop-floor reading up to a failure prediction model; infrastructure carries the platform without being on the logical path.
Try this
Section titled “Try this”- Follow the end-to-end example: a vibration reading from a fictitious shop floor crosses the field, the network, the platform, the monitoring application and the data before feeding a failure prediction model. At each step, ask which signal would warn you if it broke.
- Compare two neighboring layers: open Infrastructure, then Applications. The USE method looks at resources, the RED method looks at requests; a slow service on idle servers points the investigation somewhere else than the reverse.
- Open Field, then Data: in both cases a value can arrive on time and be wrong. The OPC UA quality code and data tests answer the same need at two different layers.
- Open both bands: security and cost are not extra layers. Each asks every other layer a question: who has access, and what it costs compared with what it brings.
The rule to remember
Section titled “The rule to remember”A symptom is read with the layer below and the layer above. The layer below often gives the cause, the layer above measures the impact. An observability stack that covers a single layer answers “does it work here?” well, and “why doesn’t the business see what it expects?” badly.
Go further: the clickable architecture for an example stack that collects and stores these signals, the trace explorer for the application layer, and the LLM observability training for the AI layer.