Observability, from signal to decision
- POST /api/ask
- invoke_agent support-assistant
- embeddings mistral-embed
- query kb_support_frtop_k=1, config changed
- chat grand-modele-v1
- execute_tool get_order_status
- chat grand-modele-v1
- evaluate faithfulness0.34 < 0.7
MTTL stands for Mean Time To Learn: the average time it takes to understand. The name echoes MTTR (mean time to recovery), the average time to restore service after an incident: you only recover fast from what you first understood fast. This site exists to shorten that learning time, for engineering teams as much as for finance, business lines and leadership (why this site).
Pick up where you left off
All my coursesWhere to start
Section titled “Where to start”If you are not sure where you stand, the self-assessment takes ten minutes and suggests an action plan. Otherwise, start from your role:
- Engineer, SRE, developerSignal fundamentals, then OpenTelemetry and the labs.The path →
- ArchitectTelemetry pipelines, cost, choosing a stack.The path →
- Team managerOn-call, SLOs, adoption by the teams.The path →
- Business line or product ownerThe questions to ask and the indicators to request.The path →
- Finance, risk, complianceUnit cost and evidence from telemetry.The path →
- Leadership, CIO, procurementBusiness case, sovereignty, reversibility.The path →
Observability in a minute and a half
Section titled “Observability in a minute and a half”Read the video text
Observability. Understanding a system from what it lets you see.
A very old idea.
- Antiquity: In Egypt, nilometers measure the Nile flood to anticipate the harvest and set taxes.
- 1788: Watt’s flyball governor regulates the steam engine’s speed on its own.
- 1868: Maxwell describes how these governors work in equations.
- 1960: Rudolf Kálmán defines observability: inferring a system’s internal state from its outputs.
- 1988: SNMP: querying network equipment.
- 1999: NetSaint, later Nagios: monitoring with checks and thresholds.
- 2010: Google describes Dapper: following a request across distributed services.
- 2012: Prometheus is born at SoundCloud: labelled metrics, queried on the fly.
- 2019: OpenTelemetry: an open standard to collect telemetry.
- Today: Observing AI systems too, which can answer fast and wrong.
From monitoring to observability. Monitoring: Is it working? Metrics and thresholds set in advance. Observability: Why is it behaving this way? Questions nobody planned for. The signals: Metrics, Logs, Traces, Profiles, Events.
The goals: Detect, Understand, Decide, Learn. Shorten the time between the incident and the lesson: Mean Time To Learn.
What is at stake:
- Business: The cost of outages and of telemetry itself.
- Organization: The evidence required by NIS2, DORA, the AI Act.
- People: Sustainable teams and on-call.
- Business lines: The trust of customers and users.
The layers of observability:
- Field: industrial OT, PLCs, sensors, IoT
- Network: links, protocols, latency
- Infrastructure: servers, storage, GPUs
- Platform: containers, cloud, high performance computing
- Applications: services, APIs, user experience
- Data: pipelines, quality, freshness
- AI and models: LLMs, agents, evaluation
Security, Cost and business value: across every layer.
MTTL, Mean Time To Learn: Observability, from signal to decision. www.meantimetolearn.com: Launching on October 7, 2026.
Music: “In the Aftermath”, Michael Rothery (Epidemic Sound).
To go further, the interactive layers diagram details each layer, from the industrial field to AI models: what to observe, the signals, the indicators and the tools.
Why this site exists and how it is written: the About page and the manifesto.
Four ways to read it
Section titled “Four ways to read it”Signals, OpenTelemetry, collection pipelines, backends, SLOs, AI observability, log integrity.
BusinessWhy, and at what cost?Telemetry cost, pricing models, business case, sovereignty, contract reversibility.
PeopleWho keeps it alive?On-call, alert fatigue, post-mortems, skills, team adoption.
OrganizationHow do we make it last?Team models, governance and data contracts, maturity, compliance (DORA, NIS2, AI Act).
Bridges, not silos
Section titled “Bridges, not silos”The same signal is worth different things depending on who reads it. Payment latency matters to on-call; the sales director wants to know how many carts were abandoned at that step.
| The technical signal | What it becomes for others |
|---|---|
| payment service latency | business line: carts abandoned at the payment step |
| tokens consumed by an AI feature | finance: cost per customer conversation |
| retained traces and logs | compliance: evidence that holds in an audit |
| night alerts per person | management: team sustainability |
The full approach is on the Bridges page, and every article ends with a box stating what it changes beyond the technical team.
Recently published
Section titled “Recently published”- What a business or product owner can ask of observabilityBusiness
- 3D diagram: AI in cross-sectionTechnical
- The executive dashboardOrganization
- Explorable diagram: an HPC AI clusterTechnical
- Team models for observabilityOrganization
- Governing telemetryOrganization
To follow new publications: the RSS feed.
Play with it instead of reading
Section titled “Play with it instead of reading”The labs published so far are freely available on the Labs & Trainings forge, the public code repository where I publish the code of the labs and the course material; the Labs page says which exist and which are in preparation. My videos, on projects close to MTTL, are on the YouTube channel; lab demos are planned (see Videos).