Skip to content

Observability, from signal to decision

An open site, in French and English, so that engineering, finance, business lines and leadership finally mean the same thing when they talk about production.
trace 7d3c…5f60200 OK2.78 s
  1. POST /api/ask
  2. invoke_agent support-assistant
  3. embeddings mistral-embed
  4. query kb_support_frtop_k=1, config changed
  5. chat grand-modele-v1
  6. execute_tool get_order_status
  7. chat grand-modele-v1
  8. evaluate faithfulness0.34 < 0.7
200 OK in 2.8 s. And a wrong answer.

MTTL stands for Mean Time To Learn: the average time it takes to understand. The name echoes MTTR (mean time to recovery), the average time to restore service after an incident: you only recover fast from what you first understood fast. This site exists to shorten that learning time, for engineering teams as much as for finance, business lines and leadership (why this site).

If you are not sure where you stand, the self-assessment takes ten minutes and suggests an action plan. Otherwise, start from your role:

Read the video text

Observability. Understanding a system from what it lets you see.

A very old idea.

  • Antiquity: In Egypt, nilometers measure the Nile flood to anticipate the harvest and set taxes.
  • 1788: Watt’s flyball governor regulates the steam engine’s speed on its own.
  • 1868: Maxwell describes how these governors work in equations.
  • 1960: Rudolf Kálmán defines observability: inferring a system’s internal state from its outputs.
  • 1988: SNMP: querying network equipment.
  • 1999: NetSaint, later Nagios: monitoring with checks and thresholds.
  • 2010: Google describes Dapper: following a request across distributed services.
  • 2012: Prometheus is born at SoundCloud: labelled metrics, queried on the fly.
  • 2019: OpenTelemetry: an open standard to collect telemetry.
  • Today: Observing AI systems too, which can answer fast and wrong.

From monitoring to observability. Monitoring: Is it working? Metrics and thresholds set in advance. Observability: Why is it behaving this way? Questions nobody planned for. The signals: Metrics, Logs, Traces, Profiles, Events.

The goals: Detect, Understand, Decide, Learn. Shorten the time between the incident and the lesson: Mean Time To Learn.

What is at stake:

  • Business: The cost of outages and of telemetry itself.
  • Organization: The evidence required by NIS2, DORA, the AI Act.
  • People: Sustainable teams and on-call.
  • Business lines: The trust of customers and users.

The layers of observability:

  1. Field: industrial OT, PLCs, sensors, IoT
  2. Network: links, protocols, latency
  3. Infrastructure: servers, storage, GPUs
  4. Platform: containers, cloud, high performance computing
  5. Applications: services, APIs, user experience
  6. Data: pipelines, quality, freshness
  7. AI and models: LLMs, agents, evaluation

Security, Cost and business value: across every layer.

MTTL, Mean Time To Learn: Observability, from signal to decision. www.meantimetolearn.com: Launching on October 7, 2026.

Music: “In the Aftermath”, Michael Rothery (Epidemic Sound).

To go further, the interactive layers diagram details each layer, from the industrial field to AI models: what to observe, the signals, the indicators and the tools.

Why this site exists and how it is written: the About page and the manifesto.

The same signal is worth different things depending on who reads it. Payment latency matters to on-call; the sales director wants to know how many carts were abandoned at that step.

The technical signalWhat it becomes for others
payment service latencybusiness line: carts abandoned at the payment step
tokens consumed by an AI featurefinance: cost per customer conversation
retained traces and logscompliance: evidence that holds in an audit
night alerts per personmanagement: team sustainability

The full approach is on the Bridges page, and every article ends with a box stating what it changes beyond the technical team.

  1. What a business or product owner can ask of observabilityBusiness
  2. 3D diagram: AI in cross-sectionTechnical
  3. The executive dashboardOrganization
  4. Explorable diagram: an HPC AI clusterTechnical
  5. Team models for observabilityOrganization
  6. Governing telemetryOrganization

To follow new publications: the RSS feed.

The labs published so far are freely available on the Labs & Trainings forge, the public code repository where I publish the code of the labs and the course material; the Labs page says which exist and which are in preparation. My videos, on projects close to MTTL, are on the YouTube channel; lab demos are planned (see Videos).