Skip to content

Cross-cutting

Bridges: taking observability beyond IT

For: engineers and SREs · architects · team managers · business and product · finance, risk and compliance · executives and CIOsPrerequisites: None.

Why observability projects rarely fail because of the tool

Section titled “Why observability projects rarely fail because of the tool”

When an observability project fails, it is rarely for technical reasons. The data is there, and so are the dashboards. What you find instead are situations like these:

  • finance sees a growing bill without understanding what it buys;
  • business lines do not recognize themselves in CPU curves and p99 latency;
  • compliance finds out at audit time that the evidence exists but is neither retained nor exportable;
  • leadership cannot say whether the investment reduced risk;
  • the technical team, the only one looking, wears itself out defending a budget nobody else owns.

As long as it stays inside IT, observability is seen as a cost center. Its standing changes when other departments use it to make decisions, because it gives them a shared language with the technical team.

You instrument a system to answer questions. So I recommend starting a project by listing those questions, before any talk of tools. In my view, most of them come from outside the technical team.

Who asksThe question rarely askedWhat observability can answer
Business line, productdo our customers succeed at what they came to do?completed journey rate, end-to-end time, per customer segment
Financewhat does a transaction, a customer, a feature cost?unit cost of infrastructure and telemetry, cost per LLM request
Risk, compliance, legalcan we prove what happened, to a third party?tamper-evident logs, retained traces, DORA, NIS2, AI Act evidence matrix
Customer service, supportdid this customer really suffer the incident they report?the trace of their request, with the goal of finding it in under a minute
Procurementare we locked in by our vendor?open standards, data portability, exit cost
HR, managementcan the team keep this up?on-call load, alerts per night per person, noise
Executive leadershipwhere should we invest to reduce risk?quantified exposure of blind spots, avoided loss
flowchart LR
  O(("Observability"))
  O --- B["1. Business<br/>value and cost"]
  O --- ORG["2. Organization<br/>who decides, who owns"]
  O --- MET["3. Business lines<br/>translated signals"]
  O --- SUP["4. Support functions<br/>finance, compliance, procurement"]
  O --- PER["5. People<br/>language and rituals"]

1. To the business, talking value and cost

Section titled “1. To the business, talking value and cost”
  • Express SLOs as promises made to customers. “99.5% of payments complete in under two seconds” means something to a product director, whereas “API p99 latency under 800 ms” tells them nothing.
  • Compute a unit cost, per transaction, per customer or per feature. It is the only figure a financial controller can compare month to month.
  • Quantify avoided loss from incident cost, exposure duration and blind spots. The exposure calculator gives an example for AI.

2. To the organization, saying who decides and who owns

Section titled “2. To the organization, saying who decides and who owns”
  • Who chooses what is collected, who pays for telemetry, who owns each alert? Until the answer is written down somewhere, nobody takes it on.
  • Data contracts and a RACI turn good intentions into rules. See the organization dimension and part VII of the GenAI method.

3. To business lines, translating the signals

Section titled “3. To business lines, translating the signals”
  • Instrument business events (order confirmed, file processed, loan granted) as carefully as technical calls, and link them to traces. This is sometimes called business observability.
  • Keep a translation dictionary between technical and business indicators (see the example below).
  • Build dashboards per audience. An operations dashboard is not built for an executive committee: in my view, it loses more attention than it gains.

4. To support functions (finance, compliance, procurement)

Section titled “4. To support functions (finance, compliance, procurement)”
  • Finance: a telemetry budget tracked like any other, with its drifts and levers. See the business dimension.
  • Compliance: evidence is a by-product of good instrumentation, provided it is designed in from the start. See log integrity and the evidence matrix.
  • Procurement: once the contract is signed, it is too late to secure reversibility. It has to be settled during negotiation.

5. To people, with a shared language and shared rituals

Section titled “5. To people, with a shared language and shared rituals”
  • A monthly review bringing engineering, product and finance together around three figures: service quality as customers see it, unit cost, and incidents with their causes.
  • Post-mortems open to the business lines concerned. They find out how the system works, and the technical team learns what the incident actually cost.
Technical signalBusiness indicatorDecision informed
payment service p95 latencycart abandonment at the payment stepprioritize optimization or a second payment provider
pricing engine error ratequotes not issued per day, deferred revenuechoose between an immediate fix and degraded mode
tokens consumed per AI featurecost per assisted customer conversationadjust the model, the cache or the offer’s price
faithfulness score of a RAG assistantcomplaints caused by wrong informationtrigger a knowledge base update
night alerts per on-call engineerteam turnover and sick leaverevisit thresholds, rotations, staffing

Each row connects three people who do not always talk to each other: whoever sees the signal, whoever owns the indicator and whoever decides.

For the first row, the chain reads as follows. It closes when the decision changes the signal.

flowchart LR
  S["Technical signal<br/>payment p95 latency"] -->|"translation"| I["Business indicator<br/>cart abandonment"]
  I -->|"trade-off"| D["Decision<br/>optimize or second provider"]
  D -.->|"measured effect"| S
  • The observability project is owned by IT alone, with no business sponsor.
  • Dashboards have never been opened by anyone outside the technical team.
  • No SLO is phrased in a customer’s words.
  • Telemetry cost is discovered on the invoice.
  • Compliance was not consulted on log retention.
  • Post-mortems never leave the operations team.

Every article and course ends with a Bridges box stating what the topic changes for the business, the organization, teams and business lines. I keep to this on every page, as the manifesto explains.