Cross-cutting
Bridges: taking observability beyond IT
For: engineers and SREs · architects · team managers · business and product · finance, risk and compliance · executives and CIOsPrerequisites: None.
Why observability projects rarely fail because of the tool
Section titled “Why observability projects rarely fail because of the tool”When an observability project fails, it is rarely for technical reasons. The data is there, and so are the dashboards. What you find instead are situations like these:
- finance sees a growing bill without understanding what it buys;
- business lines do not recognize themselves in CPU curves and p99 latency;
- compliance finds out at audit time that the evidence exists but is neither retained nor exportable;
- leadership cannot say whether the investment reduced risk;
- the technical team, the only one looking, wears itself out defending a budget nobody else owns.
As long as it stays inside IT, observability is seen as a cost center. Its standing changes when other departments use it to make decisions, because it gives them a shared language with the technical team.
Start from other people’s questions
Section titled “Start from other people’s questions”You instrument a system to answer questions. So I recommend starting a project by listing those questions, before any talk of tools. In my view, most of them come from outside the technical team.
| Who asks | The question rarely asked | What observability can answer |
|---|---|---|
| Business line, product | do our customers succeed at what they came to do? | completed journey rate, end-to-end time, per customer segment |
| Finance | what does a transaction, a customer, a feature cost? | unit cost of infrastructure and telemetry, cost per LLM request |
| Risk, compliance, legal | can we prove what happened, to a third party? | tamper-evident logs, retained traces, DORA, NIS2, AI Act evidence matrix |
| Customer service, support | did this customer really suffer the incident they report? | the trace of their request, with the goal of finding it in under a minute |
| Procurement | are we locked in by our vendor? | open standards, data portability, exit cost |
| HR, management | can the team keep this up? | on-call load, alerts per night per person, noise |
| Executive leadership | where should we invest to reduce risk? | quantified exposure of blind spots, avoided loss |
Five bridges to build
Section titled “Five bridges to build”flowchart LR
O(("Observability"))
O --- B["1. Business<br/>value and cost"]
O --- ORG["2. Organization<br/>who decides, who owns"]
O --- MET["3. Business lines<br/>translated signals"]
O --- SUP["4. Support functions<br/>finance, compliance, procurement"]
O --- PER["5. People<br/>language and rituals"]
1. To the business, talking value and cost
Section titled “1. To the business, talking value and cost”- Express SLOs as promises made to customers. “99.5% of payments complete in under two seconds” means something to a product director, whereas “API p99 latency under 800 ms” tells them nothing.
- Compute a unit cost, per transaction, per customer or per feature. It is the only figure a financial controller can compare month to month.
- Quantify avoided loss from incident cost, exposure duration and blind spots. The exposure calculator gives an example for AI.
2. To the organization, saying who decides and who owns
Section titled “2. To the organization, saying who decides and who owns”- Who chooses what is collected, who pays for telemetry, who owns each alert? Until the answer is written down somewhere, nobody takes it on.
- Data contracts and a RACI turn good intentions into rules. See the organization dimension and part VII of the GenAI method.
3. To business lines, translating the signals
Section titled “3. To business lines, translating the signals”- Instrument business events (order confirmed, file processed, loan granted) as carefully as technical calls, and link them to traces. This is sometimes called business observability.
- Keep a translation dictionary between technical and business indicators (see the example below).
- Build dashboards per audience. An operations dashboard is not built for an executive committee: in my view, it loses more attention than it gains.
4. To support functions (finance, compliance, procurement)
Section titled “4. To support functions (finance, compliance, procurement)”- Finance: a telemetry budget tracked like any other, with its drifts and levers. See the business dimension.
- Compliance: evidence is a by-product of good instrumentation, provided it is designed in from the start. See log integrity and the evidence matrix.
- Procurement: once the contract is signed, it is too late to secure reversibility. It has to be settled during negotiation.
5. To people, with a shared language and shared rituals
Section titled “5. To people, with a shared language and shared rituals”- A monthly review bringing engineering, product and finance together around three figures: service quality as customers see it, unit cost, and incidents with their causes.
- Post-mortems open to the business lines concerned. They find out how the system works, and the technical team learns what the incident actually cost.
A translation example
Section titled “A translation example”| Technical signal | Business indicator | Decision informed |
|---|---|---|
| payment service p95 latency | cart abandonment at the payment step | prioritize optimization or a second payment provider |
| pricing engine error rate | quotes not issued per day, deferred revenue | choose between an immediate fix and degraded mode |
| tokens consumed per AI feature | cost per assisted customer conversation | adjust the model, the cache or the offer’s price |
| faithfulness score of a RAG assistant | complaints caused by wrong information | trigger a knowledge base update |
| night alerts per on-call engineer | team turnover and sick leave | revisit thresholds, rotations, staffing |
Each row connects three people who do not always talk to each other: whoever sees the signal, whoever owns the indicator and whoever decides.
For the first row, the chain reads as follows. It closes when the decision changes the signal.
flowchart LR S["Technical signal<br/>payment p95 latency"] -->|"translation"| I["Business indicator<br/>cart abandonment"] I -->|"trade-off"| D["Decision<br/>optimize or second provider"] D -.->|"measured effect"| S
Signs a bridge is missing
Section titled “Signs a bridge is missing”- The observability project is owned by IT alone, with no business sponsor.
- Dashboards have never been opened by anyone outside the technical team.
- No SLO is phrased in a customer’s words.
- Telemetry cost is discovered on the invoice.
- Compliance was not consulted on log retention.
- Post-mortems never leave the operations team.
On this site
Section titled “On this site”Every article and course ends with a Bridges box stating what the topic changes for the business, the organization, teams and business lines. I keep to this on every page, as the manifesto explains.