Business
1. Building the business case
For: finance, risk and compliance · executives and CIOsPrerequisites: None.
Without figures, observability remains just another IT project, and it is the first budget to be cut when things get tight. This lesson gives a method to build a case that a finance department recognizes as serious.
The real cost of incidents
Section titled “The real cost of incidents”Every business case starts with the cost of incidents over the last twelve months. This figure is almost always missing from existing dashboards, and its absence is already an argument: you cannot steer what you do not measure.
Four indicators structure the calculation:
- MTTD (mean time to detect): time between the start of an anomaly and its detection;
- MTTA (mean time to acknowledge): time between the alert and someone taking ownership of it;
- MTTR (mean time to recovery): time between detection and recovery;
- financial impact: margin lost during downtime, contractual penalties, customer compensation.
Cost per minute: your figure, not a study’s
Section titled “Cost per minute: your figure, not a study’s”The most quoted figure is Gartner’s: $5,600 per minute of network downtime on average, or about $300,000 per hour. It comes from an estimate published in 2014 by Andrew Lerner, an analyst at Gartner. It is an average, in dollars, old, and it mixes companies of all sizes. It can be used to show that an order of magnitude exists, never to put a figure on your own case.
Your figure is calculated with management control, as lost margin, not as IT cost:
hourly downtime cost = lost hourly margin + contractual penalties per hour + catch-up costs (overtime, scrap, follow-ups)annual direct cost = sum of downtime hours x hourly cost of the system concernedAdd indirect costs separately (loss of customer trust, team hours mobilized), presenting them as an estimate along with their assumption. Mixing a validated direct cost with an assumed indirect cost weakens the whole figure.
Calculating the return on investment
Section titled “Calculating the return on investment”ROI (%) = (gains - costs) / costs x 100payback period (months) = initial investment / monthly gainGains are built from four sources, each to be quantified separately:
- lower incident costs: fewer hours of downtime, detected earlier, recovered faster;
- tooling rationalization: redundant licenses removed;
- team time recovered: less alert noise, shorter diagnostics;
- savings on the telemetry itself: see lesson 4.
Costs include licenses or infrastructure, storage, people (running the platform, instrumentation), training, change management and migration.
An executive committee is often more sensitive to the payback period than to ROI: “how many months until we get our money back?” speaks louder than a five-year percentage. To estimate it on your own assumptions, the MTTR business case calculates the return, the payback period and the sensitivity to assumptions.
The six-step method
Section titled “The six-step method”Order matters. Skipping a step or doing them out of order produces cases that do not convince.
flowchart LR E1["1. Cost of<br/>incidents"] --> E2["2. Expected<br/>gains"] --> E3["3. Full<br/>cost"] --> E4["4. ROI and<br/>payback"] --> E5["5. Three years,<br/>three scenarios"] --> E6["6. Pitch to<br/>committee"]
- Put a figure on the cost of incidents over the last twelve months: number, duration, hourly cost per system, penalties.
- Identify the expected gains, source by source. Make a downtime reduction assumption that you justify (for example the incidents you would have detected earlier with an alert on the user symptom).
- Total the full cost over three years, not just the licenses.
- Calculate the ROI and the payback period, and the net present value if your finance department uses it.
- Project over three years with three scenarios (conservative, central, favorable) and at least two architectures compared with the same criteria (see lesson 3).
- Pitch to the committee: five slides, ten minutes, one decision requested (see lesson 7).
The full cost, item by item
Section titled “The full cost, item by item”A business case that only shows licenses is disqualified at the first serious budget review. Six items are regularly underestimated:
| Item | What to count | Main lever |
|---|---|---|
| Metrics storage | volume of active series, retention, replication | control of cardinality |
| Log storage | volume ingested and indexed, retention by category | hot, warm and cold tiers |
| Traces | spans ingested, retention | tail sampling |
| People | running the platform, on-call, instrumentation | automation, versioned configuration |
| SaaS subscriptions | billed units (hosts, volume, users, modules) | negotiation, caps, usage review |
| Change management | training, support, documentation | internal champions, written procedures |
I do not give a price per gigabyte here, nor a total cost range by company size. They vary too much with architecture and contract to be useful out of context. The Collector cost simulator lets you put a figure on your own volume, with unit prices you enter yourself.
Six savings levers to plan from the start
Section titled “Six savings levers to plan from the start”These levers are decided at design time. Pulled early, they weigh on all three years of the business case.
- Smart trace sampling: keep all traces with errors or high latency, and only a fraction of nominal traffic.
- Differentiated retention: not all data has the same value over time; a debug log does not have the lifespan of an authentication log.
- Controlled cardinality: no unique identifier (user, request) as a metric label.
- An observability cost dashboard: cost per gigabyte ingested, per service, per team, with an alert on drifts.
- Contract negotiation: commitment, caps, exit clause, always with comparable offers in hand.
- Architecture chosen on criteria: self-managed open source, SaaS or hybrid, compared with the same grid (see the solution evaluation grid).
The effect of each lever depends on your volume and your contract. Rather than a generic percentage, measure it: the Collector cost simulator shows the bill before and after each lever.
The business case template
Section titled “The business case template”A template to fill in, where the “method” column matters as much as the value: that is what a CFO asks about.
Cost of incidents (last twelve months)
| Item | Calculation method | Value |
|---|---|---|
| Major incidents | priority 1 and 2 tickets | |
| Downtime duration | sum of downtime durations, per system | hours |
| Hourly cost of downtime | lost hourly margin, validated by management control | €/h |
| Direct cost | duration x hourly cost, per system | € |
| Estimated indirect cost | explicit assumption, validated or not | € |
| Regulatory exposure | separate line, not added to gains | qualitative |
Projected gains over three years
| Year | Downtime reduction assumption (to be justified) | Annual gain | Cumulative |
|---|---|---|---|
| 1 | |||
| 2 | |||
| 3 |
Summary
| Indicator | Formula | Value |
|---|---|---|
| Three-year ROI | (gains - costs) / costs x 100 | % |
| Payback period | initial investment / monthly gain | months |
| Gains to costs ratio | total gains / total costs | to 1 |
Workshop: your business case
Section titled “Workshop: your business case”Duration: 45 minutes, alone or in pairs.
- Put a figure on the incidents of the last twelve months: number, total duration, cumulative cost.
- Obtain or estimate the hourly cost of downtime for your most critical system, with its assumption.
- Formulate a downtime reduction assumption and justify it incident by incident.
- Calculate the annual gains, the three-year ROI and the payback period in the MTTR business case.
- Write down your two biggest uncertainties: they will be the committee’s first questions.
My next action: what action, by what date, with whom?
Going further
Section titled “Going further”- The MTTR business case, to test your assumptions.
- For an AI service: the quantified impact in the GenAI method and its exposure calculator.
- The business dimension.