Skip to content

BusinessPractitioner

1. Building the business case

For: finance, risk and compliance · executives and CIOsPrerequisites: None.

Reading mode

Without figures, observability remains just another IT project, and it is the first budget to be cut when things get tight. This lesson gives a method to build a case that a finance department recognizes as serious.

Every business case starts with the cost of incidents over the last twelve months. This figure is almost always missing from existing dashboards, and its absence is already an argument: you cannot steer what you do not measure.

Four indicators structure the calculation:

  • MTTD (mean time to detect): time between the start of an anomaly and its detection;
  • MTTA (mean time to acknowledge): time between the alert and someone taking ownership of it;
  • MTTR (mean time to recovery): time between detection and recovery;
  • financial impact: margin lost during downtime, contractual penalties, customer compensation.

Cost per minute: your figure, not a study’s

Section titled “Cost per minute: your figure, not a study’s”

The most quoted figure is Gartner’s: $5,600 per minute of network downtime on average, or about $300,000 per hour. It comes from an estimate published in 2014 by Andrew Lerner, an analyst at Gartner. It is an average, in dollars, old, and it mixes companies of all sizes. It can be used to show that an order of magnitude exists, never to put a figure on your own case.

Your figure is calculated with management control, as lost margin, not as IT cost:

hourly downtime cost = lost hourly margin
+ contractual penalties per hour
+ catch-up costs (overtime, scrap, follow-ups)
annual direct cost = sum of downtime hours x hourly cost of the system concerned

Add indirect costs separately (loss of customer trust, team hours mobilized), presenting them as an estimate along with their assumption. Mixing a validated direct cost with an assumed indirect cost weakens the whole figure.

ROI (%) = (gains - costs) / costs x 100
payback period (months) = initial investment / monthly gain

Gains are built from four sources, each to be quantified separately:

  1. lower incident costs: fewer hours of downtime, detected earlier, recovered faster;
  2. tooling rationalization: redundant licenses removed;
  3. team time recovered: less alert noise, shorter diagnostics;
  4. savings on the telemetry itself: see lesson 4.

Costs include licenses or infrastructure, storage, people (running the platform, instrumentation), training, change management and migration.

An executive committee is often more sensitive to the payback period than to ROI: “how many months until we get our money back?” speaks louder than a five-year percentage. To estimate it on your own assumptions, the MTTR business case calculates the return, the payback period and the sensitivity to assumptions.

Order matters. Skipping a step or doing them out of order produces cases that do not convince.

flowchart LR
  E1["1. Cost of<br/>incidents"] --> E2["2. Expected<br/>gains"] --> E3["3. Full<br/>cost"] --> E4["4. ROI and<br/>payback"] --> E5["5. Three years,<br/>three scenarios"] --> E6["6. Pitch to<br/>committee"]
  1. Put a figure on the cost of incidents over the last twelve months: number, duration, hourly cost per system, penalties.
  2. Identify the expected gains, source by source. Make a downtime reduction assumption that you justify (for example the incidents you would have detected earlier with an alert on the user symptom).
  3. Total the full cost over three years, not just the licenses.
  4. Calculate the ROI and the payback period, and the net present value if your finance department uses it.
  5. Project over three years with three scenarios (conservative, central, favorable) and at least two architectures compared with the same criteria (see lesson 3).
  6. Pitch to the committee: five slides, ten minutes, one decision requested (see lesson 7).

A business case that only shows licenses is disqualified at the first serious budget review. Six items are regularly underestimated:

ItemWhat to countMain lever
Metrics storagevolume of active series, retention, replicationcontrol of cardinality
Log storagevolume ingested and indexed, retention by categoryhot, warm and cold tiers
Tracesspans ingested, retentiontail sampling
Peoplerunning the platform, on-call, instrumentationautomation, versioned configuration
SaaS subscriptionsbilled units (hosts, volume, users, modules)negotiation, caps, usage review
Change managementtraining, support, documentationinternal champions, written procedures

I do not give a price per gigabyte here, nor a total cost range by company size. They vary too much with architecture and contract to be useful out of context. The Collector cost simulator lets you put a figure on your own volume, with unit prices you enter yourself.

These levers are decided at design time. Pulled early, they weigh on all three years of the business case.

  1. Smart trace sampling: keep all traces with errors or high latency, and only a fraction of nominal traffic.
  2. Differentiated retention: not all data has the same value over time; a debug log does not have the lifespan of an authentication log.
  3. Controlled cardinality: no unique identifier (user, request) as a metric label.
  4. An observability cost dashboard: cost per gigabyte ingested, per service, per team, with an alert on drifts.
  5. Contract negotiation: commitment, caps, exit clause, always with comparable offers in hand.
  6. Architecture chosen on criteria: self-managed open source, SaaS or hybrid, compared with the same grid (see the solution evaluation grid).

The effect of each lever depends on your volume and your contract. Rather than a generic percentage, measure it: the Collector cost simulator shows the bill before and after each lever.

A template to fill in, where the “method” column matters as much as the value: that is what a CFO asks about.

Cost of incidents (last twelve months)

ItemCalculation methodValue
Major incidentspriority 1 and 2 tickets
Downtime durationsum of downtime durations, per systemhours
Hourly cost of downtimelost hourly margin, validated by management control€/h
Direct costduration x hourly cost, per system€
Estimated indirect costexplicit assumption, validated or not€
Regulatory exposureseparate line, not added to gainsqualitative

Projected gains over three years

YearDowntime reduction assumption (to be justified)Annual gainCumulative
1
2
3

Summary

IndicatorFormulaValue
Three-year ROI(gains - costs) / costs x 100%
Payback periodinitial investment / monthly gainmonths
Gains to costs ratiototal gains / total coststo 1

Duration: 45 minutes, alone or in pairs.

  1. Put a figure on the incidents of the last twelve months: number, total duration, cumulative cost.
  2. Obtain or estimate the hourly cost of downtime for your most critical system, with its assumption.
  3. Formulate a downtime reduction assumption and justify it incident by incident.
  4. Calculate the annual gains, the three-year ROI and the payback period in the MTTR business case.
  5. Write down your two biggest uncertainties: they will be the committee’s first questions.

My next action: what action, by what date, with whom?