Skip to content

OrganizationPractitioner

6. Strategic steering

For: team managers · executives and CIOsPrerequisites: Have read lesson 1 of the course.

Reading mode

The CIO needs different indicators from those of the operations team. Knowing how to translate the technical into business impact is what separates projects that survive a change of leadership from those that get cut at the first budget trade-off.

Executive indicators and technical indicators

Section titled “Executive indicators and technical indicators”

For the executive committee

  • Service availability, as the customer sees it: the share of time the service is fully usable from the user’s point of view, not the infrastructure’s.
  • Cost per incident: mean and median, including lost margin, remediation and penalties.
  • Resolution time as the customer sees it: from the first complaint to actual restoration.
  • Meeting contractual commitments: share of service levels met over rolling 30 and 90 days.
  • Cost of observability: relative to the IT budget or, better, to business activity (see lesson 4).

For the operations teams

  • SLOs and burn rate: how fast the error budget is being consumed, which triggers the alert before the breach (see the error budget simulator);
  • MTTD and MTTR by severity level;
  • error rate by service and by route;
  • latency in percentiles (p95, p99);
  • instrumentation coverage of critical services.

Every executive indicator must be tied to at least one technical measure. That mapping is what lets the committee ask precise questions when a figure drifts.

Five principles:

  1. Five indicators at most on the home page. Everything else is available in detail, not displayed up front.
  2. Three colors. Green, orange, red, with written thresholds. Ten-level gradients survive neither projection nor a change of tool.
  3. Always a comparison: against the target, against the previous period. A figure on its own says nothing.
  4. A path to the detail: each indicator opens onto the technical measures that make it up.
  5. Continuous updates, a monthly review with commentary by the IT department. Without a review, the dashboard turns into wallpaper.

The interactive executive dashboard applies these five principles to Helinord’s fictional data: click an indicator to drill down to its technical measures.

AIOps, the application of AI techniques to operations, is one of the most marketing-wrapped topics around.

What it actually does, on good data: reduce alert noise through anomaly detection, group alerts that belong to the same incident, spot recurring patterns, anticipate some short-term saturation.

What it does not do: replace defining SLOs (without a target, it has nothing to optimize), make up for poor data, fix an architecture, or decide alone on a major incident.

Prerequisites: several months of clean history, consistent tags across services and environments, controlled cardinality, structured post-mortems that make it possible to label past incidents.

Five questions to ask a vendor: which models, and how explainable are they? What OpenTelemetry compatibility? What cost, capped in the contract? What results on your data, in a proof of concept, rather than in a demo? What exit plan?

The ADKAR model, published by Prosci, describes five stages a person goes through to adopt a change. Applied to observability:

StageWhat the person experiencesLevers
Awarenessthey understand why change is neededcost of incidents quantified, comparable cases
Desirethey want to take partinternal champions, visible quick wins, recognition
Knowledgethey know how to do ittraining, written procedures, pairing, community of practice
Abilitythey can do it day to dayplatform in service, SLOs in production, self-service, up-to-date documentation
Reinforcementthe change lastsblameless post-mortems, indicators presented to the committee, adjusted roadmap

The thresholds below are benchmarks I suggest for triggering a discussion, to be adjusted to your context.

SignalWarning thresholdWhat it means
Departures from the teamseveral departures in a year from a small teamexpertise and context leave with the people
Alert noisethe vast majority of alerts require no actionthe real signals are drowned out
Trace coveragea minority of critical services instrumenteddiagnosis happens blind
Post-mortemsnone in six monthsthe same mistakes keep repeating
Alert debthundreds of rules never reviewedobsolete rules mask the real anomalies

When several of these signals appear at the same time, the project deserves a reset led directly by the CIO.

Duration: 45 minutes, on your own.

  1. Choose five indicators for your executive committee.
  2. Tie each one to at least one existing technical measure, or note that it is missing.
  3. Set a target for each, and green, orange and red thresholds.
  4. Write the commentary you would give at the next monthly review.
  5. If an indicator concerns a business unit, plan a workshop with them before locking it in.

My next action: what action, by what date, with whom?