Organization
3. Governance and roadmap
For: architects · team managers · executives and CIOsPrerequisites: Have read lesson 1 of the course.
An observability strategy without governance quickly turns into a collection of isolated projects that drift. This lesson gives three tools: a phased roadmap, a RACI and a selection method.
The four-phase roadmap
Section titled “The four-phase roadmap”Each phase builds on the previous one. The durations are indicative and should be calibrated to your scope; the order, however, must be respected.
flowchart LR P1["Phase 1<br/>Foundations<br/>0 to 3 months"] --> P2["Phase 2<br/>Structuring<br/>3 to 6 months"] --> P3["Phase 3<br/>Optimization<br/>6 to 12 months"] --> P4["Phase 4<br/>Maturity<br/>12 months and beyond"]
Phase 1, foundations. Inventory of tools, sources and teams. Choice of the target architecture, based on the business case from lesson 1. Service level indicators (SLIs) for critical services. Two or three quick wins achievable in under six weeks, to build trust.
Phase 2, structuring. Instrumentation of priority services with OpenTelemetry. Formalized service level objectives (SLOs). Alerts based on SLOs and burn rate. Signed RACI and on-call organization.
Phase 3, optimization. Distributed traces on critical journeys. Correlation between metrics, logs and traces. Systematic blameless post-mortems on major incidents. Steering the cost of telemetry (see lesson 4).
Phase 4, maturity. Configurations, dashboards, alerts and SLOs managed as code, under version control. Targeted automation of responses to recurring incidents. AIOps on specific use cases, if the data allows it (see lesson 6). Observability of AI systems if there are any (see the GenAI method).
The observability RACI
Section titled “The observability RACI”| Activity | CIO | CISO | Platform, SRE | Development | Business |
|---|---|---|---|---|---|
| Define the observability strategy | A | C | R | I | C |
| Decide the annual budget | A | I | R | I | C |
| Instrument services | C | I | R | A | I |
| Define a service’s SLOs | C | I | C | R | A |
| Lead the response to a major incident | I | C | A | R | I |
| Report an incident to the authority | C | A | R | I | I |
| Produce management reporting | A | C | R | I | C |
| Steer the roadmap | A | C | R | C | I |
R: responsible, does the work. A: accountable, approves, only one per row. C: consulted before the decision. I: informed after the decision.
Two rules are enough to avoid most drift: only one A per row, otherwise nobody takes ownership, and a quarterly review, because the organization changes. The interactive RACI matrix checks these rules and shows the load on each role.
Self-managed, SaaS or hybrid: decide with the same criteria
Section titled “Self-managed, SaaS or hybrid: decide with the same criteria”No option is better in absolute terms. What separates them is your volume, your skills, your sovereignty constraints and the way you instrument.
| Criterion | Self-managed open source | SaaS | Hybrid |
|---|---|---|---|
| Upfront cost | infrastructure and team time | subscription, little setup effort | both, on separate scopes |
| Cost over time | mostly driven by people and storage | mostly driven by billed volume | to be tracked on both lines |
| Skills required | running the platform, on-call | administration, usage governance | both, on a smaller scale |
| Time to first results | longer | often shorter | in between |
| Sovereignty | under control if the hosting is | depends on hosting, applicable law and subcontractors | to be qualified per data flow |
| Dependency | on internal skills | on the vendor and its pricing | shared |
| Exit cost | low if instrumentation is standard; dashboards and alerts to be rebuilt | low if instrumentation uses OpenTelemetry; high with proprietary agents and query language | same, per scope |
The factor that weighs most on reversibility is not the commercial model: it is instrumentation. Collection with OpenTelemetry lets you change backend without re-instrumenting; proprietary agents, whether they come from a vendor or an open source project, tie the organization to that backend.
The solution evaluation grid: six criteria
Section titled “The solution evaluation grid: six criteria”To compare two or three options on a common basis, score each one from 1 (weak) to 4 (excellent) on six criteria, weighted according to your context. The interactive grid does the calculation and flags disqualifying scores.
| Criterion | Question to ask | Warning sign |
|---|---|---|
| Portability | does the solution receive and export data in OpenTelemetry, in open formats? | exclusive format, limited export |
| Exit cost | how much does a migration cost: history, dashboards, alerts, integrations? Is it written into the contract? | no exit clause |
| Data sovereignty | where is the data stored, under which law, with which subcontractors? | opaque location or subcontractors |
| Cost predictability | do the 12 and 36 month simulations hold on your realistic volumes? | usage-based billing with no cap |
| Support and service levels | guaranteed response times, language, dedicated contact, active community for open source? | no written commitment |
| Ecosystem | integrations with the CI/CD chain, ITSM, security; open and documented API? | integrations limited to the same vendor’s products |
eBPF: instrumenting without touching the code
Section titled “eBPF: instrumenting without touching the code”eBPF is a Linux kernel technology that runs verified programs on hook points (system calls, network, functions). In observability, it makes it possible to capture HTTP, gRPC, TCP or DNS traffic without modifying applications. Well-known projects include Cilium and Tetragon for networking and security, Pixie for Kubernetes, and Beyla, whose code Grafana Labs donated to OpenTelemetry in 2025.
For the CIO, eBPF reduces the cost of instrumentation and the time before the first data arrives, especially on applications that cannot be modified. It requires a recent Linux kernel and specific skills, to be checked tool by tool before a proof of concept.
The platform as a product, and frugality
Section titled “The platform as a product, and frugality”In an internal developer platform, observability becomes a self-service offering: teams instrument their own services with ready-to-use templates, without waiting for the operations team. Two conditions: shared OpenTelemetry standards and configurations versioned as code, which also brings auditability and reversibility.
Observability also serves frugality: it spots oversized services, inefficient queries, redundant processing. It has a footprint of its own, hence a simple rule: every metric collected must have an identified use (an alert, a dashboard, a decision). If your company publishes a sustainability report, infrastructure metrics can feed its energy section. The scope and timeline of the CSRD were revised by the EU in 2025 and 2026 and must be checked for your case.
Six CIO anti-patterns
Section titled “Six CIO anti-patterns”- The empty dashboard: tools without SLOs or business indicators. Remedy: start from the five indicators management expects, then link them to technical measurements.
- The invisible project: no business case, so a cost with no visible return, cut at the first budget review. Remedy: lesson 1.
- The big bang: instrumenting everything at once. Remedy: the phased roadmap and quick wins.
- The ghost RACI: everyone thinks it is someone else’s job. Remedy: a RACI signed in phase 2 at the latest, reviewed every quarter.
- Firefighting compliance: discovering your obligations at the first inspection. Remedy: build evidence capability into the roadmap (see lesson 2).
- The tool before the strategy: choosing a solution before defining the needs. Remedy: strategy, needs, shortlist, proof of concept, decision, in that order.
Workshop: your roadmap
Section titled “Workshop: your roadmap”Duration: 45 minutes, in pairs or in a small group.
- Place your organization on the grid below, then on the site’s maturity self-assessment.
- Identify two or three quick wins achievable in under six weeks.
- Build phases 1 and 2: objectives, main deliverable, teams, budget, dependencies, quantified success indicator.
- Fill in the RACI matrix with your actual roles, and spot the rows that spark debate.
- Compare two options in the evaluation grid.
| Dimension | Level 1 | Level 2 | Level 3 | Level 4 |
|---|---|---|---|---|
| Vision | none | implicit | documented and shared | aligned with group strategy |
| Budget | scattered | identified | structured | multi-year |
| Governance | no RACI | informal roles | signed RACI | RACI reviewed every quarter |
| Compliance | nothing | firefighting mode | built into the roadmap | evidence produced continuously |
| Communication | no reporting | occasional | monthly dashboard | regular review by the executive committee |
| Tooling | disconnected tools | partially integrated | unified OpenTelemetry collection | configuration managed as code |
The lowest dimensions point to the priorities. An organization at level 3 on tooling and level 1 on governance has invested in technology without structuring responsibilities: the RACI comes first.
My next action: what action, by what date, with whom?
Going further
Section titled “Going further”- The organization dimension and the clickable architecture, with its “who owns what” layer.
- The Collector cost simulator.