Skip to content

OrganizationPractitioner

Team models for observability

For: architects · team managers · executives and CIOsPrerequisites: Basic notions of observability and of how IT teams are organized.

Reading mode

In my view, an observability platform rarely fails for purely technical reasons. It lasts when three questions have a written answer: who builds the platform, who owns what goes into it and who pays. The team model is how an organization answers all three at once.

This article compares three models, proposes an ownership boundary and a RACI matrix, then covers funding and the signals that say it is time to change models.

The book Team Topologies by Matthew Skelton and Manuel Pais (IT Revolution, 2019) provides a useful vocabulary for this topic. It describes four team types and three interaction modes:

Team typeRole in observability
Stream-aligned teama product team that delivers and runs a service
Platform teamprovides collection, storage, standard dashboards and alerting as self-service
Enabling teamtemporarily helps product teams build skills, for example on OpenTelemetry instrumentation
Complicated-subsystem teamrarely needed for observability, except for a very specialized component

The interaction modes are collaboration, X-as-a-Service and facilitating. The mapping to observability is my reading of the book, not a definition by its authors; I recommend it because it forces you to say how teams work together, not only who does what.

The platform team. A dedicated team builds and runs the platform as a product: collection pipelines, storage, visualization, alerting, instrumentation templates. Product teams consume it as self-service. The CNCF white paper on platforms (CNCF Platforms White Paper) stresses this product approach: internal users, gathered requirements, a carefully designed experience.

The center of excellence. A small cross-functional group defines the standards (naming conventions, semantic conventions, alerting rules, post-mortem templates), trains and advises. It does not necessarily run the platform, which may be a purchased service or operated by the operations team. Its authority rests on expertise and leadership backing.

The federated model. A central core runs the shared foundation and each product team names an observability champion, who spends an explicit share of their time on their team’s instrumentation, SLOs and dashboards. Champions form a community led by the core team.

Platform teamCenter of excellenceFederated
What is centralizedthe tool and its operationstandards and expertisethe foundation and the community
Main strengthconsistency, economies of scale, self-servicelight start, spreads good practicesclose to the business, peer-driven adoption
Main riskbecoming a bottleneck or a ticket deskproducing standards nobody appliesheterogeneity, champions with no real time
When to choose it, in my viewdozens of teams, need for cost controlearly stage, purchased tool, few teamsalready decentralized organization, autonomous product teams
Success conditiona platform product owner and usage indicatorsa leadership mandate and standards checked automaticallychampion time written into objectives

These models are not mutually exclusive. A common path is to start as a center of excellence, set up a platform team when volume justifies it, then federate through champions. I recommend naming the current model and the target model, even if they are hybrids: an organization that does not know which model it is in does not know who decides.

Whatever the model, I recommend writing down a simple boundary between what the platform owns and what product teams own. The “you build it, you run it” principle, stated by Werner Vogels in a 2006 ACM Queue interview, applies to observability too: the team that delivers a service owns its signals.

Owned by the platformOwned by the product team
Collectors, pipelines, storage, default retentioninstrumentation of its code and dependencies
Availability and performance of the platform itselfSLIs and SLOs of its services, with the business
Instrumentation templates, libraries, standard dashboardsdashboards specific to its service
Generic alerting rules (platform, shared infrastructure)alerts on its symptoms and the related runbooks
Data contracts and default cardinality budgetscompliance with its data contract and requests for exceptions
Measuring and showing costs per teamcost trade-offs on its own signals

The platform has SLOs of its own: ingestion availability, delay before a signal shows up, query latency. An observability platform that does not observe itself loses the teams’ trust at the first data gap.

Lesson 3 of the CIO path proposes a leadership-level RACI. This one goes one level down, to day-to-day activities, for a platform team model with champions (illustrative example, to be adapted in a workshop).

ActivityPlatform teamProduct teamObservability championSecurityBusiness
Run the platform and its SLOsA, RIICI
Instrument a new serviceCARII
Define the SLO of a customer journeyCRCIA
Create a paging alertCARII
Add a high-cardinality attributeARCII
Collect potentially personal dataCRCAI
Decide on a telemetry budget overrunRCCIA
Evolve instrumentation standardsA, RCCCI

The two rules that matter: a single A per row, and a review at least once a year. The interactive RACI matrix checks the first one and shows each role’s load.

Three funding modes coexist, matching the FinOps progression described in lesson 4 of the CIO path:

ModePrincipleWhat it requires
Central budgetIT carries all the spendingnothing more than a budget; but no team has an incentive to cut its signals
Showbackeach team sees what it consumesteam labels enforced at ingestion and a monthly review
Chargebackeach team pays for its consumptionreliable, accepted measurement and an internal budget mechanism

I recommend starting with showback as soon as the platform is shared, even if chargeback is never introduced. Seeing one’s own consumption is often enough to change behavior. I also recommend funding the foundation (operations, standards, training) centrally: charging it back pushes teams to work around it.

Observed signalWhat it suggests
The time to get a first dashboard on a new service keeps growingthe platform has become a ticket desk; invest in self-service
Several teams have built their own stackneeds are not covered or not heard; gather requirements as for a product
Standards exist but code reviews do not check themthe center of excellence has no leverage; automate the check in continuous integration
Champions no longer have time for the rolethe federated model is not funded; write their time into objectives
Nobody can answer “who owns this signal?”the ownership boundary is not written down

The indicators for running a platform as a product (time to first dashboard, developer satisfaction, self-service rate, cost per instrumented service) are detailed in lesson 5 of the CIO path.

  • The current model and the target model are named
  • The ownership boundary between platform and product teams is written down
  • The platform has its own SLOs and publishes them
  • The operational RACI was built in a workshop and has a single A per row
  • Each team sees what it consumes
  • The shared foundation is funded centrally
  • Roles (platform engineer, product owner, champion) have a job description
  • The signals for changing models are reviewed at least once a year