Organization
Team models for observability
For: architects · team managers · executives and CIOsPrerequisites: Basic notions of observability and of how IT teams are organized.
In my view, an observability platform rarely fails for purely technical reasons. It lasts when three questions have a written answer: who builds the platform, who owns what goes into it and who pays. The team model is how an organization answers all three at once.
This article compares three models, proposes an ownership boundary and a RACI matrix, then covers funding and the signals that say it is time to change models.
A shared vocabulary: Team Topologies
Section titled “A shared vocabulary: Team Topologies”The book Team Topologies by Matthew Skelton and Manuel Pais (IT Revolution, 2019) provides a useful vocabulary for this topic. It describes four team types and three interaction modes:
| Team type | Role in observability |
|---|---|
| Stream-aligned team | a product team that delivers and runs a service |
| Platform team | provides collection, storage, standard dashboards and alerting as self-service |
| Enabling team | temporarily helps product teams build skills, for example on OpenTelemetry instrumentation |
| Complicated-subsystem team | rarely needed for observability, except for a very specialized component |
The interaction modes are collaboration, X-as-a-Service and facilitating. The mapping to observability is my reading of the book, not a definition by its authors; I recommend it because it forces you to say how teams work together, not only who does what.
Three models
Section titled “Three models”The platform team. A dedicated team builds and runs the platform as a product: collection pipelines, storage, visualization, alerting, instrumentation templates. Product teams consume it as self-service. The CNCF white paper on platforms (CNCF Platforms White Paper) stresses this product approach: internal users, gathered requirements, a carefully designed experience.
The center of excellence. A small cross-functional group defines the standards (naming conventions, semantic conventions, alerting rules, post-mortem templates), trains and advises. It does not necessarily run the platform, which may be a purchased service or operated by the operations team. Its authority rests on expertise and leadership backing.
The federated model. A central core runs the shared foundation and each product team names an observability champion, who spends an explicit share of their time on their team’s instrumentation, SLOs and dashboards. Champions form a community led by the core team.
| Platform team | Center of excellence | Federated | |
|---|---|---|---|
| What is centralized | the tool and its operation | standards and expertise | the foundation and the community |
| Main strength | consistency, economies of scale, self-service | light start, spreads good practices | close to the business, peer-driven adoption |
| Main risk | becoming a bottleneck or a ticket desk | producing standards nobody applies | heterogeneity, champions with no real time |
| When to choose it, in my view | dozens of teams, need for cost control | early stage, purchased tool, few teams | already decentralized organization, autonomous product teams |
| Success condition | a platform product owner and usage indicators | a leadership mandate and standards checked automatically | champion time written into objectives |
These models are not mutually exclusive. A common path is to start as a center of excellence, set up a platform team when volume justifies it, then federate through champions. I recommend naming the current model and the target model, even if they are hybrids: an organization that does not know which model it is in does not know who decides.
The ownership boundary
Section titled “The ownership boundary”Whatever the model, I recommend writing down a simple boundary between what the platform owns and what product teams own. The “you build it, you run it” principle, stated by Werner Vogels in a 2006 ACM Queue interview, applies to observability too: the team that delivers a service owns its signals.
| Owned by the platform | Owned by the product team |
|---|---|
| Collectors, pipelines, storage, default retention | instrumentation of its code and dependencies |
| Availability and performance of the platform itself | SLIs and SLOs of its services, with the business |
| Instrumentation templates, libraries, standard dashboards | dashboards specific to its service |
| Generic alerting rules (platform, shared infrastructure) | alerts on its symptoms and the related runbooks |
| Data contracts and default cardinality budgets | compliance with its data contract and requests for exceptions |
| Measuring and showing costs per team | cost trade-offs on its own signals |
The platform has SLOs of its own: ingestion availability, delay before a signal shows up, query latency. An observability platform that does not observe itself loses the teams’ trust at the first data gap.
An operational RACI matrix
Section titled “An operational RACI matrix”Lesson 3 of the CIO path proposes a leadership-level RACI. This one goes one level down, to day-to-day activities, for a platform team model with champions (illustrative example, to be adapted in a workshop).
| Activity | Platform team | Product team | Observability champion | Security | Business |
|---|---|---|---|---|---|
| Run the platform and its SLOs | A, R | I | I | C | I |
| Instrument a new service | C | A | R | I | I |
| Define the SLO of a customer journey | C | R | C | I | A |
| Create a paging alert | C | A | R | I | I |
| Add a high-cardinality attribute | A | R | C | I | I |
| Collect potentially personal data | C | R | C | A | I |
| Decide on a telemetry budget overrun | R | C | C | I | A |
| Evolve instrumentation standards | A, R | C | C | C | I |
The two rules that matter: a single A per row, and a review at least once a year. The interactive RACI matrix checks the first one and shows each role’s load.
Funding
Section titled “Funding”Three funding modes coexist, matching the FinOps progression described in lesson 4 of the CIO path:
| Mode | Principle | What it requires |
|---|---|---|
| Central budget | IT carries all the spending | nothing more than a budget; but no team has an incentive to cut its signals |
| Showback | each team sees what it consumes | team labels enforced at ingestion and a monthly review |
| Chargeback | each team pays for its consumption | reliable, accepted measurement and an internal budget mechanism |
I recommend starting with showback as soon as the platform is shared, even if chargeback is never introduced. Seeing one’s own consumption is often enough to change behavior. I also recommend funding the foundation (operations, standards, training) centrally: charging it back pushes teams to work around it.
Signals that it is time to change models
Section titled “Signals that it is time to change models”| Observed signal | What it suggests |
|---|---|
| The time to get a first dashboard on a new service keeps growing | the platform has become a ticket desk; invest in self-service |
| Several teams have built their own stack | needs are not covered or not heard; gather requirements as for a product |
| Standards exist but code reviews do not check them | the center of excellence has no leverage; automate the check in continuous integration |
| Champions no longer have time for the role | the federated model is not funded; write their time into objectives |
| Nobody can answer “who owns this signal?” | the ownership boundary is not written down |
The indicators for running a platform as a product (time to first dashboard, developer satisfaction, self-service rate, cost per instrumented service) are detailed in lesson 5 of the CIO path.
Checklist
Section titled “Checklist”- The current model and the target model are named
- The ownership boundary between platform and product teams is written down
- The platform has its own SLOs and publishes them
- The operational RACI was built in a workshop and has a single A per row
- Each team sees what it consumes
- The shared foundation is funded centrally
- Roles (platform engineer, product owner, champion) have a job description
- The signals for changing models are reviewed at least once a year
Going further
Section titled “Going further”- What the platform allows to be collected is decided in telemetry governance.
- The maturity self-assessment highlights gaps between dimensions, such as an advanced stack nobody owns.
- The clickable architecture offers a “who owns what” layer.
Sources
Section titled “Sources”- Matthew Skelton and Manuel Pais, Team Topologies: Organizing Business and Technology Teams for Fast Flow, IT Revolution, 2019; overview of the concepts on teamtopologies.com.
- CNCF TAG App Delivery, Platforms White Paper.
- Jim Gray, A Conversation with Werner Vogels, ACM Queue, 2006.