People
5. People and culture
For: team managers · executives and CIOsPrerequisites: None.
The platform rolls out in a few months; the culture that brings it to life takes years. That gap explains many failures: the dashboards are ready, the SLOs documented, but in a crisis the teams fall back on their old reflexes. This lesson focuses on what is specific to observability. Two guiding choices: the transformation moves forward in small cohorts, and no tool replaces the direct conversation between manager and employee.
From gatekeeper to doctor
Section titled “From gatekeeper to doctor”For a long time, operations worked on a gatekeeper model: watch health indicators, react to alerts, escalate. Mastery came from intimate knowledge of the infrastructure. That model runs into the cloud (you no longer see inside the machine, you see an API), distributed architectures (you cannot know hundreds of components intimately) and shorter delivery cycles.
The new job looks like a doctor’s: equipped with correlated traces, metrics and logs, they reason through hypotheses and elimination. They know how to read the system rather than know it by heart, and they document so the team learns. The skills that become central: reading a dashboard, querying structured logs, understanding dependencies, reasoning in percentiles and SLOs.
Resistance to this change is not a lack of will, it is a signal of meaning: a professional identity built on tacit knowledge is being called into question. Three phrases to avoid: “you need to adapt,” “it’s not that different,” “this is the future.” Three phrases that open the door: “I understand this transition is uncomfortable,” “your knowledge is still valuable, here is how we will recognize it,” “what are your concrete concerns?”
Four archetypes to support
Section titled “Four archetypes to support”| Archetype | What motivates them | Possible path | Main risk | Levers |
|---|---|---|---|---|
| Senior system administrator | recognition of their expertise | senior SRE, platform engineer | leaving and taking unwritten knowledge with them | pairing with a younger SRE, public recognition, OpenTelemetry training, a role in the SLOs of legacy systems |
| Monitoring lead | operational usefulness, less noise | production SRE, incident commander | burnout before the transformation | healthy on-call, training in reliability engineering, recognition of noise-reduction work |
| Infrastructure project manager | structure and control | platform product manager, observability project manager | ending up without a defined role | an explicit transformation mission, objectives, product management training |
| System architect | consistency and durability of choices | platform architect, staff engineer | slowing things down out of skepticism | involve them early in decisions, make them the technical arbiter, give them a public voice |
The blameless post-mortem
Section titled “The blameless post-mortem”In a healthy organization, an incident is an expensive gift: it reveals a blind spot in the system. In an unhealthy one, it is a fault to pin on someone, and teams end up hiding incidents. The blameless post-mortem is the ritual that moves you from one to the other.
Four rules:
- Look for what made the error possible, not for a culprit. A person who made a mistake in production showed that the mistake was possible.
- The account is factual: a timestamped timeline, facts kept separate from interpretations, no “they should have.”
- Actions are few and assigned: a handful of actions with an owner, a date and a success criterion are worth more than fifty with no owner.
- The document circulates internally, so other teams benefit from it.
The CIO’s role is decisive. They must state publicly that they will not sanction a good-faith error made during an incident. Then they must prove it at the first serious incident, including by defending the operator in front of the executive committee.
Standard structure: summary, timeline, root cause analysis, aggravating factors, corrective actions, lessons learned, and what went well.
Sustainable on-call
Section titled “Sustainable on-call”An exhausted on-call team makes worse decisions in a crisis, restores service more slowly, and eventually leaves. This debt can be measured just as well as technical debt.
Benchmarks. Google’s book Site Reliability Engineering, in its chapter on on-call, gives three orders of magnitude that are often cited. At most 25% of an engineer’s time spent on call. At most two incidents per twelve-hour shift. A rotation of at least eight people for a single-site team (six per site across two sites). These are benchmarks, not standards; what matters is choosing your own and measuring them.
Rules.
- Compensate. In France, the Labour Code requires on-call periods to be compensated, either financially or with time off (Article L3121-9). Beyond the law, on-call that goes unrecognized drives the best people away.
- Limit the noise. An alert at night should be exceptional and require immediate human action. The alert fatigue simulator estimates the real load of an on-call week.
- Respect the right to disconnect. Someone who is not on call does not get called; exceptions are written down and compensated.
- Track perceived load. A monthly subjective indicator, shared within the team, spots drift before burnout sets in.
- Decompress. After a major incident, the team involved takes time to recover, without having to justify it.
The CIO is permanently on crisis call. They need a trusted deputy, real time off the grid and written escalation protocols.
The platform as a product
Section titled “The platform as a product”Platform engineering treats the observability infrastructure as a product whose users are the developers. The central team stops answering tickets; it builds self-service tools. An internal developer portal (Backstage, created by Spotify and now a CNCF project, or commercial solutions) brings together service documentation, instrumentation templates, standard dashboards and procedures.
Four indicators to steer the platform, with targets to set for your own context:
| Indicator | What it measures |
|---|---|
| Time to first dashboard | time between creating a service and its first usable dashboard |
| Developer satisfaction | willingness to recommend the platform, measured each quarter |
| Self-service rate | share of requests resolved without the platform team stepping in |
| Cost per instrumented service | platform spend divided by the number of services covered |
Recruit, train, retain
Section titled “Recruit, train, retain”Recruit. The market for SRE and observability profiles remains tight. For salaries, rely on a recent compensation survey for your local job market rather than on a figure read in a training course: they change too fast. Three channels work. Employee referrals, with a clear job description and a short process. Internal reskilling of administrators trained in observability: cheaper than an external hire, it preserves knowledge of the information system. Work-study programs, provided there is real supervision.
Retain. Five factors carry weight. The first three: the quality of the tooling, a reliability culture where incidents do not lead to sanctions, a balanced on-call load. Next come written career paths (SRE, staff, principal, management, with criteria and pay). Last, recognition, including public recognition: letting people publish, speak at conferences, write.
Exercises
Section titled “Exercises”- Archetype mapping (two hours): list the ten to twenty key people in your operations, assign each one to an archetype, note risks and levers, and schedule a conversation with the two whose departure would cost the most.
- On-call audit (half a day): for each rotation, record the size, how often people are on duty, night alerts per week, the compensation, and whether there is recovery time. Compare with the benchmarks above.
- Post-mortem replay (two hours as a team): take a past incident, redo the post-mortem with the four rules, and compare with the original document. The differences measure your cultural maturity.
A two-year HR roadmap
| Period | Actions |
|---|---|
| Quarter 1 | archetype mapping, one-on-one interviews, updated job descriptions, first on-call audit |
| Quarter 2 | individual training plans for senior profiles, blameless post-mortem in two pilot teams |
| Quarter 3 | first targeted hire, first internal reskilling cycle, mentoring in both directions |
| Quarter 4 | review of departures, publication of career paths, first assessment of the blameless post-mortem |
| Year 2 | rollout to all teams, developer portal, expanded work-study program, consolidated pay scale |
CIO checklist
- I know by name the ten most critical people in my operations
- The archetypes and their paths are mapped
- The blameless post-mortem is in place in at least one team
- Every on-call rotation is sized, compensated and measured
- SRE and platform career paths are written, pay included
- Every senior profile has a training budget and plan
- Departures of SRE and platform profiles are tracked
- I spend time every month listening directly to the operations teams
Going further
Section titled “Going further”- The people dimension and the alert fatigue simulator.
- The error budget: multi-window alerts that wake people up for nothing less often.