Business
4. Observability FinOps
For: architects · finance, risk and compliance · executives and CIOsPrerequisites: Basic notions of metrics, logs and traces.
Observability was long sold as insurance. With billing based on ingested volume, it can become a line item that doubles without anyone deciding it should. This lesson gives the CIO a framework to take back control: understand the bill, pull the levers, set up governance.
Why the bill drifts
Section titled “Why the bill drifts”A model that rewards verbosity. Platforms bill on a combination of units: hosts or containers, ingested volume (logs, traces, custom metrics), indexed volume, users, modules. When the unit is volume, every engineering decision has a cost: a log level switched to debug, full tracing left on, an extra tag. An application that throws a lot of errors also emits a lot of logs: technical debt gets paid twice.
Cardinality. A metric described by route, status code and method stays reasonable. Add a user or session ID, and the number of series jumps by an order of magnitude. The cardinality simulator shows this live.
No feedback loop. Those who instrument do not see the bill, and those who pay it do not see the decisions that drive it. Without that link, drift is the natural slope.
These mechanisms apply to every solution. When self-managed, drift does not show up on a bill but in storage, compute and team time: it is just less visible.
Reading a bill line by line
Section titled “Reading a bill line by line”An observability bill usually breaks down into seven to ten lines. Their relative weight varies by vendor and by your usage, which is why it pays to measure it in your own environment.
| Line | Typical unit | What makes it drift |
|---|---|---|
| Hosts, containers | flat fee per monitored unit | autoscaling, test environments left running |
| Log ingestion | ingested volume | debug level, duplicated logs, non-production environments |
| Log indexing and retention | indexed volume, duration | indexing everything instead of archiving |
| Custom metrics | active series | cardinality |
| Traces | ingested or indexed spans | full tracing, long retention |
| Optional modules | per unit or per test | modules switched on for a trial and never revisited |
| Users | per seat | low in value, useful in negotiation |
| Security, SIEM, automation | variable | recent lines, to be shared with the CISO’s budget |
The first exercise is simple and rarely done: ask for the bill broken down by line, by team and by environment over a rolling twelve months. Without that view, no serious discussion is possible.
The technical levers
Section titled “The technical levers”Trace sampling. Head sampling decides at the start of the trace: simple, but it loses the rare traces, the ones with errors or high latency. Tail sampling decides at the end: it keeps every error, every trace above a latency threshold, and a fraction of normal traffic. It requires an intermediate collection layer, typically an OpenTelemetry Collector. The sampling simulator compares the two at equal volume.
Log retention tiers. An authentication log and a debug log have neither the same value nor the same lifespan. A three-tier scheme works well: hot (indexed, instant search, a few days to a few weeks), warm (compressed archive, retrievable within a day), cold (low-cost object storage, for compliance). The durations for each category follow the grid in lesson 2.
Controlling cardinality. No unique identifier as a metric dimension: that information belongs in traces. A published cap on series per service, and an automated check that rejects a metric above the threshold.
Upstream aggregation. Aggregate at the source, or at the Collector, rather than at the provider: one metric per second instead of one data point per request, with no loss for standard dashboards.
Measuring with intent. A few dozen well-chosen business metrics are worth more than thousands of metrics emitted by default by a framework. Every metric should answer three questions: what do we want to measure, for which decision, for whom? A metric with no audience is pure spend.
The Collector cost simulator puts a figure on these levers for your volume.
The contractual levers
Section titled “The contractual levers”Commitment. Vendors grant discounts in exchange for a multi-year commitment or a committed volume. The trap: committing to an overestimated trajectory. The prudent approach is to commit below the usage observed over the last twelve months and to negotiate flexibility, both upward and downward.
Clauses to request every time:
- a billing cap, with notification and the option to suspend ingestion rather than pay for an abnormal spike;
- a long notice period before any price increase;
- auditing of your own usage through the API, without restriction;
- documented reversibility, with export in an open format;
- unit prices held if the offering changes or the vendor is acquired;
- a ban on training models on your data without explicit consent.
Keeping negotiating power. OpenTelemetry instrumentation that sends to one or more interchangeable backends is the simplest way to stay free. It also lets you separate uses: one backend for day-to-day operations, cold storage for compliance.
Benchmarking the market. Coming to a renewal with two or three comparable offers, on the same scope and the same requirements, changes the incumbent vendor’s stance. The exercise takes a few weeks.
FinOps governance
Section titled “FinOps governance”The FinOps Foundation describes three maturity levels, Crawl, Walk and Run, assessed capability by capability; it states that reaching Run everywhere is not a goal in itself (FinOps maturity model). The mapping below is my reading of that model applied to observability, not a definition from the Foundation:
- Crawl, visibility: knowing what you spend, by team, application and environment. Prerequisite: tags enforced at ingestion (team, application, environment, criticality).
- Walk, accountability: each team sees its monthly cost (showback) and answers for it in a review.
- Run, continuous optimization: internal chargeback (chargeback), budgets per team, trade-offs handled as for the cloud.
Who does it. Three models coexist: centralized in the platform team (clear, but a bottleneck), federated with a champion per team and a central unit, or delegated to the vendor. The last one is comfortable, but the vendor has no interest in lowering its bill for good: it cannot be the main model.
The rituals. A monthly drift review (thirty minutes: the month’s biggest increases, their causes, the decisions). A quarterly committee (usage forecasts, commitments). And after any billing surprise, a blameless post-mortem with the team involved.
Four indicators.
| Indicator | What it tells you | Expected trend |
|---|---|---|
| Cost per business transaction | the cost of observability relative to activity | stable or falling |
| Committed share of the bill | what is covered by negotiated commitments | to be set according to your predictability |
| Cost per team or per service | who consumes what | compared month over month |
| Observability to infrastructure ratio | observability spend relative to infrastructure spend | tracked over time, compared across entities |
Four traps
Section titled “Four traps”- The ratchet effect: a usage tier reached during a spike stays billed when usage comes back down. Negotiate the way down in the initial contract, not just the way up.
- Non-production environments: development, test and staging are often instrumented like production and left running overnight. Measure their share; it is often surprising.
- Abandoned dashboards: they consume queries, and therefore budget. An annual cleanup review is an easy lever.
- Tracing drift: full tracing switched on for an incident and never turned back down, as in the journal above.
Exercises
Section titled “Exercises”- Quick bill audit (two hours): plot each line over twelve months, identify the three that grew the most and, for each, find the decision that explains it. If you cannot find it, visibility is your first project.
- Quick benchmark (a few weeks): ask three alternatives for an offer on a scope identical to yours, and compare with your bill.
- Governance (one quarter): write a one-page FinOps RACI (who optimizes, who answers for the budget, who is consulted, who is informed), circulate it, and launch the monthly review.
Checklist
- Costs broken down by team, application and environment
- Standardized tags mandatory at ingestion
- Traces tail-sampled rather than kept at 100%
- Logs classified into hot, warm and cold tiers, with justified durations
- Cap, reversibility and no-training clauses in the contract
- At least two comparable offers known before the next renewal
- Monthly drift review with the teams
- Cost per business transaction tracked
Going further
Section titled “Going further”- The Collector cost simulator, to put a figure on each lever for your volume.
- The business dimension.