Skip to content

OrganizationPractitioner

Governing telemetry

For: engineers and SREs · architects · team managers · finance, risk and compliancePrerequisites: Basic notions of metrics, labels and cardinality.

Reading mode

Without governance, telemetry follows a natural slope: every team adds metrics, attributes and logs, nobody removes any, and retention stays at the default. The result shows first on the bill, then in slowing queries, sometimes in an audit that finds personal data in logs kept far too long.

Governing telemetry means deciding in advance what may be collected, in what quantity, for how long and where it goes, then checking automatically that those decisions are respected. Four tools are enough: the data contract, the cardinality budget, the retention policy and a quarterly review.

The glossary defines it as a written rule, checked in continuous integration, that sets for a signal the allowed attributes, the cardinality budget, the retention and the destination. I recommend adding two fields: an owner and whether the signal may contain personal data.

Here is what a contract for a metric can look like. The format is an illustrative example, to be adapted to your tooling; it does not follow the schema of any specific tool.

signal: http.server.request.duration
type: histogram
owner: payments-team
usage: [payments-slo, payments-dashboard]
allowed_attributes:
- service.name
- http.request.method
- http.response.status_code
- http.route # normalized route, never the raw path
forbidden_attributes:
- user.id # unbounded: belongs on traces
- url.full
cardinality_budget: 5000 # maximum active series for this signal
personal_data: no
retention: operations-class
destination: hot-metrics-storage

Three choices in this contract deserve an explanation.

The usage field. Every signal must serve a purpose: an SLO, an alert, a dashboard, a decision. A signal with no declared usage is the first candidate for removal at review time. This is the frugality rule proposed in lesson 3 of the CIO path.

Attribute names. The OpenTelemetry semantic conventions provide standardized names and meanings. I recommend using them rather than inventing an in-house vocabulary: contracts stay readable from one team to another and from one tool to another.

Forbidden attributes. The Prometheus documentation explicitly advises against using labels to store high-cardinality dimensions such as user IDs or email addresses. Writing these bans into the contract makes them checkable.

A contract checked only in human review ends up ignored. I recommend three automated checks, from the simplest to the most complete:

CheckWhat it blocksPossible tools
Naming and formatnon-compliant names, missing unitspromtool check metrics for metrics in Prometheus format
Convention complianceunknown or mistyped attributesOTel Weaver, which validates a semantic convention registry
Contractforbidden attributes, signal without owner or usagean in-house test comparing declared instrumentation with the contracts

The third check is, in my view, the most useful; it is also the one a generic tool cannot provide as is, since it depends on your contracts. It can start small: a script that rejects a merge when a new signal has no contract.

Cardinality is the number of distinct time series a metric produces. It grows multiplicatively with labels, which the cardinality explosion simulator makes visible.

A cardinality budget sets a ceiling of active series per signal, per service or per team. I recommend setting it at two levels:

  • per signal, in the contract, to stop in review the attribute that would blow up a metric;
  • per team, as an overall envelope, so the team makes its own trade-offs between its signals.

The budget must also be enforced at runtime, because a mistake always gets through review one day. On the Prometheus side, the sample_limit and label_limit scrape settings make the scrape of a target fail when it exceeds a ceiling (Prometheus configuration). In an OpenTelemetry Collector, the filter and transform processors drop or rewrite forbidden attributes before export. The Collector cost simulator quantifies the effect of these levers.

The exception process. A budget with no possible exception gets bypassed. I recommend a short, written process: the team asks, justifies the usage and proposes a duration; the platform team estimates the impact; the decision is recorded in the contract with an end date. An exception without an end date becomes the rule.

Not all data has the same useful life. I recommend defining a small number of retention classes and attaching each signal to a class in its contract.

ClassExample signalsDurationStorage
Diagnosissampled nominal traces, authorized debug logsa few dayshot
Operationsoperational metrics, application logs, error tracesa few weekshot then warm
Trendaggregated metrics (per service, reduced resolution)several months to a few yearscold or compressed
Evidenceaudit logs, traces with evidential valueset by the retention policy and applicable regulationscold, immutable if needed

The durations in this table are illustrative orders of magnitude. Evidence-class durations are not chosen based on cost: they come from the company’s retention policy and sector regulations, as explained in lesson 2 of the CIO path. For the integrity of these logs, see audit log integrity.

Personal data. The GDPR sets out the principles of data minimization and storage limitation (Regulation (EU) 2016/679, article 5). Applied to telemetry, they mean that a user identifier, an IP address or request content should only be collected if it serves a declared usage, and should not be kept longer than that usage requires. The contract’s personal_data field makes this choice explicit and checkable by the DPO.

Contracts and budgets age. I recommend a one-hour quarterly review bringing together the platform team, product team champions, a FinOps representative and, at least once a year, security and the DPO.

Typical agenda.

  1. Change in volume and cardinality per team since the previous review.
  2. The ten most expensive signals and their declared usage.
  3. Signals with no observed usage over the quarter: remove or justify.
  4. Exceptions that have reached their end date: close or renew with a new date.
  5. Incidents where a signal was missing: add it, with a contract.
  6. Retention and personal data deviations found.
  7. Decisions, with an owner and a date.

Item 5 matters as much as the others. Governance is not only about cutting: a post-mortem that concludes there was an observability gap must be able to add a signal quickly. See blameless post-mortems.

IndicatorWhat it reveals
Share of signals covered by a contractthe real reach of governance
Share of signals with a declared and observed usagehow frugal collection is
Active series per team, against their budgetwhether cardinality budgets hold
Number of merges blocked by contract checkswhether rules are applied or bypassed
Open and expired exceptionsthe discipline of the exception process
Signals containing personal data outside any contractcompliance risk
Cost per team and cost per business transactionthe link with FinOps management in lesson 4
  • A data contract format is chosen and documented
  • Every new signal comes with a contract, an owner and a usage
  • Attribute names follow the OpenTelemetry semantic conventions
  • Continuous integration blocks a signal without a contract
  • A cardinality budget exists per team and is enforced at runtime
  • The exception process is written down and every exception has an end date
  • Every signal is attached to a retention class
  • The presence of personal data is declared and reviewed with the DPO
  • The quarterly review takes place and produces recorded decisions