Business
What a business or product owner can ask of observability
For: business and product · team managers · executives and CIOsPrerequisites: None.
This page is for business owners and product owners. It requires no technical knowledge. Technical terms are defined as they appear and in the glossary.
Why it concerns you
Section titled “Why it concerns you”Observability is the ability to understand what happens inside a system from the data it emits. It is often associated with server outages. Yet it first answers business questions. Do our customers achieve what they came to do? How long do they wait? What does each use cost?
These answers do not appear on their own. The engineering team instruments the system to answer the questions it is asked. If nobody asks the business questions, it instruments what it knows: processors, memory, technical errors. The Bridges page develops this idea and offers a translation dictionary between technical signals and business indicators.
Your role is therefore to frame the questions. The engineering team’s role is to make the answers possible.
The questions you can ask
Section titled “The questions you can ask”Three words are enough to read this section. A signal is a type of data emitted by the application. A trace tells the path of a request through all the services. An attribute is a label added to that trace, for example “journey step = payment”. The details are in Observability in 10 minutes.
The attribute names below are illustrative. Your team will choose its own and write them down in a shared naming rule.
“How many customers were affected by yesterday’s incident and for how long?”
Section titled ““How many customers were affected by yesterday’s incident and for how long?””- The indicator. The number of distinct customers whose requests failed or were too slow, with start and end times.
- The signals. Traces and logs, provided they carry a customer identifier or at least a segment.
- What to ask for. A customer segment attribute on traces. A pseudonymized customer identifier, if the data protection officer agrees.
- A point of caution. The customer identifier goes on traces, not on metrics. On a metric, it multiplies the storage cost: this is cardinality.
- Effort, in my view. Modest if traces already exist. Heavier if the services must be traced first.
“Where in the purchase journey do customers drop out or wait?”
Section titled ““Where in the purchase journey do customers drop out or wait?””- The indicator. For each step, the share of customers who move on to the next one and the time spent.
- The signals. Traces carrying the journey step, linked to business events such as “cart confirmed” or “order paid”.
- What to ask for. A step attribute, for example
journey.step. The emission of the key business events. A dashboard per step, readable without training. - Effort, in my view. A project of a few iterations, to run step by step, starting with the one that worries you most.
“Is the new feature slower than the old one?”
Section titled ““Is the new feature slower than the old one?””- The indicator. Response time compared between the two versions. Teams often look at p95, the time under which 95% of requests complete. The average hides the worst-served customers.
- The signals. Traces and response time metrics, tagged with the version or the active feature flag.
- What to ask for. That every piece of data carries the service version. OpenTelemetry, the open collection standard, provides the
service.versionattribute for this. For a progressive rollout, add the feature flag name. - Effort, in my view. Low if the version is already filled in. I recommend making it a condition of every production release.
“How much does our AI feature cost per conversation?”
Section titled ““How much does our AI feature cost per conversation?””- The indicator. The number of tokens consumed per conversation, multiplied by the unit price in your contract. A token is the unit of text billed by model providers.
- The signals. Traces of calls to the model. OpenTelemetry’s conventions for generative AI cover the input and output tokens of each call.
- What to ask for. A conversation identifier on each call. An attribute naming the feature concerned, to allocate the cost.
- Effort, in my view. Reasonable if the application goes through a shared layer for model calls. Otherwise, plan for it from the design stage.
“Are we within our commitment to customers?”
Section titled ““Are we within our commitment to customers?””- The indicator. The service level actually observed, compared to the promise, and the remaining margin.
- The signals. Success and response time measurements on the journeys covered by the promise.
- What to ask for. One SLO per important journey, tracked in a dashboard you can open yourself.
- Effort, in my view. The most structuring of the five, because it requires agreeing on the promise. That is the subject of the next section.
Turning a customer promise into an SLO
Section titled “Turning a customer promise into an SLO”Three acronyms come up often. An SLI measures what the customer experiences, for example the share of payments that succeed in under 2 seconds. An SLO is the internal objective set on that SLI. An SLA is the contractual commitment, with penalties if it is not met. The SLOs and alerting article covers them in detail.
I recommend working out the wording with the engineering team in four steps:
- Start from a customer sentence. For example, as an illustration: “A customer who pays sees the confirmation in under 2 seconds.”
- Ask what is measurable. The team says where it can measure and what the measurement does not see.
- Look at the current level before setting the target. A target chosen without measurement is either untenable or pointless.
- Write the final sentence. “99.9% of payments succeed in under 2 seconds, over a rolling 30 days” (illustrative example).
Aiming for 99.9% means accepting 0.1% failures. This tolerated share is called the error budget. Over 30 days, it amounts to about 43 minutes of full outage (computed value: 0.1% of 43,200 minutes).
The error budget is a shared decision tool. While some remains, the team can ship new features and take risks. When it is exhausted, priority shifts to reliability. Google’s Site Reliability Workbook publishes an example error budget policy that writes this rule down. I recommend that the product owner sign it with the engineering team, before the first incident. The error budget simulator helps discuss it on a concrete case.
Your place in incidents and post-mortems
Section titled “Your place in incidents and post-mortems”During the incident. Expect a single point of contact and regular updates, even when there is no progress. Avoid contacting the people doing the repair directly. Bring what they cannot see: the key customers affected, a commercial deadline, a communication to prepare.
After the incident. The post-mortem is the meeting that analyzes the incident to draw actions from it, without looking for someone to blame. The template in The blameless post-mortem includes a business impact established with the business concerned. Your presence is therefore not a courtesy. Google’s chapter on post-mortem culture also states that any stakeholder can request one.
I recommend coming with three things: what customers experienced, what the incident cost in your view and what you would have needed to know sooner. The last point often becomes an instrumentation action.
AI features: quality and cost, not only uptime
Section titled “AI features: quality and cost, not only uptime”An AI feature can respond quickly, without any technical error, and still respond wrongly. Uptime is therefore not enough. Chapter 1 of the AI observability guide adds two signals to the classic ones: evaluations, which score the quality of answers, and user feedback.
I recommend asking for three indicators for any AI feature:
- quality: the share of answers judged correct, with an objective, in other words a quality SLO;
- cost per use: per conversation, per document processed or per customer;
- feedback: the share of answers that users flag as wrong or unhelpful.
The quality SLO simulator shows how to set an objective on quality. Chapter 3 of the guide helps locate your current level. To put a figure on an undetected degradation, see the GenAI method, part V and its exposure calculator.
Questions to ask your engineering team
Section titled “Questions to ask your engineering team”This list fits on one page and can serve as a meeting handout.
| Question | What to ask the team for | Sign that it is in place |
|---|---|---|
| How many customers did an incident affect? | a segment or a pseudonymized customer identifier on traces | the figure is available the next day, without a manual query |
| Where does the journey stall? | a step attribute and the key business events | a dashboard per step that you open yourself |
| Is the new version slower? | the version and the feature flag on every piece of data | a before and after comparison at every release |
| How much does AI cost per use? | tokens and a conversation identifier on each call | a cost per conversation tracked every month |
| Does the AI answer correctly? | evaluations and collection of user feedback | a quality SLO with an error budget |
| Are we keeping the customer promise? | one SLO per important journey | a visible error budget and a signed policy |
| Who decides when the budget is exhausted? | a written error budget policy | a name next to each decision |
| Will I be involved in post-mortems? | a systematic invitation when customers are affected | an impact section filled in with you |
Going further
Section titled “Going further”- Bridges: the questions from other departments and the translation dictionary.
- Observability in 10 minutes: the basic vocabulary on a slow payment example.
- SLOs and alerting: SLI, SLO, SLA and error budget in detail.
- The blameless post-mortem: the full process and template.
- The executive dashboard: an example of indicators presented to leadership.
- The maturity self-assessment: locate your organization.
- The MTTR business case: what faster detection is worth.
- The AI observability guide: chapters 1 and 3 to start.
Sources
Section titled “Sources”- Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering, O’Reilly, 2016: chapter 4, Service Level Objectives and chapter 15, Post-mortem Culture.
- Beyer, Murphy, Rensin, Kawahara and Thorne (eds.), The Site Reliability Workbook, O’Reilly, 2018: chapter 2, Implementing SLOs and the appendix Example Error Budget Policy.
- OpenTelemetry, semantic conventions for service resources (
service.versionattribute) and semantic conventions for generative AI (token usage). - The 43-minute budget is computed from a 30-day period, that is 43,200 minutes.