Business
Part V. Quantified impact analysis
For: architects · finance, risk and compliance · executives and CIOsPrerequisites: Have read part I (the silent failure model).
18. The cost model of not monitoring
Section titled “18. The cost model of not monitoring”Unifying principle: the cost of a blind spot is the product of an accumulation rate and a detection delay. Without the signal, the detection delay is that of an external channel (monthly bill, customer complaint); we assume one billing cycle, that is 30 days. With the signal, it drops to minutes or hours.
signal_value ~= accumulation_rate x (delay_without_signal - delay_with_signal)Notation: V requests or tasks per day, t_in and t_out average input and output tokens, p_in and p_out price per token.
cost_per_call = t_in * p_in + t_out * p_outbase_monthly_cost = V * cost_per_call * 30Quantifying the main blind spots:
- Agent loop. Normal task at K steps, loop rate r, looping steps L. Monthly overspend
= V * r * (L - K) * cost_per_step * 30. Overnight runaway scenario: an agent stuck at R requests per minute for H hours costsR * 60 * H * cost_per_call. - Prompt bloat. Monthly growth g of input tokens undetected over M months. Cumulative overspend approximated by
base_monthly_cost_in * g * M * (M + 1) / 2, wherebase_monthly_cost_in = V * t_in * p_in * 30. - Unexploited response cache. Achievable hit rate h not captured: foregone saving
= h * base_monthly_cost. The formula holds for a response cache, where a hit avoids the whole call. A prompt cache only lowers the price of input tokens read from cache: the saving is then smaller. - Hallucination. Expected cost
= P_hallucination * N_exposed_incidents * cost_per_incident, the cost per incident including remediation, reputational harm and regulatory exposure.
19. Worked example
Section titled “19. Worked example”Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list. The example therefore sizes the blind spots in tokens and calls; to get euros, multiply input tokens by p_in and output tokens by p_out (see annex D to build your table).
Illustrative parameters:
- V = 50,000 calls per day, t_in = 1,500, t_out = 400 tokens.
- Daily volume: 50,000 x 1,500 = 75 million input tokens and 50,000 x 400 = 20 million output tokens.
- Monthly volume, the cost base: 2,250 million input tokens and 600 million output tokens, so
base_monthly_cost = 2,250 M x p_in + 600 M x p_out.
Blind spots, on the same parameters:
- Undetected prompt bloat, g = 8 percent per month over M = 4 months: cumulative overconsumption on the order of
2,250 M x 0.08 x 4 x 5 / 2 = 1,800 million input tokensover four months, to multiply by p_in, never alerted. - Achievable response cache h = 30 percent unexploited:
0.30 x 50,000 x 30 = 450,000 callsavoidable per month, that is 675 million input tokens and 180 million output tokens, or 30 percent of the base cost. - A single agent running away one night, R = 20 requests per minute for 8 hours:
20 x 60 x 8 = 9,600 callsin one night, that is 14.4 million input tokens and 3.84 million output tokens, multiplied by incident frequency.
Reading: none of these losses triggers a latency or error alert. They materialize precisely the silent failure model, and justify the observability investment in euros, once your prices are applied, not in curves.
20. Tool: the economic exposure calculator
Section titled “20. Tool: the economic exposure calculator”The model above is operationalized in an interactive calculator, intended for a discussion with a finance function or a risk committee. You enter volume, tokens, prices, the agent profile and drift assumptions, and it returns base cost, the exposure of each blind spot and the total exposure of not monitoring over twelve months, that is the value observability can avoid. The calculator simplifies this model on some items; its page details each difference.
21. Prioritization matrix
Section titled “21. Prioritization matrix”quadrantChart title Prioritization of blind spots x-axis "Loud" --> "Silent" y-axis "Low severity" --> "Critical severity" quadrant-1 "Priority 1: monitor first" quadrant-2 "Covered by classical monitoring" quadrant-3 "Low stakes" quadrant-4 "Instrument next" "Hallucination": [0.85, 0.9] "Cost drift": [0.9, 0.7] "MCP tool poisoning": [0.8, 0.95] "Agent loop": [0.75, 0.65] "Quality drift": [0.88, 0.6] "Provider rate-limit": [0.25, 0.55] "Server 500 error": [0.15, 0.5] "One-off latency": [0.3, 0.25]
Revised on 2 October 2026: prices removed; the worked example is expressed in tokens and calls.