Skip to content

OrganizationExpert

Part IX. Implementation

For: engineers and SREs · architects · team managersPrerequisites: Have read parts I to VIII of the method.

flowchart LR
  P0["Phase 0<br/>Mapping"] --> P1["Phase 1<br/>Instrumentation"] --> P2["Phase 2<br/>SLI, SLO, alerting"] --> P3["Phase 3<br/>Evaluation"] --> P4["Phase 4<br/>FinOps"] --> P5["Phase 5<br/>Compliance"] --> P6["Phase 6<br/>Continuous improvement"]
  P6 -.->|"closed loop"| P3
  • Phase 0, mapping: components, flows, sensitivity boundaries, regulatory obligations, observability users.
  • Phase 1, instrumentation: OTel GenAI everywhere, no content capture by default, collector wired to existing backends.
  • Phase 2, SLI and SLO: service, quality and cost SLIs, dashboards correlated by trace, alerting on quality and cost drift.
  • Phase 3, evaluation: golden set, judge calibration, gating in continuous integration, production sampling.
  • Phase 4, FinOps: attribution, tracking per request, session and task, budgets, drift alerts, levers.
  • Phase 5, compliance: masking, content and metadata separation, distinct retention, immutable audit, MCP escalation to the SOC, evidence matrix.
  • Phase 6, continuous improvement: production feeds the datasets, regressions flow back to tests, thresholds are revised.

To turn the method into a skill, the target is an executable environment that reproduces each silent failure and shows how it is caught. This full lab is in preparation: it is not published yet, and the configurations and rules in this method have not yet been run there.

What already exists on the public labstraining forge:

  • labs/genaiotel: a RAG application instrumented with the OTel GenAI conventions (simulated model and embeddings, no API key), a Collector that computes the cost of each span (OTTL) and derives per-model metrics (spanmetrics), VictoriaMetrics and vmalert (cost rules, alerts on latency, cost, errors, input tokens, weak retrieved context and evaluated quality), Jaeger, Grafana and an evaluation pass to Phoenix. It covers part of scenarios 2 and 3 below; it contains no agent, no MCP and no Wilson bound sidecar.
  • labs/aiobsagent: an alert investigation agent written in Go, read-only, that fences third-party content against prompt injection and keeps a chained journal. It illustrates the injection guardrails and least privilege of the agent and MCP grids in Part III, but reproduces none of the four scenarios below.

Target structure of the full lab, in preparation:

labstraining/observability-genai/
docker-compose.yml # OTel collector, VictoriaMetrics, VictoriaLogs, Grafana, Langfuse
otel-collector.yaml # reference config (part II)
rules/ # MetricsQL recording rules and alerts
dashboards/ # provisioned Grafana dashboards
agent-demo/ # OTel GenAI instrumented agent (Python)
scenarios/ # reproducible failure scenarios
evals/ # golden set, RAGAS or DeepEval harness

Planned teaching scenarios, one per silent failure:

  1. Agent loop: the agent loops, cost per task explodes. The learner enables loop detection and the cost alert, then replays it.
  2. Prompt bloat: progressive input token growth, visible only through the dedicated recording rule.
  3. Hallucination from degraded RAG context: retrieval is degraded, faithfulness drops, the sampled eval catches it where latency stays green.
  4. MCP server mutation: a tool changes definition between two calls, mutation detection and tool pinning catch it.
  • Production sampling evaluation active
  • Golden set built, judge agreement tracked, judge versioned
  • Quality SLO with confidence interval
  • Regression tests in continuous integration
  • Drift detection configured
  • finish_reasons distribution monitored
  • Service SLOs per layer
  • Health checks for inference, vector store, MCP servers
  • Tested multi-provider fallback
  • Graceful degradation on tool unavailability
  • Tokens and cost traced with an attribution dimension
  • Budgets and alerts at 80 percent
  • Prompt bloat and agent loop monitoring
  • Cache hit rate measured and levers active
  • Content not captured by default, opt-in capture controlled
  • Masking and redaction at the collector
  • Content and metadata separation, distinct retention
  • Immutable audit trail on invocations and access
  • MCP security events escalated to the SOC
  • AI Act, NIS2, GDPR evidence matrix documented
  • Self-hosted backends if sovereignty required
  • RACI established over the eight capabilities
  • Quality on-call defined, mitigation playbooks written

Revised on 2 October 2026: the full lab of the method is presented as in preparation, with a precise pointer to the two existing public labs.