Organization
Module 5: industrialize
For: architects · team managers · finance, risk and compliancePrerequisites: Have followed the previous modules and labs of the course.
The 30-, 60- and 90-day plan
Section titled “The 30-, 60- and 90-day plan”Maturity grows in three stages, in this order. There is no point automating reindexing or a model change before you have a reliable dashboard.
| Deadline | Theme | Deliverables |
|---|---|---|
| 30 days | instrument | OpenTelemetry SDK on the main chain, GenAI conventions applied, LLM Overview dashboard, breakdown by customer and feature |
| 60 days | evaluate | faithfulness measured continuously, automatic drift detection, judge on 1% of traffic, self-hosted Phoenix |
| 90 days | govern | SLOs formalized and signed, one procedure per alert, monthly quality and cost review, logs and documentation required by the AI Act in place if the system is in scope |
Mapped onto the guide’s maturity scale (levels 0 to 5):
| Deadline | Guide levels targeted |
|---|---|
| 30 days | levels 1 and 2: basic telemetry (tokens, cost, latency, errors), then structured tracing with the GenAI conventions and attributes per customer and feature |
| 60 days | level 3 (online evaluation: faithfulness measured continuously, judge on a sample) and a first element of level 4 (drift detection) |
| 90 days | consolidation of levels 3 and 4: SLOs, procedures and the monthly review belong to governance, which the guide’s scale does not measure as such; level 5 (closed loop) is beyond this plan |
These deadlines hold for a narrow scope, one main chain like the one in the labs; for a full program, the guide considers level 3 realistic in six months.
Who watches what
Section titled “Who watches what”In short, each role has its indicators and its rhythm:
- on-call follows latency, error and token budget alerts continuously;
- the product team looks at the quality score, user feedback and drift every week;
- finance tracks cost per customer and per feature every month; an AI quality committee reviews trends, incidents and change decisions (prompt, retrieval, model version) every month.
The RACI matrix and the organization of quality on-call are detailed in part VII of the method.
Organization models are covered in detail in the organization section.
What the AI Act requires
Section titled “What the AI Act requires”- Obligations depend on the role (provider, who places the system on the market; deployer, who uses it) and on the risk level: record-keeping, technical documentation and post-market monitoring cover high-risk systems, whose deployer keeps automatically generated logs for at least six months; Article 50 (transparency) covers any system that interacts with people.
- Timeline from Regulation (EU) 2026/1744: Article 50 has applied since 2 August 2026, high-risk systems of Annex III from 2 December 2027, of Annex I from 2 August 2028.
- An OpenTelemetry stack provides the technical material (logs, measurements, history); qualifying the system and your role, documenting, and maintaining the setup over time remain to be done.
The detailed table of requirements, article by article, with the role concerned and what the stack covers, is in part VI of the method.
What the course covered
Section titled “What the course covered”- The top-tier indicators change: drift, faithfulness, hallucination and time to first token come before CPU, p95 and error rate.
- GenAI conventions make instrumentation portable, provided you absorb their changes in the Collector.
- A self-hosted stack (in the kit: OpenTelemetry, VictoriaMetrics, Tempo, Grafana and Phoenix) covers these needs; other combinations, open source or commercial, do so too, and OpenTelemetry instrumentation keeps the choice reversible.
- Maturity is built in the order of the plan above: instrument, evaluate, then govern.
Going further
Section titled “Going further”- Observing an LLM system: the production stack, MCP, privacy.
- GenAI observability method: quality SLOs, quantified impact, compliance, organization.
- Beyond the LLM: GPUs and model quality: GPU infrastructure and statistical drift detection.
- GenAI conventions: the
semantic-conventions-genairepository (since version 1.42 of the conventions). - Phoenix, RAGAS, Evidently and VictoriaMetrics documentation.
Revised on 2 October 2026: AI Act section corrected (provider and deployer roles, Article 26 and log retention, Article 50 applicable to any conversational system, timeline from Regulation (EU) 2026/1744), link to the new GenAI conventions repository.
Revised on 4 October 2026: milestones of the 30-, 60- and 90-day plan mapped onto the guide’s maturity levels, “retraining” replaced by the actual changes of an LLM called through an API (prompt, retrieval, model version); “Who watches what” and the AI Act section summarized with pointers to parts VII and VI of the method, which now holds the table of requirements by role.