Skip to content

Cross-cutting

Labs

Reading a configuration is not enough to understand it. The site aims for every technical article to link to a lab that reproduces it on a workstation. That is not yet the case everywhere: here is the actual state.

What is published and what is in preparation

Section titled “What is published and what is in preparation”
LabTopicState
labs/VMLLMVictoriaMetrics as an LLM observability backend, material for the VictoriaMetrics coursepublished on the forge
labs/HPCOBSHPC observability proof of concept, with a training simulatorpublished on the forge
labs/genaiotelRAG application instrumented with OpenTelemetry, with an evaluation modulepublished on the forge
labs/cardinality-cost-labcardinality explosion and its cost, with load profilespublished on the forge
labs/llm-observabilitykit for LLM observability: the labs (OpenTelemetry, VictoriaMetrics, Tempo, Grafana, Phoenix, Qdrant, Ollama)being fixed, not yet published

Articles that do not have a lab yet say so, and a configuration published without having been run end to end carries a box that flags it.

Upcoming kits will ship a .devcontainer/ folder; the labs already published on the forge do not have one yet. Once the kit is published, the stack will open in VS Code (Dev Containers), or in GitHub Codespaces if the repository is also published on GitHub.

The real Grafana dashboard of the HPCOBS lab, filmed live (1 min 36). The data comes from the lab simulator: a 4-GPU training job, then two failures triggered on purpose. GPU utilization stays above 90% while the data loading wait rises and MFU drops; then the gradients explode while the infrastructure looks normal. The lab code is freely available on the forge (public code repository).Music: “Radar Focus”, Blue Saga (Epidemic Sound).
  1. Clone the lab repository from the forge.

  2. Start the stack with docker compose up. No account is required, but you need to download the container images and, for AI labs, the local models, which means several gigabytes on first start.

  3. Follow the scenario in README.md, which has you trigger the incident, observe it and then fix it.

  4. Check the result with the checks listed in README.md, when the lab provides them.