Skip to content

Technical

Technical dimension

The technical dimension is about getting the signal you need to troubleshoot, at the right cost. Every chapter starts from a concrete operations problem and gives the full configuration. The goal is to make it reproducible in a lab; the Labs page says which ones are published, and a chapter without a lab or with a configuration not yet run end to end says so.

If you are new to the topic, start with the first three chapters of the table, in order: they set out the vocabulary used everywhere else.

ChapterWhat you learn to doStatus
Observability in 10 minutestell monitoring from observability, metrics, logs, traces, profiles and events apart, link them, reason about cardinalitypublished, for beginners; cardinality simulator
OpenTelemetry end to endfollow a span from the application to storage, instrument, configure a minimal Collector, apply semantic conventionspublished, for beginners
SLOs and alertingdefine SLIs and SLOs, compute an error budget, alert on burn rate instead of thresholdspublished, for beginners; error budget simulator
Telemetry pipelinesdesign a Collector: filtering, transformation, sampling, routingplanned
Backends and storagechoose between Prometheus, VictoriaMetrics, Elastic, Loki, Tempo, ClickHouseplanned
AI observabilityobserve inference, RAG and MCP tools with an open source stackpublished, with a full course
Log integritybuild an audit log a third party can verifypublished
HPC and GPU infrastructuretie every GPU metric to a job and an account, catch silent degradationreference architecture, course in preparation
Agents, RAG and MCPper-component measurement grids, probabilistic quality SLOsGenAI method
VictoriaMetricsa metrics backend for high-cardinality workloadscourse
Edge and IoTobserve constrained, intermittent and heterogeneous fleetsplanned

Trace explorer, clickable architecture, animated Collector pipeline, tail sampling, quality SLO. The rest are gathered on the tools and diagrams pages.