Technical
The HPC AI observability stack
For: engineers and SREs · architectsPrerequisites: Have read chapter 2 of the course; basics of Prometheus and Linux.
Metrics, logs and traces on a training cluster
Section titled “Metrics, logs and traces on a training cluster”| Signal | Content | Storage | Use |
|---|---|---|---|
| Metrics | SM utilization, HBM bandwidth, step time, NCCL latency, InfiniBand retransmits | Prometheus-compatible time-series database: Prometheus with Thanos or Mimir, or VictoriaMetrics | alerts, trends |
| Logs | CUDA errors, NCCL warnings, Slurm interruptions, out-of-memory kills | Loki or VictoriaLogs | qualitative context metrics miss |
| Traces | end-to-end path across all nodes and processes | Tempo or Jaeger | find on which node and in which function time is lost; essential for stragglers |
A three-tier stack
Section titled “A three-tier stack”flowchart TB
subgraph C["Collection (per node)"]
direction LR
D["dcgm-exporter (NVIDIA)<br/>or AMD Device Metrics<br/>Exporter, GPU"]
NE["node-exporter<br/>CPU, RAM, disk"]
O["OTel Collector<br/>traces, logs"]
F["Falco / Tetragon<br/>kernel events"]
S["slurm-exporter<br/>jobs, queues"]
I["ibstat + scripts<br/>InfiniBand"]
end
subgraph St["Storage (central)"]
direction LR
VM["Prometheus + Thanos or Mimir,<br/>or VictoriaMetrics<br/>metrics, 90 d"]
L["Loki or VictoriaLogs<br/>logs, 30 d"]
T["Tempo / Jaeger<br/>traces, 7 d"]
AM["Alertmanager"]
end
subgraph V["Visualization (Grafana)"]
direction LR
V1["GPU cluster"]
V2["Network and storage"]
V3["Training"]
V4["Security"]
V5["Slurm jobs"]
end
C --> St --> V
Retention periods are starting values, to adjust to your audit obligations. No component is proprietary: everything is open source, self-hosted and auditable, and data never leaves the infrastructure.
At each tier, several open source components fill the same role; the figure names a few without recommending one by default. For metrics, Prometheus with Thanos or Mimir and VictoriaMetrics are compared on the same criteria: handling of cardinality (one series per GPU and per rank), long retention, high availability, operating cost and team skills. Commercial offerings such as Grafana Cloud, Datadog or Dynatrace also cover these needs; they fall outside the self-hosted scope of this architecture, but are judged with the same grid.
eBPF: observing without instrumenting
Section titled “eBPF: observing without instrumenting”eBPF attaches kernel-verified programs to thousands of instrumentation points: system calls, kernel functions, network events, file access, and through uprobes to user-space library functions. Overhead is low and training code stays untouched.
Trace CUDA memory allocations of a process (a libcuda function, so a uprobe, not a kprobe):
# The libcuda path depends on the distribution and driver.bpftrace -e 'uprobe:/usr/lib/x86_64-linux-gnu/libcuda.so.1:cuMemAlloc_v2 /pid == $1/ { printf("alloc %lu bytes\n", arg1); }' <PID>Detect unexpected outbound connections from a Python process:
bpftrace -e 'kprobe:tcp_connect /comm == "python3" || comm == "pt_main_thread"/ { $sk = (struct sock *)arg0; printf("%s -> %s\n", comm, ntop($sk->__sk_common.skc_daddr));}'PyTorch versions 2.3 and 2.4 rename the main thread pt_main_thread (added by PR 121170, removed before 2.5 by PR 134066): with these versions, a filter on python3 alone does not see training processes. Prefer filtering by PID (/pid == $1/) or by the job’s cgroup (cgroup) rather than by process name. The skc_daddr field only covers IPv4; for IPv6, read skc_v6_daddr.
Watch file opens in the checkpoint directory:
bpftrace -e 'tracepoint:syscalls:sys_enter_openat /strncmp(str(args->filename), "/checkpoints", 12) == 0/ { printf("[%s] opens %s\n", comm, str(args->filename));}'This filter only sees absolute paths: a file opened through a relative path, for example checkpoints/step_100.pt from the working directory, escapes it.
Two production tools build on eBPF:
- Falco detects and alerts in real time on suspicious behavior: unusual file access, unexpected connections, binary modifications.
- Tetragon (Cilium project) enforces per-process syscall policies and can block an unauthorized operation before it completes: CUDA ioctl, network connection, checkpoint read.
The stack’s tools
Section titled “The stack’s tools”| Tool | Role |
|---|---|
| Prometheus with Thanos or Mimir, or VictoriaMetrics | Prometheus-compatible metric storage; to be compared on handling of cardinality (per GPU, per rank), retention and operating cost |
| dcgm-exporter | NVIDIA GPU metrics via DCGM: SM, HBM, ECC, NVLink, power; the entry point for NVIDIA GPUs, with the hpc_job label to tie metrics to the job |
| AMD Device Metrics Exporter | the equivalent for AMD GPUs (ROCm), with a Slurm integration |
| OpenTelemetry | traces and logs, Python SDK for PyTorch, OTLP protocol, Collector |
| Grafana | dashboards correlating metrics, logs and traces |
| Loki or VictoriaLogs | label-indexed log aggregation, efficient for Slurm, NCCL and Python logs |
| Falco | eBPF behavioral detection |
| Tetragon | eBPF syscall policy enforcement |
| Slurm | scheduler; prolog and epilog for per-job instrumentation |
Revised on 2 October 2026: GPU metrics tied to the job through dcgm-exporter’s hpc_job label (a UUID join does not work), AMD equivalent added, metric storage options presented with the same criteria, bpftrace filter adapted to the pt_main_thread of PyTorch 2.3 and 2.4, role of the prolog and epilog scripts clarified, IPv4 and relative path limits flagged.