Skip to content

TechnicalPractitioner

The HPC AI observability stack

For: engineers and SREs · architectsPrerequisites: Have read chapter 2 of the course; basics of Prometheus and Linux.

Metrics, logs and traces on a training cluster

Section titled “Metrics, logs and traces on a training cluster”
SignalContentStorageUse
MetricsSM utilization, HBM bandwidth, step time, NCCL latency, InfiniBand retransmitsPrometheus-compatible time-series database: Prometheus with Thanos or Mimir, or VictoriaMetricsalerts, trends
LogsCUDA errors, NCCL warnings, Slurm interruptions, out-of-memory killsLoki or VictoriaLogsqualitative context metrics miss
Tracesend-to-end path across all nodes and processesTempo or Jaegerfind on which node and in which function time is lost; essential for stragglers
flowchart TB
  subgraph C["Collection (per node)"]
    direction LR
    D["dcgm-exporter (NVIDIA)<br/>or AMD Device Metrics<br/>Exporter, GPU"]
    NE["node-exporter<br/>CPU, RAM, disk"]
    O["OTel Collector<br/>traces, logs"]
    F["Falco / Tetragon<br/>kernel events"]
    S["slurm-exporter<br/>jobs, queues"]
    I["ibstat + scripts<br/>InfiniBand"]
  end
  subgraph St["Storage (central)"]
    direction LR
    VM["Prometheus + Thanos or Mimir,<br/>or VictoriaMetrics<br/>metrics, 90 d"]
    L["Loki or VictoriaLogs<br/>logs, 30 d"]
    T["Tempo / Jaeger<br/>traces, 7 d"]
    AM["Alertmanager"]
  end
  subgraph V["Visualization (Grafana)"]
    direction LR
    V1["GPU cluster"]
    V2["Network and storage"]
    V3["Training"]
    V4["Security"]
    V5["Slurm jobs"]
  end
  C --> St --> V

Retention periods are starting values, to adjust to your audit obligations. No component is proprietary: everything is open source, self-hosted and auditable, and data never leaves the infrastructure.

At each tier, several open source components fill the same role; the figure names a few without recommending one by default. For metrics, Prometheus with Thanos or Mimir and VictoriaMetrics are compared on the same criteria: handling of cardinality (one series per GPU and per rank), long retention, high availability, operating cost and team skills. Commercial offerings such as Grafana Cloud, Datadog or Dynatrace also cover these needs; they fall outside the self-hosted scope of this architecture, but are judged with the same grid.

eBPF attaches kernel-verified programs to thousands of instrumentation points: system calls, kernel functions, network events, file access, and through uprobes to user-space library functions. Overhead is low and training code stays untouched.

Trace CUDA memory allocations of a process (a libcuda function, so a uprobe, not a kprobe):

Fenêtre de terminal
# The libcuda path depends on the distribution and driver.
bpftrace -e 'uprobe:/usr/lib/x86_64-linux-gnu/libcuda.so.1:cuMemAlloc_v2
/pid == $1/ { printf("alloc %lu bytes\n", arg1); }' <PID>

Detect unexpected outbound connections from a Python process:

Fenêtre de terminal
bpftrace -e 'kprobe:tcp_connect /comm == "python3" || comm == "pt_main_thread"/ {
$sk = (struct sock *)arg0;
printf("%s -> %s\n", comm, ntop($sk->__sk_common.skc_daddr));
}'

PyTorch versions 2.3 and 2.4 rename the main thread pt_main_thread (added by PR 121170, removed before 2.5 by PR 134066): with these versions, a filter on python3 alone does not see training processes. Prefer filtering by PID (/pid == $1/) or by the job’s cgroup (cgroup) rather than by process name. The skc_daddr field only covers IPv4; for IPv6, read skc_v6_daddr.

Watch file opens in the checkpoint directory:

Fenêtre de terminal
bpftrace -e 'tracepoint:syscalls:sys_enter_openat
/strncmp(str(args->filename), "/checkpoints", 12) == 0/ {
printf("[%s] opens %s\n", comm, str(args->filename));
}'

This filter only sees absolute paths: a file opened through a relative path, for example checkpoints/step_100.pt from the working directory, escapes it.

Two production tools build on eBPF:

  • Falco detects and alerts in real time on suspicious behavior: unusual file access, unexpected connections, binary modifications.
  • Tetragon (Cilium project) enforces per-process syscall policies and can block an unauthorized operation before it completes: CUDA ioctl, network connection, checkpoint read.
ToolRole
Prometheus with Thanos or Mimir, or VictoriaMetricsPrometheus-compatible metric storage; to be compared on handling of cardinality (per GPU, per rank), retention and operating cost
dcgm-exporterNVIDIA GPU metrics via DCGM: SM, HBM, ECC, NVLink, power; the entry point for NVIDIA GPUs, with the hpc_job label to tie metrics to the job
AMD Device Metrics Exporterthe equivalent for AMD GPUs (ROCm), with a Slurm integration
OpenTelemetrytraces and logs, Python SDK for PyTorch, OTLP protocol, Collector
Grafanadashboards correlating metrics, logs and traces
Loki or VictoriaLogslabel-indexed log aggregation, efficient for Slurm, NCCL and Python logs
FalcoeBPF behavioral detection
TetragoneBPF syscall policy enforcement
Slurmscheduler; prolog and epilog for per-job instrumentation

Revised on 2 October 2026: GPU metrics tied to the job through dcgm-exporter’s hpc_job label (a UUID join does not work), AMD equivalent added, metric storage options presented with the same criteria, bpftrace filter adapted to the pt_main_thread of PyTorch 2.3 and 2.4, role of the prolog and epilog scripts clarified, IPv4 and relative path limits flagged.