Skip to content

TechnicalPractitioner

HPC AI: observability and security reference architecture

For: engineers and SREs · architectsPrerequisites: Basic HPC notions: node, scheduler, interconnect.

Reading mode

HPC (High Performance Computing) brings together hundreds to thousands of machines to solve problems too large for a single computer. Since the early 2020s, training large models has become a major workload, and the dominant one on clusters dedicated to AI. This module gives a reference architecture to observe and secure them with a self-hosted open source stack.

Read the video text

87% reported GPU utilization.

The dashboard is green. Yet the model is training slowly. Where does the time go?

A CPU has a few powerful cores: it chains varied tasks, one after another. A GPU has thousands of small cores: they all run the same operation at the same time. And a neural network is mostly matrix multiplications: billions of identical calculations. The model lives in the GPU’s memory: 80 GB, read at 3.35 TB/s on an H100.

A 70-billion-parameter model weighs about 140 GB, for its weights alone. It does not fit in a single GPU. And training it? About 1.7 million GPU hours for Llama 2 70B (Meta, 2023).

So hundreds of GPUs must work together: that is a cluster. Inside a server, 8 GPUs linked by NVLink: the internal highway, 900 GB/s. Between servers, InfiniBand: a very fast network, 400 Gb/s per port. Around them: shared storage feeding the GPUs, and Slurm handing out the work.

At each step, every GPU computes on its share of the data… … then they all pool their results before moving on. But if a single GPU is slower… … all the others wait for it. And while waiting, they count as “busy”. That is where GPU hours disappear.

Observability runs the investigation. GPU metrics, tied to the job, show which GPU slows every step. Each GPU’s traces show in which phase the time is lost. The job context says whom to talk to and what to check: data, network or hardware. Which job, which GPU, which phase: the answer fits on one screen.

But nobody can investigate every job by hand. The answer: make observability a service of the platform. Golden path: the job label is injected at launch, without asking the researcher anything. Self-service: every researcher sees their own job, without opening a ticket. Service objectives: queue wait, successful jobs, real GPU efficiency. Chargeback: the truly useful GPU hours, team by team. Default guardrails: security is provided by the platform.

A shared cluster also attracts threats specific to HPC AI. Some escape classic tools, like RDMA traffic, invisible to a firewall. They are detected with the same observability stack.

Key takeaways:

  1. A “busy” GPU is not necessarily a computing GPU.
  2. Observe the job, the GPU and the phase, not just the machine.
  3. The platform makes this observability automatic, for everyone.

Keywords:

  • Matrix: a table of numbers; a model’s weights are made of them
  • HBM: memory stacked right against the GPU processor
  • Parameters: the numbers the model learns; 2 bytes each in 16-bit
  • NVLink: direct link between the GPUs of one server
  • InfiniBand: very low latency network between servers
  • Slurm: scheduler: it assigns GPUs to jobs
  • Forward and backward: computing the prediction, then the corrections to make
  • AllReduce: pooling the results of every GPU
  • Trace: the detailed path of an operation, step by step
  • Platform engineering: a team provides other teams with self-service internal tools
  • SLO: a numeric service quality objective
  • RDMA: direct access to another server’s memory, bypassing its system

Explore the cluster in the interactive diagram: www.meantimetolearn.com.

Music: “Radar Focus”, Blue Saga (Epidemic Sound).

Reference pointOrder of magnitude
GPT-3 training compute (2020)about 3,640 petaflop/s-days, that is hundreds of GPU-years
Memory to load Llama 3 70B weights in 16 bitsabout 140 GB (70 billion x 2 bytes)
HBM bandwidth of one H100 SXM53.35 TB/s
Typical laptop16 GB RAM, 8 GB GPU memory

In 16 bits, the weights alone of a 70-billion-parameter model (140 GB) do not fit in the memory of an 80 GB GPU. They would fit on a 141 GB or 192 GB GPU, or in 8 bits, but training needs several times more memory than the weights: gradients, optimizer states, activations. The work therefore has to be spread across several GPUs and nodes, all working together at every step.

CPUGPU
Cores8 to 192 powerful cores depending on the model (192 on an AMD EPYC 9965)thousands of simple CUDA cores (16,896 on an H100 SXM, for example)
Clock2 to 5 GHz1 to 2 GHz
Strengthbranching, arbitrary memory access, orchestration, I/Othe same simple operation on thousands of values at once
Role in AIorchestrate, load data, controlthe matrix multiplications at the heart of models

Core counts are examples: the maximums rise with every generation.

  1. Cluster architecture: the five layers and what each exposes.
  2. Anatomy of a training step: data load, forward and backward passes, AllReduce, optimizer, and the straggler effect.
  3. The observability stack: collection, storage, visualization, and eBPF to observe without instrumenting.
  4. The threat surface: six vectors specific to HPC AI and how to detect them.

Revised on 2 October 2026: the place of large models on clusters correctly dated, memory needed for a 70-billion-parameter model clarified, core counts presented as examples and updated, the 87% and 19% example presented as illustrative.