Technical
HPC AI: observability and security reference architecture
For: engineers and SREs · architectsPrerequisites: Basic HPC notions: node, scheduler, interconnect.
HPC (High Performance Computing) brings together hundreds to thousands of machines to solve problems too large for a single computer. Since the early 2020s, training large models has become a major workload, and the dominant one on clusters dedicated to AI. This module gives a reference architecture to observe and secure them with a self-hosted open source stack.
On video: from the GPU to the platform
Section titled “On video: from the GPU to the platform”Read the video text
87% reported GPU utilization.
The dashboard is green. Yet the model is training slowly. Where does the time go?
A CPU has a few powerful cores: it chains varied tasks, one after another. A GPU has thousands of small cores: they all run the same operation at the same time. And a neural network is mostly matrix multiplications: billions of identical calculations. The model lives in the GPU’s memory: 80 GB, read at 3.35 TB/s on an H100.
A 70-billion-parameter model weighs about 140 GB, for its weights alone. It does not fit in a single GPU. And training it? About 1.7 million GPU hours for Llama 2 70B (Meta, 2023).
So hundreds of GPUs must work together: that is a cluster. Inside a server, 8 GPUs linked by NVLink: the internal highway, 900 GB/s. Between servers, InfiniBand: a very fast network, 400 Gb/s per port. Around them: shared storage feeding the GPUs, and Slurm handing out the work.
At each step, every GPU computes on its share of the data… … then they all pool their results before moving on. But if a single GPU is slower… … all the others wait for it. And while waiting, they count as “busy”. That is where GPU hours disappear.
Observability runs the investigation. GPU metrics, tied to the job, show which GPU slows every step. Each GPU’s traces show in which phase the time is lost. The job context says whom to talk to and what to check: data, network or hardware. Which job, which GPU, which phase: the answer fits on one screen.
But nobody can investigate every job by hand. The answer: make observability a service of the platform. Golden path: the job label is injected at launch, without asking the researcher anything. Self-service: every researcher sees their own job, without opening a ticket. Service objectives: queue wait, successful jobs, real GPU efficiency. Chargeback: the truly useful GPU hours, team by team. Default guardrails: security is provided by the platform.
A shared cluster also attracts threats specific to HPC AI. Some escape classic tools, like RDMA traffic, invisible to a firewall. They are detected with the same observability stack.
Key takeaways:
- A “busy” GPU is not necessarily a computing GPU.
- Observe the job, the GPU and the phase, not just the machine.
- The platform makes this observability automatic, for everyone.
Keywords:
- Matrix: a table of numbers; a model’s weights are made of them
- HBM: memory stacked right against the GPU processor
- Parameters: the numbers the model learns; 2 bytes each in 16-bit
- NVLink: direct link between the GPUs of one server
- InfiniBand: very low latency network between servers
- Slurm: scheduler: it assigns GPUs to jobs
- Forward and backward: computing the prediction, then the corrections to make
- AllReduce: pooling the results of every GPU
- Trace: the detailed path of an operation, step by step
- Platform engineering: a team provides other teams with self-service internal tools
- SLO: a numeric service quality objective
- RDMA: direct access to another server’s memory, bypassing its system
Explore the cluster in the interactive diagram: www.meantimetolearn.com.
Music: “Radar Focus”, Blue Saga (Epidemic Sound).
Why one machine is not enough
Section titled “Why one machine is not enough”| Reference point | Order of magnitude |
|---|---|
| GPT-3 training compute (2020) | about 3,640 petaflop/s-days, that is hundreds of GPU-years |
| Memory to load Llama 3 70B weights in 16 bits | about 140 GB (70 billion x 2 bytes) |
| HBM bandwidth of one H100 SXM5 | 3.35 TB/s |
| Typical laptop | 16 GB RAM, 8 GB GPU memory |
In 16 bits, the weights alone of a 70-billion-parameter model (140 GB) do not fit in the memory of an 80 GB GPU. They would fit on a 141 GB or 192 GB GPU, or in 8 bits, but training needs several times more memory than the weights: gradients, optimizer states, activations. The work therefore has to be spread across several GPUs and nodes, all working together at every step.
CPU and GPU: two specialties
Section titled “CPU and GPU: two specialties”| CPU | GPU | |
|---|---|---|
| Cores | 8 to 192 powerful cores depending on the model (192 on an AMD EPYC 9965) | thousands of simple CUDA cores (16,896 on an H100 SXM, for example) |
| Clock | 2 to 5 GHz | 1 to 2 GHz |
| Strength | branching, arbitrary memory access, orchestration, I/O | the same simple operation on thousands of values at once |
| Role in AI | orchestrate, load data, control | the matrix multiplications at the heart of models |
Core counts are examples: the maximums rise with every generation.
The four chapters
Section titled “The four chapters”- Cluster architecture: the five layers and what each exposes.
- Anatomy of a training step: data load, forward and backward passes, AllReduce, optimizer, and the straggler effect.
- The observability stack: collection, storage, visualization, and eBPF to observe without instrumenting.
- The threat surface: six vectors specific to HPC AI and how to detect them.
Revised on 2 October 2026: the place of large models on clusters correctly dated, memory needed for a 70-billion-parameter model clarified, core counts presented as examples and updated, the 87% and 19% example presented as illustrative.