Technical
HPC AI threat surface
For: engineers and SREs · architectsPrerequisites: Have read chapter 3 of the course.
What needs protecting
Section titled “What needs protecting”Three assets are at stake: model weights (intellectual property), training data (often confidential) and infrastructure (shared across teams or tenants). GPUs require privileged kernel access, which opens attack surfaces absent from a conventional data center.
Six vectors and their detection
Section titled “Six vectors and their detection”| Vector | Scenario | Detection and mitigation |
|---|---|---|
| Supply chain (PyPI) | a Python package ships a malicious CUDA kernel that exfiltrates weights during the backward pass | private registry, hash of every dependency, locked installs |
| Data poisoning | a file on Lustre is modified between the integrity check and the DataLoader read, to plant a backdoor | SHA-256 per shard, verified at read time |
| RDMA blind spot | reading a peer’s memory over RDMA; this traffic bypasses the kernel, no firewall, IDS or iptables sees it | InfiniBand anomaly metrics, behavioral baseline, partition isolation (P_Key) |
| Gradient injection | in federated training, crafted gradients plant a backdoor in the final model | gradient norm distribution, Kullback-Leibler divergence, robust aggregation |
| GPU memory leakage | the CUDA documentation states that allocated memory is not cleared; outside confidential computing mode, nothing guarantees that a job cannot read data left by the previous one, weights included | GPUs allocated exclusively to one job; explicit memory wipe or GPU reset (nvidia-smi --gpu-reset, with no active process) in the prolog or epilog; MIG partitioning for GPUs shared at the same time; confidential computing mode (Hopper and later), where the reset scrubs memory before the GPU goes to the next tenant |
| Container escape | NVIDIA Container Toolkit flaws: CVE-2024-0132 (fixed in 1.16.2), its bypass CVE-2025-23359 (fixed in 1.17.4), then CVE-2025-23266, known as NVIDIAScape (CVSS 9.0, fixed in 1.17.8), and CVE-2026-24260, a time-of-check time-of-use race condition (TOCTOU, CVSS 8.5, published on 1 July 2026, toolkit up to and including 1.19.0 and GPU Operator up to and including 26.3.1, fixed in 1.19.1 and 26.3.2); a malicious image can reach the host node and every workload | toolkit at version 1.19.1 or later and GPU Operator at version 26.3.2 or later, follow NVIDIA security bulletins, Falco rules, Tetragon syscall policy |
Sources: NVIDIA Container Toolkit release notes, NVD, CVE-2025-23266, NVD, CVE-2026-24260, NVIDIA security bulletin 5850 (June 2026), and, for memory scrubbing in confidential mode, ACM Queue, “Creating the First Confidential GPUs” (republished in Communications of the ACM).
flowchart LR
W(("Model weights<br/>intellectual property"))
A["PyPI supply chain"] --> W
B["Lustre poisoning"] --> W
C["RDMA blind spot"] --> W
D["Gradient injection"] --> W
E["GPU memory leakage"] --> W
F["Container escape"] --> W
The special case of RDMA
Section titled “The special case of RDMA”RDMA (Remote Direct Memory Access) lets a network card read or write another machine’s memory directly, without going through the kernel. That is what makes InfiniBand so fast, and also what makes it invisible to conventional security tools, eBPF included, since they rely on kernel hooks. The only visibility comes from adapter counters and the subnet manager: you need a per-job traffic baseline and alerts on deviations.
The detection layer
Section titled “The detection layer”| Tool | What it does |
|---|---|
| Falco | real-time detection of suspicious file access, unexpected connections, binary modifications |
| Tetragon | per-process syscall policies, blocks unauthorized CUDA ioctls and network calls |
| InfiniBand metrics | the only signal available on RDMA traffic |
| Data hashes | shard integrity at read time |
Revised on 2 October 2026: CVE-2025-23359, CVE-2025-23266 and CVE-2026-24260 added, minimum toolkit version raised to 1.19.1 and GPU Operator to 26.3.2, GPU memory leakage countermeasures made specific (reset or wipe, MIG, confidential computing).