Skip to content

TechnicalPractitioner

HPC AI threat surface

For: engineers and SREs · architectsPrerequisites: Have read chapter 3 of the course.

Three assets are at stake: model weights (intellectual property), training data (often confidential) and infrastructure (shared across teams or tenants). GPUs require privileged kernel access, which opens attack surfaces absent from a conventional data center.

VectorScenarioDetection and mitigation
Supply chain (PyPI)a Python package ships a malicious CUDA kernel that exfiltrates weights during the backward passprivate registry, hash of every dependency, locked installs
Data poisoninga file on Lustre is modified between the integrity check and the DataLoader read, to plant a backdoorSHA-256 per shard, verified at read time
RDMA blind spotreading a peer’s memory over RDMA; this traffic bypasses the kernel, no firewall, IDS or iptables sees itInfiniBand anomaly metrics, behavioral baseline, partition isolation (P_Key)
Gradient injectionin federated training, crafted gradients plant a backdoor in the final modelgradient norm distribution, Kullback-Leibler divergence, robust aggregation
GPU memory leakagethe CUDA documentation states that allocated memory is not cleared; outside confidential computing mode, nothing guarantees that a job cannot read data left by the previous one, weights includedGPUs allocated exclusively to one job; explicit memory wipe or GPU reset (nvidia-smi --gpu-reset, with no active process) in the prolog or epilog; MIG partitioning for GPUs shared at the same time; confidential computing mode (Hopper and later), where the reset scrubs memory before the GPU goes to the next tenant
Container escapeNVIDIA Container Toolkit flaws: CVE-2024-0132 (fixed in 1.16.2), its bypass CVE-2025-23359 (fixed in 1.17.4), then CVE-2025-23266, known as NVIDIAScape (CVSS 9.0, fixed in 1.17.8), and CVE-2026-24260, a time-of-check time-of-use race condition (TOCTOU, CVSS 8.5, published on 1 July 2026, toolkit up to and including 1.19.0 and GPU Operator up to and including 26.3.1, fixed in 1.19.1 and 26.3.2); a malicious image can reach the host node and every workloadtoolkit at version 1.19.1 or later and GPU Operator at version 26.3.2 or later, follow NVIDIA security bulletins, Falco rules, Tetragon syscall policy

Sources: NVIDIA Container Toolkit release notes, NVD, CVE-2025-23266, NVD, CVE-2026-24260, NVIDIA security bulletin 5850 (June 2026), and, for memory scrubbing in confidential mode, ACM Queue, “Creating the First Confidential GPUs” (republished in Communications of the ACM).

flowchart LR
  W(("Model weights<br/>intellectual property"))
  A["PyPI supply chain"] --> W
  B["Lustre poisoning"] --> W
  C["RDMA blind spot"] --> W
  D["Gradient injection"] --> W
  E["GPU memory leakage"] --> W
  F["Container escape"] --> W

RDMA (Remote Direct Memory Access) lets a network card read or write another machine’s memory directly, without going through the kernel. That is what makes InfiniBand so fast, and also what makes it invisible to conventional security tools, eBPF included, since they rely on kernel hooks. The only visibility comes from adapter counters and the subnet manager: you need a per-job traffic baseline and alerts on deviations.

ToolWhat it does
Falcoreal-time detection of suspicious file access, unexpected connections, binary modifications
Tetragonper-process syscall policies, blocks unauthorized CUDA ioctls and network calls
InfiniBand metricsthe only signal available on RDMA traffic
Data hashesshard integrity at read time

Revised on 2 October 2026: CVE-2025-23359, CVE-2025-23266 and CVE-2026-24260 added, minimum toolkit version raised to 1.19.1 and GPU Operator to 26.3.2, GPU memory leakage countermeasures made specific (reset or wipe, MIG, confidential computing).