Reading mode Engineer Decision-maker
Decision-maker reading: bridges first.
Ten questions to check what you retain from the course. Each answer is corrected immediately, with an explanation and a link to the relevant lesson. Nothing is sent: your best score stays in this browser.
Question 1 of 10 Inside a GPU node, which link do the GPUs use to talk to each other? One answer Review: Lesson 1: HPC AI cluster architecture
Question 2 of 10 Which of these signal families belong to the four families to collect on an HPC AI cluster? Several answers possible GPU: SM utilization, HBM bandwidth, ECC errors This is the first family, exposed by dcgm-exporter.
Billing: hourly cost of each node This is not a signal family in the lesson; chargeback comes from tying metrics to the job and the account.
Model: loss, gradient norm, step time, AllReduce latency This is the family most often forgotten, even though it is the one that tells you whether training is progressing.
Storage: Lustre throughput, metadata latency Storage is one of the four families, along with GPU, network and model.
Review: Lesson 1: HPC AI cluster architecture
Question 3 of 10 Bridge: business In the lesson's illustrative example (constructed values, not a measurement), a GPU showing 87% utilization computes usefully only about 19% of the step. What does GPU utilization actually measure? One answer Review: Lesson 2: Anatomy of a distributed training step
Question 4 of 10 To spot a straggler slowing down the whole training run, which indicators does the lesson recommend measuring? Several answers possible Review: Lesson 2: Anatomy of a distributed training step
Question 5 of 10 Bridge: organization How do you concretely tie each GPU metric to a job, a user and an account, the basis for chargeback? One answer Review: Lesson 3: The HPC AI observability stack
Question 6 of 10 You want to trace calls to cuMemAlloc_v2 in the libcuda library with bpftrace. Which kind of probe should you use? One answer Review: Lesson 3: The HPC AI observability stack
Question 7 of 10 Which statements about Falco and Tetragon are correct? Several answers possible Review: Lesson 3: The HPC AI observability stack
Question 8 of 10 RDMA traffic over InfiniBand bypasses the kernel. Where can visibility into this traffic come from? One answer Review: Lesson 4: HPC AI threat surface
Question 9 of 10 Bridge: business lines Your cluster is shared between several customers. Which countermeasure does the lesson give against a job reading weights left in GPU memory by the previous job? One answer Review: Lesson 4: HPC AI threat surface
Question 10 of 10 Bridge: people According to the course, why should security and operations teams work together on an HPC AI cluster? One answer Review: Lesson 4: HPC AI threat surface
Start over
Bridges What this page changes beyond the technical team.
BusinessValue, cost, risk You can tell reported GPU utilization from useful compute, and so argue hardware purchases on real efficiency.
OrganizationWho decides, who owns You know how to tie each metric to a job, a user and an account, which underpins chargeback and answers to auditors.
PeopleTeams, on-call, skills You see why security and operations gain from sharing one stack and working together.
Business linesProduct, finance, risk, customer You know how to protect each customer's weights and data on shared infrastructure, and how to show research teams where their training time goes. Why bridges matter →