Skip to content

TechnicalPractitioner

Quiz: HPC AI observability and security

For: engineers and SREs · architectsPrerequisites: Have followed the course lessons.

Reading mode

Ten questions to check what you retain from the course. Each answer is corrected immediately, with an explanation and a link to the relevant lesson. Nothing is sent: your best score stays in this browser.

  1. Question 1 of 10Inside a GPU node, which link do the GPUs use to talk to each other?One answer
  2. Question 2 of 10Which of these signal families belong to the four families to collect on an HPC AI cluster?Several answers possible
  3. Question 3 of 10Bridge: businessIn the lesson's illustrative example (constructed values, not a measurement), a GPU showing 87% utilization computes usefully only about 19% of the step. What does GPU utilization actually measure?One answer
  4. Question 4 of 10To spot a straggler slowing down the whole training run, which indicators does the lesson recommend measuring?Several answers possible
  5. Question 5 of 10Bridge: organizationHow do you concretely tie each GPU metric to a job, a user and an account, the basis for chargeback?One answer
  6. Question 6 of 10You want to trace calls to cuMemAlloc_v2 in the libcuda library with bpftrace. Which kind of probe should you use?One answer
  7. Question 7 of 10Which statements about Falco and Tetragon are correct?Several answers possible
  8. Question 8 of 10RDMA traffic over InfiniBand bypasses the kernel. Where can visibility into this traffic come from?One answer
  9. Question 9 of 10Bridge: business linesYour cluster is shared between several customers. Which countermeasure does the lesson give against a job reading weights left in GPU memory by the previous job?One answer
  10. Question 10 of 10Bridge: peopleAccording to the course, why should security and operations teams work together on an HPC AI cluster?One answer