Skip to content

TechnicalBeginner

Understanding high performance computing

For: engineers and SREs · architects · team managers · executives and CIOsPrerequisites: None.

Reading mode

Audience: beginners, non-technical profiles. Prerequisites: none. Goal: follow any HPC discussion.

  • the basic ideas, with analogies;
  • the common vocabulary;
  • the five components of a cluster;
  • the life cycle of a job, step by step;
  • the link between HPC and modern AI.

You will find neither source code nor system configuration. For monitoring a cluster, see HPC observability foundations.

Three properties characterize the problems a workstation cannot handle on its own.

PropertyMeaning
Huge amount of computationthe number of operations exceeds by orders of magnitude what one processor can do in reasonable time
Very large datadatasets do not fit in the memory of a single machine and must be distributed
Hard deadlinea weather forecast delivered two weeks late is useless

So you connect many machines and have them work together as one. That is an HPC system, also called a supercomputer, or simply a cluster.

Do many things at once, on many machines, on the same problem.

You gain speed (hours instead of weeks) and size, since you can take on problems that do not fit in a single machine.

  1. Anatomy of a cluster: the five components and the life of a job.
  2. Parallelism: three patterns and the limits of scaling.
  3. HPC and AI: two worlds that became one, and what it means for observability.