Technical
Understanding high performance computing
For: engineers and SREs · architects · team managers · executives and CIOsPrerequisites: None.
Reading mode
Audience: beginners, non-technical profiles. Prerequisites: none. Goal: follow any HPC discussion.
What the three lessons cover
Section titled “What the three lessons cover”- the basic ideas, with analogies;
- the common vocabulary;
- the five components of a cluster;
- the life cycle of a job, step by step;
- the link between HPC and modern AI.
You will find neither source code nor system configuration. For monitoring a cluster, see HPC observability foundations.
When one computer is not enough
Section titled “When one computer is not enough”Three properties characterize the problems a workstation cannot handle on its own.
| Property | Meaning |
|---|---|
| Huge amount of computation | the number of operations exceeds by orders of magnitude what one processor can do in reasonable time |
| Very large data | datasets do not fit in the memory of a single machine and must be distributed |
| Hard deadline | a weather forecast delivered two weeks late is useless |
So you connect many machines and have them work together as one. That is an HPC system, also called a supercomputer, or simply a cluster.
HPC in one sentence
Section titled “HPC in one sentence”Do many things at once, on many machines, on the same problem.
You gain speed (hours instead of weeks) and size, since you can take on problems that do not fit in a single machine.
Lessons
Section titled “Lessons”- Anatomy of a cluster: the five components and the life of a job.
- Parallelism: three patterns and the limits of scaling.
- HPC and AI: two worlds that became one, and what it means for observability.