Technical
Anatomy of a cluster
For: engineers and SREs · architects · team managers · executives and CIOsPrerequisites: None.
The five components of a cluster
Section titled “The five components of a cluster”| Component | Role |
|---|---|
| Login node | prepare and submit the work |
| Scheduler | queue, priorities, resource allocation |
| Compute nodes | the servers that compute, with processors and accelerators |
| Interconnect | the very fast network between nodes |
| Parallel storage | the shared, high-performance file system |
Work always follows the same path. The user prepares it on the login node and submits it to the scheduler, which reserves compute nodes. These nodes communicate over the interconnect and read from or write to parallel storage.
flowchart TB
U["User"] --> L["Login<br/>node"]
L -->|"submits"| S["Scheduler"]
S -->|"reserves"| C1["Compute node 1<br/>CPU, GPU"]
S -->|"reserves"| C2["Compute node 2<br/>CPU, GPU"]
C1 <--> IC["Interconnect"]
C2 <--> IC
C1 <--> FS[("Parallel<br/>storage")]
C2 <--> FS
Compute nodes and accelerators
Section titled “Compute nodes and accelerators”A node is a powerful server: processors, memory, network link. A cluster has from a few hundred to several thousand of them. Each node can work alone; the power comes from their cooperation on the same problem. If a node is a house, the cluster is the city.
Accelerators, most often GPUs, contain thousands of small specialized cores. Originally designed for graphics, they excel at repetitive mathematical computation, for simulation as for AI. On the right workload, a GPU outperforms a classic processor by a factor of ten or more. A node often carries several.
Interconnect and parallel storage
Section titled “Interconnect and parallel storage”The interconnect targets a latency in the order of a microsecond, that is, as an order of magnitude, a hundred to a thousand times lower than an office network, depending on the software stack. It is mostly InfiniBand and high-speed Ethernet.
Parallel storage is a file system shared by all nodes. Its aggregate throughput ranges from a few tens of GB/s to several TB/s depending on the size of the system. Data is spread across many disks and servers. Common solutions are Lustre, IBM Storage Scale (formerly Spectrum Scale, or GPFS) and BeeGFS.
The scheduler, the cluster’s air traffic controller
Section titled “The scheduler, the cluster’s air traffic controller”With thousands of users and jobs, no one can run whatever they want whenever they want. The scheduler:
- queues submitted jobs;
- decides priorities and applies fairness rules;
- reserves resources and starts jobs;
- accounts for usage per project or per user.
The reference tool is Slurm, where you submit a job with the sbatch command. Other solutions: PBS or OpenPBS, IBM LSF.
The life of a job in seven steps
Section titled “The life of a job in seven steps”| Step | What happens |
|---|---|
| 1. Write the script | program, requested resources, expected duration |
| 2. Submit | sbatch on Slurm |
| 3. Queue | the job joins the queue |
| 4. Decision | priority and fairness applied |
| 5. Allocation | nodes are reserved for the job |
| 6. Execution | computation, communication, writing |
| 7. Release | accounting and notification |
The same steps seen as the job’s successive states:
stateDiagram-v2 state "Script written" as Script state "Submitted" as Soumis state "Queued" as EnFile state "Prioritized" as Decide state "Allocated" as Alloue state "Running" as Exec state "Released" as Libere [*] --> Script Script --> Soumis: sbatch Soumis --> EnFile EnFile --> Decide: priority and fairness Decide --> Alloue: nodes reserved Alloue --> Exec Exec --> Libere: computation ends Libere --> [*]: accounting, notification
Users do not pick a compute node themselves to launch their work: they go through the login node, from which they submit the job to the scheduler. Many sites then allow SSH access to the nodes allocated to one of the user’s running jobs, to follow its execution, but not to other nodes.
Revised on 2 October 2026: interconnect latency presented as an order of magnitude (in the order of a microsecond, a hundred to a thousand times lower than an office network depending on the software stack). Unsourced claim about its cost removed; IBM Storage Scale instead of Spectrum Scale, parallel storage throughputs aligned with the reference course.