Skip to content

TechnicalBeginner

Anatomy of a cluster

For: engineers and SREs · architects · team managers · executives and CIOsPrerequisites: None.

ComponentRole
Login nodeprepare and submit the work
Schedulerqueue, priorities, resource allocation
Compute nodesthe servers that compute, with processors and accelerators
Interconnectthe very fast network between nodes
Parallel storagethe shared, high-performance file system

Work always follows the same path. The user prepares it on the login node and submits it to the scheduler, which reserves compute nodes. These nodes communicate over the interconnect and read from or write to parallel storage.

flowchart TB
  U["User"] --> L["Login<br/>node"]
  L -->|"submits"| S["Scheduler"]
  S -->|"reserves"| C1["Compute node 1<br/>CPU, GPU"]
  S -->|"reserves"| C2["Compute node 2<br/>CPU, GPU"]
  C1 <--> IC["Interconnect"]
  C2 <--> IC
  C1 <--> FS[("Parallel<br/>storage")]
  C2 <--> FS

A node is a powerful server: processors, memory, network link. A cluster has from a few hundred to several thousand of them. Each node can work alone; the power comes from their cooperation on the same problem. If a node is a house, the cluster is the city.

Accelerators, most often GPUs, contain thousands of small specialized cores. Originally designed for graphics, they excel at repetitive mathematical computation, for simulation as for AI. On the right workload, a GPU outperforms a classic processor by a factor of ten or more. A node often carries several.

The interconnect targets a latency in the order of a microsecond, that is, as an order of magnitude, a hundred to a thousand times lower than an office network, depending on the software stack. It is mostly InfiniBand and high-speed Ethernet.

Parallel storage is a file system shared by all nodes. Its aggregate throughput ranges from a few tens of GB/s to several TB/s depending on the size of the system. Data is spread across many disks and servers. Common solutions are Lustre, IBM Storage Scale (formerly Spectrum Scale, or GPFS) and BeeGFS.

The scheduler, the cluster’s air traffic controller

Section titled “The scheduler, the cluster’s air traffic controller”

With thousands of users and jobs, no one can run whatever they want whenever they want. The scheduler:

  • queues submitted jobs;
  • decides priorities and applies fairness rules;
  • reserves resources and starts jobs;
  • accounts for usage per project or per user.

The reference tool is Slurm, where you submit a job with the sbatch command. Other solutions: PBS or OpenPBS, IBM LSF.

StepWhat happens
1. Write the scriptprogram, requested resources, expected duration
2. Submitsbatch on Slurm
3. Queuethe job joins the queue
4. Decisionpriority and fairness applied
5. Allocationnodes are reserved for the job
6. Executioncomputation, communication, writing
7. Releaseaccounting and notification

The same steps seen as the job’s successive states:

stateDiagram-v2
  state "Script written" as Script
  state "Submitted" as Soumis
  state "Queued" as EnFile
  state "Prioritized" as Decide
  state "Allocated" as Alloue
  state "Running" as Exec
  state "Released" as Libere
  [*] --> Script
  Script --> Soumis: sbatch
  Soumis --> EnFile
  EnFile --> Decide: priority and fairness
  Decide --> Alloue: nodes reserved
  Alloue --> Exec
  Exec --> Libere: computation ends
  Libere --> [*]: accounting, notification

Users do not pick a compute node themselves to launch their work: they go through the login node, from which they submit the job to the scheduler. Many sites then allow SSH access to the nodes allocated to one of the user’s running jobs, to follow its execution, but not to other nodes.

Revised on 2 October 2026: interconnect latency presented as an order of magnitude (in the order of a microsecond, a hundred to a thousand times lower than an office network depending on the software stack). Unsourced claim about its cost removed; IBM Storage Scale instead of Spectrum Scale, parallel storage throughputs aligned with the reference course.