Skip to contact
HPC training · 2 days · open source scheduler

SLURM
basic HPC cluster administration

Two days to deploy, run and debug a SLURM Workload Manager cluster in production: installation of slurmctld, slurmd and slurmdbd, partitions and QoS, fairshare and backfill policies, CPU, GPU and memory management with GRES, accounting and troubleshooting. The open source scheduler running on MareNostrum and most of the Top500 — the free-software cousin of IBM Spectrum LSF.

Duration 2 days · 16 h
Level Intermediate · Linux admin
Hands-on Real HPC cluster
Format Online or in-person
slurm Workload Manager
SLURM Workload Manager HPC cluster sbatch · srun · squeue Fairshare & backfill GPU with GRES MUNGE auth slurmdbd accounting vs LSF · PBS · Torque slurmctld · slurmd SLURM Workload Manager HPC cluster sbatch · srun · squeue Fairshare & backfill GPU with GRES
[ 01 ]Why this course

From sbatch to running
a SLURM cluster in production.

Two days with a real SLURM cluster running in our labs: controller, backup and compute nodes, configured partitions, shared CPU and GPU pools with GRES, and real simulation workloads to exercise day-to-day operations. You leave running the system, not just with notes.

We cover what shows up every day in production — submitting and controlling jobs with sbatch, srun and squeue, partition and QoS design, fairshare and backfill policies, per-user, per-account and per-association limits, accounting with slurmdbd and debugging when something breaks. Natural companion to the IBM Spectrum LSF course — cover both schedulers that share the HPC market.

[ 02 ]Who it's for · what you take away

A course for teams administering or migrating HPC for real

It's for your team if…

  • You are Linux system administrators who will deploy or maintain a SLURM cluster in production.
  • You're coming from LSF, PBS or Torque and need to move to SLURM, or keep both running side by side during migration.
  • You manage scientific batch workloads — simulation, CFD, bioinformatics, ML, rendering — with per-project prioritisation.
  • You need to share CPU, GPU and memory across groups with clear rules, fair limits and auditable accounting.

When you finish you're operating

  • You install SLURM from scratch on RHEL or Debian with slurmctld, slurmd, slurmdbd and MUNGE.
  • You submit and control jobs with sbatch, srun, squeue, scancel and sinfo.
  • You design partitions and QoS with fairshare, backfill, limits and per-project priority.
  • You configure GPUs through GRES, per-association limits and real accounting with slurmdbd.
  • You run and debug with logs, sacct and scontrol when a job hangs or a node goes down.
[ 03 ]Syllabus · 2 days

From SLURM architecture to daily operation

Two days that build the same cluster layer by layer: start from fundamentals and deployment, then move up to partitions, fairshare policies, accounting and troubleshooting as in real production.

01 Day 1 · Fundamentals and installation

Architecture, deployment and first job submission

  • What SLURM solves and how it compares to LSF, PBS and Kubernetes for HPC.
  • Architecture: slurmctld, slurmd, slurmdbd and MUNGE; controller, backup and nodes.
  • Step-by-step installation on RHEL or Debian and cluster verification.
  • First submission with sbatch and srun; querying with squeue, sinfo and scontrol.
Hands-on real cluster
02 Day 2 · Partitions, policies and operation

Resources, QoS, accounting and troubleshooting

  • Designing partitions and QoS; per-user, per-account and per-association limits.
  • Fairshare, backfill and preemption policies with multifactor priority.
  • GPUs and generic resources (GRES); assignment per job and per partition.
  • Accounting with slurmdbd and sacct; debugging with logs, scontrol and sdiag.
Hands-on real cluster
[ 05 ]Frequently asked questions

Common questions

What prior experience do I need?

Basic Linux administration — bash, permissions, networking, shared file systems (NFS or similar). No prior experience with SLURM or other schedulers is required. If you come from LSF, PBS or Torque, the concepts will feel familiar.

When does SLURM make sense compared to LSF?

SLURM is the de facto standard in academic and scientific HPC: open source, no licence cost, huge community and presence across most of the Top500 — including the MareNostrum at BSC. LSF is still the choice when the business needs contractual SLAs, official IBM support and fine-grained management of expensive licences (EDA, CAE). We run a practical comparison in class.

What infrastructure are the labs run on?

On a real HPC cluster in our own labs with controller, backup controller and several compute nodes. Each participant works against an environment that simulates production, not an isolated virtual machine.

Do you cover GPUs and accelerators?

Yes. Managing GPUs through GRES is covered inside the resource block: how to declare them, assign them per job, share them across users and account for them. The same pattern applies to other accelerators.

Does this course include advanced scripting and tuning?

This is the basic administration course — full deployment and day-to-day operation. Fine scheduler tuning, specific integrations (Lustre, GPFS, containers) and multi-cluster architectures are delivered as custom advanced in-company training once the team has a solid foundation.

What delivery modes are available?

Get in touch to check modes and dates. We deliver training in English, Spanish and French, live online or in person, and prepare custom in-company sessions on your real cluster.

How much does it cost?

Price is calculated per group and depends on format, language and delivery mode. Tell us about your case — number of students, location and target dates — and we'll come back with a concrete proposal.

Can I take it together with the LSF course?

Yes — that's the usual pattern for teams running mixed HPC or migrating between schedulers. We can chain the two courses (2+3 days) or design an in-company path with the balance your team needs. Write to us with your scenario.
[ 06 ]   Request information

Tell us about your cluster

Let us know how many you are, in what language, which SLURM version you use (or which scheduler you're coming from) and whether it's online or in person. We'll come back with a concrete proposal — format, dates and instructor team — not a generic quote.

Phone+34 91 198 02 43
PriceOn request per group
LanguagesEN · ES · FR
SIXE