Skip to contact
Official IBM training · HPC · 3 days · H010G

IBM Spectrum LSF
basic HPC cluster administration

Official IBM course H010G — three days to deploy, run and debug a Spectrum LSF cluster in production: install, queues, fairshare and SLA policies, CPU, GPU and memory management, application integration and troubleshooting. The enterprise alternative to SLURM when the business needs vendor support, real prioritisation and HA.

Duration 3 days · 24 h
Level Intermediate · Linux admin
Hands-on Real HPC cluster
Format Online or in-person
IBM Spectrum LSF logo
IBM Spectrum LSF HPC cluster bsub · bjobs · bkill Fairshare & SLAs GPU scheduling EGO vs SLURM · PBS LIM · RES · mbatchd IBM Spectrum LSF HPC cluster bsub · bjobs · bkill Fairshare & SLAs GPU scheduling EGO vs SLURM · PBS LIM · RES · mbatchd
[ 01 ]Why this course

From bsub to running
an HPC cluster in production.

Three days with a real IBM Spectrum LSF cluster running in our labs: master, candidate masters and hosts, configured queues, shared CPU and GPU pools, and real numerical simulation workloads. The goal is that you leave with effective operational skill, not just notes.

We cover what shows up every day in production — submitting and controlling jobs with bsub, queue design, fairshare and SLA policies, per-user and per-project limits, application integration, administration through EGO and debugging when things break. All aligned with the official IBM H010G syllabus.

[ 02 ]Who it's for · what you take away

A course for teams running serious HPC workloads

It's for your team if…

  • You are Linux or UNIX system administrators who will deploy or maintain an LSF cluster in production.
  • You're coming from SLURM, PBS or LoadLeveler and need to move to LSF, or run both side by side.
  • You manage batch workloads with SLAs and per-project prioritisation: numerical simulation, EDA, ML, rendering.
  • You need to share CPU, GPU, memory and software licences across users and groups with clear, auditable rules.

When you finish you're operating

  • You install LSF from scratch on RHEL or Debian and validate master, candidate masters and hosts.
  • You submit and control jobs with bsub, bjobs, bkill, bstop and bresume.
  • You configure consumable resources (CPU, GPU, memory, licences) and per-user, per-group, per-project limits.
  • You design queues with fairshare, preemption and SLA policies for business-critical workloads.
  • You run the cluster day to day and debug with logs, traces and bconf when something breaks in production.
[ 03 ]Syllabus · 3 days

From LSF architecture to daily operation

Three days that build the same cluster layer by layer: start from fundamentals and installation, go up to queues, fairshare and SLA policies, and finish administering and debugging the system as in real production.

01 Day 1 · Fundamentals and installation

Architecture, installation and first job submission

  • What LSF solves and how it compares to SLURM, PBS or Kubernetes for HPC.
  • Internal architecture: LIM, RES, sbatchd, mbatchd and the master role.
  • Step-by-step installation on RHEL or Debian and cluster verification.
  • First submission with bsub and basic cluster query commands.
Hands-on real cluster
02 Day 2 · Resources, queues and policies

Resource management, queues and scheduling policies

  • Consumable resources: CPU, GPU, memory and software licences.
  • Queue design by user profile and workload type.
  • Fairshare, preemption and relative priority policies.
  • SLAs and guarantees, per-user, per-group, per-host and per-project limits.
Hands-on real cluster
03 Day 3 · Operations, EGO and troubleshooting

Applications, EGO administration and debugging

  • Integration and deployment of applications inside the cluster.
  • Daily administration: start, stop and adding new hosts.
  • Resource administration through EGO when the deployment calls for it.
  • Debugging with logs, traces and bconf, and fixing the most common failures.
Hands-on real cluster
[ 05 ]Frequently asked questions

Common questions

What prior experience do I need?

Basic Linux or UNIX administration — bash, permissions, networking, shared file systems (NFS or similar). No prior experience with LSF or other IBM products is required. If you come from SLURM, PBS or LoadLeveler, the concepts will feel familiar.

When does LSF make sense compared to SLURM or Kubernetes?

LSF is the enterprise choice when you need contractual SLAs, complex prioritisation between projects, fine-grained management of expensive software licences and official IBM support. SLURM covers the free scheduling side, and Kubernetes is designed for services, not for traditional HPC batch workloads. In the course we run a practical comparison and help you decide with the right criteria.

What infrastructure are the labs run on?

On a real HPC cluster in our own labs, with master, candidate masters and several hosts. Each participant works against an environment that simulates production, not an isolated virtual machine.

Do you cover GPU scheduling?

Yes. Managing GPUs as consumable resources is covered inside the resource management block: how to declare them, assign them per job and share them across users with fairshare rules or per-project limits.

How does it differ from advanced LSF courses?

This is the official basic administration course (H010G) — it covers full deployment and day-to-day operation of an LSF cluster. Advanced courses go deeper into fine tuning, specific integrations and large architectures; we deliver them separately once the team has a solid foundation.

What delivery modes are available?

Get in touch to check modes and dates. We deliver training in English, Spanish and French, live online or in person, and we can prepare custom in-company sessions on your real cluster.

How much does it cost?

Price is calculated per group and depends on format, language and delivery mode. Tell us about your case — number of students, location and target dates — and we'll come back with a concrete proposal.

Can I book it in-company for my team?

Yes. We adapt pace, examples and exercises to your organisation's real stack — OS versions, network topology, applications and SLAs. Write to us with the number of students, location (on-site or remote) and required languages.
[ 06 ]   Request information

Tell us about your cluster

Let us know how many you are, in what language, which LSF version you use (or which scheduler you're coming from) and whether it's online or in person. We'll come back with a concrete proposal — format, dates and instructor team — not a generic quote.

Phone+34 91 198 02 43
PriceOn request per group
LanguagesEN · ES · FR
SIXE