Skip to contact
Advanced training · Level 3

Ceph Production Operations | CourseWhen a 200 TB cluster fails at 3 AM and you need answers, not theory

The only distribution-agnostic course — IBM Storage Ceph, Red Hat, Ubuntu, Rocky, Alma Linux or Ceph upstream. Real production scenarios at petabyte scale that vendors don't teach, with content aligned to the Red Hat EX260 certification.

Ceph Production Operations training delivered by SIXE — instructor explaining critical scenarios for administering Ceph clusters at petabyte scale
3 daysIntensive
100%Hands-on
RealProduction scenarios
Max 10Students per group
[ 01 ]You will learn to solve

The scenarios that don't appear in the documentation

Four classes of real problems worked during the course — not slides, not textbook examples.

01

Critical failures on clusters of 200 TB+

02

Recovery of 40 TB of corrupt CephFS

03

Extreme tuning for AI/ML (500 TB/day)

04

Troubleshooting under 24/7 pressure

[ 02 ]Who it's for
Senior
level

Certified administrators or engineers with real production experience who need to master critical scenarios that vendors don't teach — the ones that show up at 3 AM when the cluster is down.

[ 03 ]Structure · 3 days

Intensive programme for facing real crises

Built to optimise production clusters at petabyte scale. Each day pairs a morning of advanced fundamentals with an afternoon of lab work on real problems.

01
Day 1

Performance Engineering and advanced Forensics

From architecture to real forensic troubleshooting.

Morning

Architectural optimisation

  • BlueStore internals: RocksDB tuning, compaction, write amplification
  • CPU optimisation: C-states impact (5× degradation labs), NUMA
  • Network: 100 GbE patterns, TCP tuning, nf_conntrack
  • NVMe-specific: reactor tuning, bdevs_per_cluster optimisation
Afternoon

Forensic troubleshooting

  • Diagnostic toolchain: blktrace, perf, objectstore-tool
  • Real case studies: NVMe degradation, OSD flapping post-upgrade
  • Advanced PG lifecycle: stuck states, manual intervention
  • Labs: cluster with real problems to diagnose
02
Day 2

Disaster Recovery, Multi-Site and Petabyte Scaling

Extreme recovery and multi-site architectures.

Morning

Advanced DR

  • Edinburgh 40 TB case: full error chain and recovery procedures
  • CephFS disasters: metadata corruption, MDS failure handling
  • RBD mirroring: pool-based vs image-based, failover automation
  • Physical DR: disk extraction, journal and whoami preservation
Afternoon

Multi-Site and Petabytes

  • RGW multisite: master zone failure, manual promotion, sync fairness
  • WAN planning: 1 GbE per 8 TB daily ingest formula
  • Petabyte challenges: CERN 30 PB (7,200 OSDs), 310 M objects
  • Labs: multi-site failover and recovery simulation
03
Day 3

Security, AI/ML Workloads and Cost Engineering

Enterprise security and optimisation for modern workloads.

Morning

Security hardening

  • Encryption: LUKS/dmcrypt OSDs, msgr2 secure, RGW SSE-S3/KMS
  • Key management: rotation (Squid 19.2.3+), Barbican integration
  • Compliance: HIPAA architecture, GDPR, audit logging
  • Threat detection: monitoring patterns and vulnerability management
Afternoon

AI/ML and ROI Engineering

  • S3 Select: Trino integration (2.5×-9× performance), analytics pushdown
  • AI/ML patterns: checkpointing, parallel access optimisation
  • TCO analysis: EC efficiency, commodity hardware savings
  • Hybrid architectures: OpenStack DCN, edge-to-core, multi-cloud
[ 04 ]Lab specifications

Realistic infrastructure on enterprise cloud

Each student works on a real cluster with pre-populated data. Lab access is kept open 7 days after the course to consolidate learning.

Infrastructure

  • Real 5-6 node cluster
  • 500 GB+ pre-populated data per student
  • 24/7 access during course + 7 days after

Real scenarios

  • Disk failures and network partitions
  • Simulated metadata corruption
  • Injected performance degradation

Tooling

  • blktrace, perf, objectstore-tool
  • Debugging symbols pre-installed
  • Real datasets with I/O patterns
[ 05 ]Distribution-agnostic

We work with your distribution and version

We align the lab with what your team runs in production. No vendor bias.

Supported distributions

  • Rocky Linux 9
  • Ubuntu 24.04 LTS
  • Red Hat Enterprise Linux
  • Alma Linux

Ceph versions

  • Ceph upstream Squid 19.2+
  • IBM Storage Ceph 7.1
  • Red Hat Ceph Storage 7.x
[ 06 ]Upcoming sessions

Three formats, small groups

Intensive 3-day training, maximum 10 participants to maximise interaction and collaborative troubleshooting.

Format · 01

On-site at SIXE

At our premises, with full access to labs and specialised equipment.

Format · 02

On-site at your offices

In-company for teams of 4 or more, with a setup customised to the client's environment.

Format · 03

Remote

Live with a dedicated cloud lab and full access to real-time practice resources.

[ 07 ]SIXE certification

Ceph Production Operations badge + EX260 alignment

On completing the course you receive the digital badge Ceph Production Operations issued by SIXE on Credly (Level 3, the most senior of the learning path). The syllabus is aligned with the areas assessed in Red Hat EX260Red Hat Certified Specialist in Ceph Cloud Storage: deployment, replicated and erasure-coded pool management, RBD/RGW/CephFS, troubleshooting, performance tuning, security and high availability/disaster recovery. If you need a specific exam simulation and step-by-step preparation guide, that's covered in the Level 2 (Advanced).

SIXE Ceph Production Operations digital badge on Credly — Level 3 (Expert) certification from the Ceph learning path

Ceph Production Operations · Level 3 · Expert

Issued by SIXE on Credly on course completion. Certifies the ability to operate Ceph clusters at petabyte scale in production: forensics, disaster recovery, multi-site, security hardening and AI/ML workloads. Verifiable credential with a unique link.

View SIXE credentials on Credly
[ 08 ]Beyond training

SIXE is an IBM Business Partner

Beyond training, we sell, implement and support IBM Storage Ceph. If you need to cover the full cycle, count on us.

Licence procurement

Technical advice to size and procure IBM Storage Ceph licences tailored to your organisation's use case.

Production deployment

Design and implementation of IBM Storage Ceph: architecture, hardware, networking, integration with OpenStack, Kubernetes and legacy systems.

Ceph technical support

Ongoing support with specialised response from our engineering team. See service →

[ 09 ]Learning path

This course is Level 3 of the Ceph learning path

Three independent courses. This is the most senior: for teams that already run Ceph and need to master the critical scenarios no vendor teaches.

01 Level 1 · Introduction

Deployment and administration

Fundamentals to working cluster Learn more
02 Level 2 · Advanced

Advanced administration

Tuning, multi-site DR and EX260 Learn more
You are here 03 Level 3 · Expert

Production Operations

Forensics, DR, AI/ML, HIPAA and EX260 See programme

Related courses: OpenStack advanced · Docker and Kubernetes · Ansible. Services: Ceph technical support.

[ 10 ]Frequently asked questions

Frequently asked questions

What student profile is expected?

Certified administrators or engineers with real production experience on Ceph. This is not an introduction or refresher course: we assume prior handling of OSDs, pools, RBD, RGW, CephFS and basic troubleshooting. If you're not there yet, start with the Level 1 or Level 2.

Does the course work with IBM Storage Ceph? And with Ceph upstream?

Yes, with both. The course is distribution-agnostic: it covers IBM Storage Ceph 7.1, Red Hat Ceph Storage 7.x, Ceph upstream Squid 19.2+ and the base distributions Rocky Linux 9, Ubuntu 24.04 LTS, RHEL and Alma Linux. The lab is adapted to the combination your team uses in production.

What is the group size?

Maximum 10 participants per session. It's a design choice: in crisis scenarios and collaborative troubleshooting, larger groups kill interaction. In-company training requires a minimum of 4 people.

How long is the lab access?

24/7 access during the 3 course days and 7 additional days after the last session. The goal is that you can repeat the exercises and consolidate learning while the material is fresh.

Does this course prepare for the Red Hat EX260 certification?

Yes. The syllabus is aligned with the areas assessed in EX260Red Hat Certified Specialist in Ceph Cloud Storage: deployment, replicated and erasure-coded pools, RBD/RGW/CephFS, troubleshooting, performance tuning, security and HA/DR. If in addition to the content you need specific exam simulation and step-by-step preparation guide, that part is in the Level 2 (Advanced). The course does not include the exam fee: it is contracted directly with Red Hat.

Are the lab scenarios real or simulated?

Both. Workloads and datasets with real I/O patterns, and on top of that controlled failures are injected: disk failures, network partitions, metadata corruption and performance degradation. Diagnosis and response are 100 % yours.

Can the syllabus be shortened to 2 days?

It can, but we don't recommend it. The DR + multi-site day needs time to run complete failovers. If you have schedule constraints, let's talk: usually we solve it by keeping the 3 days but compressing a session.

Can you provide 24/7 support for our cluster after the course?

Yes. SIXE is an IBM Business Partner and offers Ceph technical support with specialised response agreements.
[ 11 ]   Request information

Ready to remove the fear of critical scenarios?

Request information about upcoming dates, detailed programme and terms. Reply in under 24 hours.

Or call us directly +34 91 198 02 43
SIXE