SupermicroSYS-422GA-NRT-02-G2

Eight NVIDIA RTX PRO 6000 Blackwell cards with 768 GB of GPU memory, two 72-core Xeons, 1.5 TB of DDR5 and 200 and 400 GbE networking, all in 4U and without moving to the HGX form factor.

Supermicro SYS-422GA-NRT-02-G2Rack 4U · Gold Series G2

The exact configuration

Official datasheet

Front view diagram: hover over each component.
Processor2× Intel Xeon 6960P72 cores · 2.7 GHz
GPU8× NVIDIA RTX PRO 6000 Blackwell Server Edition
Memory1.5 TBDDR5-6400
Storage1× 960 GB M.2 NVMe2× 3.84 TB E1.S
Network1× dual 10 GbE RJ451× 2-port 200 GbE/NDR200 + 2× 1-port 400 GbE
Form factorRack 4U
Warranty
3 years parts and labour
On-site engineer
Optional: next business day, for 3 years

Eight GPUs for large models without the jump to HGX

With 768 GB of GPU memory you can fit models of several hundred billion parameters at 8 bits, or several mid-sized models served side by side, which is where many internal platforms end up once every department wants its own. Unlike the -01, this one comes with networking ready to reach beyond the chassis: one 200 GbE card and two 400 GbE cards for reading models and data from shared storage and for joining two or more nodes when the service needs redundancy.

The limit lies in how the GPUs talk to each other. All eight hang off PCIe 5.0 switches, without the NVLink mesh of an HGX B300; for inference that barely matters, but in distributed training, or with a huge model split across all eight, communication between cards becomes the bottleneck. If you plan to train your own models from scratch, look at the SYS-822GS-NB3RT-01-G2 first.

What we do

We build it with RHEL or Ubuntu, the GPU Operator and Kubernetes or OpenShift AI, with vLLM serving each model on its own group of GPUs, and connect the 200 and 400 GbE ports to your storage or to a Ceph cluster we can also set up. If you would rather not depend on the manufacturer for maintenance afterwards, we offer independent support with 24×7 on-call cover under contract.

What your data centre needs

Same 4U chassis and 737 mm depth as the -01. According to the official datasheet it carries four 3,200 W Titanium power supplies in 3+1 redundancy, and with all eight GPUs installed the machine runs much closer to that capacity than its sibling, so circuit layout and rack power are the first things we check. It is air-cooled by ten 8 cm fans, and a rack holding several of these needs cold-aisle containment.

Compared with the rest of the family

ModelWhat it hasWhen to choose it
SYS-822GS-NB3RT-01-G2HGX B300 8-GPU · 2 TBThe HGX B300 with Intel; the choice if your standard is Xeon.
AS-8126GS-NB3RT-01-G2HGX B300 8-GPU · 3 TBSame HGX B300 with AMD, 3 TB and a separate storage network.
AS-5126GS-TNRT-01-G22x RTX PRO 6000 · 1.5 TBTwo GPUs today, space for eight and memory to spare.
SYS-212GB-FNR-01-G22x RTX PRO 6000 · 512 GBThe most compact: 2U for a two-GPU service.
SYS-422GA-NRT-01-G24x RTX PRO 6000 · 1 TBFour GPUs that MIG splits into up to sixteen instances for several teams.
SYS-C542i-11302URTX 50-ready · 32 GBDevelopment tower without a standard GPU; you pick the card.
ARS-511GD-NB-LCC-01-G2Grace + B300 · 748 GBA GB300 node in a tower for teams that fine-tune models without booking a cluster.

See all AI and GPU servers

What the delivery includes

From the factory

  • Configuration assembled and tested by Supermicro
  • Warranty: 3 years parts and labour

From us

  • Rack mounting
  • Remote management (BMC) and up-to-date firmware
  • Operating system and the stack that runs on top
  • Testing with your workload before it goes into production

If you need it

  • 24×7 on-call support with SLA
  • On-site engineer next business day for the full 3 years

Questions about the SYS-422GA-NRT-02-G2

When is an HGX B300 worth it instead of this one?

When you are going to train large models or serve one that does not fit in 768 GB. The HGX B300 has around 2.3 TB of HBM3e and joins its eight GPUs over NVLink, much faster than PCIe for GPU-to-GPU traffic, but it needs 8U and, according to its official datasheet, six 6,600 W supplies against the four 3,200 W units here. For inference on models up to a few hundred billion parameters, this one is usually enough.

Which AI model fits in 768 GB?

Using the rule of about 2 GB per billion parameters at 16 bits and 1 GB at 8 bits, it holds a model of around 300 billion parameters at 16 bits or one of more than 500 billion at 8 bits, with room left for the context cache. We refine the numbers with your actual model and expected number of users.

What are the 200 and 400 GbE ports for?

Getting data in and out of the chassis. The 200 GbE card has two ports, also compatible with InfiniBand NDR (200 Gb/s), and the two 400 GbE cards one each: they read weights and data from shared storage such as Ceph and join two or more nodes for redundancy. The two 10 GbE ports are left for management.

Is the SYS-422GA-NRT-02-G2 right for what you're building?

Tell us about the workload and the data centre it's going into, and we'll come back with the price, a confirmed lead time and what it would take to get it running.