SupermicroSYS-422GA-NRT-01-G2

Four NVIDIA RTX PRO 6000 Blackwell cards with 384 GB of GPU memory between them, two 72-core Xeon 6960P and 1 TB of DDR5 in a 4U chassis with room for four more.

Supermicro SYS-422GA-NRT-01-G2Rack 4U · Gold Series G2

The exact configuration

Official datasheet

Front view diagram: hover over each component.
Processor2× Intel Xeon 6960P72 cores · 2.7 GHz
GPU4× NVIDIA RTX PRO 6000 Blackwell Server Edition
Memory1 TBDDR5-6400
Storage1× 960 GB M.2 NVMe2× 3.84 TB E1.S
Network1× dual 10 GbE RJ45
Form factorRack 4U
Warranty
3 years parts and labour
On-site engineer
Optional: next business day, for 3 years

GPUs shared between several teams and services

Supermicro aims this family at shared inference platforms, and that is where it performs best: one server drawn on at the same time by the data team, the people building agents and a couple of production services. Each RTX PRO 6000 can be split with MIG (NVIDIA's partitioning of a GPU into isolated instances) into up to four slices, each with its own memory, so the machine can run as sixteen small GPUs or four large ones depending on what each workload needs. The 144 cores of the two Xeon 6960P are more than enough for preprocessing, vector databases and the other services that sit around a model.

The four GPUs communicate through a PCIe 5.0 switch, without the NVLink fabric of the HGX machines, and that sets its ceiling. Serving a model split across the four works well; for training a large model, where the GPUs exchange data constantly, the answer is an HGX B300, because NVLink moves data between GPUs far faster than PCIe. Standard networking is 10 GbE, enough if the models live on the local E1.S drives and too little if you load them from shared storage.

What we do

We integrate it into OpenShift AI or Kubernetes with the NVIDIA GPU Operator, which hands out the MIG instances between projects and lets each team request the GPU it needs without touching the machine. On top we deploy vLLM for production models, plus the notebooks and data pipelines the data team uses. The full design is in AI inference on Kubernetes and Ceph.

What your data centre needs

It is 4U and 737 mm deep, so it fits almost any cabinet. The official datasheet lists four 3,200 W Titanium power supplies in 3+1 redundancy, the same as the eight-GPU version, so if you expand later the power side is already sized and what needs planning from the outset is sockets and PDU. It is air-cooled by ten 8 cm fans.

Compared with the rest of the family

ModelWhat it hasWhen to choose it
SYS-822GS-NB3RT-01-G2HGX B300 8-GPU · 2 TBThe HGX B300 with Intel; the choice if your standard is Xeon.
AS-8126GS-NB3RT-01-G2HGX B300 8-GPU · 3 TBSame HGX B300 with AMD, 3 TB and a separate storage network.
AS-5126GS-TNRT-01-G22x RTX PRO 6000 · 1.5 TBTwo GPUs today, space for eight and memory to spare.
SYS-212GB-FNR-01-G22x RTX PRO 6000 · 512 GBThe most compact: 2U for a two-GPU service.
SYS-422GA-NRT-02-G28x RTX PRO 6000 · 1.5 TBThe most PCIe can do: eight GPUs and 400 GbE networking.
SYS-C542i-11302URTX 50-ready · 32 GBDevelopment tower without a standard GPU; you pick the card.
ARS-511GD-NB-LCC-01-G2Grace + B300 · 748 GBA GB300 node in a tower for teams that fine-tune models without booking a cluster.

See all AI and GPU servers

What the delivery includes

From the factory

  • Configuration assembled and tested by Supermicro
  • Warranty: 3 years parts and labour

From us

  • Rack mounting
  • Remote management (BMC) and up-to-date firmware
  • Operating system and the stack that runs on top
  • Testing with your workload before it goes into production

If you need it

  • 24×7 on-call support with SLA
  • On-site engineer next business day for the full 3 years

Questions about the SYS-422GA-NRT-01-G2

What is the difference between the SYS-422GA-NRT-01-G2 and the -02?

They share the chassis and the Xeon 6900 series processors. The -01 has four RTX PRO 6000 cards, 1 TB of memory and 10 GbE networking only; the -02 fills the chassis with eight GPUs, goes up to 1.5 TB and adds one 200 GbE card and two 400 GbE cards for storage and for joining nodes. The -01 suits a shared platform that will grow over time, the -02 a model or user count that already calls for eight GPUs.

Can one GPU be shared between several users?

Yes. With MIG each RTX PRO 6000 splits into up to four isolated instances, each with its own memory and cores, so what one team runs does not slow down the next. In Kubernetes or OpenShift AI those instances are requested like any other resource, and we configure them around the size of the models each group will use.

Can I expand it to eight GPUs?

Yes. The chassis takes up to eight double-width GPUs and its four 3,200 W supplies in 3+1 redundancy are the same ones the -02 uses with eight, so expanding means adding the cards, updating the firmware and revalidating with your workload. What you should review is the 10 GbE networking, because with eight GPUs you will usually need 200 or 400 GbE.

Is the SYS-422GA-NRT-01-G2 right for what you're building?

Tell us about the workload and the data centre it's going into, and we'll come back with the price, a confirmed lead time and what it would take to get it running.