SupermicroAS-8126GS-NB3RT-01-G2

The eight-GPU NVIDIA HGX B300 board paired with two AMD EPYC 9575F, 3 TB of DDR5 and eight 800 GbE ports, in the build Supermicro puts together for serving large models to a lot of people at once.

Supermicro AS-8126GS-NB3RT-01-G2Rack 8U · Gold Series G2

The exact configuration

Official datasheet

Front view diagram: hover over each component.
Processor2× AMD EPYC 9575F64 cores · 3.3 GHz
GPUNVIDIA HGX B300 8-GPU
Memory3 TBDDR5-6400
Storage2× 1.9 TB M.2 NVMe
Network2× 2-port 200 GbE / NDR2008-port 800 GbE / XDR800
Form factorRack 8U
Warranty
3 years parts and labour
On-site engineer
Optional: next business day, for 3 years
Software
Supermicro Data Center Management Suite, licence per node

Serving large models to hundreds of users from a single chassis

Supermicro positions it for large-scale inference, and the configuration follows suit. The eight Blackwell Ultra GPUs add up to around 2.3 TB of HBM3e and talk to each other over NVLink (NVIDIA's direct GPU-to-GPU link) through an NVSwitch, so a model with several hundred billion parameters fits entirely inside the machine with room left for the context cache of many simultaneous conversations. The EPYC 9575F is the AMD part that favours clock speed over core count, which shows in the work that still runs on the CPU: queuing requests, tokenising and keeping the GPUs fed so they never sit waiting.

With a single node, any maintenance window takes the model offline, so if it is heading for production the design starts at two machines joined through the eight 800 GbE ports, with the two 200 GbE cards kept for storage. And if the model you plan to serve is around 70 billion parameters, two RTX PRO 6000 cards are enough: a SYS-212GB-FNR-01-G2 in 2U, or an AS-5126GS-TNRT-01-G2 if you want room to grow.

What we do

We deliver it with RHEL or Ubuntu, the NVIDIA drivers and the GPU Operator on Kubernetes or OpenShift AI, plus vLLM as the inference server, tested with the model you are going to run. The two dual-port 200 GbE cards go to storage, usually a Ceph cluster holding the model weights and the data, so that traffic never competes with the GPUs. There is more on how we build it in AI inference on Kubernetes and Ceph.

What your data centre needs

It takes 8U and the chassis is 950 mm deep, more than many 1,000 mm cabinets can take once the cabling is in. According to the official datasheet it carries six 6,600 W Titanium power supplies in 3+3 redundancy, which means six sockets split across two independent circuits if the redundancy is to mean anything. Cooling is by air, with up to twelve high-performance fans, which is why we go through rack power, PDUs and cold-aisle containment with you before anything is signed off.

Compared with the rest of the family

ModelWhat it hasWhen to choose it
SYS-822GS-NB3RT-01-G2HGX B300 8-GPU · 2 TBThe HGX B300 with Intel; the choice if your standard is Xeon.
AS-5126GS-TNRT-01-G22x RTX PRO 6000 · 1.5 TBTwo GPUs today, space for eight and memory to spare.
SYS-212GB-FNR-01-G22x RTX PRO 6000 · 512 GBThe most compact: 2U for a two-GPU service.
SYS-422GA-NRT-01-G24x RTX PRO 6000 · 1 TBFour GPUs that MIG splits into up to sixteen instances for several teams.
SYS-422GA-NRT-02-G28x RTX PRO 6000 · 1.5 TBThe most PCIe can do: eight GPUs and 400 GbE networking.
SYS-C542i-11302URTX 50-ready · 32 GBDevelopment tower without a standard GPU; you pick the card.
ARS-511GD-NB-LCC-01-G2Grace + B300 · 748 GBA GB300 node in a tower for teams that fine-tune models without booking a cluster.

See all AI and GPU servers

What the delivery includes

From the factory

  • Configuration assembled and tested by Supermicro
  • Warranty: 3 years parts and labour

From us

  • Rack mounting
  • Remote management (BMC) and up-to-date firmware
  • Operating system and the stack that runs on top
  • Testing with your workload before it goes into production

If you need it

  • 24×7 on-call support with SLA
  • On-site engineer next business day for the full 3 years

Questions about the AS-8126GS-NB3RT-01-G2

How does it differ from the SYS-822GS-NB3RT-01-G2?

Both use the same eight-GPU HGX B300 board and the same 8U chassis. This one swaps the Xeons for two AMD EPYC 9575F, raises system memory from 2 to 3 TB and adds two dual-port 200 GbE cards on top of the eight 800 GbE ports. If your estate runs on AMD, or you want a separate storage network without adding cards, this is the one.

How large a model fits on the HGX B300?

As a rule of thumb, every billion parameters takes about 2 GB at 16 bits and 1 GB at 8 bits, and on top of that you need memory for the context cache, which grows with conversation length and the number of concurrent users. With around 2.3 TB across the eight GPUs, a model of roughly 700 billion parameters fits at 8 bits with plenty of headroom to serve it widely.

What is each network port for?

The eight 800 GbE ports, which also run as InfiniBand XDR (800 Gb/s), join several GPU nodes into a cluster. The two dual-port 200 GbE cards, also compatible with InfiniBand NDR (200 Gb/s), we normally dedicate to storage, so reading weights and data does not compete with GPU-to-GPU traffic.

Is the AS-8126GS-NB3RT-01-G2 right for what you're building?

Tell us about the workload and the data centre it's going into, and we'll come back with the price, a confirmed lead time and what it would take to get it running.