Supermicro AI and GPU servers with NVIDIA RTX PRO 6000 Blackwell and HGX B300

Supermicro's Gold Series AI configurations, from a development tower to the eight-GPU HGX B300, which we install in your data centre with the software stack ready for your workload.

Supermicro AS-8126GS-NB3RT-01-G2AS-8126GS-NB3RT-01-G2

Comparison

ModelProcessorGPUMemoryStorageNetworkForm factor
SYS-822GS-NB3RT-01-G22x Intel Xeon 6768P (64C/2.4GHz)NVIDIA HGX B300 8-GPU2TB DDR5-64001x 960GB M.2 NVMe8-port 800GbE / XDR800Rack 8U
AS-8126GS-NB3RT-01-G22x AMD EPYC 9575F (64C/3.3GHz)NVIDIA HGX B300 8-GPU3TB DDR5-64002x 1.9TB M.2 NVMe2x 2-port 200GbE / NDR200 + 8-port 800GbE / XDR800Rack 8U
AS-5126GS-TNRT-01-G22x AMD EPYC 9355 (32C/3.55GHz)2x NVIDIA RTX PRO 6000 Blackwell Server Edition1.5TB DDR5-64001x 960GB M.2 NVMe1x dual 10GbE RJ45Rack 5U
SYS-212GB-FNR-01-G21x Intel Xeon 6731P (32C/2.5GHz)2x NVIDIA RTX PRO 6000 Blackwell Server Edition512GB DDR5-64001x 960GB M.2 NVMe + 2x 3.84TB E1.S1x dual 10GbE RJ45Rack 2U
SYS-422GA-NRT-01-G22x Intel Xeon 6960P (72C/2.7GHz)4x NVIDIA RTX PRO 6000 Blackwell Server Edition1TB DDR5-64001x 960GB M.2 NVMe + 2x 3.84TB E1.S1x dual 10GbE RJ45Rack 4U
SYS-422GA-NRT-02-G22x Intel Xeon 6960P (72C/2.7GHz)8x NVIDIA RTX PRO 6000 Blackwell Server Edition1.5TB DDR5-64001x 960GB M.2 NVMe + 2x 3.84TB E1.S1x dual 10GbE RJ45 + 1x 2-port 200GbE/NDR200 + 2x 1-port 400GbERack 4U
SYS-C542i-11302U1x Intel Core Ultra 7 265 (20C/5.3GHz)Ready for NVIDIA RTX 5070 / 5080 / 5090 (optional)32GB DDR5-56001TB NVMe1GbE, Thunderbolt 4 / USB4, front USB-C 20G, Wi-Fi 7, Bluetooth 5.4Tower
ARS-511GD-NB-LCC-01-G2NVIDIA Grace 72-core Arm Neoverse V2NVIDIA Blackwell Ultra (B300)748GB coherent memory (496GB LPDDR5X ECC + 252GB HBM3e ECC)2x 1.92TB M.2 NVMe + 2x 960GB M.2 NVMe2x QSFP 400GbE, NVIDIA ConnectX-8 SuperNICTower

How to choose an AI server for the model you want to run

The first figure to look at is GPU memory, because the model has to fit in it entirely with room to spare for the context cache. A useful rule: every billion parameters takes about 2 GB at 16 bits and 1 GB at 8 bits, so a 70-billion-parameter model needs about 70 GB just for its weights at 8 bits, and the total then grows with conversation length and the number of concurrent users. Each NVIDIA RTX PRO 6000 Blackwell Server Edition brings 96 GB.

The SYS-212GB-FNR-01-G2 and the AS-5126GS-TNRT-01-G2 ship with two cards and 192 GB, the first in 2U for a specific service and the second in 5U with room for eight; the SYS-422GA-NRT-01-G2 goes up to four cards and 384 GB, and the SYS-422GA-NRT-02-G2 reaches eight and 768 GB, with 400 GbE networking. Above them sit the SYS-822GS-NB3RT-01-G2 and the AS-8126GS-NB3RT-01-G2, both built on the HGX B300: eight GPUs and around 2.3 TB of HBM3e in total.

The other question is whether you will be serving models or training them. For inference, RTX PRO 6000 cards connected over PCIe perform very well and can be split into instances for several teams. For training or fine-tuning large models, the GPUs exchange data constantly, and that is where the HGX B300 comes in, because NVLink connects its eight GPUs far faster than PCIe.

Not everything has to go in a rack, either. The SYS-C542i-11302U is a tower with no GPU as standard, for a developer to try models at their desk with an RTX 5090, and the ARS-511GD-NB-LCC-01-G2, the Super AI Station with NVIDIA GB300, puts 748 GB of memory shared between CPU and GPU at the service of a team developing and fine-tuning models without relying on a cluster.

What your data centre needs

A server with several GPUs of this generation is the most demanding thing you can put in a data centre: lots of high-wattage power supplies per chassis, which need circuits and PDUs sized for them, and concentrated heat that means reviewing the cold aisle and the power available per rack. For reference, according to their official datasheets, an HGX B300 carries six 6,600 W supplies and a SYS-422GA four 3,200 W units. If your data centre cannot take an HGX, we propose the configuration that does fit. Before the order, we go through watts per rack, circuits and cooling with you.

Why buy from us

We deliver the machine installed, and if it ever fails you call the same team that put it into service.

We rack it, configure remote management (BMC) and install RHEL, SUSE, Ubuntu or Proxmox, with whatever runs on top already installed and tested: a Ceph cluster, OpenStack, Kubernetes with the NVIDIA GPU Operator, or Slurm.

When something fails you speak to the people who know that machine because they commissioned it, and if you need it there is 24×7 on-call cover with a contractual SLA. It's the same independent support we provide for IBM Power, Linux and Ceph estates.

If you need different memory, drives or networking, we start from the same base model. Tell us what you need in the custom server configurator and we'll call you back within 24 hours with a proposal.

We are an IBM Business Partner and a Red Hat, SUSE and Canonical partner, which is what counts when the server has to slot into an environment with existing systems and licences.

We work across Europe in all three languages, support included, so if your company has offices in Madrid and Brussels, the same team looks after both.

Frequently asked questions

Which server do I need to run an LLM on-premise?

It depends on the size of the model and how many people will use it. For testing and development, a tower such as the SYS-C542i-11302U with an RTX 5090 is enough; to serve a department with a model of around 70 billion parameters, a machine with two RTX PRO 6000 cards such as the SYS-212GB-FNR-01-G2; and from there you add GPUs up to the eight in the SYS-422GA-NRT-02-G2 or the HGX B300.

RTX PRO 6000 Blackwell or HGX B300?

RTX PRO 6000 Blackwell Server Edition cards connect over PCIe, can be added two at a time up to eight per chassis and cover most inference projects. The HGX B300 joins eight Blackwell Ultra GPUs over NVLink with around 2.3 TB of HBM3e in total and is the choice for training large models or serving the largest ones to many users, at the cost of 8U and a data centre prepared for it.

Can GPUs be added later?

In the RTX PRO 6000 chassis, yes: the AS-5126GS-TNRT-01-G2 and the SYS-422GA models take up to eight double-width GPUs and the SYS-212GB-FNR-01-G2 up to four. We handle the cards, the firmware and revalidating the machine. On the HGX B300 the eight GPUs are fixed to the board, so what you expand is the number of nodes.

Is liquid cooling required?

In this range, only the GB300 Super AI Station uses it, in a closed loop inside the unit itself. The rack servers, including the HGX B300, are air-cooled according to their official datasheets, although that calls for a cold aisle that works and a check of the thermal load per rack.

Which one fits what you're building?

Tell us about the workload and the data centre it's going into, and we'll come back with the price, a confirmed lead time and what it would take to get it running.