GPUs & Accelerated Computing
The processors behind NVIDIA systems: the Blackwell and Vera Rubin GPU architectures, Grace and Vera CPUs, RTX PRO GPUs for mixed AI and graphics work, and the NVLink interconnect that joins GPUs into larger systems.
Technology profiles for this category are in research.
Overview
Buyers often mix up four levels: the architecture (Blackwell, Rubin), the chip (a Rubin GPU, a Vera CPU), the system (HGX B300, DGX Station, Vera Rubin NVL72) and the platform that combines them with networking and software. This category keeps those levels apart.
Blackwell is the architecture inside GB200 and GB300 NVL72 racks and the RTX PRO 6000 Blackwell Server Edition. The Vera Rubin platform pairs Rubin GPUs with Vera CPUs, NVLink 6, ConnectX-9 SuperNICs and BlueField-4, and NVIDIA says it is ramping into full production. Grace is NVIDIA's Arm-based data center CPU used in Grace Hopper and Grace Blackwell designs; NVIDIA lists Vera as a separate CPU line next to it.
RTX PRO servers suit organizations that want AI inference, rendering and video on the same machines. NVL72 racks target large training and inference work.
Many workloads do not need the newest generation. Small models, classical analytics and batch jobs often run well on earlier or rented GPUs, and some run fine on CPUs.1234
Problems it addresses
Matching hardware to the job1
Training large models, serving many users and mixed graphics work need different memory and interconnect. NVIDIA systems range from DGX Spark on a desk to NVL72 racks.
Scaling past one server5
Large models need GPUs that exchange data quickly. NVLink and NVLink Switch connect GPUs across a rack into one domain.
Protecting data while it is processed6
Regulated workloads must keep weights and prompts private at runtime. Confidential computing on Hopper, Blackwell and Rubin GPUs runs work in a trusted execution environment.
Wasting a large GPU on small jobs4
Multi-Instance GPU on the RTX PRO 6000 Blackwell Server Edition splits one GPU into up to four isolated instances.
A typical workflow
Profile the workload
Record model size, numeric precision, batch size and latency targets before looking at hardware.
Pick the generation27
Compare Blackwell systems with Vera Rubin systems now ramping, and confirm software support, for example CUDA 13.4 for Rubin.
Choose the system form1
Decide between desktop systems (DGX Spark, DGX Station), PCIe RTX PRO Servers, HGX boards or NVL72 racks.
Plan the interconnect5
Size NVLink domains inside the rack and the scale-out network between racks.
Test before buying
Run your own model with representative traffic on a trial or rented instance of the target GPU.
AI Factory Efficiency Lab
Model token and infrastructure costs for your own numbers.
Next steps
Write down model size, target latency and expected concurrent users; these decide memory and GPU count more than peak FLOPS.
Check that your frameworks and CUDA version support the architecture you plan to buy.
Compare GPU counts and power for two hardware options in the AI Factory Efficiency Lab.
Rent a short-term cloud instance of the target GPU and test your own workload on it.
Sources
- NVIDIA Blackwell Architecture (opens in a new tab)
- NVIDIA Vera Rubin Platform (opens in a new tab)
- NVIDIA Grace CPU (opens in a new tab)
- NVIDIA RTX PRO 6000 Blackwell Server Edition (opens in a new tab)
- NVIDIA NVLink and NVLink Switch (opens in a new tab)
- NVIDIA Confidential Computing (opens in a new tab)
- NVIDIA CUDA Toolkit (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.