Skip to content

GPUs & Accelerated Computing

The processors behind NVIDIA systems: the Blackwell and Vera Rubin GPU architectures, Grace and Vera CPUs, RTX PRO GPUs for mixed AI and graphics work, and the NVLink interconnect that joins GPUs into larger systems.

Technology profiles for this category are in research.

Overview

Buyers often mix up four levels: the architecture (Blackwell, Rubin), the chip (a Rubin GPU, a Vera CPU), the system (HGX B300, DGX Station, Vera Rubin NVL72) and the platform that combines them with networking and software. This category keeps those levels apart.

Blackwell is the architecture inside GB200 and GB300 NVL72 racks and the RTX PRO 6000 Blackwell Server Edition. The Vera Rubin platform pairs Rubin GPUs with Vera CPUs, NVLink 6, ConnectX-9 SuperNICs and BlueField-4, and NVIDIA says it is ramping into full production. Grace is NVIDIA's Arm-based data center CPU used in Grace Hopper and Grace Blackwell designs; NVIDIA lists Vera as a separate CPU line next to it.

RTX PRO servers suit organizations that want AI inference, rendering and video on the same machines. NVL72 racks target large training and inference work.

Many workloads do not need the newest generation. Small models, classical analytics and batch jobs often run well on earlier or rented GPUs, and some run fine on CPUs.1234

Problems it addresses

  • Matching hardware to the job1

    Training large models, serving many users and mixed graphics work need different memory and interconnect. NVIDIA systems range from DGX Spark on a desk to NVL72 racks.

  • Scaling past one server5

    Large models need GPUs that exchange data quickly. NVLink and NVLink Switch connect GPUs across a rack into one domain.

  • Protecting data while it is processed6

    Regulated workloads must keep weights and prompts private at runtime. Confidential computing on Hopper, Blackwell and Rubin GPUs runs work in a trusted execution environment.

  • Wasting a large GPU on small jobs4

    Multi-Instance GPU on the RTX PRO 6000 Blackwell Server Edition splits one GPU into up to four isolated instances.

A typical workflow

  1. Profile the workload

    Record model size, numeric precision, batch size and latency targets before looking at hardware.

  2. Pick the generation27

    Compare Blackwell systems with Vera Rubin systems now ramping, and confirm software support, for example CUDA 13.4 for Rubin.

  3. Choose the system form1

    Decide between desktop systems (DGX Spark, DGX Station), PCIe RTX PRO Servers, HGX boards or NVL72 racks.

  4. Plan the interconnect5

    Size NVLink domains inside the rack and the scale-out network between racks.

  5. Test before buying

    Run your own model with representative traffic on a trial or rented instance of the target GPU.

AI Factory Efficiency Lab

Model token and infrastructure costs for your own numbers.

Open the lab

Next steps

  1. Write down model size, target latency and expected concurrent users; these decide memory and GPU count more than peak FLOPS.

  2. Check that your frameworks and CUDA version support the architecture you plan to buy.

  3. Compare GPU counts and power for two hardware options in the AI Factory Efficiency Lab.

  4. Rent a short-term cloud instance of the target GPU and test your own workload on it.

Sources

  1. NVIDIA Blackwell Architecture (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA Vera Rubin Platform (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA Grace CPU (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA RTX PRO 6000 Blackwell Server Edition (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA NVLink and NVLink Switch (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA Confidential Computing (opens in a new tab)NVIDIA · Vendor-reported
  7. NVIDIA CUDA Toolkit (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)