Skip to content

AI Networking & Connectivity

The networks that let many GPUs act together: NVLink inside the rack, Spectrum-X Ethernet and Quantum InfiniBand between racks, BlueField DPUs and ConnectX SuperNICs in each server, and DSX Air for testing a network design in simulation.

Technology profiles for this category are in research.

Overview

In a large AI cluster, GPUs spend much of their time exchanging data. If the fabric is slow or congested, costly accelerators wait. This category covers the networks that move that data and the tools used to design and watch them.

NVLink and NVLink Switch join GPUs inside a rack-scale domain. Between servers NVIDIA offers two fabrics: Spectrum-X Ethernet, which pairs Spectrum switches with SuperNICs for multi-tenant AI clouds, and Quantum InfiniBand, which runs some collective operations inside the switches through SHARP. BlueField DPUs take networking, storage and security work off the host CPU and are programmed with DOCA. DSX Air, formerly NVIDIA Air, creates a virtual copy of the network so teams can test configurations before hardware arrives.

The audience is network architects at cloud providers, research centers and enterprises building multi-rack clusters.

A single server or a small inference setup rarely needs a dedicated AI fabric; ordinary data center Ethernet is usually enough there.12345

Problems it addresses

  • GPUs waiting on the network3

    Collective operations in distributed training stall on a congested fabric. Quantum InfiniBand runs some of these operations in the switch with SHARP, reducing traffic.

  • Ethernet for multi-tenant AI clouds1

    Cloud operators want standard Ethernet tooling with predictable AI performance. Spectrum-X combines Spectrum switches and SuperNICs and supports Cumulus and SONiC.

  • Linking distant data centers1

    Some operators spread one AI workload over several sites. Spectrum-XGS Ethernet targets that case.

  • Host CPUs busy with infrastructure work4

    Networking, storage and security tasks consume CPU cores. BlueField DPUs take over these functions.

  • Risky configuration changes5

    Errors in a live fabric are costly. DSX Air simulates Spectrum-X switches, ConnectX SuperNICs and BlueField DPUs with software such as NetQ.

A typical workflow

  1. Size the scale-up domain2

    Decide how many GPUs share an NVLink domain, for example 8 or 72, based on model size.

  2. Choose the scale-out fabric3

    Select Spectrum-X Ethernet or Quantum InfiniBand according to tenancy, workload type and the skills of your operations team.

  3. Design and simulate5

    Build the topology in DSX Air and test provisioning and automation before go-live.

  4. Offload infrastructure services4

    Use BlueField with DOCA for networking, storage and security services at each node.

  5. Observe in production1

    Watch flows from GPU to SuperNIC with NetQ and fix hotspots.

AI Factory Efficiency Lab

Model token and infrastructure costs for your own numbers.

Open the lab

Next steps

  1. Map your planned GPU count to switch tiers, link counts and cabling, and confirm power with your facility team.

  2. Build a DSX Air simulation of the planned fabric on the free trial and rehearse your automation in it.

  3. Decide early whether your operations staff will run Ethernet or InfiniBand, since tooling and skills differ.

  4. Check the validated network topology for your GPU platform in NVIDIA networking documentation.

Sources

  1. NVIDIA Spectrum-X Ethernet Networking Platform (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA NVLink and NVLink Switch (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA Quantum InfiniBand (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA BlueField Platform (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA DSX Air Platform (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)