Skip to content

AI Factories & Infrastructure

Systems and software for building and running large GPU clusters: DGX systems and SuperPOD designs, Mission Control for operations, Run:ai for GPU scheduling, AI Enterprise for the software layer and the DSX platform for designing and powering AI factories.

Technology profiles for this category are in research.

Overview

This category covers what an organization needs once AI grows beyond a few GPUs: racks of accelerated systems, software that schedules work across them, and tools that keep the facility healthy and inside its power budget. NVIDIA calls these sites AI factories.

NVIDIA DGX ranges from the desktop DGX Spark to DGX SuperPOD, a design for on-premises and hybrid clusters. Mission Control schedules workloads, runs health checks, recovers from faults and applies power policies, and it includes Run:ai for GPU orchestration. NVIDIA AI Enterprise supplies the supported software layer. DSX is NVIDIA's platform for designing, simulating and operating AI factories, with parts for reference designs, simulation, operations, power management and grid connection.

The audience is infrastructure teams, GPU cloud providers and enterprises with large, steady AI workloads.

This layer is not needed for occasional fine-tuning or modest inference traffic. Renting GPU instances from a cloud provider, or running a single server, is usually simpler and avoids a long capital commitment.1234

Problems it addresses

  • Idle GPUs5

    Expensive GPUs sit unused when teams hold fixed allocations. Run:ai pools GPUs and uses policy-based priorities so several teams can share them.

  • Slow recovery from failures3

    In long training runs a single faulty node can stall the job. Mission Control runs continuous health checks and has a recovery engine that isolates faults and restarts jobs.

  • Power caps the cluster1

    Many sites are limited by available power rather than floor space. DSX MaxLPS manages power at GPU, rack and workload level within a fixed budget.

  • Design errors found too late1

    Mistakes in cooling, power or network layout are costly once built. DSX Sim and the Omniverse DSX Blueprint let teams test a digital twin of the facility before construction.

A typical workflow

  1. Choose a reference design12

    Start from a validated architecture such as DGX SuperPOD or a DSX Reference Design covering compute, networking, storage and facilities.

  2. Simulate facility and network6

    Model the site in DSX Sim and rehearse network configurations in DSX Air before hardware arrives.

  3. Deploy cluster operations3

    Install Mission Control, which supports Slurm and Kubernetes and includes Run:ai for workload orchestration.

  4. Add the AI software stack4

    Run models and tools from NVIDIA AI Enterprise, which bundles NIM, NeMo and Blueprints with enterprise support.

  5. Operate inside the power budget13

    Apply job-aware power policies in Mission Control, use DSX MaxLPS for power allocation and DSX Flex where the site responds to grid signals.

AI Factory Efficiency Lab

Model token and infrastructure costs for your own numbers.

Open the lab

Next steps

  1. Estimate GPU count, power draw and token cost for your workload in the AI Factory Efficiency Lab before talking to vendors.

  2. Read the DGX SuperPOD and DSX Reference Design material to see which validated architecture matches your scale.

  3. Measure current GPU utilization per team; if it is low, test Run:ai scheduling before buying more hardware.

  4. Compare an on-premises build with reserved cloud GPU capacity for the same multi-year workload.

Sources

  1. NVIDIA DSX AI Factory Platform (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA DGX Platform (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA Mission Control (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA AI Enterprise product page (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA Run:ai (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA DSX Air Platform (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)