AI Factories & Infrastructure
Systems and software for building and running large GPU clusters: DGX systems and SuperPOD designs, Mission Control for operations, Run:ai for GPU scheduling, AI Enterprise for the software layer and the DSX platform for designing and powering AI factories.
Technology profiles for this category are in research.
Overview
This category covers what an organization needs once AI grows beyond a few GPUs: racks of accelerated systems, software that schedules work across them, and tools that keep the facility healthy and inside its power budget. NVIDIA calls these sites AI factories.
NVIDIA DGX ranges from the desktop DGX Spark to DGX SuperPOD, a design for on-premises and hybrid clusters. Mission Control schedules workloads, runs health checks, recovers from faults and applies power policies, and it includes Run:ai for GPU orchestration. NVIDIA AI Enterprise supplies the supported software layer. DSX is NVIDIA's platform for designing, simulating and operating AI factories, with parts for reference designs, simulation, operations, power management and grid connection.
The audience is infrastructure teams, GPU cloud providers and enterprises with large, steady AI workloads.
This layer is not needed for occasional fine-tuning or modest inference traffic. Renting GPU instances from a cloud provider, or running a single server, is usually simpler and avoids a long capital commitment.1234
Problems it addresses
Idle GPUs5
Expensive GPUs sit unused when teams hold fixed allocations. Run:ai pools GPUs and uses policy-based priorities so several teams can share them.
Slow recovery from failures3
In long training runs a single faulty node can stall the job. Mission Control runs continuous health checks and has a recovery engine that isolates faults and restarts jobs.
Power caps the cluster1
Many sites are limited by available power rather than floor space. DSX MaxLPS manages power at GPU, rack and workload level within a fixed budget.
Design errors found too late1
Mistakes in cooling, power or network layout are costly once built. DSX Sim and the Omniverse DSX Blueprint let teams test a digital twin of the facility before construction.
A typical workflow
Choose a reference design12
Start from a validated architecture such as DGX SuperPOD or a DSX Reference Design covering compute, networking, storage and facilities.
Simulate facility and network6
Model the site in DSX Sim and rehearse network configurations in DSX Air before hardware arrives.
Deploy cluster operations3
Install Mission Control, which supports Slurm and Kubernetes and includes Run:ai for workload orchestration.
Add the AI software stack4
Run models and tools from NVIDIA AI Enterprise, which bundles NIM, NeMo and Blueprints with enterprise support.
Operate inside the power budget13
Apply job-aware power policies in Mission Control, use DSX MaxLPS for power allocation and DSX Flex where the site responds to grid signals.
AI Factory Efficiency Lab
Model token and infrastructure costs for your own numbers.
Next steps
Estimate GPU count, power draw and token cost for your workload in the AI Factory Efficiency Lab before talking to vendors.
Read the DGX SuperPOD and DSX Reference Design material to see which validated architecture matches your scale.
Measure current GPU utilization per team; if it is low, test Run:ai scheduling before buying more hardware.
Compare an on-premises build with reserved cloud GPU capacity for the same multi-year workload.
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.