Build an AI factory
Plan, build and run a dedicated GPU cluster for training, fine-tuning and inference, treating compute, network, storage, power, cooling, scheduling and day-two operations as one system rather than separate purchases.
The problem
When many teams train, fine-tune and serve models, a collection of separate GPU servers stops scaling. Jobs queue for hardware while other GPUs sit idle, the network slows multi-node work, hardware faults stop long runs, and the building reaches its power and cooling limits before it runs out of floor space.
An AI factory treats these as one design problem: compute sized to the workload, a fabric built for GPU-to-GPU traffic, storage that keeps GPUs fed, a scheduler that shares capacity fairly, and operations software that detects and recovers from faults.
The approach
Starting from a validated reference architecture rather than a blank page saves design work. NVIDIA offers DGX systems up to rack-scale DGX SuperPOD, and DSX reference designs that cover compute, networking, storage and facilities. HGX and NVIDIA-Certified partner servers are other routes to the same GPUs.
The network hardware comes next: Spectrum-X Ethernet (switches paired with SuperNICs) or Quantum InfiniBand for the compute fabric, and BlueField DPUs to offload networking, storage and security. Software then sits on top: Mission Control for provisioning, health checks and automated recovery, Run:ai for quotas and fair sharing, and AI Enterprise for supported AI software. Network designs can be tried in DSX Air before hardware arrives.
Owning a factory makes sense with sustained high utilization and data or sovereignty reasons to keep work in house. For short projects or uncertain demand, cloud GPUs, a GPU cloud provider or a colocation offer avoid facility work and up-front capital.12345678
Conceptual architecture
Diagram as a list
Applications & solutions
- Teams and workloadsTraining, fine-tuning, notebooks and inference servicesConnects to Scheduling and quotas (Run:ai or Slurm)
Operations & orchestration
- Scheduling and quotas (Run:ai or Slurm)Shares GPUs across teams with quotas, priorities and preemptionConnects to GPU systems (DGX or certified partner servers)
- Operations control plane (Mission Control)Provisioning, health checks, automated recovery and power policiesConnects to GPU systems (DGX or certified partner servers), AI compute fabric (Spectrum-X or InfiniBand)
- Power and cooling management (DSX)Plans and manages power within the site budgetConnects to GPU systems (DGX or certified partner servers)
Accelerated computing
- GPU systems (DGX or certified partner servers)Run the AI workloadsConnects to AI compute fabric (Spectrum-X or InfiniBand), High-throughput storage
- High-throughput storageFeeds datasets and stores checkpoints
Networking, power & facilities
- AI compute fabric (Spectrum-X or InfiniBand)Carries GPU-to-GPU traffic for multi-node jobs
- BlueField DPUsOffload networking, storage and security and isolate tenantsConnects to AI compute fabric (Spectrum-X or InfiniBand), High-throughput storage
Technologies and their roles
dgx3
Compute building block
DGX systems and DGX SuperPOD come with NVIDIA software, reference architectures and support as one validated stack.
spectrum-x4
Ethernet AI fabric
Pairs switches with SuperNICs for adaptive routing, congestion control and tenant performance isolation.
mission-control9
Cluster operations
Provisioning, Slurm and Kubernetes scheduling, health checks and automated recovery; required for GB200 and GB300 NVL72 systems.
run-ai10
GPU sharing across teams
Quotas, fair share and priorities so many teams use one pool instead of fixed machines.
dsx2
Facility-level design
Reference designs, simulation, operations software and power management for AI factory buildings.
bluefield6
Infrastructure offload and isolation
Runs networking, storage and security services on the DPU and separates tenants.
What you need first1
- A demand forecast for training, fine-tuning and inference over the next two to three years
- Facility power, cooling and floor loading; DGX GB200 and GB300 systems are liquid-cooled
- Network, storage and cluster operations skills (Slurm or Kubernetes)
- A security and tenancy model for shared access
- Budget for support contracts and software licenses as well as hardware
Risks and how to reduce them
- Low utilization makes owned hardware expensive
- Model demand first, start with a smaller pod and share capacity through quotas and preemption.
- Power and cooling limits at the site
- Check the power envelope and cooling design early and simulate before ordering.
- Delivery timelines and fast hardware generations
- Plan around availability dates NVIDIA and partners publish and avoid designs that depend on unannounced products.
- Hardware faults stall long jobs
- Use health checks, regular checkpoints and automated recovery, and rehearse failure procedures.
- Dependence on a single vendor stack
- Prefer open interfaces such as Kubernetes, Slurm and standards-based Ethernet where they meet requirements.
Related
Sources
- NVIDIA DGX SuperPOD (opens in a new tab)
- NVIDIA DSX AI Factory Platform (opens in a new tab)
- NVIDIA DGX Platform (opens in a new tab)
- NVIDIA Spectrum-X Ethernet Networking Platform (opens in a new tab)
- NVIDIA Quantum InfiniBand (opens in a new tab)
- NVIDIA BlueField Platform (opens in a new tab)
- NVIDIA Mission Control (opens in a new tab)
- NVIDIA DSX documentation (opens in a new tab)
- NVIDIA Mission Control documentation hub (opens in a new tab)
- NVIDIA Run:ai (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.