Skip to content

Build an AI factory

Plan, build and run a dedicated GPU cluster for training, fine-tuning and inference, treating compute, network, storage, power, cooling, scheduling and day-two operations as one system rather than separate purchases.

The problem

When many teams train, fine-tune and serve models, a collection of separate GPU servers stops scaling. Jobs queue for hardware while other GPUs sit idle, the network slows multi-node work, hardware faults stop long runs, and the building reaches its power and cooling limits before it runs out of floor space.

An AI factory treats these as one design problem: compute sized to the workload, a fabric built for GPU-to-GPU traffic, storage that keeps GPUs fed, a scheduler that shares capacity fairly, and operations software that detects and recovers from faults.

The approach

Starting from a validated reference architecture rather than a blank page saves design work. NVIDIA offers DGX systems up to rack-scale DGX SuperPOD, and DSX reference designs that cover compute, networking, storage and facilities. HGX and NVIDIA-Certified partner servers are other routes to the same GPUs.

The network hardware comes next: Spectrum-X Ethernet (switches paired with SuperNICs) or Quantum InfiniBand for the compute fabric, and BlueField DPUs to offload networking, storage and security. Software then sits on top: Mission Control for provisioning, health checks and automated recovery, Run:ai for quotas and fair sharing, and AI Enterprise for supported AI software. Network designs can be tried in DSX Air before hardware arrives.

Owning a factory makes sense with sustained high utilization and data or sovereignty reasons to keep work in house. For short projects or uncertain demand, cloud GPUs, a GPU cloud provider or a colocation offer avoid facility work and up-front capital.12345678

Conceptual architecture

Build an AI factory: conceptual architectureApplications &solutionsOperations &orchestrationAcceleratedcomputingNetworking, power &facilitiesTeams and workloads: Training, fine-tuning, notebooks and inference servicesTeams and workloadsScheduling and quotas (Run:ai or Slurm): Shares GPUs across teams with quotas, priorities and preemptionScheduling and quotas(Run:ai or Slurm)Operations control plane (Mission Control): Provisioning, health checks, automated recovery and power policiesOperations control plane(Mission Control)Power and cooling management (DSX): Plans and manages power within the site budgetPower and cooling management(DSX)GPU systems (DGX or certified partner servers): Run the AI workloadsGPU systems (DGX orcertified partner servers)High-throughput storage: Feeds datasets and stores checkpointsHigh-throughput storageAI compute fabric (Spectrum-X or InfiniBand): Carries GPU-to-GPU traffic for multi-node jobsAI compute fabric(Spectrum-X or InfiniBand)BlueField DPUs: Offload networking, storage and security and isolate tenantsBlueField DPUs
Diagram as a list
  1. Applications & solutions

    • Teams and workloadsTraining, fine-tuning, notebooks and inference servicesConnects to Scheduling and quotas (Run:ai or Slurm)
  2. Operations & orchestration

    • Scheduling and quotas (Run:ai or Slurm)Shares GPUs across teams with quotas, priorities and preemptionConnects to GPU systems (DGX or certified partner servers)
    • Operations control plane (Mission Control)Provisioning, health checks, automated recovery and power policiesConnects to GPU systems (DGX or certified partner servers), AI compute fabric (Spectrum-X or InfiniBand)
    • Power and cooling management (DSX)Plans and manages power within the site budgetConnects to GPU systems (DGX or certified partner servers)
  3. Accelerated computing

    • GPU systems (DGX or certified partner servers)Run the AI workloadsConnects to AI compute fabric (Spectrum-X or InfiniBand), High-throughput storage
    • High-throughput storageFeeds datasets and stores checkpoints
  4. Networking, power & facilities

    • AI compute fabric (Spectrum-X or InfiniBand)Carries GPU-to-GPU traffic for multi-node jobs
    • BlueField DPUsOffload networking, storage and security and isolate tenantsConnects to AI compute fabric (Spectrum-X or InfiniBand), High-throughput storage
Conceptual: one common way to arrange the parts, not a required design.8

Technologies and their roles

  • dgx3

    Compute building block

    DGX systems and DGX SuperPOD come with NVIDIA software, reference architectures and support as one validated stack.

  • spectrum-x4

    Ethernet AI fabric

    Pairs switches with SuperNICs for adaptive routing, congestion control and tenant performance isolation.

  • mission-control9

    Cluster operations

    Provisioning, Slurm and Kubernetes scheduling, health checks and automated recovery; required for GB200 and GB300 NVL72 systems.

  • run-ai10

    GPU sharing across teams

    Quotas, fair share and priorities so many teams use one pool instead of fixed machines.

  • dsx2

    Facility-level design

    Reference designs, simulation, operations software and power management for AI factory buildings.

  • bluefield6

    Infrastructure offload and isolation

    Runs networking, storage and security services on the DPU and separates tenants.

What you need first1

  • A demand forecast for training, fine-tuning and inference over the next two to three years
  • Facility power, cooling and floor loading; DGX GB200 and GB300 systems are liquid-cooled
  • Network, storage and cluster operations skills (Slurm or Kubernetes)
  • A security and tenancy model for shared access
  • Budget for support contracts and software licenses as well as hardware

Risks and how to reduce them

Low utilization makes owned hardware expensive
Model demand first, start with a smaller pod and share capacity through quotas and preemption.
Power and cooling limits at the site
Check the power envelope and cooling design early and simulate before ordering.
Delivery timelines and fast hardware generations
Plan around availability dates NVIDIA and partners publish and avoid designs that depend on unannounced products.
Hardware faults stall long jobs
Use health checks, regular checkpoints and automated recovery, and rehearse failure procedures.
Dependence on a single vendor stack
Prefer open interfaces such as Kubernetes, Slurm and standards-based Ethernet where they meet requirements.

Related

Sources

  1. NVIDIA DGX SuperPOD (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA DSX AI Factory Platform (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA DGX Platform (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA Spectrum-X Ethernet Networking Platform (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA Quantum InfiniBand (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA BlueField Platform (opens in a new tab)NVIDIA · Vendor-reported
  7. NVIDIA Mission Control (opens in a new tab)NVIDIA · Vendor-reported
  8. NVIDIA DSX documentation (opens in a new tab)NVIDIA · Vendor-reported
  9. NVIDIA Mission Control documentation hub (opens in a new tab)NVIDIA · Vendor-reported
  10. NVIDIA Run:ai (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)