Skip to content

NVIDIA Mission Control

NVIDIA Mission Control is operations software for AI factories built on DGX and GB200/GB300 NVL72 systems. It brings cluster provisioning, Slurm and Kubernetes scheduling, health checks, automated recovery, power policies and building management integration into one supported control plane.1

Also known as Mission Control, NVIDIA Mission Control 2.3

At a glance

What is it?
Mission Control is NVIDIA's management platform for running an AI factory. NVIDIA's documentation describes it as combining NVIDIA operational practices and cluster automation into a single control plane. A purchase includes several NVIDIA products through what NVIDIA calls integrated software delivery licensing, with keys or license files generated per product: Base Command Manager for provisioning and cluster management, Run:ai for workload and GPU orchestration, Unified Fabric Manager for InfiniBand, and NetQ for NVLink and Ethernet observability. On top of these, Mission Control adds its own components, such as the autonomous recovery engine, Grafana dashboards, Kubernetes security policy files and an air-gap tool. The current release line on the product page is 2.3, and the documentation lists a 2.3.1 software bill of materials.12
What does it do?
Mission Control turns racked hardware into a schedulable, monitored cluster and keeps it running. It provisions nodes, offers multi-node scheduling with Slurm or Kubernetes, and includes Run:ai for GPU orchestration. Continuous health checks validate hardware and cluster performance and can trigger automated actions based on NVIDIA's preset rules. The autonomous recovery engine detects anomalies, isolates faulty parts, restarts jobs and starts hardware remediation. Administrators can set cluster-wide, job-aware power policies through a domain power service, and building management integration covers power and cooling events such as leak detection.1
Who needs it?
Teams operating NVIDIA DGX B200/B300 or GB200/GB300 NVL72 clusters, for which NVIDIA lists Mission Control as the recommended software. Owners of GB200/GB300 NVL72 systems need it in practice, since NVIDIA's documentation states those systems require Mission Control for installation and management. It also suits AI cloud operators who want to add partner tools for tenant isolation, such as vCluster or Netris, to an NVIDIA-supported base.12
What does it need?2
  • Supported systems: DGX B200/B300 or GB200/GB300 NVL72 (including validated partner NVL72 systems)
  • A Mission Control purchase entitlement (PAK ID) and access to the NVIDIA Licensing Portal
  • A head node for Base Command Manager
  • MAC addresses of the servers that will host UFM and NetQ, for license generation
  • UFM may require additional hardware appliances
  • NVIDIA Resiliency Extension installed separately if used
What it is not
Mission Control is not presented as a general-purpose tool for any vendor's servers: NVIDIA's documentation calls it recommended software for DGX B200/B300 and GB200/GB300 NVL72 systems (required for GB200/GB300 NVL72), and Dell Technologies, HPE and Supermicro have validated it for their NVL72 systems. It is not a separate alternative to Run:ai, because Run:ai is included in the purchase. Not every resilience tool is bundled: the NVIDIA Resiliency Extension must be installed separately. NVIDIA's recovery-speed and power-efficiency figures are vendor statements, not guarantees for a given site, and the Domain Power Service is listed as an early preview.2

Availability and licensing. Commercial software sold as NVIDIA Mission Control with NVIDIA Enterprise Support. The purchase includes Base Command Manager, Run:ai, UFM and NetQ under integrated licensing. The Domain Power Service is offered as an early preview.13

The problem it solves

AI architects and HPC operators have to turn racked hardware into resources that users can safely consume. That means provisioning nodes, setting up schedulers, monitoring hundreds of components, and reacting quickly when a GPU, link or cooling loop fails in the middle of a long training run.

Doing this with separate tools for provisioning, scheduling, fabric management, monitoring and facilities leaves gaps between them. Mission Control packages those functions with NVIDIA support and adds automated recovery and power controls designed for NVIDIA's rack-scale systems.1

How it works

Mission Control is delivered as a set of NVIDIA components installed through a Mission Control workflow:

  • Base Command Manager (BCM) runs on a head node and handles provisioning and cluster management. Additional Mission Control capabilities are delivered through Base Command to licensed customers.
  • Run:ai is installed from the NGC catalog through BCM's Kubernetes setup wizard and handles workload and GPU orchestration. Slurm and bring-your-own Kubernetes are also supported.
  • UFM and NetQ manage and observe InfiniBand, NVLink and Ethernet fabrics.
  • Autonomous recovery engine covers job recovery and hardware recovery with configuration files and runbooks.
  • Grafana visualizations, downloaded from NGC, give dashboards for workloads, infrastructure and facilities.
  • Domain Power Service (early preview) applies job-aware power policies, and building management integration links the cluster to facility systems.2
NVIDIA Mission Control architecture: components by layer and how they connectOperations &orchestrationAcceleratedcomputingNetworking, power &facilitiesBase Command Manager: Provisioning and cluster management from the head nodeBase CommandManagerNVIDIA Run:ai: Kubernetes workload and GPU orchestrationNVIDIA Run:aiSlurm or bring-your-own Kubernetes: Alternative schedulersSlurm orbring-your-own…Autonomous recovery engine: Job and hardware recovery with runbooksAutonomousrecovery engineHealth checks and Grafana dashboards: Validation and monitoringHealth checksand Grafana…DGX B200/B300 and GB200/GB300 NVL72 systems: Managed hardwareDGX B200/B300 andGB200/GB300 NVL72 systemsUFM and NetQ: InfiniBand management and NVLink/Ethernet observabilityUFM and NetQDomain Power Service (early preview): Job-aware cluster power policiesDomain Power Service (earlypreview)Building management integration: Power, cooling and leak events from the facilityBuilding managementintegration
Diagram as a list
  1. Operations & orchestration

    • Base Command ManagerProvisioning and cluster management from the head nodeConnects to DGX B200/B300 and GB200/GB300 NVL72 systems
    • NVIDIA Run:aiKubernetes workload and GPU orchestrationConnects to DGX B200/B300 and GB200/GB300 NVL72 systems
    • Slurm or bring-your-own KubernetesAlternative schedulersConnects to DGX B200/B300 and GB200/GB300 NVL72 systems
    • Autonomous recovery engineJob and hardware recovery with runbooksConnects to NVIDIA Run:ai, Slurm or bring-your-own Kubernetes, DGX B200/B300 and GB200/GB300 NVL72 systems
    • Health checks and Grafana dashboardsValidation and monitoringConnects to Autonomous recovery engine
  2. Accelerated computing

    • DGX B200/B300 and GB200/GB300 NVL72 systemsManaged hardware
  3. Networking, power & facilities

    • UFM and NetQInfiniBand management and NVLink/Ethernet observabilityConnects to DGX B200/B300 and GB200/GB300 NVL72 systems
    • Domain Power Service (early preview)Job-aware cluster power policiesConnects to DGX B200/B300 and GB200/GB300 NVL72 systems
    • Building management integrationPower, cooling and leak events from the facilityConnects to Health checks and Grafana dashboards, Domain Power Service (early preview)
Components and connections as documented by NVIDIA.2

Capabilities

  • Cluster provisioning and management2

    Base Command Manager, included in the purchase, provisions and manages the cluster from a head node.

    Why it matters: One tool brings bare systems to a known configuration.

    Limits: Licensing is activated per cluster from the NVIDIA Licensing Portal.

  • Workload orchestration1

    Includes NVIDIA Run:ai for GPU orchestration and supports Slurm or bring-your-own Kubernetes.

    Why it matters: Training and inference teams can use the scheduler they already know.

    Limits: Run:ai installation needs a token for the installation artifacts.

  • Autonomous recovery engine1

    Covers anomaly detection, isolation, job restart and automated hardware remediation, now available on Blackwell DGX systems.

    Why it matters: Long training runs lose less time to hardware faults.

    Limits: NVIDIA's speed claim is a vendor figure with no stated baseline; actual recovery depends on the fault and the job's checkpointing.

  • Continuous health checks1

    Validates hardware and cluster performance through the infrastructure life cycle, with optional automated actions based on NVIDIA's preset rules.

    Why it matters: Problems are found before a user's job hits them.

    Limits: Automated actions follow NVIDIA's rules, which operators must review for their own policies.

  • Power optimization1

    Administrators can use the domain power service to set cluster-wide, dynamic, job-aware power policies.

    Why it matters: Useful in power-constrained or cost-sensitive sites.

    Limits: The Domain Power Service is listed as an early preview; NVIDIA's power and throughput figures are vendor claims.

  • Building management integration1

    Coordinates power and cooling events with facility systems, including rapid leakage detection; release 2.3 adds leak detection validation checks.

    Why it matters: Links IT operations to the facility, which matters for liquid-cooled racks.

    Limits: Depends on integration with the site's own building management system.

  • Fabric and cluster observability12

    Includes UFM for InfiniBand, NetQ for NVLink and Ethernet visibility, and ready-to-use Grafana dashboards.

    Why it matters: Network and node problems are visible in one place.

    Limits: UFM may need extra hardware appliances.

  • Air-gapped deployment12

    Mission Control 2.3 supports deployment in air-gapped environments, and NVIDIA provides an air-gap tool covering both DGX B200/B300 systems and GB200/GB300 NVL72 racks.

    Why it matters: Needed by sites without internet access, such as some public sector deployments.

    Limits: Software updates must be transferred manually into the isolated environment.

Practical use cases

A new GB200 or GB300 NVL72 deployment must go from racked hardware to a usable cluster.
Approach
Install and activate the stack through the Mission Control workflow for NVL72 systems, including BCM, Run:ai, UFM and NetQ.
Role of NVIDIA Mission Control
Mission Control is the required installation and management path for these systems.
Data, infrastructure and skills
Mission Control entitlement, head node, network plan, NVIDIA Installer Services contact.
Type of benefit
Supported, repeatable bring-up
Caveats
The DGX B200/B300 installation procedures do not apply to NVL72 systems.
First step
Read the release notes and SBOM for the Mission Control version you will install.

Sources 2

Hardware faults interrupt multi-day training jobs and engineers spend hours finding and isolating the failed part.
Approach
Use the autonomous recovery engine and continuous health checks to detect, isolate and restart affected jobs.
Role of NVIDIA Mission Control
Mission Control automates detection, isolation and job restart.
Data, infrastructure and skills
Jobs that checkpoint regularly; review of NVIDIA's preset rules.
Type of benefit
Resilience and less manual triage
Caveats
Recovery speed claims are vendor figures; results depend on checkpoint frequency.
First step
Enable health checks and review which automated actions they will trigger.

Sources 1

A site has a fixed power budget and cannot run every system at full power.
Approach
Apply job-aware, cluster-wide power policies through the domain power service.
Role of NVIDIA Mission Control
Mission Control exposes NVIDIA's power optimization in a validated workflow.
Data, infrastructure and skills
Agreement between IT and facilities on power limits.
Type of benefit
Power efficiency within a fixed budget
Caveats
The Domain Power Service is an early preview.
First step
Measure current cluster power draw per job type before setting policies.

Sources 1

Who uses it

  • Mayo Clinic · Healthcare and medical research

    Mayo Clinic: a DGX SuperPOD with DGX B200 for pathology foundation models

    Mayo Clinic deployed an NVIDIA DGX SuperPOD with DGX B200 systems in July 2025 and says its first work will be building foundation models for pathology, drug discovery and precision medicine. Mayo states the system cuts four weeks of slide analysis and model work to one; no measurement is published.

    In production

Works with

Optional integration

Complementary tools

  • NVIDIA DGXRecommended software for DGX B200/B300 and required for GB200/GB300 NVL72 installation.
  • NVIDIA AI EnterpriseNVIDIA states AI Enterprise is optimized to run on top of Mission Control.
  • NVIDIA Run:aiMission Control bundles Run:ai for workload orchestration.
  • NVIDIA Spectrum-X EthernetMission Control includes NetQ for Ethernet fabric observability.
  • NVIDIA BlueFieldPartner software listed on the Mission Control page (Netris) spans BlueField DPUs for tenant isolation.

Optional integration for

  • NVIDIA DGXMission Control operates DGX clusters and is required for GB200/GB300 NVL72 installation.

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Confirm system support

    Check that your systems are DGX B200/B300, GB200/GB300 NVL72 or a partner NVL72 system validated for Mission Control.

    Check: Your system model appears in the Mission Control documentation or partner list.

  2. Read the release notes and SBOM

    Open the latest release announcement and the 2.3.1 software bill of materials for GB200/GB300 NVL72.

    Check: You know which component versions the release installs.

  3. Activate Base Command Manager

    Generate the BCM product key from the PAK ID in the NVIDIA Licensing Portal, download the ISO, install it on the head node and run request-license.

    Check: The head node shows an active cluster license.

  4. Install Run:ai through BCM

    Request the installation token, then use the BCM Kubernetes setup wizard to download Run:ai from NGC and install it.

    Check: The Run:ai cluster reports as connected in its control plane.

  5. License fabric tools

    Generate UFM and NetQ license files from the Network Entitlements tab using each server's MAC address.

    Check: UFM and NetQ start without license errors.

  6. Add dashboards and resilience tools

    Import the Mission Control Grafana visualizations from NGC and, if needed, install the NVIDIA Resiliency Extension separately.

    Check: Dashboards show live workload and infrastructure data.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA Mission Control as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA Mission Control (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. NVIDIA Mission Control documentation hub (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. NVIDIA Run:ai (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  4. NVIDIA DGX Platform documentation hub (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)