Skip to content

NVIDIA Run:ai

NVIDIA Run:ai is a Kubernetes-based platform that pools GPUs and schedules AI workloads across teams using quotas, priorities and fair sharing, so a shared cluster can serve notebooks, training and inference without each team owning fixed hardware.1

Also known as Run:ai, Runai, Run AI, KAI Scheduler (open-source scheduler based on Run:ai)

At a glance

What is it?
Run:ai is NVIDIA's platform for AI workload and GPU orchestration on Kubernetes. It has two parts: a control plane, which NVIDIA offers either as a managed SaaS service or self-hosted, and a cluster component installed on each Kubernetes cluster, where the Run:ai Scheduler places workloads. The current release on the product page is v2.25. Run:ai is included in NVIDIA AI Enterprise and in NVIDIA Mission Control, and it is listed as a workload management component of DSX OS. NVIDIA also publishes related open-source projects: the KAI Scheduler, which is based on Run:ai, Grove for inference scheduling, and the Run:ai Model Streamer.1234
What does it do?
Run:ai lets administrators map the organization into departments and projects and give each a guaranteed GPU, CPU and memory quota per node pool. Unused capacity can be borrowed by preemptible workloads as over-quota resources and reclaimed when the owner needs it. The scheduler computes a fair share for each project, can balance usage over time with time-based fair share, and supports gang scheduling for distributed jobs. Administrators choose bin-pack or spread placement per node pool. NVIDIA also describes fractional GPU allocation and GPU memory swap for running several models on shared GPUs, and an API-first design for connecting other tools.15
Who needs it?
Organizations where several teams share one GPU cluster and argue over access, or where GPUs are statically assigned and sit idle while other jobs wait. Typical users are platform teams running Kubernetes for data science, research and inference, enterprises consolidating GPU spend, and AI cloud providers that need quotas and governance per customer.1
What does it need?3
  • A Kubernetes cluster: the v2.25 requirements list Kubernetes 1.33 through 1.35 and OpenShift 4.18 through 4.22
  • Supported distributions include vanilla Kubernetes, OpenShift, EKS, GKE, AKS, OKE and RKE2
  • NVIDIA GPU Operator (versions 25.10 to 26.3 per the v2.25 requirements page)
  • Containerd or CRI-O container runtime
  • Ingress controller, FQDN with DNS and TLS certificates when the control plane runs on a separate cluster
  • A namespace named runai
  • Kubernetes 1.32 or later for Multi-Node NVLink systems such as GB200
What it is not
Low GPU utilization is not automatically recoverable waste. Run:ai's own documentation explains that the scheduler cannot always deliver full quotas because of cluster fragmentation, topology constraints, node failures or over-subscription. In our view, idle time caused by slow data loading or very small jobs is also something a scheduler alone cannot fix. Run:ai is not an inference server or model runtime; it schedules workloads such as NVIDIA Dynamo deployments rather than serving models itself. The product page shows availability and utilization multipliers without stating a baseline, so this profile does not repeat them. NVIDIA sells Run:ai through AI Enterprise and Mission Control, and its product page lists KAI Scheduler, Grove and Model Streamer as open source projects from Run:ai.15

Availability and licensing. Commercial platform. NVIDIA states that NVIDIA AI Enterprise now includes Run:ai, and Run:ai is also included with NVIDIA Mission Control. Partners (OEMs, ISVs, cloud partners) offer Run:ai integrations. The product page lists KAI Scheduler, Grove and Model Streamer as open source projects from Run:ai.1

The problem it solves

When GPUs are handed out to teams as fixed allocations, some sit idle while other teams queue. Without central rules, the loudest team gets the hardware, and nobody can see who is using what.

Run:ai treats the cluster as a pool. Each project gets a guaranteed share, idle capacity is lent out to preemptible work, and the scheduler rebalances when owners return. Distributed jobs are placed as a group so they do not start half-scheduled, and policies standardize how workloads are submitted.1

How it works

A user submits a workload (a workspace, a training job or an inference service) through the UI, CLI or API. It is sent to the chosen Kubernetes cluster and handled by the Run:ai Scheduler.

  • Pod groups: every pod belongs to a pod group that represents the whole workload, so gang scheduling, minimum workers, priority and preemptibility apply to the group. Multi-component workloads such as NVIDIA Dynamo (router, prefill, decode) are handled as hierarchical sub-groups.
  • Queues and quotas: a queue exists for each project and department per node pool. Non-preemptible workloads run only within deserved quota; preemptible workloads can use over-quota capacity.
  • Fairness: the scheduler computes a fair share per project and department and shifts resources between queues, which may preempt over-quota work.
  • Placement: per node pool, administrators choose bin-pack (fewer GPUs and nodes used) or spread (more headroom per workload), first across nodes and then across GPU devices.

The cluster depends on the NVIDIA GPU Operator to provision GPUs.35

NVIDIA Run:ai architecture: components by layer and how they connectApplications &solutionsInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingUI, CLI and REST API: Workload submission and administrationUI, CLI and REST APIWorkspaces, training and inference workloads (including Dynamo): What users run on the poolWorkspaces, training andinference workloads…Run:ai control plane (SaaS or self-hosted): UI, API, policies, users and projectsRun:ai controlplane (SaaS or…Run:ai cluster component: Runs on each Kubernetes cluster and connects to the control planeRun:ai clustercomponentRun:ai Scheduler: Queues, quotas, fair share and placementRun:ai SchedulerKubernetes or OpenShift: Container platform the cluster component runs onKubernetes orOpenShiftNVIDIA GPU Operator: Provisions GPU drivers and resources in KubernetesNVIDIA GPUOperatorNode pools of NVIDIA GPUs: The shared GPU capacityNode pools of NVIDIA GPUs
Diagram as a list
  1. Applications & solutions

    • UI, CLI and REST APIWorkload submission and administrationConnects to Run:ai control plane (SaaS or self-hosted)
  2. Inference & runtime software

    • Workspaces, training and inference workloads (including Dynamo)What users run on the poolConnects to Run:ai Scheduler
  3. Operations & orchestration

    • Run:ai control plane (SaaS or self-hosted)UI, API, policies, users and projectsConnects to Run:ai cluster component
    • Run:ai cluster componentRuns on each Kubernetes cluster and connects to the control planeConnects to Run:ai Scheduler
    • Run:ai SchedulerQueues, quotas, fair share and placementConnects to Node pools of NVIDIA GPUs
    • Kubernetes or OpenShiftContainer platform the cluster component runs onConnects to Run:ai cluster component
    • NVIDIA GPU OperatorProvisions GPU drivers and resources in KubernetesConnects to Node pools of NVIDIA GPUs
  4. Accelerated computing

    • Node pools of NVIDIA GPUsThe shared GPU capacity
Components and connections as documented by NVIDIA.6

Capabilities

  • Hierarchical quotas5

    Departments and projects receive deserved GPU, CPU and memory quotas per node pool, with an optional maximum allocation limit.

    Why it matters: Each team gets a guaranteed share of a shared cluster.

    Limits: Quotas can be over-subscribed, in which case not every request within quota can be placed.

  • Over-quota sharing and preemption5

    Projects can use unused node pool capacity as over-quota resources, but only for preemptible workloads, which can be reclaimed by higher-ranked or in-quota work.

    Why it matters: Idle GPUs can be used without taking them away from their owners permanently.

    Limits: Work running over quota can be preempted, so jobs need checkpointing.

  • Fair share and time-based fair share5

    The scheduler computes a fair share per project and department and can balance over-quota use by historical GPU-hour consumption.

    Why it matters: Reduces long-running imbalances between teams.

    Limits: Time-based fair share applies only to over-quota resources and is enabled per node pool.

  • Gang and hierarchical scheduling5

    Workloads are scheduled as pod groups with sub-groups, which covers distributed training and disaggregated inference such as Dynamo.

    Why it matters: Distributed jobs start with all parts placed or not at all.

    Limits: Large gangs may wait longer for enough free capacity.

  • Placement strategies5

    Bin-pack or spread placement can be set per node pool, at both node and GPU-device level.

    Why it matters: Lets administrators favor consolidation or headroom.

    Limits: One strategy per node pool, so mixed needs may require separate pools.

  • Fractional GPUs and memory swap1

    NVIDIA describes fractional allocation of GPUs across inference, embedding and generation tasks and swapping inactive model memory between GPU and host.

    Why it matters: More small models can share fewer GPUs.

    Limits: Shared GPUs trade isolation and latency headroom for density.

  • SaaS or self-hosted deployment2

    The control plane is available as a fully managed, cloud-hosted service or self-hosted for on-premises and private cloud deployments.

    Why it matters: Fits both connected and restricted environments.

    Limits: Self-hosted versions follow your own upgrade cycle.

  • Open-source components1

    NVIDIA publishes the KAI Scheduler (based on Run:ai), Grove for topology-aware inference scheduling, and the Run:ai Model Streamer for faster model loading.

    Why it matters: Smaller teams can adopt the scheduling approach without the full platform.

    Limits: The open-source projects do not include the commercial control plane and UI.

Practical use cases

Several research groups share a GPU cluster, and fixed allocations leave some GPUs idle while others queue.
Approach
Map groups to departments and projects with deserved quotas, allow preemptible over-quota use, and enable fair share.
Role of NVIDIA Run:ai
Run:ai is the scheduler and policy layer for the shared pool.
Data, infrastructure and skills
Kubernetes cluster, agreed quota rules, jobs that tolerate preemption.
Type of benefit
Higher utilization with fair access
Caveats
Preemptible jobs must checkpoint; some idle capacity stems from data or job design, not scheduling.
First step
Export current per-team GPU usage to set initial quotas.

Sources 6

Many small inference models each occupy a full GPU while using a fraction of it.
Approach
Use fractional GPU allocation and GPU memory swap to run several models on shared GPUs.
Role of NVIDIA Run:ai
Run:ai allocates GPU fractions and manages memory swapping.
Data, infrastructure and skills
Models with modest memory needs and tolerance for shared hardware.
Type of benefit
Infrastructure cost efficiency
Caveats
Latency-sensitive services may need dedicated GPUs.
First step
Measure each model's GPU memory and utilization under real traffic.

Sources 1

A disaggregated inference deployment with router, prefill and decode parts must be placed together on the right nodes.
Approach
Schedule the deployment as a hierarchical pod group with gang-scheduled sub-groups.
Role of NVIDIA Run:ai
Run:ai places the whole multi-component workload consistently.
Data, infrastructure and skills
A Dynamo or similar multi-component inference deployment on Kubernetes.
Type of benefit
Consistent placement of multi-part workloads
Caveats
Large gangs may wait for capacity.
First step
Read the Scheduler concepts page on pod groups and sub-groups.

Sources 5

Works with

Optional integration

  • NVIDIA DynamoThe scheduler places Dynamo's router, prefill and decode parts as gang-scheduled sub-groups.

Complementary tools

  • NVIDIA Mission ControlMission Control bundles Run:ai for workload orchestration.
  • NVIDIA DGXRun:ai, included in Mission Control, schedules workloads on DGX GPU pools.
  • NVIDIA DSXListed as a workload management component of DSX OS.

Same family

Alternative approaches

  • NVIDIA AI WorkbenchFor sharing GPU pools across many users and jobs, a scheduler such as Run:ai fits; Workbench manages individual project environments.

Optional integration for

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Choose SaaS or self-hosted

    Use the SaaS control plane for a managed, always-current service, or self-hosted for on-premises and private clouds.

    Check: The matching documentation track (SaaS or Self-hosted) is selected.

  2. Check cluster requirements

    Confirm the Kubernetes or OpenShift version, container runtime, GPU Operator version and ingress setup against the system requirements page.

    Check: Every requirement on the page is met or has a remediation plan.

  3. Install the cluster

    Create the runai namespace and install the cluster component with Helm, then connect it to the control plane.

    Check: The cluster appears as connected in the Run:ai UI.

  4. Map the organization

    Create departments and projects that mirror your teams and assign deserved quotas per node pool.

    Check: Each team can see its project and quota.

  5. Set placement and policies

    Pick bin-pack or spread per node pool and define workload submission policies.

    Check: A test training job and a test workspace run in different projects and respect their quotas.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA Run:ai as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA Run:ai (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. NVIDIA Run:ai documentation hub (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. NVIDIA Run:ai Self-hosted: Cluster System Requirements (v2.25) (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  4. NVIDIA DSX documentation (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  5. The NVIDIA Run:ai Scheduler: Concepts and Principles (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  6. NVIDIA Run:ai Self-hosted documentation (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)