NVIDIA Dynamo
NVIDIA Dynamo is an open source framework that coordinates generative AI inference across many GPUs and nodes. It sits above engines such as vLLM, SGLang and TensorRT LLM and adds disaggregated serving, KV-cache-aware routing, cache offload and latency-driven autoscaling.1
Also known as Dynamo, ai-dynamo
At a glance
- What is it?
- Dynamo is NVIDIA's open source, distributed inference-serving framework. It sits on top of existing engines such as SGLang, TensorRT LLM and vLLM and coordinates their workers, rather than replacing them. It is written in Rust, with Python for extensibility. NVIDIA announced Dynamo 1.0 as a production release on March 16, 2026, and its developer page calls Dynamo the successor to Triton Inference Server, which is now named Dynamo-Triton.245
- What does it do?
- Dynamo splits the prefill and decode phases of LLM inference into separately scalable GPU pools, routes each request to the worker that already holds the most relevant KV cache, moves KV cache between GPU memory, CPU memory, SSD and remote storage, and resizes worker pools with a planner that targets latency goals. It also includes NIXL for fast data transfer, Grove for topology-aware scheduling on Kubernetes, AIConfigurator for configuration recommendations and AIPerf for benchmarking. Clients use an OpenAI-compatible frontend.12
- Who needs it?
- Teams serving large language or reasoning models across several GPUs or nodes, where throughput, time to first token and GPU cost matter: cloud providers, AI factory operators and platform teams running shared inference on Kubernetes or Slurm. The Dynamo README states that a single model on a single GPU is probably served well enough by the inference engine alone.2
- What does it need?23
- GPUs supported by your chosen engine (Dynamo lists NVIDIA GPUs, AMD GPUs and Intel XPUs)
- One inference engine: vLLM, SGLang or TensorRT LLM
- Kubernetes for production multi-node clusters; Slurm and local runs are also supported
- Docker for the prebuilt runtime containers, or uv and Python for PyPI installs
- Latency targets (time to first token, inter-token latency) for the planner to work toward
- Working knowledge of prefill, decode and KV cache concepts
- What it is not
- Dynamo is not an inference engine and does not compete with TensorRT LLM, vLLM or SGLang; it orchestrates them. It is not the same as NIM, which packages a model and engine as a container, although NVIDIA says NIM will include Dynamo capabilities and the NIM Operator can deploy Dynamo resources experimentally. Dynamo is not part of NVIDIA AI Enterprise today: NVIDIA's Dynamo page says AI Enterprise will include it in a future release. It is not limited to NVIDIA hardware either; its documentation lists NVIDIA and AMD GPUs and Intel XPUs.12356
Availability and licensing. Dynamo is open source on GitHub (the repository shows an Apache 2.0 license badge), with prebuilt runtime containers on NGC and packages on PyPI. NVIDIA's Dynamo page states that NVIDIA AI Enterprise will include Dynamo for production inference in a future release, so it is not part of AI Enterprise today.25
The problem it solves
Large and mixture-of-experts models often exceed the memory of one GPU or even one node, and reasoning models produce many more tokens per request. Serving them with independent engine instances wastes capacity: the compute-heavy prefill phase and the memory-bound decode phase compete for the same GPUs, repeated prompts recompute the same KV cache, and fixed-size deployments end up over or under provisioned.
Dynamo treats the cluster as one inference system. It separates the two phases, tracks where KV cache already lives and scales each pool against service-level targets, with the aim of serving more requests from the same hardware.1
How it works
Requests arrive at the Dynamo frontend, which serves an OpenAI-compatible HTTP API. The router picks a worker using load and KV cache overlap, so requests that share a prompt prefix reuse cached work. In disaggregated mode, prefill workers build the KV cache and the NIXL library transfers it to decode workers, which generate the tokens.
- Planner: profiles the workload and resizes prefill and decode pools to meet latency targets.
- KV Block Manager (KVBM): offloads KV cache from GPU memory to CPU memory, SSD or remote storage; 1.0 added S3 and Azure Blob support.
- Grove: a Kubernetes operator for topology-aware gang scheduling, for example across NVL72 racks.
- ModelExpress: streams weights from GPU to GPU so new replicas start faster.
- Fault tolerance: canary health checks and migration of in-flight requests when a worker fails.
On Kubernetes, Dynamo can own the request entry point itself or sit behind the Gateway API Inference Extension. A DynamoGraphDeploymentRequest (beta in 1.0) lets you declare a model, backend and latency targets and have Dynamo propose and apply a topology.25
Diagram as a list
Applications & solutions
- Client application or agentSends OpenAI-compatible requestsConnects to Dynamo frontend
Inference & runtime software
- Dynamo frontendOpenAI-compatible HTTP entry pointConnects to KV-aware router
- KV-aware routerChooses workers by load and KV cache overlapConnects to Prefill workers (vLLM, SGLang or TensorRT LLM), Decode workers
- Prefill workers (vLLM, SGLang or TensorRT LLM)Process prompts and build KV cacheConnects to NIXL transfer library
- Decode workersGenerate output tokensConnects to GPU cluster
Operations & orchestration
- KV Block ManagerOffloads KV cache to CPU, SSD or remote storage
- PlannerResizes prefill and decode pools against latency targetsConnects to Prefill workers (vLLM, SGLang or TensorRT LLM), Decode workers
- Grove on KubernetesTopology-aware gang scheduling of workersConnects to GPU cluster
Accelerated computing
- GPU clusterNVIDIA or other supported accelerators running the workers
Networking, power & facilities
- NIXL transfer libraryMoves KV cache between GPUs, memory and storageConnects to Decode workers, KV Block Manager
Capabilities
Disaggregated prefill and decode2
Runs the prefill and decode phases in separate GPU pools that scale independently.
Why it matters: Each phase gets hardware and parallelism suited to its workload.
Limits: Adds KV cache transfers between pools, so test it on your own network and model before relying on it.
KV-aware routing2
Routes requests by worker load and KV cache overlap to avoid recomputing shared prefixes.
Why it matters: Helps chat and agent workloads with repeated system prompts or long histories.
Limits: Benefit depends on how much prompt prefix your traffic actually shares.
KV Block Manager2
Offloads KV cache across GPU, CPU, SSD and remote storage tiers.
Why it matters: Extends effective context and cache capacity beyond GPU memory.
Limits: The feature matrix marks KVBM support for SGLang as in progress.
SLA-driven planner2
Profiles workloads and resizes worker pools; NVIDIA says the planner aims to meet latency targets while minimizing total cost of ownership.
Why it matters: Reduces over or under provisioning of GPUs.
Limits: Needs realistic latency targets and profiling data for your model.
Grove and Kubernetes integration2
Grove handles topology-aware gang scheduling; Dynamo also works with the Kubernetes Gateway API Inference Extension.
Why it matters: Places tightly coupled inference components well across racks and nodes.
Limits: The zero-config DynamoGraphDeploymentRequest is beta in 1.0.
NIXL data transfer5
A low-latency library for moving inference data such as KV cache between GPUs and across memory and storage types.
Why it matters: Makes disaggregated serving and cache offload practical.
Limits: Performance depends on the underlying network and storage hardware.
Practical use cases
A reasoning or mixture-of-experts model must serve many users across several nodes at acceptable cost.
- Approach
- Deploy the model with a disaggregated Dynamo recipe so prefill and decode scale separately, with the planner enforcing latency targets.
- Role of NVIDIA Dynamo
- Dynamo orchestrates the engine workers across the cluster.
- Data, infrastructure and skills
- A multi-GPU cluster with fast interconnect, Kubernetes and a supported engine.
- Type of benefit
- Higher throughput per GPU
- Caveats
- Gains depend on model, hardware and traffic; NVIDIA's published figures are vendor benchmarks.
- First step
- Benchmark the engine alone with AIPerf to set a baseline.
Sources 1
Chat and agent sessions resend long, shared prompts, and time to first token is too high.
- Approach
- Enable KV-aware routing so requests land on workers that already hold matching cache.
- Role of NVIDIA Dynamo
- Dynamo's router reuses existing KV cache instead of recomputing it.
- Data, infrastructure and skills
- Traffic with shared prefixes and several engine workers.
- Type of benefit
- Lower latency
- Caveats
- Little effect when prompts rarely overlap.
- First step
- Measure how much prompt prefix your requests share.
Sources 2
A platform team runs shared inference for many teams and keeps over-provisioning GPUs.
- Approach
- Use the Dynamo Platform on Kubernetes with the SLA-based planner to scale pools to demand.
- Role of NVIDIA Dynamo
- Dynamo scales workers against service-level objectives.
- Data, infrastructure and skills
- Kubernetes, latency targets per service and monitoring.
- Type of benefit
- Cost efficiency
- Caveats
- Requires profiling each model and keeping targets realistic.
- First step
- Set time-to-first-token and inter-token latency targets for one service.
Sources 2
Who uses it
Instacart (Maplebear Inc.) · Grocery retail technology
Instacart: Caper smart carts on Jetson and GPU ranking with Dynamo
Instacart runs item recognition on its Caper smart carts with NVIDIA Jetson Orin NX modules and moved online ad and item ranking to NVIDIA GPUs with Dynamo. Published results include 65 percent lower whole-page ranking latency and an incremental sales lift above 1 percent in A/B tests.
Scaling
Works with
Optional integration
- NVIDIA TensorRT LLMTensorRT LLM is one of the three engines Dynamo orchestrates, with vLLM and SGLang.
- NVIDIA NeMoDynamo 1.0 lists integration with the NeMo Agent Toolkit for agentic inference hints.
Complementary tools
- NVIDIA NemotronNVIDIA lists Dynamo among tools for running Nemotron at scale.
Same family
- NVIDIA Dynamo-TritonDynamo-Triton (formerly Triton Inference Server) now sits under the Dynamo brand; NVIDIA says Dynamo complements it for LLMs.
Optional integration for
- NVIDIA NIMThe NIM Operator can deploy Dynamo resources (experimental), and NVIDIA says NIM will include Dynamo capabilities.
- NVIDIA Run:aiThe scheduler places Dynamo's router, prefill and decode parts as gang-scheduled sub-groups.
- NVIDIA DSXDynamo and Grove provide optimized inference in DSX OS.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Run a single-node container23
Pull the Dynamo vLLM runtime container from NGC (the README used version 1.5.1 on October 9, 2026), start dynamo.frontend and a dynamo.vllm worker with a small model such as Qwen3-0.6B.
Check: A curl to localhost:8000/v1/chat/completions returns a reply.
Benchmark the baseline5
Use AIPerf to measure the engine alone and the same model behind Dynamo.
Check: You have throughput and latency figures for both setups.
Install on Kubernetes2
Install the Dynamo Platform with the installation guide, then apply a recipe or a DynamoGraphDeploymentRequest naming model, backend and latency targets.
Check: The deployment is ready and serves the OpenAI-compatible endpoint.
Turn on disaggregation and routing
Switch to a disaggregated recipe for your model and enable KV-aware routing.
Check: A repeat benchmark shows the effect against your baseline.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA Dynamo as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
- NVIDIA Dynamo product page (opens in a new tab)
- NVIDIA Dynamo GitHub repository (opens in a new tab)
- NVIDIA Dynamo documentation (opens in a new tab)
- NVIDIA Newsroom: Dynamo 1.0 production release (opens in a new tab)
- NVIDIA Dynamo developer page (opens in a new tab)
- NVIDIA NIM Operator documentation (opens in a new tab)
- NVIDIA Dynamo-Triton developer page (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.