Skip to content

NVIDIA Dynamo-Triton

NVIDIA Dynamo-Triton, formerly Triton Inference Server, is open source inference serving software that runs models from many frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL, on GPUs and CPUs behind HTTP/REST and gRPC APIs.1

Also known as NVIDIA Triton Inference Server, Triton Inference Server, Triton, tritonserver

At a glance

What is it?
Dynamo-Triton is the current name of NVIDIA Triton Inference Server, now presented under the Dynamo brand. It is a general-purpose model server: you place models in a model repository, and it loads and serves them through framework backends. It runs on NVIDIA GPUs, non-NVIDIA accelerators, x86 and Arm CPUs and AWS Inferentia, in cloud, data center, edge and embedded settings. The code and containers still use the Triton name, for example the tritonserver container on NGC and the triton-inference-server GitHub organization.123
What does it do?
It serves real-time, batched, ensemble and audio or video streaming requests. Dynamic batching groups incoming requests to keep the hardware busy, concurrent model execution runs several models or copies on the same hardware, and sequence batching keeps state for stateful models. Ensembles and Business Logic Scripting chain models with pre- and post-processing. It exposes HTTP/REST and gRPC endpoints based on the KServe protocol, C and Java APIs for in-process use, and metrics for GPU utilization, throughput and latency.13
Who needs it?
Teams that serve many kinds of models (vision, speech, recommendation, tree-based and language models) from one consistent server; MLOps teams that need Kubernetes scaling and Prometheus monitoring; and edge or embedded projects on Jetson or Windows that want the same serving software as the data center.1
What does it need?12
  • A model repository: one folder per model with its files and configuration
  • Docker and the tritonserver container from NGC (x86 or Arm), or the GitHub binary releases for Windows and Jetson
  • An NVIDIA GPU for accelerated serving; the quickstart also covers CPU-only systems
  • Client libraries or the SDK container to send requests
  • Kubernetes and Prometheus for scaled, monitored deployments (optional)
What it is not
Dynamo-Triton is not the same as Dynamo. NVIDIA positions Dynamo for distributed LLM inference with disaggregated serving and KV-cache features and says it complements Dynamo-Triton; the Dynamo developer page also calls Dynamo the successor to Triton. Dynamo-Triton is not an engine like TensorRT LLM, which it can run through a backend, and not a packaged model container like NIM. After the rename, images, repositories and commands still use the Triton name.145

Availability and licensing. Dynamo-Triton is open source; the server repository on GitHub is licensed BSD-3-Clause, and x86 and Arm containers are available on NGC; the repository also ships an NVIDIA Deep Learning Container License file. NVIDIA offers enterprise support through NVIDIA AI Enterprise, which includes Triton Inference Server for production, with a 90-day evaluation license.12

The problem it solves

Organizations rarely run just one kind of model. A single product might combine an image classifier exported to ONNX, a PyTorch recommender, a tree-based fraud model and a language model. Giving each its own serving stack multiplies the work of scaling, monitoring and securing them, and leaves GPUs underused because each server handles small, separate request streams.

Dynamo-Triton puts these models behind one server with one set of protocols and metrics. Batching and concurrent execution let several models share a GPU, and ensembles let a request pass through several models and processing steps without leaving the server.3

How it works

Dynamo-Triton reads models from a model repository, a folder where each model has its files and a configuration. At runtime:

  • Backends execute each model in its framework, such as TensorRT, PyTorch, ONNX Runtime, OpenVINO, Python or RAPIDS FIL. A Backend API lets teams add custom backends, including Python-based ones.
  • Schedulers apply dynamic batching, sequence batching for stateful models, and concurrent execution of several models or instances.
  • Ensembles and Business Logic Scripting connect models and processing steps into pipelines.
  • Protocols: clients call HTTP/REST or gRPC endpoints based on the KServe protocol, or embed the server through the C or Java API.
  • Metrics report GPU utilization, throughput and latency for Prometheus.

Containers for x86 and Arm come from NGC; Windows and Jetson builds are published on GitHub.12

NVIDIA Dynamo-Triton architecture: components by layer and how they connectApplications &solutionsModels & frameworksInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingClient applications: Send inference requestsClient applicationsModel repository: Holds model files and configurationsModel repositoryHTTP/REST and gRPC (KServe protocol): Request entry points; C and Java APIs for in-process useHTTP/REST and gRPC(KServe protocol)Schedulers and dynamic batcher: Batch requests and run models concurrentlySchedulers and dynamicbatcherEnsembles and Business Logic Scripting: Chain models and processing stepsEnsembles and BusinessLogic ScriptingFramework backends (TensorRT, PyTorch, ONNX, OpenVINO, Python, FIL): Execute each model in its frameworkFramework backends(TensorRT, PyTorch,…Metrics endpoint and Prometheus: GPU utilization, throughput and latency monitoringMetrics endpoint andPrometheusKubernetes: Scales server replicasKubernetesGPUs and CPUs: NVIDIA GPUs, other accelerators, x86 and Arm CPUsGPUs and CPUs
Diagram as a list
  1. Applications & solutions

    • Client applicationsSend inference requestsConnects to HTTP/REST and gRPC (KServe protocol)
  2. Models & frameworks

    • Model repositoryHolds model files and configurationsConnects to Framework backends (TensorRT, PyTorch, ONNX, OpenVINO, Python, FIL)
  3. Inference & runtime software

    • HTTP/REST and gRPC (KServe protocol)Request entry points; C and Java APIs for in-process useConnects to Schedulers and dynamic batcher
    • Schedulers and dynamic batcherBatch requests and run models concurrentlyConnects to Framework backends (TensorRT, PyTorch, ONNX, OpenVINO, Python, FIL), Ensembles and Business Logic Scripting
    • Ensembles and Business Logic ScriptingChain models and processing stepsConnects to Framework backends (TensorRT, PyTorch, ONNX, OpenVINO, Python, FIL)
    • Framework backends (TensorRT, PyTorch, ONNX, OpenVINO, Python, FIL)Execute each model in its frameworkConnects to GPUs and CPUs
  4. Operations & orchestration

    • Metrics endpoint and PrometheusGPU utilization, throughput and latency monitoring
    • KubernetesScales server replicasConnects to HTTP/REST and gRPC (KServe protocol)
  5. Accelerated computing

    • GPUs and CPUsNVIDIA GPUs, other accelerators, x86 and Arm CPUs
Components and connections as documented by NVIDIA.3

Capabilities

  • Multi-framework backends12

    Serves TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL models, with an API for custom backends.

    Why it matters: One server for a mixed model estate instead of one stack per framework.

    Limits: Each backend has its own supported versions and features; check the backend documentation.

  • Dynamic batching and concurrent execution3

    Groups requests into batches and runs multiple models or instances at once on the same hardware.

    Why it matters: Raises hardware utilization for many small request streams.

    Limits: Batching trades a little latency for throughput and must be tuned per model.

  • Ensembles and Business Logic Scripting3

    Chains models and pre- and post-processing steps inside the server.

    Why it matters: Keeps multi-step pipelines close to the GPU and avoids extra network hops.

    Limits: Complex logic in BLS is harder to test than ordinary application code.

  • Standard protocols and embedding3

    HTTP/REST and gRPC based on the KServe protocol, plus C and Java APIs for in-process use.

    Why it matters: Fits existing clients and edge applications that cannot run a separate server.

    Limits: Clients must use these protocols or NVIDIA's client libraries.

  • Metrics and Kubernetes integration13

    Reports GPU utilization, throughput and latency metrics; integrates with Kubernetes for scaling and Prometheus for monitoring.

    Why it matters: Fits existing DevOps and MLOps tooling.

    Limits: Scaling policies are configured in your orchestration layer, not in the server.

  • Broad hardware support3

    Runs on NVIDIA GPUs, non-NVIDIA accelerators, x86 and Arm CPUs and AWS Inferentia, with Windows and Jetson builds.

    Why it matters: The same serving software spans cloud, data center and edge.

    Limits: Feature support can differ between platforms and builds.

Practical use cases

Several computer vision and tabular models each run on their own server and GPUs sit mostly idle.
Approach
Consolidate them into one Dynamo-Triton deployment with dynamic batching and concurrent execution.
Role of NVIDIA Dynamo-Triton
Dynamo-Triton serves all models from one repository with shared hardware.
Data, infrastructure and skills
Models exported to supported formats and per-model configuration.
Type of benefit
Higher GPU utilization
Caveats
Batching settings must be tuned so latency targets still hold.
First step
Move one model into a model repository and compare utilization.

Sources 3

A request needs preprocessing, two models and postprocessing, and network hops add latency.
Approach
Build an ensemble or Business Logic Scripting pipeline inside the server.
Role of NVIDIA Dynamo-Triton
Dynamo-Triton runs the whole pipeline in one process.
Data, infrastructure and skills
Clear interfaces between steps and Python or model-based processing code.
Type of benefit
Simpler architecture
Caveats
Pipeline logic inside the server needs its own tests.
First step
Map the pipeline steps and their inputs and outputs.

Sources 3

A product must run the same models in the data center and on edge devices.
Approach
Use the NGC containers in the data center and the Jetson or Windows builds at the edge.
Role of NVIDIA Dynamo-Triton
Dynamo-Triton provides the same serving interface across sites.
Data, infrastructure and skills
Models compatible with the edge hardware and its backend support.
Type of benefit
Portability
Caveats
Not every backend or feature is available on every platform.
First step
Check which backends your edge build supports.

Sources 1

Works with

Optional integration

  • NVIDIA TensorRT LLMTensorRT LLM's API integrates with Triton, which can serve TensorRT LLM models through a backend.
  • NVIDIA JetsonNVIDIA publishes Triton binary releases for Jetson JetPack on GitHub.
  • NVIDIA CUDA-X Data ScienceThe RAPIDS FIL backend serves tree-based models from the CUDA-X Data Science (RAPIDS) stack.

Complementary tools

Same family

  • NVIDIA DynamoDynamo-Triton (formerly Triton Inference Server) now sits under the Dynamo brand; NVIDIA says Dynamo complements it for LLMs.
  • NVIDIA AI EnterpriseNVIDIA states AI Enterprise includes Triton Inference Server for production.

Optional integration for

  • NVIDIA DeepStream SDKGst-nvinferserver serves models through Dynamo-Triton; 9.1 supports Triton 26.03 on x86 and 26.04 on Jetson.
  • NVIDIA NeMoNeMo Export and Deploy can target Triton backends.

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Create the example model repository2

    Clone the server repository at branch r26.09 and run docs/examples/fetch_models.sh.

    Check: A model_repository folder with the example models exists.

  2. Launch the server2

    Run the nvcr.io/nvidia/tritonserver:26.09-py3 container with --gpus=1, mount model_repository at /models and load densenet_onnx.

    Check: The server starts and reports the model as loaded.

  3. Send a test request2

    Run image_client from the 26.09-py3-sdk container against densenet_onnx with the sample mug image.

    Check: The top result is the coffee mug class, as in NVIDIA's example.

  4. Add your own model

    Place your model and a configuration in the repository, enable dynamic batching and watch the metrics endpoint.

    Check: Metrics show requests being served for your model.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA Dynamo-Triton as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA Dynamo-Triton developer page (opens in a new tab) NVIDIA · Vendor-reported, Recommendation · link checked 9 Oct 2026
  2. Triton Inference Server GitHub repository (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. Triton Inference Server user guide (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  4. NVIDIA Dynamo developer page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  5. TensorRT LLM documentation: overview (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  6. NVIDIA NIM Microservices product page (opens in a new tab) NVIDIA · Recommendation · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)