Skip to content

NVIDIA TensorRT LLM

NVIDIA TensorRT LLM (often written TensorRT-LLM) is an open source library that speeds up large language model and visual generation inference on NVIDIA GPUs, using custom kernels, quantization, in-flight batching, paged KV cache, speculative decoding and multi-GPU parallelism behind a Python LLM API.1

Also known as TensorRT-LLM, TRT-LLM, TRT LLM

At a glance

What is it?
TensorRT LLM is NVIDIA's open source library for accelerating inference of large language models and, more recently, visual generation models on NVIDIA GPUs. It is architected on PyTorch: models are defined in native PyTorch, and a high-level Python LLM API covers setups from one GPU to multiple nodes. It ships the trtllm-serve command for an OpenAI-compatible server and trtllm-bench for benchmarking. NVIDIA's pages now write the name as TensorRT LLM, while the repository and many documents still use TensorRT-LLM.1234
What does it do?
It applies GPU-specific optimizations at the engine level: custom kernels for attention, GEMM and mixture-of-experts layers; in-flight batching and paged attention; a paged KV cache with block reuse; chunked prefill; FP8, FP4, INT4 AWQ and INT8 SmoothQuant quantization; speculative decoding with EAGLE, multi-token prediction and N-gram methods; LoRA adapters; guided decoding; and tensor, pipeline and expert parallelism across GPUs and nodes. Disaggregated serving of the context and generation phases is listed as beta.13
Who needs it?
Inference engineers who want more tokens per second or lower latency from specific NVIDIA GPUs, teams building their own serving stack on an engine, and platform builders who run TensorRT LLM as a backend inside Dynamo, Dynamo-Triton or NIM.15
What does it need?136
  • Linux on x86_64 or aarch64
  • An NVIDIA GPU of a supported architecture: Blackwell, Grace Hopper, Hopper, Ada Lovelace or Ampere (GB200 NVL72 also listed)
  • Ada Lovelace, Hopper or later for FP8 (per the support matrix); Blackwell (B200) for native FP4
  • Docker with the NGC release container, or a pip installation
  • Python skills for the LLM API
  • A Hugging Face model or checkpoint; quantized Model Optimizer checkpoints are optional
What it is not
TensorRT LLM is an engine-level library, not a full serving platform: it does not manage clusters, model catalogs or licenses. It is distinct from TensorRT, NVIDIA's general deep learning inference SDK. It does not compete with Dynamo or NIM: Dynamo orchestrates TensorRT LLM, vLLM or SGLang workers across nodes, and NIM packages engines into containers. In our assessment, the closest alternatives at the same layer are open source engines such as vLLM and SGLang, which Dynamo and NIM also support.278

Availability and licensing. TensorRT LLM is open source on GitHub (the repository shows an Apache 2.0 license badge). Release containers are freely available on NVIDIA NGC, and it can also be installed with pip on Linux.14

The problem it solves

How many users an LLM deployment can serve per GPU depends heavily on how the inference software uses GPU memory and keeps the GPU busy. NVIDIA points to two examples in its documentation: FP8 weights, which NVIDIA states can roughly halve memory use against 16-bit floating point, and in-flight batching, which handles the context and generation phases of different requests together to keep the GPU working.

TensorRT LLM targets this by matching the software to NVIDIA hardware: kernels written for the GPU architecture, low-precision formats such as FP8 and FP4 that cut memory use, batching that keeps the GPU busy while requests arrive and finish at different times, and parallelism that spreads models too large for one GPU across several.3

How it works

Developers work through the LLM API: you pass a Hugging Face model name, a local checkpoint or a quantized checkpoint from TensorRT Model Optimizer, and a single LLM object handles loading, optimization and generation. The same runtime powers trtllm-serve, which exposes OpenAI-compatible endpoints such as /v1/chat/completions.

  • Model definitions are written in native PyTorch, so new architectures can be added or modified directly.
  • The runtime schedules requests with in-flight batching, manages a paged KV cache with block reuse and splits long prompts with chunked prefill.
  • Kernels tuned for NVIDIA GPUs run attention, GEMM and MoE operations, including FP8 on Ada Lovelace, Hopper and later GPUs and native FP4 on Blackwell.
  • Parallelism spreads a model across GPUs and nodes with tensor, pipeline and expert parallelism.

Higher-level systems call this runtime: NVIDIA lists integration with Dynamo and Dynamo-Triton (formerly Triton Inference Server), and NIM can use it as an engine.310

NVIDIA TensorRT LLM architecture: components by layer and how they connectModels & frameworksInference & runtimesoftwareAcceleratedcomputingPyTorch model definitions: Native PyTorch implementations of supported architecturesPyTorch model definitionsCheckpoints (Hugging Face, Model Optimizer quantized): Weights loaded by the LLM APICheckpoints (Hugging Face,Model Optimizer quantized)Serving layer (Dynamo, Dynamo-Triton or NIM): Calls TensorRT LLM as its engineServing layer (Dynamo,Dynamo-Triton or NIM)trtllm-serve: OpenAI-compatible server for direct usetrtllm-serveLLM API (Python): Loads models and runs generationLLM API (Python)Runtime: in-flight batching, KV cache manager, parallelism: Schedules requests and manages memoryRuntime: in-flightbatching, KV cache…Optimized kernels (attention, GEMM, MoE): Run the math on NVIDIA GPUsOptimized kernels(attention, GEMM, MoE)NVIDIA GPUs (Ampere to Blackwell): Execute inferenceNVIDIA GPUs (Ampere toBlackwell)
Diagram as a list
  1. Models & frameworks

    • PyTorch model definitionsNative PyTorch implementations of supported architecturesConnects to Runtime: in-flight batching, KV cache manager, parallelism
    • Checkpoints (Hugging Face, Model Optimizer quantized)Weights loaded by the LLM APIConnects to LLM API (Python)
  2. Inference & runtime software

    • Serving layer (Dynamo, Dynamo-Triton or NIM)Calls TensorRT LLM as its engineConnects to LLM API (Python)
    • trtllm-serveOpenAI-compatible server for direct useConnects to LLM API (Python)
    • LLM API (Python)Loads models and runs generationConnects to Runtime: in-flight batching, KV cache manager, parallelism
    • Runtime: in-flight batching, KV cache manager, parallelismSchedules requests and manages memoryConnects to Optimized kernels (attention, GEMM, MoE)
  3. Accelerated computing

    • Optimized kernels (attention, GEMM, MoE)Run the math on NVIDIA GPUsConnects to NVIDIA GPUs (Ampere to Blackwell)
    • NVIDIA GPUs (Ampere to Blackwell)Execute inference
Components and connections as documented by NVIDIA.3

Capabilities

  • PyTorch-based LLM API36

    A high-level Python API that loads a model by name or path and runs single-GPU to multi-node inference; models are defined in native PyTorch.

    Why it matters: Shortens the path from a Hugging Face model to an optimized deployment and makes custom changes easier.

    Limits: Support depends on the model architecture being in the supported list or added by you.

  • In-flight batching and paged KV cache3

    Processes context and generation phases together as requests come and go, with a paged KV cache and block reuse.

    Why it matters: Keeps GPUs busy and serves more users at once.

    Limits: Gains vary with request mix, sequence lengths and memory headroom.

  • Low-precision quantization136

    Supports FP8, FP4, INT4 AWQ and INT8 SmoothQuant; NVIDIA states FP8 can roughly halve memory use against 16-bit floating point.

    Why it matters: Fits larger models or more users on the same GPUs.

    Limits: The support matrix lists FP8 for Ada Lovelace and Hopper GPUs, and native FP4 needs Blackwell; FP8 is not implemented for every model, and accuracy must be checked per model.

  • Speculative decoding3

    Includes EAGLE, multi-token prediction and N-gram methods to generate several tokens per step.

    Why it matters: Can cut response latency for interactive use.

    Limits: Benefit depends on draft acceptance rates for your prompts and model.

  • Multi-GPU and multi-node parallelism3

    Tensor, pipeline and expert parallelism across GPUs and nodes, with disaggregated serving in beta.

    Why it matters: Serves models too large for one GPU, including large mixture-of-experts models.

    Limits: Requires fast GPU interconnects and careful configuration; disaggregated serving is beta.

  • Serving and benchmarking tools110

    trtllm-serve starts an OpenAI-compatible server; trtllm-bench measures performance; Nsight Systems can profile execution.

    Why it matters: Lets teams test and measure before integrating into a larger stack.

    Limits: trtllm-serve is a single-deployment server; cluster-level routing and scaling come from Dynamo or other systems.

Practical use cases

An interactive assistant on Hopper or Blackwell GPUs responds too slowly and serves too few users per GPU.
Approach
Serve an FP8 or FP4 quantized checkpoint with trtllm-serve and enable speculative decoding.
Role of NVIDIA TensorRT LLM
TensorRT LLM provides the optimized kernels, quantization and decoding methods.
Data, infrastructure and skills
Supported GPUs, a quantized checkpoint and an accuracy test set.
Type of benefit
Lower latency and cost per token
Caveats
Quantization can change outputs; verify accuracy on your own tasks.
First step
Benchmark the 16-bit model with trtllm-bench, then the quantized version.

Sources 10

A large mixture-of-experts model does not fit on one GPU.
Approach
Use tensor, pipeline or expert parallelism across GPUs and nodes through the LLM API.
Role of NVIDIA TensorRT LLM
TensorRT LLM splits the model and coordinates its execution.
Data, infrastructure and skills
Multiple GPUs with fast interconnect and model support in the support matrix.
Type of benefit
Ability to serve larger models
Caveats
Parallel layouts need tuning per model and hardware.
First step
Check the model's deployment guide in the documentation.

Sources 3

A cluster-scale serving platform needs a fast engine under its routing and scaling layer.
Approach
Run TensorRT LLM as the backend for Dynamo workers.
Role of NVIDIA TensorRT LLM
TensorRT LLM executes inference while Dynamo routes and scales.
Data, infrastructure and skills
A Dynamo deployment and the TensorRT LLM runtime container.
Type of benefit
Higher throughput at cluster scale
Caveats
Version compatibility between Dynamo and TensorRT LLM must be checked.
First step
Start from a Dynamo recipe that uses the TensorRT LLM backend.

Sources 7

Works with

Complementary tools

Optional integration for

  • NVIDIA Dynamo-TritonTensorRT LLM's API integrates with Triton, which can serve TensorRT LLM models through a backend.
  • NVIDIA DynamoTensorRT LLM is one of the three engines Dynamo orchestrates, with vLLM and SGLang.
  • NVIDIA NIMNIM can run LLMs on TensorRT LLM, and the NIM Operator can prebuild TensorRT LLM engines.
  • NVIDIA NeMoNeMo Export and Deploy can target TensorRT LLM.
  • NVIDIA NemotronNVIDIA provides cookbooks to serve Nemotron with TensorRT LLM.

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Get the release container10

    Pull the TensorRT LLM release container from NGC, which the quick start names as the fastest option, or install with pip on Linux.

    Check: Inside the container, nvidia-smi shows your GPUs.

  2. Start a server10

    Run trtllm-serve with a small model such as TinyLlama/TinyLlama-1.1B-Chat-v1.0.

    Check: The server listens and accepts requests on its port.

  3. Send a request10

    POST a chat request to /v1/chat/completions with curl or an OpenAI client.

    Check: You receive a chat.completion response with generated text.

  4. Measure1

    Use trtllm-bench and the performance tuning guides to compare settings, such as an FP8 checkpoint on Hopper.

    Check: You have throughput and latency numbers for each configuration.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA TensorRT LLM as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA TensorRT LLM developer page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. TensorRT LLM documentation home (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. TensorRT LLM documentation: overview (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  4. TensorRT LLM GitHub repository (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  5. NVIDIA NIM for Developers (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  6. TensorRT LLM documentation: support matrix (opens in a new tab) NVIDIA · Vendor-reported, Recommendation · link checked 9 Oct 2026
  7. NVIDIA Dynamo GitHub repository (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  8. NVIDIA NIM Microservices product page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  9. NVIDIA Dynamo-Triton developer page (opens in a new tab) NVIDIA · Recommendation · link checked 9 Oct 2026
  10. TensorRT LLM documentation: quick start guide (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)