Skip to content

AI Inference & LLM Optimization

Software that serves trained models at a target latency and cost: Dynamo for distributed generative AI serving, Dynamo-Triton (formerly Triton Inference Server), TensorRT LLM for model optimization, NIM packaging and AIPerf benchmarking.

Technology profiles for this category are in research.

Overview

Once a model is live, every request costs GPU time. Inference teams trade off latency, throughput and GPU count, and large language models add their own problems, such as KV cache memory and long prompts. This category covers the serving and optimization software for that work.

TensorRT LLM speeds up LLM execution with formats such as FP8 and NVFP4, in-flight batching, paged KV caching and speculative decoding. Dynamo serves generative models across many GPUs and nodes: it splits prefill from decode, routes requests by KV cache contents and offloads cache to other memory tiers, and it works with SGLang, TensorRT LLM and vLLM. Dynamo-Triton, formerly Triton Inference Server, serves models from many frameworks and also runs on CPUs. AIPerf measures latency and throughput. NVIDIA says AI Enterprise will include Dynamo in a future release.

The audience is ML platform engineers and teams serving models to many users.

At low traffic, one GPU running an open-source server or a managed API may be enough, and a smaller quantized model often saves more than tuning a large one.1234

Problems it addresses

  • Slow responses on long prompts2

    Prefill and decode compete for the same GPUs. Dynamo separates them so each phase can be tuned on its own.

  • Recomputing the same context2

    Repeated prompts waste compute. Dynamo's KV-aware router sends requests to where cache already exists, and its KV Block Manager offloads cache to other memory tiers.

  • Model does not fit in memory1

    Lower-precision formats in TensorRT LLM, such as FP8, FP4 and INT4 AWQ, reduce memory use; accuracy must be checked after quantizing.

  • Many model types to serve3

    Teams with vision, tabular and speech models need one server. Dynamo-Triton supports TensorRT, PyTorch, ONNX, OpenVINO and Python backends.

  • Guessing capacity4

    Without measurements, GPU counts are guesses. AIPerf reports time to first token, inter-token latency, throughput and goodput.

A typical workflow

  1. Set service targets

    Define time to first token, inter-token latency and throughput targets for each use case.

  2. Optimize the model5

    Quantize and tune with TensorRT LLM, or use a NIM that already bundles an optimized engine.

  3. Choose the serving layer3

    Use Dynamo for multi-node LLM serving and Dynamo-Triton for mixed model types.

  4. Benchmark2

    Run AIPerf against the endpoint with realistic input and output lengths.

  5. Scale on Kubernetes2

    Deploy with Grove for topology-aware scheduling and let the SLO Planner adjust GPU resources.

AI Factory Efficiency Lab

Model token and infrastructure costs for your own numbers.

Open the lab

Next steps

  1. Record the real distribution of prompt and response lengths; it drives every serving decision.

  2. Benchmark your current endpoint with AIPerf to get a baseline before changing anything.

  3. Test one quantized variant of your model against an accuracy set you trust.

  4. Model cost per million tokens for two GPU options in the AI Factory Efficiency Lab.

Sources

  1. NVIDIA TensorRT LLM developer page (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA Dynamo developer page (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA Dynamo-Triton developer page (opens in a new tab)NVIDIA · Vendor-reported
  4. AIPerf (NVIDIA Dynamo project) on GitHub (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA NIM Microservices product page (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)