AI Inference & LLM Optimization
Software that serves trained models at a target latency and cost: Dynamo for distributed generative AI serving, Dynamo-Triton (formerly Triton Inference Server), TensorRT LLM for model optimization, NIM packaging and AIPerf benchmarking.
Technology profiles for this category are in research.
Overview
Once a model is live, every request costs GPU time. Inference teams trade off latency, throughput and GPU count, and large language models add their own problems, such as KV cache memory and long prompts. This category covers the serving and optimization software for that work.
TensorRT LLM speeds up LLM execution with formats such as FP8 and NVFP4, in-flight batching, paged KV caching and speculative decoding. Dynamo serves generative models across many GPUs and nodes: it splits prefill from decode, routes requests by KV cache contents and offloads cache to other memory tiers, and it works with SGLang, TensorRT LLM and vLLM. Dynamo-Triton, formerly Triton Inference Server, serves models from many frameworks and also runs on CPUs. AIPerf measures latency and throughput. NVIDIA says AI Enterprise will include Dynamo in a future release.
The audience is ML platform engineers and teams serving models to many users.
At low traffic, one GPU running an open-source server or a managed API may be enough, and a smaller quantized model often saves more than tuning a large one.1234
Problems it addresses
Slow responses on long prompts2
Prefill and decode compete for the same GPUs. Dynamo separates them so each phase can be tuned on its own.
Recomputing the same context2
Repeated prompts waste compute. Dynamo's KV-aware router sends requests to where cache already exists, and its KV Block Manager offloads cache to other memory tiers.
Model does not fit in memory1
Lower-precision formats in TensorRT LLM, such as FP8, FP4 and INT4 AWQ, reduce memory use; accuracy must be checked after quantizing.
Many model types to serve3
Teams with vision, tabular and speech models need one server. Dynamo-Triton supports TensorRT, PyTorch, ONNX, OpenVINO and Python backends.
Guessing capacity4
Without measurements, GPU counts are guesses. AIPerf reports time to first token, inter-token latency, throughput and goodput.
A typical workflow
Set service targets
Define time to first token, inter-token latency and throughput targets for each use case.
Optimize the model5
Quantize and tune with TensorRT LLM, or use a NIM that already bundles an optimized engine.
Choose the serving layer3
Use Dynamo for multi-node LLM serving and Dynamo-Triton for mixed model types.
Benchmark2
Run AIPerf against the endpoint with realistic input and output lengths.
Scale on Kubernetes2
Deploy with Grove for topology-aware scheduling and let the SLO Planner adjust GPU resources.
AI Factory Efficiency Lab
Model token and infrastructure costs for your own numbers.
Next steps
Record the real distribution of prompt and response lengths; it drives every serving decision.
Benchmark your current endpoint with AIPerf to get a baseline before changing anything.
Test one quantized variant of your model against an accuracy set you trust.
Model cost per million tokens for two GPU options in the AI Factory Efficiency Lab.
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.