Reduce inference costs
Lower the GPU cost of each answer a model serves by right-sizing the model, quantizing, batching, reusing cached work and scaling GPU pools to real demand, always checked against latency targets.
The problem
Once a model is in production, GPU hours are usually a large share of the running cost. Costs rise when GPUs sit idle between peaks, when every request recomputes the same long system prompt, when a model runs at higher precision than the task needs, and when reasoning models generate many more tokens per answer.
The hard part is cutting cost without breaking the experience. Time to first token and the gap between tokens are what users notice, so every saving has to be tested against those numbers.1
The approach
Start by measuring the real prompt mix: time to first token, inter-token latency, throughput and GPU use. Then work through the levers, cheapest first.
- Model choice: a smaller model or tier often saves more than tuning a large one.
- Precision: TensorRT LLM supports FP8, FP4, INT4 AWQ and INT8 SmoothQuant quantization; recheck accuracy after each change.
- Batching and caching: in-flight batching and paged KV cache with block reuse keep GPUs busy and reuse shared prompt prefixes.
- Distributed serving: Dynamo separates prefill and decode into independently scaled GPU pools, routes requests to workers that already hold matching KV cache, offloads cache to CPU memory or storage (support varies by engine; check the Dynamo feature matrix) and autoscales to latency targets.
- Sharing GPUs: Run:ai can allocate fractions of a GPU so small models do not each occupy a whole device.
At low or bursty traffic, a pay-per-token managed API can cost less than any GPU you keep running. Self-hosting usually pays off only with steady volume and high utilization, which you should confirm with your own numbers.23456
Conceptual architecture
Diagram as a list
Models & frameworks
- Right-sized, quantized modelRuns at the lowest precision and size that still passes accuracy checksConnects to Shared GPU pool
Inference & runtime software
- KV-aware routerSends each request to the worker with the most reusable cache and the least loadConnects to Prefill worker pool, Decode worker pool
- Prefill worker poolProcesses prompts; scaled separately from decodeConnects to KV cache tiers (GPU, CPU, SSD, remote)
- Decode worker poolGenerates tokens; scaled separately from prefillConnects to KV cache tiers (GPU, CPU, SSD, remote)
Operations & orchestration
- Benchmark harness (AIPerf)Measures latency and throughput on your own prompt mix before and after each changeConnects to KV-aware router
- KV cache tiers (GPU, CPU, SSD, remote)Keeps reusable cache without holding it all in GPU memory; offload support varies by engine (in progress for SGLang)
- SLA-driven plannerAdds or removes workers to meet latency targetsConnects to Prefill worker pool, Decode worker pool
Accelerated computing
- Shared GPU poolRuns the workers; fractional allocation for small models
Technologies and their roles
dynamo15
Distributed serving and autoscaling
Disaggregated prefill and decode, KV-aware routing, cache offload and an SLA-driven planner target cost per token at scale.
tensorrt-llm4
Engine-level efficiency
Quantization formats, in-flight batching, paged KV cache with reuse and speculative decoding reduce GPU time per request.
run-ai6
GPU utilization
Fractional GPU allocation and memory swap let several small models share devices.
nim7
Prebuilt optimized profiles
Model-specific NIMs ship validated quantization profiles selected for the detected GPU.
nemotron8
Smaller model options
Reasoning models come in Nano, Super and Ultra tiers, so a task can be matched to a smaller model.
What you need first
- Samples of production traffic or realistic synthetic prompts
- Agreed latency targets such as time to first token and inter-token latency
- An accuracy test set to check quantized or smaller models
- GPU metrics and the cost of each GPU hour
- Kubernetes skills for disaggregated serving and autoscaling
Risks and how to reduce them
- Quality drops after quantization or switching to a smaller model
- Run the same evaluation set before and after each change and keep a rollback path.
- Published speed-ups do not match your workload
- Treat vendor figures as indicative and benchmark your own prompts on your own hardware.
- Distributed serving adds operational complexity5
- Adopt it only when a single engine instance no longer meets targets.
- Aggressive scale-down causes cold starts
- Pre-cache weights and keep a minimum number of warm replicas.
Related
Sources
- NVIDIA Dynamo product page (opens in a new tab)
- NVIDIA Dynamo developer page (opens in a new tab)
- NVIDIA TensorRT LLM developer page (opens in a new tab)
- TensorRT LLM documentation: overview (opens in a new tab)
- NVIDIA Dynamo GitHub repository (opens in a new tab)
- NVIDIA Run:ai (opens in a new tab)
- NVIDIA NIM for LLM and VLM documentation: overview (opens in a new tab)
- NVIDIA Nemotron foundation models (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.