Skip to content

Reduce inference costs

Lower the GPU cost of each answer a model serves by right-sizing the model, quantizing, batching, reusing cached work and scaling GPU pools to real demand, always checked against latency targets.

The problem

Once a model is in production, GPU hours are usually a large share of the running cost. Costs rise when GPUs sit idle between peaks, when every request recomputes the same long system prompt, when a model runs at higher precision than the task needs, and when reasoning models generate many more tokens per answer.

The hard part is cutting cost without breaking the experience. Time to first token and the gap between tokens are what users notice, so every saving has to be tested against those numbers.1

The approach

Start by measuring the real prompt mix: time to first token, inter-token latency, throughput and GPU use. Then work through the levers, cheapest first.

  • Model choice: a smaller model or tier often saves more than tuning a large one.
  • Precision: TensorRT LLM supports FP8, FP4, INT4 AWQ and INT8 SmoothQuant quantization; recheck accuracy after each change.
  • Batching and caching: in-flight batching and paged KV cache with block reuse keep GPUs busy and reuse shared prompt prefixes.
  • Distributed serving: Dynamo separates prefill and decode into independently scaled GPU pools, routes requests to workers that already hold matching KV cache, offloads cache to CPU memory or storage (support varies by engine; check the Dynamo feature matrix) and autoscales to latency targets.
  • Sharing GPUs: Run:ai can allocate fractions of a GPU so small models do not each occupy a whole device.

At low or bursty traffic, a pay-per-token managed API can cost less than any GPU you keep running. Self-hosting usually pays off only with steady volume and high utilization, which you should confirm with your own numbers.23456

Conceptual architecture

Reduce inference costs: conceptual architectureModels & frameworksInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingRight-sized, quantized model: Runs at the lowest precision and size that still passes accuracy checksRight-sized, quantized modelKV-aware router: Sends each request to the worker with the most reusable cache and the least loadKV-aware routerPrefill worker pool: Processes prompts; scaled separately from decodePrefill worker poolDecode worker pool: Generates tokens; scaled separately from prefillDecode worker poolBenchmark harness (AIPerf): Measures latency and throughput on your own prompt mix before and after each changeBenchmark harness (AIPerf)KV cache tiers (GPU, CPU, SSD, remote): Keeps reusable cache without holding it all in GPU memory; offload support varies by engine (in progress for SGLang)KV cache tiers (GPU, CPU,SSD, remote)SLA-driven planner: Adds or removes workers to meet latency targetsSLA-driven plannerShared GPU pool: Runs the workers; fractional allocation for small modelsShared GPU pool
Diagram as a list
  1. Models & frameworks

    • Right-sized, quantized modelRuns at the lowest precision and size that still passes accuracy checksConnects to Shared GPU pool
  2. Inference & runtime software

    • KV-aware routerSends each request to the worker with the most reusable cache and the least loadConnects to Prefill worker pool, Decode worker pool
    • Prefill worker poolProcesses prompts; scaled separately from decodeConnects to KV cache tiers (GPU, CPU, SSD, remote)
    • Decode worker poolGenerates tokens; scaled separately from prefillConnects to KV cache tiers (GPU, CPU, SSD, remote)
  3. Operations & orchestration

    • Benchmark harness (AIPerf)Measures latency and throughput on your own prompt mix before and after each changeConnects to KV-aware router
    • KV cache tiers (GPU, CPU, SSD, remote)Keeps reusable cache without holding it all in GPU memory; offload support varies by engine (in progress for SGLang)
    • SLA-driven plannerAdds or removes workers to meet latency targetsConnects to Prefill worker pool, Decode worker pool
  4. Accelerated computing

    • Shared GPU poolRuns the workers; fractional allocation for small models
Conceptual: one common way to arrange the parts, not a required design.5

Technologies and their roles

  • dynamo15

    Distributed serving and autoscaling

    Disaggregated prefill and decode, KV-aware routing, cache offload and an SLA-driven planner target cost per token at scale.

  • tensorrt-llm4

    Engine-level efficiency

    Quantization formats, in-flight batching, paged KV cache with reuse and speculative decoding reduce GPU time per request.

  • run-ai6

    GPU utilization

    Fractional GPU allocation and memory swap let several small models share devices.

  • nim7

    Prebuilt optimized profiles

    Model-specific NIMs ship validated quantization profiles selected for the detected GPU.

  • nemotron8

    Smaller model options

    Reasoning models come in Nano, Super and Ultra tiers, so a task can be matched to a smaller model.

What you need first

  • Samples of production traffic or realistic synthetic prompts
  • Agreed latency targets such as time to first token and inter-token latency
  • An accuracy test set to check quantized or smaller models
  • GPU metrics and the cost of each GPU hour
  • Kubernetes skills for disaggregated serving and autoscaling

Risks and how to reduce them

Quality drops after quantization or switching to a smaller model
Run the same evaluation set before and after each change and keep a rollback path.
Published speed-ups do not match your workload
Treat vendor figures as indicative and benchmark your own prompts on your own hardware.
Distributed serving adds operational complexity5
Adopt it only when a single engine instance no longer meets targets.
Aggressive scale-down causes cold starts
Pre-cache weights and keep a minimum number of warm replicas.

Related

Sources

  1. NVIDIA Dynamo product page (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA Dynamo developer page (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA TensorRT LLM developer page (opens in a new tab)NVIDIA · Vendor-reported
  4. TensorRT LLM documentation: overview (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA Dynamo GitHub repository (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA Run:ai (opens in a new tab)NVIDIA · Vendor-reported
  7. NVIDIA NIM for LLM and VLM documentation: overview (opens in a new tab)NVIDIA · Vendor-reported
  8. NVIDIA Nemotron foundation models (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)