Results for “Reduce LLM Inference Costs”
15 results
Use case
- Reduce inference costs
Lower the GPU cost of each answer a model serves by right-sizing the model, quantizing, batching, reusing cached work and scaling GPU pools to real demand, always checked against latency targets.
- Deploy an LLM
Run an open or fine-tuned large language model as a reliable service on infrastructure you control, behind a standard API, with health checks and scaling, and a path from one test GPU to a production cluster.
- Logistics optimization
Plan delivery routes, fleet assignments and schedules with mathematical optimization and rerun plans quickly when orders, vehicles or road conditions change, using GPU-accelerated solvers alongside existing planning systems.
- Build an AI factory
Plan, build and run a dedicated GPU cluster for training, fine-tuning and inference, treating compute, network, storage, power, cooling, scheduling and day-two operations as one system rather than separate purchases.
Category
- AI Inference & LLM Optimization
Software that serves trained models at a target latency and cost: Dynamo for distributed generative AI serving, Dynamo-Triton (formerly Triton Inference Server), TensorRT LLM for model optimization, NIM packaging and AIPerf benchmarking.
- Data Science & Analytics
GPU acceleration for dataframes, machine learning, graph analytics and optimization: CUDA-X Data Science (formerly RAPIDS) with cuDF, cuML and cuGraph, the RAPIDS Accelerator for Apache Spark, and the cuOpt decision optimization engine.
Technology
- NVIDIA TensorRT LLM
NVIDIA TensorRT LLM (often written TensorRT-LLM) is an open source library that speeds up large language model and visual generation inference on NVIDIA GPUs, using custom kernels, quantization, in-flight batching, paged KV cache, speculative decoding and multi-GPU parallelism behind a Python LLM API.
- NVIDIA NIM
NVIDIA NIM packages an AI model, an inference engine and its runtime into a container with standard APIs, so teams can self-host models on NVIDIA GPUs in the cloud, a data center, a workstation or at the edge instead of building their own serving stack.
- NVIDIA Dynamo
NVIDIA Dynamo is an open source framework that coordinates generative AI inference across many GPUs and nodes. It sits above engines such as vLLM, SGLang and TensorRT LLM and adds disaggregated serving, KV-cache-aware routing, cache offload and latency-driven autoscaling.
- NVIDIA Dynamo-Triton
NVIDIA Dynamo-Triton, formerly Triton Inference Server, is open source inference serving software that runs models from many frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL, on GPUs and CPUs behind HTTP/REST and gRPC APIs.
- NVIDIA cuOpt
NVIDIA cuOpt is an open source, GPU-accelerated optimization engine for routing, linear programming and quadratic programming, with mixed integer and conic programming in beta. It runs as a Python or C library or as a self-hosted server, and plugs into modeling tools such as AMPL, CVXPY, PuLP, GAMSPy and JuMP.
- NVIDIA Nemotron
NVIDIA Nemotron is NVIDIA's family of open AI models for building agents: reasoning models in several sizes plus models for vision, retrieval, speech and safety. NVIDIA publishes the weights, much of the training data and the training recipes, and the models run on common open inference engines or as NIM.
- NVIDIA Run:ai
NVIDIA Run:ai is a Kubernetes-based platform that pools GPUs and schedules AI workloads across teams using quotas, priorities and fair sharing, so a shared cluster can serve notebooks, training and inference without each team owning fixed hardware.
Case study
- BMW Group: planning car plants in an Omniverse-based Virtual Factory
BMW Group plans and checks its car plants in a Virtual Factory built on NVIDIA Omniverse and OpenUSD. BMW says digital twins now cover more than 30 production sites, cut collision checks for new models from almost four weeks to about three days, and are projected to lower planning costs by up to 30 percent.
- Unilever: product digital twins for marketing imagery with Omniverse and OpenUSD
Unilever builds photoreal 3D twins of its products with NVIDIA Omniverse and OpenUSD, working with creative technology partner Collective World, and renders marketing images from them instead of running repeated photo shoots. Unilever and NVIDIA report imagery made twice as fast at half the cost.
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.