Skip to content

Results for “Reduce LLM Inference Costs”

15 results

Use case

  • Reduce inference costs

    Lower the GPU cost of each answer a model serves by right-sizing the model, quantizing, batching, reusing cached work and scaling GPU pools to real demand, always checked against latency targets.

  • Deploy an LLM

    Run an open or fine-tuned large language model as a reliable service on infrastructure you control, behind a standard API, with health checks and scaling, and a path from one test GPU to a production cluster.

  • Logistics optimization

    Plan delivery routes, fleet assignments and schedules with mathematical optimization and rerun plans quickly when orders, vehicles or road conditions change, using GPU-accelerated solvers alongside existing planning systems.

  • Build an AI factory

    Plan, build and run a dedicated GPU cluster for training, fine-tuning and inference, treating compute, network, storage, power, cooling, scheduling and day-two operations as one system rather than separate purchases.

Category

  • AI Inference & LLM Optimization

    Software that serves trained models at a target latency and cost: Dynamo for distributed generative AI serving, Dynamo-Triton (formerly Triton Inference Server), TensorRT LLM for model optimization, NIM packaging and AIPerf benchmarking.

  • Data Science & Analytics

    GPU acceleration for dataframes, machine learning, graph analytics and optimization: CUDA-X Data Science (formerly RAPIDS) with cuDF, cuML and cuGraph, the RAPIDS Accelerator for Apache Spark, and the cuOpt decision optimization engine.

Technology

  • NVIDIA TensorRT LLM

    NVIDIA TensorRT LLM (often written TensorRT-LLM) is an open source library that speeds up large language model and visual generation inference on NVIDIA GPUs, using custom kernels, quantization, in-flight batching, paged KV cache, speculative decoding and multi-GPU parallelism behind a Python LLM API.

  • NVIDIA NIM

    NVIDIA NIM packages an AI model, an inference engine and its runtime into a container with standard APIs, so teams can self-host models on NVIDIA GPUs in the cloud, a data center, a workstation or at the edge instead of building their own serving stack.

  • NVIDIA Dynamo

    NVIDIA Dynamo is an open source framework that coordinates generative AI inference across many GPUs and nodes. It sits above engines such as vLLM, SGLang and TensorRT LLM and adds disaggregated serving, KV-cache-aware routing, cache offload and latency-driven autoscaling.

  • NVIDIA Dynamo-Triton

    NVIDIA Dynamo-Triton, formerly Triton Inference Server, is open source inference serving software that runs models from many frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL, on GPUs and CPUs behind HTTP/REST and gRPC APIs.

  • NVIDIA cuOpt

    NVIDIA cuOpt is an open source, GPU-accelerated optimization engine for routing, linear programming and quadratic programming, with mixed integer and conic programming in beta. It runs as a Python or C library or as a self-hosted server, and plugs into modeling tools such as AMPL, CVXPY, PuLP, GAMSPy and JuMP.

  • NVIDIA Nemotron

    NVIDIA Nemotron is NVIDIA's family of open AI models for building agents: reasoning models in several sizes plus models for vision, retrieval, speech and safety. NVIDIA publishes the weights, much of the training data and the training recipes, and the models run on common open inference engines or as NIM.

  • NVIDIA Run:ai

    NVIDIA Run:ai is a Kubernetes-based platform that pools GPUs and schedules AI workloads across teams using quotas, priorities and fair sharing, so a shared cluster can serve notebooks, training and inference without each team owning fixed hardware.

Case study

  • BMW Group: planning car plants in an Omniverse-based Virtual Factory

    BMW Group plans and checks its car plants in a Virtual Factory built on NVIDIA Omniverse and OpenUSD. BMW says digital twins now cover more than 30 production sites, cut collision checks for new models from almost four weeks to about three days, and are projected to lower planning costs by up to 30 percent.

  • Unilever: product digital twins for marketing imagery with Omniverse and OpenUSD

    Unilever builds photoreal 3D twins of its products with NVIDIA Omniverse and OpenUSD, working with creative technology partner Collective World, and renders marketing images from them instead of running repeated photo shoots. Unilever and NVIDIA report imagery made twice as fast at half the cost.

Fill out the form below to request your copy.

Name(Required)