Skip to content

Deploy an LLM

Run an open or fine-tuned large language model as a reliable service on infrastructure you control, behind a standard API, with health checks and scaling, and a path from one test GPU to a production cluster.

The problem

A team has chosen an open model, or fine-tuned one, and needs it running as a dependable service. Model weights alone are not a service: someone must pick an inference engine, match precision and parallelism to the GPUs, expose an API, add health checks and metrics, and keep the stack patched as engines and drivers change.

Constraints vary. Some teams need data to stay on premises, some need predictable latency under load, and some must run in an air-gapped network. Each pushes toward a different level of packaging and orchestration.

The approach

Deployments usually fall into three levels.

  • Engine only: run an open source engine such as vLLM, SGLang or TensorRT LLM directly. TensorRT LLM's trtllm-serve command, for example, starts an OpenAI-compatible server.
  • Packaged container: an LLM NIM bundles the model, engine and API layer, selects a hardware profile at startup and exposes liveness and readiness endpoints.
  • Distributed serving: Dynamo sits above the engines and coordinates serving across many GPUs and nodes.

On Kubernetes, the NIM Operator caches models and runs them as managed services with autoscaling, and Run:ai can share a GPU pool between this service and other teams. For a first test, NVIDIA suggests trying the model on its hosted API before running it locally.

If data may leave your environment and volume is modest, a managed model API from a cloud provider is simpler than any self-hosted option. Dynamo's own documentation notes that a single model on a single GPU probably needs only the engine.12345

Conceptual architecture

Deploy an LLM: conceptual architectureApplications &solutionsModels & frameworksInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingApplication or agent: Sends OpenAI-style chat requestsApplication or agentAPI gateway and authentication: Authenticates callers, applies quotas and routes to model endpointsAPI gateway andauthenticationModel store and cache: Holds weights and fine-tuned adapters close to the clusterModel store and cacheServing container (NIM or engine server): Exposes chat endpoints, health checks and metricsServing container (NIM orengine server)Inference engine (vLLM, SGLang or TensorRT LLM): Runs the model on the GPU with the selected precisionInference engine (vLLM,SGLang or TensorRT LLM)Kubernetes with NIM Operator: Caches models, runs replicas and autoscales on metricsKubernetes with NIM OperatorNVIDIA GPUs: Provide the memory and compute the model profile needsNVIDIA GPUs
Diagram as a list
  1. Applications & solutions

    • Application or agentSends OpenAI-style chat requestsConnects to API gateway and authentication
    • API gateway and authenticationAuthenticates callers, applies quotas and routes to model endpointsConnects to Serving container (NIM or engine server)
  2. Models & frameworks

    • Model store and cacheHolds weights and fine-tuned adapters close to the cluster
  3. Inference & runtime software

    • Serving container (NIM or engine server)Exposes chat endpoints, health checks and metricsConnects to Inference engine (vLLM, SGLang or TensorRT LLM)
    • Inference engine (vLLM, SGLang or TensorRT LLM)Runs the model on the GPU with the selected precisionConnects to NVIDIA GPUs
  4. Operations & orchestration

    • Kubernetes with NIM OperatorCaches models, runs replicas and autoscales on metricsConnects to Serving container (NIM or engine server), Model store and cache
  5. Accelerated computing

    • NVIDIA GPUsProvide the memory and compute the model profile needs
Conceptual: one common way to arrange the parts, not a required design.6

Technologies and their roles

  • nim7

    Packaged serving

    Bundles model, engine and API in one container with hardware-aware profiles and health endpoints, so the team works against a standard API.

  • tensorrt-llm8

    Inference engine

    Open source library with quantization, in-flight batching and an OpenAI-compatible server for teams that want to run the engine directly.

  • dynamo9

    Multi-GPU and multi-node serving

    Coordinates engines across GPUs and nodes when one instance no longer meets demand.

  • run-ai10

    Shared GPU scheduling

    Applies quotas, priorities and fractional GPUs so model services and other work share one cluster.

  • ai-enterprise6711

    Production license and support

    NVIDIA directs production use of self-hosted NIM to an AI Enterprise license, which adds extended-lifetime production branches and enterprise support.

What you need first12

  • Model weights and their license terms, including any fine-tuned adapters
  • GPU memory sized to the model and precision; NVIDIA lists 24 GB as the minimum for Llama 3.1 8B Instruct as a NIM
  • Linux with NVIDIA driver 580 or later, Docker and NVIDIA Container Toolkit; Kubernetes for production
  • Latency and throughput targets for each use of the model
  • Monitoring for request latency, errors and GPU use

Risks and how to reduce them

The model does not fit or runs too slowly on available GPUs
Benchmark the real model and prompt mix first; use a quantized profile or a smaller model.
Engine and security updates fall behind
Pin versions, follow release notes, or use a supported branch with continuous CVE patching.
License mismatch for production use7
Check the model license and the serving software terms; free NIM access covers development and testing only.
Open, unmetered model endpoints
Put authentication, quotas and request logging in front of every model API.

Related

Sources

  1. TensorRT LLM documentation: quick start guide (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA NIM for LLM and VLM: quickstart (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA Dynamo GitHub repository (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA NIM Operator documentation (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA NIM for LLM and VLM: getting started (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA NIM for LLM and VLM documentation: overview (opens in a new tab)NVIDIA · Vendor-reported
  7. NVIDIA NIM Microservices product page (opens in a new tab)NVIDIA · Vendor-reported
  8. NVIDIA TensorRT LLM developer page (opens in a new tab)NVIDIA · Vendor-reported
  9. NVIDIA Dynamo product page (opens in a new tab)NVIDIA · Vendor-reported
  10. NVIDIA Run:ai (opens in a new tab)NVIDIA · Vendor-reported
  11. NVIDIA AI Enterprise product page (opens in a new tab)NVIDIA · Vendor-reported
  12. NVIDIA NIM for LLM and VLM: prerequisites (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)