NVIDIA NIM
NVIDIA NIM packages an AI model, an inference engine and its runtime into a container with standard APIs, so teams can self-host models on NVIDIA GPUs in the cloud, a data center, a workstation or at the edge instead of building their own serving stack.1
Also known as NIM microservices, NVIDIA inference microservices, NIM for LLM and VLM, NIM LLM
At a glance
- What is it?
- NVIDIA NIM is a set of prebuilt inference microservices. Each NIM is a container that holds a model (or, for model-free NIMs, loads one you choose at startup), an inference engine and the runtime pieces around it, and serves the model through industry-standard APIs such as OpenAI-compatible endpoints. NVIDIA documents NIM as part of NVIDIA AI Enterprise. The current NIM for LLM and VLM line (version 2.0) uses one backend per container and is built on the open source vLLM engine; its quickstart also refers to SGLang images.1245
- What does it do?
- A NIM starts a model server with a single docker run command, checks which GPUs are present and selects a matching model profile, downloads and caches the weights, and then answers chat completion, text completion and responses requests on a local HTTP port, with streaming. It exposes liveness and readiness endpoints and observability metrics and can load LoRA adapters. On Kubernetes, the NIM Operator adds custom resources for caching models, running NIM services, grouping them into pipelines and autoscaling them.3456
- Who needs it?
- Teams that want to run open or fine-tuned models on their own NVIDIA infrastructure, keep prompts and data inside their environment, and avoid assembling and patching a serving stack. NVIDIA's getting-started guide addresses AI and ML engineers testing pipelines, platform operators validating infrastructure and evaluators exploring capabilities. It also fits organizations that need long-lived, patched production branches, which NVIDIA ties to an AI Enterprise license.147
- What does it need?8
- An NVIDIA GPU with enough memory for the chosen model profile (NVIDIA gives 24 GB as the minimum for Llama 3.1 8B Instruct)
- Linux on AMD64 or ARM64; NVIDIA recommends Ubuntu 22.04 LTS or later as the validated choice
- NVIDIA GPU driver 580 or later and CUDA 12.9 or later
- Docker 24.0 or later with NVIDIA Container Toolkit 1.14.0 or later
- A free NVIDIA Developer Program membership or an NVIDIA AI Enterprise license
- An NGC Personal API key only for production-branch or older NIMs; a Hugging Face token when a model-free NIM pulls from Hugging Face
- What it is not
- NIM is not an inference engine and does not compete with TensorRT LLM, vLLM or SGLang: it packages an engine with a model and an API layer. It is not Dynamo-Triton (formerly Triton Inference Server), a general model server for many frameworks, and it is not Dynamo, which coordinates inference across many GPUs and nodes. NIM is mainly a container you run on infrastructure you choose; NVIDIA also offers hosted NIM APIs for prototyping, and NVIDIA partners offer hosted NIM endpoints. Free access covers development and testing; NVIDIA states that production use moves to an AI Enterprise license.13910
Availability and licensing. NVIDIA offers free hosted NIM APIs for prototyping and free download of NIM containers for development and testing through the NVIDIA Developer Program. For production, NVIDIA directs users to an NVIDIA AI Enterprise license, which offers a free 90-day evaluation license. NVIDIA partners also offer hosted NIM endpoints.18
The problem it solves
Putting a current generative model into production means choosing an inference engine, matching it to the GPU, picking quantization and parallelism settings, downloading large weights, exposing an API and then keeping the whole stack patched. NVIDIA's NIM documentation says the product is aimed at teams that do not have the time to follow how fast the model-serving ecosystem moves.
NIM answers this with validated containers that ship curated weights and tuned runtime settings, so the application team works against a standard API rather than a hand-built server. For long-lived or regulated deployments, NVIDIA describes production branches with continuous CVE patching and built-in health and observability APIs.4
How it works
NVIDIA's documentation describes a NIM for LLM and VLM container in three parts:
- Orchestration layer (nim-llm): starts the service, applies settings from CLI flags, environment variables and config files, loads LoRA adapters and runs a thin proxy for health checks, routing and TLS.
- Profile and model management (nimlib): handles model licensing, picks a hardware-aware profile and downloads the model.
- Inference engine (vLLM): runs the model and serves the OpenAI-compatible endpoints.
Model-specific NIMs ship a manifest, curated weights and validated quantization profiles for widely used models. Model-free NIMs take the model you name when the container starts, and can pull it from local storage, NGC, Hugging Face, S3 or GCS. A local cache directory keeps weights between restarts.
On Kubernetes, the NIM Operator manages NIMCache, NIMService and NIMPipeline resources, and a NIMBuild resource can prebuild TensorRT LLM engines for a target GPU so new replicas start faster.456
Diagram as a list
Applications & solutions
- Application or agentSends OpenAI-style requests to the NIM endpointConnects to Orchestration layer (nim-llm)
Models & frameworks
- Model sources (NGC, Hugging Face, S3, GCS, local)Where container images and weights come from
Inference & runtime software
- Orchestration layer (nim-llm)Starts the service, applies configuration and LoRA, handles health checks, routing and TLSConnects to Profile and model management (nimlib), Inference engine (vLLM)
- Profile and model management (nimlib)Selects the hardware profile, handles model licensing and downloadsConnects to Local cache or NIMCache, Model sources (NGC, Hugging Face, S3, GCS, local)
- Inference engine (vLLM)Runs the model and serves OpenAI-compatible endpointsConnects to NVIDIA GPU
Operations & orchestration
- Local cache or NIMCacheKeeps weights so restarts and new replicas skip the download
- NIM OperatorKubernetes resources for caching, services, pipelines, engine builds and autoscalingConnects to Orchestration layer (nim-llm), Local cache or NIMCache
Accelerated computing
- NVIDIA GPURuns inference; the driver and Container Toolkit expose it to the container
Capabilities
Standard inference APIs15
Serves /v1/chat/completions, /v1/completions and /v1/responses, each with streaming, plus a /v1/models listing.
Why it matters: Code written for OpenAI-style clients can point at a self-hosted endpoint with little change.
Limits: Supported parameters and modalities depend on the model; check the API reference for the NIM you run.
Hardware-aware profile selection58
At startup the NIM detects the GPUs and picks a matching model profile; NIM_MODEL_PROFILE overrides the choice.
Why it matters: Removes manual tuning of engine settings per GPU type.
Limits: Each model has a minimum GPU memory requirement, and NVIDIA directs users to the support matrix to see which images and profiles fit their model and GPU.
Model-specific and model-free containers4
Model-specific NIMs ship curated weights and tuned settings; model-free NIMs serve a model you configure at runtime from NGC, Hugging Face, S3, Google Cloud Storage or local storage.
Why it matters: Covers both popular models and custom fine-tuned ones with a single approved container image.
Limits: Model-free NIMs may need credentials for the model source, for example a Hugging Face token when the model comes from Hugging Face.
Health checks, metrics and LoRA34
Liveness and readiness endpoints, observability metrics for dashboards and support for LoRA adapters.
Why it matters: Lets orchestration tools know when a model is ready and lets one base model serve several adapted variants.
Limits: You still need your own monitoring stack to collect and display the metrics.
Kubernetes lifecycle with NIM Operator6
Custom resources for model caching, NIM services, pipelines and engine prebuilds, with autoscaling on GPU or NIM metrics and support for air-gapped clusters.
Why it matters: Moves NIM from single containers to managed, scalable cluster services.
Limits: Dynamo deployment and Kata Containers support in the operator are marked experimental.
Enterprise lifecycle46
Production branches with continuous CVE patching, OSRB compliance and FedRAMP-ready builds; a government-ready NIM Operator for AI Enterprise customers.
Why it matters: Supports long-running and regulated deployments.
Limits: Production-branch NIMs need an NGC API key and an AI Enterprise license.
Practical use cases
Staff need a chat assistant, but internal documents and prompts must stay on the company network.
- Approach
- Run an LLM NIM on in-house GPUs and point the assistant at its chat completions endpoint.
- Role of NVIDIA NIM
- NIM serves the model behind a standard API inside the data center.
- Data, infrastructure and skills
- Supported NVIDIA GPUs, Docker or Kubernetes, and a license path for production.
- Type of benefit
- Data control
- Caveats
- Production use requires an AI Enterprise license; GPU memory limits which models fit.
- First step
- Try the model on NVIDIA's hosted API, then run the same model as a local NIM.
Sources 1
A team has fine-tuned a model and needs to serve it without building and securing a serving stack.
- Approach
- Use a model-free NIM that loads the fine-tuned weights from Hugging Face, object storage or a local folder.
- Role of NVIDIA NIM
- NIM supplies the engine, API and health endpoints around the custom model.
- Data, infrastructure and skills
- Model weights in a supported format and credentials for the model source.
- Type of benefit
- Faster deployment
- Caveats
- The model architecture must be supported by the NIM backend.
- First step
- Check the support matrix for the model architecture, then start the model-free container.
Sources 4
A platform team must run several models with autoscaling, including on an air-gapped cluster.
- Approach
- Install the NIM Operator, pre-cache models with NIMCache and deploy them as NIMService resources with autoscaling.
- Role of NVIDIA NIM
- NIM provides the services; the operator manages their lifecycle on Kubernetes.
- Data, infrastructure and skills
- A Kubernetes cluster with the GPU Operator and shared storage for caches.
- Type of benefit
- Operational simplicity
- Caveats
- Some operator features, such as Dynamo deployment, are experimental.
- First step
- Deploy one NIMCache and one NIMService in a test namespace.
Sources 6
Works with
Optional integration
- NVIDIA TensorRT LLMNIM can run LLMs on TensorRT LLM, and the NIM Operator can prebuild TensorRT LLM engines.
- NVIDIA DynamoThe NIM Operator can deploy Dynamo resources (experimental), and NVIDIA says NIM will include Dynamo capabilities.
Complementary tools
- NVIDIA cuOptNVIDIA describes agents that use LLM NIM microservices to formulate problems for cuOpt.
Same family
- NVIDIA AI EnterpriseNVIDIA documents NIM as part of NVIDIA AI Enterprise; production use needs that license.
Optional integration for
- NVIDIA BioNeMoBioNeMo models are served as NIM microservices, and agent skills call NIM-hosted models.
- NVIDIA NeMoNeMo's deploy stage uses NIM to serve tuned models.
- NVIDIA NemotronNVIDIA offers Nemotron models as NIM microservices (AI Enterprise license required).
- NVIDIA AI WorkbenchExample projects include a downloadable NIM deployment, and the RAG Blueprint project uses NIM microservices.
- NVIDIA MetropolisVision NIM microservices and locally hosted NIMs serve models in the pipeline.
- NVIDIA CosmosCosmos 3 can be served through NVIDIA NIM microservices.
Has a reference implementation in
- NVIDIA BlueprintsBlueprints show NIM microservices in complete applications.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Prepare the host8
Install NVIDIA driver 580 or later, Docker 24.0 or later and NVIDIA Container Toolkit 1.14.0 or later, and configure Docker to use the NVIDIA runtime.
Check: docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi lists your GPUs and driver version.
Get access8
Join the free NVIDIA Developer Program or use an AI Enterprise license. Create an NGC Personal API key only if your image is a production-branch or older NIM.
Check: You can pull the NIM image from nvcr.io.
Start a model5
Set a local cache directory and run the NIM container with --gpus=all, mounting the cache and publishing port 8000.
Check: curl http://localhost:8000/v1/health/ready reports the status ready.
Send a first request5
Query /v1/models for the served model name, then post a chat completion to /v1/chat/completions.
Check: The response is a chat.completion object with the model's answer.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA NIM as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
- NVIDIA NIM Microservices product page (opens in a new tab)
- NVIDIA NIM documentation hub (opens in a new tab)
- NVIDIA NIM for Developers (opens in a new tab)
- NVIDIA NIM for LLM and VLM documentation: overview (opens in a new tab)
- NVIDIA NIM for LLM and VLM: quickstart (opens in a new tab)
- NVIDIA NIM Operator documentation (opens in a new tab)
- NVIDIA NIM for LLM and VLM: getting started (opens in a new tab)
- NVIDIA NIM for LLM and VLM: prerequisites (opens in a new tab)
- NVIDIA Dynamo-Triton developer page (opens in a new tab)
- NVIDIA Dynamo product page (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.