Skip to content

NVIDIA Nemotron

NVIDIA Nemotron is NVIDIA's family of open AI models for building agents: reasoning models in several sizes plus models for vision, retrieval, speech and safety. NVIDIA publishes the weights, much of the training data and the training recipes, and the models run on common open inference engines or as NIM.1

Also known as Nemotron 3, Nemotron 3 Nano, Nemotron 3 Super, Nemotron 3 Ultra, Nemotron 3.5 Lightning, Nemotron 3 Nano Omni, Nemotron Retriever, Nemotron Safety, Nemotron Speech

At a glance

What is it?
Nemotron is a family of open models built by NVIDIA. The core of the current generation, Nemotron 3, is a set of reasoning language models in tiers named Nano, Super and Ultra, joined by Nemotron 3.5 Lightning for fast, high-volume tasks and Nano Omni for text, image, video and audio input. Around them sit Nemotron Retriever models for document extraction, embedding and reranking, Nemotron Speech models for speech recognition, speech synthesis and translation, and Nemotron Safety models for content moderation and jailbreak detection. NVIDIA publishes model weights and training data on Hugging Face and training recipes on GitHub. The license differs by model: some current models use the NVIDIA Nemotron Open Model License and others use OpenMDW 1.1.123456
What does it do?
Nemotron models supply the language understanding, reasoning, tool calling, retrieval and moderation steps inside an agent or application. The reasoning tiers target different deployment sizes, from edge devices and PCs (Nano) through Super, which NVIDIA pitches for single-GPU use (its FP8 model card lists 2x H100-80GB as the minimum and a single GPU only on B200 or B300), to multi-GPU data center systems (Ultra). The Nemotron 3 models use a hybrid Mamba-Transformer mixture-of-experts design with a context window of up to one million tokens. Beyond the weights, NVIDIA's Nemotron repository provides reusable pipeline steps and complete recipes for data preparation, pretraining, fine-tuning, reinforcement learning and evaluation, plus deployment cookbooks for vLLM, SGLang and TensorRT LLM.235
Who needs it?
Teams that want to run capable language and multimodal models on their own GPUs, inspect how a model was trained, or adapt an open model to their domain. NVIDIA positions the family for long-running agents, with small models for specialized sub-agents and larger ones for planning and orchestration. It also suits organizations that need open retrieval or safety models to sit beside whatever main model they use.12
What does it need?125
  • NVIDIA GPUs sized for the chosen tier: NVIDIA describes Nano for edge and PC use, Super for single-GPU deployment (the FP8 model card lists 2x H100-80GB as the minimum, or one B200/B300 GPU) and Ultra
  • An inference engine such as vLLM, SGLang, TensorRT LLM, Ollama or llama.cpp, or an NVIDIA AI Enterprise license for Nemotron NIM microservices
  • Access to the model weights on Hugging Face and acceptance of the license stated on each model card
  • For retraining or fine-tuning, the NeMo libraries used by the Nemotron recipes and enough GPU capacity for the recipe you run
What it is not
Nemotron is not a framework or a serving product. NeMo is the tooling used to build and tune models, while NIM, Dynamo and engines such as vLLM, SGLang and TensorRT LLM serve them. Not every Nemotron model shares one license: the license is set per model on its model card, so check it before commercial use. NVIDIA's training recipes use only the openly released part of the training data, so a model you reproduce from them will not match the published benchmark results. Free use of the open weights does not include the NIM packaging, which NVIDIA ties to an AI Enterprise license.13

Availability and licensing. NVIDIA states that Nemotron models can be downloaded from Hugging Face and run in production free of charge, under the license on each model card (for example the NVIDIA Nemotron Open Model License for Nemotron 3 Super and OpenMDW 1.1 for Nemotron 3 Ultra). Nemotron models packaged as NIM microservices require an NVIDIA AI Enterprise license. Third-party inference providers also offer hosted Nemotron models.1

The problem it solves

Teams building agents often need several models at once: a capable planner, cheaper models for routine steps, an embedding and reranking stack for retrieval, and a safety filter. Closed APIs make it hard to keep data in house, to see how a model was trained, or to adapt it deeply.

Nemotron offers an open set of models that cover these roles, with weights, training data and recipes published so teams can evaluate the models before deployment, run them on their own GPUs and retrain or fine-tune them for a domain.2

How it works

The Nemotron family is organized by role:

  • Reasoning models: Nano, Super and Ultra tiers in Nemotron 3, plus Nemotron 3.5 Lightning for fast execution of specialized steps. NVIDIA's repository lists Ultra at 550B total and 55B active parameters, Super at about 120.6B total and 12.7B active, Nano at about 31.6B total and 3.6B active, and Lightning at 30B total and 3B active.
  • Multimodal: Nemotron 3 Nano Omni accepts text, image, video and audio for perception sub-agents.
  • Retrieval: Nemotron Retriever models extract content from documents, create embeddings and rerank results for RAG.
  • Speech and safety: speech recognition, synthesis and translation models, and safety models for moderation, PII detection and jailbreak detection that pair with NeMo Guardrails.

Models are downloaded from Hugging Face and served with open engines such as vLLM, SGLang, TensorRT LLM, Ollama or llama.cpp, or as NIM microservices. NeMo Switchyard, an open source router, can send each step of an agent workflow to a suitable model. For customization, the Nemotron repository provides a step catalog and recipes built on NeMo Curator, Megatron-Bridge, NeMo RL and NeMo Evaluator.23

NVIDIA Nemotron architecture: components by layer and how they connectApplications &solutionsModels & frameworksInference & runtimesoftwareAcceleratedcomputingAgent or application (harness): Plans tasks and calls models for reasoning, retrieval and checksAgent or application(harness)Nemotron reasoning models (Nano, Super, Ultra, Lightning): Reasoning, tool calling and generationNemotron reasoningmodels (Nano, Super,…Nemotron Retriever models: Document extraction, embeddings and reranking for RAGNemotron RetrievermodelsNemotron Safety and Speech models: Moderation, PII and jailbreak checks; speech in and outNemotron Safety andSpeech modelsOpen data, steps and recipes (NeMo libraries): Reproduce, fine-tune or retrain modelsOpen data, steps andrecipes (NeMo…NeMo Switchyard router (optional): Sends each workflow step to a suitable modelNeMo Switchyard router(optional)Serving: vLLM, SGLang, TensorRT LLM or NIM: Runs the model weights and exposes APIsServing: vLLM, SGLang,TensorRT LLM or NIMNVIDIA GPUs, from PC to data center: Execute inference and trainingNVIDIA GPUs, from PC to datacenter
Diagram as a list
  1. Applications & solutions

    • Agent or application (harness)Plans tasks and calls models for reasoning, retrieval and checksConnects to NeMo Switchyard router (optional), Nemotron reasoning models (Nano, Super, Ultra, Lightning), Nemotron Retriever models, Nemotron Safety and Speech models
  2. Models & frameworks

    • Nemotron reasoning models (Nano, Super, Ultra, Lightning)Reasoning, tool calling and generationConnects to Serving: vLLM, SGLang, TensorRT LLM or NIM
    • Nemotron Retriever modelsDocument extraction, embeddings and reranking for RAGConnects to Serving: vLLM, SGLang, TensorRT LLM or NIM
    • Nemotron Safety and Speech modelsModeration, PII and jailbreak checks; speech in and outConnects to Serving: vLLM, SGLang, TensorRT LLM or NIM
    • Open data, steps and recipes (NeMo libraries)Reproduce, fine-tune or retrain modelsConnects to Nemotron reasoning models (Nano, Super, Ultra, Lightning)
  3. Inference & runtime software

    • NeMo Switchyard router (optional)Sends each workflow step to a suitable modelConnects to Nemotron reasoning models (Nano, Super, Ultra, Lightning)
    • Serving: vLLM, SGLang, TensorRT LLM or NIMRuns the model weights and exposes APIsConnects to NVIDIA GPUs, from PC to data center
  4. Accelerated computing

    • NVIDIA GPUs, from PC to data centerExecute inference and training
Components and connections as documented by NVIDIA.12

Capabilities

  • Reasoning models in several sizes13

    Nano, Super and Ultra tiers plus Nemotron 3.5 Lightning cover small sub-agents up to multi-step planning.

    Why it matters: Lets an agent use a small model for routine steps and a larger one only where it is needed.

    Limits: Larger tiers need multi-GPU data center hardware; tier choice should be tested on your own tasks.

  • Open weights, data and recipes23

    NVIDIA publishes weights and training data on Hugging Face and complete training recipes on GitHub.

    Why it matters: Teams can inspect how a model was built and reproduce or adapt the process.

    Limits: Recipes use only the open subset of data, so results differ from NVIDIA's published benchmarks.

  • Long context and hybrid architecture27

    Nemotron 3 models combine Mamba and Transformer layers in a mixture-of-experts design with up to one million tokens of context.

    Why it matters: Long documents and long agent histories can stay in a single prompt.

    Limits: Very long contexts raise memory needs; the serving engine must support the architecture.

  • Retrieval models2

    Nemotron Retriever models extract text and tables from documents, generate embeddings and rerank passages.

    Why it matters: Covers the retrieval half of a RAG pipeline with open models.

    Limits: Retrieval quality depends on your documents; evaluate on your own corpus.

  • Safety and speech models1

    Safety models detect harmful content, PII and jailbreak attempts; speech models handle recognition, synthesis and translation.

    Why it matters: Adds moderation and voice to an agent without a separate vendor.

    Limits: Safety models reduce risk but do not replace application-level controls and review.

  • Broad deployment options12

    Models run on vLLM, SGLang, TensorRT LLM, Ollama and llama.cpp, or as NIM microservices.

    Why it matters: Teams can keep their existing serving stack.

    Limits: NIM packaging of Nemotron requires an NVIDIA AI Enterprise license.

  • Per-model licensing789

    Nemotron 3 Super and Nano 30B use the NVIDIA Nemotron Open Model License; Nemotron 3 Ultra and 3.5 Lightning use OpenMDW 1.1.

    Why it matters: License terms decide how weights and derivatives may be used and shared.

    Limits: Older and specialized Nemotron models carry other licenses; always read the model card.

Practical use cases

An agent runs many routine steps, and sending all of them to a large frontier model is slow and costly.
Approach
Route high-volume steps to a smaller Nemotron model such as Nemotron 3.5 Lightning and reserve a larger model for planning.
Role of NVIDIA Nemotron
Nemotron supplies the smaller execution model; NeMo Switchyard can decide which model handles each step.
Data, infrastructure and skills
A serving stack for the smaller model and routing rules tested on real traffic.
Type of benefit
Lower inference cost
Caveats
Routing errors can lower answer quality; measure task success before and after.
First step
Deploy Lightning with the vLLM or SGLang cookbook and compare it on a sample of your agent's steps.

Sources 2

A company needs a retrieval pipeline over internal documents and wants open models it can host itself.
Approach
Use Nemotron Retriever models for extraction, embedding and reranking, with a Nemotron reasoning model to write answers.
Role of NVIDIA Nemotron
Nemotron covers both the retrieval models and the answering model.
Data, infrastructure and skills
GPUs for the models, a vector store and a document ingestion process.
Type of benefit
Data control
Caveats
Check each model's license and test retrieval accuracy on your own documents.
First step
Index a small document set with a Nemotron embedding model and review the top results for real questions.

Sources 1

A team needs a model tuned to its domain and wants to know exactly how the base model was trained.
Approach
Start from an open Nemotron checkpoint and run fine-tuning or reinforcement learning steps from the Nemotron repository.
Role of NVIDIA Nemotron
Nemotron provides weights, datasets and recipes as a documented starting point.
Data, infrastructure and skills
Domain data, GPU capacity for training and familiarity with the NeMo libraries.
Type of benefit
Domain accuracy
Caveats
Recipes use only open data, so baseline results will differ from NVIDIA's reports.
First step
Run the getting-started example for Nemotron steps on a small configuration.

Sources 3

Works with

Optional integration

  • NVIDIA NIMNVIDIA offers Nemotron models as NIM microservices (AI Enterprise license required).
  • NVIDIA TensorRT LLMNVIDIA provides cookbooks to serve Nemotron with TensorRT LLM.

Complementary tools

  • NVIDIA TensorRT LLMThe TensorRT LLM docs include deployment guides for Nemotron 3 Ultra and Super models.
  • NVIDIA BioNeMoNemotron is listed among the core pieces of the BioNeMo Agent Toolkit.
  • NVIDIA NeMoNemotron models are a suggested starting point for NeMo training and tuning.
  • NVIDIA DynamoNVIDIA lists Dynamo among tools for running Nemotron at scale.

Same family

  • NVIDIA OpenShellBoth are part of NVIDIA Agent Toolkit, according to NVIDIA's March 2026 announcement.

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Pick a model and read its license5

    Choose a tier from the Nemotron model list (for example Nano for a single workstation or Super for one data center GPU) and open its Hugging Face model card.

    Check: You know the model's license name (NVIDIA Nemotron Open Model License or OpenMDW 1.1) and its hardware needs.

  2. Serve the model23

    Follow the usage cookbook for that model in the Nemotron repository to start it on vLLM, SGLang or TensorRT LLM.

    Check: The server answers a test prompt through its OpenAI-compatible endpoint.

  3. Add retrieval or safety models if needed2

    Add a Nemotron Retriever model for RAG or a Nemotron Safety model in front of the reasoning model, for example through NeMo Guardrails.

    Check: Retrieval returns relevant passages, and a known unsafe prompt is flagged.

  4. Customize with Nemotron steps3

    Clone the Nemotron repository and run a fine-tuning or evaluation step from the step catalog with your own data.

    Check: The evaluation step reports results for the base and the tuned model on the same benchmark.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA Nemotron as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA Nemotron foundation models (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. NVIDIA Nemotron developer page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. NVIDIA Nemotron developer repository (README) (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  4. Hugging Face model card: NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (opens in a new tab) NVIDIA (Hugging Face) · Vendor-reported · link checked 9 Oct 2026
  5. Nemotron 3 Super 120B FP8 model card (Hugging Face) (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  6. NVIDIA newsroom: NVIDIA Agent Toolkit announcement, 16 March 2026 (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  7. Nemotron 3.5 Lightning 30B model card (Hugging Face) (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  8. Hugging Face model card: NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (opens in a new tab) NVIDIA (Hugging Face) · Vendor-reported · link checked 9 Oct 2026
  9. NVIDIA RAG Blueprint GitHub repository (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)