Skip to content

NVIDIA NeMo

NVIDIA NeMo is an open suite of libraries for preparing data, training and post-training models, evaluating them and adding guardrails to AI agents. Its containerized NeMo Microservices reached their announced sunset date of October 1, 2026; NVIDIA's docs still list the open source NeMo Framework.123

Also known as NeMo, NeMo Microservices (sunset October 1, 2026)

At a glance

What is it?
NeMo is NVIDIA's software suite for the lifecycle of models and AI agents, which NVIDIA now describes as an agent-first, open suite of libraries. Two parts are often confused. The NeMo Framework is a set of open source, PyTorch-based libraries (such as Megatron-Bridge, AutoModel, NeMo RL, Curator and NeMo Speech) for pretraining, post-training and reinforcement learning of language, multimodal and speech models. NeMo Microservices was a set of containerized APIs for customization, evaluation and guardrails; its archived documentation gives October 1, 2026 as the sunset date and names NeMo Platform as where new development moved. NVIDIA has since renamed NeMo Platform to NeMo Helix.12345
What does it do?
NeMo covers the work around a model rather than serving it. Curator cleans and filters multimodal training data; Data Designer, Safe Synthesizer and Anonymizer create synthetic or privacy-protected datasets; AutoModel and Megatron-Bridge train and fine-tune models; NeMo RL and NeMo Gym run reinforcement-learning post-training; Evaluator benchmarks models and agents, including LLM-as-a-judge; Guardrails adds programmable safety and topic controls; and the NeMo Agent Toolkit profiles, evaluates and optimizes agent workflows. Export and Deploy hands models to TensorRT, TensorRT LLM, vLLM or Triton backends.13
Who needs it?
Teams that adapt models to their own data or domain through fine-tuning, distillation or reinforcement learning; researchers training large language, multimodal or speech models on multi-GPU clusters; and developers who need evaluation, guardrails and observability around production agents. Organizations that used NeMo Microservices need a migration plan, since NVIDIA asks them to contact their account team.12
What does it need?1246
  • One or more NVIDIA GPUs; multi-node GPU clusters for large pretraining or post-training jobs
  • Python and PyTorch, the framework's main interface and base
  • Training or evaluation data you are permitted to use
  • A starting model, such as a Nemotron model or another open model
  • For the NeMo Framework container: acceptance of the NVIDIA AI Product Agreement
  • For former NeMo Microservices users: a migration plan agreed with the NVIDIA account team
What it is not
NeMo is not an inference server: serving is handled by NIM, Dynamo, Dynamo-Triton or an engine such as TensorRT LLM or vLLM. It is not a model family either; Nemotron is NVIDIA's model family, while NeMo is tooling used to build and tune models. NeMo Microservices and the NeMo Framework are different things: the containerized microservices are sunset, the framework libraries continue as open source. The sunset notice names NeMo Platform as where new development moved, and the NeMo Helix v0.7.0 release notes (October 6, 2026) say that release completes the rename of NeMo Platform to NeMo Helix. So NeMo Helix, an open source control plane for production agents that is still before version 1.0, is the renamed successor.1257

Availability and licensing. The NVIDIA-NeMo GitHub organization states that its repositories are Apache 2.0 licensed, with third-party attributions listed in each repository. The NeMo Framework container is licensed under the NVIDIA AI Product Agreement. NVIDIA positions enterprise support for NeMo through NVIDIA AI Enterprise. NeMo Microservices reached their announced sunset date of October 1, 2026; their documentation is kept as an archive for reference.1246

The problem it solves

Getting a general model to perform reliably on an organization's own tasks takes more than a prompt. Teams need clean and lawful training data, a way to fine-tune or post-train at scale, repeatable evaluation and safety controls before anything reaches users. Each of these steps often means a different tool, and moving work between them is slow.

NeMo groups these steps into one family of libraries that share formats and connect to NVIDIA's serving stack. NVIDIA presents the suite around an agent lifecycle: build (prepare data and select or train a model), deploy (serve it, ground it in data and apply guardrails) and optimize (observe the agent in production, evaluate it and feed what is learned back into training as a data flywheel).1

How it works

NeMo is modular: each library can be installed and used on its own. NVIDIA groups them by lifecycle stage.

  • Data: NeMo Curator (GPU-accelerated cleaning and filtering of multimodal data), Data Designer (synthetic datasets), Anonymizer and Safe Synthesizer (privacy-preserving data).
  • Training and post-training: Megatron-Bridge (Megatron-Core parallelism with a PyTorch training loop), AutoModel (PyTorch-native training and fine-tuning of Hugging Face models), NeMo RL and NeMo Gym (reinforcement learning and simulated environments), NeMo Speech (ASR and TTS) and NeMo Run (launching and tracking jobs on local, on-premises and cloud clusters).
  • Quality and safety: NeMo Evaluator (academic benchmarks, LLM-as-a-judge and custom metrics) and NeMo Guardrails (programmable safety, policy and topic control).
  • Agents: NeMo Agent Toolkit (framework-agnostic profiling, evaluation and optimization of agent systems).
  • Hand-off to serving: Export and Deploy moves models to TensorRT, TensorRT LLM, vLLM or Triton backends; NIM can then serve them behind OpenAI-compatible APIs.

The framework is built on PyTorch with Python as its main interface. Starting with release 26.02, NeMo Framework container releases are documented through Megatron-Bridge.36

NVIDIA NeMo architecture: components by layer and how they connectApplications &solutionsModels & frameworksInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingNeMo Guardrails: Applies safety and topic controls around model callsNeMo GuardrailsAgent with NeMo Agent Toolkit: Agent workflow that is profiled and optimizedAgent with NeMo AgentToolkitCurator and Data Designer: Prepare real and synthetic training dataCurator and Data DesignerMegatron-Bridge and AutoModel: Train and fine-tune modelsMegatron-Bridge andAutoModelNeMo RL and NeMo Gym: Reinforcement-learning post-training in simulated environmentsNeMo RL and NeMo GymExport and Deploy: Hand models to TensorRT, TensorRT LLM, vLLM or Triton backendsExport and DeployServing (NIM or other runtime): Serves the tuned model behind an APIServing (NIM or otherruntime)NeMo Evaluator: Benchmark models and agents against baselinesNeMo EvaluatorNVIDIA GPU clusters (jobs via NeMo Run): Run curation, training and RL jobsNVIDIA GPU clusters (jobsvia NeMo Run)
Diagram as a list
  1. Applications & solutions

    • NeMo GuardrailsApplies safety and topic controls around model callsConnects to Agent with NeMo Agent Toolkit
    • Agent with NeMo Agent ToolkitAgent workflow that is profiled and optimizedConnects to Curator and Data Designer
  2. Models & frameworks

    • Curator and Data DesignerPrepare real and synthetic training dataConnects to Megatron-Bridge and AutoModel
    • Megatron-Bridge and AutoModelTrain and fine-tune modelsConnects to NeMo RL and NeMo Gym, NeMo Evaluator
    • NeMo RL and NeMo GymReinforcement-learning post-training in simulated environmentsConnects to NeMo Evaluator
  3. Inference & runtime software

    • Export and DeployHand models to TensorRT, TensorRT LLM, vLLM or Triton backendsConnects to Serving (NIM or other runtime)
    • Serving (NIM or other runtime)Serves the tuned model behind an APIConnects to NeMo Guardrails
  4. Operations & orchestration

    • NeMo EvaluatorBenchmark models and agents against baselinesConnects to Export and Deploy
  5. Accelerated computing

    • NVIDIA GPU clusters (jobs via NeMo Run)Run curation, training and RL jobsConnects to Megatron-Bridge and AutoModel
Components and connections as documented by NVIDIA.3

Capabilities

  • Data preparation and synthetic data1

    Curator cleans and filters multimodal data on GPUs; Data Designer, Safe Synthesizer and Anonymizer create synthetic or privacy-protected datasets.

    Why it matters: Training and evaluation quality depend on data quality, and real data is often scarce or sensitive.

    Limits: Synthetic data still needs human review for quality and bias.

  • Scalable training and fine-tuning3

    Megatron-Bridge brings Megatron-Core parallelism to a PyTorch training loop; AutoModel trains natively in PyTorch and fine-tunes Hugging Face models.

    Why it matters: One toolset scales from small fine-tunes to large multi-node runs.

    Limits: Large runs need GPU clusters and distributed-training skills.

  • Reinforcement-learning post-training3

    NeMo RL aligns models with reinforcement learning at scale; NeMo Gym builds simulated environments that generate agentic training rollouts.

    Why it matters: Lets teams improve reasoning and agent behavior beyond supervised fine-tuning.

    Limits: Reward design and environment building take significant expertise.

  • Evaluation1

    NeMo Evaluator benchmarks models and agents with academic tests, LLM-as-a-judge and custom metrics.

    Why it matters: Gives a repeatable baseline before and after every change.

    Limits: Results are only as meaningful as the chosen tasks and judges.

  • Guardrails3

    NeMo Guardrails adds programmable safety, policy and topical control to LLM and agent systems.

    Why it matters: Keeps assistants on approved topics and filters unsafe exchanges.

    Limits: We recommend treating guardrails as one layer of risk reduction and keeping testing and human review in place.

  • Agent profiling and optimization3

    The NeMo Agent Toolkit is an open source, framework-agnostic toolkit to build, profile, evaluate and optimize agentic systems.

    Why it matters: Shows where an agent spends time and tokens so it can be tuned.

    Limits: It instruments agents; it does not host or serve them.

Practical use cases

A general model misreads industry terms and internal document formats.
Approach
Curate domain data with Curator, fine-tune with AutoModel or Megatron-Bridge, compare against the base model with Evaluator, then serve the result.
Role of NVIDIA NeMo
NeMo supplies data, training and evaluation; a serving layer such as NIM runs the model.
Data, infrastructure and skills
Curated domain data, GPU time and an evaluation set built from real tasks.
Type of benefit
Domain accuracy
Caveats
Fine-tuning can reduce general ability; always compare with the base model.
First step
Build a small evaluation set from real questions before training.

Sources 1

A customer-facing assistant must stay on approved topics and refuse unsafe requests.
Approach
Wrap model calls with NeMo Guardrails and test the configuration with Evaluator.
Role of NVIDIA NeMo
NeMo Guardrails enforces policy and topic rules at runtime.
Data, infrastructure and skills
Written policies, test prompts and a deployed model endpoint.
Type of benefit
Risk reduction
Caveats
Rules need maintenance as products and policies change.
First step
Write the five policies that matter most and test each with adversarial prompts.

Sources 3

An agent in production gets slower and more expensive as usage grows.
Approach
Profile the workflow with the NeMo Agent Toolkit, collect production feedback, evaluate smaller candidate models and fine-tune them in a data flywheel.
Role of NVIDIA NeMo
NeMo provides profiling, evaluation and fine-tuning steps of the loop.
Data, infrastructure and skills
Logged agent traces, consent to reuse them and an evaluation harness.
Type of benefit
Cost and quality optimization
Caveats
NVIDIA's Data Flywheel blueprint was built on NeMo Microservices, which are now sunset; check its current status.
First step
Instrument one agent workflow with the Agent Toolkit and record a baseline.

Sources 3

Who uses it

  • Foxconn (Hon Hai Technology Group) · Electronics manufacturing

    Foxconn: digital twins for new server plants with Omniverse, Isaac and Metropolis

    Foxconn uses NVIDIA Omniverse digital twins to plan production lines, Isaac to simulate robots and Metropolis for camera-based monitoring, from Hsinchu to new server plants in Mexico and the US. Most published results are expectations, such as a forecast energy cut of over 30 percent in Mexico.

    In production

Works with

Optional integration

Complementary tools

  • NVIDIA BioNeMoNVIDIA says the Agent Toolkit is built on NIM, Parabricks, NeMo and Nemotron technologies.
  • NVIDIA NemotronNemotron models are a suggested starting point for NeMo training and tuning.

Same family

Optional integration for

  • NVIDIA DynamoDynamo 1.0 lists integration with the NeMo Agent Toolkit for agentic inference hints.

Has a reference implementation in

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Pick the library for your stage3

    Use the NeMo documentation hub to choose one library, such as Curator, AutoModel, Evaluator or Guardrails, instead of installing the whole suite.

    Check: You can name the single library that covers your first task.

  2. Set up the environment6

    Use the NeMo Framework container (documented through Megatron-Bridge from release 26.02) or install the library into a Python environment with PyTorch.

    Check: The library imports and detects your GPUs.

  3. Record a baseline3

    Run NeMo Evaluator on the base model with tasks drawn from your use case before changing anything.

    Check: You have a baseline score saved for comparison.

  4. Fine-tune and compare

    Run a small fine-tuning job with AutoModel or Megatron-Bridge and evaluate again with the same tasks.

    Check: The new score differs measurably from the baseline on your tasks.

  5. Plan the microservices migration2

    If you used NeMo Microservices, contact your NVIDIA account team as the sunset notice asks.

    Check: Each microservice you used has a documented replacement or exit plan.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA NeMo as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA NeMo product page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. NeMo Microservices documentation (archived) (opens in a new tab) NVIDIA · Vendor-reported, Recommendation · link checked 9 Oct 2026
  3. NVIDIA NeMo documentation hub (opens in a new tab) NVIDIA · Vendor-reported, Recommendation · link checked 9 Oct 2026
  4. NVIDIA-NeMo GitHub organization (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  5. NeMo Helix release notes v0.7.0 (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  6. NVIDIA NeMo Framework user guide: overview (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  7. NeMo Helix GitHub repository (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)