NVIDIA NeMo
NVIDIA NeMo is an open suite of libraries for preparing data, training and post-training models, evaluating them and adding guardrails to AI agents. Its containerized NeMo Microservices reached their announced sunset date of October 1, 2026; NVIDIA's docs still list the open source NeMo Framework.123
Also known as NeMo, NeMo Microservices (sunset October 1, 2026)
At a glance
- What is it?
- NeMo is NVIDIA's software suite for the lifecycle of models and AI agents, which NVIDIA now describes as an agent-first, open suite of libraries. Two parts are often confused. The NeMo Framework is a set of open source, PyTorch-based libraries (such as Megatron-Bridge, AutoModel, NeMo RL, Curator and NeMo Speech) for pretraining, post-training and reinforcement learning of language, multimodal and speech models. NeMo Microservices was a set of containerized APIs for customization, evaluation and guardrails; its archived documentation gives October 1, 2026 as the sunset date and names NeMo Platform as where new development moved. NVIDIA has since renamed NeMo Platform to NeMo Helix.12345
- What does it do?
- NeMo covers the work around a model rather than serving it. Curator cleans and filters multimodal training data; Data Designer, Safe Synthesizer and Anonymizer create synthetic or privacy-protected datasets; AutoModel and Megatron-Bridge train and fine-tune models; NeMo RL and NeMo Gym run reinforcement-learning post-training; Evaluator benchmarks models and agents, including LLM-as-a-judge; Guardrails adds programmable safety and topic controls; and the NeMo Agent Toolkit profiles, evaluates and optimizes agent workflows. Export and Deploy hands models to TensorRT, TensorRT LLM, vLLM or Triton backends.13
- Who needs it?
- Teams that adapt models to their own data or domain through fine-tuning, distillation or reinforcement learning; researchers training large language, multimodal or speech models on multi-GPU clusters; and developers who need evaluation, guardrails and observability around production agents. Organizations that used NeMo Microservices need a migration plan, since NVIDIA asks them to contact their account team.12
- What does it need?1246
- One or more NVIDIA GPUs; multi-node GPU clusters for large pretraining or post-training jobs
- Python and PyTorch, the framework's main interface and base
- Training or evaluation data you are permitted to use
- A starting model, such as a Nemotron model or another open model
- For the NeMo Framework container: acceptance of the NVIDIA AI Product Agreement
- For former NeMo Microservices users: a migration plan agreed with the NVIDIA account team
- What it is not
- NeMo is not an inference server: serving is handled by NIM, Dynamo, Dynamo-Triton or an engine such as TensorRT LLM or vLLM. It is not a model family either; Nemotron is NVIDIA's model family, while NeMo is tooling used to build and tune models. NeMo Microservices and the NeMo Framework are different things: the containerized microservices are sunset, the framework libraries continue as open source. The sunset notice names NeMo Platform as where new development moved, and the NeMo Helix v0.7.0 release notes (October 6, 2026) say that release completes the rename of NeMo Platform to NeMo Helix. So NeMo Helix, an open source control plane for production agents that is still before version 1.0, is the renamed successor.1257
Availability and licensing. The NVIDIA-NeMo GitHub organization states that its repositories are Apache 2.0 licensed, with third-party attributions listed in each repository. The NeMo Framework container is licensed under the NVIDIA AI Product Agreement. NVIDIA positions enterprise support for NeMo through NVIDIA AI Enterprise. NeMo Microservices reached their announced sunset date of October 1, 2026; their documentation is kept as an archive for reference.1246
The problem it solves
Getting a general model to perform reliably on an organization's own tasks takes more than a prompt. Teams need clean and lawful training data, a way to fine-tune or post-train at scale, repeatable evaluation and safety controls before anything reaches users. Each of these steps often means a different tool, and moving work between them is slow.
NeMo groups these steps into one family of libraries that share formats and connect to NVIDIA's serving stack. NVIDIA presents the suite around an agent lifecycle: build (prepare data and select or train a model), deploy (serve it, ground it in data and apply guardrails) and optimize (observe the agent in production, evaluate it and feed what is learned back into training as a data flywheel).1
How it works
NeMo is modular: each library can be installed and used on its own. NVIDIA groups them by lifecycle stage.
- Data: NeMo Curator (GPU-accelerated cleaning and filtering of multimodal data), Data Designer (synthetic datasets), Anonymizer and Safe Synthesizer (privacy-preserving data).
- Training and post-training: Megatron-Bridge (Megatron-Core parallelism with a PyTorch training loop), AutoModel (PyTorch-native training and fine-tuning of Hugging Face models), NeMo RL and NeMo Gym (reinforcement learning and simulated environments), NeMo Speech (ASR and TTS) and NeMo Run (launching and tracking jobs on local, on-premises and cloud clusters).
- Quality and safety: NeMo Evaluator (academic benchmarks, LLM-as-a-judge and custom metrics) and NeMo Guardrails (programmable safety, policy and topic control).
- Agents: NeMo Agent Toolkit (framework-agnostic profiling, evaluation and optimization of agent systems).
- Hand-off to serving: Export and Deploy moves models to TensorRT, TensorRT LLM, vLLM or Triton backends; NIM can then serve them behind OpenAI-compatible APIs.
The framework is built on PyTorch with Python as its main interface. Starting with release 26.02, NeMo Framework container releases are documented through Megatron-Bridge.36
Diagram as a list
Applications & solutions
- NeMo GuardrailsApplies safety and topic controls around model callsConnects to Agent with NeMo Agent Toolkit
- Agent with NeMo Agent ToolkitAgent workflow that is profiled and optimizedConnects to Curator and Data Designer
Models & frameworks
- Curator and Data DesignerPrepare real and synthetic training dataConnects to Megatron-Bridge and AutoModel
- Megatron-Bridge and AutoModelTrain and fine-tune modelsConnects to NeMo RL and NeMo Gym, NeMo Evaluator
- NeMo RL and NeMo GymReinforcement-learning post-training in simulated environmentsConnects to NeMo Evaluator
Inference & runtime software
- Export and DeployHand models to TensorRT, TensorRT LLM, vLLM or Triton backendsConnects to Serving (NIM or other runtime)
- Serving (NIM or other runtime)Serves the tuned model behind an APIConnects to NeMo Guardrails
Operations & orchestration
- NeMo EvaluatorBenchmark models and agents against baselinesConnects to Export and Deploy
Accelerated computing
- NVIDIA GPU clusters (jobs via NeMo Run)Run curation, training and RL jobsConnects to Megatron-Bridge and AutoModel
Capabilities
Data preparation and synthetic data1
Curator cleans and filters multimodal data on GPUs; Data Designer, Safe Synthesizer and Anonymizer create synthetic or privacy-protected datasets.
Why it matters: Training and evaluation quality depend on data quality, and real data is often scarce or sensitive.
Limits: Synthetic data still needs human review for quality and bias.
Scalable training and fine-tuning3
Megatron-Bridge brings Megatron-Core parallelism to a PyTorch training loop; AutoModel trains natively in PyTorch and fine-tunes Hugging Face models.
Why it matters: One toolset scales from small fine-tunes to large multi-node runs.
Limits: Large runs need GPU clusters and distributed-training skills.
Reinforcement-learning post-training3
NeMo RL aligns models with reinforcement learning at scale; NeMo Gym builds simulated environments that generate agentic training rollouts.
Why it matters: Lets teams improve reasoning and agent behavior beyond supervised fine-tuning.
Limits: Reward design and environment building take significant expertise.
Evaluation1
NeMo Evaluator benchmarks models and agents with academic tests, LLM-as-a-judge and custom metrics.
Why it matters: Gives a repeatable baseline before and after every change.
Limits: Results are only as meaningful as the chosen tasks and judges.
Guardrails3
NeMo Guardrails adds programmable safety, policy and topical control to LLM and agent systems.
Why it matters: Keeps assistants on approved topics and filters unsafe exchanges.
Limits: We recommend treating guardrails as one layer of risk reduction and keeping testing and human review in place.
Agent profiling and optimization3
The NeMo Agent Toolkit is an open source, framework-agnostic toolkit to build, profile, evaluate and optimize agentic systems.
Why it matters: Shows where an agent spends time and tokens so it can be tuned.
Limits: It instruments agents; it does not host or serve them.
Practical use cases
A general model misreads industry terms and internal document formats.
- Approach
- Curate domain data with Curator, fine-tune with AutoModel or Megatron-Bridge, compare against the base model with Evaluator, then serve the result.
- Role of NVIDIA NeMo
- NeMo supplies data, training and evaluation; a serving layer such as NIM runs the model.
- Data, infrastructure and skills
- Curated domain data, GPU time and an evaluation set built from real tasks.
- Type of benefit
- Domain accuracy
- Caveats
- Fine-tuning can reduce general ability; always compare with the base model.
- First step
- Build a small evaluation set from real questions before training.
Sources 1
A customer-facing assistant must stay on approved topics and refuse unsafe requests.
- Approach
- Wrap model calls with NeMo Guardrails and test the configuration with Evaluator.
- Role of NVIDIA NeMo
- NeMo Guardrails enforces policy and topic rules at runtime.
- Data, infrastructure and skills
- Written policies, test prompts and a deployed model endpoint.
- Type of benefit
- Risk reduction
- Caveats
- Rules need maintenance as products and policies change.
- First step
- Write the five policies that matter most and test each with adversarial prompts.
Sources 3
An agent in production gets slower and more expensive as usage grows.
- Approach
- Profile the workflow with the NeMo Agent Toolkit, collect production feedback, evaluate smaller candidate models and fine-tune them in a data flywheel.
- Role of NVIDIA NeMo
- NeMo provides profiling, evaluation and fine-tuning steps of the loop.
- Data, infrastructure and skills
- Logged agent traces, consent to reuse them and an evaluation harness.
- Type of benefit
- Cost and quality optimization
- Caveats
- NVIDIA's Data Flywheel blueprint was built on NeMo Microservices, which are now sunset; check its current status.
- First step
- Instrument one agent workflow with the Agent Toolkit and record a baseline.
Sources 3
Who uses it
Foxconn (Hon Hai Technology Group) · Electronics manufacturing
Foxconn: digital twins for new server plants with Omniverse, Isaac and Metropolis
Foxconn uses NVIDIA Omniverse digital twins to plan production lines, Isaac to simulate robots and Metropolis for camera-based monitoring, from Hsinchu to new server plants in Mexico and the US. Most published results are expectations, such as a forecast energy cut of over 30 percent in Mexico.
In production
Works with
Optional integration
- NVIDIA NIMNeMo's deploy stage uses NIM to serve tuned models.
- NVIDIA TensorRT LLMNeMo Export and Deploy can target TensorRT LLM.
- NVIDIA Dynamo-TritonNeMo Export and Deploy can target Triton backends.
Complementary tools
- NVIDIA BioNeMoNVIDIA says the Agent Toolkit is built on NIM, Parabricks, NeMo and Nemotron technologies.
- NVIDIA NemotronNemotron models are a suggested starting point for NeMo training and tuning.
Same family
- NVIDIA AI EnterpriseAI Enterprise lists NeMo among its production-ready components.
Optional integration for
- NVIDIA DynamoDynamo 1.0 lists integration with the NeMo Agent Toolkit for agentic inference hints.
Has a reference implementation in
- NVIDIA BlueprintsBlueprints such as AI-Q and Data Flywheel use NeMo components.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Pick the library for your stage3
Use the NeMo documentation hub to choose one library, such as Curator, AutoModel, Evaluator or Guardrails, instead of installing the whole suite.
Check: You can name the single library that covers your first task.
Set up the environment6
Use the NeMo Framework container (documented through Megatron-Bridge from release 26.02) or install the library into a Python environment with PyTorch.
Check: The library imports and detects your GPUs.
Record a baseline3
Run NeMo Evaluator on the base model with tasks drawn from your use case before changing anything.
Check: You have a baseline score saved for comparison.
Fine-tune and compare
Run a small fine-tuning job with AutoModel or Megatron-Bridge and evaluate again with the same tasks.
Check: The new score differs measurably from the baseline on your tasks.
Plan the microservices migration2
If you used NeMo Microservices, contact your NVIDIA account team as the sunset notice asks.
Check: Each microservice you used has a documented replacement or exit plan.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA NeMo as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
- NVIDIA NeMo product page (opens in a new tab)
- NeMo Microservices documentation (archived) (opens in a new tab)
- NVIDIA NeMo documentation hub (opens in a new tab)
- NVIDIA-NeMo GitHub organization (opens in a new tab)
- NeMo Helix release notes v0.7.0 (opens in a new tab)
- NVIDIA NeMo Framework user guide: overview (opens in a new tab)
- NeMo Helix GitHub repository (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.