Create AI agents
Build AI agents that plan multi-step tasks, call tools and company systems and report back, with evaluation, guardrails and tracing so their behavior can be measured, limited and improved over time.
The problem
Many business tasks are more than one question. Researching a topic across sources, triaging support tickets, preparing a report or updating records takes several steps, tool calls and checks. A chatbot cannot do this alone, and a quickly assembled agent often loops, calls the wrong tool, runs up cost or takes actions nobody approved.
The engineering problem is to give an agent the right model, tools and data while making every run observable, testable and limited in what it can touch.
The approach
A typical agent combines a reasoning model, a set of tools (search, databases, business APIs, code execution), retrieval or memory, and an orchestration loop from an agent framework or a custom harness. Teams do well to start with one narrow task with clear success criteria and add tools one at a time.
In NVIDIA's ecosystem, Nemotron provides open reasoning models in Nano, Super and Ultra tiers plus retriever and safety models; NIM serves models behind OpenAI-compatible APIs; the NeMo Agent Toolkit profiles, evaluates and optimizes agent workflows and is framework-agnostic; NeMo Guardrails adds runtime controls; and the AI-Q Deep Researcher blueprint (repository deep-researcher-agent, where the former AI-Q research assistant repository now redirects) provides starting code. OpenShell can run agents in a sandbox with policy-controlled access.
A hosted model API with a general-purpose agent framework is enough for many internal agents. NVIDIA components matter more when models must run on your own GPUs, when you want open models you can fine-tune, or when agent traffic is large enough that serving cost dominates.123456
Conceptual architecture
Diagram as a list
Applications & solutions
- User request or event triggerStarts an agent run with a goal and contextConnects to Agent orchestration loop
- Agent orchestration loopPlans steps, calls tools, keeps state and decides when to stopConnects to Reasoning model endpoint, Tools and business APIs, Retrieval and memory
- Tools and business APIsSearch, databases, ticketing and other systems the agent may call
Models & frameworks
- Retrieval and memoryFinds relevant documents and earlier context
Inference & runtime software
- Reasoning model endpointChooses the next action and writes results
Operations & orchestration
- Guardrails and policyFilters inputs, tool output and responsesConnects to Agent orchestration loop
- Evaluation and tracingRecords each step, scores runs and catches regressionsConnects to Agent orchestration loop
Technologies and their roles
nemotron1
Open reasoning and retrieval models
Model tiers for sub-agents through complex multi-step work, plus retriever and safety models, with open weights.
nim7
Model serving
Serves the agent's models behind OpenAI-compatible APIs on your own infrastructure.
nemo4
Agent evaluation and guardrails
The NeMo Agent Toolkit profiles and evaluates agent workflows; Guardrails and Evaluator add controls and scoring.
blueprints4
Starting code
The AI-Q Deep Researcher blueprint (repository deep-researcher-agent) shows a full working research-agent pattern.
openshell6
Sandboxed runtime
Runs agents in isolated sandboxes where nothing is permitted by default and every access decision is logged.
What you need first
- One narrow first task with measurable success criteria
- Service accounts for each tool, scoped to that task
- A test set of tasks with expected outcomes
- Python skills and familiarity with an agent framework
- Logging and tracing for every agent step and tool call
Risks and how to reduce them
- The agent takes an unintended action in a live system
- Start read-only, require human approval for writes, payments and messages, and scope credentials per tool.
- Prompt injection through retrieved documents or web pages
- Treat tool output as untrusted data, filter it and limit which tools a single run can reach.
- Runaway loops and cost
- Set step, time and token budgets per run and alert on outliers.
- Quality drifts as models, prompts or tools change
- Re-run the evaluation suite on every change before release.
Related
Sources
- NVIDIA Nemotron foundation models (opens in a new tab)
- NVIDIA NIM for LLM and VLM: quickstart (opens in a new tab)
- NVIDIA NeMo documentation hub (opens in a new tab)
- NVIDIA NeMo product page (opens in a new tab)
- NVIDIA AI Blueprints GitHub organization (opens in a new tab)
- NVIDIA OpenShell (opens in a new tab)
- NVIDIA NIM Microservices product page (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.