Skip to content

Create AI agents

Build AI agents that plan multi-step tasks, call tools and company systems and report back, with evaluation, guardrails and tracing so their behavior can be measured, limited and improved over time.

The problem

Many business tasks are more than one question. Researching a topic across sources, triaging support tickets, preparing a report or updating records takes several steps, tool calls and checks. A chatbot cannot do this alone, and a quickly assembled agent often loops, calls the wrong tool, runs up cost or takes actions nobody approved.

The engineering problem is to give an agent the right model, tools and data while making every run observable, testable and limited in what it can touch.

The approach

A typical agent combines a reasoning model, a set of tools (search, databases, business APIs, code execution), retrieval or memory, and an orchestration loop from an agent framework or a custom harness. Teams do well to start with one narrow task with clear success criteria and add tools one at a time.

In NVIDIA's ecosystem, Nemotron provides open reasoning models in Nano, Super and Ultra tiers plus retriever and safety models; NIM serves models behind OpenAI-compatible APIs; the NeMo Agent Toolkit profiles, evaluates and optimizes agent workflows and is framework-agnostic; NeMo Guardrails adds runtime controls; and the AI-Q Deep Researcher blueprint (repository deep-researcher-agent, where the former AI-Q research assistant repository now redirects) provides starting code. OpenShell can run agents in a sandbox with policy-controlled access.

A hosted model API with a general-purpose agent framework is enough for many internal agents. NVIDIA components matter more when models must run on your own GPUs, when you want open models you can fine-tune, or when agent traffic is large enough that serving cost dominates.123456

Conceptual architecture

Create AI agents: conceptual architectureApplications &solutionsModels & frameworksInference & runtimesoftwareOperations &orchestrationUser request or event trigger: Starts an agent run with a goal and contextUser request or eventtriggerAgent orchestration loop: Plans steps, calls tools, keeps state and decides when to stopAgent orchestration loopTools and business APIs: Search, databases, ticketing and other systems the agent may callTools and business APIsRetrieval and memory: Finds relevant documents and earlier contextRetrieval and memoryReasoning model endpoint: Chooses the next action and writes resultsReasoning model endpointGuardrails and policy: Filters inputs, tool output and responsesGuardrails and policyEvaluation and tracing: Records each step, scores runs and catches regressionsEvaluation and tracing
Diagram as a list
  1. Applications & solutions

    • User request or event triggerStarts an agent run with a goal and contextConnects to Agent orchestration loop
    • Agent orchestration loopPlans steps, calls tools, keeps state and decides when to stopConnects to Reasoning model endpoint, Tools and business APIs, Retrieval and memory
    • Tools and business APIsSearch, databases, ticketing and other systems the agent may call
  2. Models & frameworks

    • Retrieval and memoryFinds relevant documents and earlier context
  3. Inference & runtime software

    • Reasoning model endpointChooses the next action and writes results
  4. Operations & orchestration

    • Guardrails and policyFilters inputs, tool output and responsesConnects to Agent orchestration loop
    • Evaluation and tracingRecords each step, scores runs and catches regressionsConnects to Agent orchestration loop
Conceptual: one common way to arrange the parts, not a required design.

Technologies and their roles

  • nemotron1

    Open reasoning and retrieval models

    Model tiers for sub-agents through complex multi-step work, plus retriever and safety models, with open weights.

  • nim7

    Model serving

    Serves the agent's models behind OpenAI-compatible APIs on your own infrastructure.

  • nemo4

    Agent evaluation and guardrails

    The NeMo Agent Toolkit profiles and evaluates agent workflows; Guardrails and Evaluator add controls and scoring.

  • blueprints4

    Starting code

    The AI-Q Deep Researcher blueprint (repository deep-researcher-agent) shows a full working research-agent pattern.

  • openshell6

    Sandboxed runtime

    Runs agents in isolated sandboxes where nothing is permitted by default and every access decision is logged.

What you need first

  • One narrow first task with measurable success criteria
  • Service accounts for each tool, scoped to that task
  • A test set of tasks with expected outcomes
  • Python skills and familiarity with an agent framework
  • Logging and tracing for every agent step and tool call

Risks and how to reduce them

The agent takes an unintended action in a live system
Start read-only, require human approval for writes, payments and messages, and scope credentials per tool.
Prompt injection through retrieved documents or web pages
Treat tool output as untrusted data, filter it and limit which tools a single run can reach.
Runaway loops and cost
Set step, time and token budgets per run and alert on outliers.
Quality drifts as models, prompts or tools change
Re-run the evaluation suite on every change before release.

Related

Sources

  1. NVIDIA Nemotron foundation models (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA NIM for LLM and VLM: quickstart (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA NeMo documentation hub (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA NeMo product page (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA AI Blueprints GitHub organization (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA OpenShell (opens in a new tab)NVIDIA · Vendor-reported
  7. NVIDIA NIM Microservices product page (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)