Skip to content

NVIDIA Blueprints

NVIDIA Blueprints are open reference workflows for agentic and generative AI use cases such as retrieval-augmented generation, video search and summarization, and research agents. Each provides sample code, deployment files and documentation built on NIM microservices and other NVIDIA libraries, meant to be adapted.12

Also known as NVIDIA AI Blueprints, NIM Agent Blueprints

At a glance

What is it?
NVIDIA Blueprints are reference AI workflows for specific use cases. Each is a working sample application with source code, customization documentation and deployment assets such as Helm charts, built mainly on NVIDIA NIM microservices plus NeMo, Omniverse or CUDA-X libraries. NVIDIA launched them as NIM Agent Blueprints in August 2024 and renamed them NVIDIA Blueprints in October 2024. The code is published in the NVIDIA-AI-Blueprints GitHub organization, which listed 37 repositories at review time.245
What does it do?
A blueprint shows how the parts of an AI application fit together end to end. The RAG Blueprint, for example, extracts text, tables, charts and audio from documents with NeMo Retriever models, stores them in a vector database such as Milvus or Elasticsearch, answers questions grounded in that data, and can add guardrails and an agentic plan-and-execute mode. The Video Search and Summarization blueprint combines vision language models, LLMs and RAG to search, summarize and answer questions about live or recorded video. Other repositories cover research agents, PDF to podcast, LLM routing and cuOpt-based portfolio optimization.2367
Who needs it?
Developers and solution architects who want a working starting point instead of a blank page, partners and systems integrators building customer solutions, and teams that want to test whether NVIDIA's stack fits a use case before committing to a custom build.4
What does it need?36
  • NVIDIA GPUs for self-hosted mode, as listed in each blueprint's support matrix, or an NVIDIA API key for hosted model endpoints
  • Docker, or Kubernetes with Helm for cluster deployments
  • Disk space for model downloads (the RAG Blueprint asks for about 200 GB when self-hosting)
  • Acceptance of the blueprint's code license and the separate model licenses
  • Developers able to adapt Python services, prompts and configuration
What it is not
A blueprint is not a finished product or a managed service. It is sample code that you deploy, adapt and operate yourself, and it downloads third-party software and models that carry their own license terms. Blueprints are not an alternative to NIM or NeMo: they are built from them. Using the models inside a blueprint is governed by the model licenses, separately from the blueprint code. Check each repository's release notes and support matrix before relying on it.46

Availability and licensing. Blueprint source code is published on GitHub. The RAG Blueprint, for example, is licensed under Apache 2.0, while each model it uses has its own license: the README names NVIDIA community model licenses, the Llama 3.1 or 3.2 Community License for Llama-based models (NemoGuard safety models, llama-nemotron embed and rerank models) and Apache 2.0 for extraction models such as nemotron-page-elements and nemotron-ocr. Check each model before use. Production use of the underlying NIM microservices requires an NVIDIA AI Enterprise license.68

The problem it solves

Turning model APIs into a working application involves many decisions at once: how to ingest and chunk documents, which embedding and reranking models to use, where to store vectors, how to orchestrate an agent, how to add safety checks and how to deploy all of it. Teams often spend weeks wiring these parts together before they can judge whether the idea works.

Blueprints give a tested arrangement of these parts with code, configuration and deployment files, so a team can run a complete example quickly and then replace pieces with its own data, models and services.4

How it works

Each blueprint is a GitHub repository with application code, configuration and deployment assets. Using the RAG Blueprint as an example:

  • Ingestion: an ingestor service extracts multimodal content with NeMo Retriever models and writes embeddings to a vector database (Milvus or Elasticsearch).
  • Retrieval and generation: a RAG server retrieves and reranks passages and calls an LLM NIM to answer; an optional agentic mode uses a LangGraph plan-and-execute pipeline for multi-hop questions.
  • Safety and evaluation: optional NeMo Guardrails, plus guides for accuracy and performance evaluation and observability.
  • Interfaces: a reference web UI, a Python package and REST APIs.

Deployment options include Docker with self-hosted models, Docker with NVIDIA-hosted model endpoints, Helm on Kubernetes or OpenShift, the NIM Operator, and a containerless lite mode. NVIDIA also states that blueprints can be launched in one click with NVIDIA Launchables or downloaded for local PCs and workstations.13

NVIDIA Blueprints architecture: components by layer and how they connectApplications &solutionsInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingReference UI or your application: Asks questions and shows answersReference UI or yourapplicationRAG server (optional agentic mode): Retrieves context and orchestrates the answerRAG server (optionalagentic mode)Ingestor server (NeMo Retriever extraction): Extracts text, tables, charts and audio from documentsIngestor server (NeMoRetriever extraction)NeMo Guardrails (optional): Applies safety and topic rulesNeMo Guardrails(optional)Embedding and reranking NIMs: Turn content and queries into vectors and rank resultsEmbedding and reranking NIMsLLM NIM: Generates the grounded answerLLM NIMVector database (Milvus or Elasticsearch): Stores and searches embeddingsVector database (Milvus orElasticsearch)Docker, Helm or NIM Operator: Deploys the servicesDocker, Helm or NIM OperatorNVIDIA GPUs or NVIDIA-hosted endpoints: Run the modelsNVIDIA GPUs or NVIDIA-hostedendpoints
Diagram as a list
  1. Applications & solutions

    • Reference UI or your applicationAsks questions and shows answersConnects to RAG server (optional agentic mode)
    • RAG server (optional agentic mode)Retrieves context and orchestrates the answerConnects to Embedding and reranking NIMs, LLM NIM, NeMo Guardrails (optional)
    • Ingestor server (NeMo Retriever extraction)Extracts text, tables, charts and audio from documentsConnects to Embedding and reranking NIMs
    • NeMo Guardrails (optional)Applies safety and topic rules
  2. Inference & runtime software

    • Embedding and reranking NIMsTurn content and queries into vectors and rank resultsConnects to Vector database (Milvus or Elasticsearch)
    • LLM NIMGenerates the grounded answerConnects to NVIDIA GPUs or NVIDIA-hosted endpoints
  3. Operations & orchestration

    • Vector database (Milvus or Elasticsearch)Stores and searches embeddings
    • Docker, Helm or NIM OperatorDeploys the servicesConnects to Embedding and reranking NIMs, LLM NIM
  4. Accelerated computing

    • NVIDIA GPUs or NVIDIA-hosted endpointsRun the models
Components and connections as documented by NVIDIA.6

Capabilities

  • Working reference application46

    Each blueprint ships a sample application with source code, documentation and, for RAG, a reference UI.

    Why it matters: Teams see a complete, running example before writing their own code.

    Limits: We recommend hardening, testing and assigning an owner to the sample code before production use.

  • Multiple deployment paths3

    Docker with self-hosted or NVIDIA-hosted models, Helm on Kubernetes or OpenShift, the NIM Operator, and library or lite modes.

    Why it matters: The same workflow can move from a laptop test to a cluster.

    Limits: Self-hosted deployments are large: the RAG Blueprint asks for about 200 GB of disk and long first downloads.

  • Built on NIM and NVIDIA libraries5

    Blueprints combine NIM microservices with NeMo, Omniverse or CUDA-X libraries and partner microservices.

    Why it matters: Shows the intended way NVIDIA's components fit together.

    Limits: Ties the reference design to NVIDIA components; swapping them takes extra work.

  • Designed to be modified34

    Components such as models, prompts, vector databases and pipeline modes are configurable; NVIDIA says blueprints are meant to be modified and extended.

    Why it matters: Teams can replace parts with their own data and services.

    Limits: Changes move you away from the tested configuration.

  • Open source code6

    Blueprint code is public on GitHub; the RAG Blueprint is licensed under Apache 2.0.

    Why it matters: Code can be inspected, forked and reused.

    Limits: Each model used by the blueprint has its own license (for the RAG Blueprint: NVIDIA community model licenses, the Llama 3.1 or 3.2 Community License, or Apache 2.0).

Practical use cases

Employees cannot find answers spread across PDFs, slides and reports.
Approach
Deploy the RAG Blueprint, ingest a sample of company documents and test grounded question answering.
Role of NVIDIA Blueprints
The blueprint supplies ingestion, retrieval, generation and UI as a working starting point.
Data, infrastructure and skills
Documents, GPUs or hosted endpoints, and a vector database.
Type of benefit
Faster time to prototype
Caveats
Retrieval quality depends on document quality and access controls you must add.
First step
Run the Docker deployment with NVIDIA-hosted models on a small document set.

Sources 6

Staff spend hours reviewing recorded or live video for specific events.
Approach
Use the Video Search and Summarization blueprint to search, summarize and question video with vision language models.
Role of NVIDIA Blueprints
The blueprint provides the video ingestion, VLM, LLM and RAG pipeline.
Data, infrastructure and skills
Video sources, GPUs sized per the blueprint's hardware requirements and privacy review.
Type of benefit
Review automation
Caveats
Video analytics raises privacy and consent obligations.
First step
Read the VSS hardware requirements and test on archived footage.

Sources 7

Analysts need research reports that combine internal data and web sources.
Approach
Start from NVIDIA's research agent blueprint (AI-Q / deep researcher) and connect it to internal data.
Role of NVIDIA Blueprints
The blueprint provides an agent that plans, retrieves and writes reports.
Data, infrastructure and skills
Data connectors, model access and evaluation of report accuracy.
Type of benefit
Analyst productivity
Caveats
Generated reports need human review before use in decisions.
First step
Run the reference agent on a public topic and review its sources.

Sources 9

Works with

Complementary tools

  • NVIDIA AI WorkbenchAI Blueprint example projects (PDF to Podcast, RAG, AI-Q) run in Workbench.
  • NVIDIA AI EnterpriseNVIDIA presents Blueprints with AI Enterprise for moving to production.
  • NVIDIA BioNeMoNVIDIA publishes a BioNeMo Blueprint for generative protein binder design.

Reference implementation for

  • NVIDIA NIMBlueprints show NIM microservices in complete applications.
  • NVIDIA NeMoBlueprints such as AI-Q and Data Flywheel use NeMo components.
  • NVIDIA cuOptThe portfolio-optimization blueprint is built on cuOpt.

Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.

Getting started

  1. Pick a blueprint2

    Browse the NVIDIA-AI-Blueprints GitHub organization and choose the repository closest to your use case.

    Check: The blueprint's support matrix matches the hardware you can use.

  2. Get keys and choose a mode3

    Get an NVIDIA API key and decide between NVIDIA-hosted models (few local GPUs) and self-hosted models.

    Check: You have the credentials and disk space the chosen mode needs.

  3. Deploy the RAG Blueprint with Docker3

    Follow the Docker guide; a first self-hosted deployment can take 15 to 30 minutes while models download.

    Check: The reference UI loads and answers a question about the sample documents.

  4. Swap in your data

    Ingest a small set of your own documents and adjust prompts or models.

    Check: Answers cite passages from your documents.

Official resources

Could this technology help you?

Describe your project to the Solution Architect. It starts with NVIDIA Blueprints as context but recommends independently, including when you do not need it.

Check it against my project

Sources

Each statement above links to the source it comes from. Labels say who reported it.

  1. NVIDIA NIM for Developers (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  2. NVIDIA AI Blueprints GitHub organization (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  3. NVIDIA RAG Blueprint documentation (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  4. NVIDIA Blog: NVIDIA Blueprints launch post (with 2024 rename note) (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  5. NVIDIA AI Enterprise product page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  6. NVIDIA RAG Blueprint GitHub repository (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  7. NVIDIA Video Search and Summarization Blueprint repository (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  8. NVIDIA NIM Microservices product page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026
  9. NVIDIA NeMo product page (opens in a new tab) NVIDIA · Vendor-reported · link checked 9 Oct 2026

Fill out the form below to request your copy.

Name(Required)