NVIDIA Blueprints
NVIDIA Blueprints are open reference workflows for agentic and generative AI use cases such as retrieval-augmented generation, video search and summarization, and research agents. Each provides sample code, deployment files and documentation built on NIM microservices and other NVIDIA libraries, meant to be adapted.12
Also known as NVIDIA AI Blueprints, NIM Agent Blueprints
At a glance
- What is it?
- NVIDIA Blueprints are reference AI workflows for specific use cases. Each is a working sample application with source code, customization documentation and deployment assets such as Helm charts, built mainly on NVIDIA NIM microservices plus NeMo, Omniverse or CUDA-X libraries. NVIDIA launched them as NIM Agent Blueprints in August 2024 and renamed them NVIDIA Blueprints in October 2024. The code is published in the NVIDIA-AI-Blueprints GitHub organization, which listed 37 repositories at review time.245
- What does it do?
- A blueprint shows how the parts of an AI application fit together end to end. The RAG Blueprint, for example, extracts text, tables, charts and audio from documents with NeMo Retriever models, stores them in a vector database such as Milvus or Elasticsearch, answers questions grounded in that data, and can add guardrails and an agentic plan-and-execute mode. The Video Search and Summarization blueprint combines vision language models, LLMs and RAG to search, summarize and answer questions about live or recorded video. Other repositories cover research agents, PDF to podcast, LLM routing and cuOpt-based portfolio optimization.2367
- Who needs it?
- Developers and solution architects who want a working starting point instead of a blank page, partners and systems integrators building customer solutions, and teams that want to test whether NVIDIA's stack fits a use case before committing to a custom build.4
- What does it need?36
- NVIDIA GPUs for self-hosted mode, as listed in each blueprint's support matrix, or an NVIDIA API key for hosted model endpoints
- Docker, or Kubernetes with Helm for cluster deployments
- Disk space for model downloads (the RAG Blueprint asks for about 200 GB when self-hosting)
- Acceptance of the blueprint's code license and the separate model licenses
- Developers able to adapt Python services, prompts and configuration
- What it is not
- A blueprint is not a finished product or a managed service. It is sample code that you deploy, adapt and operate yourself, and it downloads third-party software and models that carry their own license terms. Blueprints are not an alternative to NIM or NeMo: they are built from them. Using the models inside a blueprint is governed by the model licenses, separately from the blueprint code. Check each repository's release notes and support matrix before relying on it.46
Availability and licensing. Blueprint source code is published on GitHub. The RAG Blueprint, for example, is licensed under Apache 2.0, while each model it uses has its own license: the README names NVIDIA community model licenses, the Llama 3.1 or 3.2 Community License for Llama-based models (NemoGuard safety models, llama-nemotron embed and rerank models) and Apache 2.0 for extraction models such as nemotron-page-elements and nemotron-ocr. Check each model before use. Production use of the underlying NIM microservices requires an NVIDIA AI Enterprise license.68
The problem it solves
Turning model APIs into a working application involves many decisions at once: how to ingest and chunk documents, which embedding and reranking models to use, where to store vectors, how to orchestrate an agent, how to add safety checks and how to deploy all of it. Teams often spend weeks wiring these parts together before they can judge whether the idea works.
Blueprints give a tested arrangement of these parts with code, configuration and deployment files, so a team can run a complete example quickly and then replace pieces with its own data, models and services.4
How it works
Each blueprint is a GitHub repository with application code, configuration and deployment assets. Using the RAG Blueprint as an example:
- Ingestion: an ingestor service extracts multimodal content with NeMo Retriever models and writes embeddings to a vector database (Milvus or Elasticsearch).
- Retrieval and generation: a RAG server retrieves and reranks passages and calls an LLM NIM to answer; an optional agentic mode uses a LangGraph plan-and-execute pipeline for multi-hop questions.
- Safety and evaluation: optional NeMo Guardrails, plus guides for accuracy and performance evaluation and observability.
- Interfaces: a reference web UI, a Python package and REST APIs.
Deployment options include Docker with self-hosted models, Docker with NVIDIA-hosted model endpoints, Helm on Kubernetes or OpenShift, the NIM Operator, and a containerless lite mode. NVIDIA also states that blueprints can be launched in one click with NVIDIA Launchables or downloaded for local PCs and workstations.13
Diagram as a list
Applications & solutions
- Reference UI or your applicationAsks questions and shows answersConnects to RAG server (optional agentic mode)
- RAG server (optional agentic mode)Retrieves context and orchestrates the answerConnects to Embedding and reranking NIMs, LLM NIM, NeMo Guardrails (optional)
- Ingestor server (NeMo Retriever extraction)Extracts text, tables, charts and audio from documentsConnects to Embedding and reranking NIMs
- NeMo Guardrails (optional)Applies safety and topic rules
Inference & runtime software
- Embedding and reranking NIMsTurn content and queries into vectors and rank resultsConnects to Vector database (Milvus or Elasticsearch)
- LLM NIMGenerates the grounded answerConnects to NVIDIA GPUs or NVIDIA-hosted endpoints
Operations & orchestration
- Vector database (Milvus or Elasticsearch)Stores and searches embeddings
- Docker, Helm or NIM OperatorDeploys the servicesConnects to Embedding and reranking NIMs, LLM NIM
Accelerated computing
- NVIDIA GPUs or NVIDIA-hosted endpointsRun the models
Capabilities
Working reference application46
Each blueprint ships a sample application with source code, documentation and, for RAG, a reference UI.
Why it matters: Teams see a complete, running example before writing their own code.
Limits: We recommend hardening, testing and assigning an owner to the sample code before production use.
Multiple deployment paths3
Docker with self-hosted or NVIDIA-hosted models, Helm on Kubernetes or OpenShift, the NIM Operator, and library or lite modes.
Why it matters: The same workflow can move from a laptop test to a cluster.
Limits: Self-hosted deployments are large: the RAG Blueprint asks for about 200 GB of disk and long first downloads.
Built on NIM and NVIDIA libraries5
Blueprints combine NIM microservices with NeMo, Omniverse or CUDA-X libraries and partner microservices.
Why it matters: Shows the intended way NVIDIA's components fit together.
Limits: Ties the reference design to NVIDIA components; swapping them takes extra work.
Designed to be modified34
Components such as models, prompts, vector databases and pipeline modes are configurable; NVIDIA says blueprints are meant to be modified and extended.
Why it matters: Teams can replace parts with their own data and services.
Limits: Changes move you away from the tested configuration.
Open source code6
Blueprint code is public on GitHub; the RAG Blueprint is licensed under Apache 2.0.
Why it matters: Code can be inspected, forked and reused.
Limits: Each model used by the blueprint has its own license (for the RAG Blueprint: NVIDIA community model licenses, the Llama 3.1 or 3.2 Community License, or Apache 2.0).
Practical use cases
Employees cannot find answers spread across PDFs, slides and reports.
- Approach
- Deploy the RAG Blueprint, ingest a sample of company documents and test grounded question answering.
- Role of NVIDIA Blueprints
- The blueprint supplies ingestion, retrieval, generation and UI as a working starting point.
- Data, infrastructure and skills
- Documents, GPUs or hosted endpoints, and a vector database.
- Type of benefit
- Faster time to prototype
- Caveats
- Retrieval quality depends on document quality and access controls you must add.
- First step
- Run the Docker deployment with NVIDIA-hosted models on a small document set.
Sources 6
Staff spend hours reviewing recorded or live video for specific events.
- Approach
- Use the Video Search and Summarization blueprint to search, summarize and question video with vision language models.
- Role of NVIDIA Blueprints
- The blueprint provides the video ingestion, VLM, LLM and RAG pipeline.
- Data, infrastructure and skills
- Video sources, GPUs sized per the blueprint's hardware requirements and privacy review.
- Type of benefit
- Review automation
- Caveats
- Video analytics raises privacy and consent obligations.
- First step
- Read the VSS hardware requirements and test on archived footage.
Sources 7
Analysts need research reports that combine internal data and web sources.
- Approach
- Start from NVIDIA's research agent blueprint (AI-Q / deep researcher) and connect it to internal data.
- Role of NVIDIA Blueprints
- The blueprint provides an agent that plans, retrieves and writes reports.
- Data, infrastructure and skills
- Data connectors, model access and evaluation of report accuracy.
- Type of benefit
- Analyst productivity
- Caveats
- Generated reports need human review before use in decisions.
- First step
- Run the reference agent on a public topic and review its sources.
Sources 9
Works with
Complementary tools
- NVIDIA AI WorkbenchAI Blueprint example projects (PDF to Podcast, RAG, AI-Q) run in Workbench.
- NVIDIA AI EnterpriseNVIDIA presents Blueprints with AI Enterprise for moving to production.
- NVIDIA BioNeMoNVIDIA publishes a BioNeMo Blueprint for generative protein binder design.
Reference implementation for
- NVIDIA NIMBlueprints show NIM microservices in complete applications.
- NVIDIA NeMoBlueprints such as AI-Q and Data Flywheel use NeMo components.
- NVIDIA cuOptThe portfolio-optimization blueprint is built on cuOpt.
Relationship labels follow NVIDIA's documentation. "Alternative approaches" does not mean one is better: each profile says when it fits.
Getting started
Pick a blueprint2
Browse the NVIDIA-AI-Blueprints GitHub organization and choose the repository closest to your use case.
Check: The blueprint's support matrix matches the hardware you can use.
Get keys and choose a mode3
Get an NVIDIA API key and decide between NVIDIA-hosted models (few local GPUs) and self-hosted models.
Check: You have the credentials and disk space the chosen mode needs.
Deploy the RAG Blueprint with Docker3
Follow the Docker guide; a first self-hosted deployment can take 15 to 30 minutes while models download.
Check: The reference UI loads and answers a question about the sample documents.
Swap in your data
Ingest a small set of your own documents and adjust prompts or models.
Check: Answers cite passages from your documents.
Official resources
Could this technology help you?
Describe your project to the Solution Architect. It starts with NVIDIA Blueprints as context but recommends independently, including when you do not need it.
Sources
Each statement above links to the source it comes from. Labels say who reported it.
- NVIDIA NIM for Developers (opens in a new tab)
- NVIDIA AI Blueprints GitHub organization (opens in a new tab)
- NVIDIA RAG Blueprint documentation (opens in a new tab)
- NVIDIA Blog: NVIDIA Blueprints launch post (with 2024 rename note) (opens in a new tab)
- NVIDIA AI Enterprise product page (opens in a new tab)
- NVIDIA RAG Blueprint GitHub repository (opens in a new tab)
- NVIDIA Video Search and Summarization Blueprint repository (opens in a new tab)
- NVIDIA NIM Microservices product page (opens in a new tab)
- NVIDIA NeMo product page (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.