Build an enterprise AI assistant
An assistant that answers staff questions from internal documents and shows its sources, while prompts and data stay under company control. The usual pattern is retrieval over your own content, a served language model, guardrails and permission-aware access.
The problem
Staff lose time searching wikis, policy PDFs, tickets and shared drives, and a general chatbot cannot answer from material it has never seen. Pasting internal documents into a public chat service raises questions about confidentiality, retention and who can see what.
The engineering problem is retrieval and control. The assistant has to find the right passages across many file formats, answer only from what the person asking is allowed to read, show where each answer came from, and decline or hand off when the sources do not cover the question.
The approach
The common pattern is retrieval-augmented generation (RAG). Documents are parsed, split into passages and converted into embeddings stored in a vector database. When someone asks a question, the system retrieves the closest passages, reranks them and asks a language model to answer from them with citations. Guardrails check questions and answers for off-topic, unsafe or policy-breaking content.
In NVIDIA's ecosystem, the RAG Blueprint is reference code for this flow: multimodal extraction, retrieval with NeMo Retriever models, generation, guardrails and a sample user interface, with Milvus or Elasticsearch as the vector store. Models run as NIM microservices on your own GPUs, or on NVIDIA-hosted endpoints during a pilot.
A simpler route exists. If your documents may be processed by a cloud provider, a managed assistant or a hosted model API with a managed search index needs no GPUs and no container operations. Self-hosting makes sense when data must stay on your network, when steady volume makes per-call pricing expensive, or when you need to control the model version.123
Conceptual architecture
Diagram as a list
Applications & solutions
- Staff chat interfaceTakes questions and shows answers with links to the source passagesConnects to Guardrails (NeMo Guardrails)
- RAG orchestratorRetrieves passages the user may see, builds the prompt and returns a cited answerConnects to Embedding and reranking models, Language model served as a NIM
- Ingestion and extractionParses PDFs, slides, tables and images into passages with permission metadataConnects to Embedding and reranking models, Vector database (Milvus or Elasticsearch)
Models & frameworks
- Embedding and reranking modelsTurn passages and questions into vectors and reorder results by relevanceConnects to Vector database (Milvus or Elasticsearch)
Inference & runtime software
- Language model served as a NIMGenerates the answer from retrieved passages behind an OpenAI-compatible APIConnects to NVIDIA GPUs on premises or in the cloud
Operations & orchestration
- Guardrails (NeMo Guardrails)Checks questions and answers against topic, safety and policy rulesConnects to RAG orchestrator
- Vector database (Milvus or Elasticsearch)Stores embeddings and access metadata for filtered search
Accelerated computing
- NVIDIA GPUs on premises or in the cloudRun the retrieval and generation models
Technologies and their roles
blueprints1
Reference implementation
The RAG Blueprint provides working ingestion, retrieval, guardrail and interface code to adapt, so the team does not start from an empty repository.
nim3
Model serving
Serves the generation and retrieval models behind standard APIs on infrastructure you choose, so prompts and documents stay in your environment.
nemotron4
Open models
Nemotron includes retriever models for document understanding, embeddings and reranking, plus reasoning models, with open weights on Hugging Face.
nemo56
Guardrails and evaluation
NeMo Guardrails adds programmable safety and topic rules, and NeMo Evaluator can score answers before and after changes.
ai-enterprise37
Production license and support
NVIDIA directs production use of NIM to an AI Enterprise license, which adds supported release branches and security patching.
What you need first2
- A curated document set with named owners and a record of who may read each source
- An identity provider so retrieval can filter results by the user's permissions
- A test set of real staff questions with expected answers and sources
- Python and container skills; Kubernetes skills for a multi-user production service
- NVIDIA GPUs with enough memory for the chosen models, or hosted endpoints for a pilot
- About 200 GB of free disk for a self-hosted RAG Blueprint deployment, per NVIDIA's documentation
Risks and how to reduce them
- Answers that sound right but are not supported by the sources
- Require citations, measure groundedness on a fixed test set and let the assistant say when it does not know.
- Users see content they are not entitled to
- Carry document permissions into the index, filter at retrieval time and test with accounts from different roles.
- Personal or confidential data in prompts and logs
- Set retention rules for prompts and logs, mask personal data where possible and keep self-hosted models inside your network.
- Outdated answers from stale content
- Schedule re-ingestion and remove retired documents from the index.
- Different license terms for code and models1
- The repository is Apache 2.0, but NVIDIA's terms of use for the blueprint and several model licenses (NVIDIA community model licenses and the Llama 3.1 and 3.2 community licenses) also apply; review all of them before production.
Related
Sources
- NVIDIA RAG Blueprint GitHub repository (opens in a new tab)
- NVIDIA RAG Blueprint documentation (opens in a new tab)
- NVIDIA NIM Microservices product page (opens in a new tab)
- NVIDIA Nemotron foundation models (opens in a new tab)
- NVIDIA NeMo documentation hub (opens in a new tab)
- NVIDIA NeMo product page (opens in a new tab)
- NVIDIA AI Enterprise product page (opens in a new tab)
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.