Skip to content

Build an enterprise AI assistant

An assistant that answers staff questions from internal documents and shows its sources, while prompts and data stay under company control. The usual pattern is retrieval over your own content, a served language model, guardrails and permission-aware access.

The problem

Staff lose time searching wikis, policy PDFs, tickets and shared drives, and a general chatbot cannot answer from material it has never seen. Pasting internal documents into a public chat service raises questions about confidentiality, retention and who can see what.

The engineering problem is retrieval and control. The assistant has to find the right passages across many file formats, answer only from what the person asking is allowed to read, show where each answer came from, and decline or hand off when the sources do not cover the question.

The approach

The common pattern is retrieval-augmented generation (RAG). Documents are parsed, split into passages and converted into embeddings stored in a vector database. When someone asks a question, the system retrieves the closest passages, reranks them and asks a language model to answer from them with citations. Guardrails check questions and answers for off-topic, unsafe or policy-breaking content.

In NVIDIA's ecosystem, the RAG Blueprint is reference code for this flow: multimodal extraction, retrieval with NeMo Retriever models, generation, guardrails and a sample user interface, with Milvus or Elasticsearch as the vector store. Models run as NIM microservices on your own GPUs, or on NVIDIA-hosted endpoints during a pilot.

A simpler route exists. If your documents may be processed by a cloud provider, a managed assistant or a hosted model API with a managed search index needs no GPUs and no container operations. Self-hosting makes sense when data must stay on your network, when steady volume makes per-call pricing expensive, or when you need to control the model version.123

Conceptual architecture

Build an enterprise AI assistant: conceptual architectureApplications &solutionsModels & frameworksInference & runtimesoftwareOperations &orchestrationAcceleratedcomputingStaff chat interface: Takes questions and shows answers with links to the source passagesStaff chat interfaceRAG orchestrator: Retrieves passages the user may see, builds the prompt and returns a cited answerRAG orchestratorIngestion and extraction: Parses PDFs, slides, tables and images into passages with permission metadataIngestion and extractionEmbedding and reranking models: Turn passages and questions into vectors and reorder results by relevanceEmbedding and rerankingmodelsLanguage model served as a NIM: Generates the answer from retrieved passages behind an OpenAI-compatible APILanguage model served as aNIMGuardrails (NeMo Guardrails): Checks questions and answers against topic, safety and policy rulesGuardrails (NeMo Guardrails)Vector database (Milvus or Elasticsearch): Stores embeddings and access metadata for filtered searchVector database (Milvus orElasticsearch)NVIDIA GPUs on premises or in the cloud: Run the retrieval and generation modelsNVIDIA GPUs on premises orin the cloud
Diagram as a list
  1. Applications & solutions

    • Staff chat interfaceTakes questions and shows answers with links to the source passagesConnects to Guardrails (NeMo Guardrails)
    • RAG orchestratorRetrieves passages the user may see, builds the prompt and returns a cited answerConnects to Embedding and reranking models, Language model served as a NIM
    • Ingestion and extractionParses PDFs, slides, tables and images into passages with permission metadataConnects to Embedding and reranking models, Vector database (Milvus or Elasticsearch)
  2. Models & frameworks

    • Embedding and reranking modelsTurn passages and questions into vectors and reorder results by relevanceConnects to Vector database (Milvus or Elasticsearch)
  3. Inference & runtime software

    • Language model served as a NIMGenerates the answer from retrieved passages behind an OpenAI-compatible APIConnects to NVIDIA GPUs on premises or in the cloud
  4. Operations & orchestration

    • Guardrails (NeMo Guardrails)Checks questions and answers against topic, safety and policy rulesConnects to RAG orchestrator
    • Vector database (Milvus or Elasticsearch)Stores embeddings and access metadata for filtered search
  5. Accelerated computing

    • NVIDIA GPUs on premises or in the cloudRun the retrieval and generation models
Conceptual: one common way to arrange the parts, not a required design.1

Technologies and their roles

  • blueprints1

    Reference implementation

    The RAG Blueprint provides working ingestion, retrieval, guardrail and interface code to adapt, so the team does not start from an empty repository.

  • nim3

    Model serving

    Serves the generation and retrieval models behind standard APIs on infrastructure you choose, so prompts and documents stay in your environment.

  • nemotron4

    Open models

    Nemotron includes retriever models for document understanding, embeddings and reranking, plus reasoning models, with open weights on Hugging Face.

  • nemo56

    Guardrails and evaluation

    NeMo Guardrails adds programmable safety and topic rules, and NeMo Evaluator can score answers before and after changes.

  • ai-enterprise37

    Production license and support

    NVIDIA directs production use of NIM to an AI Enterprise license, which adds supported release branches and security patching.

What you need first2

  • A curated document set with named owners and a record of who may read each source
  • An identity provider so retrieval can filter results by the user's permissions
  • A test set of real staff questions with expected answers and sources
  • Python and container skills; Kubernetes skills for a multi-user production service
  • NVIDIA GPUs with enough memory for the chosen models, or hosted endpoints for a pilot
  • About 200 GB of free disk for a self-hosted RAG Blueprint deployment, per NVIDIA's documentation

Risks and how to reduce them

Answers that sound right but are not supported by the sources
Require citations, measure groundedness on a fixed test set and let the assistant say when it does not know.
Users see content they are not entitled to
Carry document permissions into the index, filter at retrieval time and test with accounts from different roles.
Personal or confidential data in prompts and logs
Set retention rules for prompts and logs, mask personal data where possible and keep self-hosted models inside your network.
Outdated answers from stale content
Schedule re-ingestion and remove retired documents from the index.
Different license terms for code and models1
The repository is Apache 2.0, but NVIDIA's terms of use for the blueprint and several model licenses (NVIDIA community model licenses and the Llama 3.1 and 3.2 community licenses) also apply; review all of them before production.

Related

Sources

  1. NVIDIA RAG Blueprint GitHub repository (opens in a new tab)NVIDIA · Vendor-reported
  2. NVIDIA RAG Blueprint documentation (opens in a new tab)NVIDIA · Vendor-reported
  3. NVIDIA NIM Microservices product page (opens in a new tab)NVIDIA · Vendor-reported
  4. NVIDIA Nemotron foundation models (opens in a new tab)NVIDIA · Vendor-reported
  5. NVIDIA NeMo documentation hub (opens in a new tab)NVIDIA · Vendor-reported
  6. NVIDIA NeMo product page (opens in a new tab)NVIDIA · Vendor-reported
  7. NVIDIA AI Enterprise product page (opens in a new tab)NVIDIA · Vendor-reported

Fill out the form below to request your copy.

Name(Required)