How to build a RAG system: the pipeline, the stacks, and where NVIDIA's blueprints fit

Retrieval-augmented generation is how an organisation's own documents become answers. A practical guide to the ingestion and query pipeline, the open-source, NVIDIA and cloud stacks you can build it from, what the NVIDIA AI Blueprint for RAG actually contains, and how to do it well with Bangla documents.

CompTech Team
  • AI / ML
  • RAG
  • LLM
  • NVIDIA
The Wren Library at Trinity College, Cambridge, with its rows of bookcases and a chequered floor

Every organisation already has the knowledge it needs. It sits in circulars, contracts, manuals, audit reports, tender documents and years of email, mostly as PDF and scanned paper. A large language model on its own knows none of it, and will happily invent an answer. Retrieval-augmented generation, RAG, is the technique that fixes this: find the right passages first, then let the model answer from them, with citations.

RAG is also the first AI project most organisations attempt, and the one most often abandoned after a demo. The demo takes a day. The production system takes a plan. This post gives you that plan: the pipeline that works, the technology choices at every layer, what NVIDIA's RAG blueprint actually packages, and what changes when the documents are in Bangla.

Why RAG, and not fine-tuning or a bigger context window

Three ways exist to give a model your knowledge, and they solve different problems:

Approach What it does Best for Weakness
RAG Retrieves relevant passages at question time and puts them in the prompt Facts that change, large document sets, answers that must cite a source, permission-aware access Only as good as retrieval; needs an ingestion pipeline
Fine-tuning Trains the model's weights on your examples Style, format, domain vocabulary, a specialised task Does not reliably teach facts; stale the day after training; expensive to repeat
Long context Pastes whole documents into a model with a 128K to 1M token window Small, stable corpora; one-off analysis of a few files Cost and latency grow with every token; accuracy drops in the middle of long prompts

In practice the answer is usually RAG, sometimes with a lightly fine-tuned model behind it. A bank's policy manual changes monthly, a ministry's circulars weekly, and nobody can retrain a model on that schedule. Retrieval makes the knowledge live.

The pipeline that works

Diagram of a RAG pipeline with an ingestion row (sources, parse and OCR, chunk, embed, index) and a query row (question, input guardrail, hybrid retrieve, rerank, LLM, output guardrail) and an evaluation loop feeding back
The shape of every serious RAG system. The blue boxes are the GPU inference steps; the rest is data engineering, which is where most of the effort goes.

Ingestion runs whenever documents change, and it decides the quality ceiling of everything after it. Most teams that say "RAG doesn't work" have a parsing problem. Four rules:

  1. Parse the layout, not just the text. A scanned circular has headers, stamps, two-column text and a table of fees. A naive PDF-to-text conversion scrambles all of it. Use a layout-aware parser (Docling, Unstructured, NVIDIA's NeMo Retriever extraction, or a cloud document service) that emits headings, paragraphs and tables separately, and OCR for scans.
  2. Chunk by structure. Split on headings and paragraphs into chunks of roughly 500 to 1,000 tokens, and carry the document title and section heading into every chunk. Sentence-level chunks lose context; page-level chunks dilute it.
  3. Embed densely and sparsely. A dense embedding model captures meaning; a sparse keyword index (BM25) captures exact terms such as account numbers, section references and product codes. Store both.
  4. Label every chunk with who may see it. Access control belongs in the index, not in the prompt.
A bound newspaper volume on a Qidenus book scanner, lit for high-resolution capture
For most organisations in Bangladesh the first RAG task is digitisation. Scan quality and OCR decide what the model can ever know. Photo: Qidenus, CC BY 4.0.

The query path runs on every question and has to be fast:

  1. Guard the input. Reject off-topic or unsafe questions and strip personal data before anything is logged.
  2. Retrieve widely, then rerank. Pull the top 50 chunks by combining vector and keyword scores (reciprocal rank fusion is the usual method), then let a cross-encoder reranker choose the best five. Reranking is the single cheapest accuracy improvement available; skipping it is the most common mistake.
  3. Generate from the context only. The prompt instructs the model to answer from the retrieved passages, to cite them, and to say when the answer is not in the documents.
  4. Guard the output. Check that the answer is grounded in the sources before it reaches the user.

Evaluation is the half that separates a demo from a product. Build a test set of 100 to 300 real questions with known good answers before writing application code. Measure retrieval (did the right chunk make the top five?) separately from generation (is the answer faithful to the chunk?). Tools such as RAGAS, DeepEval and Arize Phoenix automate the scoring; Langfuse or LangSmith record every live question so the test set grows from real use. Change one thing at a time and re-run.

The stacks: open source, NVIDIA, cloud

Grid of RAG stack layers (apps, orchestration, LLM serving, embedding and rerank, vector index, document parsing, guardrails and evaluation, infrastructure) against three columns: open source self-hosted, NVIDIA AI Blueprint for RAG, and cloud managed services, with product names in each cell
Every layer has three kinds of answer. Most real deployments mix them: open-source orchestration and vector database, NVIDIA NIMs for the models, a cloud model for the hardest questions.

Open source, self-hosted is where the field moves fastest and where data never leaves the building. LangChain with LangGraph, LlamaIndex and Haystack are the three orchestration frameworks; LlamaIndex is the most retrieval-focused, Haystack the most production-minded, LangChain the broadest. For the index, Milvus, Qdrant and Weaviate are purpose-built vector databases; pgvector adds vectors to a Postgres you already run; OpenSearch and Elasticsearch give hybrid search in a system your operations team already knows. For models, vLLM or SGLang serve open weights such as Llama, Qwen, Gemma and Mistral; bge-m3 is the default multilingual embedding model and bge-reranker-v2-m3 its reranker. Dify, RAGFlow, AnythingLLM and Open WebUI package the whole thing with a user interface and are the right starting point for a pilot.

NVIDIA's AI Blueprint for RAG is a reference architecture that NVIDIA ships as containers and Helm charts, designed to run on its GPUs on-premises or in any cloud. It is worth understanding precisely, because "we use the NVIDIA blueprint" is becoming a common line in tenders:

  • NeMo Retriever extraction NIMs parse PDFs with separate models for page layout, table structure, charts and OCR, so tables and figures become searchable text. NVIDIA claims this runs about 15 times faster than CPU-based parsing.
  • NeMo Retriever embedding and reranking NIMs, currently the Llama 3.2 based NV-EmbedQA and NV-RerankQA models and their Nemotron successors, both multilingual.
  • A Llama Nemotron LLM NIM for generation, swappable for any other NIM or an external model.
  • Milvus accelerated with cuVS, NVIDIA's GPU vector-search library, as the default index, with Elasticsearch as an alternative. Hybrid dense-plus-sparse search is built in.
  • LangChain for orchestration, MinIO for object storage, Redis for caching, and a sample React interface.
  • NeMo Guardrails with content-safety and topic-control NIMs, and an optional vision-language model for questions about images.

What you get is a tested combination that deploys in an afternoon with Docker Compose or on Kubernetes with Helm, and that NVIDIA supports commercially through NVIDIA AI Enterprise. What you do not get is a finished application: your parsers, your access control, your evaluation set and your Bangla testing are still yours to do. Other NVIDIA blueprints, such as the AI virtual assistant for customer service and the AI-Q research assistant, are built on top of this one.

An NVIDIA HGX B200 eight-GPU baseboard
The blueprint runs on anything from a single L40S to an HGX B200 server like this. A department pilot needs one GPU; an organisation-wide assistant with a 70B model needs two to four. Photo: Pokiiri, CC BY-SA 4.0.

Cloud managed services remove the operations work in exchange for data residency and per-token cost. Amazon Bedrock Knowledge Bases, Azure AI Search with Azure AI Foundry, and Google's Vertex AI RAG Engine each give you ingestion, vector storage, retrieval and a model behind one API, with GraphRAG and agentic multi-step retrieval arriving on all three during 2025 and 2026. For a Bangladeshi bank or ministry the obstacle is rarely the technology; it is that the documents cannot leave the country, which is why the first two columns matter more here than elsewhere.

Enterprise platforms sit between the columns. Red Hat OpenShift AI, Dell AI Factory with NVIDIA and HPE Private Cloud AI bundle the NVIDIA blueprint, or an equivalent open-source stack, with the Kubernetes platform, support and hardware, for organisations that want one vendor accountable for the whole thing.

A long aisle of library shelves receding into the distance
Retrieval is a library problem before it is an AI problem: if the catalogue is wrong, the best reader in the world finds the wrong book. Photo: James JA Kruk, CC0.

Beyond basic RAG

Three extensions are now mature enough to use:

  • Agentic RAG lets the model decide how to search: rewrite the question, run several retrievals, consult a database as well as documents, and check its own answer before replying. LangGraph, LlamaIndex workflows and NVIDIA's NeMo Agent Toolkit are the usual frameworks. It costs more tokens per question and pays off on complex, multi-part questions.
  • GraphRAG, from Microsoft Research, with lighter variants such as LightRAG, extracts entities and relationships from the corpus into a knowledge graph and retrieves along it. It answers "what are the themes across these thousand reports" questions that chunk retrieval cannot.
  • Multimodal RAG indexes figures, charts and scanned forms as images and retrieves them with a vision-language model. NVIDIA's extraction NIMs and the VLM option in the blueprint are built for it, and it matters for engineering drawings, invoices and the stamped, handwritten forms common in government files.

Doing it in Bangla

Most of the stack is language-neutral, but four layers are not:

  • OCR. Tesseract's Bengali model is adequate for clean print and poor on stamps, handwriting and low-contrast photocopies. PaddleOCR and commercial document services do better. Budget for a correction pass on the documents that matter most.
  • Embeddings. Use a multilingual model. bge-m3 covers more than 100 languages including Bengali and leads the MIRACL multilingual retrieval benchmark among open models, ahead of multilingual-e5; Qwen3-Embedding and NVIDIA's NV-EmbedQA are also multilingual. English-only models such as the older OpenAI embeddings will silently fail on Bangla text.
  • Chunking. Token counts differ: Bangla text uses two to three times as many tokens per word as English in most tokenisers, so a 1,000-token chunk holds much less. Chunk by paragraph, not by token count, and check what fits.
  • Generation. Llama, Gemma and Qwen models answer in Bangla with varying fluency; Gemma 3 and Qwen3 are currently the strongest open options for Bangla, and the choice should be made on your own test set, not on English benchmarks. Mixed Bangla-English questions, which are how people actually type, deserve test cases of their own.
The entrance and signboard of the National Archives Building in Agargaon, Dhaka
The National Archives in Agargaon. Government records in Bangladesh are overwhelmingly paper and Bangla, which is exactly where the hard and valuable RAG work is. Photo: Wasiul Bahar, CC BY-SA 4.0.

Sizing and cost

The GPU bill for RAG is smaller than people expect, because retrieval does most of the work and the model only reads five chunks per question.

Scale Typical use Models Hardware
Pilot One department, 50 users, 10,000 documents 8B-class LLM in FP8, bge-m3, bge-reranker One L40S or RTX PRO 6000 in an existing server
Department 500 users, 200,000 documents, Bangla and English 27B to 32B LLM in FP8, same retrievers, separate extraction GPU for ingestion One two-GPU server (H100 or L40S class)
Enterprise Thousands of users, millions of pages, several assistants 70B LLM in FP8 or FP4, reranker and embeddings on their own GPU, guardrail models One eight-GPU HGX server, Kubernetes, high-availability vector database

Ingestion is bursty: parsing a million pages with GPU extraction takes days on one GPU and hours on eight, then nothing until the next batch. Plan it as a job, not a service.

What this means for Bangladesh

The organisations we talk to want the same three things: answers from their own documents, in Bangla and English, without the documents leaving their premises. RAG on local GPUs is the only architecture that delivers all three, and the pieces are mature. The work that remains is unglamorous and decisive: scanning, OCR correction, access-control labelling and a test set of real questions. Done in that order, a department-level assistant is a ten-week project, and the same pipeline scales to a ministry.

CompTech builds RAG systems on NVIDIA infrastructure for banks, government and enterprises in Bangladesh, from a single-server pilot on the NVIDIA blueprint to clustered deployments on OpenShift. If you have a document set and a list of questions people keep asking, talk to us; the first step is a two-week evaluation on your own documents, not a hardware quote.


Cover photo of the Wren Library, Trinity College, Cambridge, from Wikimedia Commons, CC BY-SA 2.0. Other photographs are credited in their captions and used under their Creative Commons licences. Diagrams by CompTech. Product names and blueprint contents as published by their vendors in October 2026.