
Every organisation already has the knowledge it needs. It sits in circulars, contracts, manuals, audit reports, tender documents and years of email, mostly as PDF and scanned paper. A large language model on its own knows none of it, and will happily invent an answer. Retrieval-augmented generation, RAG, is the technique that fixes this: find the right passages first, then let the model answer from them, with citations.
RAG is also the first AI project most organisations attempt, and the one most often abandoned after a demo. The demo takes a day. The production system takes a plan. This post gives you that plan: the pipeline that works, the technology choices at every layer, what NVIDIA's RAG blueprint actually packages, and what changes when the documents are in Bangla.
Why RAG, and not fine-tuning or a bigger context window
Three ways exist to give a model your knowledge, and they solve different problems:
| Approach | What it does | Best for | Weakness |
|---|---|---|---|
| RAG | Retrieves relevant passages at question time and puts them in the prompt | Facts that change, large document sets, answers that must cite a source, permission-aware access | Only as good as retrieval; needs an ingestion pipeline |
| Fine-tuning | Trains the model's weights on your examples | Style, format, domain vocabulary, a specialised task | Does not reliably teach facts; stale the day after training; expensive to repeat |
| Long context | Pastes whole documents into a model with a 128K to 1M token window | Small, stable corpora; one-off analysis of a few files | Cost and latency grow with every token; accuracy drops in the middle of long prompts |
In practice the answer is usually RAG, sometimes with a lightly fine-tuned model behind it. A bank's policy manual changes monthly, a ministry's circulars weekly, and nobody can retrain a model on that schedule. Retrieval makes the knowledge live.
The pipeline that works
Ingestion runs whenever documents change, and it decides the quality ceiling of everything after it. Most teams that say "RAG doesn't work" have a parsing problem. Four rules:
- Parse the layout, not just the text. A scanned circular has headers, stamps, two-column text and a table of fees. A naive PDF-to-text conversion scrambles all of it. Use a layout-aware parser (Docling, Unstructured, NVIDIA's NeMo Retriever extraction, or a cloud document service) that emits headings, paragraphs and tables separately, and OCR for scans.
- Chunk by structure. Split on headings and paragraphs into chunks of roughly 500 to 1,000 tokens, and carry the document title and section heading into every chunk. Sentence-level chunks lose context; page-level chunks dilute it.
- Embed densely and sparsely. A dense embedding model captures meaning; a sparse keyword index (BM25) captures exact terms such as account numbers, section references and product codes. Store both.
- Label every chunk with who may see it. Access control belongs in the index, not in the prompt.
The query path runs on every question and has to be fast:
- Guard the input. Reject off-topic or unsafe questions and strip personal data before anything is logged.
- Retrieve widely, then rerank. Pull the top 50 chunks by combining vector and keyword scores (reciprocal rank fusion is the usual method), then let a cross-encoder reranker choose the best five. Reranking is the single cheapest accuracy improvement available; skipping it is the most common mistake.
- Generate from the context only. The prompt instructs the model to answer from the retrieved passages, to cite them, and to say when the answer is not in the documents.
- Guard the output. Check that the answer is grounded in the sources before it reaches the user.
Evaluation is the half that separates a demo from a product. Build a test set of 100 to 300 real questions with known good answers before writing application code. Measure retrieval (did the right chunk make the top five?) separately from generation (is the answer faithful to the chunk?). Tools such as RAGAS, DeepEval and Arize Phoenix automate the scoring; Langfuse or LangSmith record every live question so the test set grows from real use. Change one thing at a time and re-run.
The stacks: open source, NVIDIA, cloud
Open source, self-hosted is where the field moves fastest and where data never leaves the building. LangChain with LangGraph, LlamaIndex and Haystack are the three orchestration frameworks; LlamaIndex is the most retrieval-focused, Haystack the most production-minded, LangChain the broadest. For the index, Milvus, Qdrant and Weaviate are purpose-built vector databases; pgvector adds vectors to a Postgres you already run; OpenSearch and Elasticsearch give hybrid search in a system your operations team already knows. For models, vLLM or SGLang serve open weights such as Llama, Qwen, Gemma and Mistral; bge-m3 is the default multilingual embedding model and bge-reranker-v2-m3 its reranker. Dify, RAGFlow, AnythingLLM and Open WebUI package the whole thing with a user interface and are the right starting point for a pilot.
NVIDIA's AI Blueprint for RAG is a reference architecture that NVIDIA ships as containers and Helm charts, designed to run on its GPUs on-premises or in any cloud. It is worth understanding precisely, because "we use the NVIDIA blueprint" is becoming a common line in tenders:
- NeMo Retriever extraction NIMs parse PDFs with separate models for page layout, table structure, charts and OCR, so tables and figures become searchable text. NVIDIA claims this runs about 15 times faster than CPU-based parsing.
- NeMo Retriever embedding and reranking NIMs, currently the Llama 3.2 based NV-EmbedQA and NV-RerankQA models and their Nemotron successors, both multilingual.
- A Llama Nemotron LLM NIM for generation, swappable for any other NIM or an external model.
- Milvus accelerated with cuVS, NVIDIA's GPU vector-search library, as the default index, with Elasticsearch as an alternative. Hybrid dense-plus-sparse search is built in.
- LangChain for orchestration, MinIO for object storage, Redis for caching, and a sample React interface.
- NeMo Guardrails with content-safety and topic-control NIMs, and an optional vision-language model for questions about images.
What you get is a tested combination that deploys in an afternoon with Docker Compose or on Kubernetes with Helm, and that NVIDIA supports commercially through NVIDIA AI Enterprise. What you do not get is a finished application: your parsers, your access control, your evaluation set and your Bangla testing are still yours to do. Other NVIDIA blueprints, such as the AI virtual assistant for customer service and the AI-Q research assistant, are built on top of this one.
Cloud managed services remove the operations work in exchange for data residency and per-token cost. Amazon Bedrock Knowledge Bases, Azure AI Search with Azure AI Foundry, and Google's Vertex AI RAG Engine each give you ingestion, vector storage, retrieval and a model behind one API, with GraphRAG and agentic multi-step retrieval arriving on all three during 2025 and 2026. For a Bangladeshi bank or ministry the obstacle is rarely the technology; it is that the documents cannot leave the country, which is why the first two columns matter more here than elsewhere.
Enterprise platforms sit between the columns. Red Hat OpenShift AI, Dell AI Factory with NVIDIA and HPE Private Cloud AI bundle the NVIDIA blueprint, or an equivalent open-source stack, with the Kubernetes platform, support and hardware, for organisations that want one vendor accountable for the whole thing.
Beyond basic RAG
Three extensions are now mature enough to use:
- Agentic RAG lets the model decide how to search: rewrite the question, run several retrievals, consult a database as well as documents, and check its own answer before replying. LangGraph, LlamaIndex workflows and NVIDIA's NeMo Agent Toolkit are the usual frameworks. It costs more tokens per question and pays off on complex, multi-part questions.
- GraphRAG, from Microsoft Research, with lighter variants such as LightRAG, extracts entities and relationships from the corpus into a knowledge graph and retrieves along it. It answers "what are the themes across these thousand reports" questions that chunk retrieval cannot.
- Multimodal RAG indexes figures, charts and scanned forms as images and retrieves them with a vision-language model. NVIDIA's extraction NIMs and the VLM option in the blueprint are built for it, and it matters for engineering drawings, invoices and the stamped, handwritten forms common in government files.
Doing it in Bangla
Most of the stack is language-neutral, but four layers are not:
- OCR. Tesseract's Bengali model is adequate for clean print and poor on stamps, handwriting and low-contrast photocopies. PaddleOCR and commercial document services do better. Budget for a correction pass on the documents that matter most.
- Embeddings. Use a multilingual model. bge-m3 covers more than 100 languages including Bengali and leads the MIRACL multilingual retrieval benchmark among open models, ahead of multilingual-e5; Qwen3-Embedding and NVIDIA's NV-EmbedQA are also multilingual. English-only models such as the older OpenAI embeddings will silently fail on Bangla text.
- Chunking. Token counts differ: Bangla text uses two to three times as many tokens per word as English in most tokenisers, so a 1,000-token chunk holds much less. Chunk by paragraph, not by token count, and check what fits.
- Generation. Llama, Gemma and Qwen models answer in Bangla with varying fluency; Gemma 3 and Qwen3 are currently the strongest open options for Bangla, and the choice should be made on your own test set, not on English benchmarks. Mixed Bangla-English questions, which are how people actually type, deserve test cases of their own.
Sizing and cost
The GPU bill for RAG is smaller than people expect, because retrieval does most of the work and the model only reads five chunks per question.
| Scale | Typical use | Models | Hardware |
|---|---|---|---|
| Pilot | One department, 50 users, 10,000 documents | 8B-class LLM in FP8, bge-m3, bge-reranker | One L40S or RTX PRO 6000 in an existing server |
| Department | 500 users, 200,000 documents, Bangla and English | 27B to 32B LLM in FP8, same retrievers, separate extraction GPU for ingestion | One two-GPU server (H100 or L40S class) |
| Enterprise | Thousands of users, millions of pages, several assistants | 70B LLM in FP8 or FP4, reranker and embeddings on their own GPU, guardrail models | One eight-GPU HGX server, Kubernetes, high-availability vector database |
Ingestion is bursty: parsing a million pages with GPU extraction takes days on one GPU and hours on eight, then nothing until the next batch. Plan it as a job, not a service.
What this means for Bangladesh
The organisations we talk to want the same three things: answers from their own documents, in Bangla and English, without the documents leaving their premises. RAG on local GPUs is the only architecture that delivers all three, and the pieces are mature. The work that remains is unglamorous and decisive: scanning, OCR correction, access-control labelling and a test set of real questions. Done in that order, a department-level assistant is a ten-week project, and the same pipeline scales to a ministry.
CompTech builds RAG systems on NVIDIA infrastructure for banks, government and enterprises in Bangladesh, from a single-server pilot on the NVIDIA blueprint to clustered deployments on OpenShift. If you have a document set and a list of questions people keep asking, talk to us; the first step is a two-week evaluation on your own documents, not a hardware quote.
Cover photo of the Wren Library, Trinity College, Cambridge, from Wikimedia Commons, CC BY-SA 2.0. Other photographs are credited in their captions and used under their Creative Commons licences. Diagrams by CompTech. Product names and blueprint contents as published by their vendors in October 2026.