A production-style, fully local Retrieval-Augmented Generation (RAG) system. Upload a PDF and ask questions about it — answers are grounded in the document, retrieved with high-quality embeddings, reranked for relevance, and generated by a local LLM. Runs entirely offline with no API keys.
Unlike a simple linear RAG, this is built as a stateful graph using LangGraph, with a self-correcting step that avoids answering when retrieved context is weak.
- 📂 Chat with any PDF, fully offline
- 🔗 LangGraph stateful pipeline:
retrieve → rerank → grade → generate - 🎯 BGE reranker (cross-encoder) re-scores results for far better relevance than vector search alone
- 🗂️ Qdrant vector database (embedded mode — no Docker required)
- 🏷️ Metadata filtering — restrict answers to a specific page
- 🛑 Self-correction — skips generation and reports "not found" when context is too weak
- 📌 Answers cite the source pages used
- 🪶 Tuned to run on low-RAM hardware
┌──────────┐ ┌────────┐ ┌────────┐ ┌──────────┐
Query → │ retrieve │ → │ rerank │ → │ grade │ → │ generate │ → Answer
│ (Qdrant) │ │ (BGE) │ │ │ │ (Ollama) │
└──────────┘ └────────┘ └───┬────┘ └──────────┘
│ (weak context)
▼
"not found"
| Stage | What it does | Tech |
|---|---|---|
| retrieve | Embeds query, fetches top candidates | BGE embeddings + Qdrant |
| rerank | Re-scores candidates by true relevance | BGE cross-encoder |
| grade | Conditional edge — proceed or bail | LangGraph |
| generate | Answers from reranked context | Ollama (local LLM) |
- LangGraph — stateful RAG orchestration
- Qdrant — vector database (embedded)
- BGE (
bge-small-en-v1.5+bge-reranker-base) — embeddings & reranking - Ollama (
qwen2.5:0.5b) — local LLM inference - Gradio — web UI
- pypdf — PDF parsing
Download Ollama, then:
ollama pull qwen2.5:0.5bgit clone https://github.com/pawanahirwa/Advanced-RAG-LangGraph.git
cd Advanced-RAG-LangGraph
python -m venv rag_env
# Windows:
rag_env\Scripts\activate
# macOS/Linux:
source rag_env/bin/activate
pip install -r requirements.txtpython rag_langgraph.pyOpen the local URL shown in the terminal (usually http://127.0.0.1:7860).
On first run, the BGE embedding and reranker models download automatically (~1.3 GB, one time).
- Upload a PDF and click Process PDF
- (Optional) Enter a page number in Filter by page to restrict retrieval
- Ask questions — answers cite the source pages used
All settings are at the top of rag_langgraph.py:
| Setting | Default | Description |
|---|---|---|
EMBED_MODEL |
bge-small-en-v1.5 |
Embedding model |
RERANK_MODEL |
bge-reranker-base |
Cross-encoder reranker |
LLM_MODEL |
qwen2.5:0.5b |
Ollama model |
FETCH_K |
10 |
Candidates retrieved before reranking |
TOP_K |
3 |
Chunks kept after reranking |
Low on RAM? Swap RERANK_MODEL to cross-encoder/ms-marco-MiniLM-L-6-v2 (~80 MB).
More RAM? Scale up to bge-base, bge-reranker-v2-m3, and qwen2.5:1.5b / llama3.2:3b.
Pawanahirwa