A $0-cost evaluation harness for Retrieval-Augmented Generation (RAG) systems across Hindi and English documents. It ingests PDFs with two chunking strategies, embeds them with two sentence-transformer models, retrieves with FAISS, generates answers with a free LLM (Ollama/Groq/HuggingFace), and scores every run against 20+20 hand-authored gold Q&A pairs using five custom metrics — no RAGAS, no paid APIs, no cloud services required.
flowchart LR
PDFs[Hindi + English PDFs] --> Ingest[ingest.py<br/>fixed + semantic chunking]
Ingest --> Embed[embeddings.py<br/>multilingual + english-only models]
Embed --> Store[(vectorstore.py<br/>4 FAISS indices)]
Queries[40 gold Q&A pairs] --> Eval[evaluator.py]
Store --> Eval
Eval --> Gen[generator.py<br/>Ollama / Groq / HF]
Gen --> Metrics[metrics.py<br/>precision, recall, faithfulness,<br/>correctness, latency]
Metrics --> Compare[compare.py<br/>CSV + summary.json]
Compare --> Dashboard[Streamlit dashboard]
The 4 index configurations evaluated are every combination of:
- Chunking: fixed-size (500 chars, 50 overlap) vs. sentence-based semantic
- Embeddings:
paraphrase-multilingual-MiniLM-L12-v2vs.all-MiniLM-L6-v2(English-only baseline)
Results from a live run are already committed under results/, so you can jump straight to the dashboard:
pip install -r requirements.txt # lean deps: streamlit, pandas, plotly, python-dotenv
streamlit run dashboard/app.pyTo re-run the full pipeline (ingestion, embeddings, retrieval, generation, eval) yourself:
pip install -r requirements-pipeline.txt # adds torch, sentence-transformers, faiss-cpu, etc.
python scripts/download_sample_data.py # regenerates the 5 Hindi + 5 English sample PDFs
python -m src.ingest && python -m src.vectorstore && python -m eval.run_eval && python -m eval.comparerequirements.txt is intentionally lean (dashboard-only) so Streamlit Community Cloud deploys don't install ~1.5GB of unused ML libraries. requirements-pipeline.txt layers the full stack on top for local development.
Or with Docker (defaults to Ollama — pull the model once with docker exec -it <ollama-container> ollama pull llama3.1:8b):
docker-compose upRun python -m eval.compare to regenerate results/comparison.csv and results/summary.json. Results below are from a live run against Groq's llama-3.1-8b-instant free tier (106 of 160 planned query/config runs — the free tier's 6000 tokens/minute cap makes a full 160-run sweep slow; see note below):
| Config | n | Context Precision | Context Recall | Faithfulness | Answer Correctness | Avg Latency (ms) | Composite |
|---|---|---|---|---|---|---|---|
| fixed_multilingual | 27 | 0.5704 | 0.9259 | 0.8144 | 0.7499 | 6218.8 | 0.7652 |
| semantic_multilingual | 26 | 0.5385 | 0.9231 | 0.7784 | 0.7150 | 6139.5 | 0.7387 |
| fixed_english | 27 | 0.4148 | 0.6667 | 0.9789 | 0.7314 | 8237.6 | 0.6979 |
| semantic_english | 26 | 0.4308 | 0.5385 | 0.9890 | 0.6776 | 6822.1 | 0.6590 |
- Best overall:
fixed_multilingual(composite 0.765) - Best for Hindi:
fixed_multilingual(composite 0.740) - Best for English:
semantic_english(composite 0.926)
Screenshots of the Streamlit dashboard: add after your first local run.
Note on run size: Groq's free tier for llama-3.1-8b-instant is capped at 6000 tokens/minute, which — once prompts include 5 retrieved chunks — throttles the sweep to roughly 4-8 calls/minute in practice (well under the advertised 30 req/min, since token budget binds first). eval/run_eval.py is resumable (resume=True by default): rerunning python -m eval.run_eval skips any (query, config) pairs already in results/raw_results.json and continues from where it left off. To finish all 160 runs, either let it run longer unattended, or switch LLM_PROVIDER=ollama in .env for unthrottled local generation.
Chunking strategies. Fixed-size chunking is the simplest baseline and is language-agnostic. Semantic (sentence-boundary) chunking respects sentence structure, which matters more for Hindi, where a hard character cutoff can split a Devanagari conjunct or clause mid-thought. Comparing both surfaces whether semantic boundaries actually improve retrieval quality enough to justify the extra complexity.
Embedding models. paraphrase-multilingual-MiniLM-L12-v2 supports Hindi and English in a shared embedding space, so it's the only real option for the Hindi documents. all-MiniLM-L6-v2 is included as an English-only baseline — it's smaller and faster, and comparing it against the multilingual model on the English test set shows how much (if any) retrieval quality is traded away by supporting multiple languages.
Eval metrics. RAGAS depends on an LLM judge, which typically means a paid API call per metric per query. All five metrics here are computed from local embeddings and token overlap instead: context precision/recall via key-term overlap with the gold answer, faithfulness via embedding similarity between generated claims and retrieved chunks, answer correctness via a blend of embedding similarity and token-F1, and latency as a simple wall-clock measurement. These are proxies, not ground truth — see "What I'd Do Differently" below.
What the failure analysis revealed. Across the 106 completed runs: 9 hallucination cases (faithfulness < 0.5) and 25 retrieval failures (context recall < 0.3) — retrieval, not generation, is the bigger weak point in this setup. The Hindi/English breakdown was the more surprising finding: the multilingual embedding model scores well on Hindi (0.80 answer correctness for fixed_multilingual) but underperforms on English queries in the same index (0.61 for fixed_multilingual vs. 0.82 for fixed_english) — supporting two languages in one embedding space costs real English retrieval quality, not just a hypothetical tradeoff. Faithfulness is consistently higher on English (0.98–0.99) than Hindi (0.78–0.81) regardless of config, suggesting the LLM is more prone to drifting from Hindi source context than English source context, independent of which embedding model retrieved it.
- Larger, human-annotated test sets. 40 queries is enough to compare configurations directionally, not enough to trust a single score. Production evaluation needs hundreds of queries per language, ideally annotated by native speakers, with inter-annotator agreement tracked.
- RAGAS or an LLM-judge pass. Token/embedding overlap metrics are cheap proxies for faithfulness and correctness. An LLM-as-judge (even a local one) catches semantic correctness that surface-level overlap misses — e.g., a paraphrased correct answer that shares no tokens with the gold answer.
- CI/CD for the eval pipeline itself. Every change to prompts, chunking, or embedding models should trigger an automated eval run with regression alerts, not a manual
eval/compare.pyinvocation. - Production deployment on Kubernetes with the vector store as a persistent service, embedding model warm pools, and monitoring on retrieval latency, cache hit rate, and generation cost/token usage.
- Support for more Indian languages (Tamil, Bengali, Marathi, Telugu, etc.), which would require validating that the multilingual embedding model's coverage and chunking rules (sentence boundary characters differ by script) generalize beyond Hindi and English.
multilingual-rag-eval/
├── data/ # Hindi/English PDFs + gold Q&A pairs
├── src/ # ingest, embeddings, vectorstore, generator, evaluator, config
├── eval/ # metrics, run_eval, compare
├── dashboard/ # Streamlit app (Overview / Query Explorer / Failure Analysis)
├── scripts/ # sample data generation + query validation
└── results/ # generated indices, raw results, comparison CSV/JSON