"Some are born great, some hack their way into greatness — one retrieval eval at a time."
Something is slithering through the walls of the contribution graph...
The Basilisk only answers to its master — and it feeds on every commit.
Every champion carries a scar. Here's what each one actually broke, and what I did about it.
| Trial | Project | Pattern | The Challenge | The Scar — What Broke → What I Did |
|---|---|---|---|---|
| 🐉 The Dragon | research-agent | Agent loop (LangGraph) | Autonomous tool-use loop over web search + arXiv, persistent memory (SQLite + FAISS), structured research briefs. The LLM routes tool calls itself — no hardcoded pipeline. | |
| 🌊 The Black Lake | multilingual-rag-eval | RAG eval | Hindi + English retrieval evaluation across 4 FAISS index configs, 40 gold Q&A pairs, $0-cost proxy metrics for faithfulness/correctness/recall. | The multilingual embedding model I picked to support Hindi quietly cost 21 points of English answer-correctness in the same index (0.61 vs. 0.82) — invisible until I broke results out by language. Full breakdown → |
| 🌀 The Maze | llm-red-team-eval | Red-team eval | Automated hallucination, injection, consistency, and refusal testing — generates a model report card. | First live run of the new rag_injection suite got a real model to leak a fake admin backdoor code from a single poisoned support doc, on a completely benign customer question. Verified with two independent retrievers so it wasn't a fluke. Transcript → |
| 📖 The Restricted Section | knowledge-mcp-server | RAG + MCP | Local MCP server turning PDFs/notes into a semantic knowledge base — hybrid BM25/dense search, cross-encoder reranking. | The Ghost Chunk — a retrieval bug that silently corrupted every search snippet returned to the model, hunted down and fixed. Full postmortem → |
| ⚔️ The Champion's Duel | search-ranking-service | Retrieval + re-ranking | Two-stage production search: Elasticsearch BM25 retrieval → XGBoost learning-to-rank re-ranking, Dockerized, MLflow-tracked. |
| Project | What it does | The Scar |
|---|---|---|
| omnifeedback-ai | Capstone: enterprise feedback-triage platform — SQL star schema, TF-IDF/K-Means clustering, PyTorch BiLSTM + transfer-learned DistilBERT urgency scoring, BERT NER/BART summarization, GenAI copilot. | Two real, publicly-logged bugs across 3 model versions: an overfitting dataset (V1→V2), then a generalization failure a live user found — the model scored obviously-critical and obviously-positive feedback almost identically — fixed with transfer learning (V2→V3). Full build log in PROGRESS.md. |
| book-recommendation-system | Content-based + collaborative filtering book recommender with Streamlit UI | — |
| garbage-classification-deep-learning | Deep learning image classifier for waste sorting — CNN with transfer learning | — |
| real-estate-investment-advisor | Real estate investment analysis tool — ROI calculation, market comparison, risk assessment | — |
| brand-visibility-project | Brand visibility analysis using NLP — sentiment analysis, media monitoring, competitive intelligence | — |
| brickview2.0 | Property analytics dashboard — SQLite backend, data visualization | — |
| Piece | Link |
|---|---|
| Women's Day 2026 | achellesheel/womens_day_2026 |
I don't build demos. I build agent loops and RAG pipelines that get evaluated, get broken, and get fixed — in public.
- 🪄 Accio Data — retrieval that pulls exactly the right context, nothing extraneous (hybrid BM25 + dense + cross-encoder reranking)
- 🧭 Point Me — agent tool-routing: letting an LLM decide when to search, fetch, summarize, or stop, instead of a hardcoded pipeline
- 🛡️ Protego — input validation, guardrails, prompt-injection & indirect RAG-injection defense
- 🧠 Legilimens — LLM & RAG evaluation — hallucination rate, faithfulness, retrieval recall, not vibes
- 🕊️ Expecto Patronum — graceful fallback under real-world load
- 🧹 Wingardium Leviosa — CI/CD, shipped past
localhost - 🌀 Confundo — adversarial red-teaming and RAG-poisoning, breaking it before someone else does
- 😄 Riddikulus — five-minute incident fix, not a five-hour fire drill
- 🔓 Alohomora — reverse-engineering APIs with no manual
Firebolt-class, not Nimbus 2000 nostalgia — if it's not fast enough to keep up, it doesn't make the team.
Endgame: agentic AI systems that are measured, monitored, and trusted in production. No checkers with side projects — pawns (retrieval eval pipelines) advancing, knights (agent tool-use loops) deep in enemy territory, queen (full production-grade agent stack, with every failure documented, not hidden) coming out soon.

