Skip to content
View achellesheel's full-sized avatar

Block or report achellesheel

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
achellesheel/README.md

Typing SVG

"Some are born great, some hack their way into greatness — one retrieval eval at a time."

🐍 The Chamber of Secrets

Something is slithering through the walls of the contribution graph...

The Basilisk devours the contribution graph

The Basilisk only answers to its master — and it feeds on every commit.

divider

🏆 The Triwizard Tournament — Agent Loops & RAG Systems

Every champion carries a scar. Here's what each one actually broke, and what I did about it.

Trial Project Pattern The Challenge The Scar — What Broke → What I Did
🐉 The Dragon research-agent Agent loop (LangGraph) Autonomous tool-use loop over web search + arXiv, persistent memory (SQLite + FAISS), structured research briefs. The LLM routes tool calls itself — no hardcoded pipeline. ⚠️ No scar logged yet. Single-commit build, no incident write-up. This is the one champion that hasn't actually been to battle — next thing I fix before I call this story done.
🌊 The Black Lake multilingual-rag-eval RAG eval Hindi + English retrieval evaluation across 4 FAISS index configs, 40 gold Q&A pairs, $0-cost proxy metrics for faithfulness/correctness/recall. The multilingual embedding model I picked to support Hindi quietly cost 21 points of English answer-correctness in the same index (0.61 vs. 0.82) — invisible until I broke results out by language. Full breakdown →
🌀 The Maze llm-red-team-eval Red-team eval Automated hallucination, injection, consistency, and refusal testing — generates a model report card. First live run of the new rag_injection suite got a real model to leak a fake admin backdoor code from a single poisoned support doc, on a completely benign customer question. Verified with two independent retrievers so it wasn't a fluke. Transcript →
📖 The Restricted Section knowledge-mcp-server RAG + MCP Local MCP server turning PDFs/notes into a semantic knowledge base — hybrid BM25/dense search, cross-encoder reranking. The Ghost Chunk — a retrieval bug that silently corrupted every search snippet returned to the model, hunted down and fixed. Full postmortem →
⚔️ The Champion's Duel search-ranking-service Retrieval + re-ranking Two-stage production search: Elasticsearch BM25 retrieval → XGBoost learning-to-rank re-ranking, Dockerized, MLflow-tracked. ⚠️ No scar logged yet. Runs and ranks correctly, but no documented incident — the second gap in the lineup.

divider

🎓 Hogwarts Coursework — GUVI-HCL Data Science (Batch DS-C-WE-E-B148)

Project What it does The Scar
omnifeedback-ai Capstone: enterprise feedback-triage platform — SQL star schema, TF-IDF/K-Means clustering, PyTorch BiLSTM + transfer-learned DistilBERT urgency scoring, BERT NER/BART summarization, GenAI copilot. Two real, publicly-logged bugs across 3 model versions: an overfitting dataset (V1→V2), then a generalization failure a live user found — the model scored obviously-critical and obviously-positive feedback almost identically — fixed with transfer learning (V2→V3). Full build log in PROGRESS.md.
book-recommendation-system Content-based + collaborative filtering book recommender with Streamlit UI —
garbage-classification-deep-learning Deep learning image classifier for waste sorting — CNN with transfer learning —
real-estate-investment-advisor Real estate investment analysis tool — ROI calculation, market comparison, risk assessment —
brand-visibility-project Brand visibility analysis using NLP — sentiment analysis, media monitoring, competitive intelligence —
brickview2.0 Property analytics dashboard — SQLite backend, data visualization —

✍️ The Daily Prophet — Articles

Piece Link
Women's Day 2026 achellesheel/womens_day_2026

divider

🪄 The Wand Chooses the Engineer

Core Wood Loyalty House Patronus

I don't build demos. I build agent loops and RAG pipelines that get evaluated, get broken, and get fixed — in public.

divider

📜 Spells & Charms Cast Daily

  • 🪄 Accio Data — retrieval that pulls exactly the right context, nothing extraneous (hybrid BM25 + dense + cross-encoder reranking)
  • 🧭 Point Me — agent tool-routing: letting an LLM decide when to search, fetch, summarize, or stop, instead of a hardcoded pipeline
  • 🛡️ Protego — input validation, guardrails, prompt-injection & indirect RAG-injection defense
  • 🧠 Legilimens — LLM & RAG evaluation — hallucination rate, faithfulness, retrieval recall, not vibes
  • 🕊️ Expecto Patronum — graceful fallback under real-world load
  • 🧹 Wingardium Leviosa — CI/CD, shipped past localhost
  • 🌀 Confundo — adversarial red-teaming and RAG-poisoning, breaking it before someone else does
  • 😄 Riddikulus — five-minute incident fix, not a five-hour fire drill
  • 🔓 Alohomora — reverse-engineering APIs with no manual

divider

🧹 The Broom — Built for Speed

Docker K8s MLflow FAISS MCP

Firebolt-class, not Nimbus 2000 nostalgia — if it's not fast enough to keep up, it doesn't make the team.

♟️ Wizard's Chess — Current Strategy

Endgame: agentic AI systems that are measured, monitored, and trusted in production. No checkers with side projects — pawns (retrieval eval pipelines) advancing, knights (agent tool-use loops) deep in enemy territory, queen (full production-grade agent stack, with every failure documented, not hidden) coming out soon.

divider

🔮 The Marauder's Map

Followers

🪄 Grimoire (Tech Stack)

Python LangChain MCP FAISS Docker Kubernetes FastAPI Streamlit MLflow SQLite

divider

🦉 Send an Owl

LinkedIn Email

"It is our choices, Harry, that show what we truly are, far more than our abilities." — and my choice is to ship, break it in public, and fix it.

Footer

Popular repositories Loading

  1. deeplearningiai101 deeplearningiai101 Public archive

    short courses on openai api

    Python

  2. achellesheel achellesheel Public

    Config files for my GitHub profile.

    Python

  3. chatpc chatpc Public archive

    JavaScript

  4. kaggle- kaggle- Public

    kaggle competitions

  5. my-ifs-tracker my-ifs-tracker Public archive

    HTML

  6. notemaking notemaking Public archive

    HTML