MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.
-
Updated
Mar 27, 2026 - Python
MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.
Self-hosted search + markdown harvester for AI agents. SearXNG (100+ engines) + FastAPI + trafilatura. Tavily-compatible /search plus /extract with size presets and pagination. One-command Docker Compose.
pebkac Chrome Nonautomation - A Local LLM-Driven Web Co-Browser using Smolagents, Zendriver, Trafilatura.
Fast, accurate web content extraction in Rust. ML page-type classification, per-type extraction, confidence scoring. F1=0.966 on ScrapingHub (#1), F1=0.859 across 2,008 annotated pages (1,497 development + 511 held-out test
Tools for LLMs to anonymously search and browse the web
Fast and accurate web content extraction
web Scrapper In Python
Keyless, uv-native web search + read for AI agents: ddgs multi-engine search with de-correlated rank fusion, Trafilatura extraction to paginated Markdown, keyless arxiv and github search. Self-hosted SearXNG without Docker, an optional Tor layer for .onion search and fetch, SOCKS5 egress, one-call init, and a doctor.
Convert any URL into clean, token-efficient Markdown for LLMs. CLI, Python API, MCP server.
ChatGPT AI Clone
Telegram Mini App that saves internet articles to read them later
Selective web content extraction for AI agents — URL + query returns only the chunks that matter (Python library + MCP server)
Fast, permissive, self-hosted web extraction service — URL → clean LLM-ready markdown. Async HTTP/2, conditional Playwright rendering, multi-strategy extraction, PDF/academic pipeline. No LLM required, no copyleft deps. Apache-2.0.
A pipe-based news article scraping and metadata extraction library for Python
Real-time AI search and chat backend with WebSocket streaming, powered by Tavily web search and Google Gemini for Flutter apps.
Safe, config-driven Python web ingestion pipeline with extraction, evidence-gated AI generation, provenance ledger, RAG chunks, data cards, and multi-provider exports.
Local-first AI trend intelligence dashboard with GitHub Radar, Research Radar, Startup Gap Finder, Scrapy/Trafilatura crawler, SQLite, Streamlit, and Ollama support
Protocole de collecte et d'analyse d'archives de la Wayback Machine pour une analyse textuelle et statistique
HTML main-content extraction for Rust — ports of Mozilla Readability, Trafilatura, and htmldate.
FastAPI service that classifies publisher websites for affiliate campaigns using an LLM pipeline (scrape → signals → RAG → scoring). Detects cashback, adult, gambling, scams. Supports OpenAI/Ollama, Redis cache, Docker.
Add a description, image, and links to the trafilatura topic page so that developers can more easily learn about it.
To associate your repository with the trafilatura topic, visit your repo's landing page and select "manage topics."