Pareto-optimal models for cleaning the web — fast, encoder-based main-content extraction from HTML.
-
Updated
Jul 1, 2026 - HTML
Pareto-optimal models for cleaning the web — fast, encoder-based main-content extraction from HTML.
Benson turns a list of URLs into mp3s of the contents of each web page - take control over your reading backlog!
WCXB: Web Content Extraction Benchmark — 2,008 pages, 7 page types, 1,613 domains. The largest open benchmark for web content extraction, boilerplate removal, and main content detection.
Heuristic text extraction from news sites in Python3
HTML main-content extraction for Rust — ports of Mozilla Readability, Trafilatura, and htmldate.
Remove boilerplate from scraped markdown before it reaches an LLM
Swift port of the jusText boilerplate removal library — extracts main article content from HTML
Elixir NIF bindings for rs-trafilatura — main content and metadata extraction for web pages, with no Python at deploy time.
Official Node.js client SDK for AI Web-to-Markdown Extract API. Convert URL to Clean JSON & Markdown for LLMs and RAG pipelines.
To associate your repository with the boilerplate-removal topic, visit your repo's landing page and select "manage topics."