Remove boilerplate from scraped markdown before it reaches an LLM.
Scraping and extract APIs like Tavily hand you "LLM-ready" markdown that still carries navigation, footers, cookie banners, promos, and link farms — 10–90% of the tokens depending on the page. The same goes for your own scraper or loader pipeline: if it produces markdown (or HTML converted to markdown), Winnow slots in right after it — any markdown in, leaner markdown out, with a receipt for every removed block.
- Subtractive-only. Winnow deletes blocks; it never rewrites a word. Zero hallucination risk by construction.
- Recall-first. Dropping real content is data loss; keeping boilerplate just costs tokens. When uncertain, Winnow keeps.
- Auditable. Every removed block comes back with a reason code and score.
- Template memory. Feed Winnow multiple pages from one site and it learns the site's template — blocks repeated across pages are boilerplate, near-certainly.
- Zero runtime dependencies. The core install adds nothing to your tree.
pip install winnow-mdOptional extras:
pip install "winnow-md[model]" # learned block-sequence scorer (numpy + model2vec)
pip install "winnow-md[tokens]" # exact token counts via tiktoken
pip install "winnow-md[model,tokens]"Python 3.9+. The package installs as winnow-md and imports as winnow. The
core is pure Python with no dependencies; if the [model] extra isn't
installed, the learned scorer is silently skipped and the heuristics run alone.
import winnow
# One page
res = winnow.clean(markdown_text)
print(res.markdown) # cleaned markdown
print(res.stats) # tokens before/after, reduction %
for r in res.removed: # the receipt
print(r.reasons, r.text[:60])
# A crawl — template memory kicks in across pages of the same domain
w = winnow.Winnow(aggressiveness=0.5)
results = w.clean_many(pages) # list of markdown strings (or (md, url) tuples)
# Streaming with a persistent per-domain template store
w = winnow.Winnow(store="winnow.db")
res = w.clean(md, url="https://example.com/post/1")winnow clean page.md # cleaned markdown to stdout
winnow clean ./crawl/ --report out.html # batch + filterable HTML audit report
winnow clean ./crawl/ --url-mode strip # also strip URL bodies from kept linksJina Reader output (Title: / URL Source: preamble) is auto-detected and the
source URL is used for template memory.
Seven generations of independently-labeled, adversarially-arbitrated exam batches (each fetched fresh, dual-labeled blind, disputes refereed) — ~19,200 hand-adjudicated blocks across 32 domains:
| Exam batch | Content recall | Junk recall | Token cut |
|---|---|---|---|
| batch 7 (newest, still converging) | 0.969 | 0.55 | −30% |
| batch 6 | 0.970 | 0.55 | −27% |
| batch 5 | 0.963 | 0.58 | −41% |
| batch 4 | 0.979 | 0.61 | −49% |
| batch 3 | 0.992 | 0.57 | −35% |
| batch 2 | 0.997 | 0.58 | −30% |
| batch 1 | 1.000 | 0.69 | −42% |
Content recall is the fraction of real content kept — the number that must never slip. Junk recall is the fraction of boilerplate actually removed; what it misses costs tokens, never correctness. Each batch was a fresh exam nothing had been tuned on when first scored, then became training data — the newest batch is always the honest one.
Block-level recall understates how much information survives, because the blocks Winnow wrongly removes are overwhelmingly short. Measured across all ~19,200 labelled blocks:
- 98.4% of content blocks kept — but 99.3% of content words.
- Of the wrongly-removed blocks, 46% have 20 characters of visible text or less, and 19% are images with no text at all. Only 13% run to 25+ words.
- The substantive misses concentrate in one shape: link-dense lists that are genuinely content (cast/credit lists, legal cross-references, curated example indexes) — the hard case of telling a nav menu from a content list. Some of the rest are duplicate blocks whose text survives elsewhere.
Every removal is in res.removed with a reason code, and res.integrity()
reports exactly which tables, links and words disappeared — so you can audit
your own corpus rather than trusting these numbers.
aggressiveness (0.0–1.0, default 0.5) trades junk removal against risk.
Measured on the full benchmark:
| aggressiveness | word recall | block recall | junk recall | token cut |
|---|---|---|---|---|
| 0.00 (safest) | 99.75% | 99.1% | 0.47 | −30% |
| 0.25 | 99.29% | 98.7% | 0.56 | −36% |
| 0.50 (default) | 99.28% | 98.4% | 0.57 | −36% |
| 0.75 | 99.26% | 98.3% | 0.66 | −37% |
| 1.00 | 96.28% | 96.0% | 0.67 | −40% |
For high-stakes corpora — legal, financial, medical — use
aggressiveness=0.0: it still cuts ~30% of tokens while keeping 99.75% of
content words. Avoid 1.00 unless you have verified it on your own pages.
Add --url-mode strip for roughly 15 additional points of token cut with zero
text loss. With the [model] extra installed, a learned block-sequence scorer
raises junk recall further; it is capped so that it can never delete a block
on its own.
The benchmark harness, labeling pipeline, and mutation self-test live in
bench/ — see ARCHITECTURE.md.
Contributions are welcome — bug fixes, new signals, docs, tests, or ideas. One especially easy and useful report: a page Winnow handles badly, since most improvements so far came from being shown a real page it got wrong. See CONTRIBUTING.md for setup, the three design rules any change must respect, and how to run the benchmark gates before opening a PR.
MIT — see LICENSE.