Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Winnow

PyPI Python License: MIT

Remove boilerplate from scraped markdown before it reaches an LLM.

Scraping and extract APIs like Tavily hand you "LLM-ready" markdown that still carries navigation, footers, cookie banners, promos, and link farms — 10–90% of the tokens depending on the page. The same goes for your own scraper or loader pipeline: if it produces markdown (or HTML converted to markdown), Winnow slots in right after it — any markdown in, leaner markdown out, with a receipt for every removed block.

  • Subtractive-only. Winnow deletes blocks; it never rewrites a word. Zero hallucination risk by construction.
  • Recall-first. Dropping real content is data loss; keeping boilerplate just costs tokens. When uncertain, Winnow keeps.
  • Auditable. Every removed block comes back with a reason code and score.
  • Template memory. Feed Winnow multiple pages from one site and it learns the site's template — blocks repeated across pages are boilerplate, near-certainly.
  • Zero runtime dependencies. The core install adds nothing to your tree.

Install

pip install winnow-md

Optional extras:

pip install "winnow-md[model]"    # learned block-sequence scorer (numpy + model2vec)
pip install "winnow-md[tokens]"   # exact token counts via tiktoken
pip install "winnow-md[model,tokens]"

Python 3.9+. The package installs as winnow-md and imports as winnow. The core is pure Python with no dependencies; if the [model] extra isn't installed, the learned scorer is silently skipped and the heuristics run alone.

Quickstart

import winnow

# One page
res = winnow.clean(markdown_text)
print(res.markdown)          # cleaned markdown
print(res.stats)             # tokens before/after, reduction %
for r in res.removed:        # the receipt
    print(r.reasons, r.text[:60])

# A crawl — template memory kicks in across pages of the same domain
w = winnow.Winnow(aggressiveness=0.5)
results = w.clean_many(pages)          # list of markdown strings (or (md, url) tuples)

# Streaming with a persistent per-domain template store
w = winnow.Winnow(store="winnow.db")
res = w.clean(md, url="https://example.com/post/1")
winnow clean page.md                      # cleaned markdown to stdout
winnow clean ./crawl/ --report out.html   # batch + filterable HTML audit report
winnow clean ./crawl/ --url-mode strip    # also strip URL bodies from kept links

Jina Reader output (Title: / URL Source: preamble) is auto-detected and the source URL is used for template memory.

Benchmark

Seven generations of independently-labeled, adversarially-arbitrated exam batches (each fetched fresh, dual-labeled blind, disputes refereed) — ~19,200 hand-adjudicated blocks across 32 domains:

Exam batch Content recall Junk recall Token cut
batch 7 (newest, still converging) 0.969 0.55 −30%
batch 6 0.970 0.55 −27%
batch 5 0.963 0.58 −41%
batch 4 0.979 0.61 −49%
batch 3 0.992 0.57 −35%
batch 2 0.997 0.58 −30%
batch 1 1.000 0.69 −42%

Content recall is the fraction of real content kept — the number that must never slip. Junk recall is the fraction of boilerplate actually removed; what it misses costs tokens, never correctness. Each batch was a fresh exam nothing had been tuned on when first scored, then became training data — the newest batch is always the honest one.

What actually gets lost

Block-level recall understates how much information survives, because the blocks Winnow wrongly removes are overwhelmingly short. Measured across all ~19,200 labelled blocks:

  • 98.4% of content blocks kept — but 99.3% of content words.
  • Of the wrongly-removed blocks, 46% have 20 characters of visible text or less, and 19% are images with no text at all. Only 13% run to 25+ words.
  • The substantive misses concentrate in one shape: link-dense lists that are genuinely content (cast/credit lists, legal cross-references, curated example indexes) — the hard case of telling a nav menu from a content list. Some of the rest are duplicate blocks whose text survives elsewhere.

Every removal is in res.removed with a reason code, and res.integrity() reports exactly which tables, links and words disappeared — so you can audit your own corpus rather than trusting these numbers.

Choosing an aggressiveness

aggressiveness (0.0–1.0, default 0.5) trades junk removal against risk. Measured on the full benchmark:

aggressiveness word recall block recall junk recall token cut
0.00 (safest) 99.75% 99.1% 0.47 −30%
0.25 99.29% 98.7% 0.56 −36%
0.50 (default) 99.28% 98.4% 0.57 −36%
0.75 99.26% 98.3% 0.66 −37%
1.00 96.28% 96.0% 0.67 −40%

For high-stakes corpora — legal, financial, medical — use aggressiveness=0.0: it still cuts ~30% of tokens while keeping 99.75% of content words. Avoid 1.00 unless you have verified it on your own pages.

Add --url-mode strip for roughly 15 additional points of token cut with zero text loss. With the [model] extra installed, a learned block-sequence scorer raises junk recall further; it is capped so that it can never delete a block on its own.

The benchmark harness, labeling pipeline, and mutation self-test live in bench/ — see ARCHITECTURE.md.

Contributing

Contributions are welcome — bug fixes, new signals, docs, tests, or ideas. One especially easy and useful report: a page Winnow handles badly, since most improvements so far came from being shown a real page it got wrong. See CONTRIBUTING.md for setup, the three design rules any change must respect, and how to run the benchmark gates before opening a PR.

License

MIT — see LICENSE.

About

Remove boilerplate from scraped markdown before it reaches an LLM

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages