The open-source operating system behind a content site that reached 11.08M impressions and 44.8K clicks in 90 days. That's 7.7× the impressions and 12× the clicks of the previous 90 days. Every number here is verifiable in RESULTS.md.
This isn't a list of tips. It's the methodology, the scripts, and the guardrails we actually run: tested, measured against controls, and including the mistakes that cost us traffic so you don't repeat them.
On our blog, human-shaped search queries convert at 0.95%. Machine-shaped ones (AI research agents, scrapers, SEO tools) convert at 0.01%, a 95× gap, and they make up almost half of the page-one impressions. So:
- Your blended CTR is no longer a title metric. Retitling a page whose impressions come from an AI agent does nothing.
human-query-splitshows you which pages are really underperforming for humans. - Those agent impressions still matter. They are AI engines reading you. Winning GEO/AEO (being the source ChatGPT, Perplexity, Claude and Google AI Overviews cite) takes extractable, dated, verified facts. Chapter 1 covers how.
- Most "wins" aren't. Measured against untouched pages, our title rewrites outgrew the control by roughly 19–48 points (depending on the window). Our snippet rewrites did nothing, and our "freshness refreshes" only looked like losses because they picked fading news. Chapter 2 shows how to tell the difference.
Founders, marketers and engineers running content, programmatic or news sites, especially with AI in the content pipeline, who want rankings and AI citations without gambling on unmeasured tactics.
Before you start: the Search Console scripts need a Google Cloud project with the Search Console API enabled and access to your property. Follow docs/setup-gsc.md first (about 10 minutes).
git clone https://github.com/TraceCohenTech/ai-seo-playbook && cd ai-seo-playbook
npm ci
# Authenticate with Search Console. Needs a Google Cloud project with the
# Search Console API enabled; see docs/setup-gsc.md for the step-by-step setup.
gcloud auth application-default login --scopes=https://www.googleapis.com/auth/webmasters.readonly,https://www.googleapis.com/auth/cloud-platform# Is your low CTR a title problem or an AI-agent problem?
node scripts/human-query-split.mjs --site sc-domain:yoursite.com
# Daily drop alarm for cron/CI (exit code 2 = alert)
node scripts/traffic-guard.mjs --site sc-domain:yoursite.com --webhook https://ntfy.sh/your-topic
# Are the URLs in your sitemaps actually indexable?
node scripts/sitemap-health.mjs --sitemap https://yoursite.com/sitemap.xmlEvery script prints its full documentation with --help.
Some scripts read your site's source files instead of (or as well as) Search Console. --dir is the folder
in your website's own code repo that holds the pages, not a cache and not a special export format:
| Your site | Point --dir at |
|---|---|
| Next.js App Router | ./src/app or ./app (blog/foo/page.tsx is read as /blog/foo) |
| Next.js Pages Router | ./pages |
| Markdown / MDX (Astro, Hugo, Jekyll, Eleventy, Gatsby…) | your content folder, e.g. ./content or ./src/content/blog |
| Static HTML | the built output folder, e.g. ./public or ./dist (most --dir scripts read .html) |
| WordPress, Webflow, Squarespace (no source files) | skip the --dir scripts and use the Search Console and URL-based ones |
Run the scripts from this repo and pass the path to your site's repo, e.g.
node scripts/thin-content-detector.mjs --dir ../my-site/src/app. Use --ext to change which file types are read. For
Markdown, a file's URL is its path inside --dir, so if content/blog/foo.md is served at /blog/foo, add --url-prefix /blog
where the script supports it.
| Needs | Scripts |
|---|---|
Search Console only (--site) |
human-query-split, weekly-report, striking-distance, ctr-audit, gsc-rewrite-candidates, cannibalization-detector, traffic-guard, matched-control-readout, rewrite-measurer, query-gap-miner (--dir optional) |
| A live URL or sitemap | sitemap-health, schema-validator (--url/--sitemap), redirect-checker, indexing-submitter |
Your source files (--dir) |
thin-content-detector, template-detector, meta-length-checker, factual-density-scorer, orphan-finder, broken-link-checker |
| Source files + Search Console | content-audit, refresh-tracker |
Setup details are in docs/setup-gsc.md.
| # | Chapter | You'll learn |
|---|---|---|
| 1 | GEO and AEO | Being cited by answer engines: extractable facts, provenance, llms.txt, MCP |
| 2 | Measurement | Matched controls, holdouts, and the traps that fake your wins |
| 3 | Incidents and guardrails | The day we noindexed our own site, and the guards that now catch it |
| 4 | Entity pages | One page per entity, indexing on substance, honest structured data |
| 5 | The content refresh system | Change fewer pages, and only the right ones |
| 6 | Automation and cost | Running cron + AI routines + CI safely and cheaply |
Plus a prompt library with sourcing rules built in, and a bot-traffic guide.
| Script | What it tells you |
|---|---|
human-query-split |
Human vs AI-agent impressions per page; the real retitle list |
weekly-report |
Week-over-week movers, trending queries, CTR triage |
striking-distance |
Pages at position 5–20 where a push pays off most |
ctr-audit |
The click gap against a CTR curve fitted to your site |
gsc-rewrite-candidates |
Title rewrite candidates (optionally human queries only) |
cannibalization-detector |
Queries where your own pages compete; flagged for review |
query-gap-miner |
Queries you rank for without a page that targets them |
refresh-tracker |
Aging pages that are losing traffic |
| Script | What it tells you |
|---|---|
matched-control-readout |
Did the change work, compared with pages you didn't touch? |
rewrite-measurer |
One change date vs a control band of similar-traffic pages; skips the change week and waits for complete data |
| Script | What it catches |
|---|---|
traffic-guard |
Daily traffic cliffs and slow recoveries (3 baselines) |
sitemap-health |
Sitemap URLs that redirect, 404, are noindexed or canonicalize elsewhere |
redirect-checker |
Redirect chains and sitemap URLs that redirect |
broken-link-checker |
Broken internal and external links |
schema-validator |
Invalid or incomplete JSON-LD (including Next.js and @graph) |
meta-length-checker |
Titles and descriptions that will be truncated |
| Script | What it does |
|---|---|
evidence-verifier |
Checks every fact's quote is really on its source page before you publish |
ai-citation-tracker |
Samples which domains AI engines cite for your queries (official APIs only) |
factual-density-scorer |
How much specific, extractable fact a page carries |
rewrite-detector |
News stories re-reporting the same event under a new URL |
| Script | What it finds |
|---|---|
content-audit |
Every page scored by traffic and quality, with review labels |
thin-content-detector |
Thin pages, with age and traffic guards before any noindex suggestion |
template-detector |
AI template phrases (lists in config/anti-ai-rules.json) |
orphan-finder |
Pages nothing links to |
| Script | Use it for |
|---|---|
websub-ping |
Notifying a WebSub hub when an RSS/Atom feed that declares it changes |
indexing-submitter |
Job posting and livestream pages only, as Google's policy requires |
To run the weekly report on a schedule, copy examples/workflows/weekly-seo-report.yml into your site's repo.
Configs (quality gates, title rules, refresh rules, schema rules…) live in config/, and JSON-LD templates in
schemas/. Sample outputs for the scripts are in samples/, generated from test fixtures.
- Measure against a control, or it didn't happen.
- Separate humans from machines before judging CTR.
- Every fact needs a source you can re-check. AI answer engines reward being right.
- Structured data mirrors visible content. No synthetic dates or invisible FAQs.
- Guardrails before automation. If a machine can publish it, a machine should also be able to catch it.
- No tricks that need hiding. Nothing here disguises automation or games a policy.
Full numbers, windows and methods are in RESULTS.md. The site behind them is valueaddvc.com, a startup and venture-capital intelligence platform with ~1,100 articles, a daily news feed and thousands of programmatic company and investor pages.
The playbook and scripts are and stay free (MIT). See ROADMAP.md for what's next, including hosted audits. Contributions with real numbers and windows are very welcome; see CONTRIBUTING.md.
Built by Trace Cohen at Value Add VC. If this helped you, a ⭐ helps others find it.
MIT, see LICENSE.