Skip to content

Repository files navigation

The AI SEO Playbook: SEO + GEO + AEO

CI License: MIT Node 20+

The open-source operating system behind a content site that reached 11.08M impressions and 44.8K clicks in 90 days. That's 7.7× the impressions and 12× the clicks of the previous 90 days. Every number here is verifiable in RESULTS.md.

This isn't a list of tips. It's the methodology, the scripts, and the guardrails we actually run: tested, measured against controls, and including the mistakes that cost us traffic so you don't repeat them.

The 2026 reality most SEO advice misses

On our blog, human-shaped search queries convert at 0.95%. Machine-shaped ones (AI research agents, scrapers, SEO tools) convert at 0.01%, a 95× gap, and they make up almost half of the page-one impressions. So:

  • Your blended CTR is no longer a title metric. Retitling a page whose impressions come from an AI agent does nothing. human-query-split shows you which pages are really underperforming for humans.
  • Those agent impressions still matter. They are AI engines reading you. Winning GEO/AEO (being the source ChatGPT, Perplexity, Claude and Google AI Overviews cite) takes extractable, dated, verified facts. Chapter 1 covers how.
  • Most "wins" aren't. Measured against untouched pages, our title rewrites outgrew the control by roughly 19–48 points (depending on the window). Our snippet rewrites did nothing, and our "freshness refreshes" only looked like losses because they picked fading news. Chapter 2 shows how to tell the difference.

Who this is for

Founders, marketers and engineers running content, programmatic or news sites, especially with AI in the content pipeline, who want rankings and AI citations without gambling on unmeasured tactics.

Quick start

Before you start: the Search Console scripts need a Google Cloud project with the Search Console API enabled and access to your property. Follow docs/setup-gsc.md first (about 10 minutes).

git clone https://github.com/TraceCohenTech/ai-seo-playbook && cd ai-seo-playbook
npm ci

# Authenticate with Search Console. Needs a Google Cloud project with the
# Search Console API enabled; see docs/setup-gsc.md for the step-by-step setup.
gcloud auth application-default login --scopes=https://www.googleapis.com/auth/webmasters.readonly,https://www.googleapis.com/auth/cloud-platform
# Is your low CTR a title problem or an AI-agent problem?
node scripts/human-query-split.mjs --site sc-domain:yoursite.com

# Daily drop alarm for cron/CI (exit code 2 = alert)
node scripts/traffic-guard.mjs --site sc-domain:yoursite.com --webhook https://ntfy.sh/your-topic

# Are the URLs in your sitemaps actually indexable?
node scripts/sitemap-health.mjs --sitemap https://yoursite.com/sitemap.xml

Every script prints its full documentation with --help.

What does --dir point at?

Some scripts read your site's source files instead of (or as well as) Search Console. --dir is the folder in your website's own code repo that holds the pages, not a cache and not a special export format:

Your site Point --dir at
Next.js App Router ./src/app or ./app (blog/foo/page.tsx is read as /blog/foo)
Next.js Pages Router ./pages
Markdown / MDX (Astro, Hugo, Jekyll, Eleventy, Gatsby…) your content folder, e.g. ./content or ./src/content/blog
Static HTML the built output folder, e.g. ./public or ./dist (most --dir scripts read .html)
WordPress, Webflow, Squarespace (no source files) skip the --dir scripts and use the Search Console and URL-based ones

Run the scripts from this repo and pass the path to your site's repo, e.g. node scripts/thin-content-detector.mjs --dir ../my-site/src/app. Use --ext to change which file types are read. For Markdown, a file's URL is its path inside --dir, so if content/blog/foo.md is served at /blog/foo, add --url-prefix /blog where the script supports it.

Needs Scripts
Search Console only (--site) human-query-split, weekly-report, striking-distance, ctr-audit, gsc-rewrite-candidates, cannibalization-detector, traffic-guard, matched-control-readout, rewrite-measurer, query-gap-miner (--dir optional)
A live URL or sitemap sitemap-health, schema-validator (--url/--sitemap), redirect-checker, indexing-submitter
Your source files (--dir) thin-content-detector, template-detector, meta-length-checker, factual-density-scorer, orphan-finder, broken-link-checker
Source files + Search Console content-audit, refresh-tracker

Setup details are in docs/setup-gsc.md.

The playbook

# Chapter You'll learn
1 GEO and AEO Being cited by answer engines: extractable facts, provenance, llms.txt, MCP
2 Measurement Matched controls, holdouts, and the traps that fake your wins
3 Incidents and guardrails The day we noindexed our own site, and the guards that now catch it
4 Entity pages One page per entity, indexing on substance, honest structured data
5 The content refresh system Change fewer pages, and only the right ones
6 Automation and cost Running cron + AI routines + CI safely and cheaply

Plus a prompt library with sourcing rules built in, and a bot-traffic guide.

The toolkit

Understand your traffic (Search Console)

Script What it tells you
human-query-split Human vs AI-agent impressions per page; the real retitle list
weekly-report Week-over-week movers, trending queries, CTR triage
striking-distance Pages at position 5–20 where a push pays off most
ctr-audit The click gap against a CTR curve fitted to your site
gsc-rewrite-candidates Title rewrite candidates (optionally human queries only)
cannibalization-detector Queries where your own pages compete; flagged for review
query-gap-miner Queries you rank for without a page that targets them
refresh-tracker Aging pages that are losing traffic

Measure changes honestly

Script What it tells you
matched-control-readout Did the change work, compared with pages you didn't touch?
rewrite-measurer One change date vs a control band of similar-traffic pages; skips the change week and waits for complete data

Protect what you've built

Script What it catches
traffic-guard Daily traffic cliffs and slow recoveries (3 baselines)
sitemap-health Sitemap URLs that redirect, 404, are noindexed or canonicalize elsewhere
redirect-checker Redirect chains and sitemap URLs that redirect
broken-link-checker Broken internal and external links
schema-validator Invalid or incomplete JSON-LD (including Next.js and @graph)
meta-length-checker Titles and descriptions that will be truncated

GEO / AEO and data quality

Script What it does
evidence-verifier Checks every fact's quote is really on its source page before you publish
ai-citation-tracker Samples which domains AI engines cite for your queries (official APIs only)
factual-density-scorer How much specific, extractable fact a page carries
rewrite-detector News stories re-reporting the same event under a new URL

Content quality

Script What it finds
content-audit Every page scored by traffic and quality, with review labels
thin-content-detector Thin pages, with age and traffic guards before any noindex suggestion
template-detector AI template phrases (lists in config/anti-ai-rules.json)
orphan-finder Pages nothing links to

Feeds and indexing

Script Use it for
websub-ping Notifying a WebSub hub when an RSS/Atom feed that declares it changes
indexing-submitter Job posting and livestream pages only, as Google's policy requires

To run the weekly report on a schedule, copy examples/workflows/weekly-seo-report.yml into your site's repo.

Configs (quality gates, title rules, refresh rules, schema rules…) live in config/, and JSON-LD templates in schemas/. Sample outputs for the scripts are in samples/, generated from test fixtures.

Principles

  1. Measure against a control, or it didn't happen.
  2. Separate humans from machines before judging CTR.
  3. Every fact needs a source you can re-check. AI answer engines reward being right.
  4. Structured data mirrors visible content. No synthetic dates or invisible FAQs.
  5. Guardrails before automation. If a machine can publish it, a machine should also be able to catch it.
  6. No tricks that need hiding. Nothing here disguises automation or games a policy.

Results

Full numbers, windows and methods are in RESULTS.md. The site behind them is valueaddvc.com, a startup and venture-capital intelligence platform with ~1,100 articles, a daily news feed and thousands of programmatic company and investor pages.

Roadmap and community

The playbook and scripts are and stay free (MIT). See ROADMAP.md for what's next, including hosted audits. Contributions with real numbers and windows are very welcome; see CONTRIBUTING.md.

Built by Trace Cohen at Value Add VC. If this helped you, a ⭐ helps others find it.

License

MIT, see LICENSE.

About

The open-source SEO + GEO/AEO playbook: tested scripts, guardrails and methodology behind a site that reached 11.08M impressions / 44.8K clicks in 90 days. Human-vs-AI-agent query analysis, matched-control measurement, evidence-verified facts.

Topics

Resources

Contributing

Stars

71 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages