Grow the readers who arrive from ChatGPT, Perplexity and the rest — starting with which of them actually send you any.
AI assistants read your site, and sometimes send a human back to it afterwards. Some vendors return readers. Some take thousands of pages and return nobody. The analytics you already run cannot tell those two apart: the crawl and the visit happen hours apart, under different names, with nothing joining them up.
Crawlytics joins them up. For every vendor it sets what it took — crawls — against what it sent back — humans arriving from that vendor's assistant — and turns that into one comparable price: one reader per so many crawls. With that in hand you can spend your effort on the assistants that return readers, and stop paying to feed the ones that never do.
It is self-hosted: it ingests your web server logs and edge or middleware events, classifies AI crawlers and AI assistant referrals, verifies known bots, and runs multi-site dashboards. Your logs stay on your server.
On your own host, behind automatic HTTPS. There is no image to pull: the installer builds from a checkout of this repository.
git clone https://github.com/metawhisp/crawlytics-oss.git
cd crawlytics-oss/deploy && ./install.shIt asks for a domain and an email, generates its own secrets, brings up the app and ClickHouse behind Caddy, schedules a daily backup (a systemd timer or cron), and issues a TLS certificate on first request. Then open the dashboard, sign in with the password it set, and add a site under Setup — the wizard hands you the snippet for your sensor and flips to a tick of its own accord when the first events arrive.
Full operations guide, upgrades and backups: deploy/README.md. Sensor options (Cloudflare Worker, Node middleware, log tailer, WordPress plugin): docs/sensors.md. The Node middleware and the log tailer install from npm; the WordPress plugin is not in the WordPress plugin directory yet and is built from the checkout. What changed in 1.0.0: CHANGELOG.md.
Crawlytics speaks MCP, so Claude (or any MCP client) can query it directly with a read-only key scoped to one site — 13 tools over the same API the dashboard uses. The Setup tab generates the key and the client config.
apps/server- Fastify app for ingest, query APIs, auth, cron jobs, and serving the SPA.apps/web- React dashboard SPA.packages/detector- pure TypeScript classification core.packages/registry- bot registry compiler.packages/ingest-cli- log import and tail CLI.packages/sensor-cloudflare- Cloudflare Worker sensor.packages/sensor-node- Next.js and Express middleware sensor.packages/sensor-wordpress- WordPress plugin sensor.packages/shared- shared schemas and types.deploy- self-host deployment assets.
pnpm install
pnpm build
pnpm typecheck
pnpm lint
pnpm test
docker compose -f deploy/compose.yml --env-file deploy/.env.example configPanel behaviour is covered by an integration suite that runs the real queries against a throwaway ClickHouse container (needs Docker; binds to 127.0.0.1 only):
pnpm --filter @crawlytics/server test:integrationEvery panel is measured from your own logs. That also means each one is bounded by what a log can prove, so the boundaries are stated rather than hidden.
- Forged bots are excluded from AI panels. Anyone can send
ChatGPT-Useras a user-agent. Where a vendor publishes IP ranges or reverse-DNS records, Crawlytics checks them and marks the requestverifiedorspoofed; spoofed traffic is kept out of retrieval, landing pages, crawl health and crawls-per-vendor. It is still visible in the Security tab, which is what that tab is for. - Some crawlers cannot be verified at all. Several vendors publish no ranges
and no PTR records, so their requests are marked
unverified— not proof of forgery, not proof of authenticity. They are counted, and labelled as such. humanmeans "a browser nobody could identify as a bot", not "a person". A user-agent the bot registry does not name counts as human only when it names a browser engine (AppleWebKit, Gecko, Trident or MSIE, KHTML, Presto); an SDK, a script or a bareMozilla/5.0is filed asother_bot. Automation that sends a real browser string still lands in human, so the dashboard labels those rows "unrecognized".- Traffic refused before it reaches a sensor is invisible. A request stopped by a WAF, a firewall, a CDN's bot protection or a challenge never reaches the Worker, middleware, plugin or log behind it. A vendor you block that way appears to take nothing — which reads the same as a vendor that never came.
- Gemini training crawls cannot be seen. Google fetches for Gemini with
plain Googlebot (
Google-Extendedis a robots.txt token, not a user-agent), which Crawlytics counts as a search engine. The Take vs Give panel says so wherever Google appears. - Take vs Give only judges what the data can prove. A vendor whose assistant's referrals Crawlytics does not track reads "clicks not tracked". "Sent nobody" is said only after a vendor has crawled enough for its silence to count, measured against the best rate of return on your site taken at its lower 95% bound; below that the panel says "not enough data", and with no provable rate at all, "no baseline".
- "AI blind spots" counts browser sessions only — sessions that also fetched a stylesheet, script or image, because that is what rendering a page looks like. If your assets are served by a CDN this instance never sees, the filter switches itself off rather than empty the panel.
- Bot classes are channels, not stages.
ai_training,ai_searchandai_fetcherlabel the bot that made a request. They are disjoint, so a page can receive AI referral clicks with zero recorded crawls; nothing in the UI presents them as a funnel. - A retrieval is not a citation. A log can show that an assistant's fetcher requested a page. Whether the answer then linked to it, quoted it, or ignored it never reaches your server, so no panel claims a citation. The one signal that a link was actually shown is a human arriving with an assistant as the referrer — and a referrer is set by the visitor's browser, so it is evidence, not proof.
- Verification is "not marked forged", not "proven genuine". Several retrieval bots are checked against a vendor IP list with no reverse-DNS fallback, and a stale list is used rather than none. A lagging list can mark a real bot as spoofed, which keeps it out of the AI panels until the list catches up.
GNU Affero General Public License v3.0 — see LICENSE.
You may use, modify and self-host this, including commercially, and the licence asks you to keep the copyright notices in NOTICE. The part that distinguishes AGPL from a permissive licence is section 13: if you run a modified version as a service other people reach over a network, you have to offer those users the source of your version.