A comparative architecture analysis of 15 open-source "Claw" agent runtimes — always-on personal AI agents that live in your chat apps — read against their source code at pinned commits, one architectural dimension at a time.
The premise: modern agentic systems are converging on a shared anatomy. This repository maps the recurring patterns and treats the points where implementations diverge as the interesting result — that is where the open problems are. It is pattern-first; individual runtimes are evidence, not protagonists.
Every count in these documents (9/15, 7/15, 0/15) resolves back to named runtimes, and every claim about a runtime resolves back to a file and line range in code you can check out yourself. How that was produced, and how much to trust it, are documented below.
- Want the findings?
cross-dimension-ranking.md— the synthesis across all eight dimensions. It takes every counted claim from the talk and resolves it back to which named runtimes make up the numerator, asks whether it is the same set every time (it is not — three runtimes supply 37 of the 60 numerator slots across the 23 rarest features, and five supply none), and records the places where the analyses disagree with each other rather than smoothing them over. - Want one subject in depth? Pick a dimension from the table below.
- Want the 30-minute version? The talk deck.
Each document covers all 15 runtimes, carries a comparison matrix, and cites file:line for its claims. Absences are recorded as findings: "none" means the excavation looked and found nothing.
| Dimension | The question | A finding |
|---|---|---|
| Execution loop | Who owns the call-model-run-tools cycle, how long may it run, what may interrupt it | One skeleton fits all 15 — but iteration budgets span a 340× range |
| Context assembly | What the model sees between "a message arrived" and "the model was called" | SOUL.md is a folk standard (9/15); not one runtime counts tokens with the provider's real tokenizer |
| Memory | What persists past a single model call, and how it is written, found and forgotten | Markdown beat the database; five runtimes independently invented sleep; 0/15 enable embedding recall by default |
| Tool integration | How the loop touches the world — catalogs, schemas, results | Universal text-repair fallback and spill-to-artifact, neither in any textbook; schemas are contracts in some runtimes and mere advertisements in others |
| Channels & gateway | One brain, many mouths: Telegram, Slack, WhatsApp, Signal, email, voice, web | Convergent folk knowledge — Markdown dialect converters, fence-aware chunking, UTF-16 length counting — reinvented independently up to seven times |
| Proactivity & scheduling | Acting when no user message arrived: cron, timers, heartbeats, wakes | All 15 can fire a future turn, 13 let the model schedule its own, exactly 1 enforces a minimum cadence |
| Governance & sandboxing | Whom the runtime obeys, and what a permitted action can actually reach | Real architecture, disabled defaults: sandboxes available in 9/15, global emergency stop in 0/15 |
| State durability | What survives power loss, and what happens to in-flight work when the process dies | Durable in pieces, globally lossy: no runtime has one transaction spanning transcript, run state, queue and external effect, so none achieves exactly-once |
Supporting files: claws.md is the roster with upstream links and star counts (recorded July 2026); claws/ holds the runtimes themselves as git submodules; analysis/raw/ holds the 120 per-runtime reports the synthesis was built from; scripts/excavate.zsh is the tool that produced them.
15 runtimes, each vendored as a git submodule pinned to a specific commit, spanning TypeScript, Python, Rust, Go, Zig and C-on-ESP32 — from a 383k-star flagship down to a 5-star microcontroller port. Pinning is what makes the citations checkable: git submodule status gives you the exact tree every claim was read against.
Inclusion required being an actual runtime — a codebase that implements its own loop, memory and channels. Integration layers that install other runtimes were excluded, since counting them would double-count the runtime underneath. NVIDIA's NemoClaw was dropped for exactly this reason after excavation showed it installing OpenClaw and Hermes from pinned upstream archives.
scripts/excavate.zsh <dimension> runs one analysis per runtime, per dimension — 15 × 8 = 120 reports in analysis/raw/. Each job is a codex exec run in a read-only sandbox with the working directory set to that one submodule, fed the dimension's prompt, writing a single Markdown report. Jobs run through a five-slot pool so a slow repository never idles the others.
The prompts — one per dimension, versioned at analysis/raw/<dimension>/PROMPT.md — are the load-bearing part of the method. Each fixes the report's section order and imposes the same four rules:
- Every claim must carry a citation like
src/memory/store.ts:120-180. No citation, no claim. - If a facet does not exist, write "none" explicitly — absence is a finding.
- Read the real code; do not trust the README's marketing.
- Distill the mechanism into language-neutral pseudocode, so implementations in six languages can be compared as designs rather than as syntax.
Models used were gpt-5.6-sol and gpt-5.6-terra, in batches recorded to analysis/raw/<dimension>/MODELS.tsv — with the caveat noted below.
The eight analysis/<dimension>.md documents are written from the 120 raw reports by an LLM under human direction — the corpus, the dimensions, the prompts and the structure were chosen by a human who iterated on the output, but the claims were not re-checked against the raw reports one by one. Each carries a matrix with a row per runtime, so a count is never asserted without the membership behind it being visible. Where the raw reports were thin or contradictory, the synthesis says so.
cross-dimension-ranking.md does a final pass over the finished analyses: it extracts every counted claim in the talk, resolves each to named runtimes, and checks whether the same handful of projects keeps supplying the numerator. Its §9 records six unresolved discrepancies — counts that two documents state differently, and members that are asserted but never named — rather than guessing to make the numbers line up. It was produced the same way as the eight analyses, and carries the same caveat.
The honest limits, since the numbers above are only worth what the method behind them is worth:
- Every document here is LLM-generated. The 120 raw reports were produced under a citation-or-no-claim rule against real source; the eight analyses and the cross-dimension pass were written from those reports the same way. A human chose the corpus, the dimensions, the prompts and the structure, and iterated on the output — but no claim has been verified line by line against the code. Treat the whole stack as a well-sourced research assistant's notes, not as a proof. The citations are there so you can check any claim that matters to you.
- This is a snapshot. Every claim describes the pinned commit, not today's upstream. These are fast-moving projects; several claims here are already historical. Star counts in
claws.mdwere recorded in July 2026. - Static reading only. Nothing was executed, benchmarked or load-tested. "Ships X" means the code implements X, not that X works well.
- "None" means the excavation found nothing, not that nothing exists. A well-hidden feature can be missed.
MODELS.tsvcoverage is partial. Batch logging was added to the script after six dimensions had already been excavated, so onlygovernance-sandboxingandstate-durabilitycarry their full model history; the rest record only a later fill-in run.
git clone --recurse-submodules https://github.com/zendegani/claw-atlas.git
# or, after a plain clone:
git submodule update --initReproducing an excavation additionally needs the codex CLI on your PATH:
zsh scripts/excavate.zsh memory # all 15 runtimes, one dimension
zsh scripts/excavate.zsh memory openclaw hermes-agent # or just a fewHow AI Butlers Actually Work — Architecture Patterns from the Claws Ecosystem, presented at the AI Düsseldorf monthly meetup, July 2026.
View the deck — 28 slides across all eight dimensions, built entirely on the analyses in this repository.
PolyForm Noncommercial License 1.0.0.
This covers this repository's own content — the documents in analysis/, the tooling in scripts/, the artwork in assets/, and the root Markdown files. It does not cover the runtimes under claws/: those are third-party projects vendored as submodules, each governed by its own upstream license.
