Skip to content

Repository files navigation

A giant red lobster leading a crew of small blue crab-bots, facing Hermes in winged sandals leading a flock of winged pink robots

claw-atlas

A comparative architecture analysis of 15 open-source "Claw" agent runtimes — always-on personal AI agents that live in your chat apps — read against their source code at pinned commits, one architectural dimension at a time.

The premise: modern agentic systems are converging on a shared anatomy. This repository maps the recurring patterns and treats the points where implementations diverge as the interesting result — that is where the open problems are. It is pattern-first; individual runtimes are evidence, not protagonists.

Every count in these documents (9/15, 7/15, 0/15) resolves back to named runtimes, and every claim about a runtime resolves back to a file and line range in code you can check out yourself. How that was produced, and how much to trust it, are documented below.

Start here

  • Want the findings? cross-dimension-ranking.md — the synthesis across all eight dimensions. It takes every counted claim from the talk and resolves it back to which named runtimes make up the numerator, asks whether it is the same set every time (it is not — three runtimes supply 37 of the 60 numerator slots across the 23 rarest features, and five supply none), and records the places where the analyses disagree with each other rather than smoothing them over.
  • Want one subject in depth? Pick a dimension from the table below.
  • Want the 30-minute version? The talk deck.

The eight dimensions

Each document covers all 15 runtimes, carries a comparison matrix, and cites file:line for its claims. Absences are recorded as findings: "none" means the excavation looked and found nothing.

Dimension The question A finding
Execution loop Who owns the call-model-run-tools cycle, how long may it run, what may interrupt it One skeleton fits all 15 — but iteration budgets span a 340× range
Context assembly What the model sees between "a message arrived" and "the model was called" SOUL.md is a folk standard (9/15); not one runtime counts tokens with the provider's real tokenizer
Memory What persists past a single model call, and how it is written, found and forgotten Markdown beat the database; five runtimes independently invented sleep; 0/15 enable embedding recall by default
Tool integration How the loop touches the world — catalogs, schemas, results Universal text-repair fallback and spill-to-artifact, neither in any textbook; schemas are contracts in some runtimes and mere advertisements in others
Channels & gateway One brain, many mouths: Telegram, Slack, WhatsApp, Signal, email, voice, web Convergent folk knowledge — Markdown dialect converters, fence-aware chunking, UTF-16 length counting — reinvented independently up to seven times
Proactivity & scheduling Acting when no user message arrived: cron, timers, heartbeats, wakes All 15 can fire a future turn, 13 let the model schedule its own, exactly 1 enforces a minimum cadence
Governance & sandboxing Whom the runtime obeys, and what a permitted action can actually reach Real architecture, disabled defaults: sandboxes available in 9/15, global emergency stop in 0/15
State durability What survives power loss, and what happens to in-flight work when the process dies Durable in pieces, globally lossy: no runtime has one transaction spanning transcript, run state, queue and external effect, so none achieves exactly-once

Supporting files: claws.md is the roster with upstream links and star counts (recorded July 2026); claws/ holds the runtimes themselves as git submodules; analysis/raw/ holds the 120 per-runtime reports the synthesis was built from; scripts/excavate.zsh is the tool that produced them.

Method

1. The corpus

15 runtimes, each vendored as a git submodule pinned to a specific commit, spanning TypeScript, Python, Rust, Go, Zig and C-on-ESP32 — from a 383k-star flagship down to a 5-star microcontroller port. Pinning is what makes the citations checkable: git submodule status gives you the exact tree every claim was read against.

Inclusion required being an actual runtime — a codebase that implements its own loop, memory and channels. Integration layers that install other runtimes were excluded, since counting them would double-count the runtime underneath. NVIDIA's NemoClaw was dropped for exactly this reason after excavation showed it installing OpenClaw and Hermes from pinned upstream archives.

2. Excavation: 120 reports

scripts/excavate.zsh <dimension> runs one analysis per runtime, per dimension — 15 × 8 = 120 reports in analysis/raw/. Each job is a codex exec run in a read-only sandbox with the working directory set to that one submodule, fed the dimension's prompt, writing a single Markdown report. Jobs run through a five-slot pool so a slow repository never idles the others.

The prompts — one per dimension, versioned at analysis/raw/<dimension>/PROMPT.md — are the load-bearing part of the method. Each fixes the report's section order and imposes the same four rules:

  • Every claim must carry a citation like src/memory/store.ts:120-180. No citation, no claim.
  • If a facet does not exist, write "none" explicitly — absence is a finding.
  • Read the real code; do not trust the README's marketing.
  • Distill the mechanism into language-neutral pseudocode, so implementations in six languages can be compared as designs rather than as syntax.

Models used were gpt-5.6-sol and gpt-5.6-terra, in batches recorded to analysis/raw/<dimension>/MODELS.tsv — with the caveat noted below.

3. Synthesis

The eight analysis/<dimension>.md documents are written from the 120 raw reports by an LLM under human direction — the corpus, the dimensions, the prompts and the structure were chosen by a human who iterated on the output, but the claims were not re-checked against the raw reports one by one. Each carries a matrix with a row per runtime, so a count is never asserted without the membership behind it being visible. Where the raw reports were thin or contradictory, the synthesis says so.

4. Cross-dimension pass

cross-dimension-ranking.md does a final pass over the finished analyses: it extracts every counted claim in the talk, resolves each to named runtimes, and checks whether the same handful of projects keeps supplying the numerator. Its §9 records six unresolved discrepancies — counts that two documents state differently, and members that are asserted but never named — rather than guessing to make the numbers line up. It was produced the same way as the eight analyses, and carries the same caveat.

How much to trust this

The honest limits, since the numbers above are only worth what the method behind them is worth:

  • Every document here is LLM-generated. The 120 raw reports were produced under a citation-or-no-claim rule against real source; the eight analyses and the cross-dimension pass were written from those reports the same way. A human chose the corpus, the dimensions, the prompts and the structure, and iterated on the output — but no claim has been verified line by line against the code. Treat the whole stack as a well-sourced research assistant's notes, not as a proof. The citations are there so you can check any claim that matters to you.
  • This is a snapshot. Every claim describes the pinned commit, not today's upstream. These are fast-moving projects; several claims here are already historical. Star counts in claws.md were recorded in July 2026.
  • Static reading only. Nothing was executed, benchmarked or load-tested. "Ships X" means the code implements X, not that X works well.
  • "None" means the excavation found nothing, not that nothing exists. A well-hidden feature can be missed.
  • MODELS.tsv coverage is partial. Batch logging was added to the script after six dimensions had already been excavated, so only governance-sandboxing and state-durability carry their full model history; the rest record only a later fill-in run.

Getting the code

git clone --recurse-submodules https://github.com/zendegani/claw-atlas.git
# or, after a plain clone:
git submodule update --init

Reproducing an excavation additionally needs the codex CLI on your PATH:

zsh scripts/excavate.zsh memory                 # all 15 runtimes, one dimension
zsh scripts/excavate.zsh memory openclaw hermes-agent   # or just a few

The talk

How AI Butlers Actually Work — Architecture Patterns from the Claws Ecosystem, presented at the AI Düsseldorf monthly meetup, July 2026.

View the deck — 28 slides across all eight dimensions, built entirely on the analyses in this repository.

License

PolyForm Noncommercial License 1.0.0.

This covers this repository's own content — the documents in analysis/, the tooling in scripts/, the artwork in assets/, and the root Markdown files. It does not cover the runtimes under claws/: those are third-party projects vendored as submodules, each governed by its own upstream license.

About

15 open-source AI-butler runtimes, compared across 8 architectural dimensions: SOUL.md is a folk standard (9/15), sandboxes ship disabled, nobody has a global emergency stop.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages