Skip to content

Latest commit

 

History

History
200 lines (150 loc) · 22.4 KB

File metadata and controls

200 lines (150 loc) · 22.4 KB

Execution Loop

Dimension 1 of the claw-atlas comparative analysis. Per-runtime evidence with full file:line citations lives in raw/execution-loop/; all citations resolve against the submodule commits pinned in this repository.

The execution loop is the heartbeat of an agent runtime: call the model, look at what came back, execute any requested tools, feed the results back in, repeat until there is a final answer. Every one of the 15 runtimes needs this cycle to exist — but they disagree profoundly about who owns it, how long it may run, and what is allowed to interrupt it.

1. The first divergence: who owns the loop?

Before comparing loop designs, a more basic question splits the field: does the runtime implement the model→tool cycle at all?

Custody Runtimes What the runtime actually does
Owned — hand-written model→tool iteration OpenClaw, IronClaw, microclaw, ZeptoClaw, PicoClaw, ApexClaw, Hermes, nanobot, MimiClaw, RT-Claw, KrillClaw, NullClaw Builds requests, parses responses, executes tools, decides when to stop
Delegated — an SDK or external agent CLI iterates nanoclaw (Claude Agent SDK), Mercury (Vercel AI SDK streamText/generateText), TinyAGI (spawns claude/codex/opencode CLIs) Supplies tools/config and observes events; the actual cycling is someone else's code

This is the single most useful classifier for the ecosystem. The "delegated" tier is architecturally honest about a trade: nanoclaw gets compaction, retries, and parallel tools for free from the Claude Agent SDK but cannot see inside a turn (container/agent-runner/src/providers/claude.ts:532-559); TinyAGI goes further and treats entire agent CLIs as opaque subprocesses, reduced to parsing their JSONL progress lines (packages/core/src/invoke.ts:76-172). NullClaw sits firmly on the other side: Agent.turn owns a synchronous bounded model→tool loop, textual/native call normalization, execution, recovery, and finalization (src/agent/root.zig:1941-2921).

2. The convergent skeleton

Every owned loop in the ecosystem — from OpenClaw's ~3,000-line TypeScript core down to MimiClaw's C loop on an ESP32 microcontroller — reduces to the same skeleton:

while not done and budget remains:
    response = call_model(messages, tools)
    if response has no tool calls:  return response.text     # final answer
    append(assistant tool-call message)
    results = execute(response.tool_calls)
    append(tool results)                                     # observations for next turn

Three further convergences hold across all 15 runtimes, and they are the strongest findings of this dimension:

  1. Nobody executes tools mid-stream. Even runtimes that stream token-by-token (OpenClaw, Hermes, KrillClaw) assemble tool-call JSON fragments during the stream but wait for the complete response before dispatching a single tool. Streaming is a presentation and liveness channel, never a control channel. KrillClaw is the cleanest illustration: its SSE parser assembles tool_use blocks incrementally (src/stream.zig:92-140) yet execution strictly follows stream completion (src/agent.zig:92-153). Hermes states the inverse motivation explicitly: it streams even when nobody is watching, because chunk arrival doubles as transport-health monitoring (agent/conversation_loop.py:1382-1392).
  2. Tool failures are observations, not exceptions. In every owned loop, a failed tool becomes an error-valued result appended to the transcript for the model to react to on the next iteration — the model, not the runtime, decides whether to retry, work around, or give up. (PicoClaw: pkg/agent/pipeline_execute.go:687-720; ZeptoClaw: src/agent/loop.rs:1522-1539; Hermes: agent/tool_executor.py:603-642; even 9.4-KLOC-total MimiClaw: main/agent/agent_loop.c:153-165.)
  3. The loop is never alone. Every long-running runtime wraps the ReAct cycle in at least one outer loop that owns message intake and dispatch — a queue-consuming worker (MimiClaw's FreeRTOS queue, RT-Claw's singleton worker, nanobot's asyncio dispatcher, PicoClaw's channel loop, TinyAGI's per-agent promise chains) or a lease-backed scheduler (IronClaw). OpenClaw stacks three control layers: streaming ReAct core, session-level retry/compaction/continuation, and a whole-attempt recovery loop with auth rotation and model failover (packages/agent-core/src/agent-loop.ts:270-450, src/agents/sessions/agent-session-prompting.ts:24-81, src/agents/embedded-agent-runner/run-loop.ts:273-622).

3. Iteration budgets: a 340× spread

How many model calls may one user message trigger? The defaults span more than two orders of magnitude:

Budget Runtime Notes
3 RT-Claw MAX_TOOL_ROUNDS 3 (claw/services/ai/ai_engine.c:46) — and if call 3 requests tools, the results are computed but never shown to the model; the turn returns an error
10 MimiClaw, KrillClaw Both hard-coded. KrillClaw has a max_turns = 50 config value that the loop never reads (src/types.zig:145, src/react.zig:12-13)
20 ZeptoClaw, ApexClaw ApexClaw grants one extra out-of-budget model call to explain the blocker (core/apexclaw.go:936-962)
25 SDK steps Mercury Halved to 12 in "saver mode" (src/core/saver-mode.ts:20-23)
50 PicoClaw Soft: pending steering keeps the loop alive past the cap (pkg/agent/turn_coord.go:92-95)
90 Hermes Transactional: compression restarts refund consumed budget (agent/conversation_loop.py:4381-4409)
100 microclaw Warns the model at 70% and 90% consumption (src/agent_engine.rs:1743-1765)
200 nanobot Plus bounded injection cycles on top (nanobot/config/schema.py:128-135)
1,000 NullClaw Config default; accepted mid-turn injections can add at most eight iterations, then exhaustion gets one tool-free summary call (src/config_types.zig:406-428, src/agent/root.zig:2170-2176, 2815-2920)
1,024 IronClaw Effectively "trust the stop strategies" (crates/ironclaw_agent_loop/src/strategies/budget.rs:25-66)
∞ OpenClaw, nanoclaw, TinyAGI OpenClaw deliberately leaves the inner loop unbounded and bounds pathology instead: run deadlines, loop-detection guards, and a 32–160 cap on whole-attempt recoveries (src/agents/embedded-agent-runner/run/helpers.ts:121-132)

The budget number is a proxy for a philosophical stance: small budgets (embedded claws) treat the model as a subroutine that had better converge fast; large or absent budgets (OpenClaw, IronClaw) treat the loop as a long-lived process whose safety comes from behavioral guards, not counters.

Two graceful-exhaustion refinements recur: several runtimes spend one final tool-free synthesis call to wring an answer out of a budget-exhausted turn (Hermes agent/turn_finalizer.py:69-112, nanobot nanobot/agent/runner.py:655-678, ZeptoClaw src/agent/loop.rs:1673-1771, ApexClaw, IronClaw's "final-answer nudge", NullClaw src/agent/root.zig:2815-2920), while the minimal claws simply return a generic error (MimiClaw main/agent/agent_loop.c:302-315).

4. Loop-detection guards: the undocumented convergence

Textbook agent-loop descriptions have a max-iteration counter and nothing else. The field disagrees: seven runtimes independently grew behavioral stuck-loop detectors, which is as strong a signal of a real production problem as this dataset offers.

  • Mercury is the most elaborate: exact-repetition windows, consecutive-failure counts, text-similarity checks, parameter-diversity "productivity analysis," optional user confirmation, and an auxiliary-LLM self-check capped at three escalations per request (src/core/agent.ts:66-291).
  • microclaw works at two levels: six identical whole tool-use turns kill the run; repeated individual (tool, args) pairs get replaced with synthetic error results (src/agent_engine.rs:1268-1304, src/agent_engine.rs:1606-1651).
  • KrillClaw hashes tool calls into an eight-entry ring; a third identical call becomes a synthetic "try another approach" observation (src/react.zig:48-80).
  • ApexClaw stops after the same tool fails first-in-batch on two consecutive iterations, then asks the model to explain the failure (core/apexclaw.go:891-932).
  • IronClaw terminates on repeated no-progress capability results or repeated tiny outputs (crates/ironclaw_agent_loop/src/strategies/stop.rs:337-388); ZeptoClaw has a block/circuit-breaker guard (src/agent/loop.rs:126-228); OpenClaw ships loop-detection guards in its unbounded design.

Note the two-tier response pattern: soft (synthesize an error observation, let the model self-correct — KrillClaw, microclaw per-call) versus hard (terminate the run — Mercury, ApexClaw, IronClaw). Several runtimes use both.

5. Streaming: four positions, one invariant

Given the universal "no tools mid-stream" invariant (§2), streaming reduces to how much of the token stream reaches the user before the turn completes:

  • Streaming-native: OpenClaw (full event stream with message_update deltas; packages/agent-core/src/agent-loop.ts:487-542), Hermes (streaming by default, as liveness), KrillClaw (text deltas straight to stdout).
  • Conditionally streaming: microclaw, nanobot, PicoClaw, Mercury, NullClaw — streaming happens only when a consumer exists (an event sink, callbacks, or a channel that supports it); headless invocations of the same loop silently use blocking calls. Streaming here is presentation-aware plumbing, not architecture. PicoClaw disables streaming whenever model-fallback candidates exist (pkg/agent/pipeline_streaming.go:287-296); NullClaw removes native tool schemas and recovers only textual JSON/XML calls after the stream completes (src/agent/root.zig:2190-2227, 2484-2544).
  • Two-pass: ZeptoClaw asks the model non-streamingly whether tools are needed, then — if the answer is final — re-issues the request as a streaming call purely to deliver the answer incrementally. A text-only "streamed" reply costs two model calls (src/agent/loop.rs:2132-2189, src/agent/loop.rs:2770-2797).
  • None: MimiClaw, RT-Claw (buffered HTTP, whole-document JSON parse), ApexClaw (accumulates SSE deltas internally, delivers complete strings; model/client.go:518-594).

6. Concurrency: sessions serialized, tools sometimes parallel

Session serialization is universal in intent, diverse in mechanism: OpenClaw uses per-session execution "lanes" plus a global concurrency lane (default 4); nanobot an asyncio.Lock per session plus a global semaphore (default 3); IronClaw a durable scope lock in the store; microclaw a per-chat Tokio mutex; ApexClaw a TryLock that rejects concurrent turns with ErrBusy instead of queueing; MimiClaw and RT-Claw get serialization for free from having exactly one worker. The soft spots are instructive: microclaw deliberately proceeds without the lock after a 60-second wait, trading exclusivity for liveness (src/chat_turn_queue.rs:121-150), and Hermes doesn't serialize at all — it just logs a warning when turns overlap (agent/turn_context.py:367-377).

Parallel tool execution within a turn is where claims and reality diverge most:

Policy Runtimes
Parallel by default, sequential opt-out OpenClaw (Promise.all, deterministic transcript order preserved)
Safety-classed: read-only/safe calls parallel, side-effecting serial microclaw (waves), ZeptoClaw, Hermes (effect-aware segments, 8 workers), nanobot (concurrency_safe batches), ApexClaw (semaphore of 4, parallel batch before sequential batch)
Sequential only PicoClaw (production path), MimiClaw, RT-Claw, KrillClaw, NullClaw
Advertised but inert IronClaw — planner labels batches Parallel, the host adapter awaits each call in a plain for loop (crates/ironclaw_loop_host/src/capability_port.rs:1729-1750); microclaw — parallel_tool_max_concurrency is accepted and never enforced (src/tool_executor.rs:143-144)

7. Mid-turn steering: the butler ecosystem's signature move

A chat-native agent has a problem a CLI coding agent doesn't: the user keeps typing while the agent works. Nine runtimes solve this by injecting messages into the running loop at safe checkpoints rather than queueing a whole new turn:

  • OpenClaw: first-class steering + follow-up queues polled at turn boundaries; steering cannot cancel already-requested tools (packages/agent-core/src/agent-loop.ts:329-337, 419-443).
  • nanoclaw: a 500 ms side-poll pushes new messages into the live SDK query; the model session stays open across user turns (container/agent-runner/src/poll-loop.ts:363-443).
  • PicoClaw: steering checkpoints before model calls, after replies, after individual tools — and pending steering overrides the iteration cap (pkg/agent/turn_coord.go:92-95).
  • nanobot: bounded injection (max 5 cycles × 3 messages) so steering can't extend a turn forever (nanobot/agent/runner.py:54-57).
  • microclaw: drains pending messages at two safe boundaries (after tools; just before honoring end_turn) (src/agent_engine.rs:1425-1459, 1767-1799).
  • IronClaw: steering drained each iteration; follow-ups can revive a would-be-terminal reply (crates/ironclaw_agent_loop/src/executor/canonical.rs:278-315).
  • NullClaw: pending input is drained before model calls and before accepting a final response; injections can extend the active loop by at most eight iterations (src/agent/root.zig:2170-2185, 2587-2601).

The minimal claws (MimiClaw, RT-Claw, KrillClaw, ApexClaw) have no mid-turn input path — a message sent while the agent works waits for the next turn or is rejected. This is the cleanest single feature separating "chat-native butler" architectures from "request/response agent" architectures.

8. Failure handling: where the engineering budget went

Model-call retry stacks range from nothing to four layers deep. KrillClaw and MimiClaw retry nothing. ApexClaw layers loop-level retries (3) × client-level retries (3) × ZAI token rotation (4), so one logical model step can produce up to nine physical requests, more with token rotation (core/apexclaw.go:709-728, model/client.go:74-133). Mercury retries nothing per-provider but fails over across providers, re-entering with the original prompt — after side effects may already have happened (src/core/agent.ts:1839-1866). Hermes combines retries, credential rotation, provider fallback, Retry-After honoring, and transport rebuilds (agent/conversation_loop.py:3319-3391).

Context overflow (the loop-visible reaction; compaction internals are the memory dimension's topic) splits three ways:

  • Reactive compact-and-retry: OpenClaw (session compaction + up to 3 outer overflow recoveries + tool-result truncation), ZeptoClaw (typed ContextOverflow error → compact → retry ×3), IronClaw (force_compact_on_next_iteration restart), Hermes (≤3 compressions, only retrying if footprint actually shrank; distinguishes input overflow from output-cap errors), PicoClaw (compress + trim, fail if protected messages won't fit), NullClaw (forced compression and retry on the blocking path only).
  • Proactive only: nanobot (pre-request governance fits every request; a reactive overflow error is deliberately non-fallbackable), KrillClaw (pre-call truncation), microclaw (message-count compaction before the loop; a mid-loop overflow just propagates), RT-Claw (near-capacity compression pre-loop).
  • None: MimiClaw, ApexClaw, Mercury, TinyAGI — overflow is indistinguishable from any other model error.

Durability is IronClaw's differentiator, unique in the dataset: checkpoints before model calls, before side effects, at blocking gates, and at exits make the loop resumable — a blocked approval exits the process entirely and a later claim resumes from checkpointed state (crates/ironclaw_agent_loop/src/executor/canonical.rs:392-490). nanoclaw gets a coarser version by persisting the SDK continuation ID at session init (not turn end), so a crash mid-turn resumes the already-created session (container/agent-runner/src/poll-loop.ts:487-495). Everyone else restarts turns from scratch.

9. Oddities worth quoting (the "config lies" collection)

Real codebases drift, and this ecosystem is young. Findings the audience should hear because they're instructive, not embarrassing:

  • KrillClaw's max_turns config is loaded and never consulted; the true limit is a hard-coded constant (src/config.zig:129-141 vs src/react.zig:12-13).
  • PicoClaw's max_tool_iterations actually bounds model iterations; one iteration may run many tools (pkg/config/config.go:435).
  • Mercury's user-facing warning says "25 calls"; the implemented ceiling is 75 — and calling use_skill resets the counter entirely (src/core/agent.ts:73-76, 283-290).
  • microclaw's parallel-concurrency cap and IronClaw's parallel batch policy are both structurally present and behaviorally inert (§6).
  • RT-Claw's third round executes tools whose results no model call will ever read (§3).
  • ApexClaw doesn't use native tool-calling at all: tools are a prompt-level protocol — the model emits fenced JSON in plain text, the harness parses it with tolerant JSON repair, and results go back as user-role messages (model/toolcall.go:36-124).
  • NullClaw's ordinary configured default is 1,000 iterations even though a separate constant and tests say 25; the latter is not what fromConfig copies (src/config_types.zig:406-413, src/agent/root.zig:51-53, 613-614).

10. Distilled pseudocode: the three shapes

Minimal owned loop (MimiClaw, RT-Claw, KrillClaw — the whole idea in 8 lines):

on message from inbound queue:
    messages = history + user_message
    for i in 0..SMALL_CAP:                       # 3–10
        resp = blocking_model_call(messages, tools)
        if no tool calls: reply(resp.text); break
        messages += resp.tool_calls
        messages += execute_sequentially(resp.tool_calls)
    else: reply(generic_error)

Layered production loop (OpenClaw, IronClaw, nanobot, Hermes — same core, wrapped in recovery):

scheduler: claim run from queue, under session lane/lock
outer attempts (bounded):
    loop:
        checkpoint / drain steering / compact if needed
        resp = stream_model(messages, tools)         # retries, failover inside
        if error/abort: classify → retry, compact-restart, or fail
        if no tools and no queued follow-up: break
        results = execute(tools, safety-classed parallelism)
        messages += results; check stop strategies & loop guards
    on overflow/timeout/auth failure: recover and retry whole attempt
finalize: tool-free synthesis call if budget died before an answer

Delegated loop (nanoclaw, Mercury, TinyAGI — custody inverted):

on message batch:
    query = sdk_or_cli.start(prompt, tools, config)   # THE loop lives here
    while events arrive:
        forward progress / heartbeat liveness
        push newly arrived user messages into live query   # nanoclaw
    on result: deliver text; keep query open for next turn
watchdog (host side): kill/respawn silent sessions

11. Comparison matrix

Runtime Custody Language Shape Cap (default) Streaming at loop Parallel tools Mid-turn injection Reactive overflow recovery Resumable mid-turn
OpenClaw owned TypeScript 3-layer streaming ReAct ∞ inner / 32–160 attempts yes (deltas) yes (default) steering + follow-ups compact + 3 recoveries no
nanoclaw delegated (Claude Agent SDK) TypeScript/Bun queue + live SDK query none set events only SDK-owned push into live query SDK auto-compact session-level
NullClaw owned Zig synchronous bounded while ReAct 1,000 (+8 injected) conditional; textual tools only while streaming sequential bounded boundary injection blocking: compress + retry; streaming: none no
TinyAGI delegated (agent CLIs) TypeScript queue + promise chains none set JSONL progress CLI-owned none none no
Mercury delegated (Vercel AI SDK) TypeScript FIFO + SDK steps 25 steps text stream SDK-owned none (slash cmds only) none no
IronClaw owned Rust staged pipeline + scheduler 1,024 opportunistic text policy yes, adapter serial steering + follow-ups compact-restart yes (checkpoints)
microclaw owned Rust bounded for ReAct 100 conditional text safety waves (cap inert) 2 safe boundaries none (proactive only) no
ZeptoClaw owned Rust queue + while ReAct 20 two-pass safety-classed limited replay compact ×3 no
PicoClaw owned Go queue + phase coordinator 50 (soft) conditional accumulated sequential steering checkpoints compress + trim no
ApexClaw owned Go for ReAct, prompt-level tools 20 (+1) none (buffered) sem-4 parallel then serial none none no
Hermes owned Python while ReAct + retry inner loop 90 (refundable) yes (as liveness) effect-aware segments interrupt only compress ×3 no
nanobot owned Python queue + state machine + for 200 conditional segments safe batches injection (bounded) none (proactive governance) checkpoint restore
MimiClaw owned C / FreeRTOS queue + while ReAct 10 none none none none no
RT-Claw owned C (RTOS/POSIX) queue + for ReAct 3 none none none none no
KrillClaw owned Zig while ReAct, single-threaded 10 (config ignored) text deltas none none none (proactive truncation) no

12. Takeaways for the talk

  1. Convergence: one skeleton fits all 15; tools never run mid-stream; tool errors are observations; a queue/dispatch loop always wraps the ReAct loop. This is the "anatomy" claim, and it holds from an 18-KLOC Rust pipeline down to a microcontroller.
  2. The custody spectrum (owned → delegated) is the ecosystem's real fault line, and it maps to a build-vs-buy decision every team in the audience will face.
  3. Budgets encode philosophy: 3 → 1,024 → ∞ is not tuning noise; it's the difference between "model as subroutine" and "loop as process with behavioral guards."
  4. Loop-detection guards are the field's folk knowledge — seven independent implementations of a mechanism textbooks don't mention.
  5. Mid-turn steering separates butlers from request/response agents — the feature that makes a chat-native agent feel alive, and the minimal claws' most visible omission.