Skip to content

Benchmark continuation-prompt clarity for a batteries-included default config #386

Description

@Shearerbeard

Placeholder

The transitive dep bug (#221) came down to continuation-prompt clarity:
workers don't need every ancestor's output, just the right context passed down
the direct line. The #249 fix passed the
full ancestor chain under a token budget - scores improved, but the
continuation prompt and history stayed just as noisy - so we closed it and are
superseding it with the restructuring prototyped in #374. Notes and evidence
live in the terminalbench-aura repo.

This ticket tracks benchmarking that approach on terminalbench against a
neutral default config for #212 (batteries-included): a generic set of worker
prompts (explore, verify, ...) instead of application-specific roles, so the
default config ships with worker prompts that earn their place on benchmark
evidence rather than SRE-specific tuning.

Details land as the daily runs produce them - this is a stub so the work is
visible on the board.

Activity

  1. self-assigned this
    on Jul 17, 2026
  2. teriyakichild commented on Aug 3, 2026

    @teriyakichild
    Contributor

    Benchmark verdict on the neutral-workers question, from the init-template round-2 campaign (full writeup in the ai-experiments repo, HANDOFF-init-template-round2-results.md; runs cataloged in aura-e2e/results-index.json, ids sonnet46-r2-*).

    The generic analyst/operator/debugger/verifier split from terminalbench-aura, ported onto the RCA mock-mcp suite with the same fan-out planning and nudges as the domain-worker arms (Sonnet 4.6/Bedrock, two runs each):

    Arm Score (adj) Tokens ~s/scenario
    Domain workers, inline guidance 143 / 141 of 144 5.68M / 5.58M 332 / 303
    Generic workers 144 / 143 of 144 12.30M / 13.03M 370 / 385

    ("adj" restores 12 points the harness's worker:log-analyst present assertion takes from any non-domain worker naming. That's a harness artifact; the assertion needs parameterizing before another arm like this.)

    Score parity, 2.2x token cost, 15-25% more wall time. The cost is structural: every generic worker carries the full 23-tool surface, and the verifier adds a fixed ~87s serial phase to every scenario, including the ones that were already right. The debugger never dispatched in 24 scenarios.

    So #505 keeps domain workers as the init default. Two pieces of the generic shape survived into the template: the self-contained-task-description prose in the planning section, and the verifier as an opt-in commented-out worker. Proposing the generic set ships as a reference config instead; noted on #241.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions