Skip to content

Commit 4aa57e0

Browse files
committed
docs: mitigation 6 efficacy review — Outcome recorded (keep + disposition lines)
The ~10-ship review that §9 committed to when the falsified-by stage shipped (PyAutoBrain#140, live 2026-07-17). Measured over the 22 review-leg ship gates since go-live: unverified-claim findings 0; gates with evidence the claim pass was exercised 2; the other ~20 recorded a bare 'review CLEAN', so a rote pass and a healthy one are indistinguishable in the ledger. Instrument validated live first (probe lifts 3/3; 349 tests pass): firing rate 13/50 Brain / 3/66 Mind merge messages since go-live, 'verified' driving 17/26 lifts, mostly of evidence sentences. Verdict: not proven rote — proven unobservable. Keep the stage, vocabulary unchanged; per-claim disposition lines in the verdict make rote visible (filed as a Mind feature prompt). Full numbers: PyAutoMind complete/2026/08/falsified-by-checkpoint-efficacy-review.md. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WH4NizvBK2jki2Uh5TMABh
1 parent 90d86b0 commit 4aa57e0

1 file changed

Lines changed: 40 additions & 1 deletion

File tree

docs/agent_failure_modes.md

Lines changed: 40 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -146,6 +146,43 @@ Each: catalogue entries caught → why it fires at the decisive moment → cost
146146
ship; failure mode: rote compliance — mitigate by keeping it to
147147
load-bearing claims only. **Trial it; measure whether it goes rote.**
148148

149+
**Outcome (2026-08-18, the committed ~10-ship efficacy review) — not
150+
proven rote; proven unobservable, which is its own finding.** Shipped
151+
2026-07-17 (PyAutoBrain#140) as the review faculty's `claims to falsify`
152+
surface + step-2a `unverified-claim` finding — reader-enforced as
153+
designed, but note that on autonomous ships the "reader" is the branch's
154+
own author (`review self-CLEAN` in the ledger). Across the 22 ship gates
155+
with a review leg since go-live (21 autonomy-log rows 2026-07-17→08-01
156+
plus one August cloud-session faculty run): `unverified-claim` findings
157+
raised: **0**; records showing the claim pass demonstrably exercised:
158+
**2** (2026-07-21 interpolator-aggregator-test-mode — the "no-op outside
159+
test mode" claim checked against its off-switch test; the 2026-08
160+
autohands-firewall-allowlist cloud run). The other ~20 verdicts are bare
161+
"review CLEAN" — a healthy pass and a rote one write the identical ledger
162+
row, exactly the failure signature this trial was supposed to look for.
163+
The instrument itself is healthy: probe claims lift 3/3 and the 5 pinning
164+
tests are green (re-validated 2026-08-18), and over the two organism
165+
repos' real merged history since go-live it fires on 13/50 (Brain) and
166+
3/66 (Mind) merge messages — neither always-empty nor saturated.
167+
`verified` does most of the lifting (17/26 lines) and mostly lifts the
168+
author's *evidence* sentence, not a bare claim: shipped messages in this
169+
window conspicuously carry "Verified by/against <probe>" inline, so the
170+
zero finding rate is at least partly deterrence, not only
171+
non-engagement. Idle-phrasing scoping holds at the level that matters (no
172+
changelog chatter lifted; residual false positives — mid-sentence wrap
173+
fragments, narrative `identical`/`unchanged` — cost seconds and carry no
174+
bypass pressure, unlike F5's refusal class). The one confirmed-wrong
175+
load-bearing claim in the window (the 2026-07-27 "5 siblings" count,
176+
falsified by an external Codex review) lived in an issue comment —
177+
outside the commit-message surface this stage reads. **Verdict: keep,
178+
vocabulary unchanged; close the observability gap** — the reviewing
179+
agent's verdict gains a one-line disposition per lifted claim
180+
(basis-cited / idle / finding) recorded in the ship evidence, so a rote
181+
pass becomes visible ledger drift per this doc's own ranking (detecting
182+
beats reminding). Filed: PyAutoMind
183+
`draft/feature/pyautobrain/review_claim_dispositions.md`; full numbers in
184+
PyAutoMind `complete/2026/08/falsified-by-checkpoint-efficacy-review.md`.
185+
149186
## 6. The memory system, attacked honestly
150187

151188
The day's evidence cuts both ways. Failures: B1, B2, and F1's note that did
@@ -193,7 +230,9 @@ mitigation 6 targets exactly the boundary-crossing moment.
193230
one line, highest ratio. 2 (Mind commit guard) and 3 (claim expiry): small
194231
tasks, file on approval. 4–5: mine the conductors, file individually.
195232
6 (falsified-by lines): trial on the next ship series, review whether it
196-
went rote after ~10 ships.
233+
went rote after ~10 ships — **done 2026-08-18**, see item 6's Outcome
234+
(verdict: keep + per-claim disposition lines; the review itself was
235+
PyAutoMind `complete/2026/08/falsified-by-checkpoint-efficacy-review.md`).
197236
- Whether F5's API-gate false-positive class warrants a scope fix (only scan
198237
arguments that are Python code, not path strings) — small, filed
199238
separately.

0 commit comments

Comments
 (0)