Project-scoped lessons. Lead with the rule, then Why: and When to apply:.
2026-06-01 — shell scripts are macOS-authored but run in Ubuntu CI: force byte locale, don't trust local green
Rule: Any awk/sed that manipulates non-ASCII bytes (em-dash, etc.) must run
under LC_ALL=C so it stays byte-oriented. A passing run on macOS's BSD awk
proves nothing about Ubuntu CI, where gawk/mawk in a UTF-8 locale treat the same
bytes as multi-byte codepoints. Also: a --no-build/offline CI path can only
verify committed artifacts — anything gitignored as a build product (e.g.
tests/smoke/output/last-answer.md) must be tracked as a baseline or the check
fails on a fresh checkout.
Why: wiki-lint-typed-relations.sh built the em-dash via
sprintf("%c%c%c",226,128,148) and matched it with index(). On BSD awk that
yields the 3 raw bytes (R7 green locally); on Ubuntu CI's UTF-8 awk each %c
became a 2-byte char, so index() never matched and verbs were miscounted (R7
red). Separately, C4 reads the captured query answer, which was gitignored — CI
had no file to assert on. Both were invisible until the PR's Ubuntu run; local
macOS green was a false signal.
When to apply: Before claiming "CI-verified" — actually push and watch the Ubuntu run, never infer it from a macOS pass. Any new shell check touching non-ASCII, or any CI assertion on a generated artifact.
Rule: When guarding scripts/body-hash.sh (and verify-extract.sh) against
malformed frontmatter, validate as line 1 is --- AND there are ≥ 2 ^---$
delimiters — never exactly 2. Markdown bodies legitimately contain ---
horizontal-rule lines, which push the count past 2 on perfectly valid files.
Why: The UX stress-test report's suggested fix ("count ^---$; if ≠ 2,
exit 1") would have rejected any extracted article/PDF containing a thematic
break. The actual bug being fixed was a missing closing delimiter (count < 2)
yielding the empty-string SHA (e3b0c442…, exit 0) → /ctx-compile stamps a
placeholder hash and idempotence skips the file forever (silent data loss).
When to apply: Any change to the hashing/idempotence core. Also: the fix is
a pre-check guard only — do NOT alter the hashing awk itself, or the
already-committed ingested_hash values change and idempotence breaks. A
legitimately-empty body (well-formed frontmatter, no content) should still pass
and honestly hash to e3b0…; that is correct, not the bug. Pinned by
scripts/verify-body-hash.sh (R6 in smoke-all.sh).
Rule: When you add an interactive [y/N] confirmation to a destructive shell
command, grep for every non-interactive caller and give them the --yes
escape hatch. Headless callers (oracles, claude -p, CI) read EOF on the prompt
and abort — which is the right safety default but silently fails automation.
Why: Adding the prune --apply confirm gate to scripts/registry.sh
(report finding #5) immediately broke scripts/verify-multi-wiki.sh's M2 check,
which calls prune --apply >/dev/null 2>&1 and now hit EOF → abort → oracle red.
Fixed by passing prune --apply --yes in the oracle.
When to apply: Any new confirmation/--force/--yes gate on a script that
other scripts or slash-command prompts invoke. Mirror the existing convention:
--yes/-y/--force skips the prompt; TTY/interactive gets [y/N]; headless
without --yes aborts unchanged (see wipe-meta-wiki.sh).
Rule: When building an eval with an LLM judge, prove it on two planted fixtures — one faithful, one unfaithful — before trusting any number. The judge can be wrong both ways: (a) too lenient if you parse its output badly, (b) too harsh if you feed it the wrong evidence.
Why (both bugs hit while building eval-citation-faithfulness.sh):
- False FAITHFUL — the grader echoed the instruction "FAITHFUL or
UNFAITHFUL?", so
grep ... | head -1grabbed the echoed word, not the verdict. Fix: demand aVERDICT=token, parse the LAST one, default-closed. Also: never pass multi-line evidence through bashreadvia hand-rolled\nencoding — bash double-quotes mangle it; base64 the fields instead. - False UNFAITHFUL — the evidence window for timestamp anchors was 8 lines from the heading, but the cited quote sat at line ~58 of a ~10-line section, just past the window. The judge correctly said "evidence doesn't support the claim" — because the evidence I extracted didn't contain it. Fix: extract the whole section (heading → next heading), not a fixed window.
When to apply: any judge-based eval. The judge is only as good as the
evidence you feed it and the parsing of its answer. A green that can't catch a
planted failure is theater; a red that flags a planted-faithful case is noise.
Both waste trust. The fixture lives at tests/eval/faithfulness-fixture/.
Rule: When authoring a command that writes (source: raw/…#anchor)
citations (e.g. /wiki-learn), the citation must (a) include the .md
extension on the raw filename — raw/session-2026-06-01.md#turn-7, not
raw/session-2026-06-01#turn-7 — and (b) point at an anchor the floor can
resolve: a heading slug, a #L<n> line range, or a #M:SS timestamp. If
the captured raw is a free-form list, give each cited item a real heading
(## turn-7 → slug turn-7) so the anchor resolves. Prove it before shipping:
python3 scripts/citation-audit.py <wiki> --raw <raw> must report 0 broken.
Why: scripts/citation-audit.py C1 checks raw/<file> exists literally
(no extension = "file missing"), and C2 checks the anchor resolves (not just
that the file exists). /wiki-learn's first draft cited
raw/session-<date>#turn-N with the fact as a prose line — so C1 failed (no
.md) and, once fixed, C2 would have failed too (no heading named turn-N).
Both are invisible to the --no-build smoke suite, which only audits the
committed fixtures, never the new command's output. The bug surfaced only by
running the loop's actual output through the audit.
When to apply: authoring or editing any command/template that emits
citations. Don't trust "smoke green" — that proves old fixtures resolve, not
that your new command's citations do. Run citation-audit.py on a scratch wiki
built from the new command's output. R9 enforces the floor repo-wide.
retr_parse_questions (scripts/lib/eval-common.sh) matches known fields and
appends everything else to the question text. cite-file-matches: was never
in that list — eval-corpus.sh reads it separately for grading and its comment
even says so — so for every A1 question the catch-all folded it into the prompt.
54 of the 66 gold questions were asking the model the question plus
cite-file-matches: ^2020-09-16-, which names the date prefix of the correct
source file. The A1 leg that field feeds is precisely "did you cite a source of
the right date." The check was handing over its own answer.
Why: the field was correctly excluded from the parser's outputs and therefore assumed to be excluded from its inputs. A catch-all makes those two different things. Every A1 figure measured before this fix is suspect, including the "16/16 with correct-episode citation" headline.
When to apply: adding ANY field to a question/fixture format. Grep the parser for a catch-all before assuming an unknown field is ignored, and assert it: parse the file and check no metadata string survives in the question text.
A2 cross-temporal questions scored 3/6, and a mechanism story was built on top of it ("median 0 reads → answering from synthesis artifacts"). A router forcing dated reads was designed, built, verified (T1-T6), and measured. It fires 30/30 and runs its gate 30/30 — and both arms score identically: 26/30 answer, 30/30 date. The baseline prompt, with no router at all, was already citing the right dated episodes every time.
The 8-question pilot's 3/7 → 6/6 read as confirmation. It was noise moving.
Why: a plausible mechanism makes a small sample feel like evidence. The premise went unchallenged until a baseline arm existed to challenge it — by which point the fix was already built.
When to apply: before building anything to close a measured gap, run the baseline arm on a properly-sized sample. The order is baseline, then diagnosis, then fix — not fix, then measure. Cheap rule: if n < 20, it is a hypothesis.
2026-08-05 — a differential gate must resolve WHICH implementation consumes each fixture, and refuse to guess
Rule: When a gate proves a property by running the real code over a real fixture, and more than one implementation could consume that fixture, the gate must resolve the binding — not assume one, and not check every implementation. Resolve empirically (which ones can actually read it), use it when exactly one can, and require the fixture to DECLARE its consumer when several can. Never default silently; a fixture that resolves to nothing is a failure, not a pass.
Why: gate-eval-prompt-purity.sh first ran retr_parse_questions over all
seven question fixtures and reported five leaks in multi-hop-questions.md —
all false. That file is consumed by eval_parse_questions, which handles
baseline-absent: correctly. The next attempt ("no fixture may leak under
EITHER parser") looked conservative and was simply wrong: a retr-format fixture
reports nine leaks under the multi-hop parser, for nine fields that parser
legitimately never sees. Both wrong answers came from skipping the binding
question. The repo makes it unavoidable — two of four grain-corpus question
files are passed to eval-corpus.sh by hand via --questions, so no call-graph
can discover them, and a hand-written map rots the day someone adds a fixture.
When to apply: any gate whose detection mechanism is "run the real thing and diff", where the repo has more than one "real thing". Also the tell: if a new gate's first run produces findings that are individually plausible but all in one file, suspect the binding before believing the findings.
2026-08-05 — a CLEAN fixture can pass for the wrong reason; only a mutation proves it is load-bearing
Rule: A gate's clean fixture passing tells you nothing until you have broken it and watched it fail. Write the fixture so the ONLY thing making it pass is the property under test — in particular, prose and comments inside a fixture must not restate the thing the gate greps for.
Why: gate-reachable.sh computes reachability by looking for a script's
basename in the text of already-reached files. Its clean fixture's own header
comments said "Reaches verify-alpha.sh directly; verify-beta.sh is reached
transitively through it." Two of five mutations — severing the transitive call,
then severing the root call — both came back GREEN, because the explanatory
comments still contained the basenames. The fixture was passing on its
documentation, not its behaviour. The fix was two-sided: strip whole-line
comments before matching (a script named only in a comment is described, not
run), and rewrite the fixtures so no comment names a basename. Mutation M5 now
covers exactly this case — the call commented out, the mention surviving.
When to apply: every gate, at the moment the clean fixture first goes green — that is the least trustworthy green in the process. Deterministic-gates §5 item 1 exists for this; the two mutations that fail are worth more than the three that pass.