Skip to content

Latest commit

 

History

History
208 lines (165 loc) · 11.9 KB

File metadata and controls

208 lines (165 loc) · 11.9 KB

LEARNINGS — context-compiler-bootstrap

Project-scoped lessons. Lead with the rule, then Why: and When to apply:.


2026-06-01 — shell scripts are macOS-authored but run in Ubuntu CI: force byte locale, don't trust local green

Rule: Any awk/sed that manipulates non-ASCII bytes (em-dash, etc.) must run under LC_ALL=C so it stays byte-oriented. A passing run on macOS's BSD awk proves nothing about Ubuntu CI, where gawk/mawk in a UTF-8 locale treat the same bytes as multi-byte codepoints. Also: a --no-build/offline CI path can only verify committed artifacts — anything gitignored as a build product (e.g. tests/smoke/output/last-answer.md) must be tracked as a baseline or the check fails on a fresh checkout.

Why: wiki-lint-typed-relations.sh built the em-dash via sprintf("%c%c%c",226,128,148) and matched it with index(). On BSD awk that yields the 3 raw bytes (R7 green locally); on Ubuntu CI's UTF-8 awk each %c became a 2-byte char, so index() never matched and verbs were miscounted (R7 red). Separately, C4 reads the captured query answer, which was gitignored — CI had no file to assert on. Both were invisible until the PR's Ubuntu run; local macOS green was a false signal.

When to apply: Before claiming "CI-verified" — actually push and watch the Ubuntu run, never infer it from a macOS pass. Any new shell check touching non-ASCII, or any CI assertion on a generated artifact.


2026-05-30 — body-hash.sh frontmatter validation must allow body --- rules

Rule: When guarding scripts/body-hash.sh (and verify-extract.sh) against malformed frontmatter, validate as line 1 is --- AND there are ≥ 2 ^---$ delimiters — never exactly 2. Markdown bodies legitimately contain --- horizontal-rule lines, which push the count past 2 on perfectly valid files.

Why: The UX stress-test report's suggested fix ("count ^---$; if ≠ 2, exit 1") would have rejected any extracted article/PDF containing a thematic break. The actual bug being fixed was a missing closing delimiter (count < 2) yielding the empty-string SHA (e3b0c442…, exit 0) → /ctx-compile stamps a placeholder hash and idempotence skips the file forever (silent data loss).

When to apply: Any change to the hashing/idempotence core. Also: the fix is a pre-check guard only — do NOT alter the hashing awk itself, or the already-committed ingested_hash values change and idempotence breaks. A legitimately-empty body (well-formed frontmatter, no content) should still pass and honestly hash to e3b0…; that is correct, not the bug. Pinned by scripts/verify-body-hash.sh (R6 in smoke-all.sh).


2026-05-30 — adding a confirm-gate to a shared script breaks its automated callers

Rule: When you add an interactive [y/N] confirmation to a destructive shell command, grep for every non-interactive caller and give them the --yes escape hatch. Headless callers (oracles, claude -p, CI) read EOF on the prompt and abort — which is the right safety default but silently fails automation.

Why: Adding the prune --apply confirm gate to scripts/registry.sh (report finding #5) immediately broke scripts/verify-multi-wiki.sh's M2 check, which calls prune --apply >/dev/null 2>&1 and now hit EOF → abort → oracle red. Fixed by passing prune --apply --yes in the oracle.

When to apply: Any new confirmation/--force/--yes gate on a script that other scripts or slash-command prompts invoke. Mirror the existing convention: --yes/-y/--force skips the prompt; TTY/interactive gets [y/N]; headless without --yes aborts unchanged (see wipe-meta-wiki.sh).


2026-05-30 — an LLM-judge eval lies in BOTH directions; validate with planted cases

Rule: When building an eval with an LLM judge, prove it on two planted fixtures — one faithful, one unfaithful — before trusting any number. The judge can be wrong both ways: (a) too lenient if you parse its output badly, (b) too harsh if you feed it the wrong evidence.

Why (both bugs hit while building eval-citation-faithfulness.sh):

  1. False FAITHFUL — the grader echoed the instruction "FAITHFUL or UNFAITHFUL?", so grep ... | head -1 grabbed the echoed word, not the verdict. Fix: demand a VERDICT= token, parse the LAST one, default-closed. Also: never pass multi-line evidence through bash read via hand-rolled \n encoding — bash double-quotes mangle it; base64 the fields instead.
  2. False UNFAITHFUL — the evidence window for timestamp anchors was 8 lines from the heading, but the cited quote sat at line ~58 of a ~10-line section, just past the window. The judge correctly said "evidence doesn't support the claim" — because the evidence I extracted didn't contain it. Fix: extract the whole section (heading → next heading), not a fixed window.

When to apply: any judge-based eval. The judge is only as good as the evidence you feed it and the parsing of its answer. A green that can't catch a planted failure is theater; a red that flags a planted-faithful case is noise. Both waste trust. The fixture lives at tests/eval/faithfulness-fixture/.


2026-06-01 — any new command that emits citations must produce resolvable anchors (with .md)

Rule: When authoring a command that writes (source: raw/…#anchor) citations (e.g. /wiki-learn), the citation must (a) include the .md extension on the raw filename — raw/session-2026-06-01.md#turn-7, not raw/session-2026-06-01#turn-7 — and (b) point at an anchor the floor can resolve: a heading slug, a #L<n> line range, or a #M:SS timestamp. If the captured raw is a free-form list, give each cited item a real heading (## turn-7 → slug turn-7) so the anchor resolves. Prove it before shipping: python3 scripts/citation-audit.py <wiki> --raw <raw> must report 0 broken.

Why: scripts/citation-audit.py C1 checks raw/<file> exists literally (no extension = "file missing"), and C2 checks the anchor resolves (not just that the file exists). /wiki-learn's first draft cited raw/session-<date>#turn-N with the fact as a prose line — so C1 failed (no .md) and, once fixed, C2 would have failed too (no heading named turn-N). Both are invisible to the --no-build smoke suite, which only audits the committed fixtures, never the new command's output. The bug surfaced only by running the loop's actual output through the audit.

When to apply: authoring or editing any command/template that emits citations. Don't trust "smoke green" — that proves old fixtures resolve, not that your new command's citations do. Run citation-audit.py on a scratch wiki built from the new command's output. R9 enforces the floor repo-wide.

2026-08-03 — An eval field that isn't in the parser's field list ends up IN the question

retr_parse_questions (scripts/lib/eval-common.sh) matches known fields and appends everything else to the question text. cite-file-matches: was never in that list — eval-corpus.sh reads it separately for grading and its comment even says so — so for every A1 question the catch-all folded it into the prompt. 54 of the 66 gold questions were asking the model the question plus cite-file-matches: ^2020-09-16-, which names the date prefix of the correct source file. The A1 leg that field feeds is precisely "did you cite a source of the right date." The check was handing over its own answer.

Why: the field was correctly excluded from the parser's outputs and therefore assumed to be excluded from its inputs. A catch-all makes those two different things. Every A1 figure measured before this fix is suspect, including the "16/16 with correct-episode citation" headline.

When to apply: adding ANY field to a question/fixture format. Grep the parser for a catch-all before assuming an unknown field is ignored, and assert it: parse the file and check no metadata string survives in the question text.

2026-08-03 — n=6 is not a diagnosis; get the baseline arm first

A2 cross-temporal questions scored 3/6, and a mechanism story was built on top of it ("median 0 reads → answering from synthesis artifacts"). A router forcing dated reads was designed, built, verified (T1-T6), and measured. It fires 30/30 and runs its gate 30/30 — and both arms score identically: 26/30 answer, 30/30 date. The baseline prompt, with no router at all, was already citing the right dated episodes every time.

The 8-question pilot's 3/7 → 6/6 read as confirmation. It was noise moving.

Why: a plausible mechanism makes a small sample feel like evidence. The premise went unchallenged until a baseline arm existed to challenge it — by which point the fix was already built.

When to apply: before building anything to close a measured gap, run the baseline arm on a properly-sized sample. The order is baseline, then diagnosis, then fix — not fix, then measure. Cheap rule: if n < 20, it is a hypothesis.

2026-08-05 — a differential gate must resolve WHICH implementation consumes each fixture, and refuse to guess

Rule: When a gate proves a property by running the real code over a real fixture, and more than one implementation could consume that fixture, the gate must resolve the binding — not assume one, and not check every implementation. Resolve empirically (which ones can actually read it), use it when exactly one can, and require the fixture to DECLARE its consumer when several can. Never default silently; a fixture that resolves to nothing is a failure, not a pass.

Why: gate-eval-prompt-purity.sh first ran retr_parse_questions over all seven question fixtures and reported five leaks in multi-hop-questions.md — all false. That file is consumed by eval_parse_questions, which handles baseline-absent: correctly. The next attempt ("no fixture may leak under EITHER parser") looked conservative and was simply wrong: a retr-format fixture reports nine leaks under the multi-hop parser, for nine fields that parser legitimately never sees. Both wrong answers came from skipping the binding question. The repo makes it unavoidable — two of four grain-corpus question files are passed to eval-corpus.sh by hand via --questions, so no call-graph can discover them, and a hand-written map rots the day someone adds a fixture.

When to apply: any gate whose detection mechanism is "run the real thing and diff", where the repo has more than one "real thing". Also the tell: if a new gate's first run produces findings that are individually plausible but all in one file, suspect the binding before believing the findings.

2026-08-05 — a CLEAN fixture can pass for the wrong reason; only a mutation proves it is load-bearing

Rule: A gate's clean fixture passing tells you nothing until you have broken it and watched it fail. Write the fixture so the ONLY thing making it pass is the property under test — in particular, prose and comments inside a fixture must not restate the thing the gate greps for.

Why: gate-reachable.sh computes reachability by looking for a script's basename in the text of already-reached files. Its clean fixture's own header comments said "Reaches verify-alpha.sh directly; verify-beta.sh is reached transitively through it." Two of five mutations — severing the transitive call, then severing the root call — both came back GREEN, because the explanatory comments still contained the basenames. The fixture was passing on its documentation, not its behaviour. The fix was two-sided: strip whole-line comments before matching (a script named only in a comment is described, not run), and rewrite the fixtures so no comment names a basename. Mutation M5 now covers exactly this case — the call commented out, the mention surviving.

When to apply: every gate, at the moment the clean fixture first goes green — that is the least trustworthy green in the process. Deterministic-gates §5 item 1 exists for this; the two mutations that fail are worth more than the three that pass.