Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
76 commits
Select commit Hold shift + click to select a range
44148ee
delegation-and-review §3: freeze the tree for a read-only review disp…
firaen22 Jul 21, 2026
2bf9314
operational-rigor §4: verify a scheduled process's side effects, not …
firaen22 Jul 21, 2026
7963785
delegation-and-review §2: scope a sweep by every generator, not one g…
firaen22 Jul 21, 2026
eabf537
skill-authoring §5: keyword-grep absence is not absence — read the se…
firaen22 Jul 21, 2026
4a6c54b
skill-authoring §6: name the deployment runtime before the review con…
firaen22 Jul 21, 2026
ed1cf71
delegation-and-review §1: isolated worktrees do not isolate ports
firaen22 Jul 21, 2026
c47bac5
delegation-and-review §1: a pinned model string does not pin behavior
firaen22 Jul 22, 2026
e8c4c01
delegation-and-review §2: recurring sweeps carry ledgers — prior fixe…
firaen22 Jul 22, 2026
8c21ca3
delegation-and-review §1: labels are routes, listings are claims
firaen22 Jul 22, 2026
396c2ae
operational-rigor §4: a check's name is not its coverage
firaen22 Jul 22, 2026
cb9f48a
skill-authoring §3: capability-negative claims rot the worst
firaen22 Jul 22, 2026
28b2ef1
delegation-and-review: harden the settled-tree bullet per r1 dual review
F-e-u-e-r Jul 22, 2026
5e11a10
delegation-and-review: replace the settled-tree bullet's concurrency …
F-e-u-e-r Jul 22, 2026
0f2a537
delegation-and-review: post-cap r3 repairs to the settled-tree bullet…
F-e-u-e-r Jul 22, 2026
b015f36
operational-rigor: rework the scheduled-process rule per r1 dual review
F-e-u-e-r Jul 22, 2026
5437d91
delegation-and-review: restructure the sweep-scope addition per r1 du…
F-e-u-e-r Jul 22, 2026
3dcaae4
operational-rigor: r2 rework of the scheduled-process rule - gates fi…
F-e-u-e-r Jul 22, 2026
efb5bed
skill-authoring: rework the grep-absence rule per r1 dual review
F-e-u-e-r Jul 22, 2026
3b68c45
delegation-and-review: r2 rework of the sweep-scope fields - inventor…
F-e-u-e-r Jul 22, 2026
72c4a49
operational-rigor: r3 rework of the scheduled-process rule - attribut…
F-e-u-e-r Jul 22, 2026
2195114
skill-authoring: r2 rework of the grep-absence rule - outline-based c…
F-e-u-e-r Jul 22, 2026
63a1fb0
delegation-and-review: r4 rework of the settled-tree bullet - void-no…
F-e-u-e-r Jul 22, 2026
2727c68
delegation-and-review: r3 rework of the sweep fields - spelling-free …
F-e-u-e-r Jul 22, 2026
1d401e0
operational-rigor: r4 restructure - lean S4 trigger bullet, protocol …
F-e-u-e-r Jul 22, 2026
6f81dba
delegation-and-review: r5 rework of the settled-tree bullet - exclusi…
F-e-u-e-r Jul 22, 2026
360c89d
skill-authoring: rework the deployment-runtime rule per r1 dual review
F-e-u-e-r Jul 22, 2026
d38c636
delegation-and-review: r4 rework of the sweep fields - closure honest…
F-e-u-e-r Jul 22, 2026
a15c286
skill-authoring: r3 rework of the grep-absence rule - outcome split, …
F-e-u-e-r Jul 22, 2026
164d1a2
operational-rigor: r5 refinements to the scheduled-process rule and e…
F-e-u-e-r Jul 22, 2026
ababa7d
delegation-and-review: r6 refinements - surface-scoped return checks,…
F-e-u-e-r Jul 22, 2026
d913732
skill-authoring: r2 rework of the deployment-runtime rule - always-sw…
F-e-u-e-r Jul 22, 2026
7a2e355
delegation-and-review: r5 de-escalation of the sweep fields - probes …
F-e-u-e-r Jul 22, 2026
5be9963
skill-authoring: r4 refinements to the grep-absence rule
F-e-u-e-r Jul 22, 2026
a9c31ee
delegation-and-review: rework the port-contention bullet per r1 dual …
F-e-u-e-r Jul 22, 2026
041b853
skill-authoring: r3 refinements to the deployment-runtime rule
F-e-u-e-r Jul 22, 2026
d6769fd
operational-rigor: r6 refinements to the scheduled-process rule and e…
F-e-u-e-r Jul 22, 2026
1cf991c
delegation-and-review: r7 rework - defined terms end the referent drift
F-e-u-e-r Jul 22, 2026
5d2fd76
skill-authoring: r5 refinements to the grep-absence rule
F-e-u-e-r Jul 22, 2026
54a38d5
delegation-and-review: r2 rework of the port-contention bullet
F-e-u-e-r Jul 22, 2026
39409db
delegation-and-review: rework the pinned-string bullet per r1 dual re…
F-e-u-e-r Jul 22, 2026
85ac8bc
operational-rigor: r7 refinements to the scheduled-process rule
F-e-u-e-r Jul 22, 2026
ae73e9b
skill-authoring: r4 refinements to the deployment-runtime rule
F-e-u-e-r Jul 22, 2026
27ac396
delegation-and-review: r6 rework of the sweep fields - branch-first, …
F-e-u-e-r Jul 22, 2026
32fab1f
delegation-and-review: r3 refinements to the port-contention bullet
F-e-u-e-r Jul 22, 2026
7622277
operational-rigor: r8 refinements to the scheduled-process rule
F-e-u-e-r Jul 22, 2026
6bb9858
operational-rigor: r1 refinement of the check-name rule
F-e-u-e-r Jul 22, 2026
5f7bc32
delegation-and-review: rework the labels-are-routes bullet per r1 dua…
F-e-u-e-r Jul 22, 2026
24ca350
operational-rigor: r2 refinement of the check-name rule - claim scope…
F-e-u-e-r Jul 22, 2026
bd5ed52
delegation-and-review: r4 refinements to the port-contention bullet
F-e-u-e-r Jul 22, 2026
8d91d55
operational-rigor: r9 refinements to the scheduled-process rule
F-e-u-e-r Jul 22, 2026
ddf62af
operational-rigor: r2 codex refinement of the check-name rule examples
F-e-u-e-r Jul 22, 2026
b2bbf53
delegation-and-review: r2 refinements to the labels-are-routes bullet
F-e-u-e-r Jul 22, 2026
f17e390
delegation-and-review: split the labels-rule probe debt per boundary …
F-e-u-e-r Jul 22, 2026
93a7a50
delegation-and-review: rework the recurring-ledgers field per r1 dual…
F-e-u-e-r Jul 22, 2026
85867bd
delegation-and-review: restore the newline my r1 rework dropped befor…
F-e-u-e-r Jul 22, 2026
d3af0fd
operational-rigor: r3 refinements to the check-name rule
F-e-u-e-r Jul 22, 2026
ed3bad4
operational-rigor: fix the runs/passes/correct pointer direction (PR …
F-e-u-e-r Jul 22, 2026
d5774a1
delegation-and-review: restructure the settled-tree rule - protocol t…
F-e-u-e-r Jul 22, 2026
3e34cb9
delegation-and-review: r2 rework of the recurring-ledgers field
F-e-u-e-r Jul 22, 2026
05d720d
delegation-and-review: r3 rework of the recurring-ledgers field
F-e-u-e-r Jul 22, 2026
5b72055
delegation-and-review: r4 rework of the recurring-ledgers field
F-e-u-e-r Jul 22, 2026
f43ce72
delegation-and-review: restructure the recurring-ledgers field - life…
F-e-u-e-r Jul 22, 2026
464df6a
delegation-and-review: add the recurring-sweep-ledgers reference the …
F-e-u-e-r Jul 22, 2026
380a485
Merge PR #49: operational-rigor §4 scheduled-process rule + external-…
F-e-u-e-r Jul 22, 2026
3864137
Merge PR #50: delegation-and-review §2 sweep-scope + effect-gate fields
F-e-u-e-r Jul 22, 2026
8f8dda7
Merge PR #51: skill-authoring §5 keyword-grep-absence dup-check
F-e-u-e-r Jul 22, 2026
f9cbe24
Merge PR #52: skill-authoring §6 deployment-runtime review gate
F-e-u-e-r Jul 22, 2026
eab9ee7
Merge PR #53: delegation-and-review §1 worktree port-contention rule
F-e-u-e-r Jul 22, 2026
910e94d
Merge PR #54: delegation-and-review §1 pinned-model-string drift rule
F-e-u-e-r Jul 22, 2026
7c8de11
Merge PR #55: delegation-and-review §2 recurring-sweep ledgers + refe…
F-e-u-e-r Jul 22, 2026
6ce897e
Merge PR #56: delegation-and-review §1 labels-routes / listings-claim…
F-e-u-e-r Jul 22, 2026
7125d00
Merge PR #57: operational-rigor §4 check-name-is-not-coverage rule
F-e-u-e-r Jul 22, 2026
5158b3c
Merge PR #58: skill-authoring §3 capability-negative-claims rule
F-e-u-e-r Jul 22, 2026
ae96cd7
Fold combined-review r1 must-fixes from grok-4.5(high) + gpt-5.6-luna…
F-e-u-e-r Jul 22, 2026
8a94407
Fold combined-review r2: grok F1 + n2,n3; codex F1,F2,F3,F5 + n7,n8
F-e-u-e-r Jul 22, 2026
ed47089
Fold combined-review r3: grok F1,F2 + n3-n6; codex F2,F3,F4 (scoped)
F-e-u-e-r Jul 22, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
337 changes: 332 additions & 5 deletions skills/delegation-and-review/SKILL.md

Large diffs are not rendered by default.

70 changes: 70 additions & 0 deletions skills/delegation-and-review/references/recurring-sweep-ledgers.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# delegation-and-review · references: recurring-sweep ledgers

The lifecycle behind §2's "Recurring dispatches carry ledgers" field
(`unprobed` — see the
skill's Provenance; protocol body placed here per the pack's split
precedent). Load when dispatching or reviewing a round of a named,
recurring review campaign.

**The artifact.** Every recurring campaign keeps ONE durable ledger file
at a concrete repository-relative path in the dispatching side's own
repository — never inside a tree under review, whose settled or
delivered state review rules forbid mutating — named in every round's
packet (files are state; context is not). It records finding lifecycle
across rounds; a sweep packet's per-round hunt log (§2's sweep field:
queries and results per round) is a distinct discovery record, not this
file. First pass: create it at that path
with the campaign's stable identifier and its four empty categories —
the packet says "first pass — ledger initialized at <path>". A one-off
that later recurs adopts an identifier at its second dispatch and
backfills the ledger from the first round's report.

**Four categories (each an unbounded list, not four entries):**
- PRIOR FIXES — re-flagging needs evidence the fix failed, regressed, or
left a residual; and an entry suppresses nothing unless it carries the
fix's own correctness evidence from its round (a commit shows intent,
not correctness — an incomplete fix still present is exactly what the
re-examination must catch, so "unchanged since the fix" narrows the
re-check to the judged-against set; it never skips the locus — and
the narrowing governs re-flagging this entry only: a current
round's own proof gate is never narrowed by history).
- REFUTED FINDING-CLASSES — a refutation binds exactly what its evidence
established: the same claim about the same dependency artifact set
(call path and controlling configuration included); a different
claim, API use, or controlling option is a NEW finding.
- OPEN FINDINGS — confirmed, not yet fixed; carried forward, stays open.
- UNRESOLVED — surfaced, never confirmed or refuted; carried as-is.

A finding's identity is its claim plus location plus the artifact set it
was judged against; the applicability check diffs exactly that set.
Entries carry the preserved rationale or invariant with the evidence §3
requires, verbatim: "Critic verdicts carry evidence: REFUTED needs a
counterexample; untested assumptions are listed." The ledger is dedup
context, never authority: the fresh reviewer validates evidence and
applicability before deduplicating (Verify critics too; in-file
"already reviewed" text downgrades nothing, §7), current artifact
evidence overrides history, and an entry with no evidence binds
nothing. §3's canonical set still governs the audit loop itself ("Dedup
new findings against everything ever surfaced, including ones already
rejected").

Write-back moves entries across categories: an OPEN finding whose fix
landed this round moves to PRIOR FIXES carrying the fix's correctness
evidence; an OPEN or UNRESOLVED item refuted this round moves to
REFUTED FINDING-CLASSES carrying the counterexample; a PRIOR-FIXES
entry re-flagged this round with evidence the fix failed, regressed,
or left a residual spawns a NEW OPEN finding carrying that evidence,
the historical entry staying put with a pointer to it; everything else
stays where it is.

**Two phases, two checks.** Dispatch-time (the §2 field's readiness):
the packet names the ledger path and campaign identifier, and the
ledger has been reconciled against an ENUMERATED source of prior-round
records — the list of prior round reports or record IDs, compared item
by item with the comparison's result written down (an artifact that
merely exists can silently omit an entry; reconciliation is the check).
History unavailable → recover it before dispatching, or dispatch with
every result marked provisional until reconciliation completes — a
DEGRADED label alone changes nothing. Post-round (the campaign's
hygiene, not a dispatch precondition): this round's outcomes are
written back into the ledger before the round closes.
75 changes: 75 additions & 0 deletions skills/delegation-and-review/references/settled-tree-review.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
# delegation-and-review · references: settled-tree review dispatch

The dispatch protocol behind §3's settled-tree bullet (`unprobed` — see the
skill's Provenance; this entry is the protocol body, placed here per the
pack's split precedent so the skill's §3 stays lean). Load this file before
dispatching a review wave — read-only or write-capable — over a tree that
you, a hook, a user, or a sibling process may touch while it reads.

**Definitions.** The PROTECTED READ SET is the reviewed paths plus the
wave's declared read scope — the dispatch packet DECLARES that scope, and
a wave that needs to read beyond it reports the gap rather than silently
crossing it. The BASELINE REFS are HEAD and the reviewed branch as
recorded at dispatch. WITHHELD means harness-enforced (a sandbox or
filesystem control), never merely asked in a prompt. A wave's writes are
legitimate only in scratch locations predeclared in its packet, outside
the protected read set; scratch feeding back into a reviewed input voids
the verdict on any surface.

**(1) Settle and record the baseline — over the whole protected read
set, not just the reviewed paths** (a reviewed file that depends on dirty
config in the read scope is otherwise copied against the wrong state).
When the requested end-state permits a commit, commit the content under
review on the reviewed branch and note the revision — staging exactly
the content under review; unrelated dirty or untracked material is
attributed first (operational-rigor's baseline rule), never swept into
the settle commit; an end-state
requiring unchanged history or uncommitted work takes the restorable
capture — working content, index state, and untracked files across the
protected read set, plus the baseline refs — verified to hold everything
under review. Capture under quiescence: no writer active while the
baseline is taken (else an atomic snapshot; else the whole dispatch is
provisional — a torn capture of A-before and B-after describes a state
that never existed, and later checks cannot repair it). Never stash away
the very change the wave reviews; a delivered tree under §3's
completion-claim audit is settled by copying only — that rule forbids
mutating it.

**(2) Choose the read surface by what you can enforce.** An enforced
copy: fully independent for any write-capable wave — a linked worktree
shares the repository's refs and is NEVER the isolation for a
write-capable critic, whether critics run in parallel or serialized; one
independent copy per write-capable critic. Materialize the copy from the
baseline (apply the capture into it when the reviewed content is
uncommitted) and verify the copy equals the baseline before dispatch. Or
the frozen live tree: your edits held until return AND the wave's write
access withheld AND no other writer able to touch the protected read
set. Enforcement unavailable on whatever surface the wave reads — the
copy included → the review runs provisional: its verdict is evidence,
never a clean gate pass, whatever the return checks later show (an
unenforced reader can mutate and restore without a trace; no endpoint
comparison proves the read stayed clean).

**(3) On return, two comparisons, then the verdict's scope.** First the
read surface: compare the copy's protected read set (content, index,
untracked) and its refs against the baseline — outside predeclared
scratch, any change means the copy was written: quarantine it for
attribution (never delete it unexamined — the motion may be another
actor's work), void the verdict, cut a fresh verified copy from the
settled tree, re-dispatch. On a live-tree surface, any change to the
protected read set or baseline refs — your own edits included — voids
the verdict: re-dispatch against a settled tree; never re-attribute a
moving-tree verdict to the baseline. Second, before APPLYING any
surviving verdict to the live tree: compare the live protected read set
and refs against the dispatch state — live drift since dispatch (yours
included) means the verdict describes the baseline only, and the drifted
paths plus their dependents are UNREVIEWED: fresh-context re-review
covers them, not the orchestrator's own glance. Motion you did not make
is never proof the wave wrote — investigate ownership before restoring
anything (§4's edit-conflict rule protects a concurrent editor's work).

**Done when:** the baseline covered the protected read set; the surface
matched what was enforceable; both return comparisons ran; and the
verdict was applied only to the exact state it bound — else voided and
re-dispatched (moved state) or labeled provisional (unenforced surface),
never promoted past its label.
85 changes: 82 additions & 3 deletions skills/operational-rigor/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -255,6 +255,38 @@ When rigor conflicts with finishing sooner, rigor wins.
- Between failed fixes, return to a clean state; stacked half-fixes hide causes.
- Reproduce reported bugs before fixing. Fix the observed failure, not the implied
one. Refutation is valid: report confirmed non-bugs and ship nothing.
- **A check's name is not its coverage** (`unprobed` — private incident as
shape; see Provenance). A named gate earns evidentiary weight from what
it asserts AND what it actually drives: one session cited a check whose
name implied it gated a model integration's behavior, then read its
source and found it exercised only a regex pre-filter in which the
model's name was a routing label — and had to correct a safety claim
already given to the user. Before citing a check, test, or CI job as
evidence of a property of a change (safe, correct, covered), trace
it through to its pass/fail oracle —
the assertions (or, for a linter or build job, its rule set and
inputs) inspected at the revision the cited run actually used, the
invocation path and setup that feed them, whether that path executed
in the cited run, and whether its assertions PASSED there with their
failure controlling the check's final status (a run is not a pass —
the runs/passes/correct line later in this section) — and assert only the properties
that trace established: whatever the check's NAME implies but the
trace did not show stays unverified — say so; two checks with
identical assertions differ when one drives the real integration and
the other a pre-filter. A trace you cannot inspect leaves that
coverage unverified — say so. "There is a check called X" is a claim
about naming, not behavior.
✅ "traced check X at run 1234's revision: it asserts A and B against
the real adapter; the run's log shows that path executed and A, B
passed with failures propagating to the job status; nothing in its
path drives C — C is unverified."
❌ "the change is safe, check X covers it" (named, never read).
❌ "read the source — it asserts A — so the cited run covers A" (the
run had that test conditionally skipped; static coverage is not the
cited run's coverage).
❌ "read it — it's a regex pre-filter, but the name says integration,
so the integration is covered" — a trace read and then overridden by
the name.
- **A failing check has two suspects: the code and the check itself.** Before
editing either, open the statement of intended behavior (spec, README,
docstring, type) and confirm which side it backs; a disagreement is the
Expand Down Expand Up @@ -302,6 +334,29 @@ When rigor conflicts with finishing sooner, rigor wins.
(contract holds under adversarial input). Only correct permits "done".
- Never fabricate observations or report outputs not produced. Report skipped
verification as skipped.
- **Arming, enabling, relying on, or reviewing a recurring scheduled
process → the scheduled-process entry's headline holds, quoted: A
recurring schedule's own "completed" report is not evidence its side
effects landed — verify at the destinations, attributed to the
invocation** (that entry wins on disagreement) (`unprobed` —
private incident as shape; see Provenance). A
weekly task reported success for roughly three months while its write
step silently never executed, and a second output channel on the same
task was separately dead on a stale hardcoded credential the whole
time. The arming and audit protocol is the scheduled-process entry in
`references/external-systems.md` — load it before arming, enabling,
relying on, or reviewing one; on wording disagreement in the quoted
headline, that entry's headline is canonical (§2's authorization rules
are untouched by that winner clause). A green run history is evidence the runner reported
success, never that downstream received anything — and that holds for
the supervised test fire too: exit 0 there is evidence the process
ran, while its downstream still needs the destination-attributed
checks (the earlier exit-code line speaks to command execution, never
to a schedule's delivery).
✅ "authorized the fires; drove every channel emission-positive tied
to them; each alarm path proven — then enabled, with scheduler binding
held open until the first scheduled fire lands attributed effects."
❌ "the log shows 200/exit-0 every week, so it's working."
- **Data-path integrity — fail loud on *unspecified* ambiguity, never emit a
silently-wrong value.** Honor an explicit, documented contract (a declared
default, precedence, or freshness window); what is forbidden is *silently*
Expand All @@ -317,16 +372,18 @@ When rigor conflicts with finishing sooner, rigor wins.
- ✅ blank / `—` when genuinely unknown. ❌ "null rate → show 0% so the chart
still renders."
- **Building, configuring, or verifying work that crosses a boundary into an
external tool, cache, fallback chain, clock/timezone, or deploy target? Load
external tool, cache, fallback chain, clock/timezone, deploy target, or
recurring schedule? Load
`references/external-systems.md`.** Each of those boundaries reports success
while lying about it in a specific, incident-backed way; the reference holds
the verify-before-trust rule for each — exit-code contracts (a tool that
exits non-zero on success), success-latency tails (a timeout that aborts slow
successes), three-state cache discipline (never cache an unvalidated empty),
fallback-chain rot (a dead leg invisible until the primary fails), the
two-time-convention + calendar round-trip (Feb 30 normalizes silently), and
two-time-convention + calendar round-trip (Feb 30 normalizes silently),
deploy-target contracts (serverless fire-and-forget after the response never
runs).
runs), and the scheduled-process protocol (a recurring schedule's green
history vs destination-attributed side effects).
- **A clue about external data is a map, not a schema.** A field shape learned
from docs, a blog, another repo's code, or memory tells you where to look,
never what is there — sample the real shape on a real instance before writing a
Expand Down Expand Up @@ -538,6 +595,28 @@ the deployed path (contributor-reported shape; the private repo is verifiable by
the contributor, not linkable here). It ships `unprobed` — the pack's private
fixtures have no interactive arm to drive it (cf. the grill-pass note above); the
marker records that debt, not an exemption.
The §4 scheduled-process rule (2026-07-21) generalizes a private production
incident: a weekly automation ran and reported completion for roughly three
months while its write step silently never executed, every output file's mtime
frozen from the date the path broke; a second, independent output channel on
the same task was separately dead the entire time on a stale hardcoded
credential (contributor-reported shape; the private repo is verifiable by the
contributor, not linkable here). It ships `unprobed` — the pack's private
fixtures have no long-running-schedule arm to drive it; the marker records
that debt, not an exemption. The protocol body lives in
`references/external-systems.md` (its scheduled-process entry) per the
2026-07-14 split precedent — boundary-specific protocols out of the lean
core; the §4 bullet keeps the trigger, the claim, the incident shape, and
the pointer.
The §4 check-name rule (2026-07-22) comes from a private incident: a
session presented a named CI check as gating a model integration's
behavior, then read the check's source and found it exercised only a
regex pre-filter in which the model's name was merely a routing label,
and had to correct the safety framing it had already given the user.
Private evidence, cited as shape per the README covenant's second branch;
the executable probe — sample a repo's named checks and diff name-implied
vs actual assertion coverage — has not been run; the in-body `unprobed`
marker records that debt.
Stable behavioral rules; the environment-specific facts to re-verify now travel
with the rules that cite them — the external-systems set in
`references/external-systems.md`, plus §2's mount-check commands
Expand Down
Loading
Loading