diff --git a/AGENTS.md b/AGENTS.md index 08c9d74f5..1a1291644 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -8,11 +8,11 @@ and evidence for independently judgeable changes. Its federated integration grap connects evidence; tracker, Git and coding agents own work, history and execution. Native execution needs zero mandatory AgentOps skills. ```text -Accepted intent -> native implementation and checks -> fresh independent judgment -> finish +Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish ``` Use the existing issue or conversation and load a skill only for a concrete -uncertainty or an explicitly selected workflow. Without fresh independent judgment over the exact subject, the experiment remains unproven. +uncertainty or an explicitly selected workflow. Checks and CI are the gate for an ordinary change; a fresh read is for the costly cases named in the charter below. Native goals supply continuity; BD keeps work and handoffs. Neither needs RPI. Creating a goal does not reset the conversation or enforce a resource limit. Delegate focused intent, scope and evidence; retrieve history only when it can change a decision. Assign final review once, preserving required review legs. @@ -159,8 +159,7 @@ outcome and judgment; explicit interim analysis names its cutoff and unknowns. Neither holds code acceptance open; the overall goal still owes both deliverables. Use cheap discriminating checks during edits, required integration checks before -final judgment, and reserve capacity for integration, validation, repair and a -truthful handoff. On a genuine causal stall (unknown cause, recurrence, no +finishing, and reserve capacity for integration, repair and a truthful handoff. On a genuine causal stall (unknown cause, recurrence, no progress or wrong objective), use at most one bounded fresh helper for that incident within authority and real remaining bounds. An unhelpful answer ends the attempt; do not build a helper chain. Known failures need direct repair. @@ -168,17 +167,20 @@ Cancellation, refusal and spent hard time/cost/quota skip help. Retry counts, compaction, helpers and new subjects never renew real limits. Preserve compact recovery state in native handoff only when needed to prevent evidence loss. -Fresh author-distinct final validation is required over the exact subject, -unchanged acceptance and all changed paths. Default to a fresh reviewer from -the author's model family; cross-model review is opt-in and every explicitly -required leg remains required. No fixed ten-minute cap applies. Risk determines -evidence depth, not mandatory specialist or model-family multiplication. -PASS needs distinct identities, attested freshness, nonempty checked scope, -evidence for every criterion and empty `not_checked`. Missing identity, -freshness, subject continuity or acceptance proof means `NOT_PROVEN`; proven -out-of-scope change or failed acceptance means `FAIL`. Repair known findings -within authority and real bounds, then revalidate the changed exact subject. -Persist machine evidence only for a caller request or declared consumer. +Spend validation where a mistake is costly. For an ordinary change the author +runs the checks and CI is the gate; no fresh review is owed. Obtain one fresh +author-distinct read only when the caller asks, a mistake cannot be cheaply +undone after it lands (a published release or instructions users will follow, +a security boundary, destroying data or tracker state, deleting a check that +protects the product), or no deterministic check covers the changed behavior. +One round: the reviewer gets the exact subject and one question and does not +re-run checks. Defects are only what fails accepted behavior or would mislead a +user, break install or the CLI, or remove protection for the product; the rest +are optional notes. Repair defects, confirm each with a check, and finish; a +repair does not start another review. Keep review cost a fraction of the cost +of the work. A requested binding verdict keeps the PASS rule in +[Validate](skills/validate/SKILL.md); report `NOT_PROVEN` with its gaps instead +of chasing it. Persist machine evidence only for a caller request or declared consumer. [Memory](skills/memory/SKILL.md) owns optional find/recall, capture/mining and curation: support, applicability, invalidation and preservation; no blind TTL/deletion. @@ -193,15 +195,14 @@ independent support/disclosure review precede Git import (ADR-0016). The [RPI skill](skills/rpi/SKILL.md) packages this charter when explicitly selected; it is not a prerequisite for native execution or independent review. The [architecture reference](docs/architecture/rpi-traversal.md) owns exact evidence -semantics. Optional outer-goal guidance and the grandfathered fixed-dispatch -reference adapter stay outside the native core. No scheduler or new AO command +semantics. Optional outer-goal guidance stays outside the native core. No scheduler or new AO command is needed for this harness. ## Product boundary AgentOps reads or refines caller-owned intent, implements authorized work and direct repairs, -establishes exact content identity, and obtains fresh independent judgment. It -can persist that judgment as standalone evidence when requested. It owns no aggregate retry controller, +establishes exact content identity, and obtains one fresh judgment where a mistake is +costly. It can persist that judgment as standalone evidence when requested. It owns no aggregate retry controller, budget, queue, work ownership, Git, closure, release, landing, or delivery transition. Consumer repositories keep their own direct-push, PR, CI, merge, rollback, and release policy. @@ -243,5 +244,5 @@ AgentOps work ownership. ## Closeout Map each acceptance criterion to evidence and disclose `checked` and -`not_checked`; any unchecked acceptance means `NOT_PROVEN`. Apply the fresh -judgment and delivery-authority rules above. +`not_checked`. An unchecked item is reported, not a reason to keep validating. +Apply the validation and delivery-authority rules above. diff --git a/CHANGELOG.md b/CHANGELOG.md index 48758aa42..dc7ce686e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,20 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Changed +- Fresh validation is no longer owed on every change. For an ordinary change + the author's checks and CI are the gate. RPI, Validate, Implement, + Orchestrate, Craft Goal and Navigate call for one fresh read only when the + caller asks, a mistake cannot be cheaply undone after it lands (a published + release or instructions users will follow, a security boundary, destroying + data or tracker state, deleting a check that protects the product), or no + deterministic check covers the changed behavior. A review is one round: the + reviewer does not re-run checks, reports as defects only what fails accepted + behavior or would mislead a user, break install or the CLI, or remove + protection for the product, and a repair does not start another review. A + requested binding verdict keeps its PASS rule; `NOT_PROVEN` is reported with + its gaps instead of being chased. +- Craft Goal, Interview and Navigate are marked stable; the README no longer + labels goals experimental. - The Codex plugin now loads `skills/` directly. `.codex-plugin/plugin.json` ships `./skills`, the same tree `ao skills link` and `npx skills` already install, instead of a generated copy. Skill names, descriptions and bodies are @@ -66,6 +80,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - `ao doctor` no longer has the `fm-skills-stale-codex-sync` failure mode, and its fixers can no longer write to `~/.codex/plugins/cache/agentops-marketplace` or `~/.codex/.agentops-codex-install.json`. +- The fixed-dispatch RPI reference adapter, which modelled repeated review + rounds: `skills/rpi/scripts/run_once.py`, its tests, its reference page and + `skills/rpi/references/rpi.feature`, plus `tests/e2e/rpi-phased-domain.sh`, a + no-op tombstone that existed only to keep that feature file's scenario link + resolving. +- The Skill Builder converter's Codex target no longer writes `prompt.md`; + Codex does not read it. - Bundled Flywheel tool skills (`account-rotation`, `agent-mail`, `cass`, `cc-hooks`, `dcg`, `ms`, `ntm`, `rch`, `sbh`, `using-flywheel`) and their generated Codex copies. @@ -102,6 +123,18 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 `~/.pi/agent/skills`, and detect Pi by `~/.pi/agent`. Links an earlier version made in `~/.pi/skills` are no longer swept by default; remove them with `ao skills unlink --dest ~/.pi/skills`. +- Skill Builder's `heal.sh --check` exits 2 with a message when `ao` cannot run. + It used to report a pass in non-strict mode. +- The Codex policy check also fails on `agents/openai.yaml` shapes Codex + silently ignores: a non-object `interface`, `dependencies` that is not a + mapping with a `tools` list, and a boolean spelled other than `true` or + `false`. Any of these made an explicit-only skill implicitly selectable. +- The conformance probe that checks the Validate helper makes no Git, tracker or + delivery calls now also intercepts calls made in-process; before, those + reached the real binaries unnoticed. +- The user-facing docs (`PRODUCT.md`, the docs index, how-it-works, architecture, + philosophy, migration and CI pages) describe validation as checks and CI plus + one fresh read where a mistake is costly, matching the skills. ## [3.8.0] - 2026-09-22 diff --git a/PRODUCT.md b/PRODUCT.md index 1e9328e00..da60af1f3 100644 --- a/PRODUCT.md +++ b/PRODUCT.md @@ -25,7 +25,7 @@ coding agents, checks, and independent judgment while existing tools retain work, execution, and delivery. The standard path is: ```text -Accepted intent -> native implementation and checks -> fresh independent judgment -> finish +Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish ``` ## Engineering practices in the workflow @@ -177,7 +177,7 @@ outcomes remain visible. Positive usefulness needs later task evidence. These skills maintain external context; they do not train weights or promise deterministic inference. [RPI traversal](docs/architecture/rpi-traversal.md) and [ADR-0017](docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md) own the amended -behavior and the explicitly optional fixed-dispatch reference adapter. +behavior. ## Evidence and claim limits diff --git a/README.md b/README.md index 4ee8ba6a4..0b98b72fe 100644 --- a/README.md +++ b/README.md @@ -21,7 +21,8 @@ AgentOps provides optional skills and a CLI (`ao`). The same `SKILL.md` skills w with coding agents (Claude Code, Codex, Cursor, OpenCode, Gemini CLI, Pi and others) and personal assistants (OpenClaw, Grok Bot). You state intent as behavior in your domain's words. The skills carry it through one change -(Plan → Implement → Validate, an **RPI**) or, for bigger work, a goal made of +(Plan → Implement → checks, with Validate where a mistake is costly: an **RPI**) +or, for bigger work, a goal made of many RPIs tracked in [Beads](https://github.com/gastownhall/beads), a dependency-aware issue tracker. @@ -33,7 +34,7 @@ many RPIs tracked in |---|---| | Builds something different from what you meant | Given/When/Then examples shared by implementation and review | | Uses three names for one concept | One domain term per concept, in intent, code and tests | -| Says “done” after a green test run | A fresh judge that didn't write the change | +| Says “done” after a green test run on a change that matters | A fresh judge that didn't write the change | | Loses the thread on work bigger than one session | A Beads graph holding intent, dependencies and verdicts | | Runs off with a half-formed goal | An interview that settles the goal before agents go autonomous | | Gives one model's answer to a hard call | A [council](skills/council/SKILL.md): judges in fresh contexts, each with the model, effort and perspective you assign (one model family or several vendors), compare, duel (score each other's ideas) or debate to your majority, keep dissent, and can answer an interview for you; [Idea Genie](skills/idea-genie/SKILL.md) brainstorms options | @@ -156,9 +157,10 @@ author never approves its own work. Merging and releasing follow your repo's rul ## Goals -[`rpi`](skills/rpi/SKILL.md) runs Plan → Implement → Validate for one outcome +[`rpi`](skills/rpi/SKILL.md) runs Plan → Implement → checks, with Validate where a +mistake is costly, for one outcome without check-ins (your agent's permission prompts still apply) and stops at -acceptance, a blocker or a spent limit. Bigger work becomes a goal (experimental; +acceptance, a blocker or a spent limit. Bigger work becomes a goal (it needs Beads: `brew install beads`, then `bd init` in your repo): 1. **[Interview](skills/interview/SKILL.md).** One question at a time, each with @@ -273,9 +275,9 @@ catalog: **[docs/SKILL-ROUTER.md](docs/SKILL-ROUTER.md)**. | Group | Skills | What it covers | |---|---|---| -| Operational loop | [`plan`](skills/plan/SKILL.md) [`implement`](skills/implement/SKILL.md) [`validate`](skills/validate/SKILL.md) | Shape, build and judge every change | +| Operational loop | [`plan`](skills/plan/SKILL.md) [`implement`](skills/implement/SKILL.md) [`validate`](skills/validate/SKILL.md) | Shape and build a change; judge it where a mistake is costly | | Autonomous | [`rpi`](skills/rpi/SKILL.md) | One outcome, end to end | -| Goals (experimental) | [`interview`](skills/interview/SKILL.md) [`craft-goal`](skills/craft-goal/SKILL.md) [`navigate`](skills/navigate/SKILL.md) | Shape, write and walk a goal over the bead graph | +| Goals | [`interview`](skills/interview/SKILL.md) [`craft-goal`](skills/craft-goal/SKILL.md) [`navigate`](skills/navigate/SKILL.md) | Shape, write and walk a goal over the bead graph | | Coordination | [`orchestrate`](skills/orchestrate/SKILL.md) [`agent-native`](skills/agent-native/SKILL.md) | Fresh workers per bead, disjoint scopes, integration | | On demand | [`research`](skills/research/SKILL.md) [`domain`](skills/domain/SKILL.md) [`test`](skills/test/SKILL.md) [`refactor`](skills/refactor/SKILL.md) [`review`](skills/review/SKILL.md) [`security`](skills/security/SKILL.md) [`doc`](skills/doc/SKILL.md) [`reverse-engineer`](skills/reverse-engineer/SKILL.md) | Reached for when a specific question comes up | | Learning | [`memory`](skills/memory/SKILL.md) | Curated `.context/` pages safe to commit | @@ -300,7 +302,8 @@ lineage; [how it works](docs/how-it-works.md) covers responsibilities. | Factories such as [Gas City](skills/using-gc/SKILL.md) | Own agent coordination and execution through their native control plane | Choose which workflow leads the task. Carry accepted behavior and evidence -into [independent judgment](skills/validate/SKILL.md). Shared practices are not +into checks, and into [independent judgment](skills/validate/SKILL.md) where a +mistake would be costly. Shared practices are not proof that every combination has been tested. @@ -355,7 +358,7 @@ Skill installation does not install tool dependencies: | Skill | Needs | Why | |---|---|---| -| `rpi` | `ao`, conditional | delegates exact-subject checks to Validate; only persists `verdict.v2` when requested, with the fixed-dispatch adapter optional | +| `rpi` | `ao`, conditional | delegates exact-subject checks to Validate; only persists `verdict.v2` when requested | | `plan` | `ao`, conditional | runs `ao provenance snapshot-intent` with an explicit evidence root when the intent source is not durable | | `implement` | `ao`, conditional | at an integration boundary whose changed paths affect bound evidence, runs `ao provenance evidence-orphans` | | `validate` | `ao` | derives exact subject identity with the helper and uses `ao provenance store-verdict` when persistence is requested; Python/schema checks are developer-only | diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 3980ea624..218bfcaca 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -1,16 +1,16 @@ # Architecture AgentOps defaults to native execution with zero mandatory skills. Its small -semantic core binds acceptance to exact content and fresh independent judgment; +semantic core binds acceptance to exact content and, where a mistake is costly, one fresh independent judgment; optional skills and adapters package guidance around that boundary. ```text existing bead or caller intent -> on-demand planning and bounded implementation -> runtime-derived subject-manifest.v1 + check receipts - -> fresh independent judgment - -> PASS | FAIL | NOT_PROVEN - -> direct repair and fresh revalidation within real bounds when needed + -> checks and CI as the gate for an ordinary change + -> one fresh judgment where a mistake is costly: PASS | FAIL | NOT_PROVEN + -> direct repair, confirmed by a check rather than another judgment -> completed acceptance or truthful unfinished result ``` @@ -33,9 +33,7 @@ existing bead or caller intent Known findings are repaired within authority and real remaining bounds. A genuine causal stall admits at most one bounded fresh helper; an unhelpful answer, cancellation, refusal or a spent real bound ends the attempt. Explicit -repair-round bounds still apply, and no invocation renews an allowance. The -[fixed-dispatch adapter](../skills/rpi/references/bounded-adapter.md) keeps its -narrower optional contract; it does not govern direct native execution. +repair-round bounds still apply, and no invocation renews an allowance. ## Hexagonal boundary diff --git a/docs/CHANGELOG.md b/docs/CHANGELOG.md index 48758aa42..dc7ce686e 100644 --- a/docs/CHANGELOG.md +++ b/docs/CHANGELOG.md @@ -9,6 +9,20 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Changed +- Fresh validation is no longer owed on every change. For an ordinary change + the author's checks and CI are the gate. RPI, Validate, Implement, + Orchestrate, Craft Goal and Navigate call for one fresh read only when the + caller asks, a mistake cannot be cheaply undone after it lands (a published + release or instructions users will follow, a security boundary, destroying + data or tracker state, deleting a check that protects the product), or no + deterministic check covers the changed behavior. A review is one round: the + reviewer does not re-run checks, reports as defects only what fails accepted + behavior or would mislead a user, break install or the CLI, or remove + protection for the product, and a repair does not start another review. A + requested binding verdict keeps its PASS rule; `NOT_PROVEN` is reported with + its gaps instead of being chased. +- Craft Goal, Interview and Navigate are marked stable; the README no longer + labels goals experimental. - The Codex plugin now loads `skills/` directly. `.codex-plugin/plugin.json` ships `./skills`, the same tree `ao skills link` and `npx skills` already install, instead of a generated copy. Skill names, descriptions and bodies are @@ -66,6 +80,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - `ao doctor` no longer has the `fm-skills-stale-codex-sync` failure mode, and its fixers can no longer write to `~/.codex/plugins/cache/agentops-marketplace` or `~/.codex/.agentops-codex-install.json`. +- The fixed-dispatch RPI reference adapter, which modelled repeated review + rounds: `skills/rpi/scripts/run_once.py`, its tests, its reference page and + `skills/rpi/references/rpi.feature`, plus `tests/e2e/rpi-phased-domain.sh`, a + no-op tombstone that existed only to keep that feature file's scenario link + resolving. +- The Skill Builder converter's Codex target no longer writes `prompt.md`; + Codex does not read it. - Bundled Flywheel tool skills (`account-rotation`, `agent-mail`, `cass`, `cc-hooks`, `dcg`, `ms`, `ntm`, `rch`, `sbh`, `using-flywheel`) and their generated Codex copies. @@ -102,6 +123,18 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 `~/.pi/agent/skills`, and detect Pi by `~/.pi/agent`. Links an earlier version made in `~/.pi/skills` are no longer swept by default; remove them with `ao skills unlink --dest ~/.pi/skills`. +- Skill Builder's `heal.sh --check` exits 2 with a message when `ao` cannot run. + It used to report a pass in non-strict mode. +- The Codex policy check also fails on `agents/openai.yaml` shapes Codex + silently ignores: a non-object `interface`, `dependencies` that is not a + mapping with a `tools` list, and a boolean spelled other than `true` or + `false`. Any of these made an explicit-only skill implicitly selectable. +- The conformance probe that checks the Validate helper makes no Git, tracker or + delivery calls now also intercepts calls made in-process; before, those + reached the real binaries unnoticed. +- The user-facing docs (`PRODUCT.md`, the docs index, how-it-works, architecture, + philosophy, migration and CI pages) describe validation as checks and CI plus + one fresh read where a mistake is costly, matching the skills. ## [3.8.0] - 2026-09-22 diff --git a/docs/CI-CD.md b/docs/CI-CD.md index fe3fdcc19..cab8c47e0 100644 --- a/docs/CI-CD.md +++ b/docs/CI-CD.md @@ -13,7 +13,7 @@ Repositories own delivery policy for local and cloud agents. ## Separation of responsibilities ```text -AgentOps: Plan -> Implement once -> fresh Validate -> bounded repair -> report +AgentOps: Plan -> Implement -> checks -> one fresh Validate where a mistake is costly -> report Repository: deterministic checks -> repository-selected Git/CI/release policy ``` diff --git a/docs/MIGRATION.md b/docs/MIGRATION.md index c62e8bbaf..e6554c5ba 100644 --- a/docs/MIGRATION.md +++ b/docs/MIGRATION.md @@ -3,11 +3,12 @@ AgentOps now owns one small product boundary: ```text -Accepted intent -> native implementation and checks -> fresh independent judgment -> finish +Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish ``` -(The 3.0 through 3.6 releases stopped after one validation; ADR-0017 added the -bounded repair phase.) +(The 3.0 through 3.6 releases stopped after one validation; ADR-0017 added a +bounded repair phase; 3.9 makes the fresh read conditional, with one round and +no re-review after a repair.) Known defects can be repaired directly and an approach can change within the accepted outcome, scope and real allowance. The caller owns changes to that diff --git a/docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md b/docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md index 13fe69d61..6ff9092c0 100644 --- a/docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md +++ b/docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md @@ -34,8 +34,9 @@ The caller approved a short RPI charter with on-demand Plan, Implement, Validate and Memory. This supersedes the once-only phase lock, mandatory anti-ceremony/premortem dispatch, mandatory entry Recall, Learn's verdict-only input/scratch-TTL limit, and a default two-round native stop. Those earlier -prescriptions below are historical; the pure fixed-dispatch reference retains -its narrower contract only for explicit adapter callers. +prescriptions below are historical. (2026-10-03: the fixed-dispatch reference +adapter, `run_once.py`, and its tests were removed; review is now one round with +no re-review after a repair, per the contract's validation rule.) Own the authorized outcome through finish. Plan may revise an approach when evidence disproves an assumption under unchanged accepted outcome and scope; diff --git a/docs/agent-workflow-reference.md b/docs/agent-workflow-reference.md index aedacffaa..20fb7d7cd 100644 --- a/docs/agent-workflow-reference.md +++ b/docs/agent-workflow-reference.md @@ -10,13 +10,19 @@ shell; the tracker, Git and runtime keep their existing authority. revise an approach when evidence disproves its assumption within unchanged outcome and scope. Changing acceptance needs caller authority. 3. Run cheap discriminating checks while editing, then required integration - checks. Reserve capacity for finishing, final judgment and needed repairs. -4. Obtain fresh author-distinct final judgment over the exact subject. Same-family - is default; cross-model is opt-in, with no fixed ten-minute cap. PASS requires - all acceptance and empty `not_checked`; evidence gaps remain visible. -5. Repair known findings within real bounds and revalidate the changed subject. - A genuine causal stall admits at most one bounded helper, never a chain. - Report the outcome and evidence truthfully when complete or actually stopped. + checks. For an ordinary change these and CI are the gate: finish. +4. Obtain one fresh author-distinct read only when the caller asks, a mistake + cannot be cheaply undone after it lands (a published release or instructions + users will follow, a security boundary, destroying data or tracker state, + deleting a check that protects the product), or no deterministic check + covers the changed behavior. Same-family is default; cross-model is opt-in. + The reviewer answers one question written in advance and does not re-run + the checks. +5. Repair what fails the accepted behavior or would mislead a user, break + install or the CLI, or remove protection for the product. Confirm each + repair with a check; a repair does not start another review. A genuine + causal stall admits at most one bounded helper, never a chain. Report the + outcome, what was checked and what was not when complete or actually stopped. [Memory](../skills/memory/SKILL.md) owns on-demand find/recall, capture/mine/learn and curate operations over caller-selected reviewed project `.context/` or @@ -32,8 +38,7 @@ Specialists, anti-ceremony audits, factories and outer-goal guidance are optiona The [RPI charter](../skills/rpi/SKILL.md) is an explicitly selected workflow; native goals and direct coding do not require it. A skill earns its context cost by resolving a task-specific need, not by occupying a phase in a sequence. -No new scheduler, command or process ledger is needed. The grandfathered pure -fixed-dispatch adapter is separately described in its own reference. +No new scheduler, command or process ledger is needed. Optional context-budget tooling (an opt-in read-budget hook plus bulk-read / code-write delegation) is described in [context-budget delegation](https://github.com/boshu2/agentops/blob/main/skills/agent-native/references/context-budget-delegation.md). diff --git a/docs/architecture/rpi-traversal.md b/docs/architecture/rpi-traversal.md index 0ae967a5f..23c0f0391 100644 --- a/docs/architecture/rpi-traversal.md +++ b/docs/architecture/rpi-traversal.md @@ -15,8 +15,8 @@ accepted intent -> Plan only for missing shape or consequential uncertainty -> Implement, direct repairs and cheap discriminating checks -> required integration checks - -> fresh author-distinct Validate over the exact subject - -> repair known findings within real bounds and revalidate changed content + -> one fresh author-distinct Validate, only where a mistake is costly + -> repair known findings and confirm each repair with a check -> completed outcome or truthful stopped result ``` @@ -72,6 +72,13 @@ They add no tracker, queue, Git delivery or runtime authority. ## Fresh final judgment (optional Validate guidance) +For an ordinary change the author's checks and CI are the gate. A fresh +judgment is used once, when the caller asks, when a mistake cannot be cheaply +undone after it lands (a published release or instructions users will follow, a +security boundary, destroying data or tracker state, deleting a check that +protects the product), or when no deterministic check covers the changed +behavior. The rest of this section describes that judgment. + A fresh author-distinct context reads the exact subject, unchanged acceptance and authorized evidence independently. A new role in the author's context is not fresh. Observe actual runtime/model/context identities; unknown or colliding @@ -98,8 +105,9 @@ bounded proof in criterion reasoning, and residual risk in the report. Classify commands before running them. Mutating checks (regeneration, sync, formatting) run against a disposable copy or committed subject, never overwrite -the judged working tree. Rerun acceptance-critical or uncertain checks; a -receipt is a claim to inspect, and structure alone cannot prove behavior. +the judged working tree. The reviewer does not re-run checks the author ran on +the exact subject or that CI will run; rerun a proof only for a risk-critical +claim that has no receipt. Structure alone cannot prove behavior. Changes to checks, fixtures, tolerances or specifications must satisfy original intent; weakening the oracle to get green is FAIL. @@ -107,8 +115,8 @@ Keep all necessary findings from selected judges with stable IDs/classes. Recurrence or unknown cause calls for causal examination; counts alone cannot prove progress, regression or a wrong design. Preserve disagreement; neither majority vote nor the author's preferred review certifies PASS. The reviewer reads -and returns judgment; implementers repair and obtain new fresh judgment over -changed exact content. +and returns judgment once; implementers repair and confirm each repair with a +check. A repair does not start another judgment unless the caller asks for one. When requested by a caller or declared consumer, the reviewer authors `verdict.v2` and uses `ao provenance store-verdict` for structural verification and atomic @@ -136,9 +144,6 @@ Keep recovery state in the native handoff only when needed to protect evidence. Optional [outer-goal guidance](../../skills/rpi/references/outer-goal.md) stays outside the core. The native controller owns aggregate enforcement and selected future outcomes. Objective text is no proof of stop or budget enforcement. -The grandfathered pure Python [fixed-dispatch adapter](../../skills/rpi/references/bounded-adapter.md) -retains its optional once-only dispatch and finite review-round contract; it is -not a native execution engine or the authority for the lean charter's repairs. Return the outcome, exact changed subject, strongest checks and material unchecked acceptance. `NOT_PLANNED` and `NOT_BUILT` are progress statuses, not verdicts. diff --git a/docs/contracts/codex-skill-api.md b/docs/contracts/codex-skill-api.md index 75f5a1ded..1b839603f 100644 --- a/docs/contracts/codex-skill-api.md +++ b/docs/contracts/codex-skill-api.md @@ -102,9 +102,9 @@ own `skills//agents/openai.yaml`. The file is hand-maintained in the source skill; nothing derives it. Without it, or when Codex cannot use it, Codex selects the skill implicitly. The conformance check fails when the file is missing, is not valid YAML, or lacks `policy.allow_implicit_invocation: false`. -It does not catch every file Codex drops: a non-object `interface` or -`dependencies.tools`, or the YAML 1.1 spelling `no` for the boolean, passes the -check and is ignored by Codex 0.156.1. +It also fails on the shapes Codex 0.156.1 drops silently: a non-object +`interface`, `dependencies` that is not a mapping with a `tools` list, and a +boolean spelled other than `true` or `false` (for example the YAML 1.1 `no`). --- diff --git a/docs/contracts/scenario-test-linkage.md b/docs/contracts/scenario-test-linkage.md index 5d2b3c1cc..f760430e3 100644 --- a/docs/contracts/scenario-test-linkage.md +++ b/docs/contracts/scenario-test-linkage.md @@ -27,10 +27,10 @@ GOALS directive ──► scenario (.feature) ──► test (executing) Tag a scenario with the test that covers it using a Gherkin tag: ```gherkin -@covered-by:tests/e2e/rpi-phased-domain.sh -Scenario: Phases run in order and never compress - When /rpi executes - Then it runs Research, then Plan, then Implement in order +@covered-by:tests/e2e/goals-trace-chain.sh +Scenario: A goal traces to its dependencies + When the goal is traced + Then each dependency appears in the trace ``` Rules: @@ -96,7 +96,7 @@ gated on changes to `skills/**`, `**/*.sh`, or `.github/**`. | Feature | Covering test | |---|---| -| `skills/rpi/references/rpi.feature` | `tests/e2e/rpi-phased-domain.sh` | +| ~~`skills/rpi/references/rpi.feature`~~ (removed 2026-10-03 with the fixed-dispatch adapter it described; the tombstone e2e went with it) | ~~`tests/e2e/rpi-phased-domain.sh`~~ | | ~~`skills/goals/references/goals.feature`~~ (stale row: the feature file was already absent from the tree before 2026-07-29; the skill is now `fitness`) | `tests/e2e/goals-measure-scenarios.sh`, `goals-trace-chain.sh` still execute against the `ao goals` CLI | | ~~`skills/scenario/references/scenario.feature`~~ (removed: `scenario` folded into `eval-outcomes`, 2026-06-12) | `tests/e2e/goals-scenarios-link.sh` | diff --git a/docs/contracts/skill-ports-and-adapters.md b/docs/contracts/skill-ports-and-adapters.md index 3fe1b697a..9c8ba10cb 100644 --- a/docs/contracts/skill-ports-and-adapters.md +++ b/docs/contracts/skill-ports-and-adapters.md @@ -85,8 +85,7 @@ is optional unless a caller or declared consumer requires it. A genuine causal stall admits at most one authorized bounded helper per incident; known failures get direct repair. Optional outer-goal guidance does not create a -scheduler, budget account or new command. The pure fixed-dispatch reference -adapter has its own narrower explicit contract and is not the native charter. +scheduler, budget account or new command. Every verdict binds the validator implementation and the verdict, report, and subject-manifest schema digests. Proof contracts advance through an explicit diff --git a/docs/first-value-path.md b/docs/first-value-path.md index 80affb8e2..9890b88de 100644 --- a/docs/first-value-path.md +++ b/docs/first-value-path.md @@ -24,8 +24,8 @@ For a small task this should be one reviewable change and one independent judgment. Persist `verdict.v2` only when a caller or declared consumer needs machine-readable evidence, using protected external non-Git storage. -Success is fresh independent judgment bound to acceptance and content -identities. Pushing or releasing that content follows repository policy. +Success is the accepted behavior, shown by checks bound to acceptance and +content identities, with one fresh judgment where a mistake would be costly. Pushing or releasing that content follows repository policy. The fix and regression check remain available for future changes; preserve useful decisions in your existing work record so the next session can continue from what was established. diff --git a/docs/how-it-works.md b/docs/how-it-works.md index fbb414910..83f8bf3c8 100644 --- a/docs/how-it-works.md +++ b/docs/how-it-works.md @@ -1,12 +1,12 @@ # How it works -AgentOps connects a requested behavior to implementation and independent -validation. Your native coding agent owns the authorized change; optional +AgentOps connects a requested behavior to implementation, checks and, where a +mistake would be costly, independent validation. Your native coding agent owns the authorized change; optional skills provide engineering guidance and `ao` supplies deterministic tools. No RPI invocation or skill installation is required for the native path. ```text -Accepted intent -> native implementation and checks -> fresh independent judgment -> finish +Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish ``` ## Agree on observable behavior diff --git a/docs/index.md b/docs/index.md index 6800ef050..9d2456cfa 100644 --- a/docs/index.md +++ b/docs/index.md @@ -15,12 +15,12 @@ examples. Given/When/Then expresses the starting situation, action, and expected result. Plain text in an issue or conversation is enough. Domain-driven design (DDD) keeps the domain's terms and rule ownership -consistent across those examples, implementation, and review. A fresh reviewer -checks the exact change against the original request; passing tests supply -evidence for that judgment. +consistent across those examples, implementation, and review. Passing tests and +CI check the change; where a mistake would be costly, a fresh reviewer also +checks the exact change against the original request. ```text -Accepted intent -> native implementation and checks -> fresh independent judgment -> finish +Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish ``` ## Leave useful improvements behind diff --git a/docs/install-day2-ops.md b/docs/install-day2-ops.md index 2a81d86de..9d4496b47 100644 --- a/docs/install-day2-ops.md +++ b/docs/install-day2-ops.md @@ -47,7 +47,7 @@ requirements. Most need nothing beyond the coding agent; these need more: | Skill | Needs | Why | |---|---|---| -| `rpi` | `ao`, conditional | delegates exact-subject checks to Validate; only persists `verdict.v2` when requested, with the fixed-dispatch adapter optional | +| `rpi` | `ao`, conditional | delegates exact-subject checks to Validate; only persists `verdict.v2` when requested | | `plan` | `ao`, conditional | runs `ao provenance snapshot-intent` with an explicit evidence root when the intent source is not durable | | `implement` | `ao`, conditional | at an integration boundary whose changed paths affect bound evidence, runs `ao provenance evidence-orphans` | | `validate` | `ao` | derives exact subject identity with the helper and uses `ao provenance store-verdict` when persistence is requested; Python/schema checks are developer-only | diff --git a/docs/philosophy.md b/docs/philosophy.md index fa7c16300..611940eab 100644 --- a/docs/philosophy.md +++ b/docs/philosophy.md @@ -10,7 +10,8 @@ own work. AgentOps therefore provides a small evidence protocol: ```text intent -> one bounded experiment -> exact subject identity - -> fresh independent judgment -> bounded repair -> report + -> checks, plus one fresh judgment where a mistake is costly + -> repair confirmed by a check -> report ``` The protocol is behavior-first. Plan expresses one behavior as normal and edge diff --git a/docs/scale-without-swarms.md b/docs/scale-without-swarms.md index c46562423..51b546212 100644 --- a/docs/scale-without-swarms.md +++ b/docs/scale-without-swarms.md @@ -3,12 +3,12 @@ AgentOps keeps its operating charter small: ```text -RPI charter -> on-demand Plan -> Implement and checks -> fresh Validate -> finish +RPI charter -> on-demand Plan -> Implement and checks -> one fresh Validate where a mistake is costly -> finish ``` Plan shapes missing intent or revises a disproved approach within accepted scope. Implement makes the change and repairs known defects directly. Validate -independently judges exact content. RPI owns the authorized outcome through +independently judges exact content when a mistake would be costly. RPI owns the authorized outcome through finish within real bounds; causal stalls admit at most one bounded fresh helper. Memory and specialists are optional, and completed acceptance ends the run. diff --git a/evals/agentops-core/rpi-behavior.json b/evals/agentops-core/rpi-behavior.json index 99be4c16e..65b05b2d7 100644 --- a/evals/agentops-core/rpi-behavior.json +++ b/evals/agentops-core/rpi-behavior.json @@ -68,29 +68,6 @@ "blocking_gate": "none" }, "cases": [ - { - "id": "bounded-repair-reference-tests", - "title": "Optional fixed-dispatch reference preserves finite supplied-round evaluation", - "kind": "command", - "objective": "Prove PASS converges, FAIL/NOT_PROVEN repair under the law, and evidence of acceptance progress and defect cause gates bounded repair (ADR-0017).", - "runtime": "shell", - "inputs": { - "cwd": "../..", - "shell": "python3 skills/rpi/tests/test_run_once.py && python3 scripts/check-cathedral-cut-conformance.py" - }, - "expectations": [ - { - "type": "exit_code", - "value": 0 - } - ], - "dimensions": [ - "correctness", - "process_adherence", - "safety" - ], - "critical": true - }, { "id": "rpi-skill-stop-contract", "title": "Lean RPI owns authorized outcome through checks and fresh judgment", diff --git a/registry.json b/registry.json index 29788b01c..5c36e8ff8 100644 --- a/registry.json +++ b/registry.json @@ -1077,7 +1077,7 @@ "has_skill_md": true, "name": "rpi", "path": "skills/rpi/", - "reference_count": 4, + "reference_count": 2, "tier": "meta" }, { diff --git a/scripts/.skill-python-grandfather b/scripts/.skill-python-grandfather index fac5948b5..64b66b125 100644 --- a/scripts/.skill-python-grandfather +++ b/scripts/.skill-python-grandfather @@ -25,7 +25,6 @@ skills/reverse-engineer/scripts/generate_feature_inventory_md.py skills/reverse-engineer/scripts/reverse_engineer.py skills/reverse-engineer/scripts/scaffold_feature_registry.py skills/reverse-engineer/scripts/validate_feature_registry.py -skills/rpi/scripts/run_once.py skills/security/scripts/prompt_redteam.py skills/security/scripts/security_suite.py skills/skill-builder/scripts/authoring_scan.py diff --git a/scripts/check-cathedral-cut-conformance.py b/scripts/check-cathedral-cut-conformance.py index 067faba8d..175deda39 100755 --- a/scripts/check-cathedral-cut-conformance.py +++ b/scripts/check-cathedral-cut-conformance.py @@ -191,110 +191,6 @@ def check_schema_index_docs() -> None: ) -def check_bounded_repair_contract() -> None: - """RPI's reference adapter stops under the convergence law (ADR-0017). - - Behavior probes only: fixture validation rounds in, stop reason out. Every - stop reason `run_once.py` declares must be reached by a probe, a round past - the caller's bound is never consumed, and the one bounded experiment - dispatches each phase once, on FAIL as on PASS. - """ - runner = ROOT / "skills" / "rpi" / "scripts" / "run_once.py" - assert runner.is_file(), "RPI has no executable reference behavior" - spec = importlib.util.spec_from_file_location("rpi_run_once_canary", runner) - assert spec and spec.loader - module = importlib.util.module_from_spec(spec) - spec.loader.exec_module(module) - dg = lambda ch: ch * 64 # noqa: E731 - def leg(status, ids, digest, evidence=("acceptance-receipt",), classes=None): - classes = classes or {} - return { - "status": status, - "findings": [ - {"id": i, "summary": i, **({"class": classes[i]} if i in classes else {})} - for i in ids - ], - "subject_digest": digest, - "evidence_refs": list(evidence), - "validator_family": "fresh", - "checked": ["acceptance"], - "not_checked": [], - } - fixed_a = {"ref": "fixed-a", "subject_digest": dg("b"), "resolves": ["a"]} - progressing = [leg("FAIL", ["a", "b"], dg("a")), leg("FAIL", ["b"], dg("b"), [fixed_a])] - canaries = { - "repair_budget_exhausted": (progressing + [leg("FAIL", ["b"], dg("c"))], {"repair_rounds": 1}), - "new_finding_requires_causal_review": ( - [leg("FAIL", ["a"], dg("a")), leg("FAIL", ["b"], dg("b"), [fixed_a])], {"repair_rounds": 2}), - "reopened_finding": (progressing + [leg("FAIL", ["a"], dg("c"))], {"repair_rounds": 3}), - "recurring_finding_class": ([ - leg("FAIL", ["a", "b"], dg("a"), classes={"a": "race", "b": "docs"}), - leg("FAIL", ["b"], dg("b"), [fixed_a], classes={"b": "docs"}), - leg("FAIL", ["c"], dg("c"), classes={"c": "race"}), - ], {"repair_rounds": 3}), - "no_acceptance_progress": ([leg("FAIL", ["a"], dg("a")), leg("PASS", [], dg("b"))], {"repair_rounds": 2}), - "introduced_regression": ([leg("FAIL", ["a"], dg("a")), leg("FAIL", ["b"], dg("b"), [ - fixed_a, {"ref": "comparison", "subject_digest": dg("b"), "introduced": ["b"]}])], {"repair_rounds": 2}), - "not_converged": ([leg("FAIL", ["a"], dg("a")), leg("FAIL", ["b", "c"], dg("b"), [ - fixed_a, {"ref": "prior-reproduction", "subject_digest": dg("a"), "preexisting": ["b", "c"]}])], - {"repair_rounds": 2}), - # A selected cross-family leg that is missing cannot converge. - "diversity_unsatisfied": ([leg("PASS", [], dg("a"))], {"cross_model": True}), - "converged": ([leg("FAIL", ["a"], dg("a")), leg("PASS", [], dg("b"), [fixed_a])], {"repair_rounds": 2}), - } - unprobed = set(module.STOP_REASONS) - set(canaries) - assert not unprobed, f"declared stop reasons without a behavior probe: {sorted(unprobed)}" - for expected, (rounds, options) in canaries.items(): - outcome = module.run_repair_phase(rounds, **options) - assert outcome["stop_reason"] == expected, ( - f"law canary {expected}: reference behavior stopped with {outcome['stop_reason']!r}" - ) - diverse = module.run_repair_phase(canaries["diversity_unsatisfied"][0], cross_model=True) - assert diverse["report"]["status"] == "NOT_PROVEN", "a single-family PASS certified a cross-family selection" - for digest in (dg("a"), dg("b")): - flip = module.run_repair_phase([leg("FAIL", ["a"], dg("a")), leg("PASS", [], digest)]) - assert flip["report"]["status"] == "NOT_PROVEN", "byte or verdict movement alone must not certify progress" - # A poison round beyond the caller's bound must never be normalized. - poison = module.run_repair_phase(progressing + [{"status": "poison-not-a-round"}], repair_rounds=1) - assert poison["stop_reason"] == "repair_budget_exhausted" and poison["rounds_used"] == 1, ( - "the repair phase consumed a round past the caller's bound" - ) - assert module.run_repair_phase([leg("FAIL", ["a"], dg("a"))], repair_rounds=0)["stop_reason"] == "repair_budget_exhausted" - - # The one bounded experiment never re-dispatches a phase after FAIL. - calls: list[str] = [] - - def phase(name: str, result: dict): - def run(*_args: object) -> dict: - calls.append(name) - return result - return run - - failed = module.invoke_once( - "fail-path probe", - phase("anti-ceremony", { - "decision": "CONTINUE", - "reason": "The probe needs one failing traversal.", - "frozen_outcome": "Observe one FAIL traversal", - "parked_process_work": [], - "remaining_proof": ["fresh validation"], - "stop_condition": "Stop after one fresh validation result.", - }), - phase("plan", {"intent_ref": "probe", "acceptance_digest": dg("d")}), - phase("implement", {"subject_manifest_digest": dg("e")}), - phase("validate", { - "verdict": "FAIL", - "acceptance_digest": dg("d"), - "subject_manifest_digest": dg("e"), - "author_context_id": "probe-author", - "validator_context_id": "probe-validator", - "freshness_attestation": {"source": "runtime", "attester_identity": "probe"}, - }), - ) - assert calls == ["anti-ceremony", "plan", "implement", "validate"], f"FAIL dispatch trace is {calls}" - assert failed["status"] == "FAIL", f"FAIL traversal reported {failed['status']!r}" - - def check_validate_helper() -> None: path = ROOT / "skills" / "validate" / "tests" / "validate.py" tree = ast.parse(path.read_text(encoding="utf-8"), filename=str(path)) @@ -367,7 +263,6 @@ def executor(packet: dict) -> str: def probe_no_substrate_calls() -> None: helper = ROOT / "skills" / "validate" / "tests" / "validate.py" - rpi_runner = ROOT / "skills" / "rpi" / "scripts" / "run_once.py" with tempfile.TemporaryDirectory() as raw: temp = Path(raw) subject = temp / "subject" @@ -393,34 +288,13 @@ def probe_no_substrate_calls() -> None: assert spec and spec.loader module = importlib.util.module_from_spec(spec) spec.loader.exec_module(module) - rpi_spec = importlib.util.spec_from_file_location("cathedral_rpi", rpi_runner) - assert rpi_spec and rpi_spec.loader - rpi = importlib.util.module_from_spec(rpi_spec) - rpi_spec.loader.exec_module(rpi) - # The intent SOURCE is bytes; the acceptance identity is sha256 of those - # bytes; the resolved mapping carries that identity as a declared fact. - # Deriving the bytes from the mapping that already contains the digest - # would be circular, and folding the two together is what let the - # RPI/Validate digest disagreement hide: this probe used to set - # `intent_bytes = canonical_bytes(resolved_intent)`, which is precisely - # the one input where a canonical-JSON digest of the mapping and a byte - # digest of the source coincide. Keeping them separate means the probe - # exercises the identity rather than a coincidence. + # Validate's acceptance identity is the sha256 of the intent source bytes. intent_source = { "intent_ref": "conversation:cathedral-probe", "acceptance": ["value.txt contains candidate"], "write_scope": {"include": ["value.txt"], "exclude": []}, } intent_bytes = module.canonical_bytes(intent_source) - resolved_intent = { - **intent_source, - "acceptance_digest": hashlib.sha256(intent_bytes).hexdigest(), - } - subject_facts = { - "subject_manifest_digest": payload["canonical_manifest_digest"], - "subject_manifest": payload, - "checks": ["manifest"], - } draft = { "acceptance_digest": "a" * 64, "subject_manifest_digest": payload["canonical_manifest_digest"], @@ -436,32 +310,12 @@ def probe_no_substrate_calls() -> None: "validated_at": "2026-07-14T00:00:00Z", } verdict_dir = temp / ".agents" / "ao" / "verdicts" / "sha256" - calls: list[str] = [] - - def anti_ceremony_guard(_intent: object) -> dict: - calls.append("anti-ceremony") - return { - "decision": "CONTINUE", - "reason": "The frozen outcome still requires implementation proof.", - "frozen_outcome": "Write and prove the non-Git candidate", - "parked_process_work": [], - "remaining_proof": ["manifest", "fresh validation"], - "stop_condition": "Stop after one fresh validation result.", - } - - def plan_phase(_intent: object) -> dict: - calls.append("plan") - return resolved_intent - - def implement_phase(received_intent: dict) -> dict: - calls.append("implement") - assert received_intent == resolved_intent - return subject_facts - - def validate_phase(received_intent: dict, received_subject: dict) -> dict: - calls.append("validate") - assert received_intent == resolved_intent - assert received_subject == subject_facts + + # The in-process call must see the fake executables too, or a Git or + # tracker call from the helper would reach the real binary unnoticed. + saved_path = os.environ.get("PATH", "") + os.environ["PATH"] = env["PATH"] + try: artifact, verdict_path, existed = module.store_verdict( draft, verdict_dir, @@ -473,30 +327,10 @@ def validate_phase(received_intent: dict, received_subject: dict) -> dict: "runtime", "non-git-validator", ) - assert not existed - return { - "verdict": artifact["verdict"], - "acceptance_digest": artifact["acceptance_digest"], - "subject_manifest_digest": artifact["subject_manifest_digest"], - "verdict_digest": artifact["artifact_digest"], - "verdict_ref": str(verdict_path), - "author_context_id": artifact["author_context_id"], - "validator_context_id": artifact["validator_context_id"], - "freshness_attestation": artifact["freshness_attestation"], - "checked": artifact["checked"], - "not_checked": artifact["not_checked"], - } - - rpi_report = rpi.invoke_once( - "temporary non-Git experiment", - anti_ceremony_guard, - plan_phase, - implement_phase, - validate_phase, - ) - verdict_path = Path(rpi_report["verdict_ref"]) - assert calls == ["anti-ceremony", "plan", "implement", "validate"], f"RPI dispatch trace is {calls}" - assert rpi_report["status"] == "PASS" and verdict_path.is_file() + finally: + os.environ["PATH"] = saved_path + assert not existed + assert artifact["verdict"] == "PASS" and verdict_path.is_file() assert verdict_path.parent == verdict_dir assert not called.exists(), "Validate helper invoked a Git, tracker, push, or delivery executable" @@ -506,7 +340,6 @@ def main() -> int: check_removed_skills, check_core_schemas, check_schema_index_docs, - check_bounded_repair_contract, check_validate_helper, check_dispatch_once, probe_no_substrate_calls, diff --git a/scripts/validate-codex-api-conformance.sh b/scripts/validate-codex-api-conformance.sh index 6fcde0bef..27e378f56 100755 --- a/scripts/validate-codex-api-conformance.sh +++ b/scripts/validate-codex-api-conformance.sh @@ -35,6 +35,7 @@ if [[ ! -d "$SKILLS_ROOT" ]]; then fi python3 - "$SKILLS_ROOT" <<'PY' +import re import sys from pathlib import Path @@ -117,6 +118,22 @@ def invocation_policy(skill: str, skill_dir: Path) -> tuple[bool, object]: if not isinstance(data, dict): fail(skill, "agents/openai.yaml must be a mapping") return False, None + # Codex 0.156.1 silently drops the whole file, policy included, when any of + # these shapes is wrong, so an explicit-only skill would become implicit. + if "interface" in data and not isinstance(data["interface"], dict): + fail(skill, "agents/openai.yaml interface must be a mapping; Codex ignores the file otherwise") + return False, None + deps = data.get("dependencies") + if deps is not None and ( + not isinstance(deps, dict) or ("tools" in deps and not isinstance(deps["tools"], list)) + ): + fail(skill, "agents/openai.yaml dependencies must be a mapping with a tools list; Codex ignores the file otherwise") + return False, None + raw_allow = re.search(r"^\s*allow_implicit_invocation:\s*([^\s#]+)", path.read_text(encoding="utf-8"), re.M) + if raw_allow and raw_allow.group(1) not in {"true", "false", "True", "False", "TRUE", "FALSE"}: + fail(skill, "agents/openai.yaml policy.allow_implicit_invocation must be spelled true or false; " + f"Codex reads {raw_allow.group(1)!r} as a string and ignores the file") + return False, None policy = data.get("policy") if policy is None: return True, None diff --git a/skills/catalog.json b/skills/catalog.json index 9c8f54bb3..57666fdef 100644 --- a/skills/catalog.json +++ b/skills/catalog.json @@ -772,7 +772,7 @@ "produces": [ "rpi-report.v1" ], - "references_count": 4, + "references_count": 2, "tier": "meta", "user_invocable": true }, diff --git a/skills/craft-goal/SKILL.md b/skills/craft-goal/SKILL.md index b97078ec3..da0146f68 100644 --- a/skills/craft-goal/SKILL.md +++ b/skills/craft-goal/SKILL.md @@ -28,7 +28,7 @@ metadata: effects: [] canonical_status: canonical disposition: keep_strategy - stability: experimental + stability: stable output_contract: 'human-readable SAFE_TO_CREATE, USE_RPI, or UNSAFE_GOAL decision; copy-paste outer-goal prompt when safe; exact budgets, assumptions, and lint findings' --- @@ -40,9 +40,9 @@ the goal selects the next useful trial, preserves what was learned, and ratchets toward a larger outcome. ```text -Goal / Mayor: observe graph → choose bounded wave → consume verdicts → ratchet +Goal / Mayor: observe graph → choose bounded wave → consume results → ratchet └─ Bead: durable experiment intent, context, scratch, evidence, and links - └─ RPI: plan → implement → fresh validate → bounded repair → verdict → report + └─ RPI: plan → implement → checks → one fresh validate where a mistake is costly → report └─ Implementation: one RED → GREEN → refactor experiment ``` @@ -106,10 +106,14 @@ outcome; do not invent one universal budget. - **Bead knowledge graph:** Use the tracker as durable memory, not a parallel goal ledger. Root epic = outer intent; child bead = one experiment/RPI. **Why:** compaction must not erase the scientific record. -- **RPI membrane:** One candidate gets one bounded RPI and an author-distinct - fresh validation result. The goal may request durable verdict evidence but - never rewrites it. - **Why:** orchestration cannot author its own proof. +- **RPI membrane:** One candidate gets one bounded RPI. Its checks and CI are + the result for an ordinary bead. A bead gets one author-distinct fresh + validation only when the caller asks, a mistake cannot be cheaply undone + after it lands, or no deterministic check covers the changed behavior; a + repair does not start another. The goal may request durable verdict evidence + but never rewrites it. + **Why:** orchestration cannot author its own proof, and a review per bead + multiplies cost across the whole graph. - **Brownian ratchet:** Continue only when a result adds non-duplicative, decision-relevant knowledge or advances acceptance. **Why:** activity without information is churn. diff --git a/skills/domain/references/standards/skill-structure.md b/skills/domain/references/standards/skill-structure.md index f8ff616d1..dfef8e736 100644 --- a/skills/domain/references/standards/skill-structure.md +++ b/skills/domain/references/standards/skill-structure.md @@ -74,7 +74,8 @@ rpi -> validate ``` These are available core operations, not mandatory worksheets or dispatches for -every edit. RPI uses Plan on demand and requires fresh final Validate. +every edit. RPI uses Plan on demand and one fresh Validate only where a mistake +is costly or the caller asks. Anti-ceremony and Memory are optional, with no hard edge. ## Body contract diff --git a/skills/domain/references/standards/test-pyramid.md b/skills/domain/references/standards/test-pyramid.md index 04e00501e..523644dfc 100644 --- a/skills/domain/references/standards/test-pyramid.md +++ b/skills/domain/references/standards/test-pyramid.md @@ -100,5 +100,6 @@ Good test evidence records: - environment assumptions that affect reproducibility; - what the check did not cover. -Green tests are factual evidence, not a semantic verdict. Validate supplies the -independent judgment against the exact candidate. +Green tests are factual evidence, not a semantic verdict. For an ordinary change +they and CI are the gate. Validate supplies an independent judgment against the +exact candidate when the caller asks for one or a mistake would be costly. diff --git a/skills/implement/SKILL.md b/skills/implement/SKILL.md index fbc256c2b..71b7c1c91 100644 --- a/skills/implement/SKILL.md +++ b/skills/implement/SKILL.md @@ -64,7 +64,8 @@ to the existing Security owner and service test design to Test. Check a representative change against existing constraints before bulk propagation; check the authored source set before broad regeneration. Repair known failures directly and verify the exact result (such as a cited - file, assertion or returned record) before another review. A disproved assumption + file, assertion or returned record) with a check. A repair does not start + another review. A disproved assumption may change the approach within scope; use Plan only for consequential uncertainty. 4. Use targeted tests and applicable repository lint/static checks before broad integration. Read the repository's actual check recipe, including @@ -137,7 +138,9 @@ Respect remaining caller/native bounds and reserve finishing capacity; retries r Return facts, not semantic PASS. An implement-only handoff does not authorize Git, tracker or delivery transitions; existing caller authority remains usable. -A full outcome request continues through fresh independent final judgment; +A full outcome request finishes on its checks and CI. It continues to one fresh +independent judgment only when the caller asks, a mistake cannot be cheaply +undone after it lands, or no deterministic check covers the changed behavior. RPI is optional and explicitly selected. Success is working behavior with usable evidence, not volume of logs or process artifacts. diff --git a/skills/interview/SKILL.md b/skills/interview/SKILL.md index 786357cad..fadadcf1a 100644 --- a/skills/interview/SKILL.md +++ b/skills/interview/SKILL.md @@ -18,7 +18,7 @@ metadata: effects: [update_intent_source] canonical_status: canonical disposition: keep_strategy - stability: experimental + stability: stable output_contract: 'one question per turn with a labeled recommendation and tradeoff; decided, terms, open and deferred lists the caller can see; settled decisions appended to the existing intent source within authority; a handoff to Craft Goal, RPI or Plan' --- diff --git a/skills/navigate/SKILL.md b/skills/navigate/SKILL.md index ef99bf1f3..567f261be 100644 --- a/skills/navigate/SKILL.md +++ b/skills/navigate/SKILL.md @@ -15,7 +15,7 @@ metadata: effects: [update_native_graph] canonical_status: canonical disposition: keep_strategy - stability: experimental + stability: stable output_contract: 'wave checkpoint in the existing handoff or root epic: acceptance matrix, frontier, wave and reasons, ratchets and churn, budget, helper use and native state, next thesis, open decisions; a single pass returns it with hygiene findings and writes nothing' --- @@ -84,12 +84,16 @@ means one bead. Prefer an early falsifier. Hand each bead to one RPI. When delegation is authorized, hand it to Orchestrate or Agent Native to dispatch, one bead per worker; otherwise the -caller's runtime runs it. Each candidate gets one fresh, author-distinct Validate. +caller's runtime runs it. A candidate's checks and CI are its result. A +candidate gets one fresh, author-distinct Validate only when the caller asks, a +mistake cannot be cheaply undone after it lands, or no deterministic check +covers the changed behavior; a repair does not start another. Done when each picked bead has a one-line reason and a named handoff. ## 3. Ratchet the graph -Record each verdict unchanged on its bead, for example +Record each result unchanged on its bead: the check facts, and the verdict when +one was obtained. For example `bd update --append-notes "verdict: FAIL; evidence: ; learned: "`. Update its matrix row, then classify each discovery: @@ -107,7 +111,7 @@ from regressions by before/after reproduction or equivalent causal evidence under the same acceptance; counts, timestamps and new ids prove no cause. Unknown cause, a reopened finding or recurrence of a closed finding class is HOLD, not proof the design is wrong. Keep necessary findings necessary; nothing -resets a total. Done when every verdict sits on its bead and every discovery +resets a total. Done when every result sits on its bead and every discovery has a class. ## 4. Checkpoint diff --git a/skills/orchestrate/SKILL.md b/skills/orchestrate/SKILL.md index 36a32d0c5..028e87bfb 100644 --- a/skills/orchestrate/SKILL.md +++ b/skills/orchestrate/SKILL.md @@ -84,21 +84,29 @@ engagement nor acceptance. ## Integrate and obtain judgment -Name the integration and final-validation responsibility before launch. Follow -the consumer repository's integration policy, include all changed paths and -generated companions, and run affected checks on the actual integrated subject. +Name the integration responsibility before launch, and decide then whether the +integrated candidate needs a fresh judgment at all. Follow the consumer +repository's integration policy, include all changed paths and generated +companions, and run affected checks on the actual integrated subject. Acceptance of a leaf does not establish the combined release. -Assign fresh author-distinct judgment of the exact candidate against unchanged -acceptance through [Validate](../validate/SKILL.md), the sole skill owner of -acceptance semantics. Its identity, freshness, complete checked scope and -evidence requirements remain authoritative; preserve every explicitly required -review leg. Advisory Review, Plan challenge and Council advice are not binding -acceptance, and must be refused when offered in place of that judgment. -Changed candidate bytes invalidate the old subject binding and require judgment -of the new exact subject. Deterministic green or convincing author rationale -cannot fill missing proof. No report format or persisted artifact is mandatory -unless the caller or an existing consumer requires one. +For an ordinary candidate the integrated checks and CI are the gate. Assign one +fresh author-distinct judgment through [Validate](../validate/SKILL.md), the +sole skill owner of acceptance semantics, only when the caller asks, when a +mistake cannot be cheaply undone after it lands (a published release or +instructions users will follow, a security boundary, destroying data or tracker +state, deleting a check that protects the product), or when no deterministic +check covers the changed behavior. Preserve every explicitly required review +leg. Advisory Review, Plan challenge and Council advice are not binding +acceptance and cannot stand in for a judgment the caller requested. + +One round. The validator does not re-run the integrated checks. Route back as +repairs only what fails the accepted behavior or would mislead a user, break +install or the CLI, or remove protection for the product; confirm each repair +with a check. Changed candidate bytes need their affected checks rerun; they +need a new judgment only when the caller asks for one. Keep review cost a +fraction of the cost of the work. No report format or persisted artifact is +mandatory unless the caller or an existing consumer requires one. ## Reconcile feedback and resume diff --git a/skills/reverse-engineer/SKILL.md b/skills/reverse-engineer/SKILL.md index e3e5e0b0b..f3c305291 100644 --- a/skills/reverse-engineer/SKILL.md +++ b/skills/reverse-engineer/SKILL.md @@ -98,7 +98,7 @@ Verdict rules (hard-won — apply them, do not skip): - **steal** — we lack it and it advances our core. Steal the *pattern*, not the storage engine: re-express in our primitives, never vendor their runtime. - **park** — real, but it's substrate we deliberately delegate (e.g. orchestration per ADR-0009) or downstream of an unproven bet. Name it, don't build it. -- **reject** — it conflicts with our doctrine (e.g. a self-reported completion edge where we require a verdict — "no verdict = not done"). +- **reject** — it conflicts with our doctrine (e.g. a completion edge with no check behind it, where we require checks and CI, plus one fresh judgment for a costly mistake). - **have** — we already do this; confirm it still holds, move on. - **gap** — we should have it and don't. These are the steal candidates. diff --git a/skills/rpi/SKILL.md b/skills/rpi/SKILL.md index 66643135e..9969c5c59 100644 --- a/skills/rpi/SKILL.md +++ b/skills/rpi/SKILL.md @@ -37,7 +37,8 @@ output_contract: 'concise human-readable result; optional rpi-report.v1 when a c Own the authorized outcome through finish. Use the native coding agent and shell. BD or the caller's tracker owns work and handoffs; Git owns content and -delivery. AgentOps supplies a small charter and fresh judgment, not a scheduler. +delivery. AgentOps supplies a small charter and one fresh judgment where a +mistake is costly, not a scheduler. ## Operating charter @@ -53,17 +54,27 @@ delivery. AgentOps supplies a small charter and fresh judgment, not a scheduler. A known test failure needs a fix and a discriminating check, not another planning phase, council or helper. 4. Use focused checks during edits and complete required integration checks - before final judgment. Reuse valid exact-input receipts; rerun affected - checks after changes. Reserve capacity for integration, review and repair. - Keep the final subject unchanged while it is being judged. -5. Obtain [Validate](../validate/SKILL.md) from one fresh author-distinct context - in the author's model family unless the caller selects required additional - legs. There is no fixed ten-minute cap; explicitly required - reviewers remain required. Risk deepens evidence, not reviewer multiplication. Repair - actionable findings within authority and remaining bounds, then revalidate - the changed exact subject. For `NOT_PROVEN` from missing evidence, gather it - without changing the subject and obtain fresh judgment again. -6. Stop at completed acceptance, cancellation, refusal, a spent real bound or + before finishing. Reuse valid exact-input receipts; rerun affected checks + after changes. Reserve capacity for integration and repair. Keep a subject + unchanged while it is being judged. +5. Spend validation where a mistake is costly. For an ordinary change the + checks and CI are the gate: finish. Obtain [Validate](../validate/SKILL.md) + from one fresh author-distinct context only when the caller asks, when a + mistake cannot be cheaply undone after it lands (a published release or + instructions users will follow, a security boundary, destroying data or + tracker state, deleting a check that protects the product), or when no + deterministic check covers the changed behavior. Use the author's model + family unless the caller selects additional legs; explicitly required + reviewers remain required. +6. One round. Give the validator the exact subject and one question written + before it starts; it does not re-run the checks. Repair what fails the + accepted behavior or would mislead a user, break install or the CLI, or + remove protection for the product; treat the rest as optional notes. Confirm + each repair with a check and finish. A repair does not start another + review, and `NOT_PROVEN` is reported with its gaps, not chased. Keep review + cost a fraction of the cost of the work; when it approaches that cost, stop + and report what is unchecked. +7. Stop at completed acceptance, cancellation, refusal, a spent real bound or an unresolved causal stall after the help below. Adjacent improvements are not permission to expand the goal. Report them briefly only when useful; do not turn them into another work batch. @@ -115,19 +126,19 @@ helper use in the native handoff. Prompt text proves no native enforcement. ## Evidence and boundaries -Bind accepted intent, complete changed paths, exact subject and factual receipts -for the fresh validator; disclose affected orphaned acceptance evidence. Use +When a validator is used, bind accepted intent, complete changed paths, exact +subject and factual receipts for it; disclose affected orphaned acceptance evidence. Use existing provenance helpers rather than a new evidence format. Requested proof uses caller-selected protected external non-Git storage; preserve legacy -`.agents/` evidence. Missing identity, freshness or proof means NOT_PROVEN; -proven failed acceptance or scope violation means FAIL. PASS needs every -criterion verified and empty `not_checked`. Authors cannot issue binding PASS. +`.agents/` evidence. For a requested binding verdict, missing identity, +freshness or proof means NOT_PROVEN; proven failed acceptance or scope +violation means FAIL; PASS needs every criterion verified and empty +`not_checked`. Authors cannot issue binding PASS. Without that request, report +what was checked and what was not, and finish. [Memory](../memory/SKILL.md), specialists and runtime adapters are on demand; no-match and no-change are valid. Read [boundaries](references/boundaries.md) -when authority, scope, evidence or delivery is at issue. The optional -[fixed-dispatch adapter](references/bounded-adapter.md) is not the native -execution engine. Do not invent a runtime, hidden machine artifact or workflow +when authority, scope, evidence or delivery is at issue. Do not invent a runtime, hidden machine artifact or workflow to finish an ordinary change. Report the result, strongest checks and material limits. Plans, activity, diff --git a/skills/rpi/references/boundaries.md b/skills/rpi/references/boundaries.md index 664e15c33..5fe64bb52 100644 --- a/skills/rpi/references/boundaries.md +++ b/skills/rpi/references/boundaries.md @@ -1,7 +1,7 @@ # Ownership boundaries for the lean RPI core RPI owns the authorized outcome through implementation, checks, direct repairs -and fresh final judgment. Plan shapes missing intent and may revise an approach +and, where a mistake is costly or the caller asks, one fresh judgment. Plan shapes missing intent and may revise an approach falsified by evidence within unchanged accepted outcome/scope. Implement edits and collects facts. Validate independently judges the exact subject and alone authors semantic `verdict.v2` when persistence is selected. Memory is optional; @@ -48,12 +48,16 @@ and repairs sessions. Concurrent writers require authorized disjoint source and regeneration scope and isolation. Pass bounded task evidence, not the author's desired verdict. Do not start another runtime merely because it exists. -The optional `run_once.py` developer adapter retains its explicitly selected -fixed-dispatch and finite-round contract in [bounded-adapter.md](bounded-adapter.md). -It does not restrict native approach revision or implement direct repair for you. +Nothing in this reference restricts native approach revision or implements +direct repair for you. ## Fresh judgment +A fresh judgment is used once, and only when the caller asks, a mistake cannot +be cheaply undone after it lands, or no deterministic check covers the changed +behavior. Otherwise the author's checks and CI are the gate. A repair is +confirmed by a check and does not start another judgment. + The author cannot issue binding PASS. Judge legs read; implementers fix. Default to a fresh author-distinct same-family reviewer. Cross-model review is opt-in; an explicitly required unavailable leg leaves NOT_PROVEN. No fixed ten-minute diff --git a/skills/rpi/references/bounded-adapter.md b/skills/rpi/references/bounded-adapter.md deleted file mode 100644 index a6a803dae..000000000 --- a/skills/rpi/references/bounded-adapter.md +++ /dev/null @@ -1,64 +0,0 @@ -# Optional fixed-dispatch reference adapter - -The grandfathered `scripts/run_once.py` is a pure developer reference for callers -that explicitly select fixed dispatch and a finite list of supplied review rounds. -It invokes an explicit anti-ceremony function, Plan and Implement at most once; -its repair evaluator consumes supplied evidence and cannot execute agents, infer -causes, fix subjects or enforce aggregate budgets. Installed native RPI follows -its operating charter, not this adapter. The old phase lock and default two-round -limit apply only to this selected adapter, never as a restriction on native -implementation's direct repairs or evidence-driven approach revision. - -Its existing deterministic tests guard exact evidence and finite consumption; -they do not prove native agent behavior or practical benefit. It stops when -converged, stopped by the law, or out of `repair_rounds` and never extends the -caller's bound. The adapter preserves these narrower admission semantics: - -## The convergence law - -A repair round is admitted only while all hold: - -1. `rounds_used < repair_rounds` (caller-declared, default 2). -2. New digest-bound evidence proves closure of a named acceptance finding or, - for `NOT_PROVEN`, resolves a named proof gap. A changed digest or a smaller - finding count alone is not useful progress. Generated-only changes qualify - only when the evidence proves that they repair required behavior or parity. - An unchanged subject previously judged FAIL cannot be repaired by a new label - or verdict flip; changed bytes still require acceptance proof. -3. No finding id closed in an earlier round reopens. No closed finding class - recurs, and no introduced regression or new finding of unknown cause is - admitted. Before/after reproduction or equivalent causal evidence under the - same acceptance must distinguish a pre-existing discovery from a regression; - neither counts, timestamps, nor a new id establish that distinction. - -Keep the union of every required judge's findings, keyed by stable -`findings[].id`; do not hide a necessary finding as optional. Newly exposed -pre-existing defects may increase the open count while another acceptance gap -is demonstrably closed. Their evidence must prove prior existence; -unknown cause stops repair for causal examination even if another gap closed. -Validators reuse a short stable `class` for each kind of defect. A reopened id -or returning class warrants causal HOLD in a selected outer goal. Recurrence -alone does not prove that the design is wrong and never auto-reopens Plan. - -Reuse existing check receipts, findings summaries, and evidence references for -this reasoning. In the pure reference, decoded receipt bindings use `ref`, -`subject_digest`, and `resolves` for ids actually closed. `preexisting` ids must -bind reproduction to the prior subject digest; `introduced` ids bind causal -comparison to the current digest and stop repair. These are supplied receipt -facts, not new persisted verdict fields or a lifecycle schema. The reference -cannot prove a receipt's truth or infer cause from wording. - -Converged: the fresh validator returns PASS and every required cross-family -validator does too, over the exact subject and all acceptance with empty -`not_checked`. On any violation RPI stops and reports the current status. -`checked` carries one line per round (`repair round N: k open findings`); open -findings ride in the result and the report. A reworded finding with the same id -is the same finding. Acceptance and its digest stay fixed. The orchestrating -context fixes; judge legs only read. RPI convenes no further judge of its own, -does not escalate, and does not auto-replan. - - -An unknown cause or recurrence returns evidence to the native caller; the -adapter dispatches no helper. The native charter decides whether a genuine -causal stall merits its single bounded consultation. This does not revive a -spent caller bound or change a completed verdict. diff --git a/skills/rpi/references/outer-goal.md b/skills/rpi/references/outer-goal.md index fb84a10d4..5505b5657 100644 --- a/skills/rpi/references/outer-goal.md +++ b/skills/rpi/references/outer-goal.md @@ -14,7 +14,7 @@ wave. The goal still selects work, and neither adds a scheduler or a ledger. Carry accepted terminal outcome and scope, measured remaining allowance and the current causal incident in the native work/handoff source. Choose the smallest acceptance-advancing action or consequential uncertainty. Reserve capacity for -integration, final fresh judgment, required repairs and a useful handoff before +integration, required repairs, a fresh judgment where one is warranted and a useful handoff before spending the whole allowance on discovery or reviews. Apply the charter's at-most-one bounded helper to a genuine causal stall. Known diff --git a/skills/rpi/references/rpi.feature b/skills/rpi/references/rpi.feature deleted file mode 100644 index 5f94c66ce..000000000 --- a/skills/rpi/references/rpi.feature +++ /dev/null @@ -1,49 +0,0 @@ -Feature: Optional fixed-dispatch reference adapter remains bounded - @covered-by:skills/rpi/tests/test_run_once.py::test_anti_ceremony_guard_runs_once_before_plan - Scenario: Guard CONTINUE preserves the core phase order - Given one intent - When the fixed-dispatch adapter is explicitly selected - Then the anti-ceremony guard is invoked exactly once before Plan - And Plan and Implement are each dispatched at most once in that order, and fresh Validate repeats only inside the bounded repair phase - And the final report contains no next action - - @covered-by:skills/rpi/tests/test_run_once.py::test_anti_ceremony_stop_dispatches_no_core_phase - Scenario: Guard STOP admits no core phase - Given the anti-ceremony guard returns STOP with its required response fields - When the fixed-dispatch adapter is explicitly selected - Then Plan, Implement, and Validate are not dispatched - And RPI reports NOT_PLANNED and stops - - @covered-by:skills/rpi/tests/test_run_once.py::test_fail_from_one_experiment_feeds_the_repair_phase - Scenario: Validation failure enters the bounded repair phase - Given Validate returns FAIL or NOT_PROVEN with findings - When the convergence law admits another round - Then the adapter evaluates supplied repair evidence, without runtime dispatch or delivery - - @covered-by:skills/rpi/tests/test_run_once.py::test_repair_stops_when_a_closed_finding_reopens - Scenario: The convergence law stops a repair spiral - Given a repair round reopens a closed finding id or has no new acceptance-relevant proof - When RPI evaluates the law - Then RPI stops and reports the current status with the open findings - - @covered-by:skills/rpi/tests/test_run_once.py::test_discovered_preexisting_defects_may_grow_count_with_real_progress - Scenario: Discovery is distinct from regression - Given a repair closes a named acceptance gap with a new digest-bound receipt - And new findings are proven to exist on the prior exact subject - When the new findings increase the open count - Then RPI retains them and admits bounded repair without declaring a regression - - @covered-by:skills/rpi/tests/test_run_once.py::test_new_finding_cannot_hide_behind_another_resolved_gap - Scenario: Unknown cause requires causal examination - Given a repair closes one acceptance gap but exposes a new finding of unknown cause - When RPI evaluates the law - Then RPI stops even if the open count did not grow - - @covered-by:skills/rpi/scripts/validate.sh - Scenario: Interactive output does not require a machine artifact - Given RPI has received one fresh validation result - When RPI responds to an interactive caller - Then the response leads with status and the caller-visible outcome - And it includes only the strongest proof and material unchecked scope - And no hidden rpi-report.v1 or verdict.v2 is created - And a machine artifact is emitted only when a caller or declared consumer requested it diff --git a/skills/rpi/scripts/run_once.py b/skills/rpi/scripts/run_once.py deleted file mode 100644 index fc55255e6..000000000 --- a/skills/rpi/scripts/run_once.py +++ /dev/null @@ -1,549 +0,0 @@ -#!/usr/bin/env python3 -"""Pure reference behavior for one RPI invocation and its bounded repair phase. - -The caller supplies one anti-ceremony guard and the three core phase functions. -This module invokes the guard once before Plan, dispatches Plan and Implement at -most once, and never chooses a retry, a budget, or a next action. - -Under ADR-0017 (loop as control flow, not knowledge) the traversal no longer -ends at the first validation result. `run_repair_phase` models the bounded -repair phase as pure data: it consumes validate rounds that already happened and -decides, under the convergence law, whether another repair round is admitted. -It performs no I/O, dispatches nothing, and owns no budget of its own — the -caller declares `repair_rounds`. -""" - -from __future__ import annotations - -from collections.abc import Callable, Mapping, Sequence -import re -from typing import Any - - -# The exact-identity property is BYTE-addressed: Validate snapshots the resolved -# intent bytes under `sha256(bytes)` and stores them as `.intent` -# (validate.py snapshot_intent), then re-derives that same digest from the same -# bytes when it binds runtime facts into the verdict. RPI is a dispatcher, not a -# second digest authority — it carries the digest Plan declares over the bytes it -# snapshotted, and cross-checks Validate's independently re-derived value against -# it. -# -# This module previously computed its own `sha256(canonical-JSON(mapping))` here -# and hard-compared that against Validate's `sha256(raw bytes)`. The two can -# never agree unless the source is byte-identical canonical JSON, so the composed -# contract was broken; both unit suites stayed green only because the RPI test -# mocked Validate with THIS module's digest function. A canonical-JSON digest is -# also the wrong identity in principle: two different source files that parse to -# the same mapping share it, which is precisely the collision exact identity -# exists to forbid. -DIGEST_PATTERN = re.compile(r"^[0-9a-f]{64}$") - - -def valid_digest(value: Any) -> bool: - """True for a lowercase hex SHA-256, the only shape an identity may take.""" - return isinstance(value, str) and bool(DIGEST_PATTERN.match(value)) - - -def valid_string_list(value: Any) -> bool: - """True for the guard contract's JSON-shaped string lists.""" - return isinstance(value, list) and all( - isinstance(item, str) and bool(item.strip()) for item in value - ) - - -def guard_result(value: Any) -> dict[str, Any]: - """Return one valid artifact-free anti-ceremony decision.""" - if not isinstance(value, Mapping): - raise ValueError("anti-ceremony guard must return a mapping") - result = dict(value) - expected = { - "decision", - "reason", - "frozen_outcome", - "parked_process_work", - "remaining_proof", - "stop_condition", - } - if set(result) != expected: - raise ValueError("anti-ceremony guard returned the wrong fields") - if result["decision"] not in {"CONTINUE", "STOP"}: - raise ValueError("anti-ceremony decision must be CONTINUE or STOP") - reason = result["reason"] - if ( - not isinstance(reason, str) - or not reason.strip() - or "\n" in reason - or reason[-1] not in ".!?" - or sum(reason.count(mark) for mark in ".!?") != 1 - ): - raise ValueError("anti-ceremony reason must be exactly one sentence") - if not isinstance(result["frozen_outcome"], str) or not result["frozen_outcome"].strip(): - raise ValueError("anti-ceremony frozen_outcome must be a nonempty string") - if not valid_string_list(result["parked_process_work"]): - raise ValueError("anti-ceremony parked_process_work must be a string list") - if not valid_string_list(result["remaining_proof"]): - raise ValueError("anti-ceremony remaining_proof must be a string list") - if not isinstance(result["stop_condition"], str) or not result["stop_condition"].strip(): - raise ValueError("anti-ceremony stop_condition must be a nonempty string") - return result - - -def report( - status: str, - *, - intent_ref: str | None = None, - acceptance_digest: str | None = None, - subject_digest: str | None = None, - verdict_ref: str | None = None, - verdict_digest: str | None = None, - checked: list[str] | None = None, - not_checked: list[str] | None = None, -) -> dict[str, Any]: - return { - "schema_version": "rpi-report.v1", - "status": status, - "intent_ref": intent_ref, - "acceptance_digest": acceptance_digest, - "subject_manifest_digest": subject_digest, - "verdict_ref": verdict_ref, - "verdict_digest": verdict_digest, - "checked": checked or [], - "not_checked": not_checked or [], - } - - -def invoke_once( - intent: Any, - anti_ceremony_guard: Callable[[Any], Mapping[str, Any]], - plan_phase: Callable[[Any], Mapping[str, Any] | None], - implement_phase: Callable[[Mapping[str, Any]], Mapping[str, Any] | None], - validate_phase: Callable[[Mapping[str, Any], Mapping[str, Any]], Mapping[str, Any]], -) -> dict[str, Any]: - """Invoke the guard once, then dispatch each core phase at most once.""" - admission = guard_result(anti_ceremony_guard(intent)) - if admission["decision"] == "STOP": - return report( - "NOT_PLANNED", - checked=[f"anti-ceremony guard: STOP — {admission['reason']}"], - not_checked=["plan", "implement", "validate"], - ) - resolved_intent = plan_phase(intent) - if resolved_intent is None: - return report("NOT_PLANNED", not_checked=["implement", "validate"]) - resolved_intent = dict(resolved_intent) - intent_ref = resolved_intent.get("intent_ref") - if not isinstance(intent_ref, str) or not intent_ref: - intent_ref = "caller" - acceptance_digest = resolved_intent.get("acceptance_digest") - if not valid_digest(acceptance_digest): - raise ValueError( - "Plan must declare acceptance_digest as the SHA-256 of the exact resolved " - "intent bytes it snapshotted (validate.py snapshot-intent emits it)" - ) - - subject = implement_phase(resolved_intent) - if subject is None: - return report( - "NOT_BUILT", - intent_ref=intent_ref, - acceptance_digest=acceptance_digest, - checked=["plan"], - not_checked=["validate"], - ) - subject = dict(subject) - - validation = dict(validate_phase(resolved_intent, subject)) - status = validation.get("verdict") - if status not in {"PASS", "FAIL", "NOT_PROVEN"}: - raise ValueError("Validate must return PASS, FAIL, or NOT_PROVEN") - # Validate re-derives this from the snapshot bytes independently; equality - # here is the composed exact-identity check, not a self-comparison. - if validation.get("acceptance_digest") != acceptance_digest: - raise ValueError("Validate verdict does not match the resolved intent digest") - subject_digest = validation.get("subject_manifest_digest") - if not valid_digest(subject_digest): - raise ValueError("Validate must return the exact subject manifest digest") - candidate_digest = subject.get("subject_manifest_digest") - if candidate_digest is not None and subject_digest != candidate_digest: - raise ValueError("Validate result does not match the implemented subject digest") - author_context_id = validation.get("author_context_id") - validator_context_id = validation.get("validator_context_id") - freshness = validation.get("freshness_attestation") - if ( - not isinstance(author_context_id, str) - or not author_context_id - or not isinstance(validator_context_id, str) - or not validator_context_id - or author_context_id == validator_context_id - or not isinstance(freshness, Mapping) - or freshness.get("source") not in {"runtime", "caller"} - or not isinstance(freshness.get("attester_identity"), str) - or not freshness.get("attester_identity") - ): - raise ValueError("Validate must return distinct context identities and explicit freshness") - verdict_digest = validation.get("verdict_digest") - verdict_ref = validation.get("verdict_ref") - if (verdict_digest is None) != (verdict_ref is None): - raise ValueError("Validate must return both verdict_ref and verdict_digest when persistence is requested") - if verdict_ref is not None and ( - not isinstance(verdict_ref, str) - or not verdict_ref - or not valid_digest(verdict_digest) - ): - raise ValueError("Persisted verdict identity is invalid") - return report( - status, - intent_ref=intent_ref, - acceptance_digest=acceptance_digest, - subject_digest=subject_digest, - verdict_ref=verdict_ref, - verdict_digest=verdict_digest, - checked=list(validation.get("checked") or []), - not_checked=list(validation.get("not_checked") or []), - ) - - -# --------------------------------------------------------------------------- -# The bounded repair phase (ADR-0017) -# --------------------------------------------------------------------------- -# -# The 2026-07-14 cathedral cut removed the iterate loop together with the -# unproven compounding claim, although ADR-0011 demoted only the latter. What -# comes back is control flow, not knowledge: a repair round is admitted only -# while every condition of the convergence law holds. Byte movement and finding -# counts are identity/accounting facts, not evidence of acceptance progress. -# -# Recurrence is checked before progress. It requires causal examination by the -# caller, not an automatic claim that the design is wrong. This pure reference -# neither diagnoses causes nor dispatches a HOLD helper. - -REPAIR_ROUNDS_DEFAULT = 2 - -#: Terminal reasons `run_repair_phase` may report. `converged` is the only -#: success; the rest are law stops the caller owns the response to. -STOP_REASONS = ( - "converged", - "diversity_unsatisfied", - "repair_budget_exhausted", - "reopened_finding", - "recurring_finding_class", - "introduced_regression", - "new_finding_requires_causal_review", - "no_acceptance_progress", - "not_converged", -) - -_STATUS_RANK = {"PASS": 0, "NOT_PROVEN": 1, "FAIL": 2} - - -def _leg_status(leg: Mapping[str, Any]) -> str: - """Read a validate leg's semantic verdict under either spelling.""" - status = leg.get("status", leg.get("verdict")) - if status not in _STATUS_RANK: - raise ValueError("each validate result must report PASS, FAIL, or NOT_PROVEN") - return str(status) - - -def normalize_round(value: Any) -> dict[str, Any]: - """Fold one validation round's legs into the facts the law reasons over. - - A round is one or more validate results (the fresh validator, plus the - cross-family validator when the caller selects one). Open findings - are the UNION of the legs' stable `findings[].id`; the round's status is the - worst leg's; the digest is the subject every leg judged. - """ - legs: list[Mapping[str, Any]] - if isinstance(value, Mapping): - legs = [value] - elif isinstance(value, Sequence) and not isinstance(value, (str, bytes)): - legs = list(value) - else: - raise ValueError("a validation round must be a validate result or a list of them") - if not legs: - raise ValueError("a validation round must contain at least one validate result") - - open_findings: dict[str, dict[str, Any]] = {} - families: list[str] = [] - evidence_refs: list[dict[str, Any]] = [] - checked: list[str] = [] - not_checked: list[str] = [] - digest: Any = None - status = "PASS" - for leg in legs: - if not isinstance(leg, Mapping): - raise ValueError("each validate result must be a mapping") - leg_status = _leg_status(leg) - if _STATUS_RANK[leg_status] > _STATUS_RANK[status]: - status = leg_status - if "findings" not in leg: - raise ValueError("each validate leg must carry a findings list (empty on PASS)") - raw_findings = leg["findings"] - if not isinstance(raw_findings, (list, tuple)): - raise ValueError("findings must be a list") - leg_ids: set[str] = set() - for finding in raw_findings: - if not isinstance(finding, Mapping): - raise ValueError("each finding must be a mapping") - finding_id = finding.get("id") - if not isinstance(finding_id, str) or not finding_id.strip(): - raise ValueError("each finding must carry a stable nonempty id") - if finding_id in leg_ids: - raise ValueError(f"finding id {finding_id!r} appears twice in one validate leg") - leg_ids.add(finding_id) - if "class" in finding and ( - not isinstance(finding["class"], str) or not finding["class"].strip() - ): - raise ValueError("finding class must be a nonempty string when present") - # Wording can differ, but another leg cannot erase a class used to - # detect recurrence or silently replace it with a conflicting one. - existing = open_findings.get(finding_id, {}) - if existing.get("class") and finding.get("class") not in {None, existing["class"]}: - raise ValueError(f"finding id {finding_id!r} has conflicting classes") - open_findings[finding_id] = {**existing, **finding} - if leg_status == "PASS" and leg_ids: - raise ValueError("a PASS leg cannot carry open findings") - if leg_status == "FAIL" and not leg_ids: - raise ValueError("a FAIL leg must name at least one finding") - family = leg.get("validator_family") - if isinstance(family, str) and family and family not in families: - families.append(family) - if "evidence_refs" not in leg: - raise ValueError("each validate leg must carry an evidence_refs list (empty if none)") - raw_evidence = leg["evidence_refs"] - if not isinstance(raw_evidence, (list, tuple)): - raise ValueError("evidence_refs must be a list") - for ref in raw_evidence: - # Evidence is either a bare label (unbound; it can never admit an - # unchanged digest) or a binding {ref, subject_digest, resolves}. - if isinstance(ref, str): - entry: dict[str, Any] = {"ref": ref} - elif isinstance(ref, Mapping): - if not isinstance(ref.get("ref"), str) or not ref["ref"].strip(): - raise ValueError("each evidence binding must carry a nonempty ref") - entry = dict(ref) - # These are decoded facts from existing check receipts, not - # additional verdict.v2 fields or a persisted receipt schema. - for key in ("resolves", "preexisting", "introduced"): - ids = entry.get(key) - if ids is not None and not valid_string_list(ids): - raise ValueError(f"evidence.{key} must be a list of finding ids") - if "subject_digest" in entry and not valid_digest(entry["subject_digest"]): - raise ValueError("evidence.subject_digest must be a valid digest") - else: - raise ValueError("each evidence ref must be a string or a binding mapping") - existing_evidence = next((e for e in evidence_refs if e["ref"] == entry["ref"]), None) - if existing_evidence is None: - evidence_refs.append(entry) - elif existing_evidence != entry: - raise ValueError(f"evidence ref {entry['ref']!r} has conflicting bindings") - leg_digest = leg.get("subject_digest", leg.get("subject_manifest_digest")) - if not valid_digest(leg_digest): - raise ValueError("each validate leg must carry a valid subject digest") - if digest is not None and leg_digest != digest: - raise ValueError("validate legs disagree about the subject digest") - digest = leg_digest - for key, sink in (("checked", checked), ("not_checked", not_checked)): - items = leg.get(key, []) - if not valid_string_list(items): - raise ValueError(f"{key} must be a list of strings") - sink.extend(items) - # Each required leg must carry its own visible proof surface; a peer's - # receipts cannot repair a deficient PASS. Exact identities and all - # criterion proofs remain Validate's upstream contract, not a claim - # that this pure reference attested or re-executed them. - if leg_status == "PASS" and ( - leg.get("not_checked") or not leg.get("checked") or not raw_evidence - ) and status == "PASS": - status = "NOT_PROVEN" - - return { - "status": status, - "open_findings": list(open_findings.values()), - "open_ids": set(open_findings), - "subject_digest": digest, - "evidence_refs": evidence_refs, - "families": families, - "checked": checked, - "not_checked": not_checked, - } - - -def law_violation( - previous: Mapping[str, Any], - current: Mapping[str, Any], - closed_ids: set[str], - closed_classes: set[str] | None = None, -) -> str | None: - """Return the violated convergence-law condition, or None when all hold. - - Condition 1 (the caller's `repair_rounds`) is a precondition on admission - and is checked by `run_repair_phase` before a round is consumed; conditions - 2 and 3 are properties of the round that was produced. The existing receipt - binding names a gap that a fresh judge actually closed; the function does - not prove that closure or infer a new finding's cause from prose or counts. - """ - reopened = current["open_ids"] & closed_ids - if reopened: - return "reopened_finding" - current_classes = {f.get("class") for f in current["open_findings"] if f.get("class")} - if current_classes & (closed_classes or set()): - return "recurring_finding_class" - new_ids = current["open_ids"] - previous["open_ids"] - introduced = { - fid - for evidence in current["evidence_refs"] - if evidence.get("subject_digest") == current["subject_digest"] - for fid in evidence.get("introduced", []) - } - if introduced & current["open_ids"]: - return "introduced_regression" - preexisting = { - fid - for evidence in current["evidence_refs"] - if evidence.get("subject_digest") == previous["subject_digest"] - for fid in evidence.get("preexisting", []) - } - if new_ids - preexisting: - return "new_finding_requires_causal_review" - previous_refs = {e["ref"] for e in previous["evidence_refs"]} - resolved = previous["open_ids"] - current["open_ids"] - # New digest-bound evidence is required even when the bytes or count moved. - # It names a gap actually closed this round, not a renamed or still-open - # finding. New findings remain visible; their count is not a regression - # diagnosis. Fresh judgment must establish acceptance relevance and cause. - binding_evidence = [ - e - for e in current["evidence_refs"] - if e["ref"] not in previous_refs - and e.get("subject_digest") == current["subject_digest"] - and resolved & set(e.get("resolves") or []) - ] - if binding_evidence and ( - current["subject_digest"] != previous["subject_digest"] - or (previous["status"] == "NOT_PROVEN" and current["status"] != "FAIL") - ): - return None - return "no_acceptance_progress" - - -def run_repair_phase( - validations: Sequence[Any], - *, - repair_rounds: int = REPAIR_ROUNDS_DEFAULT, - risky_surface: bool = False, - cross_model: bool = False, - intent_ref: str | None = None, - acceptance_digest: str | None = None, - verdict_ref: str | None = None, - verdict_digest: str | None = None, -) -> dict[str, Any]: - """Walk already-produced validation rounds under the convergence law. - - `validations[0]` is the traversal's first fresh validation; every later - element is a repair round the orchestrator produced after fixing findings. - ``cross_model`` requires a second family; risk alone does not select it. - ``risky_surface`` remains an accepted compatibility hint with no effect on - family selection. This pure reference consumes declared family facts; it - does not dispatch models or attest fresh context identities. - - Returns a mapping with: - - - ``report``: the exact nine-key `rpi-report.v1` object. `checked` opens - with one `repair round N: k open findings` line per round; open findings - never enter `not_checked`, which keeps its meaning (unverified in-scope - acceptance). - - ``open_findings``: the findings still open at the stop, deduplicated by id. - - ``rounds_used``: repair rounds actually spent (the first validation is - round 0 and spends none). - - ``stop_reason``: one of :data:`STOP_REASONS`. - """ - if not validations: - raise ValueError("the repair phase needs at least one validation round") - if not isinstance(repair_rounds, int) or isinstance(repair_rounds, bool) or repair_rounds < 0: - raise ValueError("repair_rounds must be a non-negative integer") - - checked: list[str] = [] - closed_ids: set[str] = set() - closed_classes: set[str] = set() - rounds_used = 0 - current = normalize_round(validations[0]) - previous = current - stop_reason = "not_converged" - law_stopped = False - - for index, raw_candidate in enumerate(validations): - if index > 0: - # Condition 1: the caller's bound, checked before the round is even - # normalized, so a round past the bound is never consumed. - if rounds_used >= repair_rounds: - stop_reason = "repair_budget_exhausted" - break - candidate = normalize_round(raw_candidate) - rounds_used += 1 - current = candidate - checked.append( - f"repair round {rounds_used}: {len(current['open_ids'])} open findings" - ) - violation = law_violation(previous, current, closed_ids, closed_classes) - if violation is not None: - stop_reason = violation - law_stopped = True - break - closed_ids |= previous["open_ids"] - current["open_ids"] - # A class closes only when none of its findings remain open. A - # later return, even under a new id, warrants causal examination. - closed_classes |= ( - {f.get("class") for f in previous["open_findings"] if f.get("class")} - - {f.get("class") for f in current["open_findings"] if f.get("class")} - ) - else: - checked.append(f"repair round 0: {len(current['open_ids'])} open findings") - - converged, reason = _converged(current, cross_model) - previous = current - if converged: - stop_reason = "converged" - break - if reason is not None: - stop_reason = reason - break - else: - stop_reason = "not_converged" - - if stop_reason == "not_converged" and rounds_used >= repair_rounds and current["open_ids"]: - # Findings remain and the caller's bound is spent: name it as such. - stop_reason = "repair_budget_exhausted" - - status = current["status"] - if stop_reason == "diversity_unsatisfied" or (law_stopped and status == "PASS"): - # A PASS produced by a law-violating round cannot certify anything: a - # PASS over unchanged bytes after a FAIL is a flip, not a proof. A FAIL - # that also broke the law stays a FAIL; the subject is still wrong. - status = "NOT_PROVEN" - - return { - "report": report( - status, - intent_ref=intent_ref, - acceptance_digest=acceptance_digest, - subject_digest=current["subject_digest"], - verdict_ref=verdict_ref, - verdict_digest=verdict_digest, - checked=checked + current["checked"], - not_checked=list(current["not_checked"]), - ), - "open_findings": list(current["open_findings"]), - "rounds_used": rounds_used, - "stop_reason": stop_reason, - } - - -def _converged(current: Mapping[str, Any], cross_model: bool) -> tuple[bool, str | None]: - """Converged ⇔ fresh PASS, plus a cross-family PASS when selected.""" - if current["status"] != "PASS": - return False, None - if cross_model and len(current["families"]) < 2: - # Fresh same-family judgment is valid by default, but cannot satisfy - # an explicitly selected second family. - return False, "diversity_unsatisfied" - return True, None diff --git a/skills/rpi/scripts/validate.sh b/skills/rpi/scripts/validate.sh index 6411b8add..9cbd9e47c 100755 --- a/skills/rpi/scripts/validate.sh +++ b/skills/rpi/scripts/validate.sh @@ -11,14 +11,14 @@ grep -Fq 'Acceptance changes need caller authority.' "$skill_dir/SKILL.md" grep -Fq 'at most one bounded' "$skill_dir/SKILL.md" grep -Fq 'reviewers remain required.' "$skill_dir/SKILL.md" grep -Fq 'author-distinct' "$skill_dir/SKILL.md" -grep -Fq 'no fixed ten-minute cap' "$skill_dir/SKILL.md" +grep -Fq 'A repair does not start another' "$skill_dir/SKILL.md" grep -Fq 'empty' "$skill_dir/SKILL.md" grep -Fq 'When no machine' "$skill_dir/SKILL.md" if grep -Eq 'Plan is closed for that intent|dependencies:.*anti-ceremony|plan_packet_digest' "$skill_dir/SKILL.md"; then echo 'rpi retains a retired phase lock, mandatory specialist or planning packet' >&2 exit 1 fi -for ref in boundaries bounded-adapter outer-goal; do +for ref in boundaries outer-goal; do test -s "$skill_dir/references/$ref.md" done echo 'rpi lean skill contract: PASS' diff --git a/skills/rpi/tests/test_run_once.py b/skills/rpi/tests/test_run_once.py deleted file mode 100644 index 62841f566..000000000 --- a/skills/rpi/tests/test_run_once.py +++ /dev/null @@ -1,870 +0,0 @@ -from __future__ import annotations - -import hashlib -import importlib.util -from pathlib import Path -import tempfile -import unittest - - -MODULE_PATH = Path(__file__).parents[1] / "scripts" / "run_once.py" -SPEC = importlib.util.spec_from_file_location("rpi_run_once", MODULE_PATH) -assert SPEC and SPEC.loader -MODULE = importlib.util.module_from_spec(SPEC) -SPEC.loader.exec_module(MODULE) - -# The unchanged Validate reference oracle, not a stand-in. The composed-contract test -# below drives RPI against this module's actual identity functions; that is the -# only shape that can catch a disagreement between the two skills, which is -# exactly the defect that hid here (both suites green over a broken contract -# because the fake Validate borrowed RPI's own digest function). -VALIDATE_PATH = Path(__file__).parents[2] / "validate" / "tests" / "validate.py" -VALIDATE_SPEC = importlib.util.spec_from_file_location("ao_validate", VALIDATE_PATH) -assert VALIDATE_SPEC and VALIDATE_SPEC.loader -VALIDATE = importlib.util.module_from_spec(VALIDATE_SPEC) -VALIDATE_SPEC.loader.exec_module(VALIDATE) - -# A literal, independently written digest. Fakes must never derive an expected -# identity by calling the code under test — that is how the original defect -# stayed invisible. -INTENT_DIGEST = "c" * 64 - - -def validation_round( - status, - finding_ids, - *, - digest="a" * 64, - evidence=("acceptance-receipt",), - family="fresh", - summaries=None, - checked=("acceptance",), - not_checked=(), -): - """One validate leg's result, as the repair phase consumes it (pure data).""" - summaries = summaries or {} - return { - "status": status, - "findings": [ - {"id": fid, "summary": summaries.get(fid, f"finding {fid}")} - for fid in finding_ids - ], - "subject_digest": digest, - "evidence_refs": list(evidence), - "validator_family": family, - "checked": list(checked), - "not_checked": list(not_checked), - } - - -def continue_guard(intent): - return { - "decision": "CONTINUE", - "reason": "The frozen outcome still requires implementation proof.", - "frozen_outcome": str(intent), - "parked_process_work": [], - "remaining_proof": ["implementation", "fresh validation"], - "stop_condition": "Stop after one fresh validation result.", - } - - -class RunOnceTests(unittest.TestCase): - def phases(self, verdict: str = "PASS"): - calls: list[str] = [] - - def plan(intent): - calls.append("plan") - return { - "intent_ref": "bead:agentops-test", - "intent": intent, - "acceptance": ["works"], - "acceptance_digest": INTENT_DIGEST, - } - - def implement(_plan): - calls.append("implement") - return {"subject_manifest_digest": "a" * 64, "checks": ["focused"]} - - def validate(_plan, _candidate): - calls.append("validate") - return { - "verdict": verdict, - "acceptance_digest": INTENT_DIGEST, - "subject_manifest_digest": "a" * 64, - "author_context_id": "author-ctx", - "validator_context_id": "validator-ctx", - "freshness_attestation": { - "source": "runtime", - "attester_identity": "runtime:rpi-test", - }, - "verdict_digest": "b" * 64, - "verdict_ref": "/tmp/verdict.json", - "checked": ["acceptance"], - "not_checked": [], - } - - return calls, plan, implement, validate - - def test_anti_ceremony_guard_runs_once_before_plan(self): - calls, plan, implement, validate = self.phases() - - def anti_ceremony(intent): - calls.append("anti-ceremony") - return { - "decision": "CONTINUE", - "reason": "The frozen outcome still requires implementation proof.", - "frozen_outcome": intent, - "parked_process_work": [], - "remaining_proof": ["implementation", "fresh validation"], - "stop_condition": "Stop after one fresh validation result.", - } - - result = MODULE.invoke_once( - "intent", - anti_ceremony, - plan, - implement, - validate, - ) - - self.assertEqual( - calls, - ["anti-ceremony", "plan", "implement", "validate"], - ) - self.assertEqual(result["status"], "PASS") - - def test_anti_ceremony_stop_dispatches_no_core_phase(self): - calls: list[str] = [] - reason = "The proposed traversal would create only process artifacts." - - def anti_ceremony(_intent): - calls.append("anti-ceremony") - return { - "decision": "STOP", - "reason": reason, - "frozen_outcome": "Ship the already-proved caller outcome", - "parked_process_work": ["another plan", "another audit"], - "remaining_proof": [], - "stop_condition": "Stop before Plan.", - } - - result = MODULE.invoke_once( - "intent", - anti_ceremony, - lambda _intent: calls.append("plan"), - lambda _plan: calls.append("implement"), - lambda _plan, _candidate: calls.append("validate"), - ) - - self.assertEqual(calls, ["anti-ceremony"]) - self.assertEqual(result["status"], "NOT_PLANNED") - self.assertEqual( - result["checked"], - [f"anti-ceremony guard: STOP — {reason}"], - ) - self.assertEqual(result["not_checked"], ["plan", "implement", "validate"]) - - def test_each_phase_runs_once_and_pass_reports(self): - calls, plan, implement, validate = self.phases() - result = MODULE.invoke_once("intent", continue_guard, plan, implement, validate) - self.assertEqual(calls, ["plan", "implement", "validate"]) - self.assertEqual(result["status"], "PASS") - self.assertEqual(result["intent_ref"], "bead:agentops-test") - self.assertEqual(result["acceptance_digest"], INTENT_DIGEST) - self.assertNotIn("next_action", result) - - def test_fail_from_one_experiment_feeds_the_repair_phase(self): - """Replaces the old stop-on-FAIL test (ADR-0017). - - One experiment still dispatches Plan and Implement exactly once, and the - FAIL it produces is no longer terminal by itself: it is the first round - handed to the bounded repair phase, which owns the stop decision. - """ - calls, plan, implement, validate = self.phases("FAIL") - result = MODULE.invoke_once("intent", continue_guard, plan, implement, validate) - self.assertEqual(calls, ["plan", "implement", "validate"]) - self.assertEqual(result["status"], "FAIL") - - outcome = MODULE.run_repair_phase( - [ - validation_round("FAIL", ["f1"], digest="a" * 64), - validation_round("PASS", [], digest="d" * 64, evidence=( - {"ref": "fixed-f1", "subject_digest": "d" * 64, "resolves": ["f1"]}, - )), - ], - repair_rounds=2, - intent_ref=result["intent_ref"], - acceptance_digest=result["acceptance_digest"], - ) - self.assertEqual(outcome["stop_reason"], "converged") - self.assertEqual(outcome["report"]["status"], "PASS") - self.assertEqual(outcome["rounds_used"], 1) - - def test_fresh_validation_does_not_require_persisted_verdict(self): - calls, plan, implement, validate = self.phases() - - def inline_result(resolved, subject): - result = validate(resolved, subject) - result.pop("verdict_digest") - result.pop("verdict_ref") - return result - - result = MODULE.invoke_once("intent", continue_guard, plan, implement, inline_result) - - self.assertEqual(calls, ["plan", "implement", "validate"]) - self.assertEqual(result["status"], "PASS") - self.assertEqual(result["subject_manifest_digest"], "a" * 64) - self.assertIsNone(result["verdict_ref"]) - self.assertIsNone(result["verdict_digest"]) - - def test_fresh_validation_requires_distinct_contexts_and_attestation(self): - _calls, plan, implement, validate = self.phases() - - for field, value in ( - ("validator_context_id", "author-ctx"), - ("freshness_attestation", None), - ): - with self.subTest(field=field): - def invalid(resolved, subject, field=field, value=value): - result = validate(resolved, subject) - result[field] = value - return result - - with self.assertRaisesRegex(ValueError, "distinct context identities"): - MODULE.invoke_once("intent", continue_guard, plan, implement, invalid) - - def test_missing_plan_stops_before_implement(self): - calls: list[str] = [] - result = MODULE.invoke_once( - "intent", - continue_guard, - lambda _intent: None, - lambda _plan: calls.append("implement"), - lambda _plan, _candidate: calls.append("validate"), - ) - self.assertEqual(calls, []) - self.assertEqual(result["status"], "NOT_PLANNED") - - def test_missing_candidate_stops_before_validate(self): - calls: list[str] = [] - result = MODULE.invoke_once( - "intent", - continue_guard, - lambda _intent: { - "intent_ref": "caller", - "acceptance": ["works"], - "acceptance_digest": INTENT_DIGEST, - }, - lambda _plan: None, - lambda _plan, _candidate: calls.append("validate"), - ) - self.assertEqual(calls, []) - self.assertEqual(result["status"], "NOT_BUILT") - - def test_validate_cannot_report_a_different_intent(self): - calls, plan, implement, validate = self.phases() - - def mismatched(resolved, subject): - result = validate(resolved, subject) - result["acceptance_digest"] = "f" * 64 - return result - - with self.assertRaisesRegex(ValueError, "resolved intent digest"): - MODULE.invoke_once("intent", continue_guard, plan, implement, mismatched) - - def test_plan_without_a_declared_digest_is_a_contract_error(self): - _calls, _plan, implement, validate = self.phases() - - def undeclared(_intent): - return {"intent_ref": "caller", "acceptance": ["works"]} - - with self.assertRaisesRegex(ValueError, "acceptance_digest"): - MODULE.invoke_once("intent", continue_guard, undeclared, implement, validate) - - def test_plan_digest_must_be_a_sha256(self): - _calls, _plan, implement, validate = self.phases() - - for bogus in ("", "not-a-digest", "C" * 64, "a" * 63, 12345): - with self.subTest(digest=bogus): - def undeclared(_intent, value=bogus): - return { - "intent_ref": "caller", - "acceptance": ["works"], - "acceptance_digest": value, - } - - with self.assertRaisesRegex(ValueError, "acceptance_digest"): - MODULE.invoke_once("intent", continue_guard, undeclared, implement, validate) - - -class RepairPhaseTests(unittest.TestCase): - """The bounded repair phase and its convergence law (ADR-0017). - - RPI is no longer single-pass: a `FAIL` or `NOT_PROVEN` with findings may be - repaired and re-validated while acceptance progress and the bound hold. The loop is - modelled here as pure data — a sequence of already-produced validate rounds - — so the stop semantics are executable without Git, `ao`, or a tracker. - """ - - def repair(self, rounds, **kwargs): - kwargs.setdefault("intent_ref", "bead:agentops-test") - kwargs.setdefault("acceptance_digest", INTENT_DIGEST) - return MODULE.run_repair_phase(rounds, **kwargs) - - def test_repair_rounds_zero_with_findings_is_budget_exhausted(self): - outcome = self.repair([validation_round("FAIL", ["f1"])], repair_rounds=0) - self.assertEqual(outcome["stop_reason"], "repair_budget_exhausted") - self.assertEqual(outcome["rounds_used"], 0) - self.assertEqual(outcome["report"]["status"], "FAIL") - - def test_a_pass_over_unchanged_bytes_after_a_fail_is_a_flip_not_a_proof(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1"], digest="a" * 64), - validation_round("PASS", [], digest="a" * 64), - ] - ) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - self.assertEqual(outcome["report"]["status"], "NOT_PROVEN") - - def test_new_evidence_must_resolve_a_prior_finding_to_admit_an_unchanged_digest(self): - outcome = self.repair( - [ - validation_round("NOT_PROVEN", ["gap"], digest="a" * 64), - validation_round( - "NOT_PROVEN", ["gap"], digest="a" * 64, - evidence=({"ref": "receipt-2", "subject_digest": "a" * 64, "resolves": ["gap"]},), - ), - ] - ) - # "resolves" claims gap, but gap is still open: nothing was resolved. - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - - def test_a_bare_new_evidence_label_does_not_admit_an_unchanged_digest(self): - outcome = self.repair( - [ - validation_round("NOT_PROVEN", ["gap", "other"], digest="a" * 64), - validation_round("NOT_PROVEN", ["other"], digest="a" * 64, evidence=("receipt-2",)), - ] - ) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - - def test_evidence_bound_to_another_digest_does_not_admit(self): - outcome = self.repair( - [ - validation_round("NOT_PROVEN", ["gap", "other"], digest="a" * 64), - validation_round( - "NOT_PROVEN", ["other"], digest="a" * 64, - evidence=({"ref": "receipt-2", "subject_digest": "b" * 64, "resolves": ["gap"]},), - ), - ] - ) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - - def test_missing_findings_or_evidence_keys_are_rejected(self): - for key in ("findings", "evidence_refs"): - with self.subTest(missing=key): - bad = validation_round("PASS", []) - del bad[key] - with self.assertRaisesRegex(ValueError, key): - self.repair([bad]) - bad = validation_round("PASS", []) - bad["checked"] = "acceptance" - with self.assertRaisesRegex(ValueError, "checked must be a list"): - self.repair([bad]) - - def test_scalar_evidence_refs_are_rejected(self): - bad = validation_round("FAIL", ["f1"]) - bad["evidence_refs"] = "receipt-1" - with self.assertRaisesRegex(ValueError, "evidence_refs must be a list"): - self.repair([bad]) - - def test_new_evidence_does_not_admit_a_current_fail_over_unchanged_bytes(self): - outcome = self.repair( - [ - validation_round("NOT_PROVEN", ["gap", "bug"], digest="a" * 64), - validation_round("FAIL", ["bug"], digest="a" * 64, evidence=("receipt-2",)), - ] - ) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - self.assertEqual(outcome["report"]["status"], "FAIL") - - def test_malformed_rounds_are_rejected_not_swallowed(self): - cases = { - "pass with findings": validation_round("PASS", ["f1"]), - "fail without findings": validation_round("FAIL", []), - "missing digest": validation_round("FAIL", ["f1"], digest=None), - } - for name, bad in cases.items(): - with self.subTest(case=name): - with self.assertRaises(ValueError): - self.repair([bad]) - duplicate = validation_round("FAIL", ["f1"]) - duplicate["findings"].append({"id": "f1", "summary": "again"}) - with self.assertRaisesRegex(ValueError, "twice"): - self.repair([duplicate]) - - def test_rounds_past_the_bound_are_never_normalized(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1", "f2"], digest="a" * 64), - validation_round("FAIL", ["f1"], digest="b" * 64, evidence=( - {"ref": "fixed-f2", "subject_digest": "b" * 64, "resolves": ["f2"]}, - )), - {"status": "garbage-that-would-raise"}, - ], - repair_rounds=1, - ) - self.assertEqual(outcome["stop_reason"], "repair_budget_exhausted") - - def test_a_first_round_pass_converges_without_spending_a_repair_round(self): - outcome = self.repair([validation_round("PASS", [])]) - self.assertEqual(outcome["stop_reason"], "converged") - self.assertEqual(outcome["report"]["status"], "PASS") - self.assertEqual(outcome["rounds_used"], 0) - self.assertEqual(outcome["open_findings"], []) - self.assertEqual( - outcome["report"]["checked"][0], "repair round 0: 0 open findings" - ) - - def test_repair_stops_at_the_declared_repair_rounds_budget(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1", "f2"], digest="a" * 64), - validation_round("FAIL", ["f1"], digest="b" * 64, evidence=( - {"ref": "fixed-f2", "subject_digest": "b" * 64, "resolves": ["f2"]}, - )), - validation_round("FAIL", ["f1"], digest="c" * 64), - ], - repair_rounds=1, - ) - self.assertEqual(outcome["stop_reason"], "repair_budget_exhausted") - self.assertEqual(outcome["rounds_used"], 1) - self.assertEqual(outcome["report"]["status"], "FAIL") - self.assertEqual( - outcome["report"]["checked"][:2], - ["repair round 0: 2 open findings", "repair round 1: 1 open findings"], - ) - - def test_new_finding_with_unknown_cause_requires_causal_review(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1"], digest="a" * 64), - validation_round("FAIL", ["f1", "f2"], digest="b" * 64), - ] - ) - self.assertEqual(outcome["stop_reason"], "new_finding_requires_causal_review") - self.assertEqual(outcome["report"]["status"], "FAIL") - self.assertEqual(outcome["rounds_used"], 1) - self.assertEqual( - sorted(f["id"] for f in outcome["open_findings"]), ["f1", "f2"] - ) - - def test_repair_stops_when_a_closed_finding_reopens(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1", "f2"], digest="a" * 64), - validation_round("FAIL", ["f1"], digest="b" * 64, evidence=( - {"ref": "fixed-f2", "subject_digest": "b" * 64, "resolves": ["f2"]}, - )), - validation_round("FAIL", ["f2"], digest="c" * 64), - ] - ) - self.assertEqual(outcome["stop_reason"], "reopened_finding") - self.assertEqual(outcome["rounds_used"], 2) - self.assertEqual(outcome["report"]["status"], "FAIL") - - def test_digest_and_count_movement_without_bound_proof_are_not_progress(self): - for remaining in (["f1", "f2"], ["f1"], []): - with self.subTest(remaining=remaining): - outcome = self.repair([ - validation_round("FAIL", ["f1", "f2"], digest="a" * 64), - validation_round("FAIL" if remaining else "PASS", remaining, - digest="b" * 64, evidence=("another-check-label",)), - ]) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - self.assertNotEqual(outcome["report"]["status"], "PASS") - - def test_discovered_preexisting_defects_may_grow_count_with_real_progress(self): - outcome = self.repair([ - validation_round("FAIL", ["fixed"], digest="a" * 64), - validation_round("FAIL", ["discovered-1", "discovered-2"], digest="b" * 64, - evidence=( - {"ref": "fixed-check", "subject_digest": "b" * 64, "resolves": ["fixed"]}, - {"ref": "reproduced-on-prior", "subject_digest": "a" * 64, - "preexisting": ["discovered-1", "discovered-2"]}, - )), - ]) - self.assertEqual(outcome["stop_reason"], "not_converged") - self.assertEqual(outcome["report"]["status"], "FAIL") - self.assertEqual(len(outcome["open_findings"]), 2) - self.assertEqual(outcome["rounds_used"], 1) - - def test_new_finding_cannot_hide_behind_another_resolved_gap(self): - for proof in ((), ({"ref": "wrong-baseline", "subject_digest": "c" * 64, - "preexisting": ["new"]},)): - with self.subTest(proof=proof): - outcome = self.repair([ - validation_round("FAIL", ["fixed"], digest="a" * 64), - validation_round("FAIL", ["new"], digest="b" * 64, evidence=( - {"ref": "fixed-check", "subject_digest": "b" * 64, "resolves": ["fixed"]}, - ) + proof), - ]) - self.assertEqual(outcome["stop_reason"], "new_finding_requires_causal_review") - - def test_introduced_regression_stops_even_if_another_gap_closed(self): - outcome = self.repair([ - validation_round("FAIL", ["fixed"], digest="a" * 64), - validation_round("FAIL", ["regression"], digest="b" * 64, evidence=( - {"ref": "fixed-check", "subject_digest": "b" * 64, "resolves": ["fixed"]}, - {"ref": "before-after-check", "subject_digest": "b" * 64, - "introduced": ["regression"]}, - )), - ]) - self.assertEqual(outcome["stop_reason"], "introduced_regression") - self.assertEqual(outcome["report"]["status"], "FAIL") - - def test_recurring_class_requires_causal_review_without_design_diagnosis(self): - first = validation_round("FAIL", ["f1", "f2"], digest="a" * 64) - first["findings"][0]["class"] = "deadline-bypass" - recurrence = validation_round("FAIL", ["new-id"], digest="c" * 64) - recurrence["findings"][0]["class"] = "deadline-bypass" - outcome = self.repair([ - first, - validation_round("FAIL", ["f2"], digest="b" * 64, evidence=( - {"ref": "fixed-f1", "subject_digest": "b" * 64, "resolves": ["f1"]}, - )), - recurrence, - ]) - self.assertEqual(outcome["stop_reason"], "recurring_finding_class") - self.assertEqual(outcome["rounds_used"], 2) - self.assertNotIn("design", str(outcome)) - - def test_old_receipt_does_not_prove_new_acceptance_progress(self): - receipt = {"ref": "old-check", "subject_digest": "b" * 64, "resolves": ["f2"]} - outcome = self.repair([ - validation_round("FAIL", ["f1", "f2"], digest="a" * 64, evidence=(receipt,)), - validation_round("FAIL", ["f1"], digest="b" * 64, evidence=(receipt,)), - ]) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - - def test_peer_leg_cannot_silently_override_a_receipts_causal_binding(self): - with self.assertRaisesRegex(ValueError, "conflicting bindings"): - self.repair([[ - validation_round("FAIL", ["f1"], evidence=( - {"ref": "comparison", "subject_digest": "a" * 64, "preexisting": ["f1"]}, - )), - validation_round("FAIL", ["f1"], family="other", evidence=( - {"ref": "comparison", "subject_digest": "a" * 64, "introduced": ["f1"]}, - )), - ]]) - - def test_peer_leg_cannot_silently_replace_a_recurrence_class(self): - first = validation_round("FAIL", ["f1"]) - first["findings"][0]["class"] = "deadline-bypass" - peer = validation_round("FAIL", ["f1"], family="other") - self.assertEqual(MODULE.normalize_round([first, peer])["open_findings"][0]["class"], - "deadline-bypass") - peer["findings"][0]["class"] = "cosmetic" - with self.assertRaisesRegex(ValueError, "conflicting classes"): - self.repair([[first, peer]]) - - def test_pass_cannot_converge_with_unverified_acceptance_or_missing_proof(self): - cases = ( - validation_round("PASS", [], not_checked=("required-cancellation-case",)), - validation_round("PASS", [], checked=()), - validation_round("PASS", [], evidence=()), - ) - for raw in cases: - with self.subTest(raw=raw): - for round_value in (raw, [raw, validation_round("PASS", [], family="other")]): - outcome = self.repair([round_value]) - self.assertEqual(outcome["report"]["status"], "NOT_PROVEN") - self.assertNotEqual(outcome["stop_reason"], "converged") - self.assertEqual(outcome["report"]["not_checked"], raw["not_checked"]) - - def test_repair_stops_when_the_digest_is_unchanged_and_no_new_evidence(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1"]), - validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1"]), - ] - ) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - self.assertEqual(outcome["rounds_used"], 1) - - def test_not_proven_is_resolved_by_new_evidence_with_an_unchanged_digest(self): - outcome = self.repair( - [ - validation_round( - "NOT_PROVEN", ["gap1"], digest="a" * 64, evidence=["r1"] - ), - validation_round( - "PASS", [], digest="a" * 64, - evidence=["r1", {"ref": "r2", "subject_digest": "a" * 64, "resolves": ["gap1"]}], - ), - ] - ) - self.assertEqual(outcome["stop_reason"], "converged") - self.assertEqual(outcome["report"]["status"], "PASS") - self.assertEqual(outcome["rounds_used"], 1) - self.assertEqual(outcome["report"]["subject_manifest_digest"], "a" * 64) - - def test_new_evidence_does_not_rescue_a_fail_round(self): - """Evidence alone cannot repair an unchanged subject already judged FAIL. - - A FAIL means the subject is wrong; both a changed subject and proven - acceptance progress are needed, not an extra unbound evidence label. - """ - outcome = self.repair( - [ - validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1"]), - validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1", "r2"]), - ] - ) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - - def test_a_reworded_summary_with_the_same_id_is_the_same_finding(self): - outcome = self.repair( - [ - validation_round( - "FAIL", ["f1"], digest="a" * 64, summaries={"f1": "gate fails"} - ), - validation_round( - "FAIL", - ["f1"], - digest="b" * 64, - summaries={"f1": "the deterministic gate still rejects the tree"}, - ), - ] - ) - self.assertEqual(outcome["rounds_used"], 1) - self.assertEqual(outcome["stop_reason"], "no_acceptance_progress") - self.assertEqual([f["id"] for f in outcome["open_findings"]], ["f1"]) - self.assertEqual( - outcome["open_findings"][0]["summary"], - "the deterministic gate still rejects the tree", - ) - - def test_open_findings_are_the_union_of_fresh_and_cross_family_ids(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1", "f2", "f3"], digest="a" * 64), - [ - validation_round("FAIL", ["f1"], digest="b" * 64, family="fresh", evidence=( - {"ref": "fixed-f3", "subject_digest": "b" * 64, "resolves": ["f3"]}, - )), - validation_round("FAIL", ["f2"], digest="b" * 64, family="codex"), - ], - ] - ) - self.assertEqual(outcome["rounds_used"], 1) - self.assertEqual(sorted(f["id"] for f in outcome["open_findings"]), ["f1", "f2"]) - self.assertEqual( - outcome["report"]["checked"][1], "repair round 1: 2 open findings" - ) - - def test_a_generated_only_change_needs_proven_acceptance_progress(self): - outcome = self.repair( - [ - validation_round("FAIL", ["f1"], digest="a" * 64), - validation_round("PASS", [], digest="b" * 64, evidence=( - {"ref": "projection-parity", "subject_digest": "b" * 64, "resolves": ["f1"]}, - )), - ] - ) - self.assertEqual(outcome["stop_reason"], "converged") - self.assertEqual(outcome["rounds_used"], 1) - self.assertEqual( - outcome["report"]["checked"][:2], - ["repair round 0: 1 open findings", "repair round 1: 0 open findings"], - ) - - def test_open_findings_never_land_in_not_checked(self): - outcome = self.repair( - [ - validation_round( - "FAIL", - ["f1"], - digest="a" * 64, - not_checked=["edge case acceptance"], - ) - ], - repair_rounds=0, - ) - self.assertEqual(outcome["report"]["not_checked"], ["edge case acceptance"]) - self.assertEqual([f["id"] for f in outcome["open_findings"]], ["f1"]) - self.assertNotIn("finding f1", outcome["report"]["not_checked"]) - - def test_a_risky_surface_uses_one_fresh_family_by_default(self): - for family in ("codex", "claude"): - with self.subTest(family=family): - single = self.repair( - [validation_round("PASS", [], family=family)], risky_surface=True - ) - self.assertEqual(single["stop_reason"], "converged") - self.assertEqual(single["report"]["status"], "PASS") - - def test_selected_cross_model_review_needs_a_second_family(self): - single = self.repair( - [validation_round("PASS", [], family="claude")], cross_model=True - ) - self.assertEqual(single["stop_reason"], "diversity_unsatisfied") - self.assertEqual(single["report"]["status"], "NOT_PROVEN") - - same_family = self.repair( - [[validation_round("PASS", [], family="claude"), - validation_round("PASS", [], family="claude")]], - cross_model=True, - ) - self.assertEqual(same_family["stop_reason"], "diversity_unsatisfied") - self.assertEqual(same_family["report"]["status"], "NOT_PROVEN") - - crossed = self.repair( - [ - [ - validation_round("PASS", [], family="fresh"), - validation_round("PASS", [], family="codex"), - ] - ], - cross_model=True, - ) - self.assertEqual(crossed["stop_reason"], "converged") - self.assertEqual(crossed["report"]["status"], "PASS") - - def test_selected_cross_model_disagreement_keeps_the_failure(self): - split = self.repair( - [[validation_round("PASS", [], family="claude"), - validation_round("FAIL", ["f1"], family="codex")]], - cross_model=True, - ) - self.assertEqual(split["report"]["status"], "FAIL") - self.assertEqual([f["id"] for f in split["open_findings"]], ["f1"]) - - def test_the_report_keeps_the_nine_key_rpi_report_shape(self): - outcome = self.repair([validation_round("PASS", [])]) - self.assertEqual( - sorted(outcome["report"]), - sorted( - [ - "schema_version", - "status", - "intent_ref", - "acceptance_digest", - "subject_manifest_digest", - "verdict_ref", - "verdict_digest", - "checked", - "not_checked", - ] - ), - ) - self.assertEqual(outcome["report"]["schema_version"], "rpi-report.v1") - - def test_the_repair_phase_needs_at_least_one_validation_round(self): - with self.assertRaisesRegex(ValueError, "at least one validation round"): - self.repair([]) - - -class ComposedIdentityContractTests(unittest.TestCase): - """RPI against the REAL Validate identity functions. - - This is the test the defect needed. RPI used to digest a canonical-JSON - re-serialization of the parsed intent mapping while Validate digested the raw - intent bytes, and RPI hard-compared the two. Nothing caught it because RPI's - own suite mocked Validate with RPI's digest function, so the mock agreed with - the code under test by construction. Here the digest crosses the skill - boundary in both directions with no shared helper. - """ - - # Deliberately NOT canonical JSON: real intent sources have indentation, - # trailing newlines, and key order. A canonical-JSON digest of the parsed - # mapping differs from sha256(these bytes), so this payload discriminates - # between the two implementations instead of accidentally agreeing. - INTENT_BYTES = b'{\n "acceptance": ["works"],\n "intent_ref": "bead:agentops-test"\n}\n' - - def test_rpi_carries_the_digest_validate_derives_from_the_snapshot_bytes(self): - with tempfile.TemporaryDirectory() as tmp: - intent_dir = Path(tmp) / "intents" - - def plan(_intent): - # Plan resolves the intent and snapshots the EXACT bytes through - # Validate's own store, which is what defines the identity. - path, _existed = VALIDATE.snapshot_intent(self.INTENT_BYTES, intent_dir) - return { - "intent_ref": str(path), - "acceptance": ["works"], - "acceptance_digest": hashlib.sha256(self.INTENT_BYTES).hexdigest(), - } - - def implement(_plan): - return {"subject_manifest_digest": "a" * 64} - - def validate(resolved, _candidate): - # Validate re-reads the snapshot from disk and re-derives the - # digest through its own runtime-fact binder — no value is passed - # through from Plan, so agreement is earned, not assumed. - replayed = Path(resolved["intent_ref"]).read_bytes() - bound = VALIDATE.bind_runtime_facts( - {"verdict": "PASS"}, - replayed, - None, - None, - None, - None, - None, - None, - ) - return { - "verdict": "PASS", - "acceptance_digest": bound["acceptance_digest"], - "subject_manifest_digest": "a" * 64, - "author_context_id": "author-ctx", - "validator_context_id": "validator-ctx", - "freshness_attestation": { - "source": "runtime", - "attester_identity": "runtime:composed-test", - }, - "verdict_digest": "b" * 64, - "verdict_ref": str(Path(tmp) / "verdict.json"), - "checked": ["acceptance"], - "not_checked": [], - } - - result = MODULE.invoke_once( - self.INTENT_BYTES, - continue_guard, - plan, - implement, - validate, - ) - - self.assertEqual(result["status"], "PASS") - self.assertEqual( - result["acceptance_digest"], - hashlib.sha256(self.INTENT_BYTES).hexdigest(), - ) - - def test_the_canonical_json_digest_is_not_the_intent_identity(self): - """Pins the two digests apart so the defect cannot silently return. - - If someone reintroduces a canonical-JSON digest of the parsed mapping as - the acceptance identity, this fails: the byte digest and the value digest - are different numbers for the same intent. - """ - import json - - byte_digest = hashlib.sha256(self.INTENT_BYTES).hexdigest() - value_digest = VALIDATE.digest_value(json.loads(self.INTENT_BYTES)) - self.assertNotEqual(byte_digest, value_digest) - - # And the collision the byte digest forbids: two distinct sources that - # parse to the same mapping must NOT share an acceptance identity. - reordered = b'{"intent_ref": "bead:agentops-test", "acceptance": ["works"]}' - self.assertEqual(json.loads(reordered), json.loads(self.INTENT_BYTES)) - self.assertNotEqual(byte_digest, hashlib.sha256(reordered).hexdigest()) - self.assertEqual(value_digest, VALIDATE.digest_value(json.loads(reordered))) - - -if __name__ == "__main__": - unittest.main() diff --git a/skills/skill-builder/references/converter/skill-bundle-schema.md b/skills/skill-builder/references/converter/skill-bundle-schema.md index 765053a86..66ac783e2 100644 --- a/skills/skill-builder/references/converter/skill-bundle-schema.md +++ b/skills/skill-builder/references/converter/skill-bundle-schema.md @@ -75,7 +75,7 @@ Target adapters receive the SkillBundle and decide which fields to use: | Adapter | Fields Used | Notes | |---------|-------------|-------| -| codex | name, description, body, references, scripts | Emits `SKILL.md` + `prompt.md`; modular by default (copies + links resources), `--codex-layout inline` appends them | +| codex | name, description, body, references, scripts | Emits `SKILL.md`; modular by default (copies + links resources), `--codex-layout inline` appends them | | cursor | name, description, body, references, scripts | Emits a single `.mdc` rule (+ optional `mcp.json`), budget-fitted to 100KB | | test | all | Dumps the full bundle as structured markdown for inspection | diff --git a/skills/skill-builder/scripts/converter/convert.sh b/skills/skill-builder/scripts/converter/convert.sh index 646d48bd8..a613d8116 100755 --- a/skills/skill-builder/scripts/converter/convert.sh +++ b/skills/skill-builder/scripts/converter/convert.sh @@ -357,7 +357,7 @@ convert_test() { CONVERTED_FILENAME="bundle.md" } -# Codex target: SKILL.md + prompt.md +# Codex target: SKILL.md # Codex may load these skills from ~/.codex/skills or from a native plugin cache. # Description max 1024 chars, no hooks support, tool names pass through convert_codex() { @@ -440,23 +440,9 @@ convert_codex() { fi fi - # ── Build prompt.md ── - local prompt_md="" - prompt_md+="# ${BUNDLE_NAME}"$'\n\n' - prompt_md+="${desc}"$'\n\n' - prompt_md+="## Instructions"$'\n\n' - prompt_md+="Load and follow the skill instructions from the sibling \`SKILL.md\` file for this skill."$'\n' - if [[ "$CODEX_LAYOUT" == "modular" && ( ${#REF_NAMES[@]} -gt 0 || ${#SCRIPT_NAMES[@]} -gt 0 ) ]]; then - prompt_md+="Then read local files in \`references/\` and \`scripts/\` when needed."$'\n' - fi - - # Set primary output (SKILL.md) + # Codex reads SKILL.md only; it has no prompt.md consumer. CONVERTED_OUTPUT="$skill_md" CONVERTED_FILENAME="SKILL.md" - - # Set secondary output (prompt.md) - CONVERTED_OUTPUT_2="$prompt_md" - CONVERTED_FILENAME_2="prompt.md" } # Cursor target: .mdc rule file with YAML frontmatter + optional mcp.json diff --git a/skills/skill-builder/scripts/converter/validate.sh b/skills/skill-builder/scripts/converter/validate.sh index 3512bc76b..0db3096ed 100755 --- a/skills/skill-builder/scripts/converter/validate.sh +++ b/skills/skill-builder/scripts/converter/validate.sh @@ -55,7 +55,7 @@ SRC_DIGEST="$(dir_digest "$FIX")" # does not touch the source. out="$WORK/out" if bash "$CONVERT" "$FIX" codex "$out" >/dev/null 2>&1 \ - && [[ -f "$out/SKILL.md" && -f "$out/prompt.md" && -f "$out/references/note.md" && -f "$out/scripts/tool.sh" ]] \ + && [[ -f "$out/SKILL.md" && ! -e "$out/prompt.md" && -f "$out/references/note.md" && -f "$out/scripts/tool.sh" ]] \ && [[ "$(dir_digest "$FIX")" == "$SRC_DIGEST" ]]; then pass "happy-path conversion writes target + passthrough files; source intact" else diff --git a/skills/skill-builder/scripts/heal.sh b/skills/skill-builder/scripts/heal.sh index 4db96eb58..45b37a8b6 100755 --- a/skills/skill-builder/scripts/heal.sh +++ b/skills/skill-builder/scripts/heal.sh @@ -40,6 +40,11 @@ bash "$SCRIPT_DIR/run-ao.sh" skills check-source --repo "$REPO_ROOT" --strict "$ rc=$? set -e [[ $rc -ne 2 ]] || exit 2 +# 126/127 mean ao could not run at all; that is not an advisory finding. +if [[ $rc -eq 126 || $rc -eq 127 ]]; then + echo "heal.sh: could not run 'ao skills check-source'; install ao 3.9 or later, or run from a source checkout with Go" >&2 + exit 2 +fi if [[ "$MODE" == fix && $rc -eq 0 ]]; then # Source behavior remains human-authored. Repair only owned projections. diff --git a/skills/validate/SKILL.md b/skills/validate/SKILL.md index ac0f776d7..f38ecbba2 100644 --- a/skills/validate/SKILL.md +++ b/skills/validate/SKILL.md @@ -47,6 +47,22 @@ caller wants advice or an acceptance judgment and wait for the answer. Do not issue a verdict, acceptance conclusion or readiness approval while intent is unresolved; missing intent is not a `NOT_PROVEN` verdict. +## When a fresh judgment is worth it + +Spend validation where a mistake is costly. For an ordinary change the author's +checks and CI are the gate, and no fresh judgment is owed. Use Validate when: + +- the caller asks for an acceptance verdict or independent proof; +- a mistake cannot be cheaply undone after it lands: a published release or + instructions users will follow, a security boundary, destroying data or + tracker state, deleting a check that protects the product; or +- no deterministic check covers the behavior that changed. + +Judge once. After the author repairs findings, the affected checks confirm the +repair; a second judgment happens only when the caller asks for one. Keep the +judgment's cost a fraction of the cost of the work: when it approaches that +cost, stop and return what is unchecked. + After acceptance intent is established, freshly judge the exact candidate against accepted intent, return `PASS`, `FAIL`, or `NOT_PROVEN`, and stop. The author cannot provide binding PASS. Advisory findings cannot substitute for @@ -118,11 +134,11 @@ ao provenance manifest --root "$REPO_ROOT" --include "$CHANGED_PATH" depth: acceptance, permissions, tests/gates, stopping, disclosure, hooks and executable controls warrant deeper inspection, including prose policy. Unknown risk merits examination, not automatic extra reviewers. -3. Re-execute discriminating proofs for risk-critical, uncertain or thinly - evidenced claims. Valid digest-bound receipts may establish routine facts; - do not replay every author command or full suite merely because this is a - fresh context. The repository's required integration checks still run on - the final subject. A changed subject needs new judgment and affected checks. +3. Read and reason; do not re-run checks the author ran on this exact subject + or that CI will run. Their receipts establish those facts. Re-execute a + proof only for a risk-critical claim that has no receipt. A changed subject + needs its affected checks rerun by the author; it needs a new judgment only + when the caller asks for one. 4. Classify commands before executing them. Regeneration, synchronization, formatting and `--force` are subject-mutating until proven otherwise; run them only on a disposable copy or a committed subject, never an uncommitted @@ -140,6 +156,12 @@ ao provenance manifest --root "$REPO_ROOT" --include "$CHANGED_PATH" ## Findings and report +A finding is something that fails an acceptance criterion or would mislead a +user, break install or the CLI, or remove protection for the product. Report +anything else as an optional note; notes do not change the verdict and the +author may ignore them. Report `NOT_PROVEN` with its gaps and stop; do not +request or wait for another round. + `not_checked` means in-scope acceptance that was not verified. Other limits remain in criterion reasoning, declared non-goals or residual-risk prose; never hide or delete them to obtain PASS. Keep prior findings visible. For each new diff --git a/tests/docs/mkdocs-strict-allowlist.txt b/tests/docs/mkdocs-strict-allowlist.txt index 0d90fb4ec..7607b5433 100644 --- a/tests/docs/mkdocs-strict-allowlist.txt +++ b/tests/docs/mkdocs-strict-allowlist.txt @@ -41,7 +41,6 @@ WARNING - Doc file 'releases/2026-07-17-v3.3.0-notes.md' contains a link '../.. # and the strict-build policy outside docs/, rather than generated site pages. # Every target is an existing repository file; retired documents are labeled # as historical text at their source and are not allowlisted. -WARNING - Doc file 'ARCHITECTURE.md' contains a link '../skills/rpi/references/bounded-adapter.md', but the target is not found among documentation files. WARNING - Doc file 'agent-workflow-reference.md' contains a link '../skills/memory/SKILL.md', but the target is not found among documentation files. WARNING - Doc file 'agent-workflow-reference.md' contains a link '../skills/rpi/SKILL.md', but the target is not found among documentation files. WARNING - Doc file 'agent-workflow-reference.md' contains a link '../skills/validate/references/mechanics.md', but the target is not found among documentation files. @@ -52,7 +51,6 @@ WARNING - Doc file 'architecture/hexagon-port-realness-audit.md' contains a lin WARNING - Doc file 'architecture/rpi-traversal.md' contains a link '../../skills/rpi/SKILL.md', but the target '../skills/rpi/SKILL.md' is not found among documentation files. WARNING - Doc file 'architecture/rpi-traversal.md' contains a link '../../skills/agent-native/references/model-dispatch.md', but the target '../skills/agent-native/references/model-dispatch.md' is not found among documentation files. WARNING - Doc file 'architecture/rpi-traversal.md' contains a link '../../skills/rpi/references/outer-goal.md', but the target '../skills/rpi/references/outer-goal.md' is not found among documentation files. -WARNING - Doc file 'architecture/rpi-traversal.md' contains a link '../../skills/rpi/references/bounded-adapter.md', but the target '../skills/rpi/references/bounded-adapter.md' is not found among documentation files. WARNING - Doc file 'architecture/rpi-traversal.md' contains a link '../../skills/memory/SKILL.md', but the target '../skills/memory/SKILL.md' is not found among documentation files. WARNING - Doc file 'reference/skill-system-evolution.md' contains a link '../../skills/rpi/SKILL.md', but the target '../skills/rpi/SKILL.md' is not found among documentation files. WARNING - Doc file 'releases/2026-07-30-v3.4.0-notes.md' contains a link '../../skills/using-gc/SKILL.md', but the target '../skills/using-gc/SKILL.md' is not found among documentation files. diff --git a/tests/e2e/README.md b/tests/e2e/README.md index 079fb6acd..e877b46ca 100644 --- a/tests/e2e/README.md +++ b/tests/e2e/README.md @@ -23,10 +23,9 @@ against one isolated sandbox each: ```bash bash tests/e2e/goals-measure-scenarios.sh bash tests/e2e/goals-trace-chain.sh -bash tests/e2e/rpi-phased-domain.sh ``` -Other scripts (`goals-*.sh`, `rpi-phased-domain.sh`, +Other scripts (`goals-*.sh`, …) follow the same harness contract and can be run the same way. --- @@ -84,7 +83,6 @@ The agentops snapshot (audited 2026-05-18): | Feedback rewarding | 5 | 2 | 10 | ✅ mock-free (`proof-run.sh` Phase 5) | | Nightly dream cycle | 4 | 2 | 8 | ✅ mock-free (`proof-run.sh` Phase 6) | | Goals scenarios link + lint | 4 | 2 | 8 | ✅ mock-free (`goals-scenarios-link.sh`) | -| RPI phased domain dispatch | 4 | 3 | 12 | ✅ mock-free (`rpi-phased-domain.sh`) | | `ao skills link` fresh-HOME install | 5 | 3 | 15 | ✅ mock-free (`.github/workflows/install-e2e.yml`) | | openclaw daemon API | 2 | 2 | 4 | ⚠️ httptest fixture — acceptable, internal-only | | Claude CLI skill invocation | 3 | 4 | 12 | ⚠️ real Claude, non-deterministic — acceptable as advisory | @@ -153,7 +151,6 @@ prints a `[e2e-guard] WARNING:` line to stderr. | `proof-run.sh` | forge → pool-ingest → cite-promote → lookup → feedback → nightly | 5 of the top-10 highest-risk paths in one suite | | `goals-scenarios-link.sh` | create → verify-bidirectional → lint-clean → break → lint-fails | F1 of the goals epic | | `goals-measure-scenarios.sh` | measure → assert satisfaction | F2 | -| `rpi-phased-domain.sh` | dispatch → phase trace | F3 | | `goals-trace-chain.sh` | trace → dependency assert | F4 | | `goals-steer-auto.sh` | steer → re-prioritize | F5 | | `factory-operator-canary.sh` | factory admission → operator action | factory pipeline contract | diff --git a/tests/e2e/rpi-phased-domain.sh b/tests/e2e/rpi-phased-domain.sh deleted file mode 100755 index ce661ac6d..000000000 --- a/tests/e2e/rpi-phased-domain.sh +++ /dev/null @@ -1,27 +0,0 @@ -#!/usr/bin/env bash -# tests/e2e/rpi-phased-domain.sh — RETIRED TOMBSTONE (age-ipbr, 2026-07-04). -# -# This was the F3.T2 e2e for bead soc-58nt.3.7. It exercised the `ao rpi phased -# --domain` / `--scaffold-domain` CLI surface end-to-end. That entire `ao rpi` -# command surface was REMOVED in f61c5f0e7 (ADR-0009 / age-tlj6 teardown — the -# out-of-session RPI loop was replaced by the operating loop + NTM/Agent Mail -# substrate). The test's subject no longer exists, so the real assertions were -# retired: `ao rpi phased ...` now exits 1 ('unknown command "rpi"'), which is -# what left the doctrine-proof job's F3 step red on every main push since -# 2026-06-17. -# -# WHY THE FILE STILL EXISTS (not deleted): skills/rpi/references/rpi.feature -# tags four scenarios `@covered-by:tests/e2e/rpi-phased-domain.sh`. The -# scenario->test linkage gate (scripts/check-scenario-test-linkage.sh, run in the -# skill-gates job) FAILS on a dangling @covered-by path. Fully deleting this file -# therefore requires the skills/ lane to first drop those @covered-by tags and -# allowlist rpi.feature (or retarget them at the operating-loop e2e). Until that -# skills-side retarget lands, this tombstone keeps the path resolvable so the -# linkage gate stays green while the dead `ao rpi` assertions no longer run. -# -# The F3 invocation was removed from .github/workflows/validate.yml in the same -# change. This stub is intentionally a no-op that exits 0. -set -euo pipefail - -echo "RETIRED: tests/e2e/rpi-phased-domain.sh — the 'ao rpi' surface was removed in f61c5f0e7 (ADR-0009). No-op tombstone; see file header. (age-ipbr)" -exit 0 diff --git a/tests/integration/test_skill_builder.bats b/tests/integration/test_skill_builder.bats index eb02f5984..2d90cee7a 100644 --- a/tests/integration/test_skill_builder.bats +++ b/tests/integration/test_skill_builder.bats @@ -26,6 +26,13 @@ setup_file() { export BUILD_SH INIT_SH } +@test "heal check fails instead of passing when ao cannot run" { + run env AO_SKILL_BUILDER_BIN="$BATS_FILE_TMPDIR/no-such-ao" HEAL_REPO_ROOT="$SCRATCH_ROOT" \ + bash "$REAL_REPO_ROOT/skills/skill-builder/scripts/heal.sh" --check "$SCRATCH_ROOT/skills/rpi" + [ "$status" -eq 2 ] + [[ "$output" == *"could not run 'ao skills check-source'"* ]] +} + @test "builder rejects missing and unknown modes" { run bash "$BUILD_SH" [ "$status" -eq 2 ] diff --git a/tests/scripts/agents-operating-contract.bats b/tests/scripts/agents-operating-contract.bats index d9f65ca98..293ab42cf 100644 --- a/tests/scripts/agents-operating-contract.bats +++ b/tests/scripts/agents-operating-contract.bats @@ -10,10 +10,11 @@ "federated integration graph" \ "zero mandatory AgentOps skills" \ "Lean RPI operating charter" \ - "Accepted intent -> native implementation and checks -> fresh independent judgment -> finish" \ + "Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish" \ "Persist machine evidence only for a caller request or declared consumer" \ "It owns no aggregate retry controller" \ - "fresh independent judgment" \ + "Spend validation where a mistake is costly" \ + "repair does not start another review" \ "unchanged accepted outcome and scope; acceptance changes need caller authority" \ "Implement repairs ordinary known defects directly" \ "use at most one bounded fresh helper" \ diff --git a/tests/scripts/codex-skill-conformance.bats b/tests/scripts/codex-skill-conformance.bats index d8c6e5b8a..51e8590ae 100644 --- a/tests/scripts/codex-skill-conformance.bats +++ b/tests/scripts/codex-skill-conformance.bats @@ -142,7 +142,23 @@ run_gate() { write_policy $'policy:\n allow_implicit_invocation: maybe\n' run_gate [ "$status" -ne 0 ] - [[ "$output" == *"policy.allow_implicit_invocation must be a boolean"* ]] + [[ "$output" == *"policy.allow_implicit_invocation must be spelled true or false"* ]] + + # Shapes Codex 0.156.1 drops silently, policy included. + write_policy $'policy:\n allow_implicit_invocation: no\n' + run_gate + [ "$status" -ne 0 ] + [[ "$output" == *"Codex reads 'no' as a string"* ]] + + write_policy $'interface: nope\npolicy:\n allow_implicit_invocation: false\n' + run_gate + [ "$status" -ne 0 ] + [[ "$output" == *"interface must be a mapping"* ]] + + write_policy $'dependencies:\n tools: nope\npolicy:\n allow_implicit_invocation: false\n' + run_gate + [ "$status" -ne 0 ] + [[ "$output" == *"dependencies must be a mapping with a tools list"* ]] } @test "every explicit-only source skill carries the Codex policy" {