Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 22 additions & 21 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,11 +8,11 @@ and evidence for independently judgeable changes. Its federated integration grap
connects evidence; tracker, Git and coding agents own work, history and execution. Native execution needs zero mandatory AgentOps skills.

```text
Accepted intent -> native implementation and checks -> fresh independent judgment -> finish
Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish
```

Use the existing issue or conversation and load a skill only for a concrete
uncertainty or an explicitly selected workflow. Without fresh independent judgment over the exact subject, the experiment remains unproven.
uncertainty or an explicitly selected workflow. Checks and CI are the gate for an ordinary change; a fresh read is for the costly cases named in the charter below.
Native goals supply continuity; BD keeps work and handoffs. Neither needs RPI.
Creating a goal does not reset the conversation or enforce a resource limit.
Delegate focused intent, scope and evidence; retrieve history only when it can change a decision. Assign final review once, preserving required review legs.
Expand Down Expand Up @@ -159,26 +159,28 @@ outcome and judgment; explicit interim analysis names its cutoff and unknowns.
Neither holds code acceptance open; the overall goal still owes both deliverables.

Use cheap discriminating checks during edits, required integration checks before
final judgment, and reserve capacity for integration, validation, repair and a
truthful handoff. On a genuine causal stall (unknown cause, recurrence, no
finishing, and reserve capacity for integration, repair and a truthful handoff. On a genuine causal stall (unknown cause, recurrence, no
progress or wrong objective), use at most one bounded fresh helper for that
incident within authority and real remaining bounds. An unhelpful answer ends
the attempt; do not build a helper chain. Known failures need direct repair.
Cancellation, refusal and spent hard time/cost/quota skip help. Retry counts,
compaction, helpers and new subjects never renew real limits. Preserve compact
recovery state in native handoff only when needed to prevent evidence loss.

Fresh author-distinct final validation is required over the exact subject,
unchanged acceptance and all changed paths. Default to a fresh reviewer from
the author's model family; cross-model review is opt-in and every explicitly
required leg remains required. No fixed ten-minute cap applies. Risk determines
evidence depth, not mandatory specialist or model-family multiplication.
PASS needs distinct identities, attested freshness, nonempty checked scope,
evidence for every criterion and empty `not_checked`. Missing identity,
freshness, subject continuity or acceptance proof means `NOT_PROVEN`; proven
out-of-scope change or failed acceptance means `FAIL`. Repair known findings
within authority and real bounds, then revalidate the changed exact subject.
Persist machine evidence only for a caller request or declared consumer.
Spend validation where a mistake is costly. For an ordinary change the author
runs the checks and CI is the gate; no fresh review is owed. Obtain one fresh
author-distinct read only when the caller asks, a mistake cannot be cheaply
undone after it lands (a published release or instructions users will follow,
a security boundary, destroying data or tracker state, deleting a check that
protects the product), or no deterministic check covers the changed behavior.
One round: the reviewer gets the exact subject and one question and does not
re-run checks. Defects are only what fails accepted behavior or would mislead a
user, break install or the CLI, or remove protection for the product; the rest
are optional notes. Repair defects, confirm each with a check, and finish; a
repair does not start another review. Keep review cost a fraction of the cost
of the work. A requested binding verdict keeps the PASS rule in
[Validate](skills/validate/SKILL.md); report `NOT_PROVEN` with its gaps instead
of chasing it. Persist machine evidence only for a caller request or declared consumer.

[Memory](skills/memory/SKILL.md) owns optional find/recall, capture/mining and curation:
support, applicability, invalidation and preservation; no blind TTL/deletion.
Expand All @@ -193,15 +195,14 @@ independent support/disclosure review precede Git import (ADR-0016).
The [RPI skill](skills/rpi/SKILL.md) packages this charter when explicitly selected;
it is not a prerequisite for native execution or independent review. The
[architecture reference](docs/architecture/rpi-traversal.md) owns exact evidence
semantics. Optional outer-goal guidance and the grandfathered fixed-dispatch
reference adapter stay outside the native core. No scheduler or new AO command
semantics. Optional outer-goal guidance stays outside the native core. No scheduler or new AO command
is needed for this harness.

## Product boundary

AgentOps reads or refines caller-owned intent, implements authorized work and direct repairs,
establishes exact content identity, and obtains fresh independent judgment. It
can persist that judgment as standalone evidence when requested. It owns no aggregate retry controller,
establishes exact content identity, and obtains one fresh judgment where a mistake is
costly. It can persist that judgment as standalone evidence when requested. It owns no aggregate retry controller,
budget, queue, work ownership, Git, closure, release, landing, or delivery
transition. Consumer repositories keep their own direct-push, PR, CI, merge,
rollback, and release policy.
Expand Down Expand Up @@ -243,5 +244,5 @@ AgentOps work ownership.
## Closeout

Map each acceptance criterion to evidence and disclose `checked` and
`not_checked`; any unchecked acceptance means `NOT_PROVEN`. Apply the fresh
judgment and delivery-authority rules above.
`not_checked`. An unchecked item is reported, not a reason to keep validating.
Apply the validation and delivery-authority rules above.
33 changes: 33 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,20 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Changed

- Fresh validation is no longer owed on every change. For an ordinary change
the author's checks and CI are the gate. RPI, Validate, Implement,
Orchestrate, Craft Goal and Navigate call for one fresh read only when the
caller asks, a mistake cannot be cheaply undone after it lands (a published
release or instructions users will follow, a security boundary, destroying
data or tracker state, deleting a check that protects the product), or no
deterministic check covers the changed behavior. A review is one round: the
reviewer does not re-run checks, reports as defects only what fails accepted
behavior or would mislead a user, break install or the CLI, or remove
protection for the product, and a repair does not start another review. A
requested binding verdict keeps its PASS rule; `NOT_PROVEN` is reported with
its gaps instead of being chased.
- Craft Goal, Interview and Navigate are marked stable; the README no longer
labels goals experimental.
- The Codex plugin now loads `skills/` directly. `.codex-plugin/plugin.json`
ships `./skills`, the same tree `ao skills link` and `npx skills` already
install, instead of a generated copy. Skill names, descriptions and bodies are
Expand Down Expand Up @@ -66,6 +80,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- `ao doctor` no longer has the `fm-skills-stale-codex-sync` failure mode, and
its fixers can no longer write to `~/.codex/plugins/cache/agentops-marketplace`
or `~/.codex/.agentops-codex-install.json`.
- The fixed-dispatch RPI reference adapter, which modelled repeated review
rounds: `skills/rpi/scripts/run_once.py`, its tests, its reference page and
`skills/rpi/references/rpi.feature`, plus `tests/e2e/rpi-phased-domain.sh`, a
no-op tombstone that existed only to keep that feature file's scenario link
resolving.
- The Skill Builder converter's Codex target no longer writes `prompt.md`;
Codex does not read it.

- Bundled Flywheel tool skills (`account-rotation`, `agent-mail`, `cass`, `cc-hooks`, `dcg`,
`ms`, `ntm`, `rch`, `sbh`, `using-flywheel`) and their generated Codex copies.
Expand Down Expand Up @@ -102,6 +123,18 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
`~/.pi/agent/skills`, and detect Pi by `~/.pi/agent`. Links an earlier version
made in `~/.pi/skills` are no longer swept by default; remove them with
`ao skills unlink --dest ~/.pi/skills`.
- Skill Builder's `heal.sh --check` exits 2 with a message when `ao` cannot run.
It used to report a pass in non-strict mode.
- The Codex policy check also fails on `agents/openai.yaml` shapes Codex
silently ignores: a non-object `interface`, `dependencies` that is not a
mapping with a `tools` list, and a boolean spelled other than `true` or
`false`. Any of these made an explicit-only skill implicitly selectable.
- The conformance probe that checks the Validate helper makes no Git, tracker or
delivery calls now also intercepts calls made in-process; before, those
reached the real binaries unnoticed.
- The user-facing docs (`PRODUCT.md`, the docs index, how-it-works, architecture,
philosophy, migration and CI pages) describe validation as checks and CI plus
one fresh read where a mistake is costly, matching the skills.

## [3.8.0] - 2026-09-22

Expand Down
4 changes: 2 additions & 2 deletions PRODUCT.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ coding agents, checks, and independent judgment while existing tools retain
work, execution, and delivery. The standard path is:

```text
Accepted intent -> native implementation and checks -> fresh independent judgment -> finish
Accepted intent -> native implementation and checks -> one fresh read where a mistake is costly -> finish
```

## Engineering practices in the workflow
Expand Down Expand Up @@ -177,7 +177,7 @@ outcomes remain visible. Positive usefulness needs later task evidence. These
skills maintain external context; they do not train weights or promise
deterministic inference. [RPI traversal](docs/architecture/rpi-traversal.md) and
[ADR-0017](docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md) own the amended
behavior and the explicitly optional fixed-dispatch reference adapter.
behavior.

## Evidence and claim limits

Expand Down
19 changes: 11 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,8 @@ AgentOps provides optional skills and a CLI (`ao`). The same `SKILL.md` skills w
with coding agents (Claude Code, Codex, Cursor, OpenCode, Gemini CLI, Pi and
others) and personal assistants (OpenClaw, Grok Bot). You state intent as
behavior in your domain's words. The skills carry it through one change
(Plan → Implement → Validate, an **RPI**) or, for bigger work, a goal made of
(Plan → Implement → checks, with Validate where a mistake is costly: an **RPI**)
or, for bigger work, a goal made of
many RPIs tracked in
[Beads](https://github.com/gastownhall/beads), a dependency-aware issue tracker.

Expand All @@ -33,7 +34,7 @@ many RPIs tracked in
|---|---|
| Builds something different from what you meant | Given/When/Then examples shared by implementation and review |
| Uses three names for one concept | One domain term per concept, in intent, code and tests |
| Says “done” after a green test run | A fresh judge that didn't write the change |
| Says “done” after a green test run on a change that matters | A fresh judge that didn't write the change |
| Loses the thread on work bigger than one session | A Beads graph holding intent, dependencies and verdicts |
| Runs off with a half-formed goal | An interview that settles the goal before agents go autonomous |
| Gives one model's answer to a hard call | A [council](skills/council/SKILL.md): judges in fresh contexts, each with the model, effort and perspective you assign (one model family or several vendors), compare, duel (score each other's ideas) or debate to your majority, keep dissent, and can answer an interview for you; [Idea Genie](skills/idea-genie/SKILL.md) brainstorms options |
Expand Down Expand Up @@ -156,9 +157,10 @@ author never approves its own work. Merging and releasing follow your repo's rul

## Goals

[`rpi`](skills/rpi/SKILL.md) runs Plan → Implement → Validate for one outcome
[`rpi`](skills/rpi/SKILL.md) runs Plan → Implement → checks, with Validate where a
mistake is costly, for one outcome
without check-ins (your agent's permission prompts still apply) and stops at
acceptance, a blocker or a spent limit. Bigger work becomes a goal (experimental;
acceptance, a blocker or a spent limit. Bigger work becomes a goal (it
needs Beads: `brew install beads`, then `bd init` in your repo):

1. **[Interview](skills/interview/SKILL.md).** One question at a time, each with
Expand Down Expand Up @@ -273,9 +275,9 @@ catalog: **[docs/SKILL-ROUTER.md](docs/SKILL-ROUTER.md)**.

| Group | Skills | What it covers |
|---|---|---|
| Operational loop | [`plan`](skills/plan/SKILL.md) [`implement`](skills/implement/SKILL.md) [`validate`](skills/validate/SKILL.md) | Shape, build and judge every change |
| Operational loop | [`plan`](skills/plan/SKILL.md) [`implement`](skills/implement/SKILL.md) [`validate`](skills/validate/SKILL.md) | Shape and build a change; judge it where a mistake is costly |
| Autonomous | [`rpi`](skills/rpi/SKILL.md) | One outcome, end to end |
| Goals (experimental) | [`interview`](skills/interview/SKILL.md) [`craft-goal`](skills/craft-goal/SKILL.md) [`navigate`](skills/navigate/SKILL.md) | Shape, write and walk a goal over the bead graph |
| Goals | [`interview`](skills/interview/SKILL.md) [`craft-goal`](skills/craft-goal/SKILL.md) [`navigate`](skills/navigate/SKILL.md) | Shape, write and walk a goal over the bead graph |
| Coordination | [`orchestrate`](skills/orchestrate/SKILL.md) [`agent-native`](skills/agent-native/SKILL.md) | Fresh workers per bead, disjoint scopes, integration |
| On demand | [`research`](skills/research/SKILL.md) [`domain`](skills/domain/SKILL.md) [`test`](skills/test/SKILL.md) [`refactor`](skills/refactor/SKILL.md) [`review`](skills/review/SKILL.md) [`security`](skills/security/SKILL.md) [`doc`](skills/doc/SKILL.md) [`reverse-engineer`](skills/reverse-engineer/SKILL.md) | Reached for when a specific question comes up |
| Learning | [`memory`](skills/memory/SKILL.md) | Curated `.context/` pages safe to commit |
Expand All @@ -300,7 +302,8 @@ lineage; [how it works](docs/how-it-works.md) covers responsibilities.
| Factories such as [Gas City](skills/using-gc/SKILL.md) | Own agent coordination and execution through their native control plane |

Choose which workflow leads the task. Carry accepted behavior and evidence
into [independent judgment](skills/validate/SKILL.md). Shared practices are not
into checks, and into [independent judgment](skills/validate/SKILL.md) where a
mistake would be costly. Shared practices are not
proof that every combination has been tested.

</details>
Expand Down Expand Up @@ -355,7 +358,7 @@ Skill installation does not install tool dependencies:

| Skill | Needs | Why |
|---|---|---|
| `rpi` | `ao`, conditional | delegates exact-subject checks to Validate; only persists `verdict.v2` when requested, with the fixed-dispatch adapter optional |
| `rpi` | `ao`, conditional | delegates exact-subject checks to Validate; only persists `verdict.v2` when requested |
| `plan` | `ao`, conditional | runs `ao provenance snapshot-intent` with an explicit evidence root when the intent source is not durable |
| `implement` | `ao`, conditional | at an integration boundary whose changed paths affect bound evidence, runs `ao provenance evidence-orphans` |
| `validate` | `ao` | derives exact subject identity with the helper and uses `ao provenance store-verdict` when persistence is requested; Python/schema checks are developer-only |
Expand Down
12 changes: 5 additions & 7 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -1,16 +1,16 @@
# Architecture

AgentOps defaults to native execution with zero mandatory skills. Its small
semantic core binds acceptance to exact content and fresh independent judgment;
semantic core binds acceptance to exact content and, where a mistake is costly, one fresh independent judgment;
optional skills and adapters package guidance around that boundary.

```text
existing bead or caller intent
-> on-demand planning and bounded implementation
-> runtime-derived subject-manifest.v1 + check receipts
-> fresh independent judgment
-> PASS | FAIL | NOT_PROVEN
-> direct repair and fresh revalidation within real bounds when needed
-> checks and CI as the gate for an ordinary change
-> one fresh judgment where a mistake is costly: PASS | FAIL | NOT_PROVEN
-> direct repair, confirmed by a check rather than another judgment
-> completed acceptance or truthful unfinished result
```

Expand All @@ -33,9 +33,7 @@ existing bead or caller intent
Known findings are repaired within authority and real remaining bounds. A
genuine causal stall admits at most one bounded fresh helper; an unhelpful
answer, cancellation, refusal or a spent real bound ends the attempt. Explicit
repair-round bounds still apply, and no invocation renews an allowance. The
[fixed-dispatch adapter](../skills/rpi/references/bounded-adapter.md) keeps its
narrower optional contract; it does not govern direct native execution.
repair-round bounds still apply, and no invocation renews an allowance.

## Hexagonal boundary

Expand Down
Loading
Loading