Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,13 +6,13 @@
},
"metadata": {
"description": "Engineering guidance for coding agents: behavior-driven planning, shared domain language, independent validation, and reusable improvements.",
"version": "3.9.0"
"version": "3.10.0"
},
"plugins": [
{
"name": "agentops",
"description": "Engineering guidance for coding agents: behavior-driven planning, shared domain language, independent validation, and reusable improvements.",
"version": "3.9.0",
"version": "3.10.0",
"source": "./",
"author": {
"name": "Boden Fuller",
Expand Down
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "agentops",
"version": "3.9.0",
"version": "3.10.0",
"description": "Engineering guidance for coding agents: behavior-driven planning, shared domain language, independent validation, and reusable improvements.",
"author": {
"name": "Boden Fuller",
Expand Down
2 changes: 1 addition & 1 deletion .codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "agentops",
"version": "3.9.0",
"version": "3.10.0",
"description": "Engineering guidance for coding agents: behavior-driven planning, shared domain language, independent validation, and reusable improvements.",
"skills": "./skills",
"interface": {
Expand Down
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -197,3 +197,6 @@ outputs/

# Provenance ledger advisory-lock sidecar (cross-process append lock; never committed)
docs/provenance/*.lock

# claude plugin eval run output (reports, traces); the cases are tracked
evals/plugin-eval/*/results/
89 changes: 89 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,95 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [3.10.0] - 2026-10-05

AgentOps 3.10 is a release about the skills themselves. All 28 were audited,
tested with Claude Code's plugin evaluator on Claude Opus 5.5 and edited where
the test showed a gap: descriptions now use the words a user would type, and
each skill leads with the rules a model misses unaided. On the same 28 requests
the agent met 349 of 372 practice criteria with 3.10, 292 with 3.9.0 and 269
with no plugin. On requests written blind, the matching skill loaded in 29 of 48
runs against 11 of 48 on 3.9.0; nine skills still did not load there. Claude
Exec is new, the eval cases ship in the repository, and no command or skill
name changes.

See the [curated release notes](https://github.com/boshu2/agentops/blob/main/docs/releases/2026-10-05-v3.10.0-notes.md)
for upgrade notes and known limits, and the
[evaluation report](https://github.com/boshu2/agentops/blob/main/docs/evals/2026-10-05-plugin-eval-opus-5-5.md)
for every number.

### Added

- Claude Exec runs one caller-supplied prompt through headless Claude Code
(`claude -p`) with tools and permission mode scoped to the task, one time
bound that a retry spends instead of renewing, captured output and the exit
status reported as a fact. It is caller-selected, like Codex Exec and AGY
Native. The menu is now 29 skills.
- `evals/plugin-eval/` holds two suites for `claude plugin eval`: 29 behavior
cases that compare an agent with and without the plugin, and 25 routing cases
that check whether the matching skill loads on a request that never names it.
`evals/plugin-eval/grade.py` grades a response against all of its criteria in
one judge call.
- The README gains "Make it yours", on cutting, rewriting, adding and measuring
skills, and "Evidence", with the evaluation results and their limits.
- Most skills gain a fixed output shape, such as Plan's one-slice block,
Implement's handoff, Validate's verdict skeleton, Doc's handoff template and
Memory's entry template.

### Changed

- Every skill description is rewritten in user phrasing, within 26 words and
180 characters. In the behavior suite the matching skill loaded in 59 of 72
runs on 3.10 against 22 of 72 on 3.9.0. Routing phrases that tests pin moved
to frontmatter triggers on Implement and Premortem.
- Each skill opens with the few rules an unaided model tends to miss. Maintainer
detail moved into references: Codex Exec's guarded runner, Council's modes,
Using GC's trust pre-seeding, the Research and Reverse Engineer pack and
invocation contracts, Skill Builder's build mechanics, Skill Eval's readouts,
Memory's toil evidence and Plan's resume and handoff rules.
- Validate judges from its own text and runs without `ao`. It no longer requires
a file in another skill's directory, accepts the commit and changed paths as
the subject's identity, and writes out its verdict order.
- Navigate counts a row as proven by cited evidence: the passing check that
exercises the criterion, or a Validate PASS where a fresh review was required.
- Premortem may run in the context that wrote the plan when no fresh context can
be started. It must say the independence check is missing and cannot save the
result as a durable review.
- Using GC sends a ready bead to the Mayor by mail; direct `gc sling` is only
for a city with no live Mayor.
- Security runs a scripted scan once per request and keeps the repeat-until-quiet
loop for the manual hunt. A review reports a result for every vulnerability
class and names a remediation class, with no plan, owner or ship decision.
- AGY Native stops and reports when `agy` is missing, with no silent fallback to
another runtime. The retired ban on `claude -p` is removed from it and from the
dispatch reference.
- Reverse Engineer gives each row exactly one verdict and treats a capability
known only from documentation as unverified.
- Doc returns a handoff in the response when no location is named and creates no
file. New CDLC handoffs still go only to protected non-Git storage.
- Skill Eval explains how `claude plugin eval` relates to the repository probe
runner, and says to confirm a skill loaded before reading a zero delta, to
calibrate the judge first and to pass `--no-publish`.
- The README and the install guide show `ao` as optional for Validate.

### Fixed

- AGY Native claimed a five-minute default for `--print-timeout`. The CLI
default is no limit; the skill now requires an explicit timeout.
- Security's OWASP checklist said the redteam script covered secrets, input
validation, SQL injection and XSS automatically. It scans only repository
prompt and control surfaces, and the references now name
`prompt_redteam.py scan` instead of a subcommand that does not exist.
- Craft Goal stated its stop condition five times with two different pass
counts. Interview said it creates no file while appending notes. Refactor's
reference told readers to tidy messages and to self-grade PASS or FAIL. Each
now says one thing.
- Idea Genie's validator paths resolve in an installed copy, and its portfolio
shape is shown inline.
- Codex Exec says that an installed copy has no `scripts/lib/codex-exec.sh` and
gives the direct `codex exec` form.
- Implement lists the evidence-orphan scan as not run when `ao` is absent.

## [3.9.0] - 2026-10-03

AgentOps 3.9 narrows the product to its own guidance. The ten bundled external
Expand Down
9 changes: 9 additions & 0 deletions PRODUCT.md
Original file line number Diff line number Diff line change
Expand Up @@ -188,6 +188,15 @@ experiment did not demonstrate incremental benefit. Those bounded results
inform selective use of guidance; they establish neither equivalence nor
general productivity improvement.

The [October plugin evaluation](docs/evals/2026-10-05-plugin-eval-opus-5-5.md)
ran one request per skill on Claude Opus 5.5, three times with the plugin and
three times without. With AgentOps 3.10 the agent met 363 of 387 practice
criteria; with no plugin, 272 of 387. On requests written blind, the matching
skill loaded in 31 of 50 runs. The criteria come from each skill's own rules,
and the scored requests were also used to tune the 3.10 descriptions, so the
result shows that the guidance changes behavior on those requests. It does not
establish better outcomes on real tasks, results on other models or net cost.

[Independent review caught incomplete acceptance coverage](https://github.com/boshu2/agentops/pull/1129)
in the trial readout, leading to a repair and regression test. That is a concrete
example of review producing a reusable check. It does not establish the net
Expand Down
125 changes: 110 additions & 15 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@

[Install](#quickstart) · [The loop](#the-operational-loop) ·
[Goals](#goals) · [Try it](#try-it) · [Skills](#skills-at-a-glance) ·
[Make it yours](#make-it-yours) · [Evidence](#evidence) ·
[Beyond AgentOps](#beyond-agentops)

</div>
Expand All @@ -32,6 +33,13 @@ results and project memories. [Memory](skills/memory/SKILL.md) curates supported
findings into linked project knowledge, connecting lessons to the work and
sources behind them. The next session can use that record instead of starting over.

The skills are a kit, and you are expected to change it. Each skill is one
Markdown file of instructions. Keep the ones that fit how you work, rewrite the
ones that almost fit, delete the rest, and add skills you write yourself or find
in other libraries. What you end up with is your own operating model for agents,
which is the reason this project exists. [Make it yours](#make-it-yours) shows
how.

Use these paths with Claude Code, Codex, Cursor, OpenCode, Gemini CLI, Pi and
other coding agents, or personal assistants such as OpenClaw and Grok Bot.
Start with [one useful task](#try-it); follow the [operational loop](#the-operational-loop)
Expand Down Expand Up @@ -127,7 +135,7 @@ each host has been tested for is in [host coverage and limits](docs/contracts/mu
</details>

Start a new session so the skills load. Most skills need only your coding
agent; Validate also needs the [`ao` CLI](#optional-ao-cli). Invocation names
agent; a few use the [`ao` CLI](#optional-ao-cli) when it is installed. Invocation names
vary by agent: this README shows Claude Code's `/agentops:<skill>`; Codex uses
`$agentops:<skill>`.

Expand Down Expand Up @@ -276,8 +284,9 @@ Start read-only in any repo, then swap the Job example for your own change.
<summary><strong>Validate an existing change</strong></summary>

Pick a finished change whose accepted behavior is recorded in an issue or
conversation. Run the required checks, keep the candidate unchanged, and
[install `ao`](#optional-ao-cli). Then open a **new conversation**, fill in the
conversation. Run the required checks and keep the candidate unchanged.
[Installing `ao`](#optional-ao-cli) is optional; it gives Validate a content
manifest for the change. Then open a **new conversation**, fill in the
references and paste:

```text
Expand Down Expand Up @@ -335,9 +344,85 @@ catalog: **[docs/SKILL-ROUTER.md](docs/SKILL-ROUTER.md)**.
| On demand | [`research`](skills/research/SKILL.md) [`domain`](skills/domain/SKILL.md) [`test`](skills/test/SKILL.md) [`refactor`](skills/refactor/SKILL.md) [`review`](skills/review/SKILL.md) [`security`](skills/security/SKILL.md) [`doc`](skills/doc/SKILL.md) [`reverse-engineer`](skills/reverse-engineer/SKILL.md) | Reached for when a specific question comes up |
| Learning | [`memory`](skills/memory/SKILL.md) | Curated `.context/` pages safe to commit |
| Judgment strategies | [`council`](skills/council/SKILL.md) [`premortem`](skills/premortem/SKILL.md) [`postmortem`](skills/postmortem/SKILL.md) [`reality-check`](skills/reality-check/SKILL.md) [`idea-genie`](skills/idea-genie/SKILL.md) | Multi-model councils (debates, idea duels, interview panels), idea brainstorms, plan challenges, postmortems and claim audits |
| Runtimes and factories | [`codex-exec`](skills/codex-exec/SKILL.md) [`agy-native`](skills/agy-native/SKILL.md) [`using-gc`](skills/using-gc/SKILL.md) | Selected executors and Gas City integration |
| Runtimes and factories | [`codex-exec`](skills/codex-exec/SKILL.md) [`claude-exec`](skills/claude-exec/SKILL.md) [`agy-native`](skills/agy-native/SKILL.md) [`using-gc`](skills/using-gc/SKILL.md) | Selected executors and Gas City integration |
| Skill craft | [`skill-builder`](skills/skill-builder/SKILL.md) [`skill-eval`](skills/skill-eval/SKILL.md) | Author skills and measure whether they help |

## Make it yours

Every skill here is a folder with one `SKILL.md`: plain instructions an agent
loads when a task calls for them. No skill depends on the full set, and coding
with none of them still works. That makes the library easy to take apart, and
you should. The 29 skills are a starting point for an operating model that fits
your work.

- **Start small.** Install two or three skills that match work you already do.
Add another when a real task asks for it.
- **Cut what you override.** If you or the agent keep ignoring a rule, change
the rule or remove the skill. `npx skills` lets you pick skills per agent, and
[`ao skills link --skill <name>`](docs/install-day2-ops.md#install-source-checkout)
links an exact subset from a checkout.
- **Rewrite what almost fits.** Fork this repository, edit the `SKILL.md` and
install from your fork with the same commands, or link a checkout so every
agent on your machine reads your edits. A plugin update replaces the installed
copy, so keep your changes in a repository you control.
- **Write your own.** When you have explained the same thing to an agent three
times, it is a skill. [Skill Builder](skills/skill-builder/SKILL.md) drafts the
package, and tells you when a note in an existing file is enough.
- **Mix libraries.** Run these next to your company's skills, the
[Agentic Coding Flywheel](#agentic-coding-flywheel) tools or any other
library. `ao skills link` never replaces a skill it did not install.
- **Change the workflows too.** The [operational loop](#the-operational-loop)
and the [goal workflow](#goals) are defaults. Skip the steps your work does
not need, reorder them, or write your own. The Claude Code workflow scripts in
[`workflows/`](workflows/) link into a project with
[`ao workflows link`](docs/install-day2-ops.md#workflows-claude-code-only).
- **Measure what you change.** `claude plugin eval` compares an agent with and
without a plugin on the same request. To measure one edit, run the old and
the new version as two plugins. [Skill Eval](skills/skill-eval/SKILL.md)
covers how to read the result, and [Evidence](#evidence) shows the cases this
repository runs; copy them for your own skills.

## Evidence

A skill is a page of instructions, so the test is whether an agent does anything
differently with it installed. AgentOps runs that test on Claude Code's own
evaluator, `claude plugin eval`: one realistic request per skill, three runs
with the plugin and three without, each answer graded against four or five
criteria taken from the practice the skill teaches. For Test, one criterion is
that a regression test is shown failing without the fix.

On Claude Opus 5.5, measured 2026-10-05:

| | No plugin | AgentOps 3.9.0 | AgentOps 3.10 |
|---|---:|---:|---:|
| Practice criteria met on 28 requests | 269 of 372 (72%) | 292 of 372 (78%) | 349 of 372 (94%) |
| Matching skill loaded on a blind request | | 11 of 48 runs | 29 of 48 runs |

Fourteen of the 29 skills moved their case by 0.15 or more. Craft Goal, Claude
Exec, Plan, Memory and Skill Builder gained the most. The other 15 made no
measurable difference, and for eight of those the agent already met every
criterion with no plugin.

Read the limits before you quote these numbers:

- One request per skill, three runs, one model. A difference under 0.15 is
noise.
- The 28 requests in the first row were also used to tune the 3.10
descriptions, which flatters 3.10. The blind requests in the second row were
written without sight of the descriptions.
- The criteria come from each skill's own rules. A pass shows the rule landed on
that request. It says nothing about the outcome of a real task, and an earlier
[coding pilot](PRODUCT.md#evidence-and-claim-limits) found no end-to-end
difference.
- Loading is the weak point. Nine skills did not load on a request they had
never seen, and Implement does not load on a quick fix. Name the skill when
you want its rules applied.

The [full report](docs/evals/2026-10-05-plugin-eval-opus-5-5.md) has every case,
the method and what was rerun. The cases live in
[`evals/plugin-eval/`](evals/plugin-eval/README.md): run them against your own
changes, or copy the layout to test your own skills.

## Where AgentOps fits

AgentOps grew from applying DevOps experience and established engineering
Expand All @@ -363,10 +448,11 @@ proof that every combination has been tested.

<a id="optional-ao-cli"></a>

## `ao` CLI (needed for Validate)
## `ao` CLI (optional)

Most skills need only your coding agent. Validate uses `ao` to identify the
exact change it judges.
Most skills need only your coding agent. With `ao` installed, Validate binds
the exact change it judges to a content manifest; without it, Validate names
the commit and the changed paths.

```bash
brew tap boshu2/agentops
Expand All @@ -383,20 +469,29 @@ With Go installed: `go install github.com/boshu2/agentops/cli/cmd/ao@latest`.
## Updating and advanced setup

<details>
<summary><strong>Upgrading to 3.9</strong></summary>
<summary><strong>Upgrading to 3.10</strong></summary>

<a id="upgrading-to-310"></a>
<a id="upgrading-to-39"></a>
<a id="upgrading-to-38"></a>
<a id="upgrading-to-37"></a>

Version 3.9 removes ten bundled external tool skills, retires three delivery
workflows, deletes the old curl installers and changes the Codex plugin to read
`skills/` directly. Read the
[3.9 release notes](docs/releases/2026-10-03-v3.9.0-notes.md) before updating. Use the
Version 3.10 keeps every 3.9 command and skill name. It rewrites the skill
descriptions so skills load on plain requests, adds
[`claude-exec`](skills/claude-exec/SKILL.md) and lets Validate run without `ao`.
Read the [3.10 release notes](docs/releases/2026-10-05-v3.10.0-notes.md). Use the
[plugin update instructions](docs/install-day2-ops.md#install-and-update-runtime-plugins)
or, for npx installs, `npx skills@latest update` ([update notes](docs/install-day2-ops.md#update)).
For Homebrew: `brew update && brew upgrade agentops`. Start a new session
afterward; new installs do not silently remove obsolete copies.
For Homebrew: `brew update && brew upgrade agentops`. In a source checkout, run
`git pull --ff-only`, then rerun `ao skills link` with the selectors you used
before, adding `--skill claude-exec` for the new skill; without selectors it
links every skill. Start a new session afterward; new installs do not silently
remove obsolete copies.

**Upgrading from 3.8 or earlier:** version 3.9 removed ten bundled external tool
skills, retired three delivery workflows, deleted the old curl installers and
changed the Codex plugin to read `skills/` directly. Read the
[3.9 release notes](docs/releases/2026-10-03-v3.9.0-notes.md) first.

**Upgrading from 3.6 or earlier:** read the [migration guide](docs/MIGRATION.md).
Version 3.7 removed commands and skill names, including `learn`, `codebase-recon`
Expand All @@ -417,7 +512,7 @@ Skill installation does not install tool dependencies:
| `rpi` | `ao`, conditional | delegates exact-subject checks to Validate; only persists `verdict.v2` when requested |
| `plan` | `ao`, conditional | runs `ao provenance snapshot-intent` with an explicit evidence root when the intent source is not durable |
| `implement` | `ao`, conditional | at an integration boundary whose changed paths affect bound evidence, runs `ao provenance evidence-orphans` |
| `validate` | `ao` | derives exact subject identity with the helper and uses `ao provenance store-verdict` when persistence is requested; Python/schema checks are developer-only |
| `validate` | `ao`, optional | with `ao`, derives exact subject identity from a content manifest and uses `ao provenance store-verdict` when persistence is requested; without it, names the commit and changed paths |
| `reality-check` | `ao`, conditional | inspect selected goal measurements with `ao goals` or evidence-store facts with `ao status` |
| `using-gc` | `ao` | rig prep runs `ao gc prepare` and `ao gc check` |
| `doc` | `ao`, optional | a requested continuity handoff may use `ao session handoff`/`rehydrate` |
Expand Down
Loading
Loading