Skip to content

feat(skills): sharpen all skills, add claude-exec and prepare v3.10.0 - #1194

Merged
boshu2 merged 9 commits into
mainfrom
feat/skills-310
Oct 5, 2026
Merged

boshu2 merged 9 commits into
mainfrom
feat/skills-310

Conversation

@boshu2

@boshu2 boshu2 commented Oct 5, 2026

Copy link
Copy Markdown
Owner

Every skill was audited and tested with claude plugin eval on Claude Opus 5.5, then edited where the test showed a gap. This change also adds the Claude Exec skill, ships the eval suites and their results, and prepares v3.10.0.

What changed

  • Descriptions. All 29 are written in the words a user would type, within 26 words and 180 characters. Routing phrases that tests pin moved to frontmatter triggers on implement and premortem.
  • Bodies. Each skill leads with the rules a model misses unaided, and most return a fixed output shape. Maintainer detail moved into references/.
  • Defects fixed.
    • validate no longer blocks on a file outside its own directory, and runs without ao.
    • agy-native no longer states a five-minute timeout default the CLI does not have.
    • Contradictions in security, using-gc, craft-goal and interview are resolved.
    • reverse-engineer gives each row exactly one verdict.
    • navigate's proof rule now matches the 3.9 rule that ordinary changes finish on checks and CI.
  • New skill. claude-exec runs one prompt through headless claude -p with scoped permissions, one time bound and a reported exit status.
  • Evals. evals/plugin-eval/ holds 29 behavior cases, 25 blind routing cases, a grader and the scorecard. The report is docs/evals/2026-10-05-plugin-eval-opus-5-5.md.
  • README. New "Make it yours" and "Evidence" sections. ao is shown as optional for Validate.
  • Release prep. Version 3.10.0 in the manifests and CLI, changelog section and curated notes.

Evidence

Claude Opus 5.5, three runs per case, graded by Opus:

No plugin 3.9.0 3.10
Practice criteria met on 28 requests 269 of 372 292 of 372 349 of 372
Matching skill loaded on a blind request 11 of 48 runs 29 of 48 runs

Limits: one request per skill, criteria taken from each skill's own rules, and the 28 requests were also used to tune the descriptions. Nine skills did not load on a blind request, and implement does not load on a quick fix. The report has the method and everything that was rerun.

Checks

  • ao gate check --full: 66 of 66.
  • scripts/regen-all.sh --check: all projections current.
  • go build, go vet, go test ./... and scripts/check-go-lint.sh: pass. New test: ao skills find ranks the matching headless adapter first.
  • bats tests/scripts/*.bats: pass.
  • scripts/validate-release-notes.sh v3.10.0 --since v3.9.0 and scripts/extract-release-notes.sh v3.10.0 v3.9.0: pass, also against a clean git archive.
  • scripts/ci-local-release.sh --quick --release-version 3.10.0: pass.
  • One independent read of the diff by two readers who wrote none of it: 17 defects found, 16 fixed, and one kept as a deliberate change that the notes now disclose (Premortem may run in the plan author's context when no fresh context can be started).

boshu2 added 8 commits October 5, 2026 00:19
Apply the 2026-10-04 audit and plugin-eval findings to every skill.

- Descriptions rewritten in user phrasing (26 words, 180 characters max)
  so skills load on plain requests; routing phrases that tests pin move
  to frontmatter triggers on implement and premortem.
- Each skill leads with the rules an unaided model misses and gains the
  output template its frontmatter promised; maintainer internals move to
  references.
- Defects fixed: validate no longer blocks on a file outside its own
  directory and has a Git fallback for subject identity; agy-native's
  stale timeout default; security and using-gc contradictions;
  reverse-engineer's indistinguishable verdicts; navigate's proof rule
  now matches the validate-where-costly contract.
- New claude-exec skill: one prompt through headless claude -p with a
  scoped permission posture, one time bound and a reported exit status.
- New evals/plugin-eval suites for claude plugin eval: 29 behavior cases
  (with and without the plugin) and 25 held-out routing cases.
…10 upgrade block

Unfinished: the Evidence section and the 3.10 release notes are not written yet.
…rader

- implement: without ao, the evidence-orphan scan is listed as not run
  instead of silently skipped; description no longer says "however small".
- plan: the pointer to resume-and-handoff names the retrospective and
  intent-snapshot cases.
- doc: new CDLC handoffs, drafts and proof go only to protected non-Git
  storage; a caller-named location wins for everything else (ADR-0016).
- skill-eval: say that claude plugin eval publishes its report by default
  and where it writes results.
- claude-exec: drop the claim that no turn-cap flag exists; description
  reworded so the skill loads on a scripted claude -p request.
- security references: name prompt_redteam.py scan, not a collect-redteam
  subcommand that does not exist.
- evals/plugin-eval: README and grade.py (one judge call per response).
Results on Claude Opus 5.5 for 3.10, 3.9.0 and no plugin, with the method,
limits and what was rerun; scorecard with per-criterion counts.
… caveat

ao skills find must put claude-exec first for a Claude request and
codex-exec first for a Codex one. The README and PRODUCT.md now say the
scored requests were also used to tune the descriptions.
- Report the two by-design criteria per arm and say the audit read a
  checkout from just before 3.9.0.
- Keep eval run output out of the repository in the documented commands.
- grade.py lists runs it cannot grade and exits nonzero.
- Scorecard carries the grader-agreement counts and the final rerun grades
  (349 of 372, 363 of 387).
- Upgrade notes tell source-checkout users to relink with their selectors.
- claude plugin eval compares with and without the plugin, not one skill.
@boshu2
boshu2 merged commit 4fda294 into main Oct 5, 2026
9 checks passed
@boshu2
boshu2 deleted the feat/skills-310 branch October 5, 2026 13:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant