Repository navigation
feat(skills): sharpen all skills, add claude-exec and prepare v3.10.0 - #1194
Merged
Merged
Conversation
Apply the 2026-10-04 audit and plugin-eval findings to every skill. - Descriptions rewritten in user phrasing (26 words, 180 characters max) so skills load on plain requests; routing phrases that tests pin move to frontmatter triggers on implement and premortem. - Each skill leads with the rules an unaided model misses and gains the output template its frontmatter promised; maintainer internals move to references. - Defects fixed: validate no longer blocks on a file outside its own directory and has a Git fallback for subject identity; agy-native's stale timeout default; security and using-gc contradictions; reverse-engineer's indistinguishable verdicts; navigate's proof rule now matches the validate-where-costly contract. - New claude-exec skill: one prompt through headless claude -p with a scoped permission posture, one time bound and a reported exit status. - New evals/plugin-eval suites for claude plugin eval: 29 behavior cases (with and without the plugin) and 25 held-out routing cases.
…10 upgrade block Unfinished: the Evidence section and the 3.10 release notes are not written yet.
…rader - implement: without ao, the evidence-orphan scan is listed as not run instead of silently skipped; description no longer says "however small". - plan: the pointer to resume-and-handoff names the retrospective and intent-snapshot cases. - doc: new CDLC handoffs, drafts and proof go only to protected non-Git storage; a caller-named location wins for everything else (ADR-0016). - skill-eval: say that claude plugin eval publishes its report by default and where it writes results. - claude-exec: drop the claim that no turn-cap flag exists; description reworded so the skill loads on a scripted claude -p request. - security references: name prompt_redteam.py scan, not a collect-redteam subcommand that does not exist. - evals/plugin-eval: README and grade.py (one judge call per response).
Results on Claude Opus 5.5 for 3.10, 3.9.0 and no plugin, with the method, limits and what was rerun; scorecard with per-criterion counts.
… caveat ao skills find must put claude-exec first for a Claude request and codex-exec first for a Codex one. The README and PRODUCT.md now say the scored requests were also used to tune the descriptions.
- Report the two by-design criteria per arm and say the audit read a checkout from just before 3.9.0. - Keep eval run output out of the repository in the documented commands. - grade.py lists runs it cannot grade and exits nonzero. - Scorecard carries the grader-agreement counts and the final rerun grades (349 of 372, 363 of 387). - Upgrade notes tell source-checkout users to relink with their selectors. - claude plugin eval compares with and without the plugin, not one skill.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every skill was audited and tested with
claude plugin evalon Claude Opus 5.5, then edited where the test showed a gap. This change also adds the Claude Exec skill, ships the eval suites and their results, and prepares v3.10.0.What changed
implementandpremortem.references/.validateno longer blocks on a file outside its own directory, and runs withoutao.agy-nativeno longer states a five-minute timeout default the CLI does not have.security,using-gc,craft-goalandintervieware resolved.reverse-engineergives each row exactly one verdict.navigate's proof rule now matches the 3.9 rule that ordinary changes finish on checks and CI.claude-execruns one prompt through headlessclaude -pwith scoped permissions, one time bound and a reported exit status.evals/plugin-eval/holds 29 behavior cases, 25 blind routing cases, a grader and the scorecard. The report isdocs/evals/2026-10-05-plugin-eval-opus-5-5.md.aois shown as optional for Validate.Evidence
Claude Opus 5.5, three runs per case, graded by Opus:
Limits: one request per skill, criteria taken from each skill's own rules, and the 28 requests were also used to tune the descriptions. Nine skills did not load on a blind request, and
implementdoes not load on a quick fix. The report has the method and everything that was rerun.Checks
ao gate check --full: 66 of 66.scripts/regen-all.sh --check: all projections current.go build,go vet,go test ./...andscripts/check-go-lint.sh: pass. New test:ao skills findranks the matching headless adapter first.bats tests/scripts/*.bats: pass.scripts/validate-release-notes.sh v3.10.0 --since v3.9.0andscripts/extract-release-notes.sh v3.10.0 v3.9.0: pass, also against a cleangit archive.scripts/ci-local-release.sh --quick --release-version 3.10.0: pass.