diff --git a/README.md b/README.md index 96562b7..26800dc 100644 --- a/README.md +++ b/README.md @@ -144,6 +144,7 @@ These benchmarks are especially useful when you want to compare harness quality, - [AgentKit](https://github.com/inngest/agent-kit) - Inngest's TypeScript toolkit for building durable, workflow-aware agents on top of event-driven infrastructure. - [browser-use/browser-harness](https://github.com/browser-use/browser-harness) - A thin CDP-based browser harness that lets agents extend helper functions during execution, useful for inspecting self-healing web-task workflows. - [Citadel](https://github.com/SethGammon/Citadel) - A harness for Claude Code and OpenAI Codex with isolated worktrees, multi-agent coordination, and persisted memory and campaign state. +- [completely](https://github.com/23ag1/completely) - A Claude Code plugin harness that puts an independent default-FAIL evaluator in the close path of a fresh-context task loop: tasks live in a Beads queue, deterministic hooks enforce write-zone fencing, commit-before-close, and binding reviewer findings, and a task only closes on re-run evidence rather than agent self-report — with heartbeat-based orphan recovery and a post-batch integration gate for parallel workers. - [Bring Your AI MCP](https://github.com/unitedideas/bringyour-mcp) - Public harness-migration reference for Claude Code to Codex moves, with installable auditor artifacts and explicit validation notes for hooks, MCP config, and instruction-file differences. - [Harbor](https://github.com/harbor-framework/harbor) - A generalized harness for evaluating and improving agents at scale, released alongside Terminal-Bench 2.0. - [Harness Evolver](https://github.com/raphaelchristi/harness-evolver) - Claude Code plugin that autonomously evolves LLM agent harnesses using multi-agent proposers, LangSmith-backed evaluation, and git worktree isolation. Based on Meta-Harness (Lee et al., 2026).