## 定位 这是 Merry coding agent 产品化和架构收敛的父 Issue,负责维护总体目标、分层约束、子 Issue 依赖、横切质量任务和完成标准。具体实现与验收清单放在对应子 Issue 中。 ## 当前决策(2026-08-19) Merry 采用两条清晰的验证路径: 1. Rust unit/integration tests:验证 runtime、权限、生命周期、事件、artifact、ledger、取消、恢复和失败路径。测试使用 fake provider、scripted events 和 fake process runner,不依赖 live provider、外部网络或 Harbor。 2. Harbor benchmark:验证真实 coding-agent 任务、模型质量和外部 benchmark 对比。Harbor 负责数据集、任务环境、verifier、reward 和 benchmark 结果,作为手动/nightly 的上层消费者。 不再建设独立的 Offline Harness、内部 20-task Eval Suite、local eval CLI 或围绕 `merry-eval` 的兼容层。 ## 总体目标 - CLI、Rust SDK、Python SDK 使用同一个 CodingAgentProfile。 - prompt、工具 schema、工具顺序和权限策略具有稳定契约。 - Rust SDK 成为稳定、可发布的高层 API。 - runtime、coding composition、CLI、provider 和 binding 职责清晰。 - 普通 PR 保持 deterministic;live provider 和 Harbor 只进入显式的 manual/nightly 流程。 - 外部 benchmark adapter 不改变 runtime 核心契约。 ## 子 Issue 进度 - [x] [T0:建立交付基线、ROADMAP 与架构守护](https://github.com/locez/merry/issues/10) - [x] [T1:建立评测协议与 TaskSpec(原型已完成,后续暂缓)](https://github.com/locez/merry/issues/8) - [x] [T2:独立 Offline Harness(已关闭,改由 Rust 测试覆盖)](https://github.com/locez/merry/issues/11) - [x] [T3:建立唯一 CodingAgentProfile](https://github.com/locez/merry/issues/12) - [x] [T4:内部 Coding Eval Suite(已关闭,不建立独立 runner)](https://github.com/locez/merry/issues/13) - [x] [T5:统一 CLI、Debug 和 Rust Runtime Builder](https://github.com/locez/merry/issues/14) - [x] [T6:建立生产级可靠性、安全与观测 Harness](https://github.com/locez/merry/issues/15)(2026-08-26 关闭;PR #37) - [x] [T7:发布稳定的 Rust SDK](https://github.com/locez/merry/issues/16)(2026-08-29 关闭;PR #38) - [x] [T8:统一 Python 与 Rust SDK 能力](https://github.com/locez/merry/issues/17)(2026-09-01 关闭;PR #39) - [ ] [T9:接入外部 Coding Agent 评测集](https://github.com/locez/merry/issues/18)(adapter 已实现;端到端证据待补,优先级已下调) - [x] [T10:按职责拆分 Runtime 与 Binding 大模块](https://github.com/locez/merry/issues/19)(2026-09-07 关闭;PR #45) - [ ] [T11:发布门禁与 Definition of Done 收口](https://github.com/locez/merry/issues/20)(PR 门禁已建立;nightly/release 待收口) ## 依赖顺序 ```text T0 -> T3 T3 -> T5 T5 -> T6 T5 + T6 -> T7 -> T8 T9 为独立的 Harbor/manual-nightly 路径,可与 T5/T7/T8 并行 T5 + T7 + T8 -> T10 T6 + T9 + T10 -> T11 ``` ## 横切任务 - 每个子 Issue 独立完成、独立验证、独立记录回滚方式。 - 不先做机械式大规模文件拆分。 - 不先发布未稳定的 Rust SDK。 - 不先堆积 prompt 文本,再补行为测试。 - provider-specific wire format 不进入 runtime、SDK 或 evaluation adapter。 - provider-visible request 变化必须评估 prompt/KV cache 影响。 - live provider、外部 benchmark 和网络测试不得阻塞普通 deterministic CI。 - `crates/merry-eval` 暂缓并列为候选删除;后续实现不得扩展或依赖它。 ## Definition of Done - Rust unit/integration tests 覆盖 runtime 生命周期、权限、安全、取消、恢复和失败路径。 - CLI、Rust SDK、Python SDK 共用同一个 CodingAgentProfile。 - Rust SDK 有稳定 public API 并可打包发布。 - Python workspace 配置不再处于声明支持但实际 unsupported 的状态。 - Harbor 至少有 Terminal-Bench 和 Rust SWE benchmark adapter。 - cancellation、resume、权限、安全和 prompt injection 有负向测试。 - runtime、coding composition、CLI、provider、binding 的依赖方向明确。 - PR、nightly、release 三层质量门禁建立。 - 完整测试、clippy、fmt、文档、打包和 SDK parity 检查通过。 ## 进展记录 - 2026-09-20:T6、T7、T8、T10 已按各自验收完成并关闭,T0–T8 与 T10 全部完成。 - T9(#18):Harbor adapter、数据集配置与手动 smoke 已实现;端到端结果未记录,负责人已下调优先级。 - T11(#20):PR 门禁已建立;nightly、release 门禁与版本化发布记录待完成。 - 近期合入:sandbox/权限/CLI 配置系列修复(#46–#55)、MCP 离线恢复(#44)、subagent 预算恢复(#43)、上下文压缩与窗口预算(#61)、apply_patch 诊断(#60)、README 与 sandbox 文档重构(#62)、read_text 边界对齐(#63)。
定位
这是 Merry coding agent 产品化和架构收敛的父 Issue,负责维护总体目标、分层约束、子 Issue 依赖、横切质量任务和完成标准。具体实现与验收清单放在对应子 Issue 中。
当前决策(2026-08-19)
Merry 采用两条清晰的验证路径:
不再建设独立的 Offline Harness、内部 20-task Eval Suite、local eval CLI 或围绕
merry-eval的兼容层。总体目标
子 Issue 进度
依赖顺序
横切任务
crates/merry-eval暂缓并列为候选删除;后续实现不得扩展或依赖它。Definition of Done
进展记录