Fixes the Claude Code Exec backend for issue #233 - #238
Conversation
…ace support Register claude_code_exec as a full optimizer/target backend (issue microsoft#233). --backend claude_code_exec now defaults both roles to claude_code_exec so reflection sees the agent's complete session, and the SDK message stream is parsed into structured trace steps persisted as claude_trace_steps.txt and injected into the analyst prompt. - model/claude_code_backend.py (new): chat_optimizer/chat_optimizer_messages on run_claude_code_chat, reasoning_effort threaded through, retry loop that surfaces non-JSON structured replies as RuntimeError, token tracking. - model/codex_harness.py: parse/format/persist claude trace steps (text, tool_call, tool_result; drops init/thinking_tokens; 200-char tool_result cap; total truncation) + effort override on run_claude_code_chat. - trainer.py/reflect.py: inject Claude Trace Steps gated behind REFLACT_CLAUDE_TRACE_TO_OPTIMIZER, set by the trainer only for claude_code_exec targets with model.claude_trace_to_optimizer (mirrors codex gate; default true). - config.py/default.yaml/docs: model.claude_trace_to_optimizer key + flatten mapping + config.md rows. - backend_config.py + model/__init__.py: register backend, route chat dispatch, token summary, reasoning effort, deployments. - scripts/train.py, eval_only.py: symmetric default + accurate comments. - tests: tests/test_claude_code_backend.py (10 tests: parsing, dispatch, effort, retry, trainer/reflect gating); test_role_backend_resolution.py updated to the symmetric default. Verified: 58 unit tests pass; integration smoke on searchqa improved best-on-val 0.7500 -> 0.9375 with 80 claude_trace_steps.txt written; all output files valid UTF-8 (no GBK mojibake).
|
@microsoft-github-policy-service agree |
|
Thanks for tackling #233. I rechecked the current head (
The branch also conflicts with current |
Yifan Yang (Yif-Yang)
left a comment
There was a problem hiding this comment.
Thanks for tackling #233 — the diagnosis is right and the shape of the fix (drive Claude Code as the optimizer so reflection sees the full trajectory) is what I'd want. A few things need fixing before merge.
Note up front: I checked the REFLACT_ spelling and it is correct, not a typo — REFLACT_CODEX_TRACE_TO_OPTIMIZER already exists on main (reflect.py:192) from the project's former name. Matching it was the right call. Likewise claude_code_backend.py is not duplicating claude_backend.py; it reuses _build_prompt_from_messages and wraps a genuinely different transport. Both fine.
Blocker 1 — --schema is not a real Claude CLI flag
codex_harness.py:1120 emits cmd.extend(["--schema", ...]). On the installed CLI:
$ claude -p --schema '{"type":"object"}' 'hi'
error: unknown option '--schema'
The actual flag is --json-schema. Since claude-agent-sdk isn't installed in a default environment, use_sdk: auto falls through to exactly this CLI path — so structured-output calls fail at runtime for anyone without the SDK. No test covers it, which is why it's green.
Blocker 2 — the tool_result payload is dropped, defeating the point of the PR
parse_claude_trace_steps reads part.get("content") for tool_result blocks, but Anthropic content blocks carry their payload under text. Reproduced on your branch:
raw = {"messages":[{"role":"user","content":[{"type":"tool_result","tool_use_id":"tu_1",
"content":[{"type":"text","text":"THE ACTUAL RESULT PAYLOAD"}]}]}]}
parse_claude_trace_steps(json.dumps(raw))
# -> [{'type': 'tool_result', 'summary': '', 'index': 1}]summary is empty. The reflector gets the shape of the trajectory but none of the observations — which is the specific thing #233 asked to preserve.
Should fix — stale claude_trace_steps.txt across turns
_persist_claude_artifacts writes the steps file with mode "w" from the current call's raw, while _persist_artifacts concatenates claude_raw.txt across calls with a turn separator. In the spreadsheetbench multi-turn repair loop (codegen_agent.py:604), which reuses one work_dir, the steps file ends up describing only the last turn while raw holds all of them. And when parsing yields "" the write is skipped entirely, leaving the previous turn's file for the reflector to read as if it described this prediction. Suggest formatting from combined_raw and writing unconditionally.
Please reconsider — silent optimizer-backend flip
trainer.py / scripts/train.py / scripts/eval_only.py now force optimizer_backend = "claude_code_exec" whenever --backend claude_code_exec. Your own comment states the consequence: an explicit --optimizer_backend openai_chat is silently ignored because it happens to equal a base-config default. A flag the user typed should never be silently discarded. Either track provenance, or leave the optimizer default alone and let people opt in.
Also worth noting this changes cost characteristics for existing claude_code_exec users without them asking.
Housekeeping
- The
.gitignoreentry for.mcp.jsonis good practice; I confirmed no secret was actually committed. configs/_base_/default.yamlgainsclaude_trace_to_optimizer: true— please document it indocs/reference/config.mdalongsidecodex_trace_to_optimizer.- Conflicts with main in
skillopt/engine/trainer.py; needs a rebase.
Suite on your branch: 1112 passed, 10 skipped, 130 subtests.
The two blockers are both narrow and testable — a test that feeds a realistic Anthropic message stream through parse_claude_trace_steps would have caught the second one and would be good to have regardless.
Summary
Fixes the Claude Code Exec backend for issue #233: with
--backend claude_code_exec, both optimizer and target roles nowdefault to Claude Code (symmetric default), and the agent's SDK session trace is parsed into structured
claude_trace_steps.txtfiles that are injected into the reflectionprompt — addressing context-length truncation and trajectory loss.
Implementation
backend_config.py,model/__init__.py,scripts/train.py,scripts/eval_only.pyregisterclaude_code_exec; both rolesdefault to it (a role pinned to a non-default value like
minimax_chatstill overrides).codex_harness.pyaddsparse/format/persist_claude_trace_steps: extractstext/tool_call/tool_resultsteps from the SDK messagestream, drops
init/thinking_tokensbookkeeping, capstool_resultat 200 chars, and persists one file per prediction.reflect.pyinjects#### Claude Trace Stepsonly whenREFLACT_CLAUDE_TRACE_TO_OPTIMIZER=1, whichtrainer.pysets only forclaude_code_exectargets with
model.claude_trace_to_optimizer: true(default; mirrors the existing codex gate).REASONING_EFFORTmodule global is now consumed viarun_claude_code_chat(..., effort=...)and forwarded by thedispatcher.
tests/test_claude_code_backend.py;configs/_base_/default.yaml,docs/reference/config.md, and_FLATTEN_MAPextended.Results
tests/test_claude_code_backend.py+tests/test_role_backend_resolution.py→ 58 passed (11 new cases all green).searchqarun): exit 0,accept=1, best-on-val 0.7500 → 0.9375, 80claude_trace_steps.txtwritten.merged_patch.json/config.jsonvalid; skillv0000→v0001.How to verify
PYTHONUTF8=1 uv run --extra dev pytest tests/test_claude_code_backend.py tests/test_role_backend_resolution.py -q # → 58 passedRun a small
searchqatraining with--backend claude_code_exec --optimizer_backend claude_code_exec, then check:ValueError: Unsupported optimizer backend.claude_trace_steps.txt.#### Claude Trace Steps.PYTHONUTF8=1on Windows).