Outcome
Evaluation traces contain readable, complete context needed to explain benchmark behavior.
Scope
- Capture the system prompt used for a run.
- Render worker and coordinator prompts in a readable form.
- Retain enough trace context to reconstruct evaluation inputs and key decisions.
- Bound recorded content without silently destroying diagnostic value.
General streaming backpressure, client delivery, and runtime buffering remain in #208.
Complete when
- A supported evaluation failure can be investigated from its trace without relying on terminal-only output.
- Prompt content is readable and attributed to the correct run and agent.
- Trace-content limits are explicit and tested.
Outcome
Evaluation traces contain readable, complete context needed to explain benchmark behavior.
Scope
General streaming backpressure, client delivery, and runtime buffering remain in #208.
Complete when