Skip to content

[EPIC] Eval Trace Fidelity #439

Description

@charlesjohnson

Outcome

Evaluation traces contain readable, complete context needed to explain benchmark behavior.

Scope

  • Capture the system prompt used for a run.
  • Render worker and coordinator prompts in a readable form.
  • Retain enough trace context to reconstruct evaluation inputs and key decisions.
  • Bound recorded content without silently destroying diagnostic value.

General streaming backpressure, client delivery, and runtime buffering remain in #208.

Complete when

  • A supported evaluation failure can be investigated from its trace without relying on terminal-only output.
  • Prompt content is readable and attributed to the correct run and agent.
  • Trace-content limits are explicit and tested.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions