planning latency, attempt latencies, drain time → OTel span attributes
Orchestration runs currently expose only coarse timing (total duration, per-task duration). Missing phase-level timing makes it difficult to identify bottlenecks in planning, synthesis, and evaluation. There is also no breakdown of where time is spent within an LLM call — reasoning, tool execution, and generation are opaque.
Scope:
- Phase timing on IterationComplete SSE event: planning, execution (wall-clock and aggregate), synthesis, and evaluation durations
- Per-call timing breakdown: Track time spent in reasoning vs tool calls vs generation within LLM interactions
- Surface timings through existing structures: SSE events, run manifest / artifact files, and OTel spans where tracing is wired
Design notes:
- Label "start to plan_created" as planning latency (includes retries), not routing latency
- Define both wall-clock (max of wave) and aggregate compute (sum of tasks) for parallel execution
- Currently Synthesizing → IterationComplete spans both synthesis and evaluation — needs separate durations
Acceptance criteria:
- IterationComplete SSE event includes phase timing fields
- Reasoning/tool-call/generation timing breakdown available per LLM call
- Phase timing fields documented in docs/streaming-api-guide.md
- Existing E2E and unit tests pass
planning latency, attempt latencies, drain time → OTel span attributes
Orchestration runs currently expose only coarse timing (total duration, per-task duration). Missing phase-level timing makes it difficult to identify bottlenecks in planning, synthesis, and evaluation. There is also no breakdown of where time is spent within an LLM call — reasoning, tool execution, and generation are opaque.
Scope:
Design notes:
Acceptance criteria: