Skip to content

Timing Ph1: record what's already measured #217

Description

@gjanco

planning latency, attempt latencies, drain time → OTel span attributes

Orchestration runs currently expose only coarse timing (total duration, per-task duration). Missing phase-level timing makes it difficult to identify bottlenecks in planning, synthesis, and evaluation. There is also no breakdown of where time is spent within an LLM call — reasoning, tool execution, and generation are opaque.

Scope:

  1. Phase timing on IterationComplete SSE event: planning, execution (wall-clock and aggregate), synthesis, and evaluation durations
  2. Per-call timing breakdown: Track time spent in reasoning vs tool calls vs generation within LLM interactions
  3. Surface timings through existing structures: SSE events, run manifest / artifact files, and OTel spans where tracing is wired

Design notes:

  • Label "start to plan_created" as planning latency (includes retries), not routing latency
  • Define both wall-clock (max of wave) and aggregate compute (sum of tasks) for parallel execution
  • Currently Synthesizing → IterationComplete spans both synthesis and evaluation — needs separate durations

Acceptance criteria:

  • IterationComplete SSE event includes phase timing fields
  • Reasoning/tool-call/generation timing breakdown available per LLM call
  • Phase timing fields documented in docs/streaming-api-guide.md
  • Existing E2E and unit tests pass

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions