Wave execution in Orchestrator::execute drains each wave's FuturesUnordered completely before recomputing ready_tasks(), so a task whose dependencies finish early still waits for the slowest task in the current wave (orchestrator.rs:2648-2717 as of 2026-07-17). ready_tasks() also skips tasks with a Failed dependency, stranding the downstream sub-DAG until replan (types.rs:281).
What instrumented runs showed: intra-wave skew of 12-26% of wave duration on 3-task waves, but realized barrier waste was zero in every harness run because no plan had a dependent waiting behind a partial wave. Coordinator behavior matters here. gpt-4o-mini produced only serial plans; claude coordinators parallelized.
Open question this ticket tracks: whether real workloads produce waves where the barrier actually costs something. That needs per-wave timing, which is #218. Gated on #218.
Wave execution in Orchestrator::execute drains each wave's FuturesUnordered completely before recomputing ready_tasks(), so a task whose dependencies finish early still waits for the slowest task in the current wave (orchestrator.rs:2648-2717 as of 2026-07-17). ready_tasks() also skips tasks with a Failed dependency, stranding the downstream sub-DAG until replan (types.rs:281).
What instrumented runs showed: intra-wave skew of 12-26% of wave duration on 3-task waves, but realized barrier waste was zero in every harness run because no plan had a dependent waiting behind a partial wave. Coordinator behavior matters here. gpt-4o-mini produced only serial plans; claude coordinators parallelized.
Open question this ticket tracks: whether real workloads produce waves where the barrier actually costs something. That needs per-wave timing, which is #218. Gated on #218.