Current results are n=1 per arm — direction is strong but confidence is loose.
Task: repeat the large-task A/B (solo@high / solo@max / hybrid), measure with scripts/measure-session.py <session-id>, and add a row/table to docs/AB-RESULTS.md.
Acceptance: at least one additional run per arm, with session ids and the COST-WEIGHTED numbers.
Great way to learn the workflow end to end.
Current results are n=1 per arm — direction is strong but confidence is loose.
Task: repeat the large-task A/B (solo@high / solo@max / hybrid), measure with
scripts/measure-session.py <session-id>, and add a row/table todocs/AB-RESULTS.md.Acceptance: at least one additional run per arm, with session ids and the COST-WEIGHTED numbers.
Great way to learn the workflow end to end.