fix(acp): close PromptChannelGuard reload race and stale-event leak - #6672
Merged
Conversation
PromptChannelGuard::drop restored output_rx unconditionally on every exit path, keyed only by session_id looked up fresh at drop time. If a session was reloaded/resumed mid-turn (do_load_session/do_resume_session insert a fresh SessionEntry over the same id after a prior close/delete), the guard's drop could clobber the new entry's live output_rx with a dead receiver from the superseded turn. Stamp SessionEntry with a monotonically increasing generation from make_session_entry, the sole construction site for all session-creation paths. PromptChannelGuard captures the generation at acquisition and skips the restore in Drop if the entry's current generation no longer matches. Separately, a receiver restored after task abort/cancel could carry LoopbackEvents the still-alive agent loop had already queued (or would still queue after Drop returns, e.g. a second Flush after an await point), leaking into the next prompt's drain_agent_events and causing a spurious immediate EndTurn. Drain the receiver in both Drop (cheap early filter) and acquire_prompt_channels (under the sessions lock, right after output_rx.take() succeeds, closing the full inter-turn window rather than one instant). Closes #6666 Closes #6667
bug-ops
enabled auto-merge (squash)
July 28, 2026 00:41
7 tasks
bug-ops
added a commit
that referenced
this pull request
Aug 16, 2026
* docs(readme): sync crate READMEs with commits since v0.22.3 Reconciles all 24 changed crate READMEs against the actual shipped implementation for the v0.22.3..HEAD range: new subsystems (risk-chain detection, capability scoping, plugin dependency graph, session spawn cap), several pre-existing factual errors unrelated to this release (inverted file-sandbox precedence, fabricated MCP config keys, wrong anomaly-detector defaults, stale trust-level names), and terminology/ API renames that had drifted out of sync with the code. * docs(specs): reconcile spec drift for commits since v0.22.3 Closes drift left after the skill-quarantine trust fixes (#6701, #6702, #6706, #6707, #6713), the subagent session-wide spawn cap (#6545), four post-ACP-2.0.0-migration bugfixes (#6660, #6665, #6672, #6684), the mention-picker and TUI interrupt-hint updates, the MAX_RETRY_SECS compile-time bound, the sanitizer secret-shape masking extension, the tracing-guard-flush invariants, and the VigilGate per-process pattern-compile fix. Updates specs/README.md's index to match. * docs(book): sync user docs with commits since v0.22.3 Updates the TUI keybindings and mention-picker pages for the new Ctrl+C semantics, the inline @ mention picker, and the input separator's busy indicator; documents the new [tools.shell] risk_chain_window_turns config key; corrects the ACP protocol version reference (was stale at 0.11.1); bumps the sub-agent frontmatter breaking-change note to v0.22.4. * fix(serve): give build_combined_deps_wires_policy_gate test a dedicated stack cargo nextest run --features full could crash with a stack overflow (SIGABRT) on serve::agent_factory::tests::build_combined_deps_wires_policy_gate_through_to_session_agent. Same defect class already fixed once in this file for issue #6699: building a full Agent under --features full's unboxed AnyProvider variants (Candle/Gonka/Cocoon) reaches the same VigilGate::try_new stack depth that overflows the default 2 MiB test-thread stack in an unoptimized build. The #6699 fix only wrapped the one test it was filed against, leaving this one - added in PR #6007, unrelated to any change in this release - unprotected. CI's test job never caught it because it runs the curated feature set, not full, so the deeper AnyProvider frames never materialize there. Runs the test body on a dedicated 32 MiB-stack thread instead of directly under #[tokio::test], reusing the existing TEST_THREAD_STACK_SIZE constant. * release: prepare v0.22.4 Bump version across the workspace, finalize the CHANGELOG.md [0.22.4] section, refresh the README tests badge, and re-accept the splash-screen snapshots (embed the version string).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SessionEntrynow carries agenerationstamp (assigned inmake_session_entry, the sole construction site for new/load/fork/resume).PromptChannelGuardcaptures it at acquisition and skips restoringoutput_rxinDropif the entry's current generation no longer matches — closes the race where a session reloaded/resumed mid-turn (after a prior close/delete) had its fresh, liveoutput_rxclobbered by a stale receiver from the superseded turn.acquire_prompt_channelsnow drains any events already queued on the receiver (under the sessions lock, right afteroutput_rx.take()succeeds), in addition to the existingDrop-time drain. A single point-in-time drain inDroponly catches events queued at that instant; the agent loop can keep emitting afterDropreturns (e.g. a secondFlush), so draining at acquire time closes the whole inter-turn window instead of one snapshot — preventing staleLoopbackEvents from leaking into the next prompt'sdrain_agent_eventsand causing a spurious immediateEndTurn.Both bugs were found during adversarial review of PR #6665 (fix for #6661) and build on top of the
PromptChannelGuardit introduced.Closes #6666
Closes #6667
Test plan
cargo nextest run --config-file .github/nextest.toml --workspace --features "desktop,ide,server,chat,pdf,scheduler" --lib --bins— 15116 passed, 0 failedcargo nextest run --config-file .github/nextest.toml --workspace --features "desktop,ide,server,chat,pdf,scheduler" --lib --bins -E 'package(zeph-acp)'— 197/197 passed, including 2 new regression tests for zeph-acp: PromptChannelGuard restore can clobber a reloaded session's live output_rx #6666/zeph-acp: restored output_rx after cancel/abort may carry stale queued events into the next prompt #6667 plus a double-Flushregression test for the acquire-time draincargo +nightly fmt --checkcargo clippy --profile ci --workspace --all-targets --features "desktop,ide,server,chat,pdf,scheduler,testing" -- -D warningsRUSTFLAGS="-D warnings" RUSTDOCFLAGS="--deny rustdoc::broken_intra_doc_links" cargo doc --no-deps --workspace --features "desktop,ide,server,chat,pdf,scheduler"gitleaks protect --staged --no-banner --redact