Skip to content

reload during an in-flight turn kills the turn with "This extension ctx is stale after session replacement or reload" #1172

Description

@solidlime

Summary

Calling reload on a session whose turn is in flight invalidates that session's extension context (ctx) mid-turn. When the in-flight turn resumes (e.g. right after a long tool call returns), the next turn boundary touches the dead ctx and the turn dies with:

This extension ctx is stale after session replacement or reload. Do not use a captured pi or command ctx after ctx.newSession(), ctx.fork(), ctx.switchSession(), or ctx.reload(). ...

The failure is recorded as an assistant message with stopReason: "error", content: [], usage all zero, durationMs: 0 — i.e. it dies before the model call. For a subagent this surfaces as pi-web:subagent-result { "status": "failed" }, so the user just sees "the session/subagent stopped".

Impact

Any long-running tool call (a big build, a long bash, an 84-second grep, a slow MCP call) is a window where a reload will kill the turn. Because pi-web hosts several sessions in one process, a reload can also take out sibling sessions/subagents that are mid-turn (observed: a subagent and its parent both failed in the same second).

reload is reachable from the UI (/reload, settings/extension changes) and from the extension command API, so this fires without the user intending to interrupt a running turn.

Deterministic reproduction

  1. Start a background subagent whose first turn runs a long tool call:
    bash: sleep 150 && echo SLEEP_DONE
  2. While it is sleeping, trigger a reload on that same session from outside:
    curl -X POST http://127.0.0.1:8787/api/agent/<victimSessionId> \
         -H 'Content-Type: application/json' -d '{"type":"reload"}'
    → returns {"success":true,"data":{"success":true}} — it is not rejected even though the session is running.
  3. Observe: when sleep returns and the loop advances to the next turn, the turn dies with stopReason: "error", durationMs: 0.

Observed timeline (victim session 01a12777-38f3-75a4-b5d4-64f642a7651f):

20:18:09.623  assistant stop=toolUse  → bash "sleep 150 && echo SLEEP_DONE" starts
20:18:56      POST {"type":"reload"} → 20:19:01 {"success":true}
20:19:01.037  session_start fired (reload re-activation; old runner invalidated)
20:20:41.907  assistant stopReason=error, durationMs=0, content=[], usage all zero
20:20:41.990  pi-web:subagent-result status:"failed"

The death lands exactly when the 150s tool call returns.

Root cause (call path)

reload replaces the extension runtime and invalidates the old one, but it does not wait for / guard against an in-flight turn:

  • pi-web bundle .next/server/chunks/6429.js
    • case "clone" → if (this.isSessionRunningForReplacement()) throw Error("Cannot clone while the session is running") (guarded)
    • case "fork" → computes isSessionRunningForReplacement() and skips the post-replacement shutdown when running (guarded)
    • case "reload" → await this.inner.reload() with no isSessionRunningForReplacement() check (unguarded)
  • bundled core @earendil-works/pi-coding-agent@1.1.0
    • dist/core/agent-session.js — reload() calls oldRunner.invalidate() (and dispose() also calls this._extensionRunner.invalidate(...))
    • dist/core/extensions/runner.js — invalidate(message) sets staleMessage; assertActive() then throws on any later ctx use

So the in-flight turn keeps holding the pre-reload ctx; the first extension-hook/ctx use after the tool returns throws.

Suggested fix

  1. Add the missing guard (primary). Make case "reload" behave like clone/fork: refuse or defer the reload while isSessionRunningForReplacement() is true (streaming / bash / compacting / …), or queue it until the turn settles. This alone removes the failure.
  2. Optionally, make the core guard non-fatal for a turn: if an extension's ctx is stale during turn setup, skip that hook and continue the turn instead of failing the whole turn.

Environment

  • pi-web 0.11.1
  • bundled @earendil-works/pi-coding-agent 1.1.0
  • Linux container, Node (pi-web as PID 1 on port 8787)
  • provider routing bifrost → opencode-go/mimo-v2.6-pro (not related; verified normal)

Notes

  • This is not caused by a third-party extension. pi-lens guards its handlers (probeCtxActive + isStaleExtensionCtxError catch) and reproduces nothing on its own; the minimal repro above uses no third-party extension behavior beyond a plain sleep.
  • Related latent hazard worth fixing separately: pi-lens keeps a module-global latestEventCtx and exposes it via hostPorts.getContext (pi-lens/dist/index.js:143003, 143193); its own comment at 143197-143206 notes the global can belong to "a SIBLING activation after a replacement".
  • Evidence bundle (VERDICT + victim session JSONL + driver log) can be attached on request.

Activity

  1. solidlime commented on Oct 10, 2026

    @solidlime
    Author

    Adding evidence that broadens the scope — the same failure also happens without reload.

    In the original occurrence the turn died at 20:04:53.107Z with a background subagent that had been running fine for ~20 minutes. At that moment there was no session_start re-activation in the extension logs (a reload does emit one — verified in the minimal repro above, where the reload produced session_start fired at 20:19:01.037Z). So the runner was invalidated by a session replacement (teardownCurrent → dispose() → invalidate), not by reload().

    Also in that occurrence both the subagent and its parent session failed within the same second (20:04:53.107Z and 20:04:53.405Z), each with stopReason: "error", durationMs: 0.

    So the guard should not be added only to case "reload". Every path that invalidates a runner while a turn may still be in flight is affected:

    • reload() → oldRunner.invalidate() (no isSessionRunningForReplacement() guard — proven above)
    • switchSession / newSession / fork → teardownCurrent() → session.dispose() → invalidate() (clone guards; fork only guards isBashRunning + skips the post-replacement shutdown)

    Two possible fixes, either of which removes the failure:

    1. defer or reject every replacement/reload while isSessionRunningForReplacement() is true, and/or
    2. make the stale-ctx guard non-fatal during turn setup: skip the affected extension hook and continue the turn instead of failing the whole turn.

    Environment recap: pi-web 0.11.1, bundled core 1.1.0, Linux, single process hosting multiple concurrent sessions (globalThis.__piSessions).

  2. solidlime commented on Oct 11, 2026

    @solidlime
    Author

    Fixed by #1173 — the remaining invalidation paths now go through the same isSessionRunningForReplacement() guard that clone/fork already used:

    • extension-API reload and extension-ctx reload
    • the cwd sweep (destroyRpcSessionsForCwd), which shut down every session in the target cwd including running ones
    • DELETE /api/sessions/[id], which shut down the deleted session and its subagent descendants even while they were mid-turn

    The PR also covers a related path found while reproducing this one: a background subagent outlives its parent's turn, so the parent session was eligible for idle eviction while the run was still active. start()/resume() now hold a session-liveness provider for the parent for the run's lifetime.

    Verified against a local 0.11.1 build: the reload in the reproduction above is now rejected and the in-flight turn completes; idle sessions still reload and delete normally.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions