Summary
Calling reload on a session whose turn is in flight invalidates that session's extension context (ctx) mid-turn. When the in-flight turn resumes (e.g. right after a long tool call returns), the next turn boundary touches the dead ctx and the turn dies with:
This extension ctx is stale after session replacement or reload. Do not use a captured pi or command ctx after ctx.newSession(), ctx.fork(), ctx.switchSession(), or ctx.reload(). ...
The failure is recorded as an assistant message with stopReason: "error", content: [], usage all zero, durationMs: 0 — i.e. it dies before the model call. For a subagent this surfaces as pi-web:subagent-result { "status": "failed" }, so the user just sees "the session/subagent stopped".
Impact
Any long-running tool call (a big build, a long bash, an 84-second grep, a slow MCP call) is a window where a reload will kill the turn. Because pi-web hosts several sessions in one process, a reload can also take out sibling sessions/subagents that are mid-turn (observed: a subagent and its parent both failed in the same second).
reload is reachable from the UI (/reload, settings/extension changes) and from the extension command API, so this fires without the user intending to interrupt a running turn.
Deterministic reproduction
- Start a background subagent whose first turn runs a long tool call:
bash: sleep 150 && echo SLEEP_DONE
- While it is sleeping, trigger a reload on that same session from outside:
curl -X POST http://127.0.0.1:8787/api/agent/<victimSessionId> \
-H 'Content-Type: application/json' -d '{"type":"reload"}'
→ returns {"success":true,"data":{"success":true}} — it is not rejected even though the session is running.
- Observe: when
sleep returns and the loop advances to the next turn, the turn dies with stopReason: "error", durationMs: 0.
Observed timeline (victim session 01a12777-38f3-75a4-b5d4-64f642a7651f):
20:18:09.623 assistant stop=toolUse → bash "sleep 150 && echo SLEEP_DONE" starts
20:18:56 POST {"type":"reload"} → 20:19:01 {"success":true}
20:19:01.037 session_start fired (reload re-activation; old runner invalidated)
20:20:41.907 assistant stopReason=error, durationMs=0, content=[], usage all zero
20:20:41.990 pi-web:subagent-result status:"failed"
The death lands exactly when the 150s tool call returns.
Root cause (call path)
reload replaces the extension runtime and invalidates the old one, but it does not wait for / guard against an in-flight turn:
- pi-web bundle
.next/server/chunks/6429.js
case "clone" → if (this.isSessionRunningForReplacement()) throw Error("Cannot clone while the session is running") (guarded)
case "fork" → computes isSessionRunningForReplacement() and skips the post-replacement shutdown when running (guarded)
case "reload" → await this.inner.reload() with no isSessionRunningForReplacement() check (unguarded)
- bundled core
@earendil-works/pi-coding-agent@1.1.0
dist/core/agent-session.js — reload() calls oldRunner.invalidate() (and dispose() also calls this._extensionRunner.invalidate(...))
dist/core/extensions/runner.js — invalidate(message) sets staleMessage; assertActive() then throws on any later ctx use
So the in-flight turn keeps holding the pre-reload ctx; the first extension-hook/ctx use after the tool returns throws.
Suggested fix
- Add the missing guard (primary). Make
case "reload" behave like clone/fork: refuse or defer the reload while isSessionRunningForReplacement() is true (streaming / bash / compacting / …), or queue it until the turn settles. This alone removes the failure.
- Optionally, make the core guard non-fatal for a turn: if an extension's ctx is stale during turn setup, skip that hook and continue the turn instead of failing the whole turn.
Environment
- pi-web
0.11.1
- bundled
@earendil-works/pi-coding-agent 1.1.0
- Linux container, Node (pi-web as PID 1 on port 8787)
- provider routing
bifrost → opencode-go/mimo-v2.6-pro (not related; verified normal)
Notes
- This is not caused by a third-party extension.
pi-lens guards its handlers (probeCtxActive + isStaleExtensionCtxError catch) and reproduces nothing on its own; the minimal repro above uses no third-party extension behavior beyond a plain sleep.
- Related latent hazard worth fixing separately:
pi-lens keeps a module-global latestEventCtx and exposes it via hostPorts.getContext (pi-lens/dist/index.js:143003, 143193); its own comment at 143197-143206 notes the global can belong to "a SIBLING activation after a replacement".
- Evidence bundle (VERDICT + victim session JSONL + driver log) can be attached on request.
Summary
Calling
reloadon a session whose turn is in flight invalidates that session's extension context (ctx) mid-turn. When the in-flight turn resumes (e.g. right after a long tool call returns), the next turn boundary touches the deadctxand the turn dies with:The failure is recorded as an assistant message with
stopReason: "error",content: [],usageall zero,durationMs: 0— i.e. it dies before the model call. For a subagent this surfaces aspi-web:subagent-result { "status": "failed" }, so the user just sees "the session/subagent stopped".Impact
Any long-running tool call (a big build, a long
bash, an 84-secondgrep, a slow MCP call) is a window where areloadwill kill the turn. Because pi-web hosts several sessions in one process, a reload can also take out sibling sessions/subagents that are mid-turn (observed: a subagent and its parent both failed in the same second).reloadis reachable from the UI (/reload, settings/extension changes) and from the extension command API, so this fires without the user intending to interrupt a running turn.Deterministic reproduction
bash: sleep 150 && echo SLEEP_DONE{"success":true,"data":{"success":true}}— it is not rejected even though the session is running.sleepreturns and the loop advances to the next turn, the turn dies withstopReason: "error",durationMs: 0.Observed timeline (victim session
01a12777-38f3-75a4-b5d4-64f642a7651f):The death lands exactly when the 150s tool call returns.
Root cause (call path)
reloadreplaces the extension runtime and invalidates the old one, but it does not wait for / guard against an in-flight turn:.next/server/chunks/6429.jscase "clone"→if (this.isSessionRunningForReplacement()) throw Error("Cannot clone while the session is running")(guarded)case "fork"→ computesisSessionRunningForReplacement()and skips the post-replacement shutdown when running (guarded)case "reload"→await this.inner.reload()with noisSessionRunningForReplacement()check (unguarded)@earendil-works/pi-coding-agent@1.1.0dist/core/agent-session.js—reload()callsoldRunner.invalidate()(anddispose()also callsthis._extensionRunner.invalidate(...))dist/core/extensions/runner.js—invalidate(message)setsstaleMessage;assertActive()then throws on any laterctxuseSo the in-flight turn keeps holding the pre-reload
ctx; the first extension-hook/ctx use after the tool returns throws.Suggested fix
case "reload"behave likeclone/fork: refuse or defer the reload whileisSessionRunningForReplacement()is true (streaming / bash / compacting / …), or queue it until the turn settles. This alone removes the failure.Environment
0.11.1@earendil-works/pi-coding-agent1.1.0bifrost→opencode-go/mimo-v2.6-pro(not related; verified normal)Notes
pi-lensguards its handlers (probeCtxActive+isStaleExtensionCtxErrorcatch) and reproduces nothing on its own; the minimal repro above uses no third-party extension behavior beyond a plainsleep.pi-lenskeeps a module-globallatestEventCtxand exposes it viahostPorts.getContext(pi-lens/dist/index.js:143003,143193); its own comment at143197-143206notes the global can belong to "a SIBLING activation after a replacement".