Repository navigation
fix(browser): reject promptly when a Worker dies without an error event - #104
Open
TakalaWang wants to merge 1 commit into
Open
TakalaWang wants to merge 1 commit into
TakalaWang wants to merge 1 commit into
Conversation
A browser Worker killed without an error event (the browser terminating it, or terminate() from outside) went unnoticed: interact and run waited for the wall limit and reported wall-time-limit, blaming the student's program, and a killed compiler or stage Worker waited for its build or stage timeout. Every forge Worker starts from createModuleWorker's blob bootstrap, which now takes a uniquely named Web Lock for the Worker's lifetime as it starts and reports the lock's name to the parent. The parent requests the same lock; it is granted only once the Worker's context is destroyed. If that happens before the owner called terminate() (which createModuleWorker wraps to cancel the watch), onModuleWorkerLost listeners run. BrowserRunner, BrowserCompiler, the interactive side Workers and the compiler stage Workers feed them into their existing crash paths, so the engine reports runner-failure or compiler-failure. Listeners are called directly because WebKit drops events dispatched on a Worker after terminate(). Without Web Locks nothing changes. Interactive pipe waits now use 100 ms Atomics.wait slices. WebKit never stops a terminated Worker blocked in an untimed Atomics.wait, so a killed side Worker was undetectable there. The strict-CSP suite gains a liveness section that kills side, runner, compiler and rustc stage Workers mid-operation in each browser and checks prompt rejection and recovery, plus a no-false-positive run of 20 runs, 5 interactive sessions and a build. Detection: Firefox immediately, Chromium within about 2 s for busy Workers (its forced-termination delay), WebKit within about 80 ms except for a Worker spinning in pure Wasm, which WebKit does not stop until it returns to JavaScript. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TakalaWang
force-pushed
the
fix/browser-worker-liveness
branch
from
October 9, 2026 11:07
a74974e to
2d31ac8
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
On 0.2.4, when a browser Worker dies without an
errorevent (the browser terminates it, or something callsterminate()from outside), forge does not notice:interactwaiting until the wall limit, and the session is then reported aswall-time-limit, which blames the student's program;runandinteract;Uncaught errors already reject correctly (
The interactive interactor Worker crashed: …). This PR does the same for silent deaths. It is independent of #102 and the stacked WebKit PR; it touches only TypeScript and the strict-CSP suite.Change
Liveness through Web Locks (
src/runtime/module-worker.ts). Every forge Worker is created bycreateModuleWorkerthrough its blob bootstrap. That bootstrap now:wasm-oj-worker-<uuid>) as it starts and holds it for the Worker's whole life;The parent consumes that message (owners never see it) and requests the same lock. The browser grants it only once the Worker's context is destroyed. If that happens before the owner called
worker.terminate(), the Worker died silently, and itsonModuleWorkerLost(worker, listener)listeners are called.createModuleWorkerwraps the instance'sterminate()so an owner's own termination cancels the watch first; there are no false positives. Without Web Locks, nothing changes.Owners call
onModuleWorkerLostand pass the error to their existing crash path:BrowserRunnerandBrowserCompiler: reject pending requests and install a replacement Worker, as for anerrorevent;runner.worker.ts): reject withThe interactive <role> Worker crashed: …(RUNTIME_ERROR);isolated-stage.ts: fail the stage.The engine surfaces these as
WasmOjErrorrunner-failureorcompiler-failure, a system error and never a limit.Listeners are called directly rather than through a synthetic
errorevent on the Worker, because WebKit drops every event dispatched on a Worker object afterterminate()was called on it. Chromium and Firefox still deliver them.Bounded pipe waits (
src/runtime/interactive-pipe.ts).Atomics.waitin the interactive pipes now waits in 100 ms slices. WebKit does not stop a terminated Worker that is blocked in an untimedAtomics.wait(in a probe it was still alive 4 s later). A Worker in timed waits stops within about 20 ms ofterminate(). Without this change, WebKit could not detect a killed side Worker at all. The wakeups only re-check the pipe state, so behaviour is unchanged.Docs:
docs/library-contract.md(browser execution boundary) andCHANGELOG.md.How quickly each browser detects a kill
Web Locks release semantics, measured with plain Workers (not forge) by terminating a Worker that holds a lock:
Atomics.waitloopAtomics.waitChromium terminates a busy Worker gracefully and forces it only after its 2 s termination delay.
Limitation (WebKit): a Worker killed while it runs pure Wasm keeps its lock until it returns to JavaScript. A
runof a compute-bound program in WebKit therefore still ends at its wall limit when its runner Worker is killed mid-computation (liveness-runner-run-computebelow). A program that makes WASI calls (I/O,sched_yield) returns to JavaScript and is detected at once. A WebKit Worker that does not stop when terminated looks exactly like a running one from outside, so nothing better is available.Verification
Real-browser tests, new
livenesssection inscripts/verify-browser-csp.mjs. The page records the Workers it creates, so a test kills a Worker with the nativeWorker.prototype.terminate, bypassing the owner's wrapper the way an outside kill would. Nested Workers (interactive sides in the runner, the rustc stage in the compiler) are killed from inside their parent through Playwright's Worker handle. Each case then checks that the engine recovered with a normalrun.interact(blocked on input)interactinteractrun(program callssched_yield)run(pure compute)wall-time-limitat 15 sEach rejection reads like
WasmOjError: The interactive contestant Worker crashed: The wasm-oj-interactive-contestant Worker stopped without reporting an error.The same tests againstmain(Chromium): every interactive andrunkill ends withwall-time-limitat the wall (18.5 s after the kill), the compiler kill withCompiler request exceeded the 60000 ms browser boundary, and the rustc stage kill withThe persistent rustc stage exceeded 185000 ms. Without the bounded pipe waits, WebKit's two side-Worker cases end withwall-time-limit18.5 s after the kill.Unit tests:
module-worker.test.tscovers the bootstrap source, lost-listener delivery (including a listener registered after the loss), no report after the owner's ownterminate(), and no Web Locks.client-lifecycle.test.tscovers a running execution and a build rejecting at once when their Worker is lost.Commands:
pnpm run ci:verify(typecheck, lint, tests, build): 982 tests pass.python-16mb-recursion_error, which also fails onmainon this machine (WebKit 26.5 runs out of native stack first).SIGSEGVinJSC::SharedArrayBufferContents::grow←JSC::Wasm::Memory::growShared, in a Worker running the Wasmer SDK's shared-memory module while it handled its first, bootstrap-replayed message.maincrashes the same way (2 of 3 full WebKit runs on this machine, same stack, same point in the suite), so I left it out of scope.Risks
wasm-oj-worker-prefix for its own locks would collide; that seems unlikely.Workerobject'sterminate()would see it reported as lost. forge has no such path.🤖 Generated with Claude Code