Skip to content

fix(browser): reject promptly when a Worker dies without an error event - #104

Open
TakalaWang wants to merge 1 commit into
wasm-oj:mainfrom
TakalaWang:fix/browser-worker-liveness
Open

TakalaWang wants to merge 1 commit into
wasm-oj:mainfrom
TakalaWang:fix/browser-worker-liveness

Conversation

@TakalaWang

@TakalaWang TakalaWang commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Why

On 0.2.4, when a browser Worker dies without an error event (the browser terminates it, or something calls terminate() from outside), forge does not notice:

  • a killed interactive side Worker leaves interact waiting until the wall limit, and the session is then reported as wall-time-limit, which blames the student's program;
  • a killed runner Worker does the same for run and interact;
  • a killed compiler Worker or compiler stage Worker waits for the build or stage timeout (up to two minutes).

Uncaught errors already reject correctly (The interactive interactor Worker crashed: …). This PR does the same for silent deaths. It is independent of #102 and the stacked WebKit PR; it touches only TypeScript and the strict-CSP suite.

Change

Liveness through Web Locks (src/runtime/module-worker.ts). Every forge Worker is created by createModuleWorker through its blob bootstrap. That bootstrap now:

  1. takes a uniquely named Web Lock (wasm-oj-worker-<uuid>) as it starts and holds it for the Worker's whole life;
  2. sends the lock name to the parent once the lock is held. The request does not delay loading the Worker module.

The parent consumes that message (owners never see it) and requests the same lock. The browser grants it only once the Worker's context is destroyed. If that happens before the owner called worker.terminate(), the Worker died silently, and its onModuleWorkerLost(worker, listener) listeners are called. createModuleWorker wraps the instance's terminate() so an owner's own termination cancels the watch first; there are no false positives. Without Web Locks, nothing changes.

Owners call onModuleWorkerLost and pass the error to their existing crash path:

  • BrowserRunner and BrowserCompiler: reject pending requests and install a replacement Worker, as for an error event;
  • interactive side Workers (runner.worker.ts): reject with The interactive <role> Worker crashed: … (RUNTIME_ERROR);
  • compiler stage Workers (rustc, Go, Java) in isolated-stage.ts: fail the stage.

The engine surfaces these as WasmOjError runner-failure or compiler-failure, a system error and never a limit.

Listeners are called directly rather than through a synthetic error event on the Worker, because WebKit drops every event dispatched on a Worker object after terminate() was called on it. Chromium and Firefox still deliver them.

Bounded pipe waits (src/runtime/interactive-pipe.ts). Atomics.wait in the interactive pipes now waits in 100 ms slices. WebKit does not stop a terminated Worker that is blocked in an untimed Atomics.wait (in a probe it was still alive 4 s later). A Worker in timed waits stops within about 20 ms of terminate(). Without this change, WebKit could not detect a killed side Worker at all. The wakeups only re-check the pipe state, so behaviour is unchanged.

Docs: docs/library-contract.md (browser execution boundary) and CHANGELOG.md.

How quickly each browser detects a kill

Web Locks release semantics, measured with plain Workers (not forge) by terminating a Worker that holds a lock:

Worker state when killed Chromium 151 Firefox 153 WebKit 26.5
Idle (awaiting messages) 0 ms 0 ms 0 ms
Blocked in timed Atomics.wait loop ~2000 ms 0 ms ~20 ms
Blocked in untimed Atomics.wait ~2000 ms 0 ms never, until woken
Spinning in Wasm ~2000 ms 0 ms never, until it returns to JS

Chromium terminates a busy Worker gracefully and forces it only after its 2 s termination delay.

Limitation (WebKit): a Worker killed while it runs pure Wasm keeps its lock until it returns to JavaScript. A run of a compute-bound program in WebKit therefore still ends at its wall limit when its runner Worker is killed mid-computation (liveness-runner-run-compute below). A program that makes WASI calls (I/O, sched_yield) returns to JavaScript and is detected at once. A WebKit Worker that does not stop when terminated looks exactly like a running one from outside, so nothing better is available.

Verification

Real-browser tests, new liveness section in scripts/verify-browser-csp.mjs. The page records the Workers it creates, so a test kills a Worker with the native Worker.prototype.terminate, bypassing the owner's wrapper the way an outside kill would. Nested Workers (interactive sides in the runner, the rustc stage in the compiler) are killed from inside their parent through Playwright's Worker handle. Each case then checks that the engine recovered with a normal run.

Case Chromium Firefox WebKit
Kill contestant side Worker during interact (blocked on input) rejects, 2.0 s rejects, 0 ms rejects, 71 ms
Kill interactor side Worker during interact 2.0 s 0 ms 66 ms
Kill runner Worker during interact 0 ms 0 ms 0 ms
Kill runner Worker during run (program calls sched_yield) 2.0 s 0 ms 0 ms
Kill runner Worker during run (pure compute) 2.0 s 0 ms not detectable: wall-time-limit at 15 s
Kill compiler Worker during a C++ build 0 ms 0 ms 0 ms
Kill rustc stage Worker during a Rust build 0.9 s 0 ms 0–1.1 s
No false positives: 20 runs, 5 interactive sessions, 1 build pass pass pass

Each rejection reads like WasmOjError: The interactive contestant Worker crashed: The wasm-oj-interactive-contestant Worker stopped without reporting an error. The same tests against main (Chromium): every interactive and run kill ends with wall-time-limit at the wall (18.5 s after the kill), the compiler kill with Compiler request exceeded the 60000 ms browser boundary, and the rustc stage kill with The persistent rustc stage exceeded 185000 ms. Without the bounded pipe waits, WebKit's two side-Worker cases end with wall-time-limit 18.5 s after the kill.

Unit tests: module-worker.test.ts covers the bootstrap source, lost-listener delivery (including a listener registered after the loss), no report after the owner's own terminate(), and no Web Locks. client-lifecycle.test.ts covers a running execution and a build rejecting at once when their Worker is lost.

Commands:

  • pnpm run ci:verify (typecheck, lint, tests, build): 982 tests pass.
  • Full strict-CSP suite including the new section: Chromium 113 of 113, Firefox 113 of 113, WebKit 112 of 113. The WebKit failure is python-16mb-recursion_error, which also fails on main on this machine (WebKit 26.5 runs out of native stack first).
  • Pre-existing WebKit crash, not caused by this PR: in several full WebKit runs the WebContent process crashed with SIGSEGV in JSC::SharedArrayBufferContents::grow ← JSC::Wasm::Memory::growShared, in a Worker running the Wasmer SDK's shared-memory module while it handled its first, bootstrap-replayed message. main crashes the same way (2 of 3 full WebKit runs on this machine, same stack, same point in the suite), so I left it out of scope.

Risks

  • Every forge Worker now holds one Web Lock, released when the Worker dies. Lock names are unique per Worker and live under the page's origin. A host page that uses the wasm-oj-worker- prefix for its own locks would collide; that seems unlikely.
  • An owner that terminates a Worker some other way than the Worker object's terminate() would see it reported as lost. forge has no such path.
  • The interactive pipes wake ten times a second while blocked, which is negligible.

🤖 Generated with Claude Code

A browser Worker killed without an error event (the browser terminating it, or terminate() from outside) went unnoticed: interact and run waited for the wall limit and reported wall-time-limit, blaming the student's program, and a killed compiler or stage Worker waited for its build or stage timeout.

Every forge Worker starts from createModuleWorker's blob bootstrap, which now takes a uniquely named Web Lock for the Worker's lifetime as it starts and reports the lock's name to the parent. The parent requests the same lock; it is granted only once the Worker's context is destroyed. If that happens before the owner called terminate() (which createModuleWorker wraps to cancel the watch), onModuleWorkerLost listeners run. BrowserRunner, BrowserCompiler, the interactive side Workers and the compiler stage Workers feed them into their existing crash paths, so the engine reports runner-failure or compiler-failure. Listeners are called directly because WebKit drops events dispatched on a Worker after terminate(). Without Web Locks nothing changes.

Interactive pipe waits now use 100 ms Atomics.wait slices. WebKit never stops a terminated Worker blocked in an untimed Atomics.wait, so a killed side Worker was undetectable there.

The strict-CSP suite gains a liveness section that kills side, runner, compiler and rustc stage Workers mid-operation in each browser and checks prompt rejection and recovery, plus a no-false-positive run of 20 runs, 5 interactive sessions and a build. Detection: Firefox immediately, Chromium within about 2 s for busy Workers (its forced-termination delay), WebKit within about 80 ms except for a Worker spinning in pure Wasm, which WebKit does not stop until it returns to JavaScript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant