Skip to content

fix(runtime): add an uncharged metering safepoint so WebKit stops killed Workers - #106

Open
TakalaWang wants to merge 5 commits into
wasm-oj:mainfrom
TakalaWang:fix/webkit-wasm-safepoint
Open

TakalaWang wants to merge 5 commits into
wasm-oj:mainfrom
TakalaWang:fix/webkit-wasm-safepoint

Conversation

@TakalaWang

Copy link
Copy Markdown
Contributor

Stacked on #103 (and so #102) and on #104. This branch is #103's branch with #104 merged into it, plus one commit. Review only the last commit, fix(runtime): add an uncharged metering safepoint so WebKit stops killed Workers. The overlap: #103 changes the same instrumentation pass and the generated runtime, and the WebKit regression tests here use #104's Worker liveness. I'll rebase once those land.

Why

#104 reports a Worker that dies without an error event, but WebKit has one case it can't see. A Worker that is terminated or killed while it runs pure Wasm is not actually stopped, so its Web Lock is never released and the session still ends at the wall limit.

What I established

JavaScriptCore acts on termination only at JavaScript checkpoints, never inside Wasm. Time from worker.terminate() until the Worker's Web Lock is released, plain Wasm loops outside forge, 6 s cap:

Loop body WebKit 26.5 Chromium 151 Firefox 153
pure Wasm never 2.0 s 0 ms
Wasm→Wasm call each iteration (prologues) never 2.0 s 0 ms
empty JS import each iteration 5 ms 2.0 s 1 ms
empty JS import every 4,096 iterations 3.9 s 2.0 s 0 ms
empty JS import every 65,536 / 2^20 / 2^24 iterations never 2.0 s 0 ms
JS import every 2^20 or 2^24 iterations that calls Atomics.wait(cell, 0, mismatch, 0) 0–3 ms 2.0 s 0 ms
same with Atomics.wait(cell, 0, value, 0) (timeout 0) 1–3 ms 2.0 s 0 ms
same with a tiny JS loop instead never 2.0 s 0 ms

The macOS 26.6 jsc shell's --watchdog termination, which uses the same VMTraps, behaves the same way: never in pure Wasm or Wasm→Wasm calls, only with a JS call on every iteration. An ordinary JS call boundary is caught only by chance, and Atomics.wait is a reliable checkpoint. Chromium force-terminates a busy Worker after its 2 s grace whatever it runs; Firefox stops it at once.

How long a killed Worker keeps running in WebKit (after #103): until the program returns to JavaScript, that is until it makes a WASI call or exhausts its instruction budget. Default budget (1e10), time to instruction-limit:

Program WebKit Chromium Firefox
for(;;); 0.21 s 1.22 s 2.76 s
volatile counter loop 0.61 s 0.88 s 2.32 s
division loop 0.38 s 1.04 s 1.63 s
recursive fib(30) in a loop 0.68 s 1.02 s 2.20 s
64 MiB pointer chase 37.6 s 37.1 s 40.2 s
64 KiB memcpy loop 0.20 s 1.22 s 2.79 s

So in WebKit a terminated runner keeps a core busy for a fraction of a second up to about 40 s, and a liveness check waits that long.

Real kills. A Worker cannot die alone in WebKit except through terminate(): in a Worker, growing a Wasm memory past the limit throws RangeError at 3,840 MiB and the page lives; plain ArrayBuffers reached 6 GiB without failing; and kill -9 of the WebContent process takes the whole page down (Playwright's crash event, Target crashed), with nothing left to report. So the case to fix is a Worker that forge, the embedding page or a parent Worker terminates while it computes.

Change

The meter gets an uncharged safepoint (runtime-core, final instrumentation pass):

  • The instrumented module imports wasm_oj_metering.safepoint (added after the existing function imports; defined function indices shift by one) and gets a private threshold global, initially budget - 2^20.
  • The gas charge that opens each function and each loop body is inlined: subtract the cost, then compare the counter with the threshold. Below it, a cold function traps exactly as radix's gas function does if the counter is negative, and otherwise calls the safepoint and lowers the threshold to counter - 2^20 (never below 0). Every other block keeps radix's gas function unchanged.
  • A call anywhere inside a hot loop made the whole loop slower in JavaScriptCore and V8, even though it never ran (a first version that called a refill function from radix's trap branch cost 20 to 26% on realistic code; replacing that call with unreachable removed the cost). So a loop branches out to reach the safepoint: each loop becomes block (loop (block (loop ...) br 2) call slow; refund; br 0) unreachable, every branch out of the loop skips the three new labels, and the cost charged before leaving is refunded so the re-entry charges it once. Functions that use exception or GC branch instructions (which radix rejects today anyway) and loops with parameters keep the call inside the loop.
  • The counter, its export and remaining_points are untouched, so costs and the exhaustion point are identical: the same programs report the same costs before and after, from 14,013 to 4,244,150,875 units, and a run with budget = cost exits while cost - 1 reports instruction-limit.
  • Natively, safepoint does nothing. In the browser it calls Atomics.wait on a never-matching value, which returns at once and is where JavaScriptCore acts on terminate(). It is a no-op where SharedArrayBuffer is unavailable.

A terminated runner or interactive side Worker running metered code now stops within 2^20 units in WebKit (about 40 µs of compute in the benchmarks below). #104's liveness check then reports it, and WebKit no longer keeps a zombie Worker computing after a wall-time or cancel termination.

Alternatives considered:

  • A cooperative SharedArrayBuffer cancel flag read at the safepoint. It would only shave Chromium's 2 s grace, because Atomics.wait already makes WebKit honour terminate() and Firefox stops at once. Not needed.
  • A heartbeat counter. It is redundant with fix(browser): reject promptly when a Worker dies without an error event #104's Web Lock liveness now that a killed Worker actually dies, and would need "blocked on a pipe" bookkeeping to avoid false positives.
  • Atomics.wait slices. They already cover a side blocked on a pipe (fix(browser): reject promptly when a Worker dies without an error event #104) but cannot help a program that never waits.
  • An outer watchdog that terminates and recreates the runner. It already exists (the wall-time timer replaces the Worker), but in WebKit the terminated Worker is never stopped, so it cannot reclaim the CPU or tell a dead Worker from a live one.
  • A safepoint call on every gas charge, or from radix's trap branch (a counter/reserve split). Measured 20 to 26% slower on the realistic program in WebKit and Chromium.
  • Safepoints only in JavaScriptCore. Possible (the host would ask for them), but Chromium, Firefox and native show little or no cost on realistic code, so one module for every host is simpler.

Unmetered toolchain code (clang or rustc in the compiler stage) still has no safepoint. A killed compiler stage is reported when it returns to JavaScript (30 ms to 1.2 s measured in WebKit).

Overhead

Median of five runs (after one warm-up), release build, executionDurationMs, two separate sessions each, this branch against the same stack without this commit. Costs are identical in every run (2,500,000,159 and 4,244,150,875).

Tight loop (volatile sum, 3e8 iterations) Realistic C++ (sort, unordered_map, map)
WebKit 26.5 83 → 94–95 ms (+14%) 155 → 172–173 ms (+11%)
Chromium 151 145–146 → 161–163 ms (+11%) 161–162 → 161–162 ms (0%)
Firefox 153 1106–1114 → 1166–1173 ms (+5%) 1233–1235 → 1182–1186 ms (−4%)
Native (wasmer, Cranelift; includes compile) 134–135 → 136–139 ms (+2%) 321–325 → 286 ms (−11%)

The tight loop is the worst case: one iteration is a handful of instructions plus the check.

Verification

  • runtime-core: new unit tests for cost and exhaustion across many intervals; one safepoint call per interval for a loop and for loop-free recursion (the recursion case fails without the function-entry check); branches, br_table, loop results and a loop with parameters giving the same result as the uninstrumented module (fails if the label shift is dropped); a module without imports gaining the import; all runtime-core tests; cargo fmt --check; clippy native and web with -D warnings; pnpm run runtime:check-web.
  • pnpm run ci:verify, and src/server/judge.integration.test.ts with the native runner.
  • Strict-CSP suite in Chromium, Firefox and WebKit, including:
    • liveness-runner-run-compute (kill the runner Worker during a compute-bound run), which in WebKit used to end at wall-time-limit 15 s later;
    • new liveness-interactive-compute (kill a compute-bound interactive contestant side);
    • the compiler and rustc-stage kills;
    • the no-false-positive check.

Strict-CSP suite (Playwright 1.62.1: Chromium 151, Firefox 153, WebKit 26.5): Chromium 121 of 121 and Firefox 121 of 121. The full WebKit run lost its page at interactive-ac-c to the JavaScriptCore shared-memory crash that #105 works around (it happens on main too), so I ran the liveness section in WebKit on its own twice, and the full suite on this branch combined with #105 (below).

Time from the kill until the operation rejects (ms):

Case Chromium Firefox WebKit before WebKit after (2 runs)
Kill runner during a compute-bound run 2004 0 wall-time-limit, 13,502 ms after the kill 0 / 0
Kill the compute-bound interactive contestant 2006 0 wall-time-limit, 18,558 ms after the kill 0 / 0
Kill runner during interact 0 4 0 0 / 0
Kill runner during a yielding run 2003 0 0 0 / 0
Kill interactive contestant / interactor (blocked) 2002 / 2010 0 / 0 74 / 70 94, 87 / 82, 71
Kill compiler / rustc stage 440 / 720 0 / 0 0 / 1155 0, 0 / 26, 0
20 runs + 5 interactive sessions + 1 build, no false positive pass pass pass pass

"Before" is the same stack without this commit. Combined with #105, the full WebKit suite passes everything except the pre-existing python-16mb-recursion_error (it fails the same way on main).

runtime-core_bg.wasm is rebuilt with pnpm run runtime:build and the identity pins are refreshed, which changes cost profile identifiers (not cost values).

Risks

  • The instrumented module gains one import, one global and one function, and each loop gains three labels. Hosts that instantiate WASM-OJ-instrumented modules themselves must provide wasm_oj_metering.safepoint; forge's runtimes do.
  • Metered code is up to 14% slower in WebKit (tight loops), see Overhead.
  • Atomics.wait is reached about every 2^20 units; on a main thread it throws, which is ignored.

🤖 Generated with Claude Code

TakalaWang and others added 5 commits October 9, 2026 16:55
…iling with EPIPE

Closes wasm-oj#101. When one side of an interact session exited or closed its stdin, its peer's next write failed with EPIPE on both hosts. A CPython interactor that replied to a contestant that had already exited died with BrokenPipeError and exit 120, and because the stdout wrapper records into the transcript before writing to the pipe, each retried flush added the reply again (five copies). Judges that run both programs natively keep draining each program's output until it exits, so the interactor reads EOF and exits with its own verdict.

InteractiveOutput (server) and StreamOutput (browser side Workers) now treat a broken pipe as a successful write: the bytes are already in the transcript and are dropped. Reads from a side that has exited still return the remaining buffered bytes and then EOF. The transcript records each byte a side writes exactly once, up to its output limit, including bytes written after the peer exited, so it depends only on what the writer wrote.

Tests cover both hosts: runtime-core unit tests for writes after either side exits and EOF after buffered bytes, server integration tests with C and CPython interactors exiting 42/43 after replying to an exited contestant, and strict-CSP browser fixtures for the same cases plus the reply-before-EOF race. runtime-core_bg.wasm is rebuilt and the runtime identity pins refreshed, which changes cost profiles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…timizes them

In WebKit an empty C++ for(;;); took about 25 s to reach the default instruction budget (Chromium 1.3 s, Firefox 2.9 s), so it usually hit the wall limit first and was reported as a wall-time TLE.

The meter is not optimized away; the trap does arrive, just late. JavaScriptCore (Safari 26, also the macOS 26.6 jsc shell) never enters optimized code inside a loop of a function that has no parameters and no locals. With --verboseOSR the baseline tier logs "Consider OSREntryPlan" about once per iteration (18 million times in 2e8 cost units) although the OMG entry code exists, so every iteration takes the tier-up slow path: about 14 times slower than BBQ alone (--useOMGJIT=false) and 100 times slower than the same loop in a function with one unused local. The forge meter makes no difference; any loop body behaves the same way.

The final instrumentation pass (MeterInitializer, which already re-encodes the module to set the budget) now gives every function that contains a loop but has no parameters or locals one unused i32 local. It changes neither behaviour nor cost. WebKit now stops the empty loop at the budget in about 0.2 s; Chromium and Firefox are unchanged.

A unit test covers which functions receive the local, and the strict-CSP suite gains empty-loop fixtures for run and interact that must reach instruction-limit before a 15 s wall limit. runtime-core_bg.wasm is rebuilt and the runtime identity pins refreshed, which changes cost profiles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A browser Worker killed without an error event (the browser terminating it, or terminate() from outside) went unnoticed: interact and run waited for the wall limit and reported wall-time-limit, blaming the student's program, and a killed compiler or stage Worker waited for its build or stage timeout.

Every forge Worker starts from createModuleWorker's blob bootstrap, which now takes a uniquely named Web Lock for the Worker's lifetime as it starts and reports the lock's name to the parent. The parent requests the same lock; it is granted only once the Worker's context is destroyed. If that happens before the owner called terminate() (which createModuleWorker wraps to cancel the watch), onModuleWorkerLost listeners run. BrowserRunner, BrowserCompiler, the interactive side Workers and the compiler stage Workers feed them into their existing crash paths, so the engine reports runner-failure or compiler-failure. Listeners are called directly because WebKit drops events dispatched on a Worker after terminate(). Without Web Locks nothing changes.

Interactive pipe waits now use 100 ms Atomics.wait slices. WebKit never stops a terminated Worker blocked in an untimed Atomics.wait, so a killed side Worker was undetectable there.

The strict-CSP suite gains a liveness section that kills side, runner, compiler and rustc stage Workers mid-operation in each browser and checks prompt rejection and recovery, plus a no-false-positive run of 20 runs, 5 interactive sessions and a build. Detection: Firefox immediately, Chromium within about 2 s for busy Workers (its forced-termination delay), WebKit within about 80 ms except for a Worker spinning in pure Wasm, which WebKit does not stop until it returns to JavaScript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…led Workers

JavaScriptCore acts on Worker.terminate() only at JavaScript checkpoints,
never inside Wasm, so a terminated runner or interactive side kept
computing until its instruction budget ran out and Web Lock liveness
could not see the kill.

The charge that opens each function and loop body is now inlined and
compared with a threshold. Every 2^20 units a cold function calls the
imported wasm_oj_metering.safepoint, which is a no-op natively and an
Atomics.wait that returns at once in browsers. Loops reach it through
a branch out of the loop, because a call inside a hot loop slowed the
whole loop in JavaScriptCore and V8. Costs and the exhaustion point are
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant