Repository navigation
fix(runtime): add an uncharged metering safepoint so WebKit stops killed Workers - #106
Open
TakalaWang wants to merge 5 commits into
Open
TakalaWang wants to merge 5 commits into
TakalaWang wants to merge 5 commits into
Conversation
…iling with EPIPE Closes wasm-oj#101. When one side of an interact session exited or closed its stdin, its peer's next write failed with EPIPE on both hosts. A CPython interactor that replied to a contestant that had already exited died with BrokenPipeError and exit 120, and because the stdout wrapper records into the transcript before writing to the pipe, each retried flush added the reply again (five copies). Judges that run both programs natively keep draining each program's output until it exits, so the interactor reads EOF and exits with its own verdict. InteractiveOutput (server) and StreamOutput (browser side Workers) now treat a broken pipe as a successful write: the bytes are already in the transcript and are dropped. Reads from a side that has exited still return the remaining buffered bytes and then EOF. The transcript records each byte a side writes exactly once, up to its output limit, including bytes written after the peer exited, so it depends only on what the writer wrote. Tests cover both hosts: runtime-core unit tests for writes after either side exits and EOF after buffered bytes, server integration tests with C and CPython interactors exiting 42/43 after replying to an exited contestant, and strict-CSP browser fixtures for the same cases plus the reply-before-EOF race. runtime-core_bg.wasm is rebuilt and the runtime identity pins refreshed, which changes cost profiles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…timizes them In WebKit an empty C++ for(;;); took about 25 s to reach the default instruction budget (Chromium 1.3 s, Firefox 2.9 s), so it usually hit the wall limit first and was reported as a wall-time TLE. The meter is not optimized away; the trap does arrive, just late. JavaScriptCore (Safari 26, also the macOS 26.6 jsc shell) never enters optimized code inside a loop of a function that has no parameters and no locals. With --verboseOSR the baseline tier logs "Consider OSREntryPlan" about once per iteration (18 million times in 2e8 cost units) although the OMG entry code exists, so every iteration takes the tier-up slow path: about 14 times slower than BBQ alone (--useOMGJIT=false) and 100 times slower than the same loop in a function with one unused local. The forge meter makes no difference; any loop body behaves the same way. The final instrumentation pass (MeterInitializer, which already re-encodes the module to set the budget) now gives every function that contains a loop but has no parameters or locals one unused i32 local. It changes neither behaviour nor cost. WebKit now stops the empty loop at the budget in about 0.2 s; Chromium and Firefox are unchanged. A unit test covers which functions receive the local, and the strict-CSP suite gains empty-loop fixtures for run and interact that must reach instruction-limit before a 15 s wall limit. runtime-core_bg.wasm is rebuilt and the runtime identity pins refreshed, which changes cost profiles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A browser Worker killed without an error event (the browser terminating it, or terminate() from outside) went unnoticed: interact and run waited for the wall limit and reported wall-time-limit, blaming the student's program, and a killed compiler or stage Worker waited for its build or stage timeout. Every forge Worker starts from createModuleWorker's blob bootstrap, which now takes a uniquely named Web Lock for the Worker's lifetime as it starts and reports the lock's name to the parent. The parent requests the same lock; it is granted only once the Worker's context is destroyed. If that happens before the owner called terminate() (which createModuleWorker wraps to cancel the watch), onModuleWorkerLost listeners run. BrowserRunner, BrowserCompiler, the interactive side Workers and the compiler stage Workers feed them into their existing crash paths, so the engine reports runner-failure or compiler-failure. Listeners are called directly because WebKit drops events dispatched on a Worker after terminate(). Without Web Locks nothing changes. Interactive pipe waits now use 100 ms Atomics.wait slices. WebKit never stops a terminated Worker blocked in an untimed Atomics.wait, so a killed side Worker was undetectable there. The strict-CSP suite gains a liveness section that kills side, runner, compiler and rustc stage Workers mid-operation in each browser and checks prompt rejection and recovery, plus a no-false-positive run of 20 runs, 5 interactive sessions and a build. Detection: Firefox immediately, Chromium within about 2 s for busy Workers (its forced-termination delay), WebKit within about 80 ms except for a Worker spinning in pure Wasm, which WebKit does not stop until it returns to JavaScript. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md
…led Workers JavaScriptCore acts on Worker.terminate() only at JavaScript checkpoints, never inside Wasm, so a terminated runner or interactive side kept computing until its instruction budget ran out and Web Lock liveness could not see the kill. The charge that opens each function and loop body is now inlined and compared with a threshold. Every 2^20 units a cold function calls the imported wasm_oj_metering.safepoint, which is a no-op natively and an Atomics.wait that returns at once in browsers. Loops reach it through a branch out of the loop, because a call inside a hot loop slowed the whole loop in JavaScriptCore and V8. Costs and the exhaustion point are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #103 (and so #102) and on #104. This branch is #103's branch with #104 merged into it, plus one commit. Review only the last commit,
fix(runtime): add an uncharged metering safepoint so WebKit stops killed Workers. The overlap: #103 changes the same instrumentation pass and the generated runtime, and the WebKit regression tests here use #104's Worker liveness. I'll rebase once those land.Why
#104 reports a Worker that dies without an
errorevent, but WebKit has one case it can't see. A Worker that is terminated or killed while it runs pure Wasm is not actually stopped, so its Web Lock is never released and the session still ends at the wall limit.What I established
JavaScriptCore acts on termination only at JavaScript checkpoints, never inside Wasm. Time from
worker.terminate()until the Worker's Web Lock is released, plain Wasm loops outside forge, 6 s cap:Atomics.wait(cell, 0, mismatch, 0)Atomics.wait(cell, 0, value, 0)(timeout 0)The macOS 26.6
jscshell's--watchdogtermination, which uses the same VMTraps, behaves the same way: never in pure Wasm or Wasm→Wasm calls, only with a JS call on every iteration. An ordinary JS call boundary is caught only by chance, andAtomics.waitis a reliable checkpoint. Chromium force-terminates a busy Worker after its 2 s grace whatever it runs; Firefox stops it at once.How long a killed Worker keeps running in WebKit (after #103): until the program returns to JavaScript, that is until it makes a WASI call or exhausts its instruction budget. Default budget (1e10), time to
instruction-limit:for(;;);fib(30)in a loopmemcpyloopSo in WebKit a terminated runner keeps a core busy for a fraction of a second up to about 40 s, and a liveness check waits that long.
Real kills. A Worker cannot die alone in WebKit except through
terminate(): in a Worker, growing a Wasm memory past the limit throwsRangeErrorat 3,840 MiB and the page lives; plainArrayBuffers reached 6 GiB without failing; andkill -9of the WebContent process takes the whole page down (Playwright'scrashevent,Target crashed), with nothing left to report. So the case to fix is a Worker that forge, the embedding page or a parent Worker terminates while it computes.Change
The meter gets an uncharged safepoint (runtime-core, final instrumentation pass):
wasm_oj_metering.safepoint(added after the existing function imports; defined function indices shift by one) and gets a private threshold global, initiallybudget - 2^20.counter - 2^20(never below 0). Every other block keeps radix's gas function unchanged.unreachableremoved the cost). So a loop branches out to reach the safepoint: each loop becomesblock (loop (block (loop ...) br 2) call slow; refund; br 0) unreachable, every branch out of the loop skips the three new labels, and the cost charged before leaving is refunded so the re-entry charges it once. Functions that use exception or GC branch instructions (which radix rejects today anyway) and loops with parameters keep the call inside the loop.remaining_pointsare untouched, so costs and the exhaustion point are identical: the same programs report the same costs before and after, from 14,013 to 4,244,150,875 units, and a run withbudget = costexits whilecost - 1reportsinstruction-limit.safepointdoes nothing. In the browser it callsAtomics.waiton a never-matching value, which returns at once and is where JavaScriptCore acts onterminate(). It is a no-op whereSharedArrayBufferis unavailable.A terminated runner or interactive side Worker running metered code now stops within 2^20 units in WebKit (about 40 µs of compute in the benchmarks below). #104's liveness check then reports it, and WebKit no longer keeps a zombie Worker computing after a wall-time or cancel termination.
Alternatives considered:
Atomics.waitalready makes WebKit honourterminate()and Firefox stops at once. Not needed.Atomics.waitslices. They already cover a side blocked on a pipe (fix(browser): reject promptly when a Worker dies without an error event #104) but cannot help a program that never waits.Unmetered toolchain code (clang or rustc in the compiler stage) still has no safepoint. A killed compiler stage is reported when it returns to JavaScript (30 ms to 1.2 s measured in WebKit).
Overhead
Median of five runs (after one warm-up), release build,
executionDurationMs, two separate sessions each, this branch against the same stack without this commit. Costs are identical in every run (2,500,000,159 and 4,244,150,875).volatilesum, 3e8 iterations)unordered_map,map)The tight loop is the worst case: one iteration is a handful of instructions plus the check.
Verification
br_table, loop results and a loop with parameters giving the same result as the uninstrumented module (fails if the label shift is dropped); a module without imports gaining the import; all runtime-core tests;cargo fmt --check; clippy native and web with-D warnings;pnpm run runtime:check-web.pnpm run ci:verify, andsrc/server/judge.integration.test.tswith the native runner.liveness-runner-run-compute(kill the runner Worker during a compute-boundrun), which in WebKit used to end atwall-time-limit15 s later;liveness-interactive-compute(kill a compute-bound interactive contestant side);Strict-CSP suite (Playwright 1.62.1: Chromium 151, Firefox 153, WebKit 26.5): Chromium 121 of 121 and Firefox 121 of 121. The full WebKit run lost its page at
interactive-ac-cto the JavaScriptCore shared-memory crash that #105 works around (it happens onmaintoo), so I ran the liveness section in WebKit on its own twice, and the full suite on this branch combined with #105 (below).Time from the kill until the operation rejects (ms):
runwall-time-limit, 13,502 ms after the killwall-time-limit, 18,558 ms after the killinteractrun"Before" is the same stack without this commit. Combined with #105, the full WebKit suite passes everything except the pre-existing
python-16mb-recursion_error(it fails the same way onmain).runtime-core_bg.wasmis rebuilt withpnpm run runtime:buildand the identity pins are refreshed, which changes cost profile identifiers (not cost values).Risks
wasm_oj_metering.safepoint; forge's runtimes do.Atomics.waitis reached about every 2^20 units; on a main thread it throws, which is ignored.🤖 Generated with Claude Code