Skip to content

fix(browser): keep Wasmer SDK Worker teardown away from shared memory growth - #105

Open
TakalaWang wants to merge 1 commit into
wasm-oj:mainfrom
TakalaWang:fix/webkit-sdk-worker-teardown
Open

TakalaWang wants to merge 1 commit into
wasm-oj:mainfrom
TakalaWang:fix/webkit-sdk-worker-teardown

Conversation

@TakalaWang

Copy link
Copy Markdown
Contributor

Why

Full strict-CSP runs in WebKit lose the page in 7 of 10 runs on main. Playwright reports Target page, context or browser has been closed, and the WebContent process has crashed:

EXC_BAD_ACCESS (SIGSEGV) KERN_INVALID_ADDRESS at 0x0000000000000008
JSC::SharedArrayBufferContents::grow(...)  ← JSC::Wasm::Memory::growShared  ← operationGrowMemory  ← Wasm
  ← JSEventListener::handleEvent ← EventTarget::dispatchEvent (JS)  ← JSModuleRecord::evaluate (Worker thread)

(or, in 3 of the reports, at 0x10 in JSWebAssemblyInstance::updateMatchingCachedMemoriesConcurrently under SharedArrayBufferContents::tryGrow, the same walk one frame deeper).

The same crash would take down the tab in Safari. NOJV runs its browser Test on this path.

Root cause

It's a JavaScriptCore race. All 26 crash reports I collected (from forge's suite and from the repro below) crash in Wasm::Memory::growShared, and in every one at least one other Worker thread is inside JSC::VM::~VM → Heap::lastChanceToFinalize, destroying a JSWebAssemblyModule: a terminated Worker's VM being torn down. Growing a shared memory walks every instance on every thread that imported it, to refresh their cached bounds. If one of those instances belongs to a Worker whose VM is being torn down, the walk reads freed state. Recent WebKit commits touch the same mechanism: 4002a938c8 ("Growing a shared memory reads its instances' sibling memory handles across threads") and 76b3468621 (unregister Wasm::InstanceAnchor at the start of the instance destructor).

Minimal standalone repro, plain Wasm and JS with no Wasmer SDK, serve with COOP/COEP. A 60-byte module imports a shared memory. One Worker at a time keeps growing a shared memory while two loops start six Workers that instantiate the module with the same memory and terminate them:

// module: (module (import "env" "memory" (memory 1 16384 shared)) (func (export "grow") (param i32) (result i32) local.get 0 memory.grow))
const bytes = new Uint8Array([0,97,115,109,1,0,0,0,1,6,1,96,1,127,1,127,2,18,1,3,101,110,118,6,109,101,109,111,114,121,2,3,1,128,128,1,3,2,1,0,7,8,1,4,103,114,111,119,0,0,10,8,1,6,0,32,0,64,0,11]);
const workerSource = `onmessage = ({ data }) => {
  const { exports } = new WebAssembly.Instance(data.module, { env: { memory: data.memory } });
  if (data.grow) { const sleep = new Int32Array(new SharedArrayBuffer(4)); for (let i = 0; i < 1500 && exports.grow(1) >= 0; i++) Atomics.wait(sleep, 0, 0, 0.2); }
  postMessage(0);
};`;
const workerUrl = URL.createObjectURL(new Blob([workerSource], { type: "text/javascript" }));
const start = (module, memory, grow) => new Promise((resolve) => {
  const worker = new Worker(workerUrl);
  worker.onmessage = () => resolve(worker);
  worker.postMessage({ module, memory, grow });
});
async function race(seconds) {
  const module = new WebAssembly.Module(bytes);
  const deadline = performance.now() + seconds * 1000;
  let memory = new WebAssembly.Memory({ initial: 1, maximum: 16384, shared: true });
  const grow = async () => { while (performance.now() < deadline) { memory = new WebAssembly.Memory({ initial: 1, maximum: 16384, shared: true }); (await start(module, memory, true)).terminate(); } };
  const churn = async () => { while (performance.now() < deadline) for (const worker of await Promise.all([0, 1, 2, 3, 4, 5].map(() => start(module, memory, false)))) worker.terminate(); };
  await Promise.all([grow(), churn(), churn()]);
  return "survived";
}
race(20).then(console.log);
Variant (20 s each) WebKit 26.5 (Playwright 1.62.1) Chromium 151 Firefox 153
As above crashes 5/5, within about 1 s 0/2 0/2
Dying Workers imported a different shared memory 0/4 – –
No grows 0/4 – –
Grow only after the dying Workers' Web Locks are released 4/4 – –
Grow 20 ms after the Web Locks are released 0/4 – –
Grow 100 ms after terminate() 0/4 – –

So the window is the VM teardown that follows a Worker's termination, which ends after its Web Lock is released.

Why forge hits it constantly. The Wasmer SDK terminates one of its thread Workers whenever a WASIX thread or process ends, and starts the next in a new Worker. That Worker's first message runs initSync, which allocates its thread stack and grows the memory shared by all the SDK's Workers. A C compile (driver → cc1 → wasm-ld) does this 6–9 times. On top of that, the compiler family's shared memory grows by about 148 MiB per C or C++ compile and the rustc stage's by similar amounts, so grows and teardowns overlap all the time. All crash stacks are a new thread's first-message grow.

Change

OwnedWorkerRegistry (already used by the rustc stage to own the SDK's nested Workers) now:

  • owns only Workers created from the owner's SDK thread-bootstrap URL, so forge's own stage and side Workers are untouched;
  • holds the SDK's own terminate() calls on those Workers while an operation runs (run(operation)) and flushes them when it ends;
  • starts the next operation only after 250 ms without a termination, so the flushed teardowns end before the family grows its memory again.
  • wraps the Worker constructor that is current when it is installed rather than when it is created, since each owner now creates its registry at module load and installs it when the SDK starts.

The compiler Worker (C, C++, JavaScript, TypeScript builds), the rustc stage (each compile) and the runner Worker (runs, interactions and cache clearing, where Python packages use the SDK) wrap their operations in run. Nothing changes in Chromium and Firefox except that finished SDK Workers live until the operation ends. @wasmer/sdk is not patched.

Alternatives considered:

  • A larger initial memory or a pre-reserved pool. Rust's wasm dlmalloc always calls memory.grow for new space and ignores spare initial pages; I measured 20 of 28 thread inits growing with both default and 1024-page initial memories. A malloc+realloc pool reserved right after init does stop every thread-init grow (0 of 28), but each compile grows the family by about 148 MiB, so no affordable pool covers the grows that run while threads exit.
  • Waiting for the dying Worker's Web Lock. The lock is released before the VM teardown, and the repro still crashes 4/4.
  • Recycling SDK Workers instead of terminating them. That needs SDK internals: a finished thread's guest instance and memory stay referenced from its Worker.

A proper fix belongs in JavaScriptCore (grow vs. VM teardown) or in the SDK (reuse thread Workers, or free their stacks with __wbindgen_thread_destroy, instead of terminating one Worker per WASIX thread).

Verification

Full strict-CSP suite in WebKit, 10 runs each, alternating main and this branch on the same machine:

Run main (ad1d7fe) This branch
1 page crashed at interactive-ac-cpp completed
2 completed completed
3 completed completed
4 page crashed at interactive-ac-cpp completed
5 page crashed at interactive-ac-cpp completed
6 page crashed at interactive-output-flood completed
7 page crashed at interactive-poll-timeout completed
8 completed completed
9 page crashed at interactive-ac-cpp completed
10 page crashed at interactive-interactor-exits completed
Page crashes 7 / 10 0 / 10

No run stalled on either side. Every completed run on both sides fails only python-16mb-recursion_error (exit code 1 with no output), which fails the same way on main in WebKit and is unrelated. macOS throttles crash reports, so only 2 of the 7 crashes wrote one; both have the signature above. The sum of the 82 timed checks common to all runs is 96.8 s on main (median of its 3 completed runs) and 95.7 s here (median of 10).

Also:

  • new sdk-worker-churn fixture (30 back-to-back C compiles), passes in all three browsers;
  • unit tests for deferral, flush, quiet period, URL filtering, install-time capture and terminateAll;
  • pnpm run ci:verify;
  • full strict-CSP suite: Chromium 106 of 106, Firefox 106 of 106, WebKit 105 of 106 (only the pre-existing python-16mb-recursion_error) with no page crash and no new crash report.

Timing. Back-to-back operations pay the quiet period; anything else is unchanged. Median of five rounds of compile, then run (stdin "20 22"), each compile right after the previous run:

Chromium 151 Firefox 153 WebKit 26.5
C compile 108 → 366 ms 1109 → 1220 ms 222 → 482 ms
Java compile 2577 → 2600 ms 15660 → 15843 ms 3240 → 3236 ms
Python run (SDK packages) 761 → 767 ms 4559 → 4616 ms 780 → 786 ms

Whole strict-CSP suite, sum of the 98 timed checks common to both runs: Chromium 96.3 s on main, 95.6 s here; Firefox 480.1 s and 501.9 s (the difference is in the Java checks, which vary between 217 and 247 s across runs on the same build). A first run of this branch right after a WebKit crash was about 30% slower in Chromium while macOS's crash reporter was still busy; the rerun above is on a quiet machine.

Risks

  • A finished SDK Worker lives until its operation ends, so a build with several processes holds their memory a little longer.
  • Back-to-back operations in one owner wait up to 250 ms after an operation that terminated SDK Workers (measured above: about +250 ms for a C compile that starts right after the previous one). This also applies in Chromium and Firefox, which do not need it; limiting it to JavaScriptCore would need engine detection, which forge does not do elsewhere.
  • If WebKit's teardown ever takes longer than 250 ms, the race comes back for operations that start right after one another. The repro needs about 20 ms.
  • The SDK may terminate a thread Worker that is still blocked; deferring keeps it blocked until the operation ends.

🤖 Generated with Claude Code

… growth

JavaScriptCore crashes the page (SIGSEGV in SharedArrayBufferContents::grow)
when a shared WebAssembly.Memory grows while a Worker whose instance
imported it is being torn down. The Wasmer SDK terminates a thread Worker
whenever a WASIX thread or process ends and grows its shared memory when
the next one starts, so full WebKit runs lost the page about two times in
three.

OwnedWorkerRegistry now owns only the SDK's thread Workers, holds their
terminate() calls until the running operation ends, and starts the next
operation only after 250 ms without a termination. The compiler, the
rustc stage and the runner wrap their operations in it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant