Repository navigation
fix(browser): keep Wasmer SDK Worker teardown away from shared memory growth - #105
Open
TakalaWang wants to merge 1 commit into
Open
TakalaWang wants to merge 1 commit into
TakalaWang wants to merge 1 commit into
Conversation
… growth JavaScriptCore crashes the page (SIGSEGV in SharedArrayBufferContents::grow) when a shared WebAssembly.Memory grows while a Worker whose instance imported it is being torn down. The Wasmer SDK terminates a thread Worker whenever a WASIX thread or process ends and grows its shared memory when the next one starts, so full WebKit runs lost the page about two times in three. OwnedWorkerRegistry now owns only the SDK's thread Workers, holds their terminate() calls until the running operation ends, and starts the next operation only after 250 ms without a termination. The compiler, the rustc stage and the runner wrap their operations in it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Full strict-CSP runs in WebKit lose the page in 7 of 10 runs on
main. Playwright reportsTarget page, context or browser has been closed, and the WebContent process has crashed:(or, in 3 of the reports, at
0x10inJSWebAssemblyInstance::updateMatchingCachedMemoriesConcurrentlyunderSharedArrayBufferContents::tryGrow, the same walk one frame deeper).The same crash would take down the tab in Safari. NOJV runs its browser Test on this path.
Root cause
It's a JavaScriptCore race. All 26 crash reports I collected (from forge's suite and from the repro below) crash in
Wasm::Memory::growShared, and in every one at least one other Worker thread is insideJSC::VM::~VM→Heap::lastChanceToFinalize, destroying aJSWebAssemblyModule: a terminated Worker's VM being torn down. Growing a shared memory walks every instance on every thread that imported it, to refresh their cached bounds. If one of those instances belongs to a Worker whose VM is being torn down, the walk reads freed state. Recent WebKit commits touch the same mechanism:4002a938c8("Growing a shared memory reads its instances' sibling memory handles across threads") and76b3468621(unregisterWasm::InstanceAnchorat the start of the instance destructor).Minimal standalone repro, plain Wasm and JS with no Wasmer SDK, serve with COOP/COEP. A 60-byte module imports a shared memory. One Worker at a time keeps growing a shared memory while two loops start six Workers that instantiate the module with the same memory and terminate them:
terminate()So the window is the VM teardown that follows a Worker's termination, which ends after its Web Lock is released.
Why forge hits it constantly. The Wasmer SDK terminates one of its thread Workers whenever a WASIX thread or process ends, and starts the next in a new Worker. That Worker's first message runs
initSync, which allocates its thread stack and grows the memory shared by all the SDK's Workers. A C compile (driver → cc1 → wasm-ld) does this 6–9 times. On top of that, the compiler family's shared memory grows by about 148 MiB per C or C++ compile and the rustc stage's by similar amounts, so grows and teardowns overlap all the time. All crash stacks are a new thread's first-message grow.Change
OwnedWorkerRegistry(already used by the rustc stage to own the SDK's nested Workers) now:terminate()calls on those Workers while an operation runs (run(operation)) and flushes them when it ends;Workerconstructor that is current when it is installed rather than when it is created, since each owner now creates its registry at module load and installs it when the SDK starts.The compiler Worker (C, C++, JavaScript, TypeScript builds), the rustc stage (each compile) and the runner Worker (runs, interactions and cache clearing, where Python packages use the SDK) wrap their operations in
run. Nothing changes in Chromium and Firefox except that finished SDK Workers live until the operation ends.@wasmer/sdkis not patched.Alternatives considered:
dlmallocalways callsmemory.growfor new space and ignores spare initial pages; I measured 20 of 28 thread inits growing with both default and 1024-page initial memories. A malloc+realloc pool reserved right afterinitdoes stop every thread-init grow (0 of 28), but each compile grows the family by about 148 MiB, so no affordable pool covers the grows that run while threads exit.A proper fix belongs in JavaScriptCore (grow vs. VM teardown) or in the SDK (reuse thread Workers, or free their stacks with
__wbindgen_thread_destroy, instead of terminating one Worker per WASIX thread).Verification
Full strict-CSP suite in WebKit, 10 runs each, alternating
mainand this branch on the same machine:main(ad1d7fe)interactive-ac-cppinteractive-ac-cppinteractive-ac-cppinteractive-output-floodinteractive-poll-timeoutinteractive-ac-cppinteractive-interactor-exitsNo run stalled on either side. Every completed run on both sides fails only
python-16mb-recursion_error(exit code 1 with no output), which fails the same way onmainin WebKit and is unrelated. macOS throttles crash reports, so only 2 of the 7 crashes wrote one; both have the signature above. The sum of the 82 timed checks common to all runs is 96.8 s onmain(median of its 3 completed runs) and 95.7 s here (median of 10).Also:
sdk-worker-churnfixture (30 back-to-back C compiles), passes in all three browsers;terminateAll;pnpm run ci:verify;python-16mb-recursion_error) with no page crash and no new crash report.Timing. Back-to-back operations pay the quiet period; anything else is unchanged. Median of five rounds of compile, then run (
stdin"20 22"), each compile right after the previous run:Whole strict-CSP suite, sum of the 98 timed checks common to both runs: Chromium 96.3 s on
main, 95.6 s here; Firefox 480.1 s and 501.9 s (the difference is in the Java checks, which vary between 217 and 247 s across runs on the same build). A first run of this branch right after a WebKit crash was about 30% slower in Chromium while macOS's crash reporter was still busy; the rerun above is on a quiet machine.Risks
🤖 Generated with Claude Code