Skip to content

build(deps): migrate to @wasmer/sdk 0.19.0 - #97

Open
TakalaWang wants to merge 7 commits into
wasm-oj:mainfrom
TakalaWang:feat/wasmer-sdk-0.19
Open

TakalaWang wants to merge 7 commits into
wasm-oj:mainfrom
TakalaWang:feat/wasmer-sdk-0.19

Conversation

@TakalaWang

@TakalaWang TakalaWang commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Why

@wasmer/sdk 0.10.0 is unmaintained. It also has a teardown race: Command.run() sends the exit code and closes its thread pool before the WASI runner that owns the stdout/stderr writers is dropped, so Instance.wait() can wait forever for EOF. Under concurrency this stalled about 4% of Python runtime preparations (#96 works around it by not waiting for EOF). The same race stalled TypeScript compilation in a baseline run: typescript-wasip1: Server compilation exceeded 120000 ms. The SDK was rewritten in 0.11.0 (it now lives at wasmerio/wasmer-sdk), and 0.19.0 does not have this race.

Change

  • SDK API. Server stages and browser Workers create a Wasmer client, load each pinned WEBC or raw Wasm package once, and run Clang/wasm-ld, rustc/wasm-ld, and TypeScript in a fresh sandbox per build. Processes run to exit through piped streams (runSandboxCommand). This replaces the 0.10.0 output-ready polling, so MountedOutputStabilityObserver, PackageHandleCache, the custom wasmer-thread.worker, and wasmer-runtime.ts are deleted.
    • rustc traps (unreachable) after it prints errors, which makes wait() reject. The piped streams still end, so the diagnostics are kept.
  • Guest paths. SDK sandboxes only accept files under /workspace. Clang maps the pinned /project arguments, adds -fmacro-prefix-map/-fdebug-prefix-map, and uses a -ivfsoverlay for the admitted libc++ PCH, which records /project/wasm-oj.libcxx.hpp. Rust moves its root and remaps /workspace=.. All 78 compared artifacts are byte-identical to 0.10.0 (C, C++ incl. debug, __FILE__/assert, admitted and project PCH, Rust, TypeScript). Project-PCH cache keys include the root, so PCHs persisted under /project are not reused.
  • Command.binary() replacement. src/runner/webc.ts is a small WEBC v3 reader that returns the named atom from the atoms section. Callers check it against the pinned PYTHON_COMMAND_SHA256; the bytes are identical to the old Command.binary(). Server: the runner stage's command-binary operation no longer starts the SDK. Browser: the runner Worker reads the atom from the verified package bytes.
  • Browser packaging. scripts/wasmer-sdk-runtime.mjs is a Vite plugin used by the library build and the app build:
    • It ships the SDK's browser runtime files unbundled under assets/wasmer-sdk-<digest>/ and exposes their URLs through virtual:wasm-oj/wasmer-sdk. createBrowserWasmer() imports the SDK from there and points its thread Workers at a same-origin blob bootstrap, as before.
    • It omits dist/wisp-network.js, the SDK's only path to its AGPL-3.0 @mercuryworkshop/wisp-js dependency. WASM-OJ sandboxes have networking disabled.
    • It replaces the wasm-bindgen glue's new Function("f", "return f(Array.prototype.slice.call(arguments, 1))") with the equivalent closure. Without this, 0.19.0 fails under strict CSP for every language. The build fails if that pattern changes.
    • @wasm-oj/browser no longer depends on @wasmer/sdk. Hosts that copy dist/assets/ must copy it recursively; NOJV already does.
    • The plugin lives in scripts/ because the judge image's .dockerignore admits scripts/ but not build/.
  • Reusable server compiler children. server-build-stage.mjs and rustc-stage.mjs now keep their Wasmer client and loaded packages and serve one build at a time (src/server/reusable-stage.ts), so the slower 0.19 package load is paid once per child instead of once per build. Every build still gets a fresh sandbox.
    • Cancellation, a timeout, a failure or a bad response kills the child, and the next build starts a new one.
    • The rustc child is replaced after 2 builds, the browser's existing Rust stage budget, because rustc leaves one busy SDK worker behind per run.
    • Idle children exit after 30 s and never keep the host process alive.
    • A child exits when its stdin closes, so it also exits when its parent dies. Before, an orphaned 0.10 server-build-stage.mjs kept spinning for hours after its parent was killed.
    • Python preparation, Go and Java stay one-shot: they load no SDK package per build.
  • Kill and timeout policy (upstream kill() can terminate a worker inside the allocator and hang the whole process wasmerio/wasmer-sdk#542). WASM-OJ never calls Process.kill() or Wasmer.close(). A stage that exceeds its budget is abandoned, not killed. Cancellation, timeouts, and quiescing still discard the whole server child or Worker generation. The rustc stage terminates its owned SDK Workers before closing.
  • Licenses. From 0.11.0 the SDK uses Wasmer's Modified MIT License: commercial products with more than 1M MAU or more than $1M monthly revenue must display "Wasmer" in their UI.
    • The license text, the acorn 8.18.0 notice (from the SDK snippets), and a regenerated cargo-about inventory (316 crates at wasmer-sdk-js-v0.19.0, 7e69332) replace the old material.
    • Notices, components.json, and verify-wasmer-sdk-init.mjs are updated.
  • Other. The runtime identity changes, so artifacts cached by earlier releases are rebuilt. The Sites build removes the duplicate SDK runtime directory from server output. Docs and CHANGELOG are updated. A gated stress test, WASM_OJ_RUN_CANCELLATION_STRESS=1, cancels compiles, Python preparation, and runs at random points.

The public @wasm-oj/* APIs are unchanged.

Verification

  • pnpm run ci:verify: exit 0 (typecheck, lint, 948 tests, Sites build, verify-browser-assets).
  • src/server/reusable-stage.test.ts:
    • The same child is reused; it is replaced at the limit; a failed, timed-out or terminated child is never reused; idle children exit.
    • Real Rust builds reuse the build child and replace rustc after 2 builds.
    • SIGKILLing the parent of an idle child and of a mid-build child makes both exit. Without the stdin-EOF exit, these tests fail and leave orphans.
    • These tests find children with ps; they pass on macOS, and CI will be their first Linux run.
  • All 78 compared artifacts were also built in one engine, through reused children, and are byte-identical to 0.10.0.
  • Library and license checks, all pass:
    • pnpm run library:verify (needs pnpm 10.34.5 on PATH for the NodeNext consumer, which is environmental).
    • licenses:verify, wasmer-sdk:verify, docs:verify.
  • WASM_OJ_RUN_JUDGE_INTEGRATION=1 vitest run src/server/judge.integration.test.ts: 2/2.
  • Server conformance, full suite (WASM_OJ_RUN_CONFORMANCE=1 …SUITE=full): 140/140 samples pass. Baseline 0.10.0 was 139/140 because of the TypeScript stall above. The test then fails while writing evidence because experiments/wasm-oj-contract-2-conformance/SPEC.md is missing; this is the same on main.
  • Strict-CSP real-browser suite (scripts/verify-browser-csp.mjs, 87 fixtures in all 8 languages plus capability probes):
    • Chromium 87/87, Firefox 87/87.
    • WebKit 86/87: python-16mb-recursion_error also fails on main (a JS stack limit in runtime-core).
  • App production build (vite preview, app CSP + COOP/COEP + chunk service worker, Chromium): the app's own compiler and runner Workers compiled and ran C++ (42) and Python (42) through /_next/static/wasmer-sdk-<digest>/.
  • Stall harness, both 0 stalls / 0 failures:
    • stage-stress: 300 Python runtime-file exports at concurrency 6, each archive digest verified.
    • driver: 50 iterations × 2 engines, C++ + Python.
  • Cancellation stress:
    • Server, 50 iterations: every C++, Rust, Python-preparation, CPU-loop, and output-heavy scenario was cancelled mid-flight. 0 hangs, 0 leftover child processes, and the next compile and run worked every time. This run needed the pending stdin EPIPE fix (fix(server): keep a stdin error handler after cancelling a run); without it the host crashed with write EPIPE after 42 iterations, which is a known issue on main.
    • Browser (Chromium), 40 iterations: 0 hangs, 0 failures.
  • Direct SDK kill() (#542) investigation:
    • Node: 60 fresh processes, 180 kills. No process hang, and the next command always worked. But a killed CPU-bound guest kept running on its Worker (about 0.9 core) and Wasmer.close() never resolved in 57 of 60 processes.
    • Chromium: 30 kills. close() hung for all 15 CPU-bound kills; there was no page hang.
    • The allocator-lock spin reported in #542 (Linux) did not reproduce on macOS arm64. The kill problem above is the reason forge never calls kill() or close().

Server builds with reused children (median of 5 rounds, ms; builds 1 / 2 / 3 on a fresh engine)

0.10.0 0.19.0 one-shot 0.19.0 reused
C++ 3345 / 2987 / 3010 3565 / 3428 / 3520 3516 cold / 917 / 947
Rust 3280 / 3113 / 3145 5397 / 5647 / 5277 5265 cold / 205 / 5226 cold (recycled)
TypeScript 2398 / 2275 / 2244 2826 / 2812 / 2927 2827 cold / 1816 / 1780
Python first run 2278 2046 2026 (one-shot)

Performance before reuse (median of 3, two interleaved rounds, 0.10.0 → 0.19.0, ms)

Path Server (Node) Browser (Chromium, packaged library)
Python first run (runtime preparation) 2508/2260 → 2080/2208 1688/1648 → 1659/1656
Python warm run 1151/1194 → 1164/1269 794/770 → 774/778
C++ compile (first) 3071/3257 → 3583/3692 2437/2351 → 2022/1943
C++ compile (second) 3489/3320 → 3540/3827 1011/1031 → 807/863
C++ run 100/128 → 127/124 56/56 → 56/57
Interactive session 64/61 → 65/67 n/a (fails on main: missing startupEntropyBytes)
Rust compile 3396/3299 → 5494/5792 2792/2770 → 2014/2022
TypeScript compile 2544/2416 → 2858/2869 3100/3060 → 3244/3203

Shipped browser bytes (JS + Wasm):

Before After Change
@wasm-oj/browser raw 22,157,802 20,613,144 −7.0%
@wasm-oj/browser gzip 6,059,833 5,538,900 −8.6%
App client raw 37,075,901 35,531,305
App client gzip 9,695,925 9,174,946
  • SDK assets before: a 6.60 MB Wasm, plus a thread worker, a raw SDK module, and the SDK bundled into 3 Worker chunks. After: one 5.28 MB directory (4.83 MB Wasm).
  • Worker chunks shrink: compiler 138→94 KB, runner 277→233 KB, rustc 69→24 KB.

Risks / follow-ups

  • Server package loading is slower, but each child pays it once. packages.load() hashes every WEBC and atom with SHA-256 in Wasm (Rust 1.28 → 3.56 s, Clang 0.53 → 1.51 s, Python 0.12 → 0.31 s). Only a cold child pays it:

    • the first build;
    • the first build after 30 s idle, or after a cancelled or failed build;
    • every other Rust build, because the rustc child is replaced after two.

    A burst of builds after an idle period or a cancellation still pays the cold cost once. A warm idle child holds about 250–360 MB until it exits.

  • SDK kill()/close() are unsafe for CPU-bound guests, as shown above. Forge only uses process and Worker teardown and must keep it that way.

  • rustc leaves one busy SDK worker per invocation, and Wasmer.close() then never resolves. Server children exit, and the browser Rust stage still recycles after 2 builds (MAX_OUTPUT_READY_RUST_STAGES_PER_WORKER). The Clang/Rust stage budgets and their "output-ready" names were kept as they were; with run-to-exit processes the Clang budget could probably be relaxed, but that needs measurement first.

  • Raw Clang stderr now names /workspace/.... Parsed diagnostics and artifacts are unchanged.

  • Licensing. The new SDK license condition and the AGPL wisp-js transitive install for @wasm-oj/server consumers need owner acknowledgement. Forge does not ship wisp-js in its tarballs or app.

  • Dev mode is untested. The plugin's serve branch could not be exercised. On main as well, pnpm dev fails locally: workerd supports compatibility dates up to 2026-08-08, the config uses 2026-08-09, and vinext dev returns 404 for client modules. Forge's module-Worker loader also rejects Vite dev Worker URLs because they carry a query string.

  • Conflicts with fix(runtime): stop waiting for stream EOF when exporting runtime files #96. fix(runtime): stop waiting for stream EOF when exporting runtime files #96 (no-EOF workaround) conflicts with this branch in server-runner-stage.mjs and runner.worker.ts. If it lands first, keep this branch's run() path and drop readRuntimeFilesExport.

  • Test gaps worth closing in review:

    • Only the server tests cover rustc's trap-after-diagnostics path; there is no browser Rust compile-error fixture.
    • The wisp exclusion and the CSP patch are checked against the planned file map, but not against the packed dist/assets/wasmer-sdk-*.
    • runSandboxCommand maps every wait() rejection to a failed build, so an SDK failure that is not a guest trap appears as a compile error with empty stderr.
  • Pre-existing issues found along the way (not fixed here):

    • The browser interactive request omits startupEntropyBytes, so engine.interact fails in the browser.
    • Stdin EPIPE on cancel (fix pending).
    • The conformance evidence SPEC.md is missing.

🤖 Generated with Claude Code

TakalaWang and others added 5 commits October 7, 2026 01:03
Python execution needs the CPython interpreter module so the runtime core
can meter and run it. `@wasmer/sdk` 0.10.0 exposed it through
`Command.binary()`, which the SDK removed in 0.11.0 and has no equivalent
in 0.19.0.

Add a minimal WEBC v3 reader that returns one named atom from the
container's atoms section, and pin the SHA-256 of the Python package's
`python` atom so callers can verify the extracted module. The bytes are
identical to what `Command.binary()` returned for the pinned package.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… 0.19.0

@wasmer/sdk 0.10.0 is unmaintained and has a teardown race: `Command.run`
sends the exit code and closes its thread pool before the WASI runner that
owns the stdout/stderr writers is dropped, so `Instance.wait()` can wait
forever for EOF. It stalled about 4% of concurrent Python runtime
preparations and also TypeScript builds under load. 0.11.0 rewrote the SDK
around clients, packages and sandboxes; 0.19.0 does not have the race.

Server and browser now create a `Wasmer` client, load each pinned WEBC or
raw Wasm package once, and run Clang, wasm-ld, rustc and TypeScript in a
fresh sandbox per build. Processes run to exit through piped streams:
rustc traps after printing errors, which rejects `wait()` but keeps the
captured diagnostics. This replaces the output-ready polling that 0.10.0
required, so the mounted-output stability observer is gone. Sandboxes keep
project files in `/workspace`; Clang adds macro and debug prefix maps and
a VFS overlay for the admitted libc++ PCH, and Rust remaps the new root,
so all 78 compared C, C++, Rust and TypeScript artifacts are byte-identical
to 0.10.0 builds. Project PCH cache entries are keyed by the root because a
PCH records absolute input paths.

Python reads its interpreter module from the WEBC atom and exports
runtime files with `sandbox.command().run()`; the browser runner no longer
retains SDK package handles.

WASM-OJ never calls the SDK's `kill()` or `close()`. A killed CPU-bound
guest keeps running on its worker and `close()` then waits forever, and
upstream reports that termination can hang the process inside the shared
allocator. Timeouts abandon the process, and cancellation still discards
the whole Worker generation or server child.

The browser no longer bundles the SDK. A Vite plugin ships its browser
runtime files unbundled under `assets/wasmer-sdk-<digest>/`, and Workers
import them and point the SDK's thread Workers at a same-origin blob
bootstrap. The plugin omits the WISP networking module, the SDK's only
path to its AGPL-3.0 `@mercuryworkshop/wisp-js` dependency, and replaces
the glue's `new Function` host-function trampoline with the equivalent
closure so strict CSP still works. `@wasm-oj/browser` therefore no longer
depends on `@wasmer/sdk`, and the custom thread Worker and its policy are
removed. The Sites build drops the SDK runtime duplicates from server
output, and the runtime identity records the new SDK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The SDK moved to wasmerio/wasmer-sdk and, from 0.11.0, ships under
Wasmer's Modified MIT License, which adds an attribution condition for
commercial products above 1 million monthly active users or 1 million US
dollars of monthly revenue. Replace the MIT text with the new license and
name it accordingly.

`@wasm-oj/browser` now distributes the SDK's wasm-bindgen snippets, which
bundle acorn 8.18.0, so its MIT notice joins the SDK license material.

Regenerate the cargo-about inventory for the `wasmer-sdk-js` crate at the
0.19.0 release commit (316 packages for wasm32-unknown-unknown) with the
repository toolchain, and update the notices, component manifest and the
SDK verifier pins. The notices state the one source change WASM-OJ makes
and that the AGPL-3.0 `@mercuryworkshop/wisp-js` dependency is not shipped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… processes

The SDK upgrade keeps forge's rule that cancellation and timeouts tear down
the whole server child or Worker instead of killing an SDK guest. Add a
gated integration test that repeatedly cancels C++ and Rust builds, Python
runtime preparation on a fresh cache, and CPU-bound and output-heavy native
runs at random points, then requires each cancellation to settle within a
minute and the same engine to compile and run again.

Run it with WASM_OJ_RUN_CANCELLATION_STRESS=1; WASM_OJ_CANCELLATION_ROUNDS
sets the number of rounds (default 4).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Record how compilers and Python runtime preparation use the SDK now: one
sandbox per build rooted at `/workspace`, processes that run to exit, and
no SDK kill or client shutdown inside a live Worker or child. Document the
unbundled `assets/wasmer-sdk-<digest>/` runtime directory that browser
hosts copying `@wasm-oj/browser` assets must keep, and add the Unreleased
changelog entry, including the SDK's new license.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TakalaWang and others added 2 commits October 7, 2026 02:21
@wasmer/sdk 0.19 hashes every WEBC and atom with SHA-256 in Wasm when a
package loads. With one child per build, every server build paid that again:
about +2.2 s per Rust build and +0.4 to +0.6 s per C++ or TypeScript build
compared with 0.10.

The isolated build child and its rustc stage now serve sequential requests
over stdin and answer on fd 3, keeping their Wasmer client and loaded
packages. Each build still runs in a fresh SDK sandbox, and the Clang object
graph is cleared after every build. A child serves one build at a time.
Cancellation, timeout, failure or an unreadable response discards it, and the
next build starts a new one. The rustc child is replaced after the browser's
Rust stage budget (two builds), because rustc leaves a busy SDK worker behind
after each run. Idle children exit after 30 s, are unreferenced so they never
keep the host alive, and exit on stdin EOF, so they also exit when their
parent dies instead of lingering as orphans.

Python runtime preparation stays one-shot: it runs once per cache directory,
and the command-binary path reads the WEBC atom without loading the package.
Go and Java stages do not use the SDK and stay one-shot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The architecture and library contract said every server build uses a fresh
or one-shot child. Describe the reused child, its fresh sandbox per build,
when it is discarded, and the Rust stage budget, and record the change in
the CHANGELOG.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@JacobLinCool JacobLinCool left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes needed before merging. This review covers correctness, install and ops only. The Wasmer license terms are being decided separately.

What I verified works:

  • Server builds match main. 11 successful builds produce byte-identical artifacts: C/C++ including PCH and a locked C++ dependency, Rust including a crate dependency, and TypeScript. 6 failing builds give the same parsed diagnostics as main, except the link-error message noted inline.
  • Browser Rust/C/C++ compile and link errors return full diagnostics in Chromium.
  • The Python export stall is gone: 0 stalls in 360 stress runs of the real stage script at concurrency 6, against 1 in 120 on main.

Please address (details are inline, except the README item)

  • npm install fails on linux-arm64. @wasmer/sdk@0.19.0 depends on wisp-js, which depends on bufferutil, a native install script with no arm64 Linux prebuild (packages/server/package.json:68).
  • A warm compiler child holds 0.6–1.8 GB, not 250–360 MB. Mixed Rust/C++ builds peak at about 1.8× main. The judge creates one engine per submission, so it never reuses these children, yet keeps them alive into judging (src/server/server-compiler.ts:87).
  • SDK host failures become terminal compile-error verdicts instead of retried infrastructure errors (src/compiler/sandbox-command.ts:49).
  • The README copy command skips the new SDK directory (README.md:129, not in this diff).
    • cp node_modules/@wasm-oj/browser/dist/assets/* dist/assets/ now skips wasmer-sdk-<digest>/ (macOS prints "is a directory (not copied)") and exits 1.
    • A host that ignores the error deploys a page whose first C/C++/JS/TS/Rust compile or Python run 404s on wasmer-sdk-<digest>/dist/index.js.
    • Fix:
      • Change the command to cp -R node_modules/@wasm-oj/browser/dist/assets/. dist/assets/.
      • In docs/integration-guide.md:106, drop "instead of bundling them". README:123–124 says bundlers can't see these assets, so every host has to copy them.
      • Mark the layout change as "host action required" in the CHANGELOG.

Smaller

  • fd-3 responses are decoded one chunk at a time, which can turn non-ASCII rustc output into U+FFFD (src/server/reusable-stage.ts:127).
  • The dev serve branch loads the wasm-bindgen module twice (scripts/wasmer-sdk-runtime.mjs:94).
  • The CSP guard checks the call site but not the trampoline body (scripts/wasmer-sdk-runtime.mjs:64).
  • Two CHANGELOG lines claim too much:
    • Link-error messages do change (CHANGELOG.md:26).
    • Stale cached artifacts fail once before they are rebuilt (CHANGELOG.md:28).

Rebase notes (#94, #96 and #99 are now on main)

  • Expected conflicts (git merge-tree against current main): CHANGELOG.md, docs/architecture.md, src/compiler/wasmer-engine.ts, src/core/runtime-identity.ts, src/runtime/runner.worker.ts and src/server/server-runner-stage.mjs.
  • server-runner-stage.mjs: take your version of the whole file (--theirs during git rebase). Do not resolve it hunk by hunk: #96's try { … } finally { instance.free(); } around the export can survive outside the conflict markers, which is a syntax error. Typecheck skips .mjs, but pnpm run lint reports it.
  • Delete what becomes dead:
    • #96 added src/core/process-output.ts (readProcessMessage, completeJsonObject) and readRuntimeFilesExport. They now cover both the Python export and the TypeScript compile path (transpileScriptProject). Once your sandbox run() path replaces both, delete them and their tests.
    • #96's CHANGELOG line about finishing without waiting for Instance.wait().
    • Your Python export (run()) and runSandboxCommand both read to EOF, so dropping #96's reader relies on 0.19 not having the race. The evidence: 0 stalls in 360 Python export stress runs, and your 140/140 conformance run for the compile paths.
  • runner.worker.ts:
    • Keep #99's nested-Worker code and compileRuntimeCore next to your ensureSdkClient.
    • Take your body for the runtime-file export.
    • Drop the imports that become unused (package-handle-cache, wasmer-thread.worker, readRuntimeFilesExport, …).
  • runtime-identity.ts:
    • Keep main's runtime-core pins (#99 rebuilt runtime-core) and your SDK pins, then recompute WASM_OJ_RUNTIME_IDENTITY_SHA256 over runtimeIdentityBytes(). Neither side's digest is correct after the merge, and runtime-identity.test.ts checks the value.
    • No runtime:build is needed, because #97 touches no crates, vendor or generated runtime files.
  • docs/architecture.md: keep both the reusable-child text and #99's nested-Worker paragraph.
  • Release: the new identity breaks the calibration.profiles of every published judge package, so Official Submit returns 409 cost-profile-identity until they are republished. #97 changes no metering, so ship it in the same release as #95/#99. Judge data is then recalibrated once, not twice.
  • After rebasing, run:
    • pnpm run ci:verify. It includes python-compiler.test.ts and server-compiler.test.ts, and its lint step catches the .mjs problem above.
    • scripts/verify-browser-csp.mjs by hand, since CI has no browser coverage.

"@wasm-oj/contracts": "workspace:0.2.3",
"@wasm-oj/core": "workspace:0.2.3",
"@wasmer/sdk": "0.10.0",
"@wasmer/sdk": "0.19.0",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

npm install of @wasm-oj/server, @wasm-oj/cli or @wasm-oj/sdk fails on linux-arm64 hosts without a C toolchain.

  • Dependency chain:
    • @wasmer/sdk@0.19.0 hard-depends on @mercuryworkshop/wisp-js@^0.4.1, which hard-depends on bufferutil@^4.0.9. bufferutil's install script is node-gyp-build.
    • bufferutil@4.1.0 ships prebuilds only for darwin-arm64/x64, linux-x64 (glibc) and win32-ia32/x64.
    • On any other platform it falls back to node-gyp rebuild, which needs python3, make and g++. bufferutil is not optional, so when that build fails npm aborts the whole install.
  • Repro in node:24-slim arm64:
    • @wasm-oj/cli@0.2.3 installs and woj --version prints woj 0.2.3.
    • The same package with "overrides": { "@wasmer/sdk": "0.19.0" } exits 1 with gyp ERR! … Could not find any Python installation and leaves no node_modules.
    • node:24-alpine arm64 fails the same way. Alpine amd64 installs fine.
    • win-arm64 has no prebuild either; not tested.
  • Regression:
    • On main the server's SDK dependency tree was only web-worker, with no install scripts.
    • The CLI's other native dependency, @napi-rs/keyring, ships linux-arm64 (gnu and musl) and win32-arm64 binaries.
    • So woj installs on arm64 Linux today, for example in arm64 Docker on Apple Silicon.
  • Nothing in this chain runs on Node. wisp-js is only imported from dist/wisp-network.js when network.mode === "wisp". On Node that path throws CAPABILITY_UNAVAILABLE before the import, and forge never passes network.
  • Not affected:
    • CI, the judge image and Sites. They install with pnpm 10, which skips bufferutil's script because it isn't in onlyBuiltDependencies.
    • npm install --ignore-scripts.
    • Nothing breaks until the next npm release.

Fix:

  • Ship the SDK's Node runtime files inside @wasm-oj/server's dist, without dist/wisp-network.js, and drop @wasmer/sdk from the server's runtime dependencies.
    • scripts/wasmer-sdk-runtime.mjs already walks this graph for the browser.
    • For Node the roots are dist/node.js, dist/node-worker.js and pkg/wasmer_sdk_js_bg.wasm, and the walker has to allow node: specifiers.
  • Point the three @wasmer/sdk/node imports at the shipped copy: server-compiler.ts, rustc-stage.mjs and server-runner-stage.mjs.
  • If that is too much for this PR, at minimum document pnpm add or npm install --ignore-scripts for arm64 in the README and CHANGELOG before the npm release.

const SERVER_BUILD_RESPONSE_LIMIT_BYTES = 256 * 1024 * 1024;
const SERVER_BUILD_REQUEST_LIMIT_BYTES = 768 * 1024 * 1024;
const SERVER_COMPILER_STAGE_RESPONSE_LIMIT_BYTES = 256 * 1024 * 1024;
const SERVER_STAGE_IDLE_TIMEOUT_MS = 30_000;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A warm child that just ran Clang or rustc holds about 0.6–1.8 GB, not the 250–360 MB in the PR body. The judge container gets no reuse from it.

  • Measured:
    • macOS arm64, Node 24:
      • The idle rustc-stage child holds about 1.5–1.8 GB after one hello-world Rust build.
      • The build-stage child holds 0.9–1.7 GB after C++ builds, until 30 s pass with no further build.
      • Only the build child after a Rust-only build is close to the PR's figure (about 150 MB).
    • Linux arm64 Docker, 1 vCPU: the rustc child keeps 1.3–1.4 GB and the C++ build child about 600 MB.
    • No leak: over a 100-build soak, RSS stayed between 1.4 and 1.8 GB with no upward trend, and idle CPU is about 0.
  • The two children stack. The rustc grandchild stays warm for 30 s inside the build child:
    • Rust then C++ peaks at about 3.1 GB.
    • C++ then Rust peaks at about 3.6 GB.
    • On main the children exit after every build, so the peak is one build's worth, 1.5–2.0 GB. A host sized for main's peak could OOM.
  • The judge path pays the memory cost without the benefit. container/server.mjs builds one engine per submission and disposes it after judging, so ServerCompiler compiles exactly once.
    • Every Official Submit runs cold. On Linux with 1 vCPU, Rust goes from 5.9–6.2 s to 8.4–8.8 s and C++ from 2.8 s to 3.5 s. That is within budget, and your one-shot column already shows the slowdown, but reuse never offsets it on this path.
    • The children then stay resident for up to 30 s of judging. They sit next to programs allowed up to 4 GiB (MAX_MEMORY_LIMIT_BYTES), and interactive problems run two programs at once.

Fix:

  • Make the idle timeout an option on ServerCompilerOptions and ServerEngineOptions, defaulting to 30 s. It can be @internal, like verifiedDistribution.
  • Have container/server.mjs pass 0, so the build child retires right after the compile. Its rustc grandchild exits with it on stdin EOF.
  • Correct the memory figure in the PR body.
  • Optional: retire the warm rustc stage when a non-Rust build starts, which removes the stacking.

: undefined,
readAll(child.stdout),
readAll(child.stderr),
child.wait({ check: false }).then((output) => output.ok, () => false),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SDK host failures end up as terminal compile-error verdicts.

  • What gets swallowed: () => false discards every rejection, and a resolved non-ok result is handled the same way.
  • 0.19 doesn't always reject. In a probe with a hand-built module that has an unresolvable import:
    • spawn() succeeds.
    • The SDK logs Failed to create WASI context.
    • wait() resolves { ok: false, exitCode: 45 } with empty stdout and stderr.
    • I haven't reproduced a real host-side trigger. The concern is how such a failure gets classified when it does happen.
  • How it becomes a verdict:
    • The callers return success: false with "clang exited with code 1.", "rustc failed without a diagnostic." or "TypeScript … did not return every compiled output".
    • container/server.mjs:204-206 then emits compile-error. That verdict is terminal and not retried automatically. It also counts as a contestant fault, which ICPC contests can penalize through penalizedVerdicts.
  • On main: a compiler process that produced no output and no diagnostic made the build throw, at the latest after the 55 s / 180 s output deadline.
    • The container returned 500 container-infrastructure-error.
    • The workflow retried once on a new container, and a second failure ended as infrastructure-error.
    • The PR's risk list already mentions this gap.

Fix:

  • Keep the exit code and the rejection reason.
  • Throw when the process failed without writing anything: !ok && stdout.byteLength === 0 && stderr.byteLength === 0.
    • Clang, wasm-ld and rustc report errors on stderr, including rustc's trap after diagnostics. The TypeScript wrapper answers on stdout. So real compile errors stay compile errors.
    • A throw already discards the server child or the browser Worker.
  • Add a test with a fake sandbox whose wait() resolves { ok: false } with empty streams.


private receive(chunk: Buffer): void {
if (!this.pending) return;
this.buffered += chunk.toString();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Calling chunk.toString() on each chunk corrupts a multi-byte UTF-8 character that spans two pipe chunks.

  • Where it shows up: the rustc response is one fd-3 line. It carries wasmBase64 first (0.25–0.7 MB, so several chunks), then stdout and stderr.
    • A non-ASCII identifier in a warning can arrive as unused variable: `名��` .
    • compileRust re-parses diagnostics from that stderr, so the parsed diagnostics are garbled too.
    • It needs a chunk boundary to fall inside non-ASCII stdout/stderr text, so it is rare at normal warning sizes. Builds don't fail, because the JSON stays valid.
  • Evidence:
    • A synthetic probe sent about 64 KB of CJK warnings after a 400 KB payload. 4 of 20 responses contained U+FFFD.
    • Main collected raw bytes and decoded them once.
  • Quadratic cost: each chunk also re-splits and re-measures the whole buffer.
    • A 16 MB line takes 1.35 s and a 32 MB line takes 14.3 s, against a 256 MB line limit.
    • Real Rust responses are under 1 MB, so this only matters for very large artifacts.

Fix:

  • Call responses.setEncoding("utf8") before attaching the listener, or use a StringDecoder.
  • Split only when the new chunk contains \n.
  • Track the buffered byte count incrementally.

if (id !== RESOLVED_VIRTUAL_MODULE) return null;
if (command === "serve") {
const url = (relative) => JSON.stringify(`/@fs${path.join(SDK_ROOT, relative).split(path.sep).join("/")}`);
return `export const wasmerSdkEntryUrl = ${url(ENTRY)};\nexport const wasmerSdkBindingUrl = ${url(BINDING)};\n`;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dev mode (serve) loads two instances of the wasm-bindgen module.

  • Mismatch: Vite rewrites the SDK entry's own import to /@fs/…/pkg/wasmer_sdk_js.js?v=<hash>. This line exports the binding URL without the ?v= query.
  • Result:
    • createBrowserWasmer() runs import(wasmerSdkBindingUrl) and gets a second instance that was never initialized.
    • binding.setWorkerUrl() then calls passStringToWasm0(url, wasm.__wbindgen_malloc, …) while wasm is undefined.
  • Evidence: I ran a minimal Vite 8 dev server with only this plugin. It serves the entry with the ?v= import, while the virtual module exports the binding URL without it.
  • Scope:
    • Production builds are fine, because both URLs resolve to the same emitted asset.
    • Nobody can hit this while pnpm dev is broken on main, but the branch can't work as written.

Fix: fail fast in serve with a clear "dev mode unsupported" error until dev mode is fixed and tested. Alternatively, export the same specifier Vite gives the entry's import, so both resolve to one module instance.

function cspSafeRuntimeSource(relative, source) {
if (relative !== BINDING) return source;
const text = source.toString("utf8");
if (text.split(DYNAMIC_FUNCTION).length !== 2) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: this guard pins the glue call site but not the trampoline body.

  • Failure case: a future SDK changes the body string inside the wasm, for example to rest args.
    • The build still passes.
    • __wasmOjFunction falls through to new Function.
    • Browser compiles and Python runs then fail at runtime under the strict CSP.
  • Nothing would catch it: CI doesn't run verify-browser-csp.mjs.
  • Covered today: the exact 0.19.0 pin and the per-file digests in verify-wasmer-sdk-init.mjs.

Cheap fix: in wasmerSdkRuntimeFiles(), also assert that pkg/wasmer_sdk_js_bg.wasm contains return f(Array.prototype.slice.call(arguments, 1)). It occurs 3 times in 0.19.0.

Comment thread CHANGELOG.md
- Python execution reads the interpreter module directly from the pinned WEBC package and
verifies its digest.
- Raw Clang output names `/workspace/...` instead of `/project/...` source paths. Parsed
diagnostics and emitted artifacts are unchanged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: this doesn't hold for link errors.

  • The wasm-ld diagnostic message is built from stderr.trim().
  • So wasm-ld: error: /project/build/0000.o: undefined symbol: missing becomes …/workspace/build/0000.o….
  • I saw this when comparing ServerCompiler output on main and on this PR, and in Chromium on the PR build.
  • Clang file, line and column are unchanged.

Fix, either:

  • Reword it, e.g. "Raw Clang and wasm-ld output, including link-error messages, names /workspace/...".
  • Or map /workspace/ back to /project/ in the failure message.

Comment thread CHANGELOG.md
- Raw Clang output names `/workspace/...` instead of `/project/...` source paths. Parsed
diagnostics and emitted artifacts are unchanged.
- The runtime identity changes with the SDK, so artifacts cached by earlier releases are
rebuilt.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: hosts with a persistent artifact cache get one failed compile before the rebuild.

  • Why:
    • FileSystemArtifactStore.load() rejects the stale cost profile, deletes the file and rethrows.
    • CompileCoordinator.loadCached doesn't catch that error.
    • So the first compile after the upgrade fails with Artifact cost profile '…:runtime-<previous identity>:…' does not match …. Only the next attempt rebuilds.
  • Who it hits: woj (persistent engine/artifacts) and long-lived server hosts. The judge container is not affected (artifactCache: false).
  • Not new: any identity change takes this path, including fix(runtime): meter interactive programs in-module like standalone runs #95's.

Fix, either:

  • Treat a load() failure as a cache miss in loadCached (one try/catch).
  • Or drop "rebuilt" from this line.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants