Summary
interact() in the browser runner (@wasm-oj/browser) fails on any real dialogue. On main (28cd0ac) it fails in two stages:
- The runner Worker does not send
startupEntropyBytes, so runtime-core rejects the request: invalid interactive request: missing field startupEntropyBytes. interactiveCoreProgram() in src/runtime/runner.worker.ts omits the field. InteractiveProgram.startup_entropy_bytes in crates/runtime-core/src/types.rs has no #[serde(default)]. The server path, encodeNativeInteractiveProgram, sends it.
- With the field added, runtime-core panics with
RuntimeError: unreachable. The console shows:
panicked at library/std/src/sys/sync/condvar/no_threads.rs:20:9:
condvar wait not supported
panicked at library/std/src/sys/sync/mutex/no_threads.rs:19:9:
assertion `left == right` failed: cannot recursively acquire mutex
The server (native) interact is not affected. There is no browser test for interact, so neither failure showed up in CI.
Reproduction
Chromium, strict-CSP page (the setup of scripts/verify-browser-csp.mjs), release builds:
const contestant = (await engine.compile({ language: "c", target: "wasip1", optimization: "release", entry: "main.c",
files: { "main.c": '#include <stdio.h>\nint main(void){int x;if(scanf("%d",&x)!=1)return 2;printf("%d\\n",x+1);fflush(stdout);return 0;}' } })).artifact;
const interactor = (await engine.compile({ language: "c", target: "wasip1", optimization: "release", entry: "main.c",
files: { "main.c": '#include <stdio.h>\nint main(int c,char**v){FILE*f=fopen(v[1],"r");int t,a;fscanf(f,"%d",&t);printf("%d\\n",t-1);fflush(stdout);if(scanf("%d",&a)!=1)return 2;return a==t?0:1;}' } })).artifact;
await engine.interact(contestant, interactor, { interactor: { args: ["/judge/input.txt"], files: { "/judge/input.txt": "42\n" } } });
// main: WasmOjError: invalid interactive request: Error: missing field `startupEntropyBytes`
// with the field added: WasmOjError: Uncaught RuntimeError: unreachable (panic above)
What works and what panics, with the field added (#93 and #95 merged as well, so CPython interactors and the in-module meter are in place):
| Session |
Result |
| Contestant reads first, interactor writes first |
panic |
| Contestant writes first, then reads |
panic |
| One-way: contestant writes and exits, interactor reads |
works (40 ms) |
| One-way: interactor writes, contestant reads |
panic |
| CPython interactor + C++ contestant |
panic |
C++ for(;;); contestant + interactor that writes, then reads |
contestant instruction-limit in 1.3 s; the interactor's read sees EOF |
A session only finishes when no side ever waits on an empty pipe.
Root cause
runtime-core's web build is single-threaded: wasm32-unknown-unknown without atomics, so std uses the no_threads Mutex and Condvar. interact (crates/runtime-core/src/interactive.rs) spawns both programs through WebTaskManager (crates/runtime-core/src/run/web_runtime.rs). Its task_wasm runs each program with spawn_local on the runner Worker's only thread.
The contestant starts first and runs synchronously until it reads stdin. If the interactor has not written yet, PipeRx::poll_read returns Poll::Pending. WASIX fd_read goes through __asyncify_light, which calls virtual_mio::block_on. That function parks the thread on a parking::Parker, which is a std Condvar. Without threads, Condvar::wait panics. Nothing else can run on the thread to fill the pipe anyway, so even a working park would deadlock. The second panic happens while the first one unwinds.
Proposed fix
Run each side in its own nested Worker, as a standalone metered run whose stdin and stdout are streams. The page is already cross-origin isolated, as the runner requires, so SharedArrayBuffer and Atomics are available.
- The runner Worker prepares both programs as today. It creates two
SharedArrayBuffer ring buffers, contestant→interactor and interactor→contestant, starts two side Workers, and gives each the compiled runtime-core WebAssembly.Module.
- Each side Worker calls a new runtime-core export. It runs one program through the same path as
run: same in-module meter, logical clock, filesystem quota and output budget. The difference is stdin and stdout:
- stdin is a
VirtualFile whose poll_read blocks in the Worker with Atomics.wait until bytes arrive or the writer closes. It never returns Pending to WASIX, so block_on never parks.
- stdout records into the protocol capture first, as the native
InteractiveOutput does, then writes to the peer's ring. It blocks while the ring is full and returns BrokenPipe once the peer has closed its read end.
- When a program exits, its Worker closes both of its ring ends, so the peer sees EOF or
EPIPE.
- The runner Worker turns the two side results into the existing
InteractiveRunResult shape.
- Wall time and cancellation keep working as today:
BrowserRunner terminates the runner Worker, and the side Workers are terminated with it, and also explicitly in a finally.
- Instruction and logical-time limits are per side, as in a standalone run.
- Workers stay same-origin, using the existing blob bootstrap in
module-worker.ts. No unsafe-eval.
Side effect: browser interactive costs become exactly standalone run costs, because each side is a standalone run. #95 brings the server's interactive meter to the same values, so after both land, the server and the browser agree.
Alternatives considered
- JSPI (
WebAssembly.Suspending / promising). The blocking call is deep inside WASIX (virtual_mio::block_on) with runtime-core and guest frames above it. JSPI would mean forking wasmer-wasix and wasmer's JS backend so the park becomes a suspending JS import, wrapping guest exports with promising, and scheduling the two programs as cooperating promises. As far as I know, Safari does not ship JSPI, and downstream hosts need Safari.
- Asyncify / WASIX deep sleep. WASIX can unwind a guest that exports the Asyncify entry points. Every contestant, interactor and the CPython bundle would need a Binaryen
--asyncify pass in the browser, which means a multi-MB binaryen.js. Asyncify grows code size and slows execution, and it changes the metered instruction stream. Costs would drift from the server unless the server applied the same transform.
- Threaded runtime-core (atomics,
build-std, shared memory, a WebTaskManager that spawns Workers like @wasmer/sdk). This needs nightly build-std and turns every web run into a shared-memory build, standalone runs included. It is far larger than two side Workers.
- Cooperative re-entry (when one side would block, run the other side inline). Both sides block in the middle of their own stacks. Without stack switching, nested execution deadlocks as soon as the inner program waits for the outer one.
- Running the interactor elsewhere (for example on a server). That is a product choice for hosts, not a fix for the browser runner, and every turn becomes a network round trip.
Scope estimate
crates/runtime-core:
- New web-only streaming stdio files, about 200 lines.
run/web.rs takes a stdin file and an optional stdout peer, plus a side entry point, about 80 lines.
web.rs gains an export for one side and drops interact_wasm_oj.
- A request/response type in
types.rs.
interactive.rs is untouched; the native interact stays as it is.
src/runtime: the runner.worker.ts interact path, about 100 lines, and a new interactive-side.worker.ts, about 100 lines. The startupEntropyBytes omission goes away with the old path.
- Tests: interactive fixtures in
scripts/verify-browser-csp.mjs, covering AC/WA, EOF, broken pipe, the instruction limit, an output flood, cancellation and CPython; unit tests for the ring buffer.
- Generated:
src/runner/generated/runtime-core.{js,d.ts}, runtime-core_bg.wasm (LFS), and the three pins in src/core/runtime-identity.ts: runtimeCoreWasmSha256, runtimeSourceRootSha256 and WASM_OJ_RUNTIME_IDENTITY_SHA256.
- Conflicts:
I'm working on a PR along these lines.
Summary
interact()in the browser runner (@wasm-oj/browser) fails on any real dialogue. Onmain(28cd0ac) it fails in two stages:startupEntropyBytes, so runtime-core rejects the request:invalid interactive request: missing field startupEntropyBytes.interactiveCoreProgram()insrc/runtime/runner.worker.tsomits the field.InteractiveProgram.startup_entropy_bytesincrates/runtime-core/src/types.rshas no#[serde(default)]. The server path,encodeNativeInteractiveProgram, sends it.RuntimeError: unreachable. The console shows:The server (native)
interactis not affected. There is no browser test forinteract, so neither failure showed up in CI.Reproduction
Chromium, strict-CSP page (the setup of
scripts/verify-browser-csp.mjs), release builds:What works and what panics, with the field added (#93 and #95 merged as well, so CPython interactors and the in-module meter are in place):
for(;;);contestant + interactor that writes, then readsinstruction-limitin 1.3 s; the interactor's read sees EOFA session only finishes when no side ever waits on an empty pipe.
Root cause
runtime-core's web build is single-threaded:
wasm32-unknown-unknownwithout atomics, sostduses theno_threadsMutexandCondvar.interact(crates/runtime-core/src/interactive.rs) spawns both programs throughWebTaskManager(crates/runtime-core/src/run/web_runtime.rs). Itstask_wasmruns each program withspawn_localon the runner Worker's only thread.The contestant starts first and runs synchronously until it reads stdin. If the interactor has not written yet,
PipeRx::poll_readreturnsPoll::Pending. WASIXfd_readgoes through__asyncify_light, which callsvirtual_mio::block_on. That function parks the thread on aparking::Parker, which is astdCondvar. Without threads,Condvar::waitpanics. Nothing else can run on the thread to fill the pipe anyway, so even a working park would deadlock. The second panic happens while the first one unwinds.Proposed fix
Run each side in its own nested Worker, as a standalone metered run whose stdin and stdout are streams. The page is already cross-origin isolated, as the runner requires, so
SharedArrayBufferandAtomicsare available.SharedArrayBufferring buffers, contestant→interactor and interactor→contestant, starts two side Workers, and gives each the compiled runtime-coreWebAssembly.Module.run: same in-module meter, logical clock, filesystem quota and output budget. The difference is stdin and stdout:VirtualFilewhosepoll_readblocks in the Worker withAtomics.waituntil bytes arrive or the writer closes. It never returnsPendingto WASIX, soblock_onnever parks.InteractiveOutputdoes, then writes to the peer's ring. It blocks while the ring is full and returnsBrokenPipeonce the peer has closed its read end.EPIPE.InteractiveRunResultshape.BrowserRunnerterminates the runner Worker, and the side Workers are terminated with it, and also explicitly in afinally.module-worker.ts. Nounsafe-eval.Side effect: browser interactive costs become exactly standalone
runcosts, because each side is a standalone run. #95 brings the server's interactive meter to the same values, so after both land, the server and the browser agree.Alternatives considered
WebAssembly.Suspending/promising). The blocking call is deep inside WASIX (virtual_mio::block_on) with runtime-core and guest frames above it. JSPI would mean forking wasmer-wasix and wasmer's JS backend so the park becomes a suspending JS import, wrapping guest exports withpromising, and scheduling the two programs as cooperating promises. As far as I know, Safari does not ship JSPI, and downstream hosts need Safari.--asyncifypass in the browser, which means a multi-MBbinaryen.js. Asyncify grows code size and slows execution, and it changes the metered instruction stream. Costs would drift from the server unless the server applied the same transform.build-std, shared memory, aWebTaskManagerthat spawns Workers like@wasmer/sdk). This needs nightlybuild-stdand turns every web run into a shared-memory build, standalone runs included. It is far larger than two side Workers.Scope estimate
crates/runtime-core:run/web.rstakes a stdin file and an optional stdout peer, plus a side entry point, about 80 lines.web.rsgains an export for one side and dropsinteract_wasm_oj.types.rs.interactive.rsis untouched; the nativeinteractstays as it is.src/runtime: therunner.worker.tsinteract path, about 100 lines, and a newinteractive-side.worker.ts, about 100 lines. ThestartupEntropyBytesomission goes away with the old path.scripts/verify-browser-csp.mjs, covering AC/WA, EOF, broken pipe, the instruction limit, an output flood, cancellation and CPython; unit tests for the ring buffer.src/runner/generated/runtime-core.{js,d.ts},runtime-core_bg.wasm(LFS), and the three pins insrc/core/runtime-identity.ts:runtimeCoreWasmSha256,runtimeSourceRootSha256andWASM_OJ_RUNTIME_IDENTITY_SHA256.runtime-core_bg.wasmand the pins. Whichever lands second rebuilds them withpnpm run runtime:buildand refreshes the pins.run/web.rsonly conflicts if hunks overlap; fix(runtime): meter interactive programs in-module like standalone runs #95 changes one line there.kindcheck inrunner.worker.ts; the fix keeps that region untouched.I'm working on a PR along these lines.