Skip to content

fix(browser): run interactive sides in nested Workers - #99

Merged
JacobLinCool merged 10 commits into
wasm-oj:mainfrom
TakalaWang:fix/browser-interactive-workers
Oct 8, 2026
Merged

JacobLinCool merged 10 commits into
wasm-oj:mainfrom
TakalaWang:fix/browser-interactive-workers

Conversation

@TakalaWang

Copy link
Copy Markdown
Contributor

Fixes #98.

Why

Browser Engine.interact fails on every real dialogue. The runner Worker omitted startupEntropyBytes, so runtime-core rejected the request. With the field added, the single-threaded web runtime-core panics in Condvar::wait as soon as either side reads an empty pipe. Repro and root cause are in #98.

Change

Each side of a browser interactive session now runs in its own nested Worker as a standalone metered run. Metering, the logical clock, the filesystem quota and the output budget are the same as run.

  • runtime-core
    • run/web.rs: run now delegates to execute, which takes a stdin file and an optional stdout peer and also returns the raw exit code. The standalone path behaves as before.
    • New web-only run/web_interactive.rs runs one side with streaming stdio:
      • stdin's poll_read and poll_read_ready call blocking JS callbacks, so WASIX never sees Pending.
      • stdout records into the protocol capture first, then forwards the bytes, exactly like the native InteractiveOutput. A closed peer gives BrokenPipe.
      • Exit codes are the raw WASI exit code, as in native interact.
    • New export run_interactive_side. interact_wasm_oj is removed from the web build. Native interact and interactive.rs are untouched.
  • Browser
    • src/runtime/interactive-pipe.ts: a single-producer, single-consumer SharedArrayBuffer ring. It tracks EOF and the reader closing, and uses a sequence word so no wake-up is lost. Blocking uses Atomics.wait, which only runs inside the side Workers.
    • src/runtime/interactive-side.worker.ts: initializes runtime-core from the runner's compiled WebAssembly.Module, runs one side, and closes both of its pipe ends as soon as the program exits. The peer then sees EOF, or EPIPE on write.
    • src/runtime/runner.worker.ts:
      • Compiles runtime-core once at initialization and starts the two side Workers through the existing same-origin blob bootstrap.
      • Reports running only once both sides execute, so the wall timer starts at the same point as for run.
      • Maps the two side results into the unchanged InteractiveRunResult.
      • Terminates both side Workers when the session ends or fails.
      • Sends startupEntropyBytes with each program.
    • Cancellation and wall time are unchanged: BrowserRunner terminates the runner Worker, which ends its side Workers. In Chromium, CDP shows both side Workers gone after cancel.
  • Rebuilt runtime-core_bg.wasm (pnpm run runtime:build, wasm-bindgen 0.2.127) and refreshed the three identity pins.
  • docs/architecture.md note and a CHANGELOG entry.

The public API and result shape do not change. No unsafe-eval is needed, and Workers stay same-origin.

Verification

  • New interactive fixtures in the strict-CSP suite, run with node scripts/verify-browser-csp.mjs interactive and WASM_OJ_BROWSER=chromium|firefox|webkit. All use C++ interactors.
Fixture What it checks Chromium Firefox WebKit
interactive-ac-cpp C++ binary search, 20 rounds, AC pass pass pass
interactive-ac-c C contestant, AC pass pass pass
interactive-wa-c interactor exits 1 (WA); contestant sees EOF pass pass pass
interactive-python-contestant CPython runtime-bundle contestant, AC pass pass pass
interactive-instruction-limit spinning contestant stops at instruction-limit; interactor sees EOF pass (0.94 s) pass (2.4 s) pass (0.67 s)
interactive-contestant-exits contestant exits at once; interactor's read returns EOF pass pass pass
interactive-interactor-exits interactor exits; contestant's write fails with EPIPE pass pass pass
interactive-output-flood contestant capped at outputLimitBytes (64 KiB): output-limit, transcript exactly 64 KiB pass pass pass
interactive-wall-time deadlock ends at the 2 s wall limit for both sides pass pass pass
interactive-cancel engine.cancel() mid-dialogue rejects within ~1 s, then the next interact succeeds pass pass pass
  • The whole strict-CSP suite passes in Chromium: 99 checks, exit 0.
  • Firefox: every interactive fixture passed, but no single Firefox run was fully clean. Each run had one compile that exceeded the 55-60 s compiler boundary, on a different fixture each time and once before any interact. main shows the same stall: the first compile hung in 1 of 3 runs of these fixtures, so it is not caused by this PR. It looks like the C++ compile stall that fix(runtime): stop waiting for stream EOF when exporting runtime files #96 leaves open.
  • Timings for the 20-round guess-the-number, interact only, warm. Medians of 5 runs after one cold run:
Contestant (vs C++ interactor) Chromium Firefox WebKit
C 14 ms (cold 15) 64 ms (cold 57) 16 ms (cold 17)
C++ iostream + std::endl 64 ms (cold 89) 446 ms (cold 425) 68 ms (cold ~1.0 s)
CPython 0.89 s (cold 1.8 s) 5.5 s (cold 10.9 s) 0.86 s (cold 2.8 s)

Cold CPython includes the page's one-time runtime-file export. Every CPython session also pays CPython start-up, which is about 2.4e9 metered instructions.

Risks

  • Cross-origin isolation. Pages need SharedArrayBuffer. The runner already refuses to start without cross-origin isolation, so hosts are not affected.
  • Per-session overhead. Each session starts two Workers and instantiates runtime-core in each, from a module compiled once. A warm C dialogue takes about 15-60 ms in total.
  • poll_oneoff on stdin with a timeout now waits for data instead of timing out. The native path's 1-ns readiness probe depends on timing anyway.
  • Costs.
  • Conflicts.
  • Not changed here.
    • BrowserRunner still has no runTrusted / interactTrusted, so engine.judge() cannot run interactive cases or Wasm checkers in the browser. Direct interact and run work.
    • Pre-existing on main: in WebKit, an empty for(;;); loop is not stopped by the instruction meter and runs until the wall limit, in plain run as well. A loop with a side effect stops at instruction-limit in about 0.7 s.
    • Pre-existing on main: an occasional Firefox compile stall; see Verification.

🤖 Generated with Claude Code

@JacobLinCool

Copy link
Copy Markdown
Member

Two correctness issues to address before merging:

  • Premature EOF race (src/runtime/interactive-pipe.ts:51–53): after the reader observes an empty buffer, the writer can publish its final bytes and close before the reader checks WRITER_CLOSED. wait() then returns EOF despite buffered data. I reproduced this interleaving with the PR's pipe implementation: the first read returned empty; the next returned 42\n. Recheck buffered data after observing closure and add a regression test for this interleaving.
  • Polling timeout semantics (StreamInput::poll_read_ready): calling blocking streams.wait() prevents a stdin + clock poll_oneoff from returning on its deadline. A valid timeout-driven dialogue can hang until the session wall limit. Please preserve readiness/deadline semantics and cover an empty stdin with a clock subscription.

When rebasing onto #95, rebuild the generated runtime and identity pins from the combined Rust sources.

TakalaWang and others added 2 commits October 7, 2026 23:14
Browser Engine.interact failed on every dialogue (wasm-oj#98). The runner Worker
omitted startupEntropyBytes, and with it added the single-threaded web
runtime-core panicked in Condvar::wait as soon as either side read an
empty pipe.

Each side now runs in its own nested Worker as a standalone metered run.
stdin and stdout are SharedArrayBuffer ring buffers whose reads and writes
block with Atomics.wait, so WASIX never sees Pending. A side's pipe ends
close when its program exits, giving the peer EOF or EPIPE. The result
shape, transcripts and terminations match the server's interact.

Fixes wasm-oj#98

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With runtime-bundle interactors accepted (wasm-oj#93), run a CPython guessing
interactor against a compiled C contestant through the nested-Worker
interactive path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TakalaWang
TakalaWang force-pushed the fix/browser-interactive-workers branch from 73e16da to bd3e37b Compare October 7, 2026 15:23
@TakalaWang

Copy link
Copy Markdown
Contributor Author

Rebased onto main after #93 and #95: rebuilt the runtime with pnpm run runtime:build and refreshed the identity pins (identity 1c09fb59…), and added a CPython interactor fixture to the strict-CSP suite now that #93 accepts runtime-bundle interactors. All 11 interactive fixtures (AC, WA, Python contestant and interactor, instruction-limit, early exits, output flood, wall limit, cancel) pass in Chromium, Firefox and WebKit, and cargo test, clippy (native and web, -D warnings), cargo fmt --check and pnpm run ci:verify pass locally.

…eractive sides

The pipe reader could report EOF with bytes still buffered: it saw an
empty buffer, the writer then published its last bytes and closed, and
the reader saw the close. The reader now checks the buffer again after
it sees the close.

Stdin readiness blocked in Atomics.wait, so a poll_oneoff on stdin and
a clock never returned at its deadline. When the deterministic poll
probes readiness for a clock subscription, stdin now checks the pipe
without blocking. If nothing is readable, the virtual clock advances to
the deadline, as on the server. Polls without a clock still block.

The runtime and identity pins are rebuilt from the combined sources.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TakalaWang

Copy link
Copy Markdown
Contributor Author

@JacobLinCool Both are fixed in d901464.

  • Premature EOF: InteractivePipeReader.poll(), which wait() now uses, reads the buffer again after it sees WRITER_CLOSED and returns any bytes. It reports EOF only when the writer is closed and the buffer is still empty. A new test in interactive-pipe.test.ts forces your interleaving: the writer writes 42\n and closes right after the reader's first empty buffered(). Before the fix, that read came back empty.
  • Poll deadlines: clocks here are virtual. The deterministic poll_oneoff handles stdin + clock by probing readiness, then advancing the virtual clock to the deadline. During that probe, StreamInput::poll_read_ready now checks the pipe through a new non-blocking poll callback instead of Atomics.wait; polls without a clock still block. The server behaves the same way, and the new native test interactive_empty_stdin_poll_times_out_on_the_process_clock pins that behaviour there. Two new strict-CSP fixtures cover the browser. interactive-poll-timeout is a C poll(stdin, 1000) with no input, which hit the 10 s wall limit before the fix and now returns in about 15 ms. interactive-poll-ready checks that a reply arriving later is still reported as readable.

I rebuilt the runtime and identity pins from the combined sources with pnpm run runtime:build. All 13 interactive fixtures pass in Chromium, Firefox and WebKit. cargo test, clippy (native and web, -D warnings), cargo fmt --check, runtime:check-web and pnpm run ci:verify pass.

@TakalaWang

Copy link
Copy Markdown
Contributor Author

The verify failure on d901464 is src/server/python-compiler.test.ts hitting its 300 s timeout, the server-side Python runtime-preparation stall that #96 fixes. This PR doesn't touch that path, and the previous head was green. Merging #96 first, then rebasing this one (CHANGELOG-only conflict, no runtime rebuild needed), will rerun CI with the fix in place.

@TakalaWang

Copy link
Copy Markdown
Contributor Author

@JacobLinCool Both issues are addressed in d901464. Ready for another review.

Browser and server verdicts differed for the same program pair in four
places:

- A write the reader closed partway returned the partial count, so the
  guest retried the tail and the transcript and output budget counted it
  twice. It now fails with EPIPE, as the server's pipe does.
- Closing fd 1 or fd 0 reached the peer only after the side exited.
  StreamOutput and StreamInput now close their pipe end when WASIX drops
  the last descriptor that refers to them, which is when the server
  drops its PipeTx or PipeRx.
- The fixed 64 KiB ring blocked writers that the server's unbounded pipe
  never blocks, deadlocking batch protocols. Each ring now holds its
  writer's whole output budget. Only budgeted stdout bytes enter it, so a
  write never waits.
- A poll_oneoff without a clock blocked on empty stdin even when another
  subscription was ready. Its first stdin readiness check per read
  subscription now returns Pending, so WASIX reports the ready
  subscription. Stdin blocks only when WASIX polls again because nothing
  else was ready.
Both side Workers are now created inside the try that terminates them,
so a failure starting the interactor no longer leaves the contestant
Worker running, and a side whose start message cannot be posted is
terminated at once. A side's error response now says which side failed,
as the exception path already did.
Rebuild the web runtime-core with wasm-bindgen 0.2.127 and refresh the
runtime identity pins. The refreshed identity changes cost profiles.
Add strict-CSP interactive fixtures for a contestant that closes stdout
before it reads the verdict, a batch protocol with 240 KB unread in each
direction, and a clockless poll on stdin and stdout before the first
query. All three ended in wall-time-limit on the previous runtime.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Browser Engine.interact panics as soon as either side waits for input

2 participants