Skip to content

feat(runtime): allow Python interactors in interact - #93

Merged
JacobLinCool merged 1 commit into
wasm-oj:mainfrom
TakalaWang:feat/runtime-bundle-interactor
Oct 7, 2026
Merged

JacobLinCool merged 1 commit into
wasm-oj:mainfrom
TakalaWang:feat/runtime-bundle-interactor

Conversation

@TakalaWang

@TakalaWang TakalaWang commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

interact refuses any interactor written in Python, even though Python interactors run in the same Wasm sandbox, under the same limits, as Python contestants, which interact already accepts. This PR removes that one refusal so Python interactors work. No sandbox or resource limit changes.

Background: what the guard actually checks

forge has two kinds of build artifact:

Kind What it is Example
wasm one compiled Wasm module C, C++, Rust, Go
runtime-bundle an interpreter Wasm module plus the source files it runs Python (CPython Wasm + main.py)

Both kinds execute inside the same Wasm runtime, under the same policy:

  • instruction budget and memory limit;
  • no network, no process or thread creation;
  • a virtual filesystem holding only the files the host passes in.

A runtime bundle is not less isolated than a standalone module. It is still Wasm; it just carries its interpreter with it.

ServerRunner.interact and the browser runner's interactArtifacts both start with:

if (interactorArtifact.kind !== "wasm") {
  throw new Error("Interactive judge artifacts must be standalone Wasm modules.");
}

So the check is about the artifact's packaging, not about safety. The property interaction really needs is that a program can read fd 0 as a stream while the other side is still writing. That is already checked per runtime driver, for both sides, in prepareArtifactInteraction (src/runner/artifact.ts):

if (driver.interactive !== "streaming") {
  throw new Error(`Runtime driver '${driver.id}' does not support streaming interactive execution.`);
}

The CPython driver is streaming, which is why a Python contestant already works in interact. The kind guard only blocks the same driver on the interactor side. Bundles whose driver cannot stream fd 0 are still rejected by the driver check.

Why it matters

DOMjudge-style interactors are commonly written in Python. In NOJV, which uses @wasm-oj/server to run problem interactors on the server for in-browser Test, all four interactive problems use Python interactors. With the guard in place, Test is unavailable on every one of them; with it removed, they run end to end.

Change

  • Remove the kind guard in ServerRunner.interact (src/server/server-runner.ts) and in the browser runner Worker's interactArtifacts (src/runtime/runner.worker.ts). The driver check above stays the only gate.
  • docs/library-contract.md: either side may be a standalone Wasm module or a runtime bundle; bundles that cannot stream fd 0 are still rejected.
  • CHANGELOG.md: Unreleased entry.

Verification

  • New server integration test, "runs a CPython interactor against a compiled contestant" (src/server/judge.integration.test.ts):
    • The CPython interactor reads the secret from its input file, writes a challenge and checks the C contestant's reply.
    • It expects both sides to exit 0 and the transcript 42\n / 41\n.
    • With the guard restored, it fails with "Interactive judge artifacts must be standalone Wasm modules."
  • WASM_OJ_RUN_JUDGE_INTEGRATION=1 pnpm exec vitest run src/server/judge.integration.test.ts: 3/3 pass.
  • pnpm run ci:verify passes: 194 files, 956 tests, plus build.
  • Downstream, through createServerEngine: a CPython guess-the-number interactor against a C++ binary-search contestant returns interactor exit 42 with the full transcript, in about 3 s.

Not changed

Resource policies, metering, the trusted-judge path, and the streaming capability check are untouched. A Python interactor gets exactly the limits a Python contestant gets.

🤖 Generated with Claude Code

Runner.interact rejected any interactor that was not a standalone Wasm
module, although both sides already go through the same interactive
preparation and CPython provides streaming fd 0. Drop the guard on the
server runner and the browser runner Worker so Python interactors work,
and cover it with a server integration test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TakalaWang

Copy link
Copy Markdown
Contributor Author

For context: NOJV (a downstream user) will run checkers and interactors in the browser for its Test button. It needs this PR (its four interactive problems use Python interactors), #95 and #99 (which fixes browser interact, #98) in one release. With #99 merged too, a CPython interactor and a C++ contestant give AC in Chromium, Firefox and WebKit. This PR and #99 merge cleanly apart from the CHANGELOG.

@JacobLinCool
JacobLinCool merged commit 010b2f0 into wasm-oj:main Oct 7, 2026
1 check passed
TakalaWang added a commit to TakalaWang/forge that referenced this pull request Oct 7, 2026
With runtime-bundle interactors accepted (wasm-oj#93), run a CPython guessing
interactor against a compiled C contestant through the nested-Worker
interactive path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TakalaWang added a commit to NOJV-TW/NOJV that referenced this pull request Oct 9, 2026
0.2.4 ships wasm-oj/forge#93, #94, #95, #96 and #99: runtime-bundle
interactors, in-module metering of interactive programs, the runtime-file
export stall fix and interactive sides in nested Workers. Its core and
contracts move to 0.2.4 with it. The toolchain packages stay at 0.2.0,
which forge versions independently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TakalaWang added a commit to NOJV-TW/NOJV that referenced this pull request Oct 10, 2026
…645)

* docs: plan browser-run checker and interactive Test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: remove server-side Test judging

Test will run checkers and interactors in the student's browser, so the
server half of #641 goes: the test-judge API route, limiter, Redis lock
and storage keys, the application test-judge domain and its build
dispatch after a judge-config save, the test-judge queue, workflows and
WORKER_MODE=test, the WASM-OJ runtime layers and toolchains in the
worker image, the worker-test chart Deployment, web's TEST_JUDGE_ENABLED,
and @wasm-oj/server with its patch. Core drops the request, response and
record schemas, the judge-program cache key and server identity, and the
artifact wire format. The edit page no longer shows a judge-program build
status or offers "check samples with the checker". Renovate updates
@wasm-oj/* and the rust image again.

Core keeps what the browser will use: the checker and interactive verdict
mapping, truncateUtf8, the WASM-OJ termination verdict, the DOMjudge
Python wrappers, judgeProgramCompileInput and
interactiveContestantSupported. The Test capability is no longer a
server field; the editor disables Test on special_env problems itself.

Until browser judging lands, checker and interactive Test report that
Test isn't available for the problem. The contest participation and
window checks on code drafts keep their coverage on listCodeDrafts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(problems): read a problem's checker or interactor source

GET /api/problems/[id]/judge-program?context=… returns the role,
language, source and SHA-256 of a checker or interactive problem's judge
program, read through its verified storage pointer, so browser Test can
compile and run it. Standard and special_env problems, and a problem
without a stored program, answer 404.

Access is the problem page's view access for the context: a context
grants its problem while the page for it would render (course staff and
contest organizers always; students in an open assignment, a running
contest they joined, their own virtual run, or a running exam session
that passes the proctoring gate), and otherwise the practice rules
apply, so the program stays readable after an exam or contest ends for
the students who can still view the problem. A page-locked exam session
keeps the request inside that exam's problems. The route uses the
standard API limiter and is exam-scoped; /api/drafts shares its context
query parser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: format the browser Test plan and spec

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: add the Test stopgap to the dead-code checklist

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(problems): match the judge-program window to its pages

A contest grants its problem's judge program only while the contest is
published, organizers included, because the contest workspace never
renders a draft contest. A running contest stays open through its end
instant, as the contest problem page redirects only after endsAt.

The JSON context query parser is now parseJsonContextParam, so it no
longer reads like parseContextQuery.

Integration cases cover the denials: a contest before its start, a draft
contest for its participant and organizer, an assignment before it
opens, in an archived course or as a draft, an exam session on a
non-whitelisted IP, a problem outside a non-locked exam, and an ended
virtual run. A closed assignment stays readable to its enrolled student
through the ended-assignment view rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): prepare a checker in the browser when the editor opens

On a checker problem the editor fetches the checker's source from
/api/problems/[id]/judge-program for its context and builds it with the
shared browser engine through judgeProgramCompileInput: a Python checker
is packaged with the DOMjudge wrapper, a C++ checker compiles with the
libc++ PCH shim. Its toolchain preloads next to the editor language's.
Each editor mount refetches the source; a build is kept for the page
session per problem and SHA-256, so it reruns only when the source
changes. Source requests retry like the toolchain preload, except for a
refused request.

While the checker prepares, Test stays clickable and shows the
toolchain download or "Preparing checker...". A failed build disables
Test with "This problem's checker failed to build." and the Test Result
panel shows the compiler output; a checker that cannot be loaded
disables Test with a reload hint. Interactive problems keep the
unavailable notice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): judge checker samples in the browser

Test on a checker problem waits for the prepared checker, compiles and
runs the selected cases as before, then runs the checker in the browser
on every case that exited normally and matches a sample by input:
/judge/input and /judge/answer hold the sample's input and output, the
student's output is its stdin, and its team message is read from
/judge/feedback/teammessage.txt. The checker gets official judging's
validator time, 512 MiB and the same 2x wall deadline. Exit 42 is AC,
43 is WA and anything else is SE, through checkerCaseVerdict. Custom
cases still only execute.

The result panel loses the "Judged on the server" badge and the server
notice; it explains a system error only on a case the browser checker
judged. Interactive problems keep the unavailable notice until browser
interaction lands.

browser-local-run no longer exports its internal run and compile
outcome types.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(problems): drop the hidden workspace file visibility

Hidden workspace files never stayed secret: official judging and student
code could read them, and browser Test would ship them. Production had none.
A migration turns any hidden file into a readonly one and recreates the
WorkspaceFileVisibility enum without hidden. Stored judge snapshots read a
legacy hidden file as readonly, so old submissions still rejudge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): queue browser engine work and keep judge-program failures visible

The browser engine allows one foreground compile at a time, but a
checker build kept running after the student left the problem, so the
next problem's checker failed to load or its Test reported a system
error. Every engine compile, run and interaction now goes through one
module-level queue. A judge-program build that nobody waits for any
more keeps running, because the per-digest memo wants its result. A
caller that aborts while waiting leaves the queue; the engine is
cancelled only when the aborting caller's own operation is running.

- The source response is parsed with judgeProgramSourceViewSchema from
  core, which the application's source view now shares.
- Fetch, toolchain and engine failures are logged with console.warn.
- A 403 or 404 source response disables Test without the reload hint.
- The build projectId and the memo key include the judge program's
  language.
- The checker gets the same output and filesystem limits as student
  runs.
- The Test reason is a live region, and a failed build switches the
  panel to Test Result so its diagnostics are visible.

The toolchain-preload and judge-program unit tests imported the module
graph inside beforeEach after vi.resetModules, so a cold import under a
loaded ci:verify ran into the 10 s hook timeout. They now import once at
the top and isolate state with a distinct language or problem per test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): judge interactive samples and custom cases in the browser

Test on an interactive problem now prepares the problem's interactor
when the editor opens, like a checker, and shows "Preparing
interactor...", a build failure with its diagnostics, or a load failure.
Pressing Test compiles the contestant and runs every case through
runBrowserInteraction, which calls engine.interact through the engine
queue:

- the interactor gets /judge/input /judge/answer /judge/feedback, the
  case's input as /judge/input, an empty answer and the feedback
  directory;
- the contestant gets the language-factored time limit and the
  problem's memory; the interactor gets the validator timeout and the
  problem's memory plus official judging's 64 MiB headroom (capped at
  1536 MiB); both share a wall stop of max(3 s, 3 x the factored limit)
  and student runs' output and filesystem limits;
- the verdict comes from interactiveCaseVerdict, except that a
  contestant stopped at a time or memory limit stays TLE or MLE: a
  CPython interactor that then writes to the closed pipe exits with
  120, which would otherwise turn a student's infinite loop into a
  judge error depending on timing;
- each case keeps both transcript directions (64 KiB each), the
  contestant's stderr and its time, and is marked judged so a judge
  error is explained.

The case panel edits interactor inputs on interactive problems, so
students can add their own cases. Every judge type now allows custom
cases, so the customCasesAllowed switch is gone, and so is the
"no interactor samples" state. The interactive stopgap
(client_test_judge_program, TEST_DISABLING_CODES, the controller's
disabled reason and editor_testUnavailableForProblem) is removed.
JavaScript and TypeScript contestants stay disabled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): merge interactive Test verdicts like official judging

Official interactive judging (resolveInteractiveStage) checks the
interactor first: an interactor that exits with anything but 42 or 43,
or is stopped, is SE even when the contestant hit a limit; only then does
the contestant's TLE, MLE or RE win over the interactor's AC or WA.
Core's interactiveCaseVerdict already follows that order, so browser
Test now uses it as is and drops its own rule that kept a contestant's
TLE or MLE over a failed interactor.

Official judging never breaks the interactor's pipe: its channel keeps
reading the interactor's output after the contestant ends, so the
wrapper's read() sees EOF and exits 43, and the case stays TLE. The
browser engine closes the pipe when the contestant stops, so a Python
interactor that writes after that exits 120 and the case shows SE in
Test, as does an interaction where both sides reach the shared wall stop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(decisions): Test runs entirely in the browser, judge programs included

JDG-15 now records that Test runs the problem's checker or interactor in
the student's browser, that students may read judge programs and authors
own their robustness, that the server compiles and executes nothing for
Test, and that Test never receives non-sample testcase data. It rejects
server-side judge programs (#641), server compilation, execution-only
Test and exposure toggles, keeps the earlier rejections that still hold,
and drops the rejections of shipping judge programs and compiled judge
Wasm to the browser. Generic runtime fixes land in wasm-oj/forge and are
consumed as pinned releases.

JDG-26, SEC-15 and OPS-21 are withdrawn and point to JDG-15. JDG-05 is
scoped to official judging. JDG-03 keeps the byte-identical wrapper rule
for the browser copy. JDG-12, PRB-22 and OPS-18 return to their text
before #641; PRB-03 and WEB-05 describe browser Test and the
judge-program endpoint's view access. PRB-01 and PRB-09 drop the hidden
visibility, and PRB-01 rejects it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe browser-run checker and interactive Test

The living docs drop the server test judge: the test-judge queue, worker,
workflows, route, limiter, Redis lock, storage keys, chart values,
WASM-OJ image layers and their upgrade steps, the local test-judge setup
and its runbook section, and the Test failure mode in Reliability.

The Judge Pipeline's Browser Test section now covers the judge-program
endpoint and its view access, preparation and preload when the editor
opens, the checker's args, files and limits, interactive limits and
custom interactor inputs, the engine queue, verdict mapping, the
transcript cap, and where Test differs from official judging after a
broken pipe. Architecture gets a browser Test flow; Frontend, Security,
the Threat Model and Product Sense describe readable judge programs as an
accepted risk owned by authors, with hidden testcases never leaving the
server. Workspace visibility is editable or readonly everywhere.

The Problem Test spec is rewritten as Given/When/Then for checker and
interactive Test, custom cases, build failures, JavaScript and
TypeScript on interactive problems and special_env. The Quality Ledger
drops the server-Test items, lists the forge release NOJV needs and the
upstream gaps, and adds a Wasm fast path for official judging to
evaluate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: delete the browser Test plan and spec

The work they planned has shipped on this branch, and the decision log
and living docs now hold the result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): let a contestant limit stop decide interactive Test

Official judging keeps draining the interactor's output after the
contestant ends, so when the contestant hits a time or memory limit the
interactor reads EOF, exits 43, and the case is TLE or MLE. The browser
engine closes a stopped side's pipes instead, so a Python interactor that
writes after that exits 120 and core's interactiveCaseVerdict gave SE
where Submit gives TLE or MLE; a deadlock reaching the shared wall stop
did the same.

interactiveCaseVerdict now returns TLE or MLE first when the contestant
was stopped by logical time, the instruction budget, the wall stop or the
memory limit. Every other case keeps official judging's order: an
interactor failure is SE, then a contestant RE, then the interactor's AC
or WA.

JDG-15 records the rule and why, the Judge Pipeline and Problem Test spec
describe it, and the Quality Ledger item now covers only what is left: a
contestant that exits before the interactor's next write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): tell authors students can read their checker or interactor

Test runs the checker or interactor in the student's browser, so the
judge tab now says "Students can read this program when they press Test."
under the checker and the interactor language. The tab is only on the
edit page, so only editors see it.

The checker and interactor help no longer say the program only runs in
an isolated container: it runs in the judge sandbox on Submit and in the
student's browser on Test. The interactor help and the examples drop
set_score and score.txt. Neither Python wrapper defines set_score and no
judge path reads score.txt (JDG-03 removed partial credit), so the Python
interactor example died with a NameError on every correct guess.

A component test covers the note on checker and interactive problems and
its absence on standard ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): describe judge_log and the validator timeout accurately

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): drop a queued judge-program build when its editor closes

Each editor mount queued an uncancellable checker or interactor build, so a
student who clicked through several C++-checker problems waited behind every
one of them before their own Test ran.

The editor now passes an abort signal that fires on unmount.
compileBrowserJudgeProgram uses it only while waiting in the engine queue:
a build that has not started leaves the queue and the page-session memo,
so a later visit rebuilds it, and a build that has started keeps running
so the memo still gets its result. The memo key now includes the role,
because the Python wrapper differs between checker and interactor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: align Test docs, interactor help and verdict helper with the code

- interactiveCaseVerdict drops its teamMessage parameter: no caller passed
  it, and interactive Test never shows an interactor's teammessage.
- The interactor help says Test passes an empty judge_answer, so the
  interactor should read its secret from judge_input.
- tests/tsconfig.json includes unit/worker/mailer-startup.test.ts again,
  so typecheck:tests covers it.
- The judge pipeline doc says contest organisers read the judge program
  only while the contest is published, and that Test on checker and
  interactive problems waits while the judge program prepares. It and the
  problem-test spec say the result shows the run's longest logical time,
  not a time per case.
- JDG-15 says exit codes map as in official judging but the interactive
  merge order differs, and that teammessage is shown for checkers only.
- The Quality Ledger drops the pin-bump item, which must land before this
  merges, keeping the JavaScript and TypeScript clause, and drops the
  unrelated Wasm fast path item.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: link decision sources to #645

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* build(web): pin @wasm-oj/browser 0.2.4

0.2.4 ships wasm-oj/forge#93, #94, #95, #96 and #99: runtime-bundle
interactors, in-module metering of interactive programs, the runtime-file
export stall fix and interactive sides in nested Workers. Its core and
contracts move to 0.2.4 with it. The toolchain packages stay at 0.2.0,
which forge versions independently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants