Repository navigation
feat(runtime): allow Python interactors in interact - #93
Merged
JacobLinCool merged 1 commit intoOct 7, 2026
Merged
Conversation
Runner.interact rejected any interactor that was not a standalone Wasm module, although both sides already go through the same interactive preparation and CPython provides streaming fd 0. Drop the guard on the server runner and the browser runner Worker so Python interactors work, and cover it with a server integration test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 6, 2026
This was referenced Oct 7, 2026
Contributor
Author
|
For context: NOJV (a downstream user) will run checkers and interactors in the browser for its Test button. It needs this PR (its four interactive problems use Python interactors), #95 and #99 (which fixes browser |
6 tasks done
TakalaWang
added a commit
to TakalaWang/forge
that referenced
this pull request
Oct 7, 2026
With runtime-bundle interactors accepted (wasm-oj#93), run a CPython guessing interactor against a compiled C contestant through the nested-Worker interactive path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TakalaWang
added a commit
to NOJV-TW/NOJV
that referenced
this pull request
Oct 9, 2026
0.2.4 ships wasm-oj/forge#93, #94, #95, #96 and #99: runtime-bundle interactors, in-module metering of interactive programs, the runtime-file export stall fix and interactive sides in nested Workers. Its core and contracts move to 0.2.4 with it. The toolchain packages stay at 0.2.0, which forge versions independently. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TakalaWang
added a commit
to NOJV-TW/NOJV
that referenced
this pull request
Oct 10, 2026
…645) * docs: plan browser-run checker and interactive Test Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: remove server-side Test judging Test will run checkers and interactors in the student's browser, so the server half of #641 goes: the test-judge API route, limiter, Redis lock and storage keys, the application test-judge domain and its build dispatch after a judge-config save, the test-judge queue, workflows and WORKER_MODE=test, the WASM-OJ runtime layers and toolchains in the worker image, the worker-test chart Deployment, web's TEST_JUDGE_ENABLED, and @wasm-oj/server with its patch. Core drops the request, response and record schemas, the judge-program cache key and server identity, and the artifact wire format. The edit page no longer shows a judge-program build status or offers "check samples with the checker". Renovate updates @wasm-oj/* and the rust image again. Core keeps what the browser will use: the checker and interactive verdict mapping, truncateUtf8, the WASM-OJ termination verdict, the DOMjudge Python wrappers, judgeProgramCompileInput and interactiveContestantSupported. The Test capability is no longer a server field; the editor disables Test on special_env problems itself. Until browser judging lands, checker and interactive Test report that Test isn't available for the problem. The contest participation and window checks on code drafts keep their coverage on listCodeDrafts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(problems): read a problem's checker or interactor source GET /api/problems/[id]/judge-program?context=… returns the role, language, source and SHA-256 of a checker or interactive problem's judge program, read through its verified storage pointer, so browser Test can compile and run it. Standard and special_env problems, and a problem without a stored program, answer 404. Access is the problem page's view access for the context: a context grants its problem while the page for it would render (course staff and contest organizers always; students in an open assignment, a running contest they joined, their own virtual run, or a running exam session that passes the proctoring gate), and otherwise the practice rules apply, so the program stays readable after an exam or contest ends for the students who can still view the problem. A page-locked exam session keeps the request inside that exam's problems. The route uses the standard API limiter and is exam-scoped; /api/drafts shares its context query parser. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: format the browser Test plan and spec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: add the Test stopgap to the dead-code checklist Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(problems): match the judge-program window to its pages A contest grants its problem's judge program only while the contest is published, organizers included, because the contest workspace never renders a draft contest. A running contest stays open through its end instant, as the contest problem page redirects only after endsAt. The JSON context query parser is now parseJsonContextParam, so it no longer reads like parseContextQuery. Integration cases cover the denials: a contest before its start, a draft contest for its participant and organizer, an assignment before it opens, in an archived course or as a draft, an exam session on a non-whitelisted IP, a problem outside a non-locked exam, and an ended virtual run. A closed assignment stays readable to its enrolled student through the ended-assignment view rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): prepare a checker in the browser when the editor opens On a checker problem the editor fetches the checker's source from /api/problems/[id]/judge-program for its context and builds it with the shared browser engine through judgeProgramCompileInput: a Python checker is packaged with the DOMjudge wrapper, a C++ checker compiles with the libc++ PCH shim. Its toolchain preloads next to the editor language's. Each editor mount refetches the source; a build is kept for the page session per problem and SHA-256, so it reruns only when the source changes. Source requests retry like the toolchain preload, except for a refused request. While the checker prepares, Test stays clickable and shows the toolchain download or "Preparing checker...". A failed build disables Test with "This problem's checker failed to build." and the Test Result panel shows the compiler output; a checker that cannot be loaded disables Test with a reload hint. Interactive problems keep the unavailable notice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): judge checker samples in the browser Test on a checker problem waits for the prepared checker, compiles and runs the selected cases as before, then runs the checker in the browser on every case that exited normally and matches a sample by input: /judge/input and /judge/answer hold the sample's input and output, the student's output is its stdin, and its team message is read from /judge/feedback/teammessage.txt. The checker gets official judging's validator time, 512 MiB and the same 2x wall deadline. Exit 42 is AC, 43 is WA and anything else is SE, through checkerCaseVerdict. Custom cases still only execute. The result panel loses the "Judged on the server" badge and the server notice; it explains a system error only on a case the browser checker judged. Interactive problems keep the unavailable notice until browser interaction lands. browser-local-run no longer exports its internal run and compile outcome types. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(problems): drop the hidden workspace file visibility Hidden workspace files never stayed secret: official judging and student code could read them, and browser Test would ship them. Production had none. A migration turns any hidden file into a readonly one and recreates the WorkspaceFileVisibility enum without hidden. Stored judge snapshots read a legacy hidden file as readonly, so old submissions still rejudge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): queue browser engine work and keep judge-program failures visible The browser engine allows one foreground compile at a time, but a checker build kept running after the student left the problem, so the next problem's checker failed to load or its Test reported a system error. Every engine compile, run and interaction now goes through one module-level queue. A judge-program build that nobody waits for any more keeps running, because the per-digest memo wants its result. A caller that aborts while waiting leaves the queue; the engine is cancelled only when the aborting caller's own operation is running. - The source response is parsed with judgeProgramSourceViewSchema from core, which the application's source view now shares. - Fetch, toolchain and engine failures are logged with console.warn. - A 403 or 404 source response disables Test without the reload hint. - The build projectId and the memo key include the judge program's language. - The checker gets the same output and filesystem limits as student runs. - The Test reason is a live region, and a failed build switches the panel to Test Result so its diagnostics are visible. The toolchain-preload and judge-program unit tests imported the module graph inside beforeEach after vi.resetModules, so a cold import under a loaded ci:verify ran into the 10 s hook timeout. They now import once at the top and isolate state with a distinct language or problem per test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): judge interactive samples and custom cases in the browser Test on an interactive problem now prepares the problem's interactor when the editor opens, like a checker, and shows "Preparing interactor...", a build failure with its diagnostics, or a load failure. Pressing Test compiles the contestant and runs every case through runBrowserInteraction, which calls engine.interact through the engine queue: - the interactor gets /judge/input /judge/answer /judge/feedback, the case's input as /judge/input, an empty answer and the feedback directory; - the contestant gets the language-factored time limit and the problem's memory; the interactor gets the validator timeout and the problem's memory plus official judging's 64 MiB headroom (capped at 1536 MiB); both share a wall stop of max(3 s, 3 x the factored limit) and student runs' output and filesystem limits; - the verdict comes from interactiveCaseVerdict, except that a contestant stopped at a time or memory limit stays TLE or MLE: a CPython interactor that then writes to the closed pipe exits with 120, which would otherwise turn a student's infinite loop into a judge error depending on timing; - each case keeps both transcript directions (64 KiB each), the contestant's stderr and its time, and is marked judged so a judge error is explained. The case panel edits interactor inputs on interactive problems, so students can add their own cases. Every judge type now allows custom cases, so the customCasesAllowed switch is gone, and so is the "no interactor samples" state. The interactive stopgap (client_test_judge_program, TEST_DISABLING_CODES, the controller's disabled reason and editor_testUnavailableForProblem) is removed. JavaScript and TypeScript contestants stay disabled. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): merge interactive Test verdicts like official judging Official interactive judging (resolveInteractiveStage) checks the interactor first: an interactor that exits with anything but 42 or 43, or is stopped, is SE even when the contestant hit a limit; only then does the contestant's TLE, MLE or RE win over the interactor's AC or WA. Core's interactiveCaseVerdict already follows that order, so browser Test now uses it as is and drops its own rule that kept a contestant's TLE or MLE over a failed interactor. Official judging never breaks the interactor's pipe: its channel keeps reading the interactor's output after the contestant ends, so the wrapper's read() sees EOF and exits 43, and the case stays TLE. The browser engine closes the pipe when the contestant stops, so a Python interactor that writes after that exits 120 and the case shows SE in Test, as does an interaction where both sides reach the shared wall stop. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(decisions): Test runs entirely in the browser, judge programs included JDG-15 now records that Test runs the problem's checker or interactor in the student's browser, that students may read judge programs and authors own their robustness, that the server compiles and executes nothing for Test, and that Test never receives non-sample testcase data. It rejects server-side judge programs (#641), server compilation, execution-only Test and exposure toggles, keeps the earlier rejections that still hold, and drops the rejections of shipping judge programs and compiled judge Wasm to the browser. Generic runtime fixes land in wasm-oj/forge and are consumed as pinned releases. JDG-26, SEC-15 and OPS-21 are withdrawn and point to JDG-15. JDG-05 is scoped to official judging. JDG-03 keeps the byte-identical wrapper rule for the browser copy. JDG-12, PRB-22 and OPS-18 return to their text before #641; PRB-03 and WEB-05 describe browser Test and the judge-program endpoint's view access. PRB-01 and PRB-09 drop the hidden visibility, and PRB-01 rejects it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: describe browser-run checker and interactive Test The living docs drop the server test judge: the test-judge queue, worker, workflows, route, limiter, Redis lock, storage keys, chart values, WASM-OJ image layers and their upgrade steps, the local test-judge setup and its runbook section, and the Test failure mode in Reliability. The Judge Pipeline's Browser Test section now covers the judge-program endpoint and its view access, preparation and preload when the editor opens, the checker's args, files and limits, interactive limits and custom interactor inputs, the engine queue, verdict mapping, the transcript cap, and where Test differs from official judging after a broken pipe. Architecture gets a browser Test flow; Frontend, Security, the Threat Model and Product Sense describe readable judge programs as an accepted risk owned by authors, with hidden testcases never leaving the server. Workspace visibility is editable or readonly everywhere. The Problem Test spec is rewritten as Given/When/Then for checker and interactive Test, custom cases, build failures, JavaScript and TypeScript on interactive problems and special_env. The Quality Ledger drops the server-Test items, lists the forge release NOJV needs and the upstream gaps, and adds a Wasm fast path for official judging to evaluate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: delete the browser Test plan and spec The work they planned has shipped on this branch, and the decision log and living docs now hold the result. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): let a contestant limit stop decide interactive Test Official judging keeps draining the interactor's output after the contestant ends, so when the contestant hits a time or memory limit the interactor reads EOF, exits 43, and the case is TLE or MLE. The browser engine closes a stopped side's pipes instead, so a Python interactor that writes after that exits 120 and core's interactiveCaseVerdict gave SE where Submit gives TLE or MLE; a deadlock reaching the shared wall stop did the same. interactiveCaseVerdict now returns TLE or MLE first when the contestant was stopped by logical time, the instruction budget, the wall stop or the memory limit. Every other case keeps official judging's order: an interactor failure is SE, then a contestant RE, then the interactor's AC or WA. JDG-15 records the rule and why, the Judge Pipeline and Problem Test spec describe it, and the Quality Ledger item now covers only what is left: a contestant that exits before the interactor's next write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): tell authors students can read their checker or interactor Test runs the checker or interactor in the student's browser, so the judge tab now says "Students can read this program when they press Test." under the checker and the interactor language. The tab is only on the edit page, so only editors see it. The checker and interactor help no longer say the program only runs in an isolated container: it runs in the judge sandbox on Submit and in the student's browser on Test. The interactor help and the examples drop set_score and score.txt. Neither Python wrapper defines set_score and no judge path reads score.txt (JDG-03 removed partial credit), so the Python interactor example died with a NameError on every correct guess. A component test covers the note on checker and interactive problems and its absence on standard ones. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): describe judge_log and the validator timeout accurately Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): drop a queued judge-program build when its editor closes Each editor mount queued an uncancellable checker or interactor build, so a student who clicked through several C++-checker problems waited behind every one of them before their own Test ran. The editor now passes an abort signal that fires on unmount. compileBrowserJudgeProgram uses it only while waiting in the engine queue: a build that has not started leaves the queue and the page-session memo, so a later visit rebuilds it, and a build that has started keeps running so the memo still gets its result. The memo key now includes the role, because the Python wrapper differs between checker and interactor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: align Test docs, interactor help and verdict helper with the code - interactiveCaseVerdict drops its teamMessage parameter: no caller passed it, and interactive Test never shows an interactor's teammessage. - The interactor help says Test passes an empty judge_answer, so the interactor should read its secret from judge_input. - tests/tsconfig.json includes unit/worker/mailer-startup.test.ts again, so typecheck:tests covers it. - The judge pipeline doc says contest organisers read the judge program only while the contest is published, and that Test on checker and interactive problems waits while the judge program prepares. It and the problem-test spec say the result shows the run's longest logical time, not a time per case. - JDG-15 says exit codes map as in official judging but the interactive merge order differs, and that teammessage is shown for checkers only. - The Quality Ledger drops the pin-bump item, which must land before this merges, keeping the JavaScript and TypeScript clause, and drops the unrelated Wasm fast path item. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: link decision sources to #645 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * build(web): pin @wasm-oj/browser 0.2.4 0.2.4 ships wasm-oj/forge#93, #94, #95, #96 and #99: runtime-bundle interactors, in-module metering of interactive programs, the runtime-file export stall fix and interactive sides in nested Workers. Its core and contracts move to 0.2.4 with it. The toolchain packages stay at 0.2.0, which forge versions independently. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
interactrefuses any interactor written in Python, even though Python interactors run in the same Wasm sandbox, under the same limits, as Python contestants, whichinteractalready accepts. This PR removes that one refusal so Python interactors work. No sandbox or resource limit changes.Background: what the guard actually checks
forge has two kinds of build artifact:
wasmruntime-bundlemain.py)Both kinds execute inside the same Wasm runtime, under the same policy:
A runtime bundle is not less isolated than a standalone module. It is still Wasm; it just carries its interpreter with it.
ServerRunner.interactand the browser runner'sinteractArtifactsboth start with:So the check is about the artifact's packaging, not about safety. The property interaction really needs is that a program can read fd 0 as a stream while the other side is still writing. That is already checked per runtime driver, for both sides, in
prepareArtifactInteraction(src/runner/artifact.ts):The CPython driver is
streaming, which is why a Python contestant already works ininteract. The kind guard only blocks the same driver on the interactor side. Bundles whose driver cannot stream fd 0 are still rejected by the driver check.Why it matters
DOMjudge-style interactors are commonly written in Python. In NOJV, which uses
@wasm-oj/serverto run problem interactors on the server for in-browser Test, all four interactive problems use Python interactors. With the guard in place, Test is unavailable on every one of them; with it removed, they run end to end.Change
ServerRunner.interact(src/server/server-runner.ts) and in the browser runner Worker'sinteractArtifacts(src/runtime/runner.worker.ts). The driver check above stays the only gate.docs/library-contract.md: either side may be a standalone Wasm module or a runtime bundle; bundles that cannot stream fd 0 are still rejected.CHANGELOG.md: Unreleased entry.Verification
src/server/judge.integration.test.ts):42\n/41\n.WASM_OJ_RUN_JUDGE_INTEGRATION=1 pnpm exec vitest run src/server/judge.integration.test.ts: 3/3 pass.pnpm run ci:verifypasses: 194 files, 956 tests, plus build.createServerEngine: a CPython guess-the-number interactor against a C++ binary-search contestant returns interactor exit 42 with the full transcript, in about 3 s.Not changed
Resource policies, metering, the trusted-judge path, and the streaming capability check are untouched. A Python interactor gets exactly the limits a Python contestant gets.
🤖 Generated with Claude Code