Skip to content

feat(test): run checker and interactive Test entirely in the browser - #645

Merged
TakalaWang merged 23 commits into
mainfrom
feat/browser-test-judge
Oct 10, 2026
Merged

TakalaWang merged 23 commits into
mainfrom
feat/browser-test-judge

Conversation

@TakalaWang

@TakalaWang TakalaWang commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Summary

Test on checker and interactive problems now runs entirely in the student's browser, the problem's checker or interactor included. The server half of #641 (merged, never released) is removed: the server compiles and executes nothing for Test and only serves the judge program's source. The hidden workspace file visibility is removed too.

Why

#641 ran TA-authored checkers and interactors, and on interactive problems the student's Wasm, in a server test worker. The only thing between that code and the container's object-storage keys (every problem's hidden testcases) and Redis URL was the WASM-OJ runtime. Test is the student's own run, so it belongs on the student's machine. Judge programs become readable by students; that is an accepted risk, and authors own judge programs that hold up when read (JDG-15).

What changed

Removed

  • The WORKER_MODE=test worker, the worker-test Deployment, the test-judge queue and its partitions, the Test and judge-program build workflows and the build dispatch after a judge-config save.
  • POST /api/problems/[id]/test-judge, its limiter and Redis lock, the test-judge-requests/ and test-judge-programs/ storage keys, and TEST_JUDGE_ENABLED, TEST_JUDGE_SLOTS and WASM_OJ_*.
  • @wasm-oj/server and its patch, the worker image's WASM-OJ layers, infra/docker/wasm-oj-toolchains/ and the Renovate exclusions.
  • Core's request, response and record schemas, the cache key and server identity, and the artifact wire format.
  • The edit page's judge-program build status and "check samples with the checker" action, and the "Judged on the server" badge and busy notice.
  • hidden workspace visibility: a migration turns any hidden file into a readonly one (production had none on 2026-10-07), and stored judge snapshots read a legacy hidden as readonly.

Added

  • GET /api/problems/[id]/judge-program?context=… returns { role, language, source, sha256 }. Access is the problem page's view access for the context, and a page-locked exam session is confined to that exam. The route is exam-scoped and uses the standard API limiter.
  • The editor fetches and builds the judge program when it opens, preloads its toolchain, and keeps the build per role, language and sha256 for the page session. A build still queued when the editor closes leaves the queue. A build failure disables Test and shows the diagnostics.
  • Checker Test: each sample case that exits normally is checked in the browser with the sample's output as the answer, and teammessage is shown. Custom cases stay execution-only.
  • Interactive Test: interact runs every sample and custom case, with the case's input as the interactor's input, and shows the transcript (64 KiB per direction) and contestant stderr. JavaScript and TypeScript contestants stay disabled.
  • One FIFO queue for every browser engine operation.
  • Verdicts come from core's checkerCaseVerdict and interactiveCaseVerdict only, and give the outcome Submit gives, except the broken-pipe case below. Interactive rule: a contestant stopped by a time limit (logical time, instruction budget or the wall stop) or the memory limit is TLE or MLE whatever the interactor does; otherwise an interactor failure is SE; otherwise a contestant RE wins; otherwise the interactor's AC or WA stands. The limit check comes first because the browser engine closes a stopped side's pipes, while official judging keeps draining the interactor's output and gives it EOF, so Submit gives TLE or MLE there.
  • The judge tab tells editors "Students can read this program when they press Test." under the checker and interactor language. The checker and interactor help now say the program runs in the judge sandbox on Submit and in the student's browser on Test, and the interactor help says Test passes an empty judge_answer, so the secret belongs in judge_input. The interactor help and the judge examples drop set_score and score.txt: neither Python wrapper defines set_score and no judge path reads score.txt (JDG-03), so the Python interactor example died with a NameError on every correct guess.

Kept from #641: interactorInput and interactionFormat, the seed fixes, "Executed" for cases without expected output, the contest code-draft check (WEB-05), the PCH-only-with-<bits/stdc++.h> rule, and core's verdict helpers, Python wrappers, C++ shim and judgeProgramCompileInput.

Docs: JDG-15 is rewritten. JDG-26, SEC-15 and OPS-21 are withdrawn. JDG-03, JDG-05, JDG-12, PRB-01, PRB-03, PRB-09, PRB-22, WEB-05 and OPS-18 are updated. The living docs, the threat model and the runbooks are updated, and docs/features/problem-test.md is rewritten. New Quality Ledger items, all upstream: the broken-pipe verdict gap, duplicated transcript lines after a broken pipe, and WebKit's instruction meter not stopping an empty loop. The spec and plan are deleted.

Pre-merge checklist

Known difference from official judging: broken pipe

When the contestant exits, normally or with an error, before the interactor's next write, that write fails in the browser (a Python interactor exits 120) and Test shows SE where Submit gives RE or the interactor's verdict. It is a Quality Ledger item for upstream.

After deploy (manual admin edits)

Releases do not update seed rows. In the editor:

  1. any-two-sum, course-order, shortest-route-plan: upload the updated checker from packages/db/prisma/seeds/problems.ts.
  2. guess-the-number:
    • fix the literal \n\n in the body;
    • add interaction notes;
    • sample 1: first transcript line 1 1000000, new explanation, interactor input 42;
    • sample 2: interactor input 500000.
  3. multi-interactive-bisect: interaction notes; interactor input 375000.
  4. noisy-oracle-hunt:
    • interaction notes;
    • sample 1: interactor input 500000;
    • sample 2: replace with the seed's new transcript, interactor input 671875.
  5. interactive-peak: interaction notes; interactor input 5\n1 3 2 5 4\n.

Nothing needs a backfill: nothing is precompiled.

Testing

  • pnpm ci:verify: green. It covers format, lint:repo, build, typecheck, lint and typecheck:tests; unit is 431 files (4,143 passed, 2 skipped) and component is 59 files (196 passed).
  • pnpm lint:helm and pnpm install --frozen-lockfile: green.
  • Integration (integration project) against isolated Postgres 18 and Redis 8 containers: 110 files, 789 passed, 6 skipped.
  • Dead-code checklist greps: nothing left outside history. Exceptions: migration SQL, the legacy-snapshot mapping and its tests, the withdrawn decision notes, and the kept core modules test-judge-verdict.ts and test-judge-program.ts.
  • Browser checks on seeded local stacks, with locally built forge (screenshots taken during development):
    • any-two-sum with Python and C++ checkers: valid and invalid answers;
    • shortest-route-plan: shortest and longer routes;
    • a broken C++ checker disabling Test and showing its diagnostics;
    • leaving a problem while its checker builds;
    • the four interactive problems with a C++ contestant: AC;
    • guess-the-number: WA, an infinite loop, a Python contestant, a custom interactor input, and a JavaScript contestant disabled.
  • Release checks on 0.2.4, seed student, crossOriginIsolated true: Chromium 153, WebKit 26.6 and Firefox 155 each pass the 4 interactive problems AC with transcript, WA, TLE in 4–6 s, a Python contestant, a custom interactor input, JavaScript disabled, both checkers AC and WA with teammessage, a standard problem, leaving a checker build, and engine crashes (SE).

🤖 Generated with Claude Code

TakalaWang and others added 23 commits October 7, 2026 16:22
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Test will run checkers and interactors in the student's browser, so the
server half of #641 goes: the test-judge API route, limiter, Redis lock
and storage keys, the application test-judge domain and its build
dispatch after a judge-config save, the test-judge queue, workflows and
WORKER_MODE=test, the WASM-OJ runtime layers and toolchains in the
worker image, the worker-test chart Deployment, web's TEST_JUDGE_ENABLED,
and @wasm-oj/server with its patch. Core drops the request, response and
record schemas, the judge-program cache key and server identity, and the
artifact wire format. The edit page no longer shows a judge-program build
status or offers "check samples with the checker". Renovate updates
@wasm-oj/* and the rust image again.

Core keeps what the browser will use: the checker and interactive verdict
mapping, truncateUtf8, the WASM-OJ termination verdict, the DOMjudge
Python wrappers, judgeProgramCompileInput and
interactiveContestantSupported. The Test capability is no longer a
server field; the editor disables Test on special_env problems itself.

Until browser judging lands, checker and interactive Test report that
Test isn't available for the problem. The contest participation and
window checks on code drafts keep their coverage on listCodeDrafts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GET /api/problems/[id]/judge-program?context=… returns the role,
language, source and SHA-256 of a checker or interactive problem's judge
program, read through its verified storage pointer, so browser Test can
compile and run it. Standard and special_env problems, and a problem
without a stored program, answer 404.

Access is the problem page's view access for the context: a context
grants its problem while the page for it would render (course staff and
contest organizers always; students in an open assignment, a running
contest they joined, their own virtual run, or a running exam session
that passes the proctoring gate), and otherwise the practice rules
apply, so the program stays readable after an exam or contest ends for
the students who can still view the problem. A page-locked exam session
keeps the request inside that exam's problems. The route uses the
standard API limiter and is exam-scoped; /api/drafts shares its context
query parser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A contest grants its problem's judge program only while the contest is
published, organizers included, because the contest workspace never
renders a draft contest. A running contest stays open through its end
instant, as the contest problem page redirects only after endsAt.

The JSON context query parser is now parseJsonContextParam, so it no
longer reads like parseContextQuery.

Integration cases cover the denials: a contest before its start, a draft
contest for its participant and organizer, an assignment before it
opens, in an archived course or as a draft, an exam session on a
non-whitelisted IP, a problem outside a non-locked exam, and an ended
virtual run. A closed assignment stays readable to its enrolled student
through the ended-assignment view rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a checker problem the editor fetches the checker's source from
/api/problems/[id]/judge-program for its context and builds it with the
shared browser engine through judgeProgramCompileInput: a Python checker
is packaged with the DOMjudge wrapper, a C++ checker compiles with the
libc++ PCH shim. Its toolchain preloads next to the editor language's.
Each editor mount refetches the source; a build is kept for the page
session per problem and SHA-256, so it reruns only when the source
changes. Source requests retry like the toolchain preload, except for a
refused request.

While the checker prepares, Test stays clickable and shows the
toolchain download or "Preparing checker...". A failed build disables
Test with "This problem's checker failed to build." and the Test Result
panel shows the compiler output; a checker that cannot be loaded
disables Test with a reload hint. Interactive problems keep the
unavailable notice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Test on a checker problem waits for the prepared checker, compiles and
runs the selected cases as before, then runs the checker in the browser
on every case that exited normally and matches a sample by input:
/judge/input and /judge/answer hold the sample's input and output, the
student's output is its stdin, and its team message is read from
/judge/feedback/teammessage.txt. The checker gets official judging's
validator time, 512 MiB and the same 2x wall deadline. Exit 42 is AC,
43 is WA and anything else is SE, through checkerCaseVerdict. Custom
cases still only execute.

The result panel loses the "Judged on the server" badge and the server
notice; it explains a system error only on a case the browser checker
judged. Interactive problems keep the unavailable notice until browser
interaction lands.

browser-local-run no longer exports its internal run and compile
outcome types.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hidden workspace files never stayed secret: official judging and student
code could read them, and browser Test would ship them. Production had none.
A migration turns any hidden file into a readonly one and recreates the
WorkspaceFileVisibility enum without hidden. Stored judge snapshots read a
legacy hidden file as readonly, so old submissions still rejudge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…isible

The browser engine allows one foreground compile at a time, but a
checker build kept running after the student left the problem, so the
next problem's checker failed to load or its Test reported a system
error. Every engine compile, run and interaction now goes through one
module-level queue. A judge-program build that nobody waits for any
more keeps running, because the per-digest memo wants its result. A
caller that aborts while waiting leaves the queue; the engine is
cancelled only when the aborting caller's own operation is running.

- The source response is parsed with judgeProgramSourceViewSchema from
  core, which the application's source view now shares.
- Fetch, toolchain and engine failures are logged with console.warn.
- A 403 or 404 source response disables Test without the reload hint.
- The build projectId and the memo key include the judge program's
  language.
- The checker gets the same output and filesystem limits as student
  runs.
- The Test reason is a live region, and a failed build switches the
  panel to Test Result so its diagnostics are visible.

The toolchain-preload and judge-program unit tests imported the module
graph inside beforeEach after vi.resetModules, so a cold import under a
loaded ci:verify ran into the 10 s hook timeout. They now import once at
the top and isolate state with a distinct language or problem per test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Test on an interactive problem now prepares the problem's interactor
when the editor opens, like a checker, and shows "Preparing
interactor...", a build failure with its diagnostics, or a load failure.
Pressing Test compiles the contestant and runs every case through
runBrowserInteraction, which calls engine.interact through the engine
queue:

- the interactor gets /judge/input /judge/answer /judge/feedback, the
  case's input as /judge/input, an empty answer and the feedback
  directory;
- the contestant gets the language-factored time limit and the
  problem's memory; the interactor gets the validator timeout and the
  problem's memory plus official judging's 64 MiB headroom (capped at
  1536 MiB); both share a wall stop of max(3 s, 3 x the factored limit)
  and student runs' output and filesystem limits;
- the verdict comes from interactiveCaseVerdict, except that a
  contestant stopped at a time or memory limit stays TLE or MLE: a
  CPython interactor that then writes to the closed pipe exits with
  120, which would otherwise turn a student's infinite loop into a
  judge error depending on timing;
- each case keeps both transcript directions (64 KiB each), the
  contestant's stderr and its time, and is marked judged so a judge
  error is explained.

The case panel edits interactor inputs on interactive problems, so
students can add their own cases. Every judge type now allows custom
cases, so the customCasesAllowed switch is gone, and so is the
"no interactor samples" state. The interactive stopgap
(client_test_judge_program, TEST_DISABLING_CODES, the controller's
disabled reason and editor_testUnavailableForProblem) is removed.
JavaScript and TypeScript contestants stay disabled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Official interactive judging (resolveInteractiveStage) checks the
interactor first: an interactor that exits with anything but 42 or 43,
or is stopped, is SE even when the contestant hit a limit; only then does
the contestant's TLE, MLE or RE win over the interactor's AC or WA.
Core's interactiveCaseVerdict already follows that order, so browser
Test now uses it as is and drops its own rule that kept a contestant's
TLE or MLE over a failed interactor.

Official judging never breaks the interactor's pipe: its channel keeps
reading the interactor's output after the contestant ends, so the
wrapper's read() sees EOF and exits 43, and the case stays TLE. The
browser engine closes the pipe when the contestant stops, so a Python
interactor that writes after that exits 120 and the case shows SE in
Test, as does an interaction where both sides reach the shared wall stop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cluded

JDG-15 now records that Test runs the problem's checker or interactor in
the student's browser, that students may read judge programs and authors
own their robustness, that the server compiles and executes nothing for
Test, and that Test never receives non-sample testcase data. It rejects
server-side judge programs (#641), server compilation, execution-only
Test and exposure toggles, keeps the earlier rejections that still hold,
and drops the rejections of shipping judge programs and compiled judge
Wasm to the browser. Generic runtime fixes land in wasm-oj/forge and are
consumed as pinned releases.

JDG-26, SEC-15 and OPS-21 are withdrawn and point to JDG-15. JDG-05 is
scoped to official judging. JDG-03 keeps the byte-identical wrapper rule
for the browser copy. JDG-12, PRB-22 and OPS-18 return to their text
before #641; PRB-03 and WEB-05 describe browser Test and the
judge-program endpoint's view access. PRB-01 and PRB-09 drop the hidden
visibility, and PRB-01 rejects it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The living docs drop the server test judge: the test-judge queue, worker,
workflows, route, limiter, Redis lock, storage keys, chart values,
WASM-OJ image layers and their upgrade steps, the local test-judge setup
and its runbook section, and the Test failure mode in Reliability.

The Judge Pipeline's Browser Test section now covers the judge-program
endpoint and its view access, preparation and preload when the editor
opens, the checker's args, files and limits, interactive limits and
custom interactor inputs, the engine queue, verdict mapping, the
transcript cap, and where Test differs from official judging after a
broken pipe. Architecture gets a browser Test flow; Frontend, Security,
the Threat Model and Product Sense describe readable judge programs as an
accepted risk owned by authors, with hidden testcases never leaving the
server. Workspace visibility is editable or readonly everywhere.

The Problem Test spec is rewritten as Given/When/Then for checker and
interactive Test, custom cases, build failures, JavaScript and
TypeScript on interactive problems and special_env. The Quality Ledger
drops the server-Test items, lists the forge release NOJV needs and the
upstream gaps, and adds a Wasm fast path for official judging to
evaluate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The work they planned has shipped on this branch, and the decision log
and living docs now hold the result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Official judging keeps draining the interactor's output after the
contestant ends, so when the contestant hits a time or memory limit the
interactor reads EOF, exits 43, and the case is TLE or MLE. The browser
engine closes a stopped side's pipes instead, so a Python interactor that
writes after that exits 120 and core's interactiveCaseVerdict gave SE
where Submit gives TLE or MLE; a deadlock reaching the shared wall stop
did the same.

interactiveCaseVerdict now returns TLE or MLE first when the contestant
was stopped by logical time, the instruction budget, the wall stop or the
memory limit. Every other case keeps official judging's order: an
interactor failure is SE, then a contestant RE, then the interactor's AC
or WA.

JDG-15 records the rule and why, the Judge Pipeline and Problem Test spec
describe it, and the Quality Ledger item now covers only what is left: a
contestant that exits before the interactor's next write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Test runs the checker or interactor in the student's browser, so the
judge tab now says "Students can read this program when they press Test."
under the checker and the interactor language. The tab is only on the
edit page, so only editors see it.

The checker and interactor help no longer say the program only runs in
an isolated container: it runs in the judge sandbox on Submit and in the
student's browser on Test. The interactor help and the examples drop
set_score and score.txt. Neither Python wrapper defines set_score and no
judge path reads score.txt (JDG-03 removed partial credit), so the Python
interactor example died with a NameError on every correct guess.

A component test covers the note on checker and interactive problems and
its absence on standard ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each editor mount queued an uncancellable checker or interactor build, so a
student who clicked through several C++-checker problems waited behind every
one of them before their own Test ran.

The editor now passes an abort signal that fires on unmount.
compileBrowserJudgeProgram uses it only while waiting in the engine queue:
a build that has not started leaves the queue and the page-session memo,
so a later visit rebuilds it, and a build that has started keeps running
so the memo still gets its result. The memo key now includes the role,
because the Python wrapper differs between checker and interactor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- interactiveCaseVerdict drops its teamMessage parameter: no caller passed
  it, and interactive Test never shows an interactor's teammessage.
- The interactor help says Test passes an empty judge_answer, so the
  interactor should read its secret from judge_input.
- tests/tsconfig.json includes unit/worker/mailer-startup.test.ts again,
  so typecheck:tests covers it.
- The judge pipeline doc says contest organisers read the judge program
  only while the contest is published, and that Test on checker and
  interactive problems waits while the judge program prepares. It and the
  problem-test spec say the result shows the run's longest logical time,
  not a time per case.
- JDG-15 says exit codes map as in official judging but the interactive
  merge order differs, and that teammessage is shown for checkers only.
- The Quality Ledger drops the pin-bump item, which must land before this
  merges, keeping the JavaScript and TypeScript clause, and drops the
  unrelated Wasm fast path item.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0.2.4 ships wasm-oj/forge#93, #94, #95, #96 and #99: runtime-bundle
interactors, in-module metering of interactive programs, the runtime-file
export stall fix and interactive sides in nested Workers. Its core and
contracts move to 0.2.4 with it. The toolchain packages stay at 0.2.0,
which forge versions independently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TakalaWang
TakalaWang marked this pull request as ready for review October 9, 2026 08:14
@TakalaWang
TakalaWang merged commit 1a2ecea into main Oct 10, 2026
16 checks passed
@TakalaWang
TakalaWang deleted the feat/browser-test-judge branch October 10, 2026 03:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant