Agent benchmarks for Stagehand — act, extract, observe, agent, plus dataset-backed suites (WebVoyager, OnlineMind2Web, WebTailBench, Odysseys, HardBench).
Driven by an interactive TUI (evals) or single-shot CLI (evals run …). Tasks are auto-discovered from tasks/bench/<category>/ — no registration step.
From the stagehand repo root:
pnpm install
pnpm build:cli # also: pnpm build, if you haven't built the workspace yetThis links an evals binary on your PATH. Launch the REPL:
evalsOr run a single target:
evals run extract -t 3 -c 5
evals run b:webvoyager -l 10A .env in packages/evals/ is loaded automatically. Provide whichever provider keys (OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_GENERATIVE_AI_API_KEY, …) and BROWSERBASE_API_KEY / BROWSERBASE_PROJECT_ID you need.
Inside the REPL (or as evals <command> from your shell):
| Command | What it does |
|---|---|
run [target] [options] |
Run evals. Target can be a tier, category, task, or benchmark shorthand. |
list [tier] [--detailed] |
List discovered tasks and categories. |
new <tier> <category> <name> |
Scaffold a new task file. |
config [set|reset|path] |
Read or write defaults (env, trials, concurrency, model, …). |
experiments |
Inspect and compare Braintrust experiment runs. |
help |
Show command help. Append --help to any command for details. |
Use Esc to abort an in-flight run without exiting the REPL.
evals welcome runs a guided first-run flow built on the agent benchmarks: an animated intro (the Stagehand mark, what evals measures, EVALS), then a deterministic replay of a real WebVoyager task — three models in lanes, named generically (Opus, Grok, Sol — the fastest and cheapest fails on a wrong condition filter; timings and costs are illustrative), a podium by accuracy · speed · cost, and a chat-style look inside the winning run. It ends on a real run b:webvoyager -l 3 --harness claude_code -e local (--harness codex when the key is OpenAI's; -e browserbase when that's the available browser) when an Anthropic or OpenAI key and a browser exist, or hands off to evals setup, a guided flow that asks only for what's missing (Anthropic/OpenAI key, browser), writes packages/evals/.env, and offers the first real run.
Set EVALS_WELCOME_WIZARD=1 to auto-run the flow on the first REPL launch; EVALS_NO_WELCOME=1 suppresses the first-run welcome. Any key advances the intro, Esc skips ahead, Ctrl+C cancels.
evals run accepts any of these shapes:
| Target | Meaning |
|---|---|
(none) / all |
All bench tasks |
bench |
Entire bench tier |
act / extract / observe / agent |
A category |
extract/extract_text |
A specific task |
b:webvoyager / b:onlineMind2Web / b:webtailbench |
Dataset-backed benchmark suite |
evals list shows everything that's been discovered:
| Flag | Purpose |
|---|---|
-e, --env <local|browserbase> |
Where the browser runs |
-t, --trials <n> |
Trials per task |
-c, --concurrency <n> |
Max parallel sessions |
-m, --model <id> |
Override the model matrix |
--api |
Run via the Stagehand API instead of the SDK |
--harness <stagehand|claude_code|codex|mastra|pi> |
Which agent harness drives the bench task |
-l, --limit <n> / -s, --sample <n> / -f, --filter key=value |
Suite shaping for benchmark targets |
--preview |
Print the resolved plan and exit — no browser, no LLM calls |
Defaults live in evals.config.json and can be edited via evals config set ….
--preview is useful for sanity-checking the plan before paying for a run:
A live run paints an in-place progress table, then prints a final summary with a per-model breakdown:
See the harness contract for tool surfaces, prompt policy, budget units, session diagnostics, verification, and usage accounting. See HardBench for corpus selection and rubric v1.2.
evals new bench extract my_new_taskThis drops a defineBenchTask-based file into tasks/bench/extract/. It will show up in evals list on next launch — no config edit needed.
// tasks/bench/extract/my_new_task.ts
import { defineBenchTask } from "../../../framework/defineTask.js";
export default defineBenchTask({
name: "my_new_task",
tags: ["regression"],
run: async ({ stagehand, logger }) => {
// ... drive stagehand, return { _success: boolean, ... }
},
});Runs stream into Braintrust when BRAINTRUST_API_KEY is set; otherwise a local summary prints to stdout. Use evals experiments to inspect and diff past Braintrust runs.



