Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

Stagehand Evals

Agent benchmarks for Stagehand — act, extract, observe, agent, plus dataset-backed suites (WebVoyager, OnlineMind2Web, WebTailBench, Odysseys, HardBench).

Driven by an interactive TUI (evals) or single-shot CLI (evals run …). Tasks are auto-discovered from tasks/bench/<category>/ — no registration step.

Quickstart

From the stagehand repo root:

pnpm install
pnpm build:cli   # also: pnpm build, if you haven't built the workspace yet

This links an evals binary on your PATH. Launch the REPL:

evals

REPL with help output

Or run a single target:

evals run extract -t 3 -c 5
evals run b:webvoyager -l 10

A .env in packages/evals/ is loaded automatically. Provide whichever provider keys (OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_GENERATIVE_AI_API_KEY, …) and BROWSERBASE_API_KEY / BROWSERBASE_PROJECT_ID you need.

TUI commands

Inside the REPL (or as evals <command> from your shell):

Command What it does
run [target] [options] Run evals. Target can be a tier, category, task, or benchmark shorthand.
list [tier] [--detailed] List discovered tasks and categories.
new <tier> <category> <name> Scaffold a new task file.
config [set|reset|path] Read or write defaults (env, trials, concurrency, model, …).
experiments Inspect and compare Braintrust experiment runs.
help Show command help. Append --help to any command for details.

Use Esc to abort an in-flight run without exiting the REPL.

Onboarding

evals welcome runs a guided first-run flow built on the agent benchmarks: an animated intro (the Stagehand mark, what evals measures, EVALS), then a deterministic replay of a real WebVoyager task — three models in lanes, named generically (Opus, Grok, Sol — the fastest and cheapest fails on a wrong condition filter; timings and costs are illustrative), a podium by accuracy · speed · cost, and a chat-style look inside the winning run. It ends on a real run b:webvoyager -l 3 --harness claude_code -e local (--harness codex when the key is OpenAI's; -e browserbase when that's the available browser) when an Anthropic or OpenAI key and a browser exist, or hands off to evals setup, a guided flow that asks only for what's missing (Anthropic/OpenAI key, browser), writes packages/evals/.env, and offers the first real run.

Set EVALS_WELCOME_WIZARD=1 to auto-run the flow on the first REPL launch; EVALS_NO_WELCOME=1 suppresses the first-run welcome. Any key advances the intro, Esc skips ahead, Ctrl+C cancels.

Run targets

evals run accepts any of these shapes:

Target Meaning
(none) / all All bench tasks
bench Entire bench tier
act / extract / observe / agent A category
extract/extract_text A specific task
b:webvoyager / b:onlineMind2Web / b:webtailbench Dataset-backed benchmark suite

evals list shows everything that's been discovered:

evals list output

Common options

Flag Purpose
-e, --env <local|browserbase> Where the browser runs
-t, --trials <n> Trials per task
-c, --concurrency <n> Max parallel sessions
-m, --model <id> Override the model matrix
--api Run via the Stagehand API instead of the SDK
--harness <stagehand|claude_code|codex|mastra|pi> Which agent harness drives the bench task
-l, --limit <n> / -s, --sample <n> / -f, --filter key=value Suite shaping for benchmark targets
--preview Print the resolved plan and exit — no browser, no LLM calls

Defaults live in evals.config.json and can be edited via evals config set ….

--preview is useful for sanity-checking the plan before paying for a run:

evals run --preview output

A live run paints an in-place progress table, then prints a final summary with a per-model breakdown:

Live bench run

Shared harness behavior

See the harness contract for tool surfaces, prompt policy, budget units, session diagnostics, verification, and usage accounting. See HardBench for corpus selection and rubric v1.2.

Adding a bench task

evals new bench extract my_new_task

This drops a defineBenchTask-based file into tasks/bench/extract/. It will show up in evals list on next launch — no config edit needed.

// tasks/bench/extract/my_new_task.ts
import { defineBenchTask } from "../../../framework/defineTask.js";

export default defineBenchTask({
  name: "my_new_task",
  tags: ["regression"],
  run: async ({ stagehand, logger }) => {
    // ... drive stagehand, return { _success: boolean, ... }
  },
});

Tracing / Observability

Runs stream into Braintrust when BRAINTRUST_API_KEY is set; otherwise a local summary prints to stdout. Use evals experiments to inspect and diff past Braintrust runs.