Repository navigation
feat(crewai): add Stagehand code-mode MCP example #2628
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: shrey/stg-2765-codemode-mastra
Are you sure you want to change the base?
Changes from 12 commits
eb1cc57
6959607
b4013ab
82d6c0d
6f6668d
1353b83
1143cd5
40f8b45
93ab5b1
0d7cfd7
d3abc46
b9172e4
214526c
60e1c45
f1431f6
0d1ce95
e1b860e
7b468a8
73911f3
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,73 @@ | ||
| # CrewAI with Stagehand code mode | ||
|
|
||
| This example gives a CrewAI agent one browser tool, `code_execute`, without running generated | ||
| JavaScript in the Python agent process. A small trusted Python lease owner starts the | ||
| package-installed [Vercel Sandbox example](../vercel-sandbox) and receives its `{ url, token }` | ||
| connection. CrewAI then connects directly through authenticated Streamable HTTP: | ||
|
|
||
| ```text | ||
| CrewAI MCPServerAdapter -> authenticated HTTPS -> Vercel Sandbox -> Stagehand code-mode MCP | ||
| `-> generated JavaScript | ||
| ``` | ||
|
|
||
| The adapter discovers the canonical tool description and uses it as the agent's backstory. It does | ||
| not copy the executor, schema, or code-mode skill. | ||
|
|
||
| ## Setup | ||
|
|
||
| Create a Python 3.12 environment and install the pinned CrewAI example dependencies: | ||
|
|
||
| ```bash | ||
| python3.12 -m venv .venv | ||
| . .venv/bin/activate | ||
| python -m pip install -r packages/integrations/examples/crewai/requirements.txt | ||
| ``` | ||
|
|
||
| Build and pack the exact Stagehand packages under review: | ||
|
|
||
| ```bash | ||
| pnpm install | ||
| pnpm exec turbo run build --filter @browserbasehq/stagehand-codemode | ||
|
cubic-dev-ai[bot] marked this conversation as resolved.
Outdated
|
||
| pnpm --filter @browserbasehq/stagehand-integrations-example-vercel-sandbox pack:artifacts | ||
| ``` | ||
|
|
||
| ## Run the end-to-end proof | ||
|
|
||
| ```bash | ||
| STAGEHAND_SANDBOX_ARTIFACTS="$PWD/packages/integrations/examples/vercel-sandbox/.artifacts" \ | ||
| BROWSERBASE_API_KEY=<api-key> \ | ||
| BROWSERBASE_PROJECT_ID=<project-id> \ | ||
| VERCEL_OIDC_TOKEN=<oidc-token> \ | ||
| OPENAI_API_KEY=<openai-key> \ | ||
| python packages/integrations/examples/crewai/e2e.py | ||
| ``` | ||
|
|
||
| For external CI, replace `VERCEL_OIDC_TOKEN` with `VERCEL_TEAM_ID`, `VERCEL_PROJECT_ID`, and | ||
| `VERCEL_TOKEN`. Set `CREWAI_MODEL` in your own wrapper if you want to pass a model other than the | ||
| example's `openai/gpt-5-mini` default. | ||
|
|
||
| The proof uses one live package-installed sandbox and one context-managed CrewAI MCP adapter. It: | ||
|
|
||
| 1. invokes `code_execute` twice directly and requires the same page ID and DOM marker; | ||
| 2. records CrewAI's tool-usage event from a real model-selected `code_execute` call; | ||
| 3. invokes the tool again to independently verify the model's browser-side change; | ||
| 4. proves the model key and a host-only marker are absent inside generated code; and | ||
| 5. closes the CrewAI MCP adapter before ending the sandbox lease, emitting `PASS` only afterward. | ||
|
|
||
| ## Use the agent | ||
|
|
||
| ```python | ||
| from agent import run_stagehand_agent | ||
|
cubic-dev-ai[bot] marked this conversation as resolved.
|
||
| from sandbox import StagehandSandboxLease | ||
|
|
||
| with StagehandSandboxLease() as connection: | ||
| result = run_stagehand_agent( | ||
| connection, | ||
| "Open https://example.com and return its title and URL.", | ||
| ) | ||
| print(result) | ||
| ``` | ||
|
|
||
| The lease subprocess receives only Browserbase, Vercel, artifact, and runtime variables from its | ||
| allowlist. Outer model-provider credentials are intentionally excluded, and the sandbox foundation | ||
| brokers the Browserbase credential at its egress boundary. | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,80 @@ | ||
| from __future__ import annotations | ||
|
|
||
| import os | ||
| from builtins import BaseExceptionGroup | ||
| from collections.abc import Iterator, Sequence | ||
| from contextlib import contextmanager | ||
| from typing import Any, Protocol | ||
|
|
||
| from crewai import Agent | ||
| from crewai.tools import BaseTool | ||
| from crewai_tools import MCPServerAdapter | ||
|
|
||
| DEFAULT_STAGEHAND_LLM = os.environ.get("CREWAI_MODEL", "openai/gpt-5-mini") | ||
|
|
||
|
|
||
| class StagehandSandboxConnection(Protocol): | ||
| url: str | ||
| token: str | ||
|
|
||
|
|
||
| @contextmanager | ||
| def stagehand_code_tools( | ||
| connection: StagehandSandboxConnection, | ||
| ) -> Iterator[list[BaseTool]]: | ||
| """Keep one authenticated remote MCP client open for a complete CrewAI run.""" | ||
| adapter = MCPServerAdapter( | ||
| { | ||
| "url": connection.url, | ||
| "transport": "streamable-http", | ||
| "headers": {"Authorization": f"Bearer {connection.token}"}, | ||
| }, | ||
| connect_timeout=60, | ||
| ) | ||
| try: | ||
| tools = list(adapter.tools) | ||
| names = [tool.name for tool in tools] | ||
| if names != ["code_execute"]: | ||
| raise RuntimeError( | ||
|
cubic-dev-ai[bot] marked this conversation as resolved.
Outdated
|
||
| f"Expected only code_execute from Stagehand MCP, got {names!r}." | ||
| ) | ||
| if "# Stagehand V4 code-mode syntax" not in tools[0].description: | ||
| raise RuntimeError("code_execute did not include the canonical guidance") | ||
| yield tools | ||
| except BaseException as primary_error: | ||
| try: | ||
| adapter.stop() | ||
| except BaseException as cleanup_error: | ||
| raise BaseExceptionGroup( | ||
|
cubic-dev-ai[bot] marked this conversation as resolved.
|
||
| "CrewAI run and MCP cleanup both failed", | ||
| [primary_error, cleanup_error], | ||
| ) | ||
| raise | ||
| else: | ||
| adapter.stop() | ||
|
|
||
|
|
||
| def build_stagehand_agent( | ||
| tools: Sequence[BaseTool], | ||
| llm: str | Any = DEFAULT_STAGEHAND_LLM, | ||
| ) -> Agent: | ||
| if len(tools) != 1 or tools[0].name != "code_execute": | ||
| raise ValueError("CrewAI Stagehand agent requires exactly code_execute") | ||
| return Agent( | ||
| role="Stagehand browser agent", | ||
| goal="Complete browser tasks by writing compact, correct Stagehand V4 JavaScript.", | ||
| backstory=tools[0].description, | ||
| llm=llm, | ||
| tools=list(tools), | ||
| max_iter=8, | ||
| verbose=False, | ||
| ) | ||
|
|
||
|
|
||
| def run_stagehand_agent( | ||
| connection: StagehandSandboxConnection, | ||
| prompt: str, | ||
| llm: str | Any = DEFAULT_STAGEHAND_LLM, | ||
| ) -> str: | ||
| with stagehand_code_tools(connection) as tools: | ||
| return str(build_stagehand_agent(tools, llm).kickoff(prompt)) | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,135 @@ | ||
| from __future__ import annotations | ||
|
|
||
| import json | ||
| import os | ||
| from typing import Any | ||
| from uuid import uuid4 | ||
|
|
||
| from crewai.events import ToolUsageFinishedEvent, crewai_event_bus | ||
|
|
||
| from agent import build_stagehand_agent, stagehand_code_tools | ||
| from sandbox import StagehandSandboxLease | ||
|
|
||
|
|
||
| def successful_result(raw_result: Any) -> dict[str, Any]: | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. P3: Malformed MCP-result handling is only covered by the secret-backed live smoke, making parser/contract regressions hard to diagnose or exercise locally. Add focused tests for accepted results and invalid JSON, failed, and non-object value payloads. (Based on your team's feedback about adding unit tests for new behavior.) Prompt for AI agents
Collaborator
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Addressed in |
||
| result = json.loads(str(raw_result)) | ||
| assert result["ok"] is True, result | ||
| value = result.get("value") | ||
| assert isinstance(value, dict), result | ||
| return value | ||
|
|
||
|
|
||
| def main() -> None: | ||
| direct_marker = f"crewai-direct-{uuid4()}" | ||
| model_marker = f"crewai-model-{uuid4()}" | ||
| os.environ["CREWAI_HOST_ONLY_MARKER"] = f"host-{uuid4()}" | ||
| model_tool_calls: list[str] = [] | ||
| final_state: dict[str, Any] | ||
|
|
||
| with StagehandSandboxLease() as connection: | ||
| with stagehand_code_tools(connection) as tools: | ||
| code_execute = tools[0] | ||
| first = successful_result( | ||
| code_execute.run( | ||
| code=f""" | ||
| await page.goto("https://example.com", {{ waitUntil: "domcontentloaded" }}); | ||
| await page.evaluate((marker) => {{ | ||
| document.documentElement.dataset.crewaiDirectMarker = marker; | ||
| }}, {json.dumps(direct_marker)}); | ||
| return {{ | ||
| pageId: page.pageId, | ||
| title: await page.title(), | ||
| directMarker: await page.evaluate( | ||
| () => document.documentElement.dataset.crewaiDirectMarker, | ||
| ), | ||
| modelKeyVisible: process.env.OPENAI_API_KEY ?? null, | ||
| hostMarkerVisible: process.env.CREWAI_HOST_ONLY_MARKER ?? null, | ||
| }}; | ||
| """ | ||
| ) | ||
| ) | ||
| assert first["title"] == "Example Domain" | ||
| assert first["directMarker"] == direct_marker | ||
| assert first["modelKeyVisible"] is None | ||
| assert first["hostMarkerVisible"] is None | ||
|
|
||
| second = successful_result( | ||
| code_execute.run( | ||
| code=""" | ||
| return { | ||
| pageId: page.pageId, | ||
| title: await page.title(), | ||
| directMarker: await page.evaluate( | ||
| () => document.documentElement.dataset.crewaiDirectMarker, | ||
| ), | ||
| }; | ||
| """ | ||
| ) | ||
| ) | ||
| assert second["title"] == "Example Domain" | ||
| assert second["pageId"] == first["pageId"] | ||
| assert second["directMarker"] == direct_marker | ||
|
|
||
| @crewai_event_bus.on(ToolUsageFinishedEvent) | ||
| def record_model_tool(_source: Any, event: ToolUsageFinishedEvent) -> None: | ||
| if event.tool_name == "code_execute": | ||
| model_tool_calls.append(event.tool_name) | ||
|
|
||
| try: | ||
| agent = build_stagehand_agent(tools) | ||
| agent.kickoff( | ||
| " ".join( | ||
| ( | ||
| "Use code_execute to modify the already-open page.", | ||
| "Set document.documentElement.dataset.crewaiModelMarker to", | ||
| f"{json.dumps(model_marker)}.", | ||
| "Then read that dataset value and the current pageId and report them.", | ||
| "You must call code_execute; do not merely describe JavaScript.", | ||
| ) | ||
| ) | ||
| ) | ||
| assert crewai_event_bus.flush(), "CrewAI tool events did not finish" | ||
| finally: | ||
| crewai_event_bus.off(ToolUsageFinishedEvent, record_model_tool) | ||
| assert model_tool_calls, "the real CrewAI model must select code_execute" | ||
|
|
||
| final_state = successful_result( | ||
| code_execute.run( | ||
| code=""" | ||
| return { | ||
| pageId: page.pageId, | ||
| title: await page.title(), | ||
| directMarker: await page.evaluate( | ||
| () => document.documentElement.dataset.crewaiDirectMarker, | ||
| ), | ||
| modelMarker: await page.evaluate( | ||
| () => document.documentElement.dataset.crewaiModelMarker, | ||
| ), | ||
| }; | ||
| """ | ||
| ) | ||
| ) | ||
| assert final_state["title"] == "Example Domain" | ||
| assert final_state["pageId"] == first["pageId"] | ||
| assert final_state["directMarker"] == direct_marker | ||
| assert final_state["modelMarker"] == model_marker | ||
|
|
||
| print( | ||
| json.dumps( | ||
| { | ||
| "status": "PASS", | ||
| "framework": "crewai", | ||
| "directToolCalls": 3, | ||
| "modelToolCalls": len(model_tool_calls), | ||
| "sessionPersisted": True, | ||
| "modelCredentialIsolated": True, | ||
| "finalState": final_state, | ||
| "cleanup": ["crewai-mcp", "vercel-sandbox"], | ||
| }, | ||
| sort_keys=True, | ||
| ) | ||
| ) | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| main() | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,2 @@ | ||
| crewai==1.15.13 | ||
| crewai-tools[mcp]==1.15.13 |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,21 @@ | ||
| """Expose the shared Vercel Sandbox lease from the CrewAI example directory.""" | ||
|
|
||
| from __future__ import annotations | ||
|
|
||
| import sys | ||
| from pathlib import Path | ||
|
|
||
| SHARED_EXAMPLES_DIRECTORY = Path(__file__).resolve().parents[1] / "shared" | ||
| sys.path.insert(0, str(SHARED_EXAMPLES_DIRECTORY)) | ||
|
|
||
| from vercel_sandbox_lease import ( # noqa: E402 | ||
| StagehandSandboxConnection, | ||
| StagehandSandboxLease, | ||
| StagehandSandboxLeaseError, | ||
| ) | ||
|
|
||
| __all__ = [ | ||
| "StagehandSandboxConnection", | ||
| "StagehandSandboxLease", | ||
| "StagehandSandboxLeaseError", | ||
| ] |
Uh oh!
There was an error while loading. Please reload this page.