Skip to content
Open
Show file tree
Hide file tree
Changes from 12 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
63 changes: 63 additions & 0 deletions .github/workflows/codemode-framework-examples.yml
Original file line number Diff line number Diff line change
Expand Up @@ -143,3 +143,66 @@ jobs:
VERCEL_PROJECT_ID: ${{ secrets.VERCEL_PROJECT_ID }}
VERCEL_TOKEN: ${{ secrets.VERCEL_TOKEN }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

crewai:
name: CrewAI
if: >-
github.event_name == 'push' ||
github.event.pull_request.head.repo.full_name == github.repository ||
contains(github.event.pull_request.labels.*.name, 'safe-to-test')
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6.1.0

- uses: ./.github/actions/setup-node-pnpm
with:
use-prebuilt-artifacts: "false"

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
with:
python-version: "3.12"
cache: pip
cache-dependency-path: packages/integrations/examples/crewai/requirements.txt

- run: python -m pip install -r packages/integrations/examples/crewai/requirements.txt
- run: >-
python -m py_compile
packages/integrations/examples/shared/vercel_sandbox_lease.py
packages/integrations/examples/shared/test_vercel_sandbox_lease.py
packages/integrations/examples/crewai/agent.py
packages/integrations/examples/crewai/sandbox.py
packages/integrations/examples/crewai/e2e.py
- run: python packages/integrations/examples/shared/test_vercel_sandbox_lease.py
- run: pnpm exec turbo run build --filter @browserbasehq/stagehand-codemode
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
Outdated
- run: pnpm --filter @browserbasehq/stagehand-integrations-example-vercel-sandbox pack:artifacts
- name: Detect CrewAI live test credentials
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
id: crewai-live-credentials
env:
BROWSERBASE_API_KEY: ${{ secrets.BROWSERBASE_API_KEY }}
BROWSERBASE_PROJECT_ID: ${{ secrets.BROWSERBASE_PROJECT_ID }}
VERCEL_OIDC_TOKEN: ${{ secrets.VERCEL_OIDC_TOKEN }}
VERCEL_TEAM_ID: ${{ secrets.VERCEL_TEAM_ID }}
VERCEL_PROJECT_ID: ${{ secrets.VERCEL_PROJECT_ID }}
VERCEL_TOKEN: ${{ secrets.VERCEL_TOKEN }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: |
if [[ -n "$BROWSERBASE_API_KEY" && -n "$BROWSERBASE_PROJECT_ID" && -n "$OPENAI_API_KEY" ]] && \
[[ -n "$VERCEL_OIDC_TOKEN" || ( -n "$VERCEL_TEAM_ID" && -n "$VERCEL_PROJECT_ID" && -n "$VERCEL_TOKEN" ) ]]; then
echo "available=true" >> "$GITHUB_OUTPUT"
else
echo "available=false" >> "$GITHUB_OUTPUT"
fi
- name: Run CrewAI live sandbox proof
if: steps.crewai-live-credentials.outputs.available == 'true'
run: python packages/integrations/examples/crewai/e2e.py
env:
STAGEHAND_SANDBOX_ARTIFACTS: ${{ github.workspace }}/packages/integrations/examples/vercel-sandbox/.artifacts
BROWSERBASE_API_KEY: ${{ secrets.BROWSERBASE_API_KEY }}
BROWSERBASE_PROJECT_ID: ${{ secrets.BROWSERBASE_PROJECT_ID }}
VERCEL_OIDC_TOKEN: ${{ secrets.VERCEL_OIDC_TOKEN }}
VERCEL_TEAM_ID: ${{ secrets.VERCEL_TEAM_ID }}
VERCEL_PROJECT_ID: ${{ secrets.VERCEL_PROJECT_ID }}
VERCEL_TOKEN: ${{ secrets.VERCEL_TOKEN }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
2 changes: 2 additions & 0 deletions packages/integrations/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,3 +81,5 @@ implementation modules and an in-process arbitrary-code executor are not public
- [Vercel Sandbox](./examples/vercel-sandbox) installs the exact packed artifact inside a
Firecracker microVM and returns a framework-neutral, bearer-authenticated MCP connection.
- [Mastra](./examples/mastra) consumes that connection with one persistent remote MCP client.
- [CrewAI](./examples/crewai) keeps its context-managed MCP adapter open across every tool call in
one crew execution.
73 changes: 73 additions & 0 deletions packages/integrations/examples/crewai/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
# CrewAI with Stagehand code mode

This example gives a CrewAI agent one browser tool, `code_execute`, without running generated
JavaScript in the Python agent process. A small trusted Python lease owner starts the
package-installed [Vercel Sandbox example](../vercel-sandbox) and receives its `{ url, token }`
connection. CrewAI then connects directly through authenticated Streamable HTTP:

```text
CrewAI MCPServerAdapter -> authenticated HTTPS -> Vercel Sandbox -> Stagehand code-mode MCP
`-> generated JavaScript
```

The adapter discovers the canonical tool description and uses it as the agent's backstory. It does
not copy the executor, schema, or code-mode skill.

## Setup

Create a Python 3.12 environment and install the pinned CrewAI example dependencies:

```bash
python3.12 -m venv .venv
. .venv/bin/activate
python -m pip install -r packages/integrations/examples/crewai/requirements.txt
```

Build and pack the exact Stagehand packages under review:

```bash
pnpm install
pnpm exec turbo run build --filter @browserbasehq/stagehand-codemode
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
Outdated
pnpm --filter @browserbasehq/stagehand-integrations-example-vercel-sandbox pack:artifacts
```

## Run the end-to-end proof

```bash
STAGEHAND_SANDBOX_ARTIFACTS="$PWD/packages/integrations/examples/vercel-sandbox/.artifacts" \
BROWSERBASE_API_KEY=<api-key> \
BROWSERBASE_PROJECT_ID=<project-id> \
VERCEL_OIDC_TOKEN=<oidc-token> \
OPENAI_API_KEY=<openai-key> \
python packages/integrations/examples/crewai/e2e.py
```

For external CI, replace `VERCEL_OIDC_TOKEN` with `VERCEL_TEAM_ID`, `VERCEL_PROJECT_ID`, and
`VERCEL_TOKEN`. Set `CREWAI_MODEL` in your own wrapper if you want to pass a model other than the
example's `openai/gpt-5-mini` default.

The proof uses one live package-installed sandbox and one context-managed CrewAI MCP adapter. It:

1. invokes `code_execute` twice directly and requires the same page ID and DOM marker;
2. records CrewAI's tool-usage event from a real model-selected `code_execute` call;
3. invokes the tool again to independently verify the model's browser-side change;
4. proves the model key and a host-only marker are absent inside generated code; and
5. closes the CrewAI MCP adapter before ending the sandbox lease, emitting `PASS` only afterward.

## Use the agent

```python
from agent import run_stagehand_agent
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
from sandbox import StagehandSandboxLease

with StagehandSandboxLease() as connection:
result = run_stagehand_agent(
connection,
"Open https://example.com and return its title and URL.",
)
print(result)
```

The lease subprocess receives only Browserbase, Vercel, artifact, and runtime variables from its
allowlist. Outer model-provider credentials are intentionally excluded, and the sandbox foundation
brokers the Browserbase credential at its egress boundary.
80 changes: 80 additions & 0 deletions packages/integrations/examples/crewai/agent.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
from __future__ import annotations

import os
from builtins import BaseExceptionGroup
from collections.abc import Iterator, Sequence
from contextlib import contextmanager
from typing import Any, Protocol

from crewai import Agent
from crewai.tools import BaseTool
from crewai_tools import MCPServerAdapter

DEFAULT_STAGEHAND_LLM = os.environ.get("CREWAI_MODEL", "openai/gpt-5-mini")


class StagehandSandboxConnection(Protocol):
url: str
token: str


@contextmanager
def stagehand_code_tools(
connection: StagehandSandboxConnection,
) -> Iterator[list[BaseTool]]:
"""Keep one authenticated remote MCP client open for a complete CrewAI run."""
adapter = MCPServerAdapter(
{
"url": connection.url,
"transport": "streamable-http",
"headers": {"Authorization": f"Bearer {connection.token}"},
},
connect_timeout=60,
)
try:
tools = list(adapter.tools)
names = [tool.name for tool in tools]
if names != ["code_execute"]:
raise RuntimeError(
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
Outdated
f"Expected only code_execute from Stagehand MCP, got {names!r}."
)
if "# Stagehand V4 code-mode syntax" not in tools[0].description:
raise RuntimeError("code_execute did not include the canonical guidance")
yield tools
except BaseException as primary_error:
try:
adapter.stop()
except BaseException as cleanup_error:
raise BaseExceptionGroup(
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
"CrewAI run and MCP cleanup both failed",
[primary_error, cleanup_error],
)
raise
else:
adapter.stop()


def build_stagehand_agent(
tools: Sequence[BaseTool],
llm: str | Any = DEFAULT_STAGEHAND_LLM,
) -> Agent:
if len(tools) != 1 or tools[0].name != "code_execute":
raise ValueError("CrewAI Stagehand agent requires exactly code_execute")
return Agent(
role="Stagehand browser agent",
goal="Complete browser tasks by writing compact, correct Stagehand V4 JavaScript.",
backstory=tools[0].description,
llm=llm,
tools=list(tools),
max_iter=8,
verbose=False,
)


def run_stagehand_agent(
connection: StagehandSandboxConnection,
prompt: str,
llm: str | Any = DEFAULT_STAGEHAND_LLM,
) -> str:
with stagehand_code_tools(connection) as tools:
return str(build_stagehand_agent(tools, llm).kickoff(prompt))
135 changes: 135 additions & 0 deletions packages/integrations/examples/crewai/e2e.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,135 @@
from __future__ import annotations

import json
import os
from typing import Any
from uuid import uuid4

from crewai.events import ToolUsageFinishedEvent, crewai_event_bus

from agent import build_stagehand_agent, stagehand_code_tools
from sandbox import StagehandSandboxLease


def successful_result(raw_result: Any) -> dict[str, Any]:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: Malformed MCP-result handling is only covered by the secret-backed live smoke, making parser/contract regressions hard to diagnose or exercise locally. Add focused tests for accepted results and invalid JSON, failed, and non-object value payloads.

(Based on your team's feedback about adding unit tests for new behavior.)

View Feedback

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At packages/integrations/examples/crewai/e2e.py, line 14:

<comment>Malformed MCP-result handling is only covered by the secret-backed live smoke, making parser/contract regressions hard to diagnose or exercise locally. Add focused tests for accepted results and invalid JSON, failed, and non-object value payloads.

(Based on your team's feedback about adding unit tests for new behavior.) </comment>

<file context>
@@ -0,0 +1,135 @@
+from sandbox import StagehandSandboxLease
+
+
+def successful_result(raw_result: Any) -> dict[str, Any]:
+    result = json.loads(str(raw_result))
+    assert result["ok"] is True, result
</file context>

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 60e1c451. Added four secret-free parser tests covering an accepted result, invalid JSON, a failed result, and a non-object value. Failures now use a typed fixed-message error without reflecting raw MCP payloads.

result = json.loads(str(raw_result))
assert result["ok"] is True, result
value = result.get("value")
assert isinstance(value, dict), result
return value


def main() -> None:
direct_marker = f"crewai-direct-{uuid4()}"
model_marker = f"crewai-model-{uuid4()}"
os.environ["CREWAI_HOST_ONLY_MARKER"] = f"host-{uuid4()}"
model_tool_calls: list[str] = []
final_state: dict[str, Any]

with StagehandSandboxLease() as connection:
with stagehand_code_tools(connection) as tools:
code_execute = tools[0]
first = successful_result(
code_execute.run(
code=f"""
await page.goto("https://example.com", {{ waitUntil: "domcontentloaded" }});
await page.evaluate((marker) => {{
document.documentElement.dataset.crewaiDirectMarker = marker;
}}, {json.dumps(direct_marker)});
return {{
pageId: page.pageId,
title: await page.title(),
directMarker: await page.evaluate(
() => document.documentElement.dataset.crewaiDirectMarker,
),
modelKeyVisible: process.env.OPENAI_API_KEY ?? null,
hostMarkerVisible: process.env.CREWAI_HOST_ONLY_MARKER ?? null,
}};
"""
)
)
assert first["title"] == "Example Domain"
assert first["directMarker"] == direct_marker
assert first["modelKeyVisible"] is None
assert first["hostMarkerVisible"] is None

second = successful_result(
code_execute.run(
code="""
return {
pageId: page.pageId,
title: await page.title(),
directMarker: await page.evaluate(
() => document.documentElement.dataset.crewaiDirectMarker,
),
};
"""
)
)
assert second["title"] == "Example Domain"
assert second["pageId"] == first["pageId"]
assert second["directMarker"] == direct_marker

@crewai_event_bus.on(ToolUsageFinishedEvent)
def record_model_tool(_source: Any, event: ToolUsageFinishedEvent) -> None:
if event.tool_name == "code_execute":
model_tool_calls.append(event.tool_name)

try:
agent = build_stagehand_agent(tools)
agent.kickoff(
" ".join(
(
"Use code_execute to modify the already-open page.",
"Set document.documentElement.dataset.crewaiModelMarker to",
f"{json.dumps(model_marker)}.",
"Then read that dataset value and the current pageId and report them.",
"You must call code_execute; do not merely describe JavaScript.",
)
)
)
assert crewai_event_bus.flush(), "CrewAI tool events did not finish"
finally:
crewai_event_bus.off(ToolUsageFinishedEvent, record_model_tool)
assert model_tool_calls, "the real CrewAI model must select code_execute"

final_state = successful_result(
code_execute.run(
code="""
return {
pageId: page.pageId,
title: await page.title(),
directMarker: await page.evaluate(
() => document.documentElement.dataset.crewaiDirectMarker,
),
modelMarker: await page.evaluate(
() => document.documentElement.dataset.crewaiModelMarker,
),
};
"""
)
)
assert final_state["title"] == "Example Domain"
assert final_state["pageId"] == first["pageId"]
assert final_state["directMarker"] == direct_marker
assert final_state["modelMarker"] == model_marker

print(
json.dumps(
{
"status": "PASS",
"framework": "crewai",
"directToolCalls": 3,
"modelToolCalls": len(model_tool_calls),
"sessionPersisted": True,
"modelCredentialIsolated": True,
"finalState": final_state,
"cleanup": ["crewai-mcp", "vercel-sandbox"],
},
sort_keys=True,
)
)


if __name__ == "__main__":
main()
2 changes: 2 additions & 0 deletions packages/integrations/examples/crewai/requirements.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
crewai==1.15.13
crewai-tools[mcp]==1.15.13
21 changes: 21 additions & 0 deletions packages/integrations/examples/crewai/sandbox.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
"""Expose the shared Vercel Sandbox lease from the CrewAI example directory."""

from __future__ import annotations

import sys
from pathlib import Path

SHARED_EXAMPLES_DIRECTORY = Path(__file__).resolve().parents[1] / "shared"
sys.path.insert(0, str(SHARED_EXAMPLES_DIRECTORY))

from vercel_sandbox_lease import ( # noqa: E402
StagehandSandboxConnection,
StagehandSandboxLease,
StagehandSandboxLeaseError,
)

__all__ = [
"StagehandSandboxConnection",
"StagehandSandboxLease",
"StagehandSandboxLeaseError",
]
Loading
Loading