You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 1b3d8e9
Browse filesBrowse the repository at this point in the historyBrowse files
feat(mcp): route hosted AI through OpenRouter with gpt-5.6-luna@high (#1213)
## Summary
Route hosted MCP AI through OpenRouter and default to
`openai/gpt-5.6-luna` with `reasoningEffort=high`.
### Key Changes
- Chat playground: hardcoded OpenAI `gpt-4o` → shared OpenRouter
provider
- Default OpenRouter model: `openai/gpt-5` → `openai/gpt-5.6-luna`
- Default OpenRouter reasoning effort: `high` (override with
`OPENROUTER_REASONING_EFFORT`)
- Chat `maxOutputTokens`: `2000` → `16000` so reasoning does not starve
text/tool output
- Worker/env docs/tests accept `OPENROUTER_*` +
`EMBEDDED_AGENT_PROVIDER`
## Why
Prod is moving off a direct OpenAI key onto OpenRouter. Luna@high keeps
the cheap GPT-5.6 tier while favoring query-translation quality over the
lower latency of medium effort. We can bump effort or model via env if
quality dips.
## Deploy notes
After merge, hosted worker should have:
- `OPENROUTER_API_KEY`
- `EMBEDDED_AGENT_PROVIDER=openrouter`
- optional `OPENROUTER_MODEL=openai/gpt-5.6-luna` (also the code
default)
- optional `OPENROUTER_REASONING_EFFORT=high` (also the code default;
lower it without redeploy if latency is too high)
- `OPENAI_API_KEY` can be removed once this is live
## How we'll measure
- Sentry error rate on embedded agents / chat (`callEmbeddedAgent`, LLM
provider failures, `NoObjectGeneratedError` rescues)
- Tool failure rate for `search_events` / `search_issues` /
`search_issue_events` (7d baseline ~11–12% on agent trio)
- Existing alert `444138` (global tool failure rate, critical 18%)
- Latency + OpenRouter spend vs prior OpenAI baseline
## Risks / follow-ups
- Luna may underperform on harder NL→query cases vs gpt-5 / sol / terra;
easy model override if needed
- Chat now requires `OPENROUTER_API_KEY`
- If both OpenAI and OpenRouter keys remain set without
`EMBEDDED_AGENT_PROVIDER`, provider selection still errors by design
- Did not run live OpenRouter evals in this PR
## Checks run
- `pnpm exec vitest run
packages/mcp-core/src/internal/agents/openrouter-provider.test.ts`
- `pnpm run tsc`
- `pnpm run lint`
- `pnpm run test`
<!-- junior-request-attribution:start -->
Requested by **David Cramer** via Junior.
<!-- junior-request-attribution:end -->
<!-- junior-session-footer:start -->
<!-- junior-conversation-id:slack%3AC08J1NSPU6S%3A1785362296.874629 -->
--
[View Junior
Session](https://junior-prod.sentry.dev/conversations/slack%3AC08J1NSPU6S%3A1785362296.874629)
<!-- junior-session-footer:end -->
---------
Co-authored-by: sentry-junior[bot] <264270552+sentry-junior[bot]@users.noreply.github.com>
Co-authored-by: David Cramer <david@sentry.io>
system: `You are an AI assistant designed EXCLUSIVELY for testing the Sentry MCP service. Your sole purpose is to help users test MCP functionality with their real Sentry account data - nothing more, nothing less.
@@ -332,8 +370,11 @@ Start conversations by exploring what's available in their account. Use tools li
332
370
- \`get_sentry_resource\` to dive deep into a specific issue, event, or trace
333
371
334
372
Remember: You're a test assistant, not a general-purpose helper. Stay focused on testing the MCP integration with their real data.`,
335
-
maxOutputTokens: 2000,
373
+
// Reasoning effort can consume completion budget before visible text/tool
374
+
// calls, so keep headroom above the old non-reasoning 2k cap.
0 commit comments