diff --git a/.gitignore b/.gitignore index eabebd36..579b48e4 100644 --- a/.gitignore +++ b/.gitignore @@ -1,3 +1,4 @@ .claude/* !.claude/settings.json .DS_Store +.idea/ diff --git a/README.md b/README.md index 133a8453..fabe9186 100644 --- a/README.md +++ b/README.md @@ -129,8 +129,8 @@ library that the build hydrates into the skill: - **`src/references/sdks//`** — per-platform install and per-signal code, one directory per supported platform. - **`src/references/concepts/`** — per-signal strategy: errors, tracing, logging, - metrics, profiling, session replay, user feedback, crons, releases, data scrubbing, - and choosing-a-signal. + metrics, profiling, session replay, user feedback, crons, uptime, releases, data + scrubbing, and choosing-a-signal. > Superseded per-SDK “wizard” skills are frozen under `skills-legacy/`, excluded from > the plugin build. diff --git a/src/SKILL_TREE.md b/src/SKILL_TREE.md index 6ba6d564..1d629a09 100644 --- a/src/SKILL_TREE.md +++ b/src/SKILL_TREE.md @@ -30,7 +30,7 @@ Each one is self-contained and named for the job it does. If you're not sure wha | [`sentry-debug-issue`](skills/sentry-debug-issue/SKILL.md) | Debug and fix a Sentry issue — find it (by link, ID, or search), pull full context (stack trace, breadcrumbs, trace, logs), optionally run Seer root-cause / autofix, apply the code fix, and resolve it via a `Fixes PROJECT-NAME-12A` commit/PR. Use when working a known error or hunting one down to fix. | | [`sentry-fix-stack-traces`](skills/sentry-fix-stack-traces/SKILL.md) | Make Sentry stack traces readable — upload source maps for JavaScript/TypeScript, or debug files for native and mobile (dSYM, ProGuard/R8, NDK symbols, Dart obfuscation maps, .NET PDBs). Use when frames in Sentry show minified names, bundled paths, hex addresses, "unknown", or method names with no file/line, instead of your original source. | | [`sentry-get-started`](skills/sentry-get-started/SKILL.md) | Guided entry point for using Sentry through your agent. Orients you to your current setup and, for a new project, sets up Sentry end to end with sane defaults — provision a project, install the SDK (errors, tracing, and whatever it enables by default), and confirm real telemetry reaches Sentry. Routes other intents (adding more signals, fixing issues) to the right skill. | -| [`sentry-instrument`](skills/sentry-instrument/SKILL.md) | Instrument an application with Sentry — detect the platform, install and initialize the SDK if needed, and wire up any signal — error monitoring, tracing/performance, logging, metrics, profiling, session replay, user feedback, cron check-ins, and AI/LLM monitoring (agent runs, token cost, and conversations for OpenAI, Anthropic, Vercel AI, LangChain, Google GenAI, Pydantic AI, Laravel AI, Eve, Flue, the Cloudflare Agents SDK, and Workers AI). Use to add Sentry to a project or to capture more than errors. | +| [`sentry-instrument`](skills/sentry-instrument/SKILL.md) | Instrument an application with Sentry — detect the platform, install and initialize the SDK if needed, and wire up any signal — error monitoring, tracing/performance, logging, metrics, profiling, session replay, user feedback, cron check-ins, uptime monitors for the deployed app, and AI/LLM monitoring (agent runs, token cost, and conversations for OpenAI, Anthropic, Vercel AI, LangChain, Google GenAI, Pydantic AI, Laravel AI, Eve, Flue, the Cloudflare Agents SDK, and Workers AI). Use to add Sentry to a project or to capture more than errors. | | [`sentry-otel-exporter-setup`](skills/sentry-otel-exporter-setup/SKILL.md) | Configure the OpenTelemetry Collector with Sentry Exporter for multi-project routing and automatic project creation. Use when setting up OTel with Sentry, configuring collector pipelines for traces and logs, or routing telemetry from multiple services to Sentry projects. | | [`sentry-setup-releases`](skills/sentry-setup-releases/SKILL.md) | Set up Sentry releases and deploy tracking — tag events with a version and environment, create the release in CI with its commits, and wire up suspect commits and code mappings, so Sentry can show which release introduced an issue, which commit is responsible, and release health. Use when asked to set up releases, track deploys, see what changed, or when issues show an unknown release or no suspect commit. | | [`sentry-snapshots-cocoa`](skills/sentry-snapshots-cocoa/SKILL.md) | Full Sentry Snapshots setup for Apple/Cocoa projects. Use when asked to "setup SnapshotPreviews", "setup Apple snapshot testing", "upload Apple snapshots to Sentry", "setup Apple snapshot GitHub Actions", or "setup Apple selective snapshot testing". | diff --git a/src/references/concepts/choosing-a-signal.md b/src/references/concepts/choosing-a-signal.md index eaa7a41c..1321ac6a 100644 --- a/src/references/concepts/choosing-a-signal.md +++ b/src/references/concepts/choosing-a-signal.md @@ -15,6 +15,7 @@ Decide by the question you are trying to answer. | *What did the user actually see and do?* | **Session Replay** | A video-like reproduction of a frontend/mobile session around an error or UX problem. | | *What does the user think went wrong?* | **User Feedback** | A qualitative report from a human, linked to the surrounding context. | | *Did my scheduled job run on time?* | **Cron monitor** | Check-ins that detect missed, late, or failed recurring jobs. | +| *Is my site or API up right now?* | **Uptime monitor** | Sentry requests a public URL on an interval and opens an issue when it stops answering. No SDK code. | Most of these signals carry the same **trace ID**, so once one surfaces a problem you can pivot to the others in the same request — the trace is the connective tissue that @@ -55,6 +56,8 @@ ties errors, spans, logs, replays, and metrics together for debugging. - **Replay:** frontend (and mobile) only; high sampling on errors, low on normal sessions. - **Crons:** every scheduled job whose silent failure would hurt. +- **Uptime:** every public endpoint whose downtime users would notice, once it is + deployed — usually one monitor per service, on a health route or the site root. When the user is unsure, ask what question they’re trying to answer and map it with the table above. When they say “set it up properly” / “you pick the defaults,” lean on the diff --git a/src/references/concepts/monitors.md b/src/references/concepts/monitors.md index b85e0f01..d82a6b62 100644 --- a/src/references/concepts/monitors.md +++ b/src/references/concepts/monitors.md @@ -30,7 +30,7 @@ monitor can feed several alerts. a prior window, or **dynamic anomaly detection**. Often created straight from a saved Discover or Metrics-Explorer query. - **Cron Monitor** — a scheduled-job watch via check-ins ([`crons.md`](crons.md)). - - **Uptime Monitor** — periodic HTTP checks against a URL. + - **Uptime Monitor** — periodic HTTP checks against a URL ([`uptime.md`](uptime.md)). - **Mobile Builds Monitor** — app-size thresholds across iOS/Android builds. **Monitor config also sets issue attributes at creation** — priority, auto-resolve, and @@ -61,18 +61,19 @@ An alert is **sources → triggers → filters → actions**: ## Coverage honesty -Alert creation is automatable via Sentry’s workflow-engine API; several monitor types -(uptime, dashboards) are heavier UI/API hand-offs today — be upfront about what the -agent can do end-to-end vs. -where it walks the user through the UI. The MCP is **read-only** here: it can inspect -alert rules (`find_alert_rules`, `get_alert_rule`), cron monitors and their check-ins -(`find_monitors`, `get_monitor_details`), and dashboards — useful for verifying after -creation — but there is no create or update path for any of them, and uptime monitors -have no MCP surface at all. +Alert creation is automatable via Sentry’s workflow-engine API, and **uptime monitors +can be created end-to-end through the MCP** (`create_uptime_monitor` and its siblings — +see [`uptime.md`](uptime.md)). Other monitor types and dashboards are heavier UI/API +hand-offs today — be upfront about what the agent can do end-to-end vs. +where it walks the user through the UI. For those the MCP is **read-only**: it can +inspect alert rules (`find_alert_rules`, `get_alert_rule`), cron monitors and their +check-ins (`find_monitors`, `get_monitor_details`), and dashboards — useful for +verifying after creation — but there is no create or update path for them. ## Related - [`crons.md`](crons.md) +- [`uptime.md`](uptime.md) - [`metrics.md`](metrics.md) - [`releases.md`](releases.md) - [`search-query-language.md`](../search-query-language.md) diff --git a/src/references/concepts/uptime.md b/src/references/concepts/uptime.md new file mode 100644 index 00000000..2dfd4c84 --- /dev/null +++ b/src/references/concepts/uptime.md @@ -0,0 +1,104 @@ +# Uptime Monitoring — What & Why + +Monitoring for whether a public URL is up. +Sentry sends an HTTP request to the URL on a fixed interval from its own checker +regions, and opens an issue when the URL stops responding with a success status. +Nothing runs in the app: there is no SDK code to add. +The monitor is server-side config, created through the MCP or the Sentry UI. + +Reach for it for **every public endpoint whose downtime users would notice** — the +production site, a public API, a health endpoint. +Errors and traces only arrive while the app is running and receiving traffic; an app +that is down, unreachable, or failing at the edge sends nothing, and uptime is what +notices. + +## Creating a monitor + +- **The MCP can create, update, and delete uptime monitors.** The tools are + `create_uptime_monitor`, `update_uptime_monitor`, `delete_uptime_monitor`, + `find_uptime_monitors`, and `get_uptime_monitor_details`. They are catalog tools: + reach them through `search_sentry_tools` / `execute_sentry_tool` if they aren’t + exposed directly. Creating one needs `project:write`. +- **Check first.** Call `find_uptime_monitors` before creating, and update the existing + monitor for a URL instead of adding a second one. +- **Create it in the project that owns the service**, with `environment` set to the + production environment name, so uptime issues land next to that service’s errors. +- **Defaults are usually right:** `intervalSeconds=60`, `timeoutMs=5000`, method `GET`. + An issue opens after 3 consecutive failed checks and resolves after 1 success; raise + `downtimeThreshold` only for a URL that is known to be flaky. + +## Picking the URL + +Do this as soon as the app has a production host — usually during setup, since most apps +are already deployed when Sentry is added. +Don’t wait for a later deploy step; the session may end before it. + +- **Find the production host, in this order:** + 1. The project’s own production events in Sentry — `search_events` for recent requests + in the production environment, and read the host from the request URL. This is the + host real traffic uses. + 2. The deploy setup — the hosting config, the production domain in env vars (`*_URL`, + `*_SITE_URL`, `*_BASE_URL`), the README, or a deploy section in + `AGENTS.md`/`CLAUDE.md`. + 3. Ask the user. + + Never monitor `localhost`, a preview deployment, or a host you guessed. + +- **Build the URL from the host plus the route.** Account for a `basePath` or path + prefix, rewrites and proxies, and an API served from a different host than the site. + +- **Pick a cheap, public `GET`.** An existing health route (`/health`, `/healthz`, + `/api/health`) is best: it answers fast, needs no auth, and doesn’t render a page. + If there isn’t one, the site root is fine for a site; for an API, offer to add a + minimal health route rather than pointing the check at an expensive or side-effecting + endpoint. + +- **Check it before creating.** Send a `GET` without credentials and confirm a `2xx`. + Show the user the URL and the result, and create the monitor once they confirm. + - A health route you just added isn’t deployed yet: monitor a URL that answers today + (the site root), and suggest switching the monitor to the health route after the + next deploy. + - If nothing on the host answers yet, don’t create the monitor — it would open a + downtime issue right away. + Tell the user it is the first thing to add once the app is live. + - If you can’t find or check a URL, ask; don’t guess. + +- **One monitor per public service, not per route.** Uptime answers “is it reachable”; + per-route failures are already errors and spans. + Add a second monitor only for a separately deployed service (an API on its own host, a + marketing site next to the app). + +## Alerts + +- **Usually nothing to set up.** An uptime issue opens as **high priority**, and every + new project starts with an alert that emails the issue owners (or all active members) + for new high-priority issues. + So a new monitor already notifies someone when the URL goes down. +- **Check that the default alert is still there.** Older projects may have edited or + deleted it: look with `find_alert_rules`. If no alert covers high-priority issues, + offer to add one. +- **For a different destination** — Slack, PagerDuty, a specific team — use the + `sentry-create-alert` skill. + To alert only on downtime, filter on `issue_category` `10` (Outage, which covers + uptime and cron issues). + The MCP’s uptime tools don’t configure alerts themselves. + +## Why an uptime issue often isn’t in the code + +- **Auth and redirects look like downtime.** A URL that returns `401`/`403`, or + redirects to a login page, fails the check while the app is healthy. + Monitor a URL that answers anonymously. +- **Firewalls, WAFs, and bot protection can block the checker.** If checks fail while + the site works in a browser, the requests are likely being blocked; see the + [troubleshooting docs](https://docs.sentry.io/product/monitors-and-alerts/monitors/uptime-monitoring/troubleshooting/). +- **Look at the trace.** When the app runs a Sentry SDK with tracing, a failing check + can carry a trace into the app, so the error behind a `5xx` is often on the same trace + ([uptime tracing](https://docs.sentry.io/product/monitors-and-alerts/monitors/uptime-monitoring/uptime-tracing/)). + +## Related + +- [`monitors.md`](monitors.md) — an uptime monitor is one kind of Monitor that creates + issues. +- [`crons.md`](crons.md) — the scheduled-job counterpart: crons notice a job that didn’t + run, uptime notices a service that isn’t answering. +- [Uptime Monitoring docs](https://docs.sentry.io/product/monitors-and-alerts/monitors/uptime-monitoring/) diff --git a/src/references/setup-verification.md b/src/references/setup-verification.md index ca948782..be20a52a 100644 --- a/src/references/setup-verification.md +++ b/src/references/setup-verification.md @@ -47,6 +47,8 @@ So exercise the real code path, not a standalone script: - **Metrics** → exercise the code that emits the metric. - **Crons** → find a way to invoke the job so the cron instrumentation triggers (its check-in fires). + - **Uptime** → nothing to trigger: after creating the monitor, wait for its first + check and confirm it passed with `get_uptime_monitor_details`. **Decide who boots the app — do not assume.** If you can tell how to start it, offer to start it and trigger the path yourself. diff --git a/src/skills/sentry-create-alert/SKILL.md b/src/skills/sentry-create-alert/SKILL.md index 7dd42115..0ec1c88e 100644 --- a/src/skills/sentry-create-alert/SKILL.md +++ b/src/skills/sentry-create-alert/SKILL.md @@ -97,7 +97,7 @@ Use `logicType: "all"`, `"any-short"`, or `"none"`. | `assigned_to` | `{"targetType": "Member", "targetIdentifier": 123}` | Issue assigned to target | | `level` | `{"level": 40, "match": "gte"}` | Event level (fatal=50, error=40, warning=30) | | `age_comparison` | `{"time": "hour", "value": 24, "comparisonType": "older"}` | Issue age | -| `issue_category` | `{"value": 1}` | Category (1=Error, 6=Feedback) | +| `issue_category` | `{"value": 1}` | Category (1=Error, 6=Feedback, 10=Outage: uptime and cron) | | `issue_occurrences` | `{"value": 100}` | Total occurrence count | **Interval options:** `"1min"`, `"5min"`, `"15min"`, `"1hr"`, `"1d"`, `"1w"`, `"30d"` diff --git a/src/skills/sentry-get-started/SKILL.md b/src/skills/sentry-get-started/SKILL.md index ebba14ae..70179355 100644 --- a/src/skills/sentry-get-started/SKILL.md +++ b/src/skills/sentry-get-started/SKILL.md @@ -102,7 +102,7 @@ Lead with what Sentry is, then transition into orienting: > request, and exact line that caused it — so you spend less time reproducing bugs and > more time fixing them. > Beyond errors it does tracing & performance, logs, metrics, profiling, session replay, -> cron monitoring, and AI/LLM monitoring — plus Seer, its AI debugging agent. +> cron and uptime monitoring, and AI/LLM monitoring — plus Seer, its AI debugging agent. > Right here in your agent I can set most of this up in your code and confirm it’s > actually working end to end — and once it’s running, investigate errors, dig into > performance problems, read your logs, and pull whatever Sentry telemetry we need to @@ -169,8 +169,8 @@ flag it. You’ll also want to immediately read and the baseline-signal context in hand before you start. When it’s done, surface other options — chiefly the **`sentry-instrument`** skill to add -more telemetry (logging, profiling, session replay, crons, …), and releases so issues -tie to the deploy that introduced them. +more telemetry (logging, profiling, session replay, crons, uptime, …), and releases so +issues tie to the deploy that introduced them. As in the existing-user path, only name a skill you’ve confirmed is available in your harness’s skill list; otherwise offer the docs fallback. Don’t auto-run them. @@ -192,8 +192,8 @@ to the honest docs offer below. Present the relevant options with your interactive prompt; the user can also just say what they want: -- **Add a signal** — tracing, logging, metrics, crons, profiling, session replay, user - feedback, AI/LLM monitoring. +- **Add a signal** — tracing, logging, metrics, crons, uptime, profiling, session + replay, user feedback, AI/LLM monitoring. → the **`sentry-instrument`** skill. - **Set up Sentry properly** (recommended defaults across several signals). → the **`sentry-instrument`** skill. diff --git a/src/skills/sentry-instrument/SKILL.md b/src/skills/sentry-instrument/SKILL.md index fcc030e2..b8d278a9 100644 --- a/src/skills/sentry-instrument/SKILL.md +++ b/src/skills/sentry-instrument/SKILL.md @@ -1,6 +1,6 @@ --- name: sentry-instrument -description: Instrument an application with Sentry — detect the platform, install and initialize the SDK if needed, and wire up any signal — error monitoring, tracing/performance, logging, metrics, profiling, session replay, user feedback, cron check-ins, and AI/LLM monitoring (agent runs, token cost, and conversations for OpenAI, Anthropic, Vercel AI, LangChain, Google GenAI, Pydantic AI, Laravel AI, Eve, Flue, the Cloudflare Agents SDK, and Workers AI). Use to add Sentry to a project or to capture more than errors. +description: Instrument an application with Sentry — detect the platform, install and initialize the SDK if needed, and wire up any signal — error monitoring, tracing/performance, logging, metrics, profiling, session replay, user feedback, cron check-ins, uptime monitors for the deployed app, and AI/LLM monitoring (agent runs, token cost, and conversations for OpenAI, Anthropic, Vercel AI, LangChain, Google GenAI, Pydantic AI, Laravel AI, Eve, Flue, the Cloudflare Agents SDK, and Workers AI). Use to add Sentry to a project or to capture more than errors. license: Apache-2.0 --- # Sentry Instrument @@ -35,7 +35,7 @@ Decide what you’re actually doing; it gates how much you run. | --- | --- | --- | | **First error** | Brand-new install, no Sentry yet | Detect setup ownership, then provision and install the selected base. Verify a real error when the path supports it; disclose any trace-only limitation. Defer *additional* signals (logging, profiling, replay, metrics, …). | | **Add a signal** | Sentry already installed; user wants one more signal | Preserve the base install, run setup-ownership detection, then wire only that signal. | -| **Full setup** | “Set it up properly / sensible defaults” | Run the ownership-aware base setup, then propose the rest of a baseline (releases, source maps, and any signals that fit the app) and add what the user accepts. | +| **Full setup** | “Set it up properly / sensible defaults” | Run the ownership-aware base setup, then propose the rest of a baseline (releases, source maps, an uptime monitor once the app has a production URL, and any signals that fit the app) and add what the user accepts. | Never over-instrument — wiring up logging, session replay, profiling, metrics, etc. upfront when the user only asked to get Sentry working is doing more than they asked @@ -101,8 +101,12 @@ For each signal the scope calls for: apply the code. Signals this skill wires up: error monitoring, tracing/performance, profiling (requires -tracing), logging, metrics, cron check-in code, session replay, user feedback, and -AI/LLM monitoring. +tracing), logging, metrics, cron check-in code, session replay, user feedback, uptime +monitors, and AI/LLM monitoring. + +**Uptime has no SDK code.** Instead of fetching a docs page, read +[`references/concepts/uptime.md`](references/concepts/uptime.md), confirm the production +URL with the user, and create the monitor with the MCP’s `create_uptime_monitor`. For AI/LLM monitoring, keep input and output capture enabled by default because the Agent Tracing transcript and debugging workflow rely on prompts, responses, tool @@ -204,6 +208,10 @@ auto-running them: Do not offer this JavaScript SDK option for Python, PHP, unknown SDKs, or framework-owned OTLP setups without a JavaScript Sentry SDK. - Ship it to production. +- If the app already has a production host and no uptime monitor, offer one now so + Sentry notices when the app stops answering — don’t wait for a later deploy step. + [`references/concepts/uptime.md`](references/concepts/uptime.md) covers finding the + real URL (production events in Sentry first) and checking it before creating. - Add a signal — logging, session replay, or profiling are common next steps (tracing is already in the base `init`). - Harden the setup — readable stack traces (source maps for JS, debug symbols for