diff --git a/.claude/skills/board-ops/SKILL.md b/.claude/skills/board-ops/SKILL.md index c34f1c1652..38294391b6 100644 --- a/.claude/skills/board-ops/SKILL.md +++ b/.claude/skills/board-ops/SKILL.md @@ -57,25 +57,19 @@ a fix exists. The flow is `/security-advisory`. ⚠️ **A draft card has no repository and no issue number, so the issue-side lookup below cannot find one**, and `item-add --url` has no URL to be given. -Look it up by **title** in the full listing instead, then feed that item id to -`item-edit` or `item-delete` exactly as usual: +Look it up by **title** with the script (`scripts/board-draft-find.mjs`, +#2558), then feed the printed item id to `item-edit` or `item-delete` exactly +as usual: ```sh -GHSA=GHSA-xxxx-yyyy-zzzz # the advisory's real id -ITEM_ID= # never let an earlier lookup's id survive a failed one -BOARD=$(gh project item-list 28 --owner modelcontextprotocol --format json --limit 2000) -if jq -e '(.items | length) == .totalCount' <<<"$BOARD" >/dev/null; then - ITEM_ID=$(jq -r '.items[] | select(.content.type=="DraftIssue") - | select(.content.title | startswith("['"$GHSA"']")) | .id' <<<"$BOARD") - [ -n "$ITEM_ID" ] || echo "no draft card titled [$GHSA] on #28" >&2 -else - echo "item-list incomplete or failed — raise --limit; not concluding anything" >&2 -fi +npm run board:find-draft -- --ghsa GHSA-xxxx-yyyy-zzzz # prints ITEM= ``` -Match on the **bracketed GHSA id**, not on words from the summary — a summary is -free text and two advisories can share one. Advisory drafts live on #28 only; -`/issue-triage`'s audit reports one found anywhere else. +It matches on the **bracketed GHSA id**, not on words from the summary — a +summary is free text and two advisories can share one — and it trusts the +listing only when complete, so a truncated dump reads as an error rather than +as "no draft card". Advisory drafts live on #28 only; `/issue-triage`'s audit +reports one found anywhere else. ## V2 board (#28) IDs @@ -142,34 +136,40 @@ Don't try to set one here; the field id doesn't exist. ### Add a card and set its fields +Use the script (`scripts/board-card-add.mjs`, #2558) — it adds the card, +resolves every field and option id by name, sets Status (and Priority when +given), and verifies each by reading it back before printing `card: …`: + ```sh -# Prints the item id (PVTI_…); capture it. -ITEM_ID=$(gh project item-add 28 --owner modelcontextprotocol --url <issue-url> --format json --jq '.id') +# An issue you filed through the create flow is approved by definition → Todo. +npm run board:add -- --issue <N> --status Todo --priority Medium +``` -# Status → Todo (an issue you filed through the create flow is approved by definition) -gh project item-edit --project-id PVT_kwDOCt2Azc4BJVxt --id "$ITEM_ID" \ - --field-id PVTSSF_lADOCt2Azc4BJVxtzg5iI8c --single-select-option-id fbdaf21e +For an issue swept in at triage, the differences are `--status Incoming`, a +`--priority` still scored with the rubric in `/issue-triage` (every v2 card +carries one — `board:audit` flags a card without it), and that you do **not** +set a milestone. -# Priority → Medium -gh project item-edit --project-id PVT_kwDOCt2Azc4BJVxt --id "$ITEM_ID" \ - --field-id PVTSSF_lADOCt2Azc4BJVxtzg5iJE4 --single-select-option-id da944a9c -``` +For **v1**, the same against board #11 — and **no `--priority`**, which that +board has no field for: -Each `item-edit` sets **one** field, so setting both takes two calls — there is -no combined form. +```sh +npm run board:add -- --issue <N> --status Todo --board 11 +``` -For an issue swept in at triage, the only difference is Status → **Incoming** -(`721a3d4c`) and that you do **not** set a milestone. +### Move an existing card -For **v1**, the same shape against board #11: +**For a plain Status move, use the script** (`scripts/board-card-status.mjs`, +#2558) — it does everything this recipe describes (name-resolved ids, +issue-side lookup, edit, verify re-read) and prints `card: <Status>` only on a +confirmed move: ```sh -ITEM_ID=$(gh project item-add 11 --owner modelcontextprotocol --url <issue-url> --format json --jq '.id') -gh project item-edit --project-id PVT_kwDOCt2Azc4BA5sz --id "$ITEM_ID" \ - --field-id PVTSSF_lADOCt2Azc4BA5szzgzkS-g --single-select-option-id f75ad846 +npm run board:status -- --issue <N> --status "In Review" # --board 11 for a v1 issue ``` -### Move an existing card +The manual recipe below remains for what the scripts do not do — adapting the +lookup for another field — and as the record of how the lookup works. Look the item id up **from the issue** rather than re-adding it. An issue's `projectItems` lists the cards it has on every board, so the lookup does not @@ -219,18 +219,20 @@ fi **`Done` means the work shipped.** An issue closed as duplicate / won't fix / not planned / obsolete / superseded shipped nothing, so its card is **deleted**, -not parked in Done: +not parked in Done. Use the script (`scripts/board-card-delete.mjs`, #2558) — +it looks the card up from the issue, deletes it, and verifies it is gone: ```sh -# ITEM_ID from the issue-side LOOKUP block in "Move an existing card" above — -# the lookup only, not the item-edit that follows it. -if [ -n "$ITEM_ID" ]; then - gh project item-delete 28 --owner modelcontextprotocol --id "$ITEM_ID" -else - echo "no ITEM_ID — nothing deleted" >&2 -fi +npm run board:delete -- --issue <N> # --board 11 for a v1 card +npm run board:delete -- --issue <N> --reason duplicate # …and close the issue ``` +An absent card always fails the run — including with `--reason`, so a wrong +`--board` or a typo'd issue number cannot close an issue whose real card +survives. The one legitimate absent-card case is retrying a run that deleted +the card and then failed the close; declare it with `--allow-missing-card` to +proceed to the close anyway. + Deleting the card removes it from the board only — **the issue itself is untouched**, keeps its labels and comments, and stays searchable and linkable forever. Nothing is lost; the board simply stops claiming the work was @@ -240,15 +242,9 @@ later. The close **reason** is the machine-readable form of the same distinction. `gh issue close --reason` accepts only `completed` and `not planned`, so -**`duplicate` must be set through the API**: - -```sh -gh api repos/modelcontextprotocol/inspector/issues/<N> -X PATCH \ - -f state=closed -f state_reason=duplicate -``` - -(or "Mark as duplicate" in the web UI, which additionally records a -duplicate-of link). +`--reason duplicate` goes through the API (a PATCH setting +`state_reason=duplicate`) — the script does that for you. "Mark as duplicate" +in the web UI additionally records a duplicate-of link. ## ⚠️ The option-deletion hazard @@ -272,10 +268,9 @@ Safe alternatives, in order of preference: `id`s**, then call `updateProjectV2Field` echoing back every existing option **including its `id`**, appending only the new one. `ProjectV2SingleSelectFieldOptionInput.id` is an optional `String`, so a mixed - list works. Verify afterward that no card lost its value — snapshot - `gh project item-list … --format json --limit 2000` before and after, check - each is complete the way the snapshot below does, and diff; don't just - spot-check. Send those dumps to `$BOARD_TMP` too, for the reason above. + list works. Verify afterward that no card lost its value — take a + `npm run board:snapshot` before and after and diff the two dumps; don't just + spot-check. The script keeps both out of the worktree, for the reason below. Both the `Incoming` Status option and the Urgent/High/Medium/Low Priority options were added this way (#1891), with the before/after diff confirming all @@ -294,15 +289,13 @@ so a snapshot is a full dump of item IDs and every card's Status and Priority. Left in the working tree it is one `git add -A` away from being published in a PR (Copilot). +The script (`scripts/board-snapshot.mjs`, #2558) enforces both hazards: it +writes to a fresh temp dir by default, refuses a `--dir` inside the working +tree, and writes nothing from a truncated listing — a truncated snapshot +cannot restore the cards it dropped: + ```sh -BOARD_TMP=$(mktemp -d) -gh project item-list 28 --owner modelcontextprotocol --format json --limit 2000 \ - > "$BOARD_TMP/board-snapshot.json" -# A truncated snapshot cannot restore the cards it dropped — refuse to proceed on one. -jq -e '(.items | length) == .totalCount' "$BOARD_TMP/board-snapshot.json" >/dev/null \ - && echo "snapshot: $BOARD_TMP/board-snapshot.json" \ - || { echo "SNAPSHOT INCOMPLETE — raise --limit and retake it before editing options" >&2 - rm -f "$BOARD_TMP/board-snapshot.json"; false; } +npm run board:snapshot # prints snapshot: <path> (<count> items) ``` Note the printed path; you need it to recover. @@ -313,55 +306,41 @@ This has happened twice — once via the API (~197 items, reconstructed by inference) and once via the UI (the `Done` column, 247 items, restored from a snapshot in minutes). With a snapshot the recovery is mechanical. -The recipe below is written for a deleted **Status** option. For a deleted -**Priority** option it is the same three steps with two substitutions: read -`.priority` instead of `.status` (`gh project item-list --format json` exposes -each single-select field under its lowercased name, so both keys are present), -and pass the Priority field id `PVTSSF_lADOCt2Azc4BJVxtzg5iJE4`. +The recipe is three steps; the two mechanical ones are the script +(`scripts/board-recover.mjs`, #2558), written for Status by default — pass +`--field Priority` for a deleted **Priority** option. Step 2 — the one that +edits the field schema, which is what the hazard above is about — stays a +deliberate human act. ```sh -# 0. Same temp dir the snapshot went to — keep every dump out of the worktree. -BOARD_TMP=${BOARD_TMP:-$(mktemp -d)} - -# 1. Which cards lost their value, and what did they hold? lost-ids.json is -# kept ONLY when the dump is complete AND the snapshot reports what those cards -# held — step 3 refuses to run without it, so neither a truncated dump nor a -# missing snapshot can turn into a silent no-op or an unconfirmed re-apply. -rm -f "$BOARD_TMP/lost-ids.json" -gh project item-list 28 --owner modelcontextprotocol --format json --limit 2000 \ - > "$BOARD_TMP/board-broken.json" -if jq -e '(.items | length) == .totalCount' "$BOARD_TMP/board-broken.json" >/dev/null; then - jq -r '[.items[]|select(.status==null)|.id]' "$BOARD_TMP/board-broken.json" \ - > "$BOARD_TMP/lost-ids.json" || rm -f "$BOARD_TMP/lost-ids.json" - jq -r --slurpfile L "$BOARD_TMP/lost-ids.json" '($L[0]) as $lost - | [.items[] | select(.id as $i | $lost|index($i)) | .status // "(none)"] - | group_by(.) | map({s:.[0],c:length}) | .[] | "was \(.s): \(.c)"' \ - "$BOARD_TMP/board-snapshot.json" \ - || { echo "no usable snapshot — cannot confirm what these cards held; not re-applying" >&2 - rm -f "$BOARD_TMP/lost-ids.json"; } -else - echo "board-broken.json INCOMPLETE — raise --limit and re-run step 1" >&2 - rm -f "$BOARD_TMP/board-broken.json" -fi +# 1. Which cards lost their value, and what did they hold? Writes +# lost-ids.json beside the snapshot, from a complete dump only, and prints +# "was <value>: <count>" from the snapshot. +npm run board:recover -- --phase diff --snapshot <path-from-board:snapshot> -# 2. Recreate the option, echoing every surviving option's id (see above). -# NOTE: the recreated option gets a NEW id — the deleted one never comes back. +# 2. Recreate the option — in the web UI, or echoing every surviving option's +# id (see above). NOTE: the recreated option gets a NEW id — the deleted +# one never comes back. -# 3. Re-apply it to the orphaned cards. -if [ -s "$BOARD_TMP/lost-ids.json" ]; then - for id in $(jq -r '.[]' "$BOARD_TMP/lost-ids.json"); do - gh project item-edit --project-id PVT_kwDOCt2Azc4BJVxt --id "$id" \ - --field-id PVTSSF_lADOCt2Azc4BJVxtzg5iI8c --single-select-option-id <NEW_OPTION_ID> - sleep 0.4 - done -else - echo "no lost-ids.json — step 1 did not complete; nothing re-applied" >&2 -fi +# 3. Re-apply the new option id to the orphaned cards (paced). +npm run board:recover -- --phase reapply --lost <dir>/lost-ids.json --option-id <NEW_OPTION_ID> ``` -Step 1's grouping is the safety check: confirm the orphaned set is exactly the -cards that held the deleted option, so you don't overwrite a card someone -legitimately moved in the meantime. +Step 1's grouping is the safety check, and the script enforces it: a card is +counted as lost only when it is blank now **and** held a value in the snapshot +(a card blank before the deletion, or added since, is excluded), and +`lost-ids.json` is written only when every lost card held the **same** value — +a mixed grouping is printed and refused, since one option id cannot restore +two. Step 3 refuses to run without step 1's file, so neither a truncated dump +nor a missing snapshot can turn into a silent no-op or an unconfirmed +re-apply. The file also records the board, the field and the value the lost +cards held, and step 3 verifies `--option-id` against them — an option id that +is valid on the field but is not the recreated option for that value is +refused rather than rewriting every lost card to the wrong one. Finally, step +3 re-reads each card immediately before editing it and aborts — naming how far +it got — if any card was deleted or set in the meantime, so a stale lost list +never overwrites a value a maintainer legitimately set; re-run step 1 and +retry with the remainder. Because the recreated option carries a **new id**, the tables above and every reference to it must be updated in the same change — `grep` the old id across diff --git a/.claude/skills/issue-triage/SKILL.md b/.claude/skills/issue-triage/SKILL.md index ba34eca105..db9d4c5292 100644 --- a/.claude/skills/issue-triage/SKILL.md +++ b/.claude/skills/issue-triage/SKILL.md @@ -46,29 +46,13 @@ work a maintainer had already scheduled. Diff the open issues against **both boards**. Diffing against #28 alone is wrong: a `v1` issue correctly carded on #11 is reported as unboarded and gets double-boarded (a real defect a past sweep introduced — #1929 reproduced it). +The script (`scripts/board-sweep.mjs`, #2558) does the union, trusts each dump +only when complete (a truncated listing makes a carded issue read as +unboarded), and prints each unboarded issue with its destination — milestoned +already → Todo, otherwise → Incoming: ```sh -D=$(mktemp -d) -gh issue list --repo modelcontextprotocol/inspector --state open --limit 1000 \ - --json number,milestone > "$D/open.json" -# item-list truncates SILENTLY past --limit (and a failed call writes nothing), and a -# missing card reads as an "unboarded" issue that then gets double-carded — so an -# incomplete dump is deleted, and the steps below fail on the missing file. -for P in 28 11; do - gh project item-list $P --owner modelcontextprotocol --format json --limit 2000 > "$D/b$P.json" - jq -e '(.items | length) == .totalCount' "$D/b$P.json" >/dev/null \ - || { echo "board #$P listing INCOMPLETE — raise --limit and re-run" >&2; rm -f "$D/b$P.json"; false; } -done -# Union of BOTH boards, filtered to this repo — org boards can hold other repos' issues. -jq -s '[.[].items[] | select(.content.type=="Issue" - and .content.repository=="modelcontextprotocol/inspector") - | .content.number]' "$D/b28.json" "$D/b11.json" > "$D/boarded.json" \ - || rm -f "$D/boarded.json" -# Prints the destination too: milestoned already → Todo, otherwise → Incoming. -jq -r --slurpfile b "$D/boarded.json" \ - '.[] | select(.number as $n | ($b[0]|index($n))|not) - | "#\(.number)\t→ \(if .milestone then "Todo (has milestone \(.milestone.title))" else "Incoming" end)"' \ - "$D/open.json" +npm run board:sweep # exits non-zero when any unboarded issue exists ``` ## Pass 2 — approve what should ship @@ -219,79 +203,15 @@ count means the board contradicts a rule, not that the rule needs revisiting. | Open, but carded `Done` | A card in Done ⇒ its issue is closed | Close the issue, or move the card back | ```sh -D=$(mktemp -d); R=modelcontextprotocol/inspector -# --limit must exceed the repo's TOTAL issue count (884 as of 2026-08-05), not just the open ones — -# the last check below reads closed issues' state reasons. -gh issue list --repo $R --state all --limit 2000 \ - --json number,state,stateReason,labels,milestone > "$D/i.json" -# item-list truncates SILENTLY past --limit (and a failed call writes nothing); an -# incomplete dump would make every check below lie, so it is deleted and the audit -# fails on the missing file instead. -for P in 28 11; do gh project item-list $P --owner modelcontextprotocol \ - --format json --limit 2000 > "$D/b$P.json" - jq -e '(.items | length) == .totalCount' "$D/b$P.json" >/dev/null \ - || { echo "board #$P listing INCOMPLETE — raise --limit and re-run" >&2; rm -f "$D/b$P.json"; false; } -done -jq -nr --slurpfile o "$D/i.json" --slurpfile a "$D/b28.json" --slurpfile b "$D/b11.json" --arg R "$R" ' - ($o[0] | map({key:(.number|tostring), value:{st:.state, sr:(.stateReason // ""), - lab:[.labels[].name], ms:(.milestone.title // null)}}) | from_entries) as $M - | def own($s): [$s[].items[] - # A DRAFT card has no `.content.repository`, so filtering on equality - # alone drops the very items the "non-Issue" check exists to find. - | select((.content.repository // null) == null or .content.repository==$R)]; - def I($n): ($M[($n|tostring)] // null); - def ms($n): (I($n).ms // null); - def lab($n): (I($n).lab // []); - def isopen($n): (I($n).st == "OPEN"); - def shipped($n): (I($n).sr == "COMPLETED"); - [own($a)[] | select(.content.type=="Issue") | {n:.content.number, s:.status, p:.priority}] as $B28 - | [own($b)[] | select(.content.type=="Issue") | {n:.content.number, s:.status}] as $B11 - | { - "double-boarded": [$B28[].n | select(. as $n | [$B11[].n]|index($n))], - # An advisory draft card is the ONE legitimate non-Issue item (see AGENTS.md). - # The exemption is narrowed three ways, and each one matters: DRAFTS only - # (a GHSA-titled PR is still reported), board #28 ONLY (an advisory has no - # business on #11), and the `[GHSA-` title prefix (a stray draft is still - # reported). Reports the TITLE, since a draft has no number. - "non-Issue on a board": [(own($a)[] | select(.content.type!="Issue" - and ((.content.type=="DraftIssue" - and ((.content.title // "") | startswith("[GHSA-"))) | not))), - (own($b)[] | select(.content.type!="Issue"))] - | map(.content.title // "(untitled)"), - # $B28/$B11 hold only Issue items, so the Status and Priority checks below - # cannot see an advisory draft. Exempting drafts from the check above would - # therefore have made a half-made advisory card invisible to the whole - # audit; this is the narrow replacement. - "GHSA draft missing Status/Priority": - [own($a)[] | select(.content.type=="DraftIssue" - and ((.content.title // "") | startswith("[GHSA-"))) - | select(.status==null or .priority==null) - | (.content.title[0:24])], - "no Status": [($B28[], $B11[]) | select(.s==null) | .n], - "Incoming w/ milestone": [$B28[] | select(.s=="Incoming" and ms(.n)!=null) | .n], - "past Incoming, no ms": [$B28[] | select(.s!=null and .s!="Incoming" and .s!="Done" - and isopen(.n) and ms(.n)==null) | .n], - "v1 label on #28": [$B28[] | select(isopen(.n) and (lab(.n)|index("v1"))) | .n], - "v2 label on #11": [$B11[] | select(isopen(.n) and (lab(.n)|index("v2"))) | .n], - "open, not exactly 1 version label": - [$o[0][] | select(.state=="OPEN") - | select(([.labels[].name] | map(select(IN("v1","v2"))) | length) != 1) - | .number], - "open, not exactly 1 type label": - [$o[0][] | select(.state=="OPEN") - | select(([.labels[].name] - | map(select(IN("bug","enhancement","documentation","chore","question"))) - | length) != 1) - | .number], - "#28 open, no Priority": [$B28[] | select(.p==null and isopen(.n)) | .n], - "closed unshipped, still carded": - [($B28[], $B11[]) | select(I(.n)!=null and (isopen(.n)|not) - and (shipped(.n)|not)) | .n], - "open, but carded Done": [($B28[], $B11[]) | select(.s=="Done" and isopen(.n)) | .n] - } | to_entries[] | "\(.value|length)\t\(.key)\t\(.value[0:10])"' +npm run board:audit # one line per check; exits non-zero when any is dirty ``` -Two things the queries must account for, both learned the hard way: +The script (`scripts/board-audit.mjs`, #2558) prints +`<count>\t<check>\t<first offenders>` for every invariant in the table above, +over complete dumps only — a truncated listing would make every check lie, so +it refuses one instead. + +Two things its queries account for, both learned the hard way: - **Filter by repository.** These are **org** projects and can hold cards from any repo in the org — board #11 currently carries one diff --git a/.claude/skills/local-dev/SKILL.md b/.claude/skills/local-dev/SKILL.md index 5236cb2a05..707d3d108b 100644 --- a/.claude/skills/local-dev/SKILL.md +++ b/.claude/skills/local-dev/SKILL.md @@ -21,8 +21,8 @@ npm install # at the REPO ROOT v2 is **not** an npm workspace — each client under `clients/*` keeps its own `package.json` and `node_modules`. A single root `npm install` is still all you need: the root `postinstall` (`scripts/install-clients.mjs`) cascades -`npm install` into `clients/web`, `clients/cli`, `clients/tui`, and -`clients/launcher`. +`npm install` into `clients/web`, `clients/cli`, `clients/mcpdo`, +`clients/tui`, and `clients/launcher`. - **Fresh clone:** `npm install` at the root. - **After a pull that changes a client's dependencies:** re-run `npm install` at @@ -52,21 +52,24 @@ The launcher-driven scripts run the **built** launcher, so `npm run build` first: ```sh -npm run build # web → cli → tui → launcher +npm run build # web → cli → mcpdo → tui → launcher npm run web # prod web launcher against clients/web/dist npm run web:dev # web launcher in --dev mode (Vite) ``` -Individual builds: `build:web`, `build:cli`, `build:tui`, `build:launcher`. The +Individual builds: `build:web`, `build:cli`, `build:mcpdo`, `build:tui`, +`build:launcher`. The web build produces both the browser SPA (`clients/web/dist`, Vite) and the Node prod-server runner (`clients/web/build`, tsup). To run the CLI or TUI: `node clients/launcher/build/index.js --cli …` / -`--tui …`. +`--tui …`. The connection CLI (`mcpdo`) has its own bin: +`node clients/mcpdo/build/mcp-bin.js …` (or `npm link` from +`clients/mcpdo` for a global `mcpdo`). ## The `@inspector/core` alias -`core/` holds the logic shared by all three clients and intentionally has **no +`core/` holds the logic shared by all five clients and intentionally has **no `package.json`** — it is not published on its own. Each client bundles it via a build-time alias: @@ -82,7 +85,7 @@ build-time alias: ## Where a dependency goes **The rules are in [`AGENTS.md`](../../../AGENTS.md) → Dependency placement, and -they are not restated here.** Read them there and come back for the *why* — what +they are not restated here.** Read them there and come back for the _why_ — what each rule is defending against, what it looked like when it was violated, and how to tell you have hit one. @@ -107,8 +110,8 @@ Keep two distinctions straight, because AGENTS.md's rules split on them: - **Root-declared is not the same as `core/`-imported.** `commander`, `open` and `@hono/node-server` are root `dependencies` too, but they are reached only - from client code. Only the `core/` set has to appear in *all three* bundler - `external` lists. + from client code. Only the `core/` set has to appear in _every_ client's + bundler `external` list. - **Root-declared is not the same as aliased.** The `vitest.shared.mts` pins and the `clients/web/tsconfig.*.json` `paths` cover the packages whose resolution is genuinely ambiguous, which is two different situations: the importer is @@ -163,25 +166,25 @@ do not need to be: `npm run` prepends **every ancestor** `node_modules/.bin` to all still resolves the root's copy. `clients/launcher` declares no `devDependencies` whatsoever and its `validate` is unchanged. -What a per-client declaration *does* buy is a second copy free to drift, and it +What a per-client declaration _does_ buy is a second copy free to drift, and it had (#2196): `globals` sat at `^17.7.0` at the root against `^17.4.0` in all four clients, and `typescript-eslint` at `^8.65.0` against `^8.56.1`. Nothing failed — which is the point. A lint or format tool that differs per client makes the gate's -verdict a function of *where you ran it*, and the exact `prettier` pin (#1790) +verdict a function of _where you ran it_, and the exact `prettier` pin (#1790) only means something when there is one of it. ⚠️ The line is **used by every client**, not "used by one" and not "is it toolchain". `tsx`, `playwright`, `storybook`, `happy-dom`, `ink-testing-library`, `vite-node` and each client's own `@types/*` are toolchain too and stay where -they are — hoisting them would make every client install the union of all four. +they are — hoisting them would make every client install the union of all five. So do the ones **more than one** client declares without all of them doing so: `tsup` sits in web, cli and tui, and `vite` in web and tui on top of the root -*runtime* `dependency` that `--web --dev` needs. Neither is in scope here; +_runtime_ `dependency` that `--web --dev` needs. Neither is in scope here; whether to consolidate them is a separate call with a separate rationale (`vite` especially, since its root declaration is a `dependency`, not a `devDependency`). -#### What the walk-up does *not* buy you +#### What the walk-up does _not_ buy you ⚠️ **Deleting a client's declaration does not always delete the copy** — and where a copy survives, it is the one that wins. Two mechanisms put one back, @@ -195,7 +198,7 @@ neither of which the manifest mentions: - **A hoisted transitive.** `@types/express` brings `@types/node` into web and cli's trees on its own. -Those copies sit *nearer* than the root's, so `clients/web/node_modules/.bin` +Those copies sit _nearer_ than the root's, so `clients/web/node_modules/.bin` precedes the root bin directory on `PATH` and TypeScript resolves the nearest `node_modules/@types`. Verify with `npm exec -- which eslint` from the client rather than assuming — the assumption is what made the first cut of #2196 claim @@ -217,10 +220,10 @@ on disk. The two mechanisms are **not** equally safe, and neither is a guarantee #2226. ✅ **`verify:dep-lockstep` gates both of those since #2226.** Its second tier -compares every package **any** install *declares* — `dependencies`, +compares every package **any** install _declares_ — `dependencies`, `devDependencies` and `optionalDependencies`, unioned across the root and all -four clients — against every **top-level** copy in every install, independent of -what a `tsc` program resolves. So a tool *binary* that no program loads +five clients — against every **top-level** copy in every install, independent of +what a `tsc` program resolves. So a tool _binary_ that no program loads (`eslint`, `typescript`, `vitest`) and a transitive copy that no single program meets (the cli `@types/node` above) are both in scope now, as is a skew between two **clients** with no root copy involved (`@types/react`, web against tui). @@ -270,7 +273,7 @@ Two live examples worth knowing: tsup and Vite externalise what the **client's** `package.json` declares, and a root-only dependency is in none of them — so it is **bundled**, silently. For a CJS package inlined into an ESM bundle that is fatal: esbuild leaves a -`Dynamic require of "path" is not supported` shim that throws at *import* time, +`Dynamic require of "path" is not supported` shim that throws at _import_ time, so the binary dies before it parses a flag (`proper-lockfile`, #2082). `undici` (#2067) is the worse variant, because it is `import()`ed lazily: the @@ -293,7 +296,7 @@ file. ### Why React-rendering packages are the exception An externalised package resolves its own `react` from wherever npm placed -**it** — beside a React satisfying *that package's* peer range, which is looser +**it** — beside a React satisfying _that package's_ peer range, which is looser than ours in every case here. `ink-form` and `ink-scroll-view` declare `">=18"`, so a consumer's React 18 satisfies them and hoists them while our React 19 nests underneath: the bundle renders through one React, those packages call hooks on @@ -303,8 +306,8 @@ another, and the TUI crashes on the first hook (#1952). a `createRequire` banner). ⚠️ **Never justify that exemption by a peer range** — it briefly read "its `">=19"` peer keeps npm honest", which is false: a consumer pinning React 19.0 satisfies `">=19"` while a narrower range of ours nests -underneath. What makes it safe is the *root `react` range staying open to the -whole major*, so npm can dedupe. `clients/tui/__tests__/tsupConfig.test.ts` +underneath. What makes it safe is the _root `react` range staying open to the +whole major_, so npm can dedupe. `clients/tui/__tests__/tsupConfig.test.ts` enforces the whole split, the exemption included. ### Why a version skew is worth aligning rather than working around @@ -332,7 +335,7 @@ an `overrides` entry in that install (see the next section). ### Why `overrides` beats `npm audit fix` `tsup@8.5.1` declares `esbuild: ^0.27.0`, and the advisory covers -`0.27.3 - 0.28.0` with `0.27.7` the last 0.27.x — so there is no *upward* escape +`0.27.3 - 0.28.0` with `0.27.7` the last 0.27.x — so there is no _upward_ escape inside that range, and `npm audit fix` "resolves" it by silently **downgrading** to `0.27.2` across three installs (~700 lines of lockfile churn for a low-severity dev-only advisory; tried and reverted in #2058). The override forces one deduped diff --git a/.claude/skills/pr-flow/SKILL.md b/.claude/skills/pr-flow/SKILL.md index 6aa402a1db..fb17c0ee76 100644 --- a/.claude/skills/pr-flow/SKILL.md +++ b/.claude/skills/pr-flow/SKILL.md @@ -43,46 +43,23 @@ answer "who has this?", and an assigned issue whose card still says `Todo` tells the board nobody has started. `@me` resolves to whoever `gh` is authenticated as, so an agent assigns the maintainer it is working for. -Run the whole block. It is the assignment, the card move, and a check; **the -step is done only when the last line prints `card: In Progress`.** +Run the chained command. **The step is done only when it prints +`card: In Progress`** — the `&&` makes that line unreachable when the +assignment fails: ```sh -N=<ISSUE_NUMBER>; STATUS="In Progress" -BOARD=28 # 11 for a v1 issue — board #11 has the same column names -ASSIGNED= -gh issue edit "$N" --repo modelcontextprotocol/inspector --add-assignee @me \ - && ASSIGNED=1 || echo "assignment failed — this step is NOT done" >&2 - -# Every id is resolved BY NAME at run time, so none is copied from /board-ops -# and an option recreated after a deletion (its hazard) still resolves. -PROJECT_ID= FIELD_ID= OPTION_ID= ITEM_ID= # no id survives a failed lookup -PROJECT_ID=$(gh project view "$BOARD" --owner modelcontextprotocol --format json --jq .id) -FIELDS=$(gh project field-list "$BOARD" --owner modelcontextprotocol --format json) && - FIELD_ID=$(jq -r '.fields[] | select(.name=="Status") | .id' <<<"$FIELDS") && - OPTION_ID=$(jq -r --arg s "$STATUS" '.fields[] | select(.name=="Status") - | .options[] | select(.name==$s) | .id' <<<"$FIELDS") -# The card is found from the issue, not from a board listing (see /board-ops). -card() { - gh api graphql -F n="$N" -f query='query($n:Int!){ - repository(owner:"modelcontextprotocol",name:"inspector"){issue(number:$n){ - projectItems(first:100){nodes{id project{id} - fieldValueByName(name:"Status"){... on ProjectV2ItemFieldSingleSelectValue{name}}}}}}}' \ - | jq -r --arg p "$PROJECT_ID" '.data.repository.issue.projectItems.nodes[] - | select(.project.id==$p) | "\(.id) \(.fieldValueByName.name // "(none)")"' -} -ITEM_ID=$(card | cut -d' ' -f1) -if [ -n "$PROJECT_ID" ] && [ -n "$FIELD_ID" ] && [ -n "$OPTION_ID" ] && [ -n "$ITEM_ID" ]; then - gh project item-edit --project-id "$PROJECT_ID" --id "$ITEM_ID" \ - --field-id "$FIELD_ID" --single-select-option-id "$OPTION_ID" >/dev/null -else - echo "lookup failed (project='$PROJECT_ID' field='$FIELD_ID' option='$OPTION_ID' item='$ITEM_ID') — nothing edited" >&2 -fi -NOW=$(card | cut -d' ' -f2-) -[ "$NOW" = "$STATUS" ] && [ -n "$ASSIGNED" ] && echo "card: $NOW" \ - || echo "card is '$NOW', assigned='${ASSIGNED:-no}' — this step is NOT done" >&2 +gh issue edit <ISSUE_NUMBER> --repo modelcontextprotocol/inspector --add-assignee @me \ + && npm run board:status -- --issue <ISSUE_NUMBER> --status "In Progress" # add --board 11 for a v1 issue ``` -An issue with no card on board `$BOARD` fails the lookup; board it there first with +The script (`scripts/board-card-status.mjs`, #2558) resolves every id by name +at run time — so nothing is copied from `/board-ops` and an option recreated +after a deletion (its hazard) still resolves — finds the card from the issue +rather than a board listing, edits it, re-reads it, and prints `card: <Status>` +only when the re-read confirms the move. Any other output means the step is +NOT done. + +An issue with no card on that board fails the lookup; board it there first with `/issue-create`'s card step rather than skipping the move. ## 2. Branch @@ -116,12 +93,47 @@ previous one, not all cut from `v2/main`. ## 3. Sign off every commit -**The DCO check is a hard merge gate.** The [probot DCO -app](https://probot.github.io/apps/dco/) fails the PR unless each commit carries -a `Signed-off-by: Name <email>` trailer whose name **and** email match either the -commit's author or its committer. Its only exemptions are merge commits and +**The `DCO` check fails the PR on any unsigned commit.** It is this repo's own +job (`.github/workflows/dco.yml` → `scripts/dco-check.mjs`, #2566), run on every +v2 PR, stacked PRs included, and it requires each commit to carry a `Signed-off-by: Name <email>` +trailer whose name **and** email match either the commit's author or its +committer (case-insensitively). One relaxation: a commit GitHub itself +committed (`GitHub <noreply@github.com>`, as on a squash merge, whose author +name GitHub takes from the profile) passes on an author **email** match alone. +Its only exemptions are merge commits and bot-authored commits; there is no partial credit — one unsigned commit out of six -fails the whole check. +fails the whole check, and the job's output names each offending commit and the +repair below. + +⚠️ **It is a merge gate because it is a _required_ status check** — a +ruleset setting, not something the workflow file can declare. The +`v2/main - DCO` repository ruleset (#2621) requires `DCO` on every PR into +`v2/main`, pinned to the GitHub Actions app (integration `15368`) so a commit +status someone posts by hand under the same name cannot satisfy it; it also +blocks deleting or force-pushing `v2/main`, which the push backstop below +relies on. Repository admins can bypass it. A **stacked** PR's check runs but +gates nothing until the PR is retargeted to `v2/main`. The job runs on +`pull_request`, from the PR's own ref, so it reports on every v2 PR — stacked +ones included, any `v2/**` base — with no wait for a +milestone merge (#2616). A second job, `DCO (v2/main push)`, +re-runs the check over every push that lands on `v2/main` — a backstop for +anything that merged without a passing PR check — so a red there means an +unsigned commit is already on the branch. The probot DCO app +it replaced was never required, so when the app was suspended its check simply +stopped appearing (after #1981) and nothing went red for two months. If the +`DCO` check is ever missing from a PR, treat that as the outage it is. + +**Check before you push.** `npm run local:gate` already does it: its +`local:dco` stage runs the same script over `origin/v2/main..HEAD` (minus +anything already on `origin/main`), right after +`local:validate`. On a **stacked** branch that range covers the whole stack, +parents included, which is stricter than the child PR's own check (that one +runs against the parent's branch). To check on its own, against the range the +PR will show, pass the PR's actual base: + +```sh +npm run dco:check -- --base origin/v2/main # or origin/<parent branch> when stacked +``` **Prevent it with `git commit -s`.** Two things that look like automation and are not: @@ -138,16 +150,22 @@ not: **Repairing already-pushed commits** means rewriting them: ```sh -git rebase HEAD~<n> --signoff +git rebase --rebase-merges --signoff origin/v2/main # the base the PR targets git push --force-with-lease ``` -Use `--force-with-lease` rather than `--force`, and only rewrite when you are the -sole author and nobody else has based work on the branch. The two apparent -alternatives are not alternatives: the app's empty "remediation commit" flow -requires `allowRemediationCommits.individual` and this repo ships no -`.github/dco.yml`, so it runs disabled; and the override button anyone with write -access sees only silences the check without anyone certifying anything. +⚠️ **On a stacked branch, rebase against the parent branch, never +`origin/v2/main`.** A `--signoff` rebase onto `origin/v2/main` rewrites every +parent commit too, which forks the child from its parent and breaks the stack. +If the unsigned commit is in the parent, repair the parent's own branch first, +then rebase the child onto the repaired parent. + +`--rebase-merges` keeps any merge commit on the branch — without it the rebase +flattens them, silently dropping a conflict resolution that lives only in the +merge. Use `--force-with-lease` rather than `--force`, and only rewrite when you are the +sole author and nobody else has based work on the branch. There is no +remediation-commit or override path: the check reads each commit's own message, +so a later commit cannot certify an earlier one. The signoff is a [Developer Certificate of Origin](https://developercertificate.org/) assertion made in **your own name**. It @@ -241,25 +259,21 @@ you. Re-shoot rather than shipping one that "mostly" shows the change. ### 5c. Upload -To host them, upload to GitHub's attachment endpoint with your `gh` token. Two -mechanics, both of which bite: - -- The parameters go in the **query string**, with the raw bytes as the body. A - JSON body fails with a misleading "Invalid name for request". -- ⚠️ **Do not put the token in argv.** `-H "Authorization: token $(gh auth -token)"` puts your credential in curl's command line, where any local user or - process can read it off the process table while the upload runs (Copilot). - Feed it through `--config -` instead: curl reads its options from stdin, so - the token never becomes an argument. +To host them, upload to GitHub's attachment endpoint with the script +(`scripts/pr-upload-screenshot.mjs`, #2558); it prints the hosted URL to embed: ```sh -printf 'header = "Authorization: token %s"\n' "$(gh auth token)" | curl -sS --config - \ - -X POST --data-binary @pr-screenshots/tools-tab-after.png \ - "https://uploads.github.com/user-attachments/assets?repository_id=<REPO_ID>&name=tools-tab-after.png&content_type=image/png" +npm run pr:upload -- --file pr-screenshots/tools-tab-after.png ``` -(The token is still in the shell's environment and in `printf`'s _stdin_, which -is not world-readable the way `/proc/<pid>/cmdline` is.) +Two mechanics it handles, both of which bite when done by hand: + +- The parameters go in the **query string**, with the raw bytes as the body. A + JSON body fails with a misleading "Invalid name for request". +- ⚠️ **The token never goes in argv.** A `-H "Authorization: token $(gh auth +token)"` puts your credential in a command line, where any local user or + process can read it off the process table while the upload runs (Copilot). + The script sends it only as a request header. ## 6. Open the PR @@ -282,65 +296,33 @@ is only a cross-reference — it will **not** create a hard link or close the is on merge. Keep it anyway, so the issues close if/when `v2/main` reaches `main`. **So link the PR to its issue explicitly, right after creating it.** The -`addCloseIssueReferences` GraphQL mutation adds a manual closing reference, the -same link as the UI's **Development** sidebar, and it works whatever the base -branch. It is what puts the PR in the card's **Linked pull requests** field, -which the board shows as a column in table views and as a chip on kanban cards. -Without it a v2 card shows no PR at all. +script (`scripts/pr-link-issue.mjs`, #2558) runs the `addCloseIssueReferences` +GraphQL mutation — a manual closing reference, the same link as the UI's +**Development** sidebar, working whatever the base branch — and verifies it by +reading the PR's `closingIssuesReferences` back. It is what puts the PR in the +card's **Linked pull requests** field, which the board shows as a column in +table views and as a chip on kanban cards. Without it a v2 card shows no PR at +all. ```sh -ISSUE_ID=$(gh api graphql -F n=<ISSUE_NUMBER> -f query='query($n:Int!){ - repository(owner:"modelcontextprotocol",name:"inspector"){issue(number:$n){id}}}' \ - --jq .data.repository.issue.id) -PR_ID=$(gh pr view <N> --repo modelcontextprotocol/inspector --json id --jq .id) -gh api graphql -f query='mutation($i:ID!,$p:[ID!]!){ - addCloseIssueReferences(input:{issueId:$i, pullRequestIds:$p}){clientMutationId}}' \ - -f i="$ISSUE_ID" -f p="$PR_ID" - -# Verify: the PR should list the issue. -gh api graphql -F n=<N> -f query='query($n:Int!){ - repository(owner:"modelcontextprotocol",name:"inspector"){pullRequest(number:$n){ - closingIssuesReferences(first:10){nodes{number}}}}}' \ - --jq '[.data.repository.pullRequest.closingIssuesReferences.nodes[].number]' +npm run pr:link -- --pr <N> --issue <ISSUE_NUMBER> # prints linked: … only on a verified link ``` The link does not change how the issue closes on a v2 merge; that is still -step 9. `removeCloseIssueReferences` takes the same input and undoes the link. +step 9. The `removeCloseIssueReferences` mutation takes the same input and +undoes the link. **Then move the card to In Review. Step 6 is done only when the PR is linked -_and_ the card says `In Review`.** It is step 1's block with a different -column and no assignment. Run it in full and check that the last line prints -`card: In Review`: +_and_ the card says `In Review`.** Same script as step 1, different column — +and it takes the **issue** number, not the PR's: ```sh -N=<ISSUE_NUMBER>; STATUS="In Review" # the ISSUE number, not the PR's -BOARD=28 # 11 for a v1 issue — board #11 has the same column names - -PROJECT_ID= FIELD_ID= OPTION_ID= ITEM_ID= # no id survives a failed lookup -PROJECT_ID=$(gh project view "$BOARD" --owner modelcontextprotocol --format json --jq .id) -FIELDS=$(gh project field-list "$BOARD" --owner modelcontextprotocol --format json) && - FIELD_ID=$(jq -r '.fields[] | select(.name=="Status") | .id' <<<"$FIELDS") && - OPTION_ID=$(jq -r --arg s "$STATUS" '.fields[] | select(.name=="Status") - | .options[] | select(.name==$s) | .id' <<<"$FIELDS") -card() { - gh api graphql -F n="$N" -f query='query($n:Int!){ - repository(owner:"modelcontextprotocol",name:"inspector"){issue(number:$n){ - projectItems(first:100){nodes{id project{id} - fieldValueByName(name:"Status"){... on ProjectV2ItemFieldSingleSelectValue{name}}}}}}}' \ - | jq -r --arg p "$PROJECT_ID" '.data.repository.issue.projectItems.nodes[] - | select(.project.id==$p) | "\(.id) \(.fieldValueByName.name // "(none)")"' -} -ITEM_ID=$(card | cut -d' ' -f1) -if [ -n "$PROJECT_ID" ] && [ -n "$FIELD_ID" ] && [ -n "$OPTION_ID" ] && [ -n "$ITEM_ID" ]; then - gh project item-edit --project-id "$PROJECT_ID" --id "$ITEM_ID" \ - --field-id "$FIELD_ID" --single-select-option-id "$OPTION_ID" >/dev/null -else - echo "lookup failed (project='$PROJECT_ID' field='$FIELD_ID' option='$OPTION_ID' item='$ITEM_ID') — nothing edited" >&2 -fi -NOW=$(card | cut -d' ' -f2-) -[ "$NOW" = "$STATUS" ] && echo "card: $NOW" || echo "card is '$NOW', not '$STATUS' — this step is NOT done" >&2 +npm run board:status -- --issue <ISSUE_NUMBER> --status "In Review" # add --board 11 for a v1 issue ``` +It prints `card: In Review` only when the post-edit re-read confirms the move; +any other output means this step is NOT done. + Then go straight to step 7. ## 7. Run the Copilot review loop — immediately, every PR @@ -354,71 +336,40 @@ request again if anything was pushed. It stops only on one of the exits in 7c. Only the GraphQL `requestReviews` mutation with the Copilot **bot id** works — REST, `gh pr edit --add-reviewer`, `userIds`, and `copilot-swe-agent` all fail or -silently drop. +silently drop. `scripts/pr-review-request.mjs` (#2558) owns that mutation and +the bot id: ```sh -PR_ID=$(gh pr view <N> --repo modelcontextprotocol/inspector --json id --jq .id) -gh api graphql -f query=' - mutation($pr:ID!,$bot:[ID!]!) { - requestReviews(input:{pullRequestId:$pr, botIds:$bot, union:true}) { - pullRequest { id } - } - }' -f pr="$PR_ID" -f bot='BOT_kgDOCnlnWA' +npm run pr:review-request -- --pr <N> ``` +It prints `requested: Copilot review on PR #<N>` on success and the +`pr:review-wait` invocation to run next. + ### 7b. Wait for it — review posted, or session ended A round ends one of two ways: Copilot **posts a review**, or its **pending request disappears without one** — it failed, or occasionally has nothing to say and posts nothing. Waiting only for the review hangs forever on the second -case, so the wait watches both, plus a hard cap. **Put it in one backgrounded -loop that exits when the round resolves, and wait for its notification** rather -than re-fetching once per turn; a review is remote state the harness cannot -observe, which is exactly the exception described in [Waiting on long-running -work](../../../AGENTS.md#waiting-on-long-running-work). +case, so the wait watches both, plus a hard cap. `scripts/pr-review-wait.mjs` +(#2558) implements exactly that loop. **Background it and wait for its +notification** rather than re-fetching once per turn; a review is remote state +the harness cannot observe, which is exactly the exception described in +[Waiting on long-running work](../../../AGENTS.md#waiting-on-long-running-work). ```sh -EXPECTED=1 # the review COUNT you are waiting to reach — see below -DEADLINE=$(( $(date +%s) + 1500 )) # 25 min; rounds normally land in 2–10 -count() { - # Capture first, so a gh failure stops the loop instead of being swallowed by - # a pipeline. --slurp cannot be combined with --jq, hence the separate jq. - raw=$(gh api --paginate --slurp \ - repos/modelcontextprotocol/inspector/pulls/<N>/reviews) || { - echo "gh api failed ($?) — not retrying blind" >&2; exit 1; } - n=$(jq '[.[][] | select(.user.login | startswith("copilot-pull-request-reviewer"))] | length' <<<"$raw") || { - echo "jq failed ($?) on an unexpected response shape" >&2; exit 1; } - case $n in '' | *[!0-9]*) echo "not a count: '$n'" >&2; exit 1 ;; esac -} -pending() { - p=$(gh api graphql -f query='{repository(owner:"modelcontextprotocol",name:"inspector"){pullRequest(number:<N>){reviewRequests(first:20){nodes{requestedReviewer{... on Bot{login} ... on User{login}}}}}}}' \ - --jq '[.data.repository.pullRequest.reviewRequests.nodes[].requestedReviewer.login // empty | select(test("copilot";"i"))] | length') || { - echo "gh graphql failed ($?)" >&2; exit 1; } -} -while :; do - count; [ "$n" -ge "$EXPECTED" ] && { echo "ROUND=posted"; break; } - pending - if [ "$p" = 0 ]; then - sleep 30; count # the request can clear a beat before the review is visible - [ "$n" -ge "$EXPECTED" ] && echo "ROUND=posted" || echo "ROUND=ended-without-review" - break - fi - [ "$(date +%s)" -ge "$DEADLINE" ] && { echo "ROUND=timed-out"; break; } - sleep 30 -done +npm run pr:review-wait -- --pr <N> --expected <K> # --timeout-minutes 25 is the default ``` -`EXPECTED` is the review **count** you are waiting to reach, so it is `1` only -on the first round — on round two the first round's review is still there and an -existence check returns immediately. `sleep 30` is the remote-API floor the rule -above sets. **Every step that can fail exits the loop rather than -retrying.** Piping the count straight into `awk` would make an auth or API error -read as a count of `0`; and a `jq` failure on an unexpected shape leaves `n` -empty, whereupon `[ "" -ge 1 ]` exits non-zero, `break` never fires, and the job -sleeps and retries forever — the same unbounded wait, reached from the other -end. A background task that can never succeed is worse than one that never -started, because it looks like progress. On `ROUND=posted`, give the inline -comments a further ~60s; they arrive late (see step 8). +Its last line is the outcome: `ROUND=posted`, `ROUND=ended-without-review`, or +`ROUND=timed-out` (all exit 0). A nonzero exit means the wait itself failed — +a `gh` failure (the script never retries blind on one, for the reason its +header records), a malformed response, or a bad argument — not a round outcome. + +`--expected` is the review **count** to reach, so it is `1` only on the first +round — on round two the first round's review is still there and an existence +check would return immediately. On `ROUND=posted`, give the inline comments a +further ~60s; they arrive late (see step 8). ### 7c. Decide: another round, or stop @@ -463,18 +414,20 @@ why (which exit fired), and report the same in your reply to the user. reply is what makes resolving it defensible. ```sh - # Fetch the round's comments by REVIEW id — the unpaginated /reviews listing - # hides later rounds behind your own replies. - # --paginate: this endpoint returns 30 per page, and a round you only half - # fetch is a round you only half answer. - gh api --paginate repos/modelcontextprotocol/inspector/pulls/<N>/reviews/<REVIEW_ID>/comments \ - --jq '.[]|"\(.id) \(.path):\(.line)\n\(.body)"' - - # Reply into one thread, keyed by the comment id from above. + # Fetch the latest Copilot round — review header + body, then every inline + # comment as `COMMENT=<id> <path>:<line>` with its body. Pass + # --review <REVIEW_ID> to fetch an earlier round instead. + npm run pr:review-fetch -- --pr <N> + + # Reply into one thread, keyed by the COMMENT= id from above. gh api repos/modelcontextprotocol/inspector/pulls/<N>/comments/<COMMENT_ID>/replies \ -f body='Fixed in <sha> — …' ``` + The script fetches by **review id** and paginates, because the unpaginated + `/reviews` listing hides later rounds behind your own replies, and a round + you only half fetch is a round you only half answer (#2558). + - ⚠️ **Then mirror the round at PR level, in addition — never instead.** Inline replies go hidden once the fix is pushed, because the threads become outdated, so a summary comment is what keeps the round readable afterwards. It does diff --git a/.claude/skills/pre-push-gate/SKILL.md b/.claude/skills/pre-push-gate/SKILL.md index 2ea25cd055..cb31fc1cde 100644 --- a/.claude/skills/pre-push-gate/SKILL.md +++ b/.claude/skills/pre-push-gate/SKILL.md @@ -98,18 +98,36 @@ stale install otherwise passes every check and fails later as a behavioral test reporting the *old* dependency's behavior as a product bug (#2494). Don't "fix" that test. +### `local:dco` + +A commit in `origin/v2/main..HEAD` (minus anything already on `origin/main`) +has no `Signed-off-by:` trailer matching its +author or committer (#2616). The output names each commit and prints the +repair. Repair and prevention (`git commit -s`, the `--signoff` rebase) are +step 3 of `/pr-flow`. + +If it lists commits you never made, check which case you are in: + +- **A stacked branch** (cut from another feature branch): `origin/v2/main..HEAD` + covers the whole stack, parent commits included. Fetching changes nothing. + An unsigned parent commit is repaired on the parent's own branch first, and + the child is then rebased onto the repaired parent (`/pr-flow` step 3). +- **Not stacked**: your `origin/v2/main` is probably behind what the branch was + cut from. `git fetch origin v2/main` and re-run. + ### `verify:action-pins` A job that holds a credential (`id-token`/`packages: write`, a non-default secret, or it builds an artifact such a job downloads) runs an action that is not SHA-pinned (#2484). Pin it the way its neighbours are — -`owner/repo@<40-hex sha> # vX.Y.Z` — resolving both from one lookup: +`owner/repo@<40-hex sha> # vX.Y.Z` — with the script +(`scripts/action-pin-resolve.mjs`, #2558), which resolves the SHA and the +exact-version comment from the **same** tag lookup (the guard is offline and +cannot check the two agree): ```sh -REPO=actions/checkout; TAG=v7 -SHA=$(gh api "repos/$REPO/commits/$TAG" --jq .sha) -gh api --paginate "repos/$REPO/tags?per_page=100" \ - --jq ".[] | select(.commit.sha==\"$SHA\") | .name" | grep -E '^v[0-9]+\.[0-9]+\.[0-9]+$' | sort -V | tail -1 +npm run action:resolve-pin -- --repo actions/checkout --tag v7 +# → uses: actions/checkout@<sha> # v7.x.y ``` If a job started failing because it gained a secret or a scope, that is the diff --git a/.claude/skills/project-structure/SKILL.md b/.claude/skills/project-structure/SKILL.md index c7e0e54bce..8928a71a65 100644 --- a/.claude/skills/project-structure/SKILL.md +++ b/.claude/skills/project-structure/SKILL.md @@ -24,7 +24,9 @@ inspector/ │ └── launcher/ The `mcp-inspector` bin; dispatches to web/cli/tui in-process ├── core/ Shared code, consumed via the `@inspector/core` alias (no package.json) ├── test-servers/ Composable MCP test servers + JSON configs used by tests and by hand -├── scripts/ Root build/verify tooling: install cascade, smokes, the verify:* guards +├── scripts/ Root build/verify tooling (install cascade, smokes, the verify:* guards) +│ and maintainer-workflow helpers (the pr:*, board:*, advisory:*, +│ release:notes, release:tag and action:resolve-pin aliases) ├── docs/ Task-oriented guides (see docs/README-style index in the root README) ├── specification/ Design/build specifications └── AGENTS.md The rules contract — read this before changing anything @@ -37,22 +39,26 @@ an MCP server, the request/response lifecycle, and a set of state stores. | Directory | Owns | | --- | --- | -| `core/mcp/` | `InspectorClient`, transports, state stores, config import, URI templates, task/subscription/App-elicitation protocol helpers | +| `core/mcp/` | `InspectorClient`, transports, state stores, config import, URI templates, subscription/App-elicitation protocol helpers | | `core/mcp/node/` | Node stdio transport factory; `proxyFetch.ts` (the shared HTTPS_PROXY/NO_PROXY fetch) | | `core/mcp/remote/` | Browser HTTP/SSE transport + remote logger/fetch, and (under `node/`) the Hono backend it talks to | | `core/mcp/state/` | The stores `core/react/` hooks read | | `core/auth/` | OAuth end to end — providers, discovery, storage, endpoint overrides, scopes, revocation, mid-session recovery — split into isomorphic logic plus `browser/`, `node/` and `remote/` backends | | `core/auth/node/` | Node OAuth storage + loopback callback server, **and** the `SecretStore` backends (keychain / file / memory) and their selection policy | +| `core/extension/<name>/` | Host-side adapters for MCP extensions whose protocol an upstream SDK owns. `tasks/` wraps `@modelcontextprotocol/ext-tasks` with what that package leaves to its host: the raw `rawDispatch` channel, progress routing, error identity, and task-view conversions. Imports nothing from `InspectorClient`; host state arrives through narrow interfaces | +| `core/cli/` | Node-only surface the one-shot CLI and `mcpdo` both run on (#2461): `error-handler` (exit codes, `CliExitCodeError`, the error envelope), the method handlers and their option/format types (`handlers/`), the interactive OAuth connect flow (`cliOAuth`, `cli-oauth-navigation`) and output helpers (`style`, `utils/awaitable-log`). One-shot-only output (`emit-result`, `consume-outcome`, `servers-write`, …) stays in `clients/cli/src` | | `core/client/` | Install-level client config (`client.json`): browser-safe parse plus Node load/save, remote backend, secrets, runner | | `core/json/` | JSON + parameter/argument conversion; the schema normalizations all three form builders share (nullable unions, root composition) and the tool-schema portability lint | | `core/react/` | React hooks over the state stores — consumed by both the web and TUI React trees. Every subscription reads its snapshot **during render** via `useSyncExternalStore` (#1955); `useStoreSnapshot.ts` caches the fresh-value-per-read getters | -| `core/node/` | Node-only helpers: version reader, host normalization/detection | +| `core/node/` | Node-only helpers: version reader, host normalization/detection, browser opener (`openUrl`) | | `core/storage/` | File I/O helpers used by the OAuth persist backends | | `core/logging/` | Silent pino logger singleton | `core/` is isomorphic (browser + Node) and has **no `package.json`** — it is not published on its own. Its tests live in `clients/web/src/test/core/`, and its -browser-consumed runtime is inside the web coverage gate. +browser-consumed runtime is inside the web coverage gate. **`core/cli/` is the +exception**: it is Node-only, its tests are `clients/cli/__tests__/`, and it is +gated by `clients/cli`'s coverage run instead. ## `clients/web/server/` — the Node backend diff --git a/.claude/skills/release/SKILL.md b/.claude/skills/release/SKILL.md index ceb73d77f9..d076c308e1 100644 --- a/.claude/skills/release/SKILL.md +++ b/.claude/skills/release/SKILL.md @@ -1,6 +1,6 @@ --- name: release -description: "Cut an Inspector v2 release — two PRs and then a GitHub Release. PR 1 puts the npm audit, any fixes it forces, and the version bump on v2/main; PR 2 merges v2/main into main and is smoke-tested from the production build with a ledger artifact for the maintainers; the maintainer then tags and publishes through the GitHub UI. Also covers the v1 line and what the publish jobs gate on." +description: "Cut an Inspector v2 release — two PRs and then a GitHub Release. PR 1 puts the npm audit, any fixes it forces, and the version bump on v2/main; PR 2 merges v2/main into main and is smoke-tested from the production build with a ledger artifact for the maintainers; the maintainer then drafts the Release with a script and publishes it. Also covers the v1 line and what the publish jobs gate on." disable-model-invocation: true --- @@ -206,42 +206,175 @@ the artifact the maintainers approve the merge on. ## 3. Tag and publish the Release -**Normally this is done by a maintainer through the GitHub UI**, after PR 2 has -merged: *Releases → Draft a new release → Choose a tag → type the bare `x.y.z` -→ Create new tag on publish*, with **Target: `main`**, then generate the notes -and publish. Publishing the Release is what fires the `publish` and -`publish-github-container-registry` jobs. +### 3a. Assemble the notes and draft the Release -The equivalent by hand, for when the UI is not an option — derive the tag from -the version that just landed rather than typing one, since a hard-coded tag is -either already taken (so `git tag` aborts) or, worse, wrong: +A release's notes have four parts, in this order: + +1. **What's Changed.** GitHub's generated list of every PR since the previous tag. +2. **The smoke-ledger line**, linking the artifact from 2b. +3. **`## Known issue`** (or `issues`), only when there is one. It names the + issue, who is affected and the workaround. Deciding what counts as a known + issue is a maintainer judgment, so it is written by hand and passed in, + never generated. +4. **`## Thanks for helping us improve`.** Credit to the community members whose + issues the release addresses. GitHub adds everyone `@`-mentioned in a + release body to that release's **Contributors** avatar strip, so the people + credited here appear there too (confirmed on 2.9.0). + +**`npm run release:notes` assembles all four and creates the Release** +(`scripts/release-notes.mjs`, #2550). It derives the version from `origin/main` +after an explicit fetch, picks the previous **stable** tag itself (this repo +also carries `-rc.N`, `-hotfix`, `-amended` and `v2-alpha-1` tags, any of which +would drop changes from the list), and asks the same `releases/generate-notes` +API the UI's *Generate release notes* button uses. Run it twice, then publish in 3b: + +```sh +ARGS=(--merge-branch "v2/chore/milestone-merge-vX.Y.Z" --ledger-url "https://claude.ai/artifact/<id>") +# Add one --known-issue "<markdown paragraph>" per known issue, if any. + +npm run release:notes -- "${ARGS[@]}" # 1. preview: prints the notes, creates nothing +npm run release:notes -- "${ARGS[@]}" --draft # 2. creates the Release as a DRAFT +``` + +Read the preview, then create the draft and read it again on the Releases page. +The preview prints the notes on stdout and its progress on stderr, so +`> notes.md` captures just the notes. + +The rules the helper applies: + +- **An issue counts when a listed PR closes it**, through either the manual + closing link (`closingIssuesReferences`, followed across every page) or a + closing keyword (`close`/`fix`/`resolve` in any tense, optional colon) before + a bare `#N` in the PR body. So an issue older than the release still counts + when this release closed it. A cross-repo `owner/repo#N` does not count. +- **Maintainers and bots are excluded by permission, not by name.** A + maintainer is anyone with `admin`, `maintain` or `write` on the repo. Bot + authors are dropped, which covers the issues the SDK-watch and Dependabot + sweeps file. On a public repo, anyone without a role reads as `read`, so they + are credited. +- **One line per person, most issues first:** `* @user (#1, #2, …)`. +- **"Addresses", not "fixes."** The credited issues include feature requests. +- **The section is left out** when no community reporter remains, as with 2.1.0. +- ⚠️ **Any API failure aborts the whole run.** A failed permission lookup never + reads as "community", which could credit a maintainer, and a rate-limited PR + or issue lookup is never skipped. A partial Thanks list is never published, so + on a rate limit, wait for the reset and run it again. + +To regenerate an **older** release's notes (to check them, or after editing a +PR body), pass `--version x.y.z`, and optionally `--previous-tag`. This works +for preview only. `--draft` and `--publish` refuse `--previous-tag`, and any +version other than the one on `origin/main`, because the Release would carry +the wrong range of changes or attach to the wrong tree. +Previewing `--version 2.9.0` with 2.9.0's ledger and known issue reproduces its +published notes exactly. + +### 3b. Publish the Release + +**Publishing is the deliberate maintainer action.** It fires `package` → +`publish` (npm) and `publish-github-container-registry`, so the helper never +does it implicitly. Publish the draft from 3a on the Releases page (*Edit → +Publish release*, with **Set as the latest release** checked), or from the CLI: ```sh -git fetch origin main -VERSION=$(git show origin/main:package.json | node -p "JSON.parse(require('fs').readFileSync(0)).version") -echo "$VERSION" # sanity-check before tagging -git tag "$VERSION" origin/main && git push origin "$VERSION" -# then draft & publish a GitHub Release for that tag → triggers `publish` +git fetch origin main # FETCH_HEAD, not origin/main: the tracking ref can lag (see release-tag.mjs) +VERSION=$(git show FETCH_HEAD:package.json | node -p "JSON.parse(require('fs').readFileSync(0)).version") +gh release edit "$VERSION" --repo modelcontextprotocol/inspector --draft=false --latest ``` -⚠️ **Tag `origin/main`, not your local `HEAD`.** `git checkout main && git pull` -resolves through whatever merge-or-rebase strategy you have configured, so a -divergent local `main` can quietly produce or replay local commits. Tagging -`HEAD` there tags a commit that is not on `origin/main`, and `git push origin -<tag>` pushes only the tag — leaving a release whose commit was never published. -The UI path avoids this by construction: the target is `main` itself. +A draft has no tag yet. GitHub creates the bare `x.y.z` tag **on publish**, at +`main`'s head at that moment, so publish only while `main` is still the merge +commit PR 2 landed. `npm run release:notes -- "${ARGS[@]}" --publish` creates and +publishes in one step, for when the notes were already reviewed in a preview. + +The helper passes the bare `$VERSION` as the tag and `--target main`, so the tag +name and the target are right by construction. + +If the tag has to exist before the Release (for example, to pin the commit +before drafting), push it first with `scripts/release-tag.mjs` (#2558). It +derives the tag from the version that just landed rather than taking one as +input, since a hard-coded tag is either already taken (so `git tag` aborts) or, +worse, wrong. A Release created afterwards attaches to that existing tag: + +```sh +npm run release:tag # dry run: prints what would be tagged +npm run release:tag -- --push # tags origin/main's SHA and pushes the tag +# then 3a and 3b; the Release attaches to this tag +``` + +⚠️ **It tags `origin/main`, not your local `HEAD`.** `git checkout main && git +pull` resolves through whatever merge-or-rebase strategy you have configured, +so a divergent local `main` can quietly produce or replay local commits. +Tagging `HEAD` there tags a commit that is not on `origin/main`, and `git push +origin <tag>` pushes only the tag — leaving a release whose commit was never +published. The script resolves the SHA from `origin/main` after an explicit +fetch; the UI path avoids this by construction, since the target is `main` +itself. ⚠️ **No `v` prefix.** This repo's release tags are bare `x.y.z` — which is why -the command above tags `$VERSION` and not `v$VERSION`, and why the tag typed -into the UI carries no prefix either. npm's own `tag-version-prefix` defaults to -`v` and the repo sets no `.npmrc`, so a bare `npm version` would have produced a -mismatched tag; tagging by hand is what keeps it right. (The workflow's assert -step strips a leading `v` before comparing, so a `v`-prefixed tag would still -publish — it would just be inconsistent with every previous release.) +the script tags `$VERSION` and not `v$VERSION`, and why the tag typed into the +UI carries no prefix either. npm's own `tag-version-prefix` defaults to `v` and +the repo sets no `.npmrc`, so a bare `npm version` would have produced a +mismatched tag. (The workflow's assert step strips a leading `v` before +comparing, so a `v`-prefixed tag would still publish — it would just be +inconsistent with every previous release.) The release's target commit selects which workflow runs, so this only publishes when a release is cut from a commit carrying the v2 workflow. +**Editing a published Release's notes is safe.** Every tag's `main.yml` +triggers only on `release: types: [published]` (checked for every tag from 2.0.0 +through 2.9.0), so an `edited` event never re-runs publishing. Fixing a typo or +adding a known issue after the fact needs no ceremony. + +### 3c. If the release run fails + +**Check npm for the exact version before anything else**, and never infer +from `dist-tags`. A freshly published version sits in **Validating** (npm's +automated review, shown on npmjs.com) for a few minutes, and during that window +`dist-tags` still shows the previous `latest`. That happened on 2.9.0. So wait +out validation, then query the version itself: +`npm view @modelcontextprotocol/inspector@$VERSION version --prefer-online`. +**Only an exact-version 404 means nothing was published.** Re-cutting a version +npm already owns cannot succeed, because the version number is immutable. + +⚠️ **A release event runs the workflow from the tag's commit, not from +`main`.** Re-running a failed job therefore re-runs the same broken step. The +fix has to reach `main`, and the Release has to be re-cut at that commit. This +is what 2.9.0 needed (#2551): the first run's `publish` passed a bare +`release-tarball/…tgz` path, which npm read as a GitHub `owner/repo` shorthand. + +1. **Fix it on `v2/main`** through an ordinary PR, never on the merge branch. +2. **Merge `v2/main` into `main`** in a new milestone-merge PR. Its only diff + against the released `main` should be the fix. +3. **Delete the Release *and* its tag.** ⚠️ Deleting a Release in the UI leaves + the tag behind. While the old tag exists, GitHub reuses it, so the re-cut + attaches to the broken commit again, and `generate-notes` reads that commit + too. Delete it on the remote **and locally**: `git push origin :refs/tags/$VERSION` + and `git tag -d "$VERSION"`. A stale local tag (the 3b manual path creates one) + makes the re-tag abort, and later makes `git fetch --tags` refuse to clobber it. + Confirm the remote tag is gone with `gh api repos/$REPO/git/ref/tags/$VERSION` + (expect a 404). +4. **Recreate the Release at the new `main`.** Re-run `npm run release:notes` (3a): + What's Changed now includes the fix PRs, and a fix PR can close a + community-reported issue, so the Thanks section can change too. Pass the + same `--merge-branch`, `--ledger-url` and `--known-issue` arguments as before. +5. **Verify the fix before publishing again, wherever it can be verified.** + Run the gate, and the smoke rows the fix touches, on the new tree. Only a path + that exists solely inside a release run (like #2551's publish step) has the + re-cut run as its first real evidence. Get as close as you can beforehand (a + `--dry-run` with the pinned tool version), and record in the ledger which kind + of evidence each fix has. + +If npm *did* publish and a downstream job failed (the GHCR image, for example), +**do not re-cut**: that would try to publish the same npm version again. +Instead: + +- **A transient failure** (a registry hiccup, a runner fault): re-run **only + the failed job** from the run page. It re-runs at the same tagged commit and + does not touch the npm job. +- **A defect in the tagged workflow** that cannot pass on retry: fix it on + `v2/main` and let it ship with the next release. + ## Why the bump goes on `v2/main` first (#2010) It used to happen on the milestone-merge branch, which is cut from `main` — so diff --git a/.claude/skills/security-advisory/SKILL.md b/.claude/skills/security-advisory/SKILL.md index 404feca08c..f35f7d845d 100644 --- a/.claude/skills/security-advisory/SKILL.md +++ b/.claude/skills/security-advisory/SKILL.md @@ -217,17 +217,14 @@ advisory: a private repo named `<repo>-<ghsa-id>` in the org. ⚠️ **Read `private_fork` FIRST. The POST is not a probe — it CREATES one.** Calling it to "check whether a fork exists" makes one, in the org, which then -needs cleaning up. This was learned the hard way. +needs cleaning up. This was learned the hard way. The script +(`scripts/advisory-fork.mjs`, #2558) encodes that ordering: a bare run only +reads, and it creates a fork only with the explicit `--create` flag and only +when none exists: ```sh -# Idempotency check — does one already exist? -gh api repos/modelcontextprotocol/inspector/security-advisories/<GHSA_ID> \ - --jq '.private_fork // "none"' - -# Only if that printed "none": -gh api -X POST \ - repos/modelcontextprotocol/inspector/security-advisories/<GHSA_ID>/forks -# → 202 Accepted; the fork appears shortly afterwards. +npm run advisory:fork -- --ghsa <GHSA_ID> # read-only probe +npm run advisory:fork -- --ghsa <GHSA_ID> --create # create only if absent ``` ⚠️ **Deleting a private fork needs the `delete_repo` OAuth scope, which a diff --git a/.claude/skills/test-servers/SKILL.md b/.claude/skills/test-servers/SKILL.md index 94cd272644..97ebfd68c2 100644 --- a/.claude/skills/test-servers/SKILL.md +++ b/.claude/skills/test-servers/SKILL.md @@ -177,6 +177,15 @@ Three consequences: server on its **default** config, so the showcase-config table and the protocol-era guidance below do not apply. If the case needs a specific tool set or the modern handler, it is an in-process HTTP test, not this. +- **The one variant is crashing it.** `getCrashableTestMcpServerCommand()` + starts the same entry with `--crashable`, which adds a `crash_server` tool + (plus `collect_elicitation` / `collect_sample`, to have a peer request + pending at the time) so a test can kill the process at a point it picks — + `respond: false` exits with the call in flight, `respond: true` answers + first and exits on an idle session. It is the fixture for mid-session crash + reconciliation (`inspectorClient-crash-reconciliation.test.ts`, #2437). + ⚠️ It exists **only** on this spawned entry: the tool calls `process.exit`, + so wired into an in-process server it would end the test runner instead. - **It is still the build.** The path comes from the module's own resolved location under the alias, so it is `test-servers/build/`, with the same staleness hazard. diff --git a/.claude/skills/testing/SKILL.md b/.claude/skills/testing/SKILL.md index 9f6fd256bf..519e64e2a4 100644 --- a/.claude/skills/testing/SKILL.md +++ b/.claude/skills/testing/SKILL.md @@ -1,6 +1,6 @@ --- name: testing -description: Write, run, place and fix tests in this repo. Use when adding a test, or end-to-end or integration coverage of an MCP operation (listing, paginating or calling tools); when choosing which npm command runs a given suite (web unit, web integration, Storybook, cli, tui, launcher, scripts); when deciding where a new test file belongs — beside its source, under src/test/, or in a client's __tests__/; when a per-file coverage check fails or a v8 ignore is in question; when asking which test tier spawns the built binary rather than importing it; or when rendering, mounting or asserting on Mantine components and their transitions in a test. +description: Write, run, place and fix tests in this repo. Use when adding a test, or end-to-end or integration coverage of an MCP operation (listing, paginating or calling tools); when choosing which npm command runs a given suite (web unit, web integration, Storybook, cli, mcpdo, tui, launcher, scripts); when deciding where a new test file belongs — beside its source, under src/test/, or in a client's __tests__/; when a per-file coverage check fails or a v8 ignore is in question; when asking which test tier spawns the built binary rather than importing it; or when rendering, mounting or asserting on Mantine components and their transitions in a test. disable-model-invocation: false --- @@ -20,7 +20,7 @@ choosing a location or writing a line.** end-to-end or integration coverage of an MCP operation — listing tools, paginating a list, calling a tool, reading a resource — almost always stands a fixture up, so treat that phrasing as the answer to the question above and load -`test-servers` *first*. Grepping for an existing test to copy is not a +`test-servers` _first_. Grepping for an existing test to copy is not a substitute: the fixture you find that way (a config under `test-servers/configs/`) does not tell you which of the three shapes below drives it, or that it can be stale. If the skill then shows the case needs no @@ -45,7 +45,7 @@ ways to depend on one, and they need different halves of that skill: config and no era table apply.** - **An integration or CLI test where stdio is the point → spawned stdio.** `getTestMcpServerCommand()` handed to a stdio transport or to the built CLI, - which spawns it. A subprocess *is* started, but it runs the stdio fixture's + which spawns it. A subprocess _is_ started, but it runs the stdio fixture's **default** config, so there is still nothing to pick — and nothing to override, so if the case needs a specific tool set it is an in-process HTTP test instead. @@ -64,31 +64,32 @@ ways to depend on one, and they need different halves of that skill: What applies to all three is that section's build warning. ⚠️ **Connecting is a strong hint, not the rule.** A few integration tests deliberately hand-roll a JSON-RPC server because the composable fixture - *cannot* produce what they assert on — `inspectorClient-malformed-list.test.ts` + _cannot_ produce what they assert on — `inspectorClient-malformed-list.test.ts` and `listSalvage-era.test.ts` need wire shapes the SDK's own server refuses to emit. Real transport, real client, no `test-servers/` dependency. Check whether a fixture can express the case before reaching for one. + - **It names or runs the built fixture without connecting.** `smoke:tui` boots - the TUI against a catalog whose stdio command *is* the built fixture, then + the TUI against a catalog whose stdio command _is_ the built fixture, then asserts it survives. No transport is driven and no protocol era applies, but the **build and staleness** half lands on it in full. ⚠️ **"A build ran" is not the dependency — using the artefact is.** -`clients/web`'s `pretest` runs `test-servers:build` before *every* unit run, so +`clients/web`'s `pretest` runs `test-servers:build` before _every_ unit run, so the fixture is on disk for tests that never reference it. What counts is whether the test **starts, spawns, configures, or hands a built entry to the subject under test**. That last clause is what covers `smoke:tui`, which drives no transport at all and still depends on the fixture — see the build-only bullet above. -⚠️ **And *importing* the package is not the dependency either.** The barrel +⚠️ **And _importing_ the package is not the dependency either.** The barrel exports plain functions as well as server factories, so a test can import from it and never stand a server up — `src/test/core/mcp/test-server-scope.test.ts` imports `createScopeCheckMiddleware` and friends to unit-test the scope middleware as a pure function, with no `start()` anywhere in the file. None of the procedure applies to it — no config, no era, no lifecycle — it is an ordinary unit test that happens to import its subject from that package. Ask -whether a *server* runs, not whether the import line is present. +whether a _server_ runs, not whether the import line is present. So the condition does **not** hold when the test renders a component from fixture props, exercises a pure function or a parser, or is a smoke that touches @@ -100,7 +101,7 @@ holds `storage/store-id.test.ts`, which validates a string, and `mcp/import/*`, which parses config files, right beside the tests that drive a live connection. They sit there for the node env and the 30s timeout, not because they connect — placement is the project manifest, so it cannot also be the fixture trigger. -Ask what the test *does*, not where it lives. +Ask what the test _does_, not where it lives. **In the connecting case**, the test drives a **real server over a real transport, never a mock**, and picking the fixture, building it, and connecting @@ -122,17 +123,19 @@ the Node clients are different.** Components, hooks, `lib/`, `utils/`. This is the overwhelming majority; a web-owned test living under `src/test/` instead is a bug. -`clients/web/src/test/` is for the three things that *cannot* be co-located: +`clients/web/src/test/` is for the three things that _cannot_ be co-located: 1. **Tests of the repo-root `core/` package** → `src/test/core/…`, mirroring the `core/` folder layout. `core/` physically lives outside `clients/web/`, is consumed via the `@inspector/core` alias, and has no test harness of its own. - This includes `core/json/*` and `core/client/*`. + This includes `core/json/*` and `core/client/*`. **Except `core/cli/`** — + the Node-only surface the one-shot CLI and mcpdo share (#2461): its tests + live in `clients/cli/__tests__/` and `clients/cli`'s coverage run gates it. 2. **The `integration` project** → `src/test/integration/…`, mirroring the `core/` source layout (`mcp/`, `mcp/node/`, `mcp/remote/`, `auth/`, `auth/node/`, `storage/`). **Placement is the manifest** — any file under that folder is picked up by the integration project (node env, 30s timeouts) via a - folder glob; there is no enumeration to keep in sync. ⚠️ Placement is *not* + folder glob; there is no enumeration to keep in sync. ⚠️ Placement is _not_ the fixture trigger, though — this folder holds pure parser and storage tests alongside the connecting ones. If the test you are adding here **needs a fixture from `test-servers/`, load that skill first**; the fixture is half of @@ -141,7 +144,7 @@ web-owned test living under `src/test/` instead is a bug. 3. **Shared test infrastructure** — `renderWithMantine.tsx`, `setup.ts`, `fixtures/`, `scrollAreaStoryAssertions.ts`. -### `clients/cli`, `clients/tui`, `clients/launcher` — a top-level `__tests__/` +### `clients/cli`, `clients/mcpdo`, `clients/tui`, `clients/launcher` — a top-level `__tests__/` **All** their tests, not beside their source. Their `tsconfig.json` excludes `**/*.test.*` and their `tsconfig.test.json` includes `__tests__/**/*`, so a @@ -157,17 +160,18 @@ file its glob misses and still exits 0. ## Running them -| Scope | From | Command | -| --- | --- | --- | -| Web unit | `clients/web` | `npm run test` (`test:watch` while iterating) | -| Web integration | `clients/web` | `npm run test:integration` | -| Web Storybook play fns | `clients/web` | `npm run test:storybook` | -| CLI | `clients/cli` | `npm run test` (`pretest` builds test-servers + the bin) | -| TUI | `clients/tui` | `npm run test` | -| Launcher | `clients/launcher` | `npm run test` | -| Root tooling | repo root | `npm run test:scripts` | -| Everything, fast | repo root | `npm run validate` | -| The coverage gate | repo root | `npm run coverage` | +| Scope | From | Command | +| ---------------------- | -------------------- | -------------------------------------------------------- | +| Web unit | `clients/web` | `npm run test` (`test:watch` while iterating) | +| Web integration | `clients/web` | `npm run test:integration` | +| Web Storybook play fns | `clients/web` | `npm run test:storybook` | +| CLI | `clients/cli` | `npm run test` (`pretest` builds test-servers + the bin) | +| Connection CLI (mcpdo) | `clients/mcpdo` | `npm run test` (`pretest` builds test-servers + the bin) | +| TUI | `clients/tui` | `npm run test` | +| Launcher | `clients/launcher` | `npm run test` | +| Root tooling | repo root | `npm run test:scripts` | +| Everything, fast | repo root | `npm run validate` | +| The coverage gate | repo root | `npm run coverage` | There is **no aggregate root `test` script** — each client self-validates. @@ -201,8 +205,8 @@ inside the `coverage` gate. CI therefore has no separate `test:integration` step ## The coverage gate -**Per-file ≥90 on all four dimensions**, CI-enforced, across web, cli, tui and -launcher. New code must clear 90 on every dimension. +**Per-file ≥90 on all four dimensions**, CI-enforced, across web, cli, +mcpdo, tui and launcher. New code must clear 90 on every dimension. Scope notes: @@ -216,23 +220,19 @@ Scope notes: (a composition root at ~42% branch coverage — gating it is a dedicated decomposition effort) and the `src/main.tsx` / `src/index.ts` bootstraps. - **CLI** tests run **in-process** by importing `runCli()` - (`__tests__/helpers/cli-runner.ts`) so `src` is measured; `src/index.ts` is the - only exclusion. `commander` uses `.exitOverride()` so a parse error throws + (`__tests__/helpers/cli-runner.ts`) so `src` **and the shared `core/cli/`** are + measured (the latter via `allowExternal`, since it sits outside the project + root); `src/index.ts` is the only exclusion. `commander` uses `.exitOverride()` so a parse error throws instead of tearing down the test worker. - **TUI** covers **all of `src/**`, React surface included**. Components mount - through `__tests__/helpers/renderTui.tsx` — `ink-testing-library`'s `render` - with every frame ANSI-stripped — alongside the passthrough doubles in the same - directory; keypresses are driven through stdin. The only exclusion is - `src/tui-servers.ts` (a pure re-export, excluded so it doesn't surface as a - misleading 0/0 row). - ⚠️ **Import `render` from that helper, not from `ink-testing-library`.** Ink - writes styling *inside* the styled run, so `<Text underline>I</Text>nfo` - reaches the frame buffer with escapes between `I` and `nfo` and a plain - `toContain("Info")` fails against a component that is rendering correctly. It - only shows up where chalk emits color — a developer whose shell exports - `FORCE_COLOR` — so CI, which has no TTY, stays green on a suite that is red - for them (#2207). If a frame assertion fails on a string you can plainly see - in the printed diff, that is the tell. Reach `stdout.lastFrame()` on the +through `**tests**/helpers/renderTui.tsx`—`ink-testing-library`'s `render`with every frame ANSI-stripped — alongside the passthrough doubles in the same +directory; keypresses are driven through stdin. The only exclusion is`src/tui-servers.ts`(a pure re-export, excluded so it doesn't surface as a +misleading 0/0 row). +⚠️ **Import`render`from that helper, not from`ink-testing-library`.** Ink +writes styling *inside* the styled run, so `<Text underline>I</Text>nfo`reaches the frame buffer with escapes between`I`and`nfo`and a plain`toContain("Info")`fails against a component that is rendering correctly. It +only shows up where chalk emits color — a developer whose shell exports`FORCE_COLOR`— so CI, which has no TTY, stays green on a suite that is red +for them (#2207). If a frame assertion fails on a string you can plainly see +in the printed diff, that is the tell. Reach`stdout.lastFrame()` on the returned instance for the raw bytes. ### When a `v8 ignore` is justified @@ -295,7 +295,7 @@ the skill and use all of it**: which showcase config covers the feature, which protocol era to connect with, how to add a combination that does not exist yet, and why a fixture can keep serving stale code after an edit. -**A test that only *names* the built fixture needs that skill too, for a +**A test that only _names_ the built fixture needs that skill too, for a narrower reason.** `smoke:tui` boots the TUI against a catalog whose stdio command is the build output and asserts it survives — it opens no transport, so config choice and protocol era do not apply to it, but **building the fixture diff --git a/.github/workflows/dco.yml b/.github/workflows/dco.yml new file mode 100644 index 0000000000..360fe4f17c --- /dev/null +++ b/.github/workflows/dco.yml @@ -0,0 +1,161 @@ +# DCO signoff check (#2566), replacing the probot DCO app. +# +# The app was suspended and its check silently stopped appearing after #1981 +# (2026-08-12). Nothing went red, because it was never a required check — so +# for two months "the DCO check is a hard merge gate" was true only by habit. +# This workflow is the repo-owned replacement: `scripts/dco-check.mjs` fails +# unless every commit in a range carries a `Signed-off-by:` matching its +# author or committer (merge and bot-authored commits exempt — the app's rule). +# +# Two jobs, one script: +# +# - `DCO` — on every v2 PR, over the PR's commits. This is a REQUIRED +# status check on `v2/main`, set by the `v2/main - DCO` repository ruleset +# (#2621), not by this file — no workflow can make itself required. The +# ruleset pins the check to GitHub Actions (integration 15368), so a commit +# status posted by hand under the name `DCO` cannot satisfy it. ⚠️ Renaming +# this job renames the check, and the ruleset would then wait forever on a +# name nothing reports: change both together. +# - `DCO (v2/main push)` — a backstop over every push that lands on +# `v2/main`, for whatever reached the branch without a passing PR check +# while this file and the script stayed intact: an admin merge, a direct +# push, a merge made while the PR check was red or missing. NOT a PR that +# edited the check itself (below). +# It reports when the commit lands, not one milestone later — it cannot +# un-merge anything, but an unsigned commit found the same day is still +# cheap to deal with. +# +# ⚠️ `pull_request`, not `pull_request_target` (#2616). #2603 shipped on +# `pull_request_target` so a PR could not rewrite its own check to pass — but +# GitHub reads `pull_request_target` workflows from the DEFAULT branch +# (`main`), so the check reported on no v2 PR at all until a milestone merge +# carried this file there, and unsigned commits would have surfaced only when +# the milestone PR into `main` was opened. Under `pull_request` the workflow +# runs from the PR's own ref and works the moment it is on `v2/main`, and an +# edit to this file takes effect on the PR that makes it. +# +# The trade-off, accepted deliberately: a PR can now edit this file to pass +# itself. That is tolerable here because PRs are opened by maintainers only +# (AGENTS.md, Contributing) and the check exists to catch a FORGOTTEN signoff, +# not a forged one — the trailer is self-asserted text either way (see the +# header of `scripts/dco-check.mjs`). ⚠️ The push job does NOT cover that +# case: a `push` run also reads this file, and runs the script, from the +# pushed revision — so a PR that edits the check edits the backstop with it. +# Nothing in a workflow the repo's own branches carry can close that; only a +# check read from elsewhere could, which is what `pull_request_target` was and +# what this change gives up. The control for it is review: an edit to this +# file or to `scripts/dco-check.mjs` is in the PR's diff. The PR job keeps the +# trigger safe anyway: it checks out the PR's BASE for the script and only +# FETCHES the PR's commits, reading their metadata with `git log` — nothing +# from the PR is checked out, installed, built or run. +# +# The PR job runs on every `v2/**` base — `v2/main` and, since branch names +# carry their version segment, every v2 feature branch — so a stacked PR is +# checked on its own range rather than only once its parent merges (the +# servers repo's #5055 checks stacks too). A positive filter rather than a +# list of exclusions, so the scope is the v2 line by construction: `main` +# (milestone PRs carry GitHub's squash-merge commits and reach back before the +# check existed) and the whole v1 line, `v1/main` and its stacks alike (no copy +# of the script), are out of it without being named. A stacked PR's base must +# be a branch cut after #2603, since the script is read from it. `edited` re-runs the PR job when a PR is retargeted, +# since the checked range changes with the base even when the head does not. +# +# Before either job, `npm run local:gate` runs the same script over +# `origin/v2/main..HEAD`, minus anything already on `origin/main` so it holds +# on a milestone-merge branch too (`local:dco`), so an unsigned commit is normally +# caught before it is pushed at all. +# +# Only `contents: read`, no credential persisted, no secret — so it stays on +# moving action tags (#2235; see verify:action-pins, #2484). +name: DCO + +on: + pull_request: + types: [opened, synchronize, reopened, edited] + branches: ['v2/**'] + push: + branches: [v2/main] + +permissions: + contents: read + +jobs: + dco: + name: DCO + if: github.event_name == 'pull_request' + runs-on: ubuntu-latest + # A newer push to the same PR supersedes the range being checked. Job-level + # and PR-keyed on purpose: the push job below must NOT share it, since a + # cancelled push run is a range that is never checked. + concurrency: + group: dco-pr-${{ github.event.pull_request.number }} + cancel-in-progress: true + # A git fetch, a git log and a node script; expected to finish well inside + # a minute. Not yet observed on a runner — revisit with the measured range + # per the rule in main.yml (#2333) once it has a history. + timeout-minutes: 5 + + steps: + - name: Checkout the base branch with full history + uses: actions/checkout@v7 + with: + ref: ${{ github.base_ref }} + fetch-depth: 0 + persist-credentials: false + + - name: Fetch the PR's commits (read as data, never checked out) + # `refs/pull/<N>/head` serves fork PRs too. Values pass through `env:` + # rather than being interpolated into the script. + env: + PR_NUMBER: ${{ github.event.pull_request.number }} + run: git fetch --no-tags origin "refs/pull/$PR_NUMBER/head" + + - name: Setup Node.js + uses: actions/setup-node@v7 + with: + node-version: '22.x' + + - name: Check every commit is signed off + # The base is the branch as it stands now (`origin/<base_ref>`), not + # the event's `base.sha`, so commits that have since landed on the + # base are excluded exactly as the PR's own commit list excludes them. + env: + BASE_REF: ${{ github.base_ref }} + HEAD_SHA: ${{ github.event.pull_request.head.sha }} + run: node scripts/dco-check.mjs --base "origin/$BASE_REF" --head "$HEAD_SHA" + + dco-push: + name: DCO (v2/main push) + if: github.event_name == 'push' + runs-on: ubuntu-latest + # Same work as the PR job; same unmeasured budget. + timeout-minutes: 5 + + steps: + - name: Checkout the pushed commit with full history + uses: actions/checkout@v7 + with: + fetch-depth: 0 + persist-credentials: false + + - name: Setup Node.js + uses: actions/setup-node@v7 + with: + node-version: '22.x' + + - name: Check every pushed commit is signed off + # `before..after` is exactly what the push added: for a PR merge, the + # merge commit (exempt) plus every commit it brought in. A zero + # `before` means the branch was just created, which leaves no range to + # check. A `before` missing from the history (a force-push that + # rewrote it away) makes `git log` fail and the script exit 2 — loud + # on purpose, since `v2/main` is never meant to be rewritten. + env: + BEFORE: ${{ github.event.before }} + AFTER: ${{ github.event.after }} + run: | + if [ "$BEFORE" = "0000000000000000000000000000000000000000" ]; then + echo "Branch creation push: no prior commit, nothing to check." + exit 0 + fi + node scripts/dco-check.mjs --base "$BEFORE" --head "$AFTER" diff --git a/.github/workflows/sdk-watch.yml b/.github/workflows/sdk-watch.yml index ec1ff15a1c..aae206d52c 100644 --- a/.github/workflows/sdk-watch.yml +++ b/.github/workflows/sdk-watch.yml @@ -280,8 +280,9 @@ jobs: # an encoded value passes it. # # What makes that acceptable is WHAT THIS JOB READS, not the controls - # around it. Both upstreams — `modelcontextprotocol/typescript-sdk` and - # `modelcontextprotocol/ext-apps` — are in this repository's own org, so + # around it. Every upstream — `modelcontextprotocol/typescript-sdk`, + # `modelcontextprotocol/ext-apps`, and `modelcontextprotocol/ext-tasks` + # — is in this repository's own org, so # their release notes are first-party content, and anyone able to plant # an injection in them already holds release rights here. Closing the # channel properly means workload identity federation instead of a diff --git a/AGENTS.md b/AGENTS.md index 4e8e03be84..76db4c1123 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,7 +1,8 @@ # Inspector V2 -This is an application for inspecting MCP servers. It has three incarnations — -Web, TUI, and CLI — over a shared `core/`. +This is an application for inspecting MCP servers. It has four client +surfaces — Web, TUI, one-shot CLI, and the experimental connection CLI (`mcpdo`) — +over a shared `core/`. **This file holds the _rules_: the conventions a reviewer cites against a diff.** It is loaded in full on every turn, so it stays resident and must stay complete @@ -24,7 +25,7 @@ users invoke them by name. | [`board-ops`](.claude/skills/board-ops/SKILL.md) | `gh project` recipes and the field/option IDs for boards #28 and #11; the option-deletion hazard and its recovery | Model-invoked, or `/board-ops` | | [`pr-flow`](.claude/skills/pr-flow/SKILL.md) | Branch naming, DCO signoff, screenshots, opening the PR, requesting a Copilot review, responding, closing out | Model-invoked, or `/pr-flow` | | [`pre-push-gate`](.claude/skills/pre-push-gate/SKILL.md) | Running `npm run local:gate` and diagnosing a failing stage | Model-invoked, or `/pre-push-gate` | -| [`release`](.claude/skills/release/SKILL.md) | Cutting a release: bump on `v2/main`, milestone merge, tag `origin/main`, publish | `/release` | +| [`release`](.claude/skills/release/SKILL.md) | Cutting a release: bump on `v2/main`, milestone merge, release notes (incl. reporter thanks), tag `origin/main`, publish, re-cut on failure | `/release` | | [`security-advisory`](.claude/skills/security-advisory/SKILL.md) | A privately reported vulnerability end to end: the draft card, who owns the code path, which release lines are affected, accepting, the private fork, publishing, public tracking per line (v2 converts; v1 files) | Model-invoked, or `/security-advisory` | | [`test-servers`](.claude/skills/test-servers/SKILL.md) | Picking and running a showcase test server; the stale-build hazard | Model-invoked, or `/test-servers` | @@ -41,22 +42,35 @@ inspector/ │ │ ├── server/ Node-only dev/prod backend wiring │ │ └── static/ sandbox_proxy.html — served for the MCP Apps tab │ ├── cli/ Scriptable CLI (tsup bundle, @inspector/core alias) +│ ├── mcpdo/ Experimental connection CLI (`mcpdo` bin — connect once, many +│ │ commands; implicit Unix-socket connection daemon). Bundled +│ │ into the published package — see clients/mcpdo/README.md │ ├── tui/ Ink + React terminal UI (tsup bundle) │ └── launcher/ The `mcp-inspector` bin; dispatches to web/cli/tui in-process ├── core/ Shared code, consumed via the `@inspector/core` alias (no package.json) │ ├── auth/ OAuth end to end + the per-server SecretStore backends +│ ├── cli/ Node-only surface the one-shot CLI and mcpdo share (method handlers, +│ │ exit codes, interactive OAuth flow, output styling); tested + gated by clients/cli │ ├── client/ Install-level client config (`client.json`) +│ ├── extension/ Host-side adapters for MCP extensions whose protocol an upstream SDK owns +│ │ └── tasks/ The Inspector's glue around `@modelcontextprotocol/ext-tasks` (raw dispatch, progress routing) │ ├── json/ JSON/schema utilities shared by all three form builders │ ├── logging/ Silent pino logger singleton │ ├── mcp/ InspectorClient, transports, state stores, config import -│ ├── node/ Node-only helpers (version reader, host normalization) +│ ├── node/ Node-only helpers (version reader, host normalization, browser opener) │ ├── react/ React hooks over the state stores (read during render — see React instructions) │ └── storage/ File I/O helpers for the OAuth persist backends ├── test-servers/ Composable MCP test servers + JSON configs ├── scripts/ Root build/verify tooling (install cascade, smokes, verify:* guards) │ plus repo automation run from CI (the dependency, alert + SDK sweeps) +│ and maintainer-workflow helpers (the pr:*, board:*, advisory:*, +│ release:notes, release:tag + action:resolve-pin aliases) ├── docs/ Task-oriented guides ├── specification/ Design/build specifications +├── skills/ End-user agent skills (skills/mcpdo teaches an agent to drive +│ the `mcpdo` CLI; shipped in the published tarball). Distinct +│ from .claude/skills/ (repo procedures): not indexed above +│ and not checked by verify:skills └── .claude/skills/ The procedures (see the index above) ``` @@ -89,14 +103,14 @@ The reasoning behind each of these, and what breaks when it is ignored, is the `local-dev` skill. The rules themselves: - **Every runtime dependency `core/` imports is declared in the repo-root `package.json` and nowhere else.** That is the MCP SDK packages (`@modelcontextprotocol/client`, `core`, `server`, `ext-apps`) and, since #2195, the rest of what `core/` reaches: `ajv`, `atomically`, `chokidar`, `hono`, `@napi-rs/keyring`, `pino`, `proper-lockfile`, `react`, `undici`, `zod`. So is anything reached only through root-owned code with no manifest of its own (`test-servers/src`, `core/`). `@modelcontextprotocol/server-legacy` is that case: only `test-servers/src` imports it (the legacy SSE transport), so it is a root **`devDependency`** — declared as a runtime one, every `npx` install fetched it and printed its npm deprecation warning (#2519). The v1 SDK (`@modelcontextprotocol/sdk`) is **not** a dependency of this repo and must not become one. -- **A root declaration is not by itself a claim that `core/` imports it.** `commander`, `open`, `@hono/node-server`, `vite` and `@vitejs/plugin-react` are root `dependencies` reached only from _client_ code, for the runtime-consumption reason below: a published install resolves every externalized import from the root manifest, so a client's runtime import has to be declared there whether or not `core/` also reaches it. Those need naming only in the `external` list of the client that actually imports them, not in all three. -- **A client declares only what that client alone consumes** — its own UI stack, its bundler-inlined packages, its dev tooling. `clients/cli` and `clients/launcher` therefore declare **no** runtime dependencies at all, and that is the expected steady state, not an omission: everything they run on is root-declared and resolves by walk-up from the client directory. Re-adding a root-declared package to a client manifest re-creates the second copy this rule exists to make impossible (#1896), so a missing module at runtime is a signal to check the **root** manifest and the client's `external` list, never to add it back. +- **A root declaration is not by itself a claim that `core/` imports it.** `commander`, `open`, `@hono/node-server`, `vite` and `@vitejs/plugin-react` are root `dependencies` reached only from _client_ code, for the runtime-consumption reason below: a published install resolves every externalized import from the root manifest, so a client's runtime import has to be declared there whether or not `core/` also reaches it. Those need naming only in the `external` list of the client that actually imports them, not in all four. +- **A client declares only what that client alone consumes** — its own UI stack, its bundler-inlined packages, its dev tooling. `clients/cli`, `clients/mcpdo` and `clients/launcher` therefore declare **no** runtime dependencies at all, and that is the expected steady state, not an omission: everything they run on is root-declared and resolves by walk-up from the client directory. Re-adding a root-declared package to a client manifest re-creates the second copy this rule exists to make impossible (#1896), so a missing module at runtime is a signal to check the **root** manifest and the client's `external` list, never to add it back. - **A package that moves to the root moves its `vitest.shared.mts` pin with it.** Left pointing at `<client>/node_modules` a pin resolves to a directory that no longer exists — or, where a transitive copy happens to sit there (`chokidar` under `vite`, `react` as a peer of `react-dom` and `ink`), to the very duplicate the pin list exists to prevent. **`react` and `react-dom` are the deliberate exception** and stay pinned per client, so a client's renderer and the React it calls into come from one install; every other root-owned pin resolves from the repo root. - **`dependencies` vs `devDependencies` follows from who consumes it at runtime**, not from where it is declared. Anything `core/` imports at runtime must be a root **`dependency`** — the client builds externalize npm packages and a published install resolves them from the root manifest, where devDependencies are absent. - **The shared toolchain is declared once, at the repo root, and in no client manifest.** `eslint`, `@eslint/js`, `typescript-eslint`, `globals`, `prettier`, `typescript`, `vitest`, `@vitest/coverage-v8` and `@types/node` are used by every client's own scripts, and a client that declares none of them still resolves the root copy by walk-up — `npm run` puts each ancestor `node_modules/.bin` on `PATH`, and Node and TypeScript walk parent `node_modules` / `node_modules/@types` the same way. `clients/launcher` declares no `devDependencies` at all and its `validate` is unchanged. A client-side declaration buys nothing and installs a second copy free to drift, as `globals` (`^17.7.0` root / `^17.4.0` clients) and `typescript-eslint` (`^8.65.0` / `^8.56.1`) had before #2196. These stay **`devDependencies`** — none is consumed at runtime and the tarball ships only each client's `build/`. The boundary is **used by every client**, not "used by one": anything narrower stays where it is, whether one client declares it (`tsx`, `playwright`, `storybook`, `happy-dom`, `ink-testing-library`, `vite-node`, each client's own `@types/*`) or several do — `tsup` is declared in web, cli and tui, and `vite` in web and tui on top of the root **runtime** `dependency` that `--web --dev` needs. Those are out of scope here; consolidating them is a different call with a different rationale. - ⚠️ **Deleting the declaration does not always delete the copy, and the local copy still wins.** npm auto-installs an unmet **peer** into the install that needs it, and it has no visibility into the root's tree — so a client-only ESLint plugin drags a client-local `eslint` in (`eslint-plugin-react-refresh`/`-storybook` in web, `eslint-plugin-react-hooks` in tui), and web's Storybook/Vitest stack drags in a local `typescript` and `vitest`. A hoisted transitive does the same: `@types/express` puts an `@types/node` in web and cli. Those copies sit _nearer_ than the root's and take precedence. The consolidation is therefore about **one declaration and one place to bump**, not about a single copy on disk. ⚠️ **Nothing keeps the surviving copies aligned automatically — but since #2226 the guard rejects the drift.** A **peer** copy is at least constrained by its holder's peer range — tightly for `vitest` (an exact peer, hence the pin below), loosely for `eslint` (`^9 || ^10`), where the copies agree only because npm resolves the same latest in both installs. A **transitive** copy is constrained by nothing of ours at all, and cli's `@types/node` (`24.13.1` against the root's `24.13.3`) diverged on exactly that. **That is detection, not alignment: `verify:dep-lockstep` fails on this class since #2226, and you still do the bump by hand.** Its second tier compares every package any install _declares_ (`dependencies`, `devDependencies`, `optionalDependencies`; not peers) against every top-level copy across all five installs, independent of what a `tsc` program loads, so a transitive drift and a peer shadow (`eslint`, `typescript`, `vitest`) are both in scope now. Two limits remain: the tier reads lockfiles, so a tool binary you installed by hand and never committed is still invisible; and it only compares names some manifest declares, so a purely transitive package no manifest names is out of scope in both tiers unless a `tsc` program loads both copies. Aligning a stale install is `npm update <pkg>` there; a transitive copy that will not move takes an `overrides` entry in that install (`clients/cli` pins `@types/node` this way). - ⚠️ **`vitest`, `@vitest/coverage-v8` and web's `@vitest/browser-playwright` are pinned exactly, and move together.** `@vitest/browser-playwright` declares an **exact** peer on `vitest`, so it — not the root range — decides which `vitest` web installs. Left to float, the root resolves a newer patch and web's tests then run on one `vitest` while loading a coverage provider built against another. Bumping means editing all three in one change, the same discipline the exact `prettier` pin (#1790) exists for. ⚠️ **Editing the three is necessary but not sufficient — `clients/web` also carries a `vitest` `overrides` entry that has to move with them.** Web does not declare `vitest`, so its copy is the peer shadow above; its lockfile pins that copy at the old patch, and the exact peer plus the lockfile form a knot `npm install` resolves by refusing outright (`Conflicting peer dependency: vitest@<new>`), while `npm update` will not move it either. Deleting web's lockfile clears the error and re-resolves every caret range in the tree at once — an uncontrolled dependency update wearing a security patch's clothes. The `overrides` entry is the controlled alternative, the same mechanism `clients/cli` uses for `@types/node`: it moves the shadowed copy and nothing else, keeping the churn inside the vitest constellation. So a vitest bump is **four** edits, and the override's version is an exact pin like the other three (#2301). -- **A root-declared package that `core/` imports at runtime must also be named in all three bundler `external` lists** (`clients/{cli,tui}/tsup.config.ts`, `clients/web/tsup.runner.config.ts`), since which client reaches it is a function of what `core/` imports rather than of what the client's own code names. `npm run verify:bundle-externals` enforces this against the **built output**. +- **A root-declared package that `core/` imports at runtime must also be named in all four bundler `external` lists** (`clients/{cli,mcpdo,tui}/tsup.config.ts`, `clients/web/tsup.runner.config.ts`), since which client reaches it is a function of what `core/` imports rather than of what the client's own code names. `npm run verify:bundle-externals` enforces this against the **built output**. - **A dependency that renders React components must be bundled** into the client that uses it (`noExternal`) and declared only there — an externalized one resolves its own `react` and splits the tree. `ink` is the single exemption, on cost, and it is only safe while the root `react` range stays open to the whole major (`^19.0.0`). - **One version per install-crossing dependency.** When bumping a dependency the shared sources pull in, bump it in every install that declares it. Consolidating to the root is what makes most of these unbumpable in two places at once, but it does not retire the rule — a client's `devDependencies`, and any package that arrives transitively into a client install, can still skew against the root. Never raise the tsc heap to work around one. `npm run verify:dep-lockstep` enforces this in two tiers: packages that reach one `tsc` **program** from two installs (the #1896 heap-exhaustion class), and — since #2226 — every package any install **declares** that more than one install holds a top-level copy of, whether or not a program ever sees both. - **Pin a transitive dependency with an `overrides` entry**, not with `npm audit fix` — which "resolves" an advisory with no upward escape by silently downgrading. @@ -132,8 +146,8 @@ The board write needs `organization projects: write`, which `GITHUB_TOKEN` canno **`.github/workflows/sdk-watch.yml` → `scripts/sdk-watch.mjs` runs nightly (#1063) and files one issue per MCP SDK release we are behind**, labeled `v2` + `chore` + `dependencies`. It is not a Dependabot replacement — it exists because SDK churn, OAuth especially, was being tracked by habit rather than by mechanism — but it obeys the same rule as the two above: **it files an issue, never a PR.** -- **Two upstreams, two issues.** `client`/`core`/`server`/`server-legacy` ship from `modelcontextprotocol/typescript-sdk` in lockstep and share one issue; `ext-apps` ships from its own repo and gets its own. A fifth `@modelcontextprotocol/*` package added to the root manifest and not added to `SDK_GROUPS` **fails the sweep loudly** rather than going unwatched — that guard is the point, since a hardcoded group table is otherwise a silent blind spot. -- **It compares the INSTALLED version, not the declared range.** The four SDK packages are pinned exactly, so the two agree for them; `ext-apps` is a caret range, so its lockfile can already resolve higher than the manifest's floor, and comparing the declared string would file an issue for a bump `npm install` has already taken. +- **One issue per upstream.** `client`/`core`/`server`/`server-legacy` ship from `modelcontextprotocol/typescript-sdk` in lockstep and share one issue; `ext-apps` and `ext-tasks` each ship from their own repo and get their own. A `@modelcontextprotocol/*` package added to the root manifest and not added to `SDK_GROUPS` **fails the sweep loudly** rather than going unwatched — that guard is the point, since a hardcoded group table is otherwise a silent blind spot. +- **It compares the INSTALLED version, not the declared range.** The four SDK packages are pinned exactly, so the two agree for them; `ext-apps` and `ext-tasks` are caret ranges, so a lockfile can already resolve higher than the manifest's floor, and comparing the declared string would file an issue for a bump `npm install` has already taken. - **The target is the LOWEST `latest` across a group — the version the whole group has reached — not the highest.** npm publishes a lockstep release one package at a time, so a sweep landing mid-publish sees one package ahead of its three siblings. Targeting the highest would name a version three of them do not have _and_ write a marker that suppresses the real filing once the publication completes, so the release would never be tracked at all. Taking the minimum keeps the issue actionable and lets the completed release file its own. - **It never boards, like the monthly sweep** — no `PROJECT_TOKEN` exists in this org — so the issue arrives labeled and milestoned and `/issue-triage` places it. - **It never closes an issue either.** A further release files its own issue and leaves a **supersession comment** on the older one; closing is a maintainer act, since the card may already have moved. An issue closed for the same target keeps suppressing it, so a maintainer's "not planned" is not re-argued nightly. @@ -165,7 +179,7 @@ So the model returns its write-up as structured output (`--json-schema`) and nev `WebFetch` is denied and the model has no shell, so it controls no outbound channel; release notes are prefetched with `gh api` by a deterministic step. **Keep every one of those properties when editing this job**, and note that a test pins the scan-before-upload ordering specifically, because ordering is the whole control. -⚠️ **One channel is still open, and it is accepted deliberately rather than closed.** `claude-code-action` copies the action's environment into the model's, so `ANTHROPIC_API_KEY` is readable by a `Read` the analysis genuinely needs, and `analysis` is model-controlled text this workflow publishes. The verbatim-credential scan catches only the naive shape; an encoded value passes it. **What makes that acceptable is the threat model, not the controls: both upstreams this sweep reads (`modelcontextprotocol/typescript-sdk`, `modelcontextprotocol/ext-apps`) are in this repository's own org**, so the release notes are first-party content, and anyone able to plant an injection in them already holds release rights here. Closing it properly means workload identity federation instead of a long-lived key (`anthropic_federation_rule_id` + `id-token: write`) — possible, but org-admin work on the Anthropic organization, and disproportionate against our own changelogs (#2269, closed as not planned, records the full reasoning). +⚠️ **One channel is still open, and it is accepted deliberately rather than closed.** `claude-code-action` copies the action's environment into the model's, so `ANTHROPIC_API_KEY` is readable by a `Read` the analysis genuinely needs, and `analysis` is model-controlled text this workflow publishes. The verbatim-credential scan catches only the naive shape; an encoded value passes it. **What makes that acceptable is the threat model, not the controls: every upstream this sweep reads (`modelcontextprotocol/typescript-sdk`, `modelcontextprotocol/ext-apps`, `modelcontextprotocol/ext-tasks`) is in this repository's own org**, so the release notes are first-party content, and anyone able to plant an injection in them already holds release rights here. Closing it properly means workload identity federation instead of a long-lived key (`anthropic_federation_rule_id` + `id-token: write`) — possible, but org-admin work on the Anthropic organization, and disproportionate against our own changelogs (#2269, closed as not planned, records the full reasoning). **That assessment is what to re-open if the inputs change.** Point the analysis at an upstream outside this org, or feed it third-party content, and the trade changes — at which point federation, or a dedicated CI-scoped key with a spend cap, is the next step rather than another grant to narrow. @@ -377,7 +391,7 @@ skills; the rules are here. - **`Incoming` ⇔ no milestone; everything past it ⇔ milestoned — on board #28.** Board #11 is exempt for the reason above: a v1 issue has no bucket to take, so its Status is set on its own and the audit's milestone checks do not apply to it. A `[GHSA-` **advisory draft** on #28 is exempt too, for a different reason: a draft card cannot carry a milestone, so its approval act is **accepting the advisory**, which moves it `Incoming` → `Todo`; its milestone arrives with the public issue after publication. The rest of the invariant is unchanged: assigning the milestone _is_ the approval act, so the two always go together. `Todo` asserts a maintainer signed off, so never park an unreviewed issue there — that erases the distinction and quietly promotes unreviewed work into the queue. An issue created through the documented flow skips `Incoming` entirely, because filing it _was_ the approval. - **`Done` means the work shipped.** Exactly two things earn a card a place in Done: its **PR merged**, or it is a **parent whose last sub-issue closed**. Anything else — duplicate, won't fix, not planned, obsolete, superseded — means nothing shipped, so the card is **deleted**. Done is read as the record of what a milestone actually delivered; a duplicate sitting there makes that record wrong in a way nobody can detect later. Deleting a card touches the board only — the issue keeps its labels and comments and stays searchable forever. - **When work begins**, assign the issue to yourself, create a feature branch and set Status to **In Progress**. **Branch names start with the target version segment** — `v2/fix/2071-oauth-resource-metadata`, `v1/fix/proxy-ssrf-pin` — matching the base branches themselves. -- **When work is complete**, run `npm run format` then `npm run local:gate`, **sign off every commit** (`git commit -s` — the DCO check is a hard merge gate with no partial credit), open a PR against the matching base branch with **`Closes #<ISSUE_NUMBER>` as the body's first line**, and set Status to **In Review**. +- **When work is complete**, run `npm run format` then `npm run local:gate`, **sign off every commit** (`git commit -s` — the repo-owned `DCO` check, `.github/workflows/dco.yml`, fails a PR on any unsigned commit with no partial credit; it gates merges into `v2/main` as a **required** status check, set by the `v2/main - DCO` repository ruleset (#2621) and pinned to GitHub Actions so only the workflow can satisfy it; a PR showing the check as missing or "expected" is an outage, not a pass), open a PR against the matching base branch with **`Closes #<ISSUE_NUMBER>` as the body's first line**, and set Status to **In Review**. - **After opening a PR, run a Copilot review loop to exhaustion — unprompted.** Request a review, wait for the round to post _or_ for Copilot's session to end without one, answer every comment, and request again whenever a fix was pushed. Stop on the **first** clean round (no confirming round "just to be sure"), a round holding only out-of-scope findings, or two rounds in a row where Copilot's session ends without posting. **Weigh each finding against the issue the PR closes and decline scope expansion** — pre-existing behavior, new capabilities, and hardening the issue did not ask for — because that is what turns a review cycle into overbuilding. The recipe is the `pr-flow` skill, step 7. - **Attach screenshots as proof of functionality** for any web-UI or TUI change. Put them in a **`pr-screenshots/`** folder off the repo root — it is **gitignored**, so the images are staged for upload and never committed — and name them for what they show. - ⚠️ Closing keywords only auto-link and auto-close for PRs targeting the **default branch** (`main`). A v2 PR targets `v2/main`, so `Closes #N` there is only a cross-reference and the card shows no linked PR. **Link it explicitly** with the `addCloseIssueReferences` GraphQL mutation right after opening the PR (recipe in `pr-flow`, step 6). **On merge, manually close the issue and move the card to Done.** Keep the line anyway, so the issues close if/when `v2/main` reaches `main`. @@ -398,14 +412,14 @@ When asked to respond to a code review of a PR: The _procedure_ — where a given test file goes, which command runs it, how to diagnose a failing gate — is the `testing` skill. These are the rules. -- **Ensure all code has corresponding tests.** New code must clear **≥ 90 on all four dimensions** — lines, statements, functions, and branches — per file. This gate is enforced by each client's `test:coverage` across `clients/web`, `clients/cli`, `clients/tui` and `clients/launcher`, and **CI enforces it**: a PR that drops any file below 90 on any dimension fails. +- **Ensure all code has corresponding tests.** New code must clear **≥ 90 on all four dimensions** — lines, statements, functions, and branches — per file. This gate is enforced by each client's `test:coverage` across `clients/web`, `clients/cli`, `clients/tui`, `clients/launcher`, and (experimentally) `clients/mcpdo`, and **CI enforces it**: a PR that drops any file below 90 on any dimension fails. **mcpdo** excludes only true bootstraps from the gate (`src/mcp-bin.ts`, `src/daemon/run.ts` — see `clients/mcpdo/vitest.config.ts`). The node-only surface mcpdo shares with the one-shot CLI (`core/cli/` — handlers, error-handler, OAuth helpers; #2461) is gated by `clients/cli`'s coverage run, whose suites exercise it; see `clients/cli/vitest.config.ts`. - **A genuinely-unreachable branch is annotated at the source, never waved through by lowering the gate.** Use a justified `/* v8 ignore … -- <reason> */`. Acceptable reasons: happy-dom-inherent paths (Mantine portal mount points, `useMediaQuery` fallbacks, `typeof window` SSR guards); React StrictMode effect-replay blocks; and provably-dead defensive guards (a `?? fallback` for a value the types guarantee non-null, a `Select.onChange` receiving a value outside the allowed list). Reach for it only when the branch is genuinely impossible to exercise. - **In unit tests that expect error output, suppress it from the console.** - **Test placement — side-by-side by default, `src/test/` only for what can't be co-located, and the Node clients are different.** - - **`clients/web`**: `<Name>.test.tsx` **next to the source** — components, hooks, `lib/`, `utils/`. A web-owned test living under `src/test/` instead is a bug. `src/test/` is for the three things that cannot be co-located: tests of the repo-root **`core/`** package (`src/test/core/…`, mirroring the `core/` layout — it lives outside `clients/web/` and has no harness of its own); the **`integration`** project (`src/test/integration/…` — _placement is the manifest_, picked up by a folder glob, with no enumeration to keep in sync); and **shared test infrastructure** (`renderWithMantine.tsx`, `setup.ts`, `fixtures/`). - - **`clients/cli`, `clients/tui`, `clients/launcher`**: **all** tests in a top-level **`__tests__/`**, not beside their source. Their `tsconfig.json` excludes `**/*.test.*`, so a co-located test lands in **no** tsconfig project and fails `npm run verify:typecheck-coverage`. + - **`clients/web`**: `<Name>.test.tsx` **next to the source** — components, hooks, `lib/`, `utils/`. A web-owned test living under `src/test/` instead is a bug. `src/test/` is for the three things that cannot be co-located: tests of the repo-root **`core/`** package (`src/test/core/…`, mirroring the `core/` layout — it lives outside `clients/web/` and has no harness of its own) — **except `core/cli/`**, the Node-only surface the CLI and mcpdo share, whose tests stay in `clients/cli/__tests__/` and whose coverage `clients/cli` gates (#2461); the **`integration`** project (`src/test/integration/…` — _placement is the manifest_, picked up by a folder glob, with no enumeration to keep in sync); and **shared test infrastructure** (`renderWithMantine.tsx`, `setup.ts`, `fixtures/`). + - **`clients/cli`, `clients/mcpdo`, `clients/tui`, `clients/launcher`**: **all** tests in a top-level **`__tests__/`**, not beside their source. Their `tsconfig.json` excludes `**/*.test.*`, so a co-located test lands in **no** tsconfig project and fails `npm run verify:typecheck-coverage`. - **Root tooling**: a `scripts/*.mjs` helper with pure logic gets a sibling `*.test.mjs`. Keep that exact filename — `node --test` silently _skips_ a file its glob misses and still exits 0. -- **Render Ink components through the TUI's own `render`** (`clients/tui/__tests__/helpers/renderTui.tsx`), never `ink-testing-library`'s directly. It is the same function with every frame ANSI-stripped, which is what keeps an assertion on styled text from depending on the ambient environment: Ink writes styling *inside* the styled run, so `<Text underline>I</Text>nfo` reaches the frame buffer with escapes between `I` and `nfo` and `toContain("Info")` fails. It only bites where chalk emits color — a developer whose shell exports `FORCE_COLOR` — so CI is green on a suite that is broken for them (#2207). A test that genuinely needs the raw bytes reads `stdout.lastFrame()` off the returned instance. +- **Render Ink components through the TUI's own `render`** (`clients/tui/__tests__/helpers/renderTui.tsx`), never `ink-testing-library`'s directly. It is the same function with every frame ANSI-stripped, which is what keeps an assertion on styled text from depending on the ambient environment: Ink writes styling _inside_ the styled run, so `<Text underline>I</Text>nfo` reaches the frame buffer with escapes between `I` and `nfo` and `toContain("Info")` fails. It only bites where chalk emits color — a developer whose shell exports `FORCE_COLOR` — so CI is green on a suite that is broken for them (#2207). A test that genuinely needs the raw bytes reads `stdout.lastFrame()` off the returned instance. - **Render React components through `renderWithMantine`** (`src/test/renderWithMantine.tsx`); do not hand-roll a bare `MantineProvider`, which skips the project theme and the helper's options and drifts from every other test. Pass the `colorScheme` option to exercise a forced scheme rather than hand-rolling `defaultColorScheme`. Use `renderWithMantineTransitions` **only** when a test must assert mid-flight transition state, and read the long comment on the helper before changing anything about it. - **The web coverage `include` is a whitelist.** It names `components`/`hooks`/`theme`/`lib`/`utils`/`server` plus the browser-consumed `core/*` runtime, so a module placed **outside** those directories falls out of the gate entirely, silently. Place new modules inside a gated directory. The documented exceptions — `src/App.tsx` and the `src/main.tsx` / `src/index.ts` bootstraps — are called out in a comment on the `include` array itself. diff --git a/README.md b/README.md index 3b78e4558b..53e031ab64 100644 --- a/README.md +++ b/README.md @@ -52,22 +52,29 @@ inspector/ ├── clients/ │ ├── web/ Web client (Vite + React + Mantine). src/ = browser app; server/ = Node backend │ ├── cli/ CLI client (tsup bundle, @inspector/core alias) +│ ├── mcpdo/ Experimental connection CLI (`mcpdo` bin) — bundled into the +│ │ published package; see clients/mcpdo/README.md │ ├── tui/ TUI client (Ink + React, tsup bundle) │ └── launcher/ Shared launcher — provides the `mcp-inspector` bin, dispatches to web/cli/tui ├── core/ Shared code consumed via the `@inspector/core` alias (no package.json) ├── test-servers/ Composable MCP test servers + fixtures used by integration and smoke tests ├── scripts/ Root build/verify tooling (install cascade, smokes, the verify:* guards), -│ repo automation run from CI (the dependency, Dependabot-alert and SDK sweeps) +│ repo automation run from CI (the dependency, Dependabot-alert and SDK sweeps), +│ maintainer-workflow helpers (the pr:*, board:*, advisory:*, +│ release:notes, release:tag and action:resolve-pin aliases) │ and the Docker image's HEALTHCHECK probe ├── docs/ Task-oriented guides — see below ├── specification/ Design/build specifications +├── skills/ End-user agent skills (e.g. skills/mcpdo teaches an agent to +│ drive the `mcpdo` CLI) — distinct from .claude/skills/, +│ which holds this repo's own procedures ├── .claude/skills/ Agent skills: the repo's procedures, invokable by name ├── AGENTS.md Contribution rules for agents AND humans └── README.md You are here ``` Each client has its own README with client-specific detail: -[web](./clients/web/README.md) · [cli](./clients/cli/README.md) · [tui](./clients/tui/README.md) · [launcher](./clients/launcher/README.md). +[web](./clients/web/README.md) · [cli](./clients/cli/README.md) · [mcpdo](./clients/mcpdo/README.md) · [tui](./clients/tui/README.md) · [launcher](./clients/launcher/README.md). ## Documentation diff --git a/clients/cli/README.md b/clients/cli/README.md index 78abe7a46b..16803fa023 100644 --- a/clients/cli/README.md +++ b/clients/cli/README.md @@ -10,7 +10,7 @@ You can run the CLI via `npx`: npx @modelcontextprotocol/inspector --cli node build/index.js ``` -Supports tools, resources, and prompts (plus `--method servers/list` / `servers/show` for catalog entries without connecting). +Supports tools, resources, and prompts (plus `--method servers/list` / `servers/show` to read catalog entries, and `servers/add` / `servers/edit` / `servers/remove` to change them, without connecting). > Coming from the v1 CLI? See the [v1 → v2 migration guide](../../docs/v1-to-v2-migration.md) — every v1 flag still exists, but exit codes, argument ordering, and the `--` separator changed. @@ -98,6 +98,25 @@ Because undici's `Response` is a different class from `globalThis.Response`, the `undici` is declared **only** in the root `package.json`, and every client's tsup config lists it as `external`. Both halves matter: tsup auto-externalizes what the _nearest_ manifest declares, so without the explicit entry the web and TUI bundles inlined it — and a CommonJS package inlined into an ESM bundle throws `Dynamic require of "assert" is not supported` the first time it is used. `npm run verify:bundle-externals` is the durable guard. +### Shell completion + +`--completion <bash|zsh|fish>` prints a completion script for `mcp-inspector` on stdout and exits without connecting to anything. The first word completes the launcher's mode flags (`--cli`, `--web`, `--tui`); after `--cli` it completes every CLI flag, the `--method` names (`tools/list`, `tools/call`, `servers/list`, …) and the finite values of `--transport`, `--log-level`, `--format` and `--protocol-era`. Path flags (`--catalog`, `--config`, `--cwd`, `--client-config`) and the stdio target command fall back to file completion. + +```bash +# bash — current shell, or persist it +source <(mcp-inspector --cli --completion bash) +mcp-inspector --cli --completion bash > ~/.local/share/bash-completion/completions/mcp-inspector + +# zsh — current shell (after compinit), or save it on your $fpath +source <(mcp-inspector --cli --completion zsh) +mcp-inspector --cli --completion zsh > "${fpath[1]}/_mcp-inspector" + +# fish +mcp-inspector --cli --completion fish > ~/.config/fish/completions/mcp-inspector.fish +``` + +The flag list is read from the CLI's own commander definition when the script is generated, so it cannot drift from `--help`; regenerate the script after upgrading to pick up new flags. Only `--cli` mode is completed — web and TUI flags are not. + ## Options ### MCP server (which server to connect to) @@ -108,7 +127,8 @@ Options that specify the MCP server (catalog/config file, ad-hoc command/URL, en | Option | Description | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `--method <method>` | MCP method to invoke. Supports `initialize` (connect-only probe → `{serverInfo, protocolVersion, capabilities, instructions}`), `tools/list`, `tools/call`, `resources/list`, `resources/read`, `resources/templates/list`, `prompts/list`, `prompts/get`, `logging/setLevel`, `skills/list`, `skills/get`, `resources/directory/read`, plus catalog-only `servers/list` / `servers/show` (no MCP connect). Stream / session-only methods (e.g. `logging/tail`) are rejected. | +| `--method <method>` | MCP method to invoke. Supports `initialize` (connect-only probe → `{serverInfo, protocolVersion, capabilities, instructions}`), `tools/list`, `tools/call`, `resources/list`, `resources/read`, `resources/templates/list`, `prompts/list`, `prompts/get`, `logging/setLevel`, `skills/list`, `skills/get`, `resources/directory/read`, plus catalog-only `servers/list` / `servers/show` and `servers/add` / `servers/edit` / `servers/remove` (no MCP connect; see [Editing the server catalog](#editing-the-server-catalog)). Stream / session-only methods (e.g. `logging/tail`) are rejected. | +| `--rename <name>` | New name for the entry, with `--method servers/edit` only. Secrets stored under the old name move with it. | | `--tool-name <name>` | Tool name (for `tools/call`). | | `--tool-arg <key=value>` | Tool argument; repeat for multiple. Use `key='{"json":true}'` for JSON. Values are coerced (JSON-parsed, so `count=1` becomes a number). | | `--tool-args-json <json>` | Tool arguments as a single JSON object (e.g. `'{"zip":"10001"}'`). Passed verbatim — no `key=value` coercion, so `"012"` stays a string. Mutually exclusive with `--tool-arg`. | @@ -125,10 +145,16 @@ Options that specify the MCP server (catalog/config file, ad-hoc command/URL, en | `--strict` | With `--method tools/list`: report tool-schema portability problems in full (path, issue, suggested fix) on stderr, and exit `6` if any is error-severity. Without it, a one-line count is printed instead. See [Schema portability](#schema-portability---strict). | | `--verify` | With `--method skills/list` or `--method skills/get`: run the SEP-2640 conformance, digest and frontmatter checks over the skills returned, emit one JSON report per skill on stdout, and exit `7` if any fails. See [Skill verification](#skill-verification---verify). | | `--require-digests` | With `--verify`: exit `9` when a skill advertises no digests (`resources: "dynamic"`), instead of reporting it `unverifiable` and exiting `0`. See [Skill verification](#skill-verification---verify). | +| `--skill-catalog-max-skills <n>` | With `--verify` (required; rejected without it, since it would bound nothing): the skills catalog budget, the most skills whose files one run reads (default `256`). Positive integer; anything else is rejected before connecting. Overrides the server's [`skillCatalogMaxSkills`](../../docs/mcp-server-configuration.md#inspector-specific-per-server-fields) from the catalog/config file — the same setting the web client exposes per server in Server Settings. Skills past the budget are reported `incomplete` (exit `8`). | +| `--skill-catalog-max-bytes <n>` | With `--verify` (required): the skills catalog budget, the most bytes one run reads across all skills (default `67108864`, 64 MiB). Positive integer. Overrides the server's [`skillCatalogMaxBytes`](../../docs/mcp-server-configuration.md#inspector-specific-per-server-fields). | | `--format <text\|json>` | Output format. `text` (default) pretty-prints the result. `json` emits a single JSON object on stdout (`{ "result": … }`, plus `{ "appInfo": … }` as a sibling key for App tools) with no banners, so the whole output pipes cleanly into `jq`. | +| `-q`, `--quiet` | Print only the result payload on stdout, or the error envelope on stderr on failure. Drops status lines, warnings, advisory summaries and a stdio server's own stderr; OAuth prompts a human must answer still print. See [Quiet output](#quiet-output---quiet). | +| `--output <path>` | Write the result to a file **instead of** stdout. The parent directory must exist; an existing file is replaced. See [Saving a result to a file](#saving-a-result-to-a-file---output). | +| `--output-format <raw\|json>` | Encoding of the `--output` file. `json` (default) is the whole result, pretty-printed. `raw` is the result's text, or the decoded bytes of its single image/audio/blob block; `tools/call` and `resources/read` only. | | `--relogin` | Delete stored OAuth for this server URL from the shared store before connect; interactive login still only runs if the server requires auth. Requires an HTTP/SSE URL (rejected for stdio). Conflicts with `--stored-auth-only` / `--use-stored-auth` / `--wait-for-auth` / catalog short-circuits. | | `--no-revoke` | With `--relogin`, skip the [RFC 7009](https://datatracker.ietf.org/doc/html/rfc7009) revocation request that would otherwise end the grant at the authorization server when the local state is deleted. The per-server `oauth.revokeOnClear` setting is the persistent form of the same opt-out; either one is enough to skip it. See [Revoking on `--relogin`](#revoking-on---relogin). | | `--stored-auth-only` | **CI / non-interactive safe:** never start interactive OAuth / step-up (and never auto-open a browser); use the shared store if present, otherwise fail immediately with `auth_required`. Prefer this over a bare pipe/CI run that would otherwise attempt interactive login. | +| `--completion <shell>` | Print a `bash`, `zsh` or `fish` completion script and exit (no server connection). See [Shell completion](#shell-completion). | #### Revoking on `--relogin` @@ -149,6 +175,58 @@ mcp-inspector --cli --server-url https://example.com/mcp --relogin --no-revoke - `servers/show` redacts secret-bearing fields (`env` values, sensitive headers, sensitive `settings.metadata` keys whose whole value is replaced whether or not it is structured, `requestInit` / `eventSourceInit` headers, `oauthClientSecret`). It does **not** scrub credentials embedded in a server `url` (userinfo or query tokens) or in stdio `args` — treat `detail` / raw URL fields as potentially sensitive before pasting into issues. +#### Editing the server catalog + +`servers/add`, `servers/edit` and `servers/remove` write the same catalog file the web and TUI clients use — `--catalog <path>`, else `MCP_CATALOG_PATH`, else `~/.mcp-inspector/mcp.json` — through the web backend's own `/api/servers` write path, run in-process. So the file comes out exactly as the web UI would write it: the atomic write, name validation, and the split of secret values (stdio `env` values) out of `mcp.json` into the secret store (OS keychain, or the [fallback store](../../docs/secret-storage.md)). A web backend that is already running picks the change up through its file watcher. The server entry is named with `--server <name>` and described with the same flags an ad-hoc run takes: a positional command or URL, `--server-url`, `--transport`, `-e`, `--cwd`, `--header`, `--protocol-era`. + +```bash +# Add a stdio server (its -e values go to the secret store, not mcp.json) and an HTTP one. +mcp-inspector --cli node build/index.js --method servers/add --server my-server -e API_KEY=… +mcp-inspector --cli --method servers/add --server remote --server-url https://example.com/mcp --header "X-Team: a" + +# Edit: -e merges into the env and --cwd replaces it; a new command/URL replaces the +# transport; --header replaces the headers; --protocol-era sets the era; --rename renames. +# Anything not named (OAuth, timeouts, metadata, roots) is kept. +mcp-inspector --cli --method servers/edit --server my-server -e OTHER=1 --rename my-server-2 + +# Remove the entry and its stored secrets. +mcp-inspector --cli --method servers/remove --server remote +``` + +Each prints `{ ok, action, server, catalog }` (plus `previousName` on a rename); `--format json` wraps it in `{ "result": … }`. `--config` is refused, since it names a read-only session file. `servers/add` fails on a name that already exists, and `servers/edit` / `servers/remove` fail on one that does not. `servers/remove` rejects the entry-describing flags (a command/URL, `--transport`, `-e`, `--cwd`, `--header`, `--protocol-era`, `--rename`), and `servers/edit` with nothing to change is an error. When the secret store is in-memory only (a container with no keychain and nothing durable mounted), a write that supplies `-e` values is refused rather than losing them when the CLI exits, and so is `--rename`, since this process cannot see secrets held in another process's memory — set `MCP_INSPECTOR_SECRET_STORE=file` to keep them. + +#### Quiet output (`--quiet`) + +`-q` / `--quiet` reduces a run to its result: the payload on stdout on success, and the +single-line [error envelope](#exit-codes--error-envelopes) on stderr on failure. It +composes with `--format`: `--format` shapes stdout, and `--quiet` strips stderr of everything non-essential (the exceptions are in the table below). + +```bash +mcp-inspector --cli node build/index.js -q --method tools/list | jq '.tools[].name' +``` + +What it suppresses: + +| Output | Under `--quiet` | +| ------------------------------------------------------------------------------ | --------------- | +| A stdio server's own stderr (startup banners, logs) | Dropped | +| `Schema portability: … Re-run with --strict for details.` (`tools/list`) | Dropped | +| The `--verify` one-line summary | Dropped — a failing run's envelope carries the same text | +| `Authorization complete.` / `Authorization complete. Retrying…` | Dropped | +| `Warning: could not revoke the OAuth grant …` (`--relogin`) | Dropped | +| Advisory warnings from shared Inspector code: the `[mcp-inspector] …` secret-store notice (keychain fallback, `memory` caveat), ignored `roots` / OAuth-endpoint settings, lock and persistence trouble | Dropped (all of `console.warn` is muted for the run). See [secret storage](../../docs/secret-storage.md) for what the store notice would have said. | +| The result payload / NDJSON on stdout | Kept | +| The error envelope on a non-zero exit | Kept | +| The `--strict` report | Kept — you asked for it, and it is the detail behind exit `6` | +| The OAuth authorization URL, the step-up `[y/N]` prompt, "open it by hand" | Kept — interactive login cannot finish without them | + +So a quiet run that needs an interactive login still shows what it must; for a run +that must never prompt, combine `--quiet` with `--stored-auth-only`. + +⚠️ Dropping the server's stderr and the advisory warnings also drops their explanation when something goes wrong. +The CLI still exits non-zero with an envelope, but if the reason is not obvious, +re-run without `--quiet`. + #### App probing (`--app-info`) and machine-readable output (`--format json`) `--app-info` inspects a tool's [MCP App](https://modelcontextprotocol.io) UI posture **without calling the tool**, so a pipeline can decide whether to open a browser before touching one: @@ -181,6 +259,32 @@ mcp-inspector --cli <server> --method tools/call --tool-name my_app_tool --forma A `tools/call` that returns `isError:true` still prints its payload but exits `5` (`tool_is_error`) so `&&` chains don't proceed on a failed call. +#### Saving a result to a file (`--output`) + +`--output <path>` writes the method's result to a file instead of stdout, so a script can archive or diff results without shell redirection ([#2431](https://github.com/modelcontextprotocol/inspector/issues/2431)). `--output-format` picks how the file is encoded: + +| `--output-format` | File contents | +| ----------------- | ------------- | +| `json` (default) | The whole result, pretty-printed with two-space indentation — the same shape the web client's exports download. | +| `raw` | The result's own payload: the text of every text-bearing block (text blocks and embedded text resources), joined by newlines. A result with no text and exactly **one** binary block (image, audio, or resource blob) is written as that block's decoded bytes, so an image tool saves as a viewable file. Anything else has no single raw form and fails with `output_not_raw`. Only `tools/call` and `resources/read` have a raw form. | + +```bash +# Archive the whole result as JSON. +mcp-inspector --cli <server> --method tools/call --tool-name echo --tool-arg message=hi --output echo.json + +# Save just the text a tool returned. +mcp-inspector --cli <server> --method tools/call --tool-name summarize --output summary.txt --output-format raw + +# Save an image tool's output as the image itself. +mcp-inspector --cli <server> --method tools/call --tool-name render_chart --output chart.png --output-format raw +``` + +The flag is `--output-format` rather than a new `--format` value because `--format` already shapes **stdout**, and the two compose: in text mode stdout stays empty and a one-line `Wrote N bytes (json) to <path>` confirmation goes to stderr; under `--format json` stdout still carries exactly one envelope, with `output` in place of `result` — `{"output":{"path":"echo.json","format":"json","bytes":123}}`, plus `appInfo` / `schemaFindings` when they would otherwise appear. + +Exit codes are unchanged: a `tools/call` that returns `isError:true` is still written and still exits `5`. A file that cannot be written (missing directory, no permission) exits `1` with envelope code `output_write_failed`. + +`--output` is rejected where there is no single result to write — `--app-info`, `--verify`, `--list-stored-auth`, `--print-handoff`, and `servers/list` / `servers/show` — and `--output-format` is rejected without `--output`, rather than either being silently ignored. + #### Schema portability (`--strict`) A tool schema can be perfectly legal JSON Schema and still be refused by the diff --git a/clients/cli/__tests__/README.md b/clients/cli/__tests__/README.md index 0e341bbf11..122aea452b 100644 --- a/clients/cli/__tests__/README.md +++ b/clients/cli/__tests__/README.md @@ -3,8 +3,9 @@ Tests live under `__tests__/` and run via Vitest. - Most tests import `runCli()` **in-process** (see `helpers/cli-runner.ts`) so - `clients/cli/src` is measured under the coverage gate. Suite-wide - `helpers/mock-open-url.ts` (vitest `setupFiles`) mocks `open-url` so an armed + `clients/cli/src` and the shared `core/cli` surface are measured under the + coverage gate. Suite-wide + `helpers/mock-open-url.ts` (vitest `setupFiles`) mocks core's `openUrl` so an armed interactive OAuth path cannot launch a real browser. - `e2e.test.ts` (and root `scripts/smoke-cli.mjs`) spawn the built binary for shebang / `process.exit` paths — `pretest` builds `test-servers` + the CLI @@ -24,40 +25,41 @@ npm run validate # format:check && lint && typecheck && test ## Test files -| File | Focus | -| --------------------------------------- | ------------------------------------------- | -| `cli.test.ts` | Core connect / method / config matrix | -| `tools.test.ts` | `tools/call` argument coercion | -| `headers.test.ts` | `--header` merging | -| `metadata.test.ts` | `--metadata` / `--tool-metadata` | -| `methods.test.ts` | Method allow-list / rejection | -| `app-info.test.ts` | `--app-info` probe paths | -| `format-json.test.ts` | `--format json` envelopes | -| `format-output.test.ts` | Text/json writers | -| `emit-result.test.ts` | Result emission helpers | -| `method-types.test.ts` | `ONE_SHOT_METHODS` / guards | -| `run-method.test.ts` | Handler dispatch against a real test server | -| `run-method-mocks.test.ts` | Handler edge cases with mocks | -| `servers-list.test.ts` | `servers/list` / `servers/show` + redaction | -| `cliOAuth.test.ts` | Connect / mid-RPC OAuth recovery | -| `cli-oauth-navigation.test.ts` | OSC 8, arm/disarm, `MCP_AUTO_OPEN_ENABLED` | -| `open-url.test.ts` | `open` package wrapper | -| `oauth-runner.test.ts` | Runner client-config / CIMD flags | -| `oauth-interactive.test.ts` | Loopback callback OAuth (in-process) | -| `stored-auth.test.ts` | `--use-stored-auth` / handoff / wait | -| `clear-stored-auth-for-relogin.test.ts` | Dual-key `--relogin` clear | -| `programmatic-ergonomics.test.ts` | Flag conflicts, timeouts, ergonomics | -| `error-handler.test.ts` | Exit codes / error shaping | -| `style.test.ts` | ANSI / OSC 8 helpers | -| `e2e.test.ts` | Out-of-process binary smoke | -| `helpers/assertions.test.ts` | Assertion helpers | +| File | Focus | +| --------------------------------------- | --------------------------------------------------------------------------------------- | +| `cli.test.ts` | Core connect / method / config matrix | +| `tools.test.ts` | `tools/call` argument coercion | +| `headers.test.ts` | `--header` merging | +| `metadata.test.ts` | `--metadata` / `--tool-metadata` | +| `methods.test.ts` | Method allow-list / rejection | +| `app-info.test.ts` | `--app-info` probe paths | +| `format-json.test.ts` | `--format json` envelopes | +| `format-output.test.ts` | Text/json writers | +| `emit-result.test.ts` | Result emission helpers | +| `output-file.test.ts` | `--output` / `--output-format` (#2431) | +| `method-types.test.ts` | `ONE_SHOT_METHODS` / guards | +| `run-method.test.ts` | Handler dispatch against a real test server | +| `run-method-mocks.test.ts` | Handler edge cases with mocks | +| `servers-list.test.ts` | `servers/list` / `servers/show` + redaction | +| `servers-write.test.ts` | `servers/add` / `servers/edit` / `servers/remove` against a temp catalog + secret store | +| `cliOAuth.test.ts` | Connect / mid-RPC OAuth recovery | +| `cli-oauth-navigation.test.ts` | OSC 8, arm/disarm, `MCP_AUTO_OPEN_ENABLED` | +| `oauth-runner.test.ts` | Runner client-config / CIMD flags | +| `oauth-interactive.test.ts` | Loopback callback OAuth (in-process) | +| `stored-auth.test.ts` | `--use-stored-auth` / handoff / wait | +| `clear-stored-auth-for-relogin.test.ts` | Dual-key `--relogin` clear | +| `programmatic-ergonomics.test.ts` | Flag conflicts, timeouts, ergonomics | +| `error-handler.test.ts` | Exit codes / error shaping | +| `style.test.ts` | ANSI / OSC 8 helpers | +| `e2e.test.ts` | Out-of-process binary smoke | +| `helpers/assertions.test.ts` | Assertion helpers | ## Helpers | Helper | Role | | -------------------------- | --------------------------------------------------------- | | `helpers/cli-runner.ts` | In-process `runCli` with stdout/stderr capture | -| `helpers/mock-open-url.ts` | Suite-wide `open-url` mock (vitest `setupFiles`) | +| `helpers/mock-open-url.ts` | Suite-wide `openUrl` mock (vitest `setupFiles`) | | `helpers/assertions.ts` | `expectCliSuccess` / `expectCliFailure` / output matchers | | `helpers/fixtures.ts` | Temp config / client.json factories | diff --git a/clients/cli/__tests__/cli-oauth-navigation.test.ts b/clients/cli/__tests__/cli-oauth-navigation.test.ts index 1932d5b625..a9d0c9e49f 100644 --- a/clients/cli/__tests__/cli-oauth-navigation.test.ts +++ b/clients/cli/__tests__/cli-oauth-navigation.test.ts @@ -3,12 +3,12 @@ import { createCliOAuthNavigation, isCliAutoOpenForced, resolveCliAutoOpenEnabled, -} from "../src/cli-oauth-navigation.js"; -import { openUrl } from "../src/open-url.js"; +} from "@inspector/core/cli/cli-oauth-navigation.js"; +import { openUrl } from "@inspector/core/node/openUrl.js"; // Local mock (in addition to suite-wide setupFiles) so this file owns a // `vi.mocked(openUrl)` handle and does not depend on the global mock shape. -vi.mock("../src/open-url.js", () => ({ +vi.mock("@inspector/core/node/openUrl.js", () => ({ openUrl: vi.fn().mockResolvedValue(undefined), })); diff --git a/clients/cli/__tests__/cli.test.ts b/clients/cli/__tests__/cli.test.ts index 020f14466c..d085709891 100644 --- a/clients/cli/__tests__/cli.test.ts +++ b/clients/cli/__tests__/cli.test.ts @@ -1,5 +1,5 @@ import { afterAll, beforeAll, describe, it, expect, vi } from "vitest"; -import { mkdtempSync } from "node:fs"; +import { mkdtempSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; import { join } from "node:path"; import { runCli } from "./helpers/cli-runner.js"; @@ -85,6 +85,109 @@ describe("CLI Tests", () => { expect(toolNames).toContain("get_annotated_message"); }); + // #2435. In-process, a stdio server's own stderr goes straight to the + // worker's fd 2 rather than through the captured `process.stderr.write`, + // so the child-stderr half is asserted out of process in e2e.test.ts. + it.each(["-q", "--quiet"])( + "%s prints only the result payload", + async (flag) => { + const { command, args } = getTestMcpServerCommand(); + const result = await runCli([ + command, + ...args, + flag, + "--method", + "tools/list", + ]); + + expectCliSuccess(result); + expect(expectValidJson(result)).toHaveProperty("tools"); + expect(result.stderr).toBe(""); + }, + ); + + // A usage error makes Commander print its own `error: …` line during + // `parse()`, before the envelope. Under `--quiet` that line is dropped and + // stderr is the envelope alone (Copilot on #2576). + it("drops Commander's usage diagnostic under --quiet, keeping only the envelope", async () => { + const loud = await runCli([NO_SERVER_SENTINEL, "--method"]); + expectCliFailure(loud); + expect(loud.stderr).toMatch(/^error: option '--method <method>'/); + + const quiet = await runCli([NO_SERVER_SENTINEL, "-q", "--method"]); + expectCliFailure(quiet); + const lines = quiet.stderr.trimEnd().split("\n"); + expect(lines).toHaveLength(1); + expect(JSON.parse(lines[0]!)).toMatchObject({ + error: { message: expect.stringMatching(/--method <method>/) }, + }); + }); + + // Core reports advisories with `console.warn` from many places — here + // `cleanRoots()` on a hand-edited entry. Vitest replaces `console`, so the + // captured stderr cannot show it; the spy is what observes the channel. + // Under `--quiet` the run swaps `console.warn` out and restores it after + // (Copilot on #2576). + it("mutes core's console.warn advisories for a --quiet run, and restores console.warn after", async () => { + const { command, args } = getTestMcpServerCommand(); + const dir = mkdtempSync(join(tmpdir(), "cli-quiet-roots-")); + const configPath = join(dir, "mcp.json"); + // A root with no `uri` is dropped with a warning (core/mcp/serverList.ts). + writeFileSync( + configPath, + JSON.stringify({ + mcpServers: { + s: { type: "stdio", command, args, roots: [{ name: "no-uri" }] }, + }, + }), + ); + const run = (extra: string[]) => + runCli(["--config", configPath, ...extra, "--method", "tools/list"]); + try { + const loud = vi.spyOn(console, "warn").mockImplementation(() => {}); + expectCliSuccess(await run([])); + expect(loud).toHaveBeenCalledWith( + "Dropping root without a string `uri`:", + expect.anything(), + ); + loud.mockRestore(); + + const quiet = vi.spyOn(console, "warn").mockImplementation(() => {}); + expectCliSuccess(await run(["-q"])); + expect(quiet).not.toHaveBeenCalled(); + expect(console.warn).toBe(quiet); + quiet.mockRestore(); + } finally { + rmSync(dir, { recursive: true, force: true }); + } + }); + + // Quiet is read from Commander's parsed state, so a combined `-qe` counts + // and a `-q` that is another option's value does not (Copilot on #2576). + it("applies the pre-parse --quiet handling when -q is combined with -e", async () => { + const result = await runCli([ + NO_SERVER_SENTINEL, + "-qe", + "A=1", + "--method", + ]); + expectCliFailure(result); + const lines = result.stderr.trimEnd().split("\n"); + expect(lines).toHaveLength(1); + expect(JSON.parse(lines[0]!)).toHaveProperty("error"); + }); + + it("does not treat -q as the flag when it is another option's value", async () => { + const result = await runCli([ + NO_SERVER_SENTINEL, + "--client-secret", + "-q", + "--method", + ]); + expectCliFailure(result); + expect(result.stderr).toMatch(/^error: option '--method <method>'/); + }); + it("should fail with nonexistent method", async () => { const result = await runCli([ NO_SERVER_SENTINEL, diff --git a/clients/cli/__tests__/cliOAuth.test.ts b/clients/cli/__tests__/cliOAuth.test.ts index 70738d784b..c7eb242657 100644 --- a/clients/cli/__tests__/cliOAuth.test.ts +++ b/clients/cli/__tests__/cliOAuth.test.ts @@ -10,7 +10,7 @@ import { assertInteractiveOAuthAllowed, withCliAuthRecoveryRetry, STEP_UP_PIPE_TIMEOUT_MS, -} from "../src/cliOAuth.js"; +} from "@inspector/core/cli/cliOAuth.js"; import type { MCPServerConfig } from "@inspector/core/mcp/types.js"; import { createInterface } from "node:readline/promises"; import { @@ -1184,4 +1184,127 @@ describe("cliOAuth", () => { } } }); + + // #2435: `--quiet` drops the status lines, never the prompts a human has to + // act on. The step-up [y/N] still prints, because the flow cannot finish + // without an answer to it. + describe("--quiet", () => { + const QUIET_CALLBACK = CALLBACK_URL_CONFIG; + function fakeClient(connect = vi.fn().mockResolvedValue(undefined)) { + return { + connect, + disconnect: vi.fn().mockResolvedValue(undefined), + authenticate: vi.fn(), + beginInteractiveAuthorization: vi.fn(), + completeOAuthFlow: vi.fn(), + checkAuthChallengeSatisfied: vi.fn().mockResolvedValue(false), + }; + } + + it("runCliInteractiveOAuth writes no success line", async () => { + vi.spyOn( + runnerInteractive, + "runRunnerInteractiveOAuth", + ).mockResolvedValue({ kind: "success" }); + const stderrSpy = vi + .spyOn(process.stderr, "write") + .mockImplementation(() => true); + + await runCliInteractiveOAuth( + fakeClient(), + new MutableRedirectUrlProvider(), + QUIET_CALLBACK, + { quiet: true }, + ); + + expect(stderrSpy).not.toHaveBeenCalled(); + }); + + it("withCliAuthRecoveryRetry writes no retry line after AuthRecoveryRequired, but keeps the step-up prompt", async () => { + vi.spyOn( + runnerInteractive, + "runRunnerInteractiveOAuth", + ).mockResolvedValue({ kind: "success" }); + const stderrSpy = vi + .spyOn(process.stderr, "write") + .mockImplementation(() => true); + const fn = vi + .fn() + .mockRejectedValueOnce( + new AuthRecoveryRequiredError( + new URL("https://as.example/authorize"), + { reason: "insufficient_scope", requiredScopes: ["weather:read"] }, + ), + ) + .mockResolvedValueOnce("ok"); + + const result = await withCliAuthRecoveryRetry( + fakeClient(), + OAUTH_HTTP_CONFIG, + new MutableRedirectUrlProvider(), + QUIET_CALLBACK, + makeFakeServerSettings(), + fn, + { ...INTERACTIVE, confirmStepUp: async () => true, quiet: true }, + ); + + expect(result).toBe("ok"); + const written = stderrSpy.mock.calls.map((c) => String(c[0])).join(""); + expect(written).toContain("Proceed with step-up authorization? [y/N]"); + expect(written).not.toContain("Authorization complete"); + }); + + it("withCliAuthRecoveryRetry writes no retry line after a plain 401", async () => { + vi.spyOn( + runnerInteractive, + "runRunnerInteractiveOAuth", + ).mockResolvedValue({ kind: "success" }); + const stderrSpy = vi + .spyOn(process.stderr, "write") + .mockImplementation(() => true); + const fn = vi + .fn() + .mockRejectedValueOnce(new Error("RPC failed (401)")) + .mockResolvedValueOnce("ok"); + + const result = await withCliAuthRecoveryRetry( + fakeClient(), + OAUTH_HTTP_CONFIG, + new MutableRedirectUrlProvider(), + QUIET_CALLBACK, + undefined, + fn, + { ...INTERACTIVE, quiet: true }, + ); + + expect(result).toBe("ok"); + expect(stderrSpy).not.toHaveBeenCalled(); + }); + + it("connectInspectorWithOAuth writes no success line after a 401 on connect", async () => { + vi.spyOn( + runnerInteractive, + "runRunnerInteractiveOAuth", + ).mockResolvedValue({ kind: "success" }); + const stderrSpy = vi + .spyOn(process.stderr, "write") + .mockImplementation(() => true); + const connect = vi + .fn() + .mockRejectedValueOnce(new Error("RPC failed (401)")) + .mockResolvedValue(undefined); + + await connectInspectorWithOAuth( + fakeClient(connect), + OAUTH_HTTP_CONFIG, + new MutableRedirectUrlProvider(), + QUIET_CALLBACK, + undefined, + { ...INTERACTIVE, quiet: true }, + ); + + expect(connect).toHaveBeenCalledTimes(2); + expect(stderrSpy).not.toHaveBeenCalled(); + }); + }); }); diff --git a/clients/cli/__tests__/completion.test.ts b/clients/cli/__tests__/completion.test.ts new file mode 100644 index 0000000000..4c7043428d --- /dev/null +++ b/clients/cli/__tests__/completion.test.ts @@ -0,0 +1,411 @@ +/** + * Shell completion (#2434): the generated scripts are derived from the real + * commander program, offer the method names, and actually complete in each + * shell. The shell-level tests spawn bash / zsh / fish when installed and are + * skipped otherwise; the structural tests below always run. + */ +import { describe, it, expect, afterAll } from "vitest"; +import { spawnSync } from "node:child_process"; +import { mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { Command, Option } from "commander"; +import { runCli } from "./helpers/cli-runner.js"; +import { + CATALOG_METHODS, + COMPLETION_SHELLS, + VALUE_CHOICES, + collectCompletionFlags, + emitCompletionIfRequested, + isCompletionShell, + parseCompletionShell, + registerCompletionOption, + renderCompletion, + renderZsh, + type CompletionShell, +} from "../src/completion.js"; +import { ONE_SHOT_METHODS } from "@inspector/core/cli/handlers/method-types.js"; +import { CATALOG_WRITE_METHODS } from "../src/handlers/servers-write.js"; + +async function script(shell: CompletionShell): Promise<string> { + const result = await runCli(["--completion", shell]); + expect(result.exitCode).toBe(0); + expect(result.stderr).toBe(""); + return result.stdout; +} + +const tmp = mkdtempSync(join(tmpdir(), "cli-completion-")); +afterAll(() => rmSync(tmp, { recursive: true, force: true })); + +function hasShell(shell: string): boolean { + return spawnSync(shell, ["-c", "exit 0"]).status === 0; +} + +function writeScript(shell: CompletionShell, body: string): string { + const path = join(tmp, `completion.${shell}`); + writeFileSync(path, body); + return path; +} + +describe("--completion", () => { + it.each(COMPLETION_SHELLS)("prints a %s script and exits 0", async (s) => { + const out = await script(s); + expect(out).toContain("mcp-inspector"); + expect(out).toMatch(/--method|-l method/); + }); + + it("rejects an unsupported shell", async () => { + const result = await runCli(["--completion", "pwsh"]); + expect(result.exitCode).toBe(1); + expect(result.stderr).toContain( + "Invalid shell: pwsh. Supported shells are: bash, zsh, fish", + ); + }); + + it("short-circuits ahead of method validation", async () => { + // No --method and no server: a normal run would fail "Method is required". + const result = await runCli(["--completion", "fish"]); + expect(result.exitCode).toBe(0); + }); + + it("lists every CLI flag, including ones it was not told about", async () => { + const out = await script("bash"); + for (const flag of [ + "--catalog", + "--config", + "--server", + "-e", + "--tool-arg", + "--format", + "--relogin", + "--no-revoke", + "--print-handoff", + "--completion", + "--help", + "-h", + ]) { + expect(out).toContain(` ${flag}`); + } + // Spot-check one flag from the far end of the definition too. + expect(out).toContain(" --wait-for-auth"); + }); + + it("offers every --method the CLI accepts", async () => { + const out = await script("bash"); + for (const method of [...ONE_SHOT_METHODS, ...CATALOG_METHODS]) { + expect(out).toContain(method); + } + }); + + it("offers the catalog write methods the CLI implements (#2629)", async () => { + // Read from servers-write itself, not from CATALOG_METHODS, so a write + // method the completion list forgets fails here. + expect(CATALOG_METHODS).toEqual( + expect.arrayContaining([...CATALOG_WRITE_METHODS]), + ); + const out = await script("bash"); + for (const method of CATALOG_WRITE_METHODS) { + expect(out).toContain(method); + } + }); + + it("offers no --method the CLI would reject as unsupported", async () => { + for (const method of VALUE_CHOICES["--method"]!) { + const result = await runCli(["--method", method]); + expect(result.stderr, method).not.toContain("Unsupported method"); + } + // Control: an unknown method is rejected by the same check. + const bogus = await runCli(["--method", "servers/bogus"]); + expect(bogus.stderr).toContain("Unsupported method: servers/bogus"); + }); +}); + +describe("collectCompletionFlags", () => { + it("derives value, path and choice metadata from the commander definition", () => { + const program = new Command() + .option("--plain", "A boolean switch value. Second sentence") + .option("--terse", "Needs --x. Then does a thing. Third") + .option("--file <path>", "A path") + .option("--name <value>", "Free text (e.g. foo).") + .option("--method <m>", "Method") + .option("-x <v>", "Short only") + .option("--secret", "Hidden"); + program.options.find((o) => o.long === "--secret")!.hidden = true; + registerCompletionOption(program); + + const flags = collectCompletionFlags(program); + const byName = (n: string) => + flags.find((f) => f.long === n || f.short === n); + + expect(byName("--plain")).toEqual({ + long: "--plain", + takesValue: false, + description: "A boolean switch value", + }); + expect(byName("--terse")?.description).toBe("Needs --x. Then does a thing"); + expect(byName("--file")).toMatchObject({ takesValue: true, path: true }); + expect(byName("--name")).toMatchObject({ + takesValue: true, + description: "Free text (e.g. foo)", + }); + expect(byName("--name")?.path).toBeUndefined(); + expect(byName("--method")?.choices).toBe(VALUE_CHOICES["--method"]); + expect(byName("-x")).toMatchObject({ takesValue: true }); + expect(byName("-x")?.long).toBeUndefined(); + expect(byName("--secret")).toBeUndefined(); + expect(byName("--completion")?.choices).toEqual(COMPLETION_SHELLS); + expect(byName("--help")).toMatchObject({ short: "-h" }); + }); + + it("prefers commander choices when an option declares them", () => { + const program = new Command().addOption( + new Option("--color <c>", "Color").choices(["red", "blue"]), + ); + expect(collectCompletionFlags(program)[0]!.choices).toEqual([ + "red", + "blue", + ]); + }); + + it("handles an empty description", () => { + const program = new Command().option("--bare"); + expect(collectCompletionFlags(program)[0]!.description).toBe(""); + }); + + it("every stated value choice is accepted by the real option parser", async () => { + // Drift guard: VALUE_CHOICES restates sets the CLI validates in custom + // parsers. A value the parser would reject must not be offered. + for (const [flag, values] of Object.entries(VALUE_CHOICES)) { + if (flag === "--method" || flag === "--completion") continue; + for (const value of values) { + const result = await runCli([flag, value, "--completion", "bash"]); + expect(result.exitCode, `${flag} ${value}`).toBe(0); + } + } + }); +}); + +describe("shell helpers", () => { + it("escapes backslashes before colons in a zsh _describe name (CodeQL #78)", () => { + const out = renderZsh([ + { long: "--a\\b:c", takesValue: false, description: "Desc" }, + ]); + // Name --a\b:c → --a\\b\:c, then ":" and the description. + expect(out).toContain("'--a\\\\b\\:c:Desc'"); + }); + + it("parseCompletionShell / isCompletionShell", () => { + expect(isCompletionShell("zsh")).toBe(true); + expect(isCompletionShell("csh")).toBe(false); + expect(parseCompletionShell("fish")).toBe("fish"); + expect(() => parseCompletionShell("csh")).toThrow(/Invalid shell: csh/); + }); + + it("emitCompletionIfRequested is a no-op without --completion", async () => { + const program = new Command(); + registerCompletionOption(program); + program.parse([], { from: "user" }); + expect(await emitCompletionIfRequested(program)).toBe(false); + }); + + it("quotes descriptions containing quotes and backslashes", () => { + const flags = [ + { + long: "--q", + takesValue: false, + description: `it's a \\ "test"`, + }, + ]; + expect(renderCompletion("bash", flags)).toContain("--q"); + expect(renderCompletion("zsh", flags)).toContain( + `'--q:it'\\''s a \\ "test"'`, + ); + expect(renderCompletion("fish", flags)).toContain( + `-d 'it\\'s a \\\\ "test"'`, + ); + }); +}); + +describe.skipIf(!hasShell("bash"))("bash script", () => { + async function complete(words: string[]): Promise<string[]> { + const path = writeScript("bash", await script("bash")); + const quoted = words.map((w) => `'${w}'`).join(" "); + const res = spawnSync( + "bash", + [ + "--norc", + "-c", + `source '${path}'; COMP_WORDS=(${quoted}); COMP_CWORD=$((\${#COMP_WORDS[@]}-1)); _mcp_inspector; printf '%s\\n' "\${COMPREPLY[@]}"`, + ], + { encoding: "utf8" }, + ); + expect(res.stderr).toBe(""); + return res.stdout.split("\n").filter(Boolean); + } + + it("is valid bash", async () => { + const path = writeScript("bash", await script("bash")); + expect(spawnSync("bash", ["-n", path]).status).toBe(0); + }); + + it("completes mode flags, CLI flags and method names", async () => { + expect(await complete(["mcp-inspector", "--c"])).toEqual(["--cli"]); + expect(await complete(["mcp-inspector", "--cli", "--meth"])).toEqual([ + "--method", + ]); + expect( + await complete(["mcp-inspector", "--cli", "--method", "tools/"]), + ).toEqual(["tools/list", "tools/call"]); + expect( + await complete(["mcp-inspector", "--cli", "--transport", ""]), + ).toEqual(["stdio", "sse", "http"]); + expect( + await complete(["mcp-inspector", "--cli", "--method", "servers/"]), + ).toEqual([ + "servers/list", + "servers/show", + "servers/add", + "servers/edit", + "servers/remove", + ]); + expect( + await complete(["mcp-inspector", "--cli", "--output-format", ""]), + ).toEqual(["raw", "json"]); + // A free-form value offers nothing (the shell falls back to files). + expect( + await complete(["mcp-inspector", "--cli", "--tool-name", "--"]), + ).toEqual([]); + // Flags still complete after a stdio target command. + expect( + await complete(["mcp-inspector", "--cli", "node", "s.js", "--comp"]), + ).toEqual(["--completion"]); + // Non-CLI modes are out of scope. + expect(await complete(["mcp-inspector", "--web", "--meth"])).toEqual([]); + }); + + it("completes the value of --opt=value (split on '=' by COMP_WORDBREAKS)", async () => { + // Cursor right after `--method=`: the current word is "=". + expect( + await complete(["mcp-inspector", "--cli", "--method", "="]), + ).toContain("tools/list"); + // `--method=tools/`: the value is the current word, "=" the previous. + expect( + await complete(["mcp-inspector", "--cli", "--method", "=", "tools/"]), + ).toEqual(["tools/list", "tools/call"]); + expect( + await complete(["mcp-inspector", "--cli", "--format", "=", "j"]), + ).toEqual(["json"]); + }); +}); + +describe.skipIf(!hasShell("zsh"))("zsh script", () => { + /** + * Drive `_mcp_inspector` with the completion builtins stubbed to print what + * they were handed, so the test needs no interactive shell or compinit. + */ + async function complete(words: string[]): Promise<string> { + const path = writeScript("zsh", await script("zsh")); + const quoted = words.map((w) => `'${w}'`).join(" "); + const stubs = [ + 'compdef() { print -r -- "compdef $*" }', + 'compadd() { shift; print -r -- "compadd $*" }', + "_files() { print -r -- _files }", + '_message() { print -r -- "_message $*" }', + '_describe() { print -r -- "_describe ${(P)4}" }', + 'compset() { print -r -- "compset $*" }', + ].join("\n"); + const res = spawnSync( + "zsh", + [ + "-f", + "-c", + `${stubs}\nsource '${path}'\nwords=(${quoted}); CURRENT=\${#words}; _mcp_inspector`, + ], + { encoding: "utf8" }, + ); + expect(res.stderr).toBe(""); + return res.stdout; + } + + it("is valid zsh", async () => { + const path = writeScript("zsh", await script("zsh")); + expect(spawnSync("zsh", ["-n", path]).status).toBe(0); + }); + + it("completes mode flags, CLI flags and method names", async () => { + expect(await complete(["mcp-inspector", "--c"])).toContain( + "--cli:Run the CLI", + ); + const flags = await complete(["mcp-inspector", "--cli", "--meth"]); + expect(flags).toContain("--method:Method to invoke"); + expect(flags).toContain("-e:Environment variables"); + expect( + await complete(["mcp-inspector", "--cli", "--method", ""]), + ).toContain("compadd initialize tools/list tools/call"); + expect(await complete(["mcp-inspector", "--cli", "--config", ""])).toBe( + "compdef _mcp_inspector mcp-inspector\n_files\n", + ); + expect(await complete(["mcp-inspector", "--cli", "--uri", ""])).toContain( + "_message value", + ); + expect(await complete(["mcp-inspector", "--cli", "node"])).toContain( + "_files", + ); + expect(await complete(["mcp-inspector", "--tui", "--x"])).toContain( + "_files", + ); + }); + + it("completes the value of --opt=value, moving the option into IPREFIX", async () => { + const method = await complete([ + "mcp-inspector", + "--cli", + "--method=tools/", + ]); + expect(method).toContain("compset -P *="); + expect(method).toContain("compadd initialize tools/list tools/call"); + expect(method).not.toContain("_describe"); + expect(await complete(["mcp-inspector", "--cli", "--config=./"])).toContain( + "_files", + ); + // A boolean flag given `=`: nothing to offer, and no option list either. + expect( + await complete(["mcp-inspector", "--cli", "--strict="]), + ).not.toContain("_describe"); + }); +}); + +describe.skipIf(!hasShell("fish"))("fish script", () => { + async function complete(line: string): Promise<string[]> { + const path = writeScript("fish", await script("fish")); + const res = spawnSync( + "fish", + ["--no-config", "-c", `source '${path}'; complete -C '${line}'`], + { encoding: "utf8" }, + ); + expect(res.stderr).toBe(""); + return res.stdout + .split("\n") + .filter(Boolean) + .map((l) => l.split("\t")[0]!); + } + + it("is valid fish", async () => { + const path = writeScript("fish", await script("fish")); + expect(spawnSync("fish", ["-n", path]).status).toBe(0); + }); + + it("completes mode flags, CLI flags and method names", async () => { + expect(await complete("mcp-inspector --c")).toEqual(["--cli"]); + expect(await complete("mcp-inspector --cli --meth")).toEqual(["--method"]); + expect( + (await complete("mcp-inspector --cli --method tools/")).sort(), + ).toEqual(["tools/call", "tools/list"]); + expect(await complete("mcp-inspector --cli -")).toContain("-e"); + expect(await complete("mcp-inspector --web --meth")).toEqual([]); + // fish splits `--opt=value` natively. + expect( + (await complete("mcp-inspector --cli --method=tools/")).sort(), + ).toEqual(["--method=tools/call", "--method=tools/list"]); + }); +}); diff --git a/clients/cli/__tests__/consume-outcome-stream.test.ts b/clients/cli/__tests__/consume-outcome-stream.test.ts index f527c0e172..bbcb937bb3 100644 --- a/clients/cli/__tests__/consume-outcome-stream.test.ts +++ b/clients/cli/__tests__/consume-outcome-stream.test.ts @@ -1,6 +1,6 @@ import { describe, it, expect, afterEach, vi } from "vitest"; import { consumeMethodOutcome } from "../src/handlers/consume-outcome.js"; -import type { MethodOutcome } from "../src/handlers/method-types.js"; +import type { MethodOutcome } from "@inspector/core/cli/handlers/method-types.js"; /** * The long-lived stream path's stdout error handling (#2412). A reader that diff --git a/clients/cli/__tests__/e2e.test.ts b/clients/cli/__tests__/e2e.test.ts index e12932c8f1..e8e485f607 100644 --- a/clients/cli/__tests__/e2e.test.ts +++ b/clients/cli/__tests__/e2e.test.ts @@ -1,7 +1,9 @@ -import { describe, it, expect } from "vitest"; +import { describe, it, expect, beforeEach, afterEach } from "vitest"; import { spawn } from "node:child_process"; -import { resolve, dirname } from "node:path"; +import { resolve, dirname, join } from "node:path"; import { fileURLToPath } from "node:url"; +import { mkdtempSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; import { getTestMcpServerCommand } from "@modelcontextprotocol/inspector-test-server"; const here = dirname(fileURLToPath(import.meta.url)); @@ -42,10 +44,14 @@ interface SpawnResult { * coverage gate; this only asserts the binary boots and exits correctly. The * binary is built by the `pretest` / `test:coverage` scripts before tests run. */ -function spawnCli(args: string[]): Promise<SpawnResult> { +function spawnCli( + args: string[], + env?: Record<string, string>, +): Promise<SpawnResult> { return new Promise((resolvePromise, reject) => { const child = spawn("node", [BIN, ...args], { stdio: ["pipe", "pipe", "pipe"], + ...(env && { env: { ...process.env, ...env } }), detached: process.platform !== "win32", }); let stdout = ""; @@ -91,6 +97,125 @@ describe("CLI binary (out-of-process E2E)", () => { E2E_SPAWN_MS, ); + // #2435: a stdio server's stderr is inherited by default, so its banners and + // logs land in the CLI's stderr. A `--import` preload makes the real test + // server write one such line before it starts; `--` ends the target, since + // the preload flag would otherwise end it early. + describe("--quiet and a stdio server's own stderr", () => { + const NOISE = "SERVER_STDERR_NOISE_2435"; + const noisyTarget = [ + command, + "--import", + `data:text/javascript,${encodeURIComponent( + `process.stderr.write(${JSON.stringify(NOISE + "\n")});`, + )}`, + ...args, + "--", + ]; + + it( + "passes it through without --quiet", + async () => { + const result = await spawnCli([ + ...noisyTarget, + "--method", + "tools/list", + ]); + + expect(result.exitCode).toBe(0); + expect(result.stderr).toContain(NOISE); + }, + E2E_SPAWN_MS, + ); + + it( + "suppresses it under -q, leaving only the result on stdout", + async () => { + const result = await spawnCli([ + ...noisyTarget, + "-q", + "--method", + "tools/list", + ]); + + expect(result.exitCode).toBe(0); + expect(result.stderr).toBe(""); + expect(Array.isArray(JSON.parse(result.stdout).tools)).toBe(true); + }, + E2E_SPAWN_MS, + ); + }); + + // #2435 (Copilot on #2576): core announces the secret store it picked with + // `console.warn` on first use. That is out of reach of the in-process + // runner — Vitest replaces `console`, so the notice never passes through the + // patched `process.stderr.write` — hence the real binary. `--relogin` is the + // cheapest run that reaches the store; the cli project pins + // MCP_INSPECTOR_SECRET_STORE=memory (inherited here), whose caveat is the + // notice. Port 9 (discard) refuses at once, so the run fails at connect and + // the envelope is the only thing `--quiet` should leave. The state path is a + // temp file so the relogin never touches the real store. + describe("--quiet and the secret-store notice", () => { + const relogin = [ + "--relogin", + "--server-url", + "http://127.0.0.1:9/mcp", + "--method", + "tools/list", + ]; + let dir: string; + const env = () => ({ + MCP_INSPECTOR_OAUTH_STATE_PATH: join(dir, "oauth.json"), + HTTP_PROXY: "", + HTTPS_PROXY: "", + http_proxy: "", + https_proxy: "", + }); + beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), "cli-e2e-quiet-store-")); + }); + afterEach(() => { + rmSync(dir, { recursive: true, force: true }); + }); + + it( + "prints it without --quiet", + async () => { + const result = await spawnCli(relogin, env()); + + expect(result.exitCode).not.toBe(0); + expect(result.stderr).toContain("[mcp-inspector]"); + }, + E2E_SPAWN_MS, + ); + + it( + "drops it under -q, leaving only the envelope", + async () => { + const result = await spawnCli(["-q", ...relogin], env()); + + expect(result.exitCode).not.toBe(0); + const lines = result.stderr.trimEnd().split("\n"); + expect(lines).toHaveLength(1); + expect(JSON.parse(lines[0]!)).toHaveProperty("error"); + }, + E2E_SPAWN_MS, + ); + }); + + it( + "writes only the error envelope on a usage error under -q", + async () => { + const result = await spawnCli([command, ...args, "-q", "--method"]); + + expect(result.exitCode).not.toBe(0); + const lines = result.stderr.trimEnd().split("\n"); + expect(lines).toHaveLength(1); + expect(JSON.parse(lines[0]!)).toHaveProperty("error"); + }, + E2E_SPAWN_MS, + ); + it( "exits non-zero when required --method is missing", async () => { diff --git a/clients/cli/__tests__/emit-result.test.ts b/clients/cli/__tests__/emit-result.test.ts index dfa16b379d..1776e385ce 100644 --- a/clients/cli/__tests__/emit-result.test.ts +++ b/clients/cli/__tests__/emit-result.test.ts @@ -1,6 +1,6 @@ import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; import { emitResult, collectAppInfo } from "../src/cli.js"; -import { CliExitCodeError } from "../src/error-handler.js"; +import { CliExitCodeError } from "@inspector/core/cli/error-handler.js"; import type { InspectorClient } from "@inspector/core/mcp/index.js"; /** diff --git a/clients/cli/__tests__/error-handler.test.ts b/clients/cli/__tests__/error-handler.test.ts index 342a7b4b95..b08a9d28d4 100644 --- a/clients/cli/__tests__/error-handler.test.ts +++ b/clients/cli/__tests__/error-handler.test.ts @@ -5,7 +5,7 @@ import { classifyError, formatErrorOutput, handleError, -} from "../src/error-handler.js"; +} from "@inspector/core/cli/error-handler.js"; import { UnauthorizedError } from "@modelcontextprotocol/client"; import { SecretFileLockHeldError, diff --git a/clients/cli/__tests__/format-output.test.ts b/clients/cli/__tests__/format-output.test.ts index 14d633207a..3764bce340 100644 --- a/clients/cli/__tests__/format-output.test.ts +++ b/clients/cli/__tests__/format-output.test.ts @@ -1,5 +1,5 @@ import { describe, it, expect, afterEach } from "vitest"; -import { writeFormattedResult } from "../src/handlers/format-output.js"; +import { writeFormattedResult } from "@inspector/core/cli/handlers/format-output.js"; describe("writeFormattedResult", () => { let originalWrite: typeof process.stdout.write; diff --git a/clients/cli/__tests__/helpers/cli-runner.ts b/clients/cli/__tests__/helpers/cli-runner.ts index d5c8abe66e..af0cdb1220 100644 --- a/clients/cli/__tests__/helpers/cli-runner.ts +++ b/clients/cli/__tests__/helpers/cli-runner.ts @@ -1,5 +1,5 @@ import { runCli as invokeCli } from "../../src/cli.js"; -import { formatErrorOutput } from "../../src/error-handler.js"; +import { formatErrorOutput } from "@inspector/core/cli/error-handler.js"; export interface CliResult { exitCode: number | null; @@ -45,7 +45,7 @@ function captureWrite(append: (text: string) => void) { /** * Run the CLI **in-process** by importing and invoking `runCli` directly, so - * its source (`clients/cli/src/**`) is measured under vitest's coverage + * its source (`clients/cli/src/**`, plus the shared `core/cli/**`) is measured under vitest's coverage * instrumentation. The previous implementation spawned `build/index.js` as a * subprocess, which left CLI source invisible to coverage (#1484). * diff --git a/clients/cli/__tests__/helpers/mock-open-url.ts b/clients/cli/__tests__/helpers/mock-open-url.ts index 0b8daebb30..f5c1334d53 100644 --- a/clients/cli/__tests__/helpers/mock-open-url.ts +++ b/clients/cli/__tests__/helpers/mock-open-url.ts @@ -5,6 +5,6 @@ import { vi } from "vitest"; * shell out to a real browser. Registered via vitest `setupFiles` so it does * not depend on import order in individual test files. */ -vi.mock("../../src/open-url.js", () => ({ +vi.mock("@inspector/core/node/openUrl.js", () => ({ openUrl: vi.fn().mockResolvedValue(undefined), })); diff --git a/clients/cli/__tests__/helpers/oauth-test-fakes.ts b/clients/cli/__tests__/helpers/oauth-test-fakes.ts index 8c5e91f9eb..46a8a00837 100644 --- a/clients/cli/__tests__/helpers/oauth-test-fakes.ts +++ b/clients/cli/__tests__/helpers/oauth-test-fakes.ts @@ -4,7 +4,7 @@ import { DEFAULT_TASK_TTL_MS, } from "@inspector/core/mcp/types.js"; import type { InspectorServerSettings } from "@inspector/core/mcp/types.js"; -import type { CliOAuthClient } from "../../src/cliOAuth.js"; +import type { CliOAuthClient } from "@inspector/core/cli/cliOAuth.js"; /** * Typed mock factories for the CLI OAuth tests. They exist so a test can supply diff --git a/clients/cli/__tests__/method-types.test.ts b/clients/cli/__tests__/method-types.test.ts index 85230043d8..3b40bab0fb 100644 --- a/clients/cli/__tests__/method-types.test.ts +++ b/clients/cli/__tests__/method-types.test.ts @@ -2,16 +2,18 @@ import { describe, it, expect } from "vitest"; import { isOneShotMethod, ONE_SHOT_METHODS, - SESSION_RPC_METHODS, -} from "../src/handlers/method-types.js"; + CONNECTION_RPC_METHODS, +} from "@inspector/core/cli/handlers/method-types.js"; -describe("SESSION_RPC_METHODS", () => { +describe("CONNECTION_RPC_METHODS", () => { it("lists the full RPC method set supported by runMethod", () => { - expect(SESSION_RPC_METHODS).toContain("tools/list"); - expect(SESSION_RPC_METHODS).toContain("tools/call"); - expect(SESSION_RPC_METHODS).toContain("logging/tail"); - expect(SESSION_RPC_METHODS).toContain("roots/set"); - expect(new Set(SESSION_RPC_METHODS).size).toBe(SESSION_RPC_METHODS.length); + expect(CONNECTION_RPC_METHODS).toContain("tools/list"); + expect(CONNECTION_RPC_METHODS).toContain("tools/call"); + expect(CONNECTION_RPC_METHODS).toContain("logging/tail"); + expect(CONNECTION_RPC_METHODS).toContain("roots/set"); + expect(new Set(CONNECTION_RPC_METHODS).size).toBe( + CONNECTION_RPC_METHODS.length, + ); }); }); diff --git a/clients/cli/__tests__/oauth-interactive.test.ts b/clients/cli/__tests__/oauth-interactive.test.ts index abc5198264..b73a725cef 100644 --- a/clients/cli/__tests__/oauth-interactive.test.ts +++ b/clients/cli/__tests__/oauth-interactive.test.ts @@ -28,7 +28,7 @@ import { NodeOAuthStorage } from "@inspector/core/auth/node/index.js"; import { connectInspectorWithOAuth, withCliAuthRecoveryRetry, -} from "../src/cliOAuth.js"; +} from "@inspector/core/cli/cliOAuth.js"; import type { MCPServerConfig } from "@inspector/core/mcp/types.js"; import { makeFakeServerSettings } from "./helpers/oauth-test-fakes.js"; diff --git a/clients/cli/__tests__/output-file.test.ts b/clients/cli/__tests__/output-file.test.ts new file mode 100644 index 0000000000..d78d2a0b5c --- /dev/null +++ b/clients/cli/__tests__/output-file.test.ts @@ -0,0 +1,379 @@ +import { afterAll, beforeAll, describe, expect, it } from "vitest"; +import { mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { getTestMcpServerCommand } from "@modelcontextprotocol/inspector-test-server"; +import { runCli } from "./helpers/cli-runner.js"; +import { NO_SERVER_SENTINEL } from "./helpers/fixtures.js"; +import { + parseOutputFileFormat, + renderRaw, + renderResultForFile, + validateOutputOptions, + writeResultFile, +} from "@inspector/core/cli/handlers/output-file.js"; +import { CliExitCodeError } from "@inspector/core/cli/error-handler.js"; +import { emitResult } from "../src/handlers/emit-result.js"; + +/** + * `--output <path>` / `--output-format raw|json` (#2431): the pure rules and + * rendering are unit-tested directly; the wiring through `parseArgs` and + * `emitResult` is driven end to end against the stdio test server. + */ + +let dir: string; + +beforeAll(() => { + dir = mkdtempSync(join(tmpdir(), "inspector-cli-output-")); +}); + +afterAll(() => { + rmSync(dir, { recursive: true, force: true }); +}); + +describe("parseOutputFileFormat", () => { + it("accepts raw and json", () => { + expect(parseOutputFileFormat("raw")).toBe("raw"); + expect(parseOutputFileFormat("json")).toBe("json"); + }); + + it("rejects anything else", () => { + expect(() => parseOutputFileFormat("text")).toThrow( + "--output-format must be 'raw' or 'json'.", + ); + }); +}); + +describe("validateOutputOptions", () => { + it("accepts no --output at all", () => { + expect(() => validateOutputOptions({ method: "tools/call" })).not.toThrow(); + }); + + it("rejects --output-format without --output", () => { + expect(() => + validateOutputOptions({ method: "tools/call", outputFormat: "raw" }), + ).toThrow("--output-format requires --output <path>."); + }); + + it("rejects an empty path", () => { + expect(() => + validateOutputOptions({ method: "tools/call", output: " " }), + ).toThrow("--output requires a non-empty file path."); + }); + + it.each([ + { listStoredAuth: true }, + { printHandoff: true }, + { method: "servers/list" }, + { method: "servers/show" }, + ])("rejects a path that never calls a server (%o)", (extra) => { + expect(() => validateOutputOptions({ output: "x.json", ...extra })).toThrow( + "--output requires a method that calls a server", + ); + }); + + it.each([{ appInfo: true }, { verify: true }])( + "rejects a report-emitting flag (%o)", + (extra) => { + expect(() => + validateOutputOptions({ + output: "x.json", + method: "tools/call", + ...extra, + }), + ).toThrow("--output cannot be combined with --app-info or --verify"); + }, + ); + + it("rejects raw for a method with no raw payload", () => { + expect(() => + validateOutputOptions({ + output: "x.txt", + outputFormat: "raw", + method: "tools/list", + }), + ).toThrow( + "--output-format raw requires --method tools/call or resources/read", + ); + }); + + it("accepts raw for tools/call and resources/read, json for anything", () => { + for (const method of ["tools/call", "resources/read"]) { + expect(() => + validateOutputOptions({ output: "x", outputFormat: "raw", method }), + ).not.toThrow(); + } + expect(() => + validateOutputOptions({ + output: "x", + outputFormat: "json", + method: "tools/list", + }), + ).not.toThrow(); + }); +}); + +describe("renderRaw", () => { + const png = Buffer.from([0x89, 0x50, 0x4e, 0x47]); + + it("joins the text of every text-bearing block, skipping the rest", () => { + const out = renderRaw( + { + content: [ + { type: "text", text: "one" }, + { type: "image", data: png.toString("base64"), mimeType: "x" }, + { type: "resource", resource: { uri: "a://b", text: "two" } }, + { type: "resource_link", uri: "a://c", name: "c" }, + "not-a-block", + ], + }, + "tools/call", + ); + expect(out).toBe("one\ntwo"); + }); + + it("decodes a single binary block when there is no text", () => { + const out = renderRaw( + { + content: [ + { type: "image", data: png.toString("base64"), mimeType: "x" }, + ], + }, + "tools/call", + ); + expect(Buffer.isBuffer(out) && out.equals(png)).toBe(true); + }); + + it("reads resources/read contents (text and blob)", () => { + expect( + renderRaw({ contents: [{ uri: "a://b", text: "hi" }] }, "resources/read"), + ).toBe("hi"); + const blob = renderRaw( + { contents: [{ uri: "a://b", blob: png.toString("base64") }] }, + "resources/read", + ); + expect(Buffer.isBuffer(blob) && blob.equals(png)).toBe(true); + }); + + it("refuses a result with no payload", () => { + try { + renderRaw({ structuredContent: { a: 1 } }, "tools/call"); + expect.unreachable(); + } catch (err) { + expect(err).toBeInstanceOf(CliExitCodeError); + expect((err as CliExitCodeError).exitCode).toBe(1); + expect((err as CliExitCodeError).envelope?.code).toBe("output_not_raw"); + expect((err as Error).message).toContain("no text or binary content"); + } + }); + + it("refuses a single binary block that is not valid base64", () => { + try { + renderRaw( + { content: [{ type: "image", data: "not!base64!?", mimeType: "x" }] }, + "tools/call", + ); + expect.unreachable(); + } catch (err) { + expect(err).toBeInstanceOf(CliExitCodeError); + expect((err as CliExitCodeError).envelope?.code).toBe("output_not_raw"); + expect((err as Error).message).toContain("not valid base64"); + } + }); + + it("refuses several binaries with no text", () => { + expect(() => + renderRaw( + { + content: [ + { type: "audio", data: "AA==", mimeType: "x" }, + { type: "audio", data: "AQ==", mimeType: "x" }, + ], + }, + "tools/call", + ), + ).toThrow("has 2 binary blocks and no text"); + }); +}); + +describe("renderResultForFile / writeResultFile", () => { + it("renders json as the pretty-printed result with a trailing newline", () => { + expect(renderResultForFile({ a: [1] }, "tools/list", "json")).toBe( + '{\n "a": [\n 1\n ]\n}\n', + ); + }); + + it("defaults to json and reports what it wrote", async () => { + const path = join(dir, "default.json"); + const written = await writeResultFile({ ok: true }, "tools/list", path); + expect(written).toEqual({ path, format: "json", bytes: 17 }); + expect(JSON.parse(readFileSync(path, "utf8"))).toEqual({ ok: true }); + }); + + it("counts bytes of a binary payload", async () => { + const path = join(dir, "bin.dat"); + const written = await writeResultFile( + { content: [{ type: "image", data: "AAEC", mimeType: "x" }] }, + "tools/call", + path, + "raw", + ); + expect(written.bytes).toBe(3); + expect([...readFileSync(path)]).toEqual([0, 1, 2]); + }); + + it("maps a failed write to output_write_failed", async () => { + const path = join(dir, "missing-dir", "x.json"); + const promise = writeResultFile({}, "tools/list", path); + await expect(promise).rejects.toBeInstanceOf(CliExitCodeError); + await promise.catch((err: CliExitCodeError) => { + expect(err.exitCode).toBe(1); + expect(err.envelope?.code).toBe("output_write_failed"); + expect(err.message).toContain(`Could not write --output file ${path}`); + }); + }); +}); + +describe("--output end to end", () => { + const { command, args } = getTestMcpServerCommand(); + const echo = (...extra: string[]) => + runCli([ + command, + ...args, + "--cli", + "--method", + "tools/call", + "--tool-name", + "echo", + "--tool-arg", + "message=hello", + ...extra, + ]); + + it("writes the whole result as json, keeping stdout empty", async () => { + const path = join(dir, "echo.json"); + const result = await echo("--output", path); + expect(result.exitCode).toBe(0); + expect(result.stdout).toBe(""); + expect(result.stderr).toContain(`(json) to ${path}`); + const saved = JSON.parse(readFileSync(path, "utf8")); + expect(saved.content[0].text).toContain("hello"); + }); + + it("writes raw text", async () => { + const path = join(dir, "echo.txt"); + const result = await echo("--output", path, "--output-format", "raw"); + expect(result.exitCode).toBe(0); + const text = readFileSync(path, "utf8"); + expect(text).toContain("hello"); + expect(() => JSON.parse(text)).toThrow(); + }); + + it("puts an { output } envelope on stdout under --format json", async () => { + const path = join(dir, "echo-envelope.json"); + const result = await echo("--output", path, "--format", "json"); + expect(result.exitCode).toBe(0); + const envelope = JSON.parse(result.stdout.trim()); + expect(envelope).not.toHaveProperty("result"); + expect(envelope.output).toMatchObject({ path, format: "json" }); + expect(envelope.output.bytes).toBe(readFileSync(path).length); + }); + + it("writes only the text when a result mixes text and an image", async () => { + const path = join(dir, "mixed.txt"); + const result = await runCli([ + command, + ...args, + "--cli", + "--method", + "tools/call", + "--tool-name", + "get_annotated_message", + "--tool-arg", + "messageType=success", + "includeImage=true", + "--output", + path, + "--output-format", + "raw", + ]); + // The tool also returns a text block, so raw writes the text — the + // binary-only path is pinned by the unit tests above. + expect(result.exitCode).toBe(0); + expect(readFileSync(path, "utf8").length).toBeGreaterThan(0); + }); + + it("writes resources/read text raw", async () => { + const path = join(dir, "env.json"); + const result = await runCli([ + command, + ...args, + "--cli", + "--method", + "resources/read", + "--uri", + "test://env", + "--output", + path, + "--output-format", + "raw", + ]); + expect(result.exitCode).toBe(0); + // The resource's own text is a JSON object of the env, written verbatim. + expect(typeof JSON.parse(readFileSync(path, "utf8"))).toBe("object"); + }); + + it("rejects --output on a catalog method before connecting", async () => { + const result = await runCli([ + NO_SERVER_SENTINEL, + "--cli", + "--method", + "servers/list", + "--output", + join(dir, "never.json"), + ]); + expect(result.exitCode).toBe(1); + expect(result.stderr).toContain("--output requires a method that calls"); + }); + + it("fails with output_write_failed when the directory is missing", async () => { + const result = await echo("--output", join(dir, "nope", "x.json")); + expect(result.exitCode).toBe(1); + expect(result.stderr).toContain("output_write_failed"); + }); +}); + +describe("emitResult with --output", () => { + it("folds appInfo beside { output } and still exits TOOL_ERROR on isError", async () => { + const path = join(dir, "is-error.json"); + let stdout = ""; + const original = process.stdout.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]): boolean => { + stdout += String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + try { + const promise = emitResult( + { content: [{ type: "text", text: "boom" }], isError: true }, + { hasApp: true, toolName: "t", resourceUri: "ui://t" }, + { toolName: "t", format: "json", output: path }, + ); + await expect(promise).rejects.toBeInstanceOf(CliExitCodeError); + await promise.catch((err: CliExitCodeError) => { + expect(err.exitCode).toBe(5); + }); + } finally { + process.stdout.write = original; + } + // The result is written before the exit, and method-less args fall back + // to the json rendering. + expect(JSON.parse(readFileSync(path, "utf8")).isError).toBe(true); + const envelope = JSON.parse(stdout.trim()); + expect(envelope.output.path).toBe(path); + expect(envelope.appInfo).toMatchObject({ hasApp: true }); + }); +}); diff --git a/clients/cli/__tests__/programmatic-ergonomics.test.ts b/clients/cli/__tests__/programmatic-ergonomics.test.ts index 8459644e81..83da27d129 100644 --- a/clients/cli/__tests__/programmatic-ergonomics.test.ts +++ b/clients/cli/__tests__/programmatic-ergonomics.test.ts @@ -175,7 +175,7 @@ describe("CLI --relogin flag conflicts", () => { it("rejects --relogin with catalog-only methods", async () => { const result = await runCli(["--relogin", "--method", "servers/list"]); expectCliFailure(result); - expect(result.stderr).toMatch(/servers\/list or servers\/show/); + expect(result.stderr).toMatch(/servers\/\* catalog command/); }); it("rejects --relogin with --list-stored-auth", async () => { @@ -299,6 +299,86 @@ describe("--protocol-era", () => { }); }); +describe("--skill-catalog-max-skills / --skill-catalog-max-bytes (#2420)", () => { + // The flags must reach `verifySkills`' run-level budget, not just parse. A + // budget of one skill leaves every later skill unread, which the report + // records as an `incomplete` outcome — without the flag the fixture's whole + // catalog fits the default budget and no skill is incomplete. + it("caps the --verify run at the flag's skill budget", async () => { + const server = createTestServerHttp({ + serverInfo: createTestServerInfo(), + skills: true, + }); + try { + await server.start(); + const verify = [ + server.url, + "--transport", + "http", + "--method", + "skills/list", + "--verify", + ]; + + const unbounded = await runCli(verify); + expect(reportOutcomes(unbounded.stdout)).not.toContain("incomplete"); + + const capped = await runCli([ + ...verify, + "--skill-catalog-max-skills", + "1", + ]); + expect(reportOutcomes(capped.stdout)).toContain("incomplete"); + } finally { + await server.stop(); + } + }); + + it.each([ + ["--skill-catalog-max-skills", "0"], + ["--skill-catalog-max-bytes", "1.5"], + ])("rejects %s %s before connecting", async (flag, value) => { + const { command, args } = getTestMcpServerCommand(); + const result = await runCli([ + command, + ...args, + flag, + value, + "--method", + "tools/list", + ]); + expectCliFailure(result); + expect(result.stderr).toContain( + `Invalid ${flag}: ${value}. Expected a positive integer written as plain decimal digits, at most 9007199254740991.`, + ); + }); + + it.each(["--skill-catalog-max-skills", "--skill-catalog-max-bytes"])( + "rejects %s without --verify, where it would bound nothing", + async (flag) => { + const { command, args } = getTestMcpServerCommand(); + const result = await runCli([ + command, + ...args, + flag, + "5", + "--method", + "tools/list", + ]); + expectCliFailure(result); + expect(result.stderr).toContain(`${flag} requires --verify.`); + }, + ); +}); + +/** The `outcome` of each NDJSON `--verify` report line. */ +function reportOutcomes(stdout: string): string[] { + return stdout + .trim() + .split("\n") + .map((line) => (JSON.parse(line) as { outcome: string }).outcome); +} + describe("MCP_CATALOG_PATH with an ad-hoc target", () => { it("does not conflict with an ad-hoc target (env catalog is ignored)", async () => { const { command, args } = getTestMcpServerCommand(); diff --git a/clients/cli/__tests__/relogin-revocation.test.ts b/clients/cli/__tests__/relogin-revocation.test.ts index eb1698ead4..46025c2a84 100644 --- a/clients/cli/__tests__/relogin-revocation.test.ts +++ b/clients/cli/__tests__/relogin-revocation.test.ts @@ -224,4 +224,23 @@ describe("--relogin token revocation", () => { expect(result.stderr).toMatch(/could not revoke the OAuth grant/i); }); + + // #2435: the warning is advisory, so `--quiet` drops it. The relogin itself + // (and the revocation attempt) still happen. + it("drops the revocation warning under --quiet", async () => { + seedStore(); + const fetchSpy = stubFetch(500); + + const result = await runCli([ + "--relogin", + "--quiet", + "--server-url", + SERVER_URL, + "--method", + "tools/list", + ]); + + expect(revocationCalls(fetchSpy)).toHaveLength(1); + expect(result.stderr).not.toMatch(/could not revoke/i); + }); }); diff --git a/clients/cli/__tests__/run-method-mocks.test.ts b/clients/cli/__tests__/run-method-mocks.test.ts index 4fda0c5197..e5df08b2b5 100644 --- a/clients/cli/__tests__/run-method-mocks.test.ts +++ b/clients/cli/__tests__/run-method-mocks.test.ts @@ -1,5 +1,5 @@ import { describe, it, expect, vi } from "vitest"; -import { runMethod } from "../src/handlers/run-method.js"; +import { runMethod } from "@inspector/core/cli/handlers/run-method.js"; import type { InspectorClient } from "@inspector/core/mcp/index.js"; function mockClient(overrides: Partial<InspectorClient> = {}): InspectorClient { @@ -12,6 +12,7 @@ function mockClient(overrides: Partial<InspectorClient> = {}): InspectorClient { getRequestorTask: vi.fn().mockResolvedValue({ taskId: "t1" }), cancelRequestorTask: vi.fn().mockResolvedValue(undefined), getRequestorTaskResult: vi.fn().mockResolvedValue({ content: [] }), + updateRequestorTask: vi.fn().mockResolvedValue(undefined), getRoots: vi.fn().mockReturnValue([]), setRoots: vi.fn().mockResolvedValue(undefined), setLoggingLevel: vi.fn().mockResolvedValue(undefined), @@ -65,6 +66,192 @@ vi.mock("@inspector/core/mcp/state/index.js", async (importOriginal) => { }); describe("runMethod (mocked client)", () => { + it("reference-counts same-URI subscribe streams", async () => { + const client = mockClient(); + const s1 = await runMethod(client, { + method: "resources/subscribe", + uri: "test://x", + }); + const s2 = await runMethod(client, { + method: "resources/subscribe", + uri: "test://x", + }); + // Both streams share one core subscription. + expect(client.subscribeToResource).toHaveBeenCalledTimes(1); + expect(s1.kind).toBe("stream"); + expect(s2.kind).toBe("stream"); + if (s1.kind === "stream" && s2.kind === "stream") { + const stop1 = s1.start(() => {}); + const stop2 = s2.start(() => {}); + stop1(); + // The survivor keeps the subscription alive. + expect(client.unsubscribeFromResource).not.toHaveBeenCalled(); + stop2(); + expect(client.unsubscribeFromResource).toHaveBeenCalledTimes(1); + } + + // A rejected unsubscribe (e.g. after daemon disconnectAll) is caught at + // the source instead of surfacing as an unhandled rejection. + const failing = mockClient({ + unsubscribeFromResource: vi + .fn() + .mockRejectedValue(new Error("client closed")), + } as Partial<InspectorClient>); + const s3 = await runMethod(failing, { + method: "resources/subscribe", + uri: "test://y", + }); + if (s3.kind === "stream") { + s3.start(() => {})(); + } + await new Promise((resolve) => setImmediate(resolve)); + expect(failing.unsubscribeFromResource).toHaveBeenCalledTimes(1); + }); + + it("buffers updates that land between subscribe and start", async () => { + const listeners = new Set<(ev: Event) => void>(); + const client = mockClient({ + addEventListener: vi.fn((_type: string, fn: (ev: Event) => void) => + listeners.add(fn), + ), + removeEventListener: vi.fn((_type: string, fn: (ev: Event) => void) => + listeners.delete(fn), + ), + } as unknown as Partial<InspectorClient>); + const outcome = await runMethod(client, { + method: "resources/subscribe", + uri: "test://early", + }); + // The listener is live before any consumer starts the stream... + expect(listeners.size).toBe(1); + // ...so an update in the subscribe→start window is captured, not lost. + const dispatch = (uri: string) => { + for (const fn of listeners) + fn(new CustomEvent("resourceUpdated", { detail: { uri } })); + }; + dispatch("test://early"); + dispatch("test://other"); // different URI: filtered out + const lines: unknown[] = []; + expect(outcome.kind).toBe("stream"); + if (outcome.kind !== "stream") return; + const stop = outcome.start((obj) => lines.push(obj)); + expect(lines).toEqual([ + { type: "subscribed", uri: "test://early" }, + { type: "resources/updated", uri: "test://early" }, + ]); + // Post-start events flow straight through. + dispatch("test://early"); + expect(lines).toHaveLength(3); + stop(); + expect(listeners.size).toBe(0); + }); + + it("detaches the early listener when the subscribe fails", async () => { + const listeners = new Set<(ev: Event) => void>(); + const client = mockClient({ + addEventListener: vi.fn((_type: string, fn: (ev: Event) => void) => + listeners.add(fn), + ), + removeEventListener: vi.fn((_type: string, fn: (ev: Event) => void) => + listeners.delete(fn), + ), + subscribeToResource: vi.fn().mockRejectedValue(new Error("nope")), + } as unknown as Partial<InspectorClient>); + await expect( + runMethod(client, { method: "resources/subscribe", uri: "test://f" }), + ).rejects.toThrow("nope"); + expect(listeners.size).toBe(0); + }); + + it("concurrent same-URI subscribes share one in-flight subscription", async () => { + let release!: () => void; + const gate = new Promise<void>((resolve) => (release = resolve)); + const client = mockClient({ + subscribeToResource: vi.fn().mockImplementation(() => gate), + } as Partial<InspectorClient>); + // Both setups race before the subscribe resolves; the reservation is + // synchronous, so they must join one in-flight subscribe rather than + // each subscribing and writing a count of 1. + const p1 = runMethod(client, { + method: "resources/subscribe", + uri: "test://race", + }); + const p2 = runMethod(client, { + method: "resources/subscribe", + uri: "test://race", + }); + release(); + const [s1, s2] = await Promise.all([p1, p2]); + expect(client.subscribeToResource).toHaveBeenCalledTimes(1); + const stop1 = s1.kind === "stream" ? s1.start(() => {}) : () => {}; + const stop2 = s2.kind === "stream" ? s2.start(() => {}) : () => {}; + stop1(); + stop1(); // double-stop must not corrupt the shared count + expect(client.unsubscribeFromResource).not.toHaveBeenCalled(); + stop2(); + expect(client.unsubscribeFromResource).toHaveBeenCalledTimes(1); + }); + + it("rolls back reservations when the shared subscribe fails, allowing retry", async () => { + const subscribe = vi + .fn() + .mockRejectedValueOnce(new Error("subscribe boom")) + .mockResolvedValue(undefined); + const client = mockClient({ + subscribeToResource: subscribe, + } as Partial<InspectorClient>); + const p1 = runMethod(client, { + method: "resources/subscribe", + uri: "test://fail", + }); + const p2 = runMethod(client, { + method: "resources/subscribe", + uri: "test://fail", + }); + await expect(p1).rejects.toThrow("subscribe boom"); + await expect(p2).rejects.toThrow("subscribe boom"); + // Both joined the same failed attempt… + expect(subscribe).toHaveBeenCalledTimes(1); + // …and both rolled back, so a retry issues a fresh subscribe. + const s3 = await runMethod(client, { + method: "resources/subscribe", + uri: "test://fail", + }); + expect(subscribe).toHaveBeenCalledTimes(2); + expect(s3.kind).toBe("stream"); + }); + + it("rejects explicit unsubscribe while subscribe streams share the URI", async () => { + const client = mockClient(); + const sub = await runMethod(client, { + method: "resources/subscribe", + uri: "test://shared", + }); + expect(sub.kind).toBe("stream"); + const stop = sub.kind === "stream" ? sub.start(() => {}) : () => {}; + + // Tearing down the shared subscription out from under the open stream + // (and double-unsubscribing later) is refused with guidance. + await expect( + runMethod(client, { + method: "resources/unsubscribe", + uri: "test://shared", + }), + ).rejects.toThrow(/active resources\/subscribe stream/); + expect(client.unsubscribeFromResource).not.toHaveBeenCalled(); + + // Once the last stream closes, its cleanup unsubscribes and an explicit + // unsubscribe is allowed again. + stop(); + expect(client.unsubscribeFromResource).toHaveBeenCalledTimes(1); + const out = await runMethod(client, { + method: "resources/unsubscribe", + uri: "test://shared", + }); + expect(out.kind).toBe("result"); + expect(client.unsubscribeFromResource).toHaveBeenCalledTimes(2); + }); + it("covers subscribe stream, tasks, complete, and app-info call", async () => { const client = mockClient({ callTool: vi.fn().mockResolvedValue({ @@ -104,11 +291,14 @@ describe("runMethod (mocked client)", () => { listener?.( new CustomEvent("resourceUpdated", { detail: { uri: "test://x" } }), ); - expect( - lines.some( - (l) => (l as { type?: string }).type === "resources/updated", - ), - ).toBe(true); + // Updates for other URIs on the same connection are filtered out. + listener?.( + new CustomEvent("resourceUpdated", { detail: { uri: "test://other" } }), + ); + const updated = lines.filter( + (l) => (l as { type?: string }).type === "resources/updated", + ); + expect(updated).toEqual([{ type: "resources/updated", uri: "test://x" }]); stop(); } @@ -133,6 +323,19 @@ describe("runMethod (mocked client)", () => { }); expect(result.kind).toBe("result"); + const updated = await runMethod(client, { + method: "tasks/update", + taskId: "t1", + inputResponsesJson: '{"confirm":{"approved":true}}', + }); + expect(updated.kind).toBe("result"); + if (updated.kind === "result") { + expect(updated.result).toMatchObject({ updated: true, taskId: "t1" }); + } + expect(client.updateRequestorTask).toHaveBeenCalledWith("t1", { + confirm: { approved: true }, + }); + const complete = await runMethod(client, { method: "prompts/complete", completeRefType: "ref/prompt", @@ -191,6 +394,36 @@ describe("runMethod (mocked client)", () => { /tasks\/result/, ); + await expect(runMethod(client, { method: "tasks/update" })).rejects.toThrow( + /tasks\/update/, + ); + await expect( + runMethod(client, { method: "tasks/update", taskId: "t1" }), + ).rejects.toThrow(/--input-responses/); + await expect( + runMethod(client, { + method: "tasks/update", + taskId: "t1", + inputResponsesJson: "not-json", + }), + ).rejects.toThrow(/--input-responses is invalid/); + await expect( + runMethod(client, { + method: "tasks/update", + taskId: "t1", + inputResponsesJson: "[1,2,3]", + }), + ).rejects.toThrow(/--input-responses is invalid/); + // 1e999 parses as Infinity, which serialization would silently send as + // null — reject it instead of answering with a different value. + await expect( + runMethod(client, { + method: "tasks/update", + taskId: "t1", + inputResponsesJson: '{"a":{"b":[1e999]}}', + }), + ).rejects.toThrow(/no JSON representation/); + await expect( runMethod(client, { method: "roots/set", diff --git a/clients/cli/__tests__/run-method-skills.test.ts b/clients/cli/__tests__/run-method-skills.test.ts index 9931995db2..7b1d239848 100644 --- a/clients/cli/__tests__/run-method-skills.test.ts +++ b/clients/cli/__tests__/run-method-skills.test.ts @@ -1,10 +1,10 @@ import { describe, it, expect, vi } from "vitest"; -import { runMethod } from "../src/handlers/run-method.js"; +import { runMethod } from "@inspector/core/cli/handlers/run-method.js"; import { skillVerificationExitCode, summarizeSkillVerification, -} from "../src/handlers/skills-verify.js"; -import { EXIT_CODES } from "../src/error-handler.js"; +} from "@inspector/core/cli/handlers/skills-verify.js"; +import { EXIT_CODES } from "@inspector/core/cli/error-handler.js"; import type { InspectorClient } from "@inspector/core/mcp/index.js"; import type { SkillEntry } from "@inspector/core/mcp/skillsSchemas.js"; import { sha256Digest } from "@inspector/core/mcp/skills.js"; diff --git a/clients/cli/__tests__/run-method.test.ts b/clients/cli/__tests__/run-method.test.ts index d859922366..1c27fac5e3 100644 --- a/clients/cli/__tests__/run-method.test.ts +++ b/clients/cli/__tests__/run-method.test.ts @@ -2,7 +2,7 @@ import { describe, it, expect, afterEach, vi } from "vitest"; import { getTestMcpServerCommand } from "@modelcontextprotocol/inspector-test-server"; import { InspectorClient } from "@inspector/core/mcp/index.js"; import { createTransportNode } from "@inspector/core/mcp/node/index.js"; -import { runMethod } from "../src/handlers/run-method.js"; +import { runMethod } from "@inspector/core/cli/handlers/run-method.js"; import { consumeMethodOutcome } from "../src/handlers/consume-outcome.js"; describe("runMethod", () => { diff --git a/clients/cli/__tests__/schema-lint-report.test.ts b/clients/cli/__tests__/schema-lint-report.test.ts index 997e98fea6..b063752339 100644 --- a/clients/cli/__tests__/schema-lint-report.test.ts +++ b/clients/cli/__tests__/schema-lint-report.test.ts @@ -1,6 +1,9 @@ import { describe, it, expect, beforeEach, afterEach } from "vitest"; import { emitResult, runCli } from "../src/cli.js"; -import { CliExitCodeError, EXIT_CODES } from "../src/error-handler.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; import { lintListResult, toolsFromResult, @@ -85,6 +88,16 @@ describe("writeSchemaLintReport", () => { expect(stderr).toContain("Re-run with --strict"); }); + it("writes nothing for the one-line hint under --quiet (#2435)", async () => { + await writeSchemaLintReport(lintListResult(DIRTY_LIST), false, true); + expect(stderr).toBe(""); + }); + + it("still writes the --strict report under --quiet — it was asked for", async () => { + await writeSchemaLintReport(lintListResult(DIRTY_LIST), true, true); + expect(stderr).toContain("Path: outputSchema.properties.data"); + }); + it("writes the full report under --strict", async () => { await writeSchemaLintReport(lintListResult(DIRTY_LIST), true); expect(stderr).toContain('Error: tool "info"'); @@ -200,6 +213,16 @@ describe("emitResult — schema lint wiring (#1005)", () => { expect(stderr).not.toContain("Suggestion:"); }); + it("prints only the result under --quiet without --strict (#2435)", async () => { + await emitResult(DIRTY_LIST, undefined, { + method: "tools/list", + format: "text", + quiet: true, + }); + expect(stderr).toBe(""); + expect(JSON.parse(stdout)).toEqual(DIRTY_LIST); + }); + it("folds findings into the --format json envelope under --strict", async () => { await emitResult(DIRTY_LIST, undefined, { method: "tools/list", diff --git a/clients/cli/__tests__/servers-list.test.ts b/clients/cli/__tests__/servers-list.test.ts index daa1c44418..41d34447cd 100644 --- a/clients/cli/__tests__/servers-list.test.ts +++ b/clients/cli/__tests__/servers-list.test.ts @@ -7,13 +7,14 @@ import { } from "./helpers/fixtures.js"; import { expectCliSuccess } from "./helpers/assertions.js"; import { - annotateServerEntriesWithSessions, + annotateServerEntriesWithConnections, listServerEntries, + resolveServerListSource, sanitizeServerConfig, sanitizeServerSettings, showServerEntry, summarizeServerConfig, -} from "../src/handlers/servers-list.js"; +} from "@inspector/core/cli/handlers/servers-list.js"; import type { InspectorServerSettings, MCPServerConfig, @@ -60,38 +61,71 @@ describe("summarizeServerConfig", () => { }); }); -describe("annotateServerEntriesWithSessions", () => { +describe("annotateServerEntriesWithConnections", () => { const entries = [ { name: "a", type: "stdio", detail: "node a" }, { name: "b", type: "stdio", detail: "node b" }, ]; - it("returns entries unchanged when there are no sessions", () => { - expect(annotateServerEntriesWithSessions(entries, [])).toBe(entries); + it("returns entries unchanged when there are no connections", () => { + expect(annotateServerEntriesWithConnections(entries, [])).toBe(entries); }); it("marks matching entry names and MRU", () => { expect( - annotateServerEntriesWithSessions(entries, [ + annotateServerEntriesWithConnections(entries, [ { name: "b", isMru: true }, { name: "other" }, ]), ).toEqual([ { name: "a", type: "stdio", detail: "node a" }, - { name: "b", type: "stdio", detail: "node b", session: "b", isMru: true }, + { + name: "b", + type: "stdio", + detail: "node b", + connection: "b", + isMru: true, + }, ]); }); - it("omits isMru when the session is not MRU", () => { + it("omits isMru when the connection is not MRU", () => { expect( - annotateServerEntriesWithSessions(entries, [{ name: "a", isMru: false }]), + annotateServerEntriesWithConnections(entries, [ + { name: "a", isMru: false }, + ]), ).toEqual([ - { name: "a", type: "stdio", detail: "node a", session: "a" }, + { name: "a", type: "stdio", detail: "node a", connection: "a" }, { name: "b", type: "stdio", detail: "node b" }, ]); }); }); +describe("resolveServerListSource", () => { + it("reports catalog (writable) vs config (read-only) with the resolved path", () => { + expect(resolveServerListSource({ catalogPath: "/tmp/cat.json" })).toEqual({ + kind: "catalog", + path: "/tmp/cat.json", + }); + expect(resolveServerListSource({ configPath: "/tmp/conf.json" })).toEqual({ + kind: "config", + path: "/tmp/conf.json", + }); + }); + + it("falls back to the default writable catalog when no source is given", () => { + const source = resolveServerListSource({}); + expect(source?.kind).toBe("catalog"); + expect(source?.path).toMatch(/mcp\.json$/); + }); + + it("is null for ad-hoc targets (no list source)", () => { + expect( + resolveServerListSource({ target: ["https://example.com/mcp"] }), + ).toBeNull(); + }); +}); + describe("listServerEntries / --method servers/list", () => { let configPath: string | undefined; diff --git a/clients/cli/__tests__/servers-write.test.ts b/clients/cli/__tests__/servers-write.test.ts new file mode 100644 index 0000000000..7d1f40ff25 --- /dev/null +++ b/clients/cli/__tests__/servers-write.test.ts @@ -0,0 +1,742 @@ +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import { + InMemorySecretStore, + SessionSecretStore, +} from "@inspector/core/auth/node/secret-store.js"; +import { getDefaultMcpConfigPath } from "@inspector/core/storage/store-io.js"; +import { + addCatalogServer, + editCatalogServer, + isCatalogWriteMethod, + removeCatalogServer, + resolveWritableCatalogPath, + runCatalogWrite, +} from "../src/handlers/servers-write.js"; +import { runCli } from "./helpers/cli-runner.js"; +import { expectCliFailure, expectCliSuccess } from "./helpers/assertions.js"; + +let dir: string; +let catalog: string; +let store: InMemorySecretStore; + +function readCatalog(): Record<string, Record<string, unknown>> { + return ( + JSON.parse(fs.readFileSync(catalog, "utf-8")) as { + mcpServers: Record<string, Record<string, unknown>>; + } + ).mcpServers; +} + +function writeCatalog(mcpServers: unknown): void { + fs.writeFileSync(catalog, JSON.stringify({ mcpServers }, null, 2)); +} + +beforeEach(() => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "cli-servers-write-")); + catalog = path.join(dir, "mcp.json"); + store = new InMemorySecretStore(); + // The route layer logs normalize/smuggle warnings through console.warn when + // no file logger is wired; keep the test output clean. + vi.spyOn(console, "warn").mockImplementation(() => {}); +}); + +afterEach(() => { + vi.restoreAllMocks(); + fs.rmSync(dir, { recursive: true, force: true }); +}); + +describe("isCatalogWriteMethod", () => { + it("recognises only the three write methods", () => { + expect(isCatalogWriteMethod("servers/add")).toBe(true); + expect(isCatalogWriteMethod("servers/edit")).toBe(true); + expect(isCatalogWriteMethod("servers/remove")).toBe(true); + expect(isCatalogWriteMethod("servers/list")).toBe(false); + expect(isCatalogWriteMethod(undefined)).toBe(false); + }); +}); + +describe("resolveWritableCatalogPath", () => { + it("prefers --catalog, then MCP_CATALOG_PATH, then the default", () => { + expect(resolveWritableCatalogPath({ catalog: catalog }, {})).toBe(catalog); + expect( + resolveWritableCatalogPath( + { catalog: " " }, + { MCP_CATALOG_PATH: catalog }, + ), + ).toBe(catalog); + expect(resolveWritableCatalogPath({}, {})).toBe( + path.resolve(getDefaultMcpConfigPath()), + ); + }); + + it("resolves a relative path against the current directory", () => { + expect(resolveWritableCatalogPath({ catalog: "rel.json" }, {})).toBe( + path.resolve(process.cwd(), "rel.json"), + ); + }); + + it("refuses a read-only --config", () => { + expect(() => resolveWritableCatalogPath({ config: catalog }, {})).toThrow( + /read-only/, + ); + }); + + it("refuses an explicitly blank --config too", () => { + expect(() => resolveWritableCatalogPath({ config: "" }, {})).toThrow( + /read-only/, + ); + }); +}); + +describe("addCatalogServer", () => { + it("adds a stdio entry with its env value split into the secret store", async () => { + const result = await addCatalogServer({ + server: "demo", + catalog, + target: ["node", "server.js"], + env: { API_KEY: "sekrit" }, + cwd: "/work", + secretStore: store, + }); + expect(result).toEqual({ + ok: true, + action: "added", + server: "demo", + catalog, + }); + expect(readCatalog()).toEqual({ + demo: { + type: "stdio", + command: "node", + args: ["server.js"], + env: { API_KEY: "" }, + cwd: "/work", + }, + }); + expect(await store.get("demo", "env:API_KEY")).toBe("sekrit"); + }); + + it("adds an HTTP entry with headers and a protocol era", async () => { + await addCatalogServer({ + server: "web", + catalog, + serverUrl: "https://example.com/mcp", + headers: { "X-Team": "a" }, + protocolEra: "modern", + secretStore: store, + }); + expect(readCatalog()).toEqual({ + web: { + type: "streamable-http", + url: "https://example.com/mcp", + headers: { "X-Team": "a" }, + protocolEra: "modern", + }, + }); + }); + + it("sets only the protocol era when no headers are given", async () => { + await addCatalogServer({ + server: "web", + catalog, + target: ["https://example.com/sse"], + protocolEra: "auto", + secretStore: store, + }); + expect(readCatalog().web).toEqual({ + type: "sse", + url: "https://example.com/sse", + protocolEra: "auto", + }); + }); + + it("preserves the other entries in the file", async () => { + writeCatalog({ keep: { type: "sse", url: "https://keep.example/sse" } }); + await addCatalogServer({ + server: "demo", + catalog, + target: ["node"], + secretStore: store, + }); + expect(Object.keys(readCatalog())).toEqual(["keep", "demo"]); + }); + + it("rejects a duplicate name with the route's message", async () => { + writeCatalog({ demo: { type: "stdio", command: "node" } }); + await expect( + addCatalogServer({ + server: "demo", + catalog, + target: ["node"], + secretStore: store, + }), + ).rejects.toThrow("Server 'demo' already exists"); + }); + + it("rejects an invalid name", async () => { + await expect( + addCatalogServer({ + server: "bad name", + catalog, + target: ["node"], + secretStore: store, + }), + ).rejects.toThrow(/Invalid id/); + }); + + it("requires --server, a target, and no --rename", async () => { + await expect( + addCatalogServer({ catalog, target: ["node"], secretStore: store }), + ).rejects.toThrow("servers/add requires --server <name>."); + await expect( + addCatalogServer({ server: "x", catalog, secretStore: store }), + ).rejects.toThrow(/requires a command or URL/); + await expect( + addCatalogServer({ + server: "x", + rename: "y", + catalog, + target: ["node"], + secretStore: store, + }), + ).rejects.toThrow(/--rename is only valid/); + }); + + it("rejects -e / --cwd for a URL target instead of dropping them", async () => { + await expect( + addCatalogServer({ + server: "web", + catalog, + serverUrl: "https://example.com/mcp", + env: { A: "1" }, + cwd: "/x", + secretStore: store, + }), + ).rejects.toThrow( + "-e and --cwd apply to stdio servers; this target is streamable-http.", + ); + expect(fs.existsSync(catalog)).toBe(false); + }); + + it("refuses -e values on a store that dies with the process", async () => { + await expect( + addCatalogServer({ + server: "demo", + catalog, + target: ["node"], + env: { API_KEY: "sekrit" }, + secretStore: new SessionSecretStore(), + }), + ).rejects.toThrow(/MCP_INSPECTOR_SECRET_STORE/); + expect(fs.existsSync(catalog)).toBe(false); + }); + + it("writes an entry without secrets on a session store", async () => { + await addCatalogServer({ + server: "demo", + catalog, + target: ["node"], + secretStore: new SessionSecretStore(), + }); + expect(readCatalog().demo).toEqual({ type: "stdio", command: "node" }); + }); +}); + +describe("editCatalogServer", () => { + beforeEach(async () => { + await addCatalogServer({ + server: "demo", + catalog, + target: ["node", "server.js"], + env: { API_KEY: "sekrit" }, + headers: { "X-Team": "a" }, + secretStore: store, + }); + }); + + it("merges -e into the env and keeps untouched secrets", async () => { + const result = await editCatalogServer({ + server: "demo", + catalog, + env: { OTHER: "1" }, + secretStore: store, + }); + expect(result).toEqual({ + ok: true, + action: "updated", + server: "demo", + catalog, + }); + expect(readCatalog().demo).toEqual({ + type: "stdio", + command: "node", + args: ["server.js"], + env: { API_KEY: "", OTHER: "" }, + headers: { "X-Team": "a" }, + }); + expect(await store.get("demo", "env:API_KEY")).toBe("sekrit"); + expect(await store.get("demo", "env:OTHER")).toBe("1"); + }); + + it("sets --cwd on its own", async () => { + await editCatalogServer({ + server: "demo", + catalog, + cwd: " /work ", + secretStore: store, + }); + expect(readCatalog().demo).toMatchObject({ cwd: "/work" }); + }); + + it("replaces the transport when given a new target, keeping settings", async () => { + await editCatalogServer({ + server: "demo", + catalog, + serverUrl: "https://example.com/sse", + transport: "sse", + secretStore: store, + }); + expect(readCatalog().demo).toEqual({ + type: "sse", + url: "https://example.com/sse", + headers: { "X-Team": "a" }, + }); + // The env key is no longer part of the entry, so its secret is retired. + expect(await store.get("demo", "env:API_KEY")).toBeNull(); + }); + + it("replaces headers and sets the era, preserving the env", async () => { + await editCatalogServer({ + server: "demo", + catalog, + headers: { "X-Other": "b" }, + protocolEra: "modern", + secretStore: store, + }); + expect(readCatalog().demo).toEqual({ + type: "stdio", + command: "node", + args: ["server.js"], + env: { API_KEY: "" }, + headers: { "X-Other": "b" }, + protocolEra: "modern", + }); + expect(await store.get("demo", "env:API_KEY")).toBe("sekrit"); + }); + + it("renames, carrying the secrets to the new name", async () => { + const result = await editCatalogServer({ + server: "demo", + rename: "demo2", + catalog, + secretStore: store, + }); + expect(result).toEqual({ + ok: true, + action: "updated", + server: "demo2", + previousName: "demo", + catalog, + }); + expect(Object.keys(readCatalog())).toEqual(["demo2"]); + expect(await store.get("demo2", "env:API_KEY")).toBe("sekrit"); + expect(await store.get("demo", "env:API_KEY")).toBeNull(); + }); + + it("rejects a rename onto an existing name", async () => { + await addCatalogServer({ + server: "other", + catalog, + target: ["node"], + secretStore: store, + }); + await expect( + editCatalogServer({ + server: "demo", + rename: "other", + catalog, + secretStore: store, + }), + ).rejects.toThrow("Server 'other' already exists"); + }); + + it("requires something to change", async () => { + await expect( + editCatalogServer({ + server: "demo", + rename: "demo", + catalog, + secretStore: store, + }), + ).rejects.toThrow(/needs something to change/); + }); + + it("requires --server", async () => { + await expect( + editCatalogServer({ catalog, cwd: "/x", secretStore: store }), + ).rejects.toThrow("servers/edit requires --server <name>."); + }); + + it("rejects --transport without a new target", async () => { + await expect( + editCatalogServer({ + server: "demo", + catalog, + transport: "sse", + cwd: "/x", + secretStore: store, + }), + ).rejects.toThrow(/--transport on servers\/edit/); + }); + + it("reports an unknown name without touching a missing catalog", async () => { + const missing = path.join(dir, "absent.json"); + await expect( + editCatalogServer({ + server: "demo", + catalog: missing, + cwd: "/x", + secretStore: store, + }), + ).rejects.toThrow(/Server 'demo' not found/); + expect(fs.existsSync(missing)).toBe(false); + }); + + it("reports a name the route layer drops on read", async () => { + fs.writeFileSync( + catalog, + '{"mcpServers":{"__proto__":{"type":"stdio","command":"node"}}}', + ); + await expect( + editCatalogServer({ + server: "__proto__", + catalog, + cwd: "/x", + secretStore: store, + }), + ).rejects.toThrow(/Server '__proto__' not found/); + }); + + it("rejects -e / --cwd on a non-stdio entry", async () => { + await addCatalogServer({ + server: "web", + catalog, + serverUrl: "https://example.com/mcp", + secretStore: store, + }); + await expect( + editCatalogServer({ + server: "web", + catalog, + cwd: "/x", + secretStore: store, + }), + ).rejects.toThrow(/apply to stdio servers; 'web' is streamable-http/); + }); + + it("rejects a blank --rename", async () => { + await expect( + editCatalogServer({ + server: "demo", + rename: " ", + catalog, + cwd: "/x", + secretStore: store, + }), + ).rejects.toThrow("--rename requires a non-empty name."); + }); + + it("rejects -e / --cwd alongside a new URL target", async () => { + await expect( + editCatalogServer({ + server: "demo", + catalog, + serverUrl: "https://example.com/mcp", + cwd: "/x", + secretStore: store, + }), + ).rejects.toThrow( + "--cwd apply to stdio servers; this target is streamable-http.", + ); + }); + + it("refuses a rename on a session store", async () => { + await expect( + editCatalogServer({ + server: "demo", + rename: "demo2", + catalog, + secretStore: new SessionSecretStore(), + }), + ).rejects.toThrow(/--rename needs a durable secret store/); + expect(Object.keys(readCatalog())).toEqual(["demo"]); + }); + + it("allows a non-rename edit on a session store", async () => { + await editCatalogServer({ + server: "demo", + catalog, + cwd: "/x", + secretStore: new SessionSecretStore(), + }); + expect(readCatalog().demo).toMatchObject({ cwd: "/x" }); + }); + + it("refuses new -e values on a session store", async () => { + await expect( + editCatalogServer({ + server: "demo", + catalog, + env: { NEW: "v" }, + secretStore: new SessionSecretStore(), + }), + ).rejects.toThrow(/in-memory only/); + }); +}); + +describe("removeCatalogServer", () => { + it("removes the entry and its secrets", async () => { + await addCatalogServer({ + server: "demo", + catalog, + target: ["node"], + env: { API_KEY: "sekrit" }, + secretStore: store, + }); + const result = await removeCatalogServer({ + server: "demo", + catalog, + secretStore: store, + }); + expect(result).toEqual({ + ok: true, + action: "removed", + server: "demo", + catalog, + }); + expect(readCatalog()).toEqual({}); + expect(await store.get("demo", "env:API_KEY")).toBeNull(); + }); + + it("reports an unknown name", async () => { + writeCatalog({}); + await expect( + removeCatalogServer({ server: "nope", catalog, secretStore: store }), + ).rejects.toThrow(/Server 'nope' not found/); + }); + + it("reports a malformed (non-object) entry as not found", async () => { + writeCatalog({ demo: null, other: "x" }); + await expect( + removeCatalogServer({ server: "demo", catalog, secretStore: store }), + ).rejects.toThrow(/Server 'demo' not found/); + }); + + it("treats a file without an mcpServers map as empty", async () => { + fs.writeFileSync(catalog, "{}"); + await expect( + removeCatalogServer({ server: "nope", catalog, secretStore: store }), + ).rejects.toThrow(/not found/); + fs.writeFileSync(catalog, '{"mcpServers":null}'); + await expect( + removeCatalogServer({ server: "nope", catalog, secretStore: store }), + ).rejects.toThrow(/not found/); + }); + + it("rejects every flag it would otherwise ignore", async () => { + await expect( + removeCatalogServer({ + server: "demo", + catalog, + target: ["node"], + transport: "stdio", + env: { A: "1" }, + cwd: "/x", + headers: { A: "1" }, + protocolEra: "modern", + rename: "x", + secretStore: store, + }), + ).rejects.toThrow( + "servers/remove takes only --server <name>; remove a command/URL, --transport, -e, --cwd, --header, --protocol-era, --rename.", + ); + }); + + it("requires --server", async () => { + await expect( + removeCatalogServer({ catalog, secretStore: store }), + ).rejects.toThrow("servers/remove requires --server <name>."); + }); +}); + +describe("runCatalogWrite", () => { + it("dispatches each method", async () => { + await runCatalogWrite("servers/add", { + server: "demo", + catalog, + target: ["node"], + secretStore: store, + }); + await runCatalogWrite("servers/edit", { + server: "demo", + catalog, + cwd: "/x", + secretStore: store, + }); + expect(readCatalog().demo).toMatchObject({ cwd: "/x" }); + await runCatalogWrite("servers/remove", { + server: "demo", + catalog, + secretStore: store, + }); + expect(readCatalog()).toEqual({}); + }); +}); + +describe("--method servers/add|edit|remove", () => { + // A file store under the temp dir keeps the in-process CLI off the host + // keychain while still being durable, which `-e` requires. + const env = (): Record<string, string> => ({ + MCP_INSPECTOR_SECRET_STORE: "file", + MCP_INSPECTOR_SECRET_FILE: path.join(dir, "secrets.json"), + }); + + it("adds, edits, and removes through the CLI", async () => { + const added = await runCli( + [ + "node", + "server.js", + "--catalog", + catalog, + "--method", + "servers/add", + "--server", + "demo", + "-e", + "API_KEY=sekrit", + "--format", + "json", + ], + { env: env() }, + ); + expectCliSuccess(added); + expect(JSON.parse(added.stdout)).toEqual({ + result: { ok: true, action: "added", server: "demo", catalog }, + }); + + const edited = await runCli( + [ + "--catalog", + catalog, + "--method", + "servers/edit", + "--server", + "demo", + "--rename", + "demo2", + "--header", + "X-Team: a", + ], + { env: env() }, + ); + expectCliSuccess(edited); + expect(JSON.parse(edited.stdout)).toMatchObject({ + action: "updated", + server: "demo2", + previousName: "demo", + }); + expect(readCatalog()).toEqual({ + demo2: { + type: "stdio", + command: "node", + args: ["server.js"], + env: { API_KEY: "" }, + headers: { "X-Team": "a" }, + }, + }); + + const removed = await runCli( + ["--catalog", catalog, "--method", "servers/remove", "--server", "demo2"], + { env: env() }, + ); + expectCliSuccess(removed); + expect(readCatalog()).toEqual({}); + }); + + it("honours MCP_CATALOG_PATH even alongside a target", async () => { + const result = await runCli( + [ + "--method", + "servers/add", + "--server", + "web", + "--server-url", + "https://example.com/mcp", + ], + { env: { ...env(), MCP_CATALOG_PATH: catalog } }, + ); + expectCliSuccess(result); + expect(readCatalog().web).toEqual({ + type: "streamable-http", + url: "https://example.com/mcp", + }); + }); + + it("refuses --config", async () => { + const result = await runCli( + ["node", "--config", catalog, "--method", "servers/add", "--server", "x"], + { env: env() }, + ); + expectCliFailure(result); + expect(result.stderr).toMatch(/read-only/); + }); + + it("rejects --rename outside servers/edit", async () => { + const result = await runCli([ + "--catalog", + catalog, + "--method", + "servers/list", + "--rename", + "x", + ]); + expectCliFailure(result); + expect(result.stderr).toMatch(/--rename is only valid/); + }); + + it("rejects --relogin and --advertise-apps with a write method", async () => { + const relogin = await runCli([ + "--catalog", + catalog, + "--method", + "servers/remove", + "--server", + "x", + "--relogin", + ]); + expectCliFailure(relogin); + expect(relogin.stderr).toMatch(/--relogin cannot be combined/); + const advertise = await runCli([ + "--catalog", + catalog, + "--method", + "servers/remove", + "--server", + "x", + "--advertise-apps", + ]); + expectCliFailure(advertise); + expect(advertise.stderr).toMatch(/--advertise-apps requires/); + }); + + it("lists the write methods among the supported ones", async () => { + const result = await runCli(["--catalog", catalog, "--method", "nope"]); + expectCliFailure(result); + expect(result.stderr).toMatch( + /servers\/add, servers\/edit, servers\/remove/, + ); + }); +}); diff --git a/clients/cli/__tests__/skills-verify-cli.test.ts b/clients/cli/__tests__/skills-verify-cli.test.ts index 4179c04da6..704154294e 100644 --- a/clients/cli/__tests__/skills-verify-cli.test.ts +++ b/clients/cli/__tests__/skills-verify-cli.test.ts @@ -1,7 +1,17 @@ import { describe, it, expect } from "vitest"; +import { + createTestServerHttp, + createTestServerInfo, +} from "@modelcontextprotocol/inspector-test-server"; +import type { SkillVerifyReport } from "@inspector/core/mcp/skillsVerification.js"; import { runCli } from "../src/cli.js"; import { consumeMethodOutcome } from "../src/handlers/consume-outcome.js"; -import { EXIT_CODES } from "../src/error-handler.js"; +import { EXIT_CODES } from "@inspector/core/cli/error-handler.js"; +import { + runCli as runCliCaptured, + type CliResult, +} from "./helpers/cli-runner.js"; +import { createTestConfig, deleteConfigFile } from "./helpers/fixtures.js"; /** * `--verify`'s argument validation and its NDJSON consumption path (#2248). @@ -140,6 +150,33 @@ describe("consumeMethodOutcome NDJSON summary and exit code (#2248)", () => { expect(streams.stderr).toBe("all good\n"); }); + it("drops the summary under --quiet but keeps the exit code and its message (#2435)", async () => { + const streams = captureStreams(); + let thrown: unknown; + try { + await consumeMethodOutcome( + { + kind: "ndjson", + lines: [{ ok: false }], + summary: "one failed", + exitCode: EXIT_CODES.SKILL_NONCONFORMANT, + }, + { quiet: true }, + ); + } catch (err) { + thrown = err; + } finally { + streams.restore(); + } + expect(streams.stdout.trim()).toBe('{"ok":false}'); + expect(streams.stderr).toBe(""); + // The error envelope still carries the verdict, so nothing is lost. + expect(thrown).toMatchObject({ + exitCode: EXIT_CODES.SKILL_NONCONFORMANT, + message: "one failed", + }); + }); + it("throws the exit code AFTER writing the report", async () => { // The report is the output a CI job reads; failing before writing it would // give the reader an exit code and nothing to act on. @@ -231,3 +268,99 @@ describe("consumeMethodOutcome NDJSON summary and exit code (#2248)", () => { expect(streams.stdout.trim()).toBe('{"a":1}'); }); }); + +/** + * The catalog-budget escape hatch round-trips (#2428). + * + * A skill past the run's catalog budget is reported `incomplete` with a message + * naming the command that verifies it on its own. That text is the only route a + * user has to a verdict for the skipped skill, so it is run back through the + * real argument parser against a real server rather than trusted as prose. The + * first version named `--method skills/get --uri` alone, which fetches the + * skill and checks nothing — a command that parsed, succeeded, and gave no + * verdict. + */ +describe("the catalog-budget escape hatch (#2428)", () => { + /** + * The backticked command in a report's `incomplete` message, as argv, with + * its `<uri>` placeholder filled from the report's own `uri`. Substituted as + * one argv element — the way a user's quoting would — because the message + * deliberately never splices the server-controlled URI into the command. + */ + function suggestedArgs(report: SkillVerifyReport): string[] { + const command = /`([^`]+)`/.exec(report.incomplete ?? "")?.[1]; + if (!command) throw new Error(`no command in: ${report.incomplete}`); + const args = command.split(/\s+/); + if (!args.includes("<uri>")) + throw new Error(`no <uri> placeholder in: ${command}`); + return args.map((arg) => (arg === "<uri>" ? report.uri : arg)); + } + + /** + * The NDJSON report lines of a `--verify` run. Anything else — a usage + * error, or a fetched skill printed as one JSON document because the + * command lacked `--verify` — fails here naming what the run printed. + */ + function reportsOf(result: CliResult): SkillVerifyReport[] { + try { + return result.stdout + .trim() + .split("\n") + .map((line) => JSON.parse(line) as SkillVerifyReport); + } catch { + throw new Error( + `not a --verify report (exit ${result.exitCode}):\n${result.output}`, + ); + } + } + + it("verifies a skill skipped for budget with the command the message names", async () => { + const server = createTestServerHttp({ + serverInfo: createTestServerInfo("skills-budget", "1.0.0"), + skills: true, + }); + let catalogPath: string | undefined; + try { + await server.start(); + // A budget of one skill, so every skill after the first is skipped — + // and the SAME budget applies to the follow-up run, which is what proves + // the command works for a skill this configuration skipped. + catalogPath = createTestConfig({ + mcpServers: { + skills: { + type: "streamable-http", + url: server.url, + skillCatalogMaxSkills: 1, + }, + }, + }); + const target = ["--catalog", catalogPath, "--server", "skills", "--cli"]; + + const listed = await runCliCaptured([ + ...target, + "--method", + "skills/list", + "--verify", + ]); + const skipped = reportsOf(listed).find((report) => + report.incomplete?.includes("catalog budget"), + ); + if (!skipped) throw new Error(`no skipped skill in: ${listed.stdout}`); + expect(skipped.files).toHaveLength(0); + + const suggested = suggestedArgs(skipped); + const got = await runCliCaptured([...target, ...suggested]); + + // A verdict for exactly the skipped skill, from files actually read — + // not the skill echoed back, and not the budget message again. + const reports = reportsOf(got); + expect(reports).toHaveLength(1); + expect(reports[0].uri).toBe(skipped.uri); + expect(reports[0].incomplete ?? "").not.toMatch(/catalog budget/); + expect(reports[0].files.length).toBeGreaterThan(0); + } finally { + await server.stop(); + if (catalogPath) deleteConfigFile(catalogPath); + } + }); +}); diff --git a/clients/cli/__tests__/stored-auth.test.ts b/clients/cli/__tests__/stored-auth.test.ts index 0965b093ef..1755cf5b70 100644 --- a/clients/cli/__tests__/stored-auth.test.ts +++ b/clients/cli/__tests__/stored-auth.test.ts @@ -18,6 +18,7 @@ import { deepLinkTransport, refreshStoredAuthToken, waitForStoredToken, + resolveStoredCredentials, type StoredServers, } from "../src/cli.js"; import { SecretFileLockHeldError } from "@inspector/core/auth/node/secret-store.js"; @@ -82,7 +83,9 @@ describe("refreshStoredAuthToken", () => { // Persisted writes split tokens into the process-wide (in-memory, per // vitest.config.ts) secret store, and joined reads prefer the store over // file plaintext — so purge the entry between tests or one test's rotated - // tokens would leak into the next test's fixture. + // tokens would leak into the next test's fixture. Post-#2549 each fixture + // file mints its own secrets namespace, so namespaced entries can't collide + // across tests; this purge covers the legacy (un-namespaced) id. afterEach(async () => { await defaultSecretStore().deleteAllForServer(oauthSecretServerId(SERVER)); }); @@ -1149,3 +1152,367 @@ describe("refreshStoredAuthToken discovery compatibility (#2172)", () => { } }); }); + +/** + * #2517: since the store was re-keyed per authorization server (#1625), an + * acquired token lives at `servers[url].byIssuer[activeIssuer].tokens`, not at + * the server level. Every stored-auth lookup must read that shape — the suites + * above use the legacy top-level shape, which still has to keep working. + */ +describe("issuer-keyed stored auth (#2517)", () => { + const ISSUER = "https://as.example"; + const OTHER = "https://old-as.example"; + + /** A server entry whose credentials live under `issuer`'s active slot. */ + const issuerKeyed = (issuer: string, tokens: Record<string, string>) => ({ + activeIssuer: issuer, + byIssuer: { + [issuer]: { + tokens: { token_type: "Bearer", ...tokens }, + clientInformation: { client_id: "cid", client_secret: "sec" }, + }, + }, + }); + + describe("resolveStoredCredentials", () => { + it("answers with the active issuer's slot, tokens and client together", () => { + expect( + resolveStoredCredentials({ + activeIssuer: ISSUER, + byIssuer: { + [OTHER]: { + tokens: { access_token: "other", token_type: "Bearer" }, + clientInformation: { client_id: "other-client" }, + }, + [ISSUER]: { + tokens: { access_token: "active", token_type: "Bearer" }, + clientInformation: { client_id: "active-client" }, + }, + }, + tokens: { access_token: "legacy", token_type: "Bearer" }, + clientInformation: { client_id: "legacy-client" }, + }), + ).toEqual({ + issuer: ISSUER, + tokens: { access_token: "active", token_type: "Bearer" }, + clientInformation: { client_id: "active-client" }, + }); + }); + + it("never borrows the legacy client for an issuer slot's tokens", () => { + expect( + resolveStoredCredentials({ + activeIssuer: ISSUER, + byIssuer: { + [ISSUER]: { + tokens: { access_token: "active", token_type: "Bearer" }, + }, + }, + clientInformation: { client_id: "legacy-client" }, + }).clientInformation, + ).toBeUndefined(); + }); + + it("prefers a preregistered (static) client over the slot's", () => { + expect( + resolveStoredCredentials({ + activeIssuer: ISSUER, + byIssuer: { + [ISSUER]: { + tokens: { access_token: "active", token_type: "Bearer" }, + clientInformation: { client_id: "dcr-client" }, + }, + }, + preregisteredClientInformation: { client_id: "static-client" }, + }).clientInformation, + ).toEqual({ client_id: "static-client" }); + }); + + it("falls back to the legacy fields when the active slot holds no tokens", () => { + expect( + resolveStoredCredentials({ + activeIssuer: ISSUER, + byIssuer: { [ISSUER]: { clientInformation: { client_id: "c" } } }, + tokens: { access_token: "legacy", token_type: "Bearer" }, + clientInformation: { client_id: "legacy-client" }, + }), + ).toEqual({ + tokens: { access_token: "legacy", token_type: "Bearer" }, + clientInformation: { client_id: "legacy-client" }, + }); + }); + + it("never picks an arbitrary issuer when none is active", () => { + expect( + resolveStoredCredentials({ + byIssuer: { + [ISSUER]: { + tokens: { access_token: "orphan", token_type: "Bearer" }, + }, + }, + }), + ).toEqual({ tokens: undefined, clientInformation: undefined }); + }); + + it("does not resolve a __proto__ active issuer to the prototype", () => { + expect( + resolveStoredCredentials({ activeIssuer: "__proto__", byIssuer: {} }) + .tokens, + ).toBeUndefined(); + }); + }); + + describe("refreshStoredAuthToken", () => { + const SERVER = "https://issuer-keyed.example/mcp"; + + afterEach(async () => { + await defaultSecretStore().deleteAllForServer( + oauthSecretServerId(SERVER), + ); + }); + + it("refreshes with the active issuer's credentials and persists the rotation under it", async () => { + const path = writeOAuthFixture({ + [SERVER]: { + activeIssuer: ISSUER, + byIssuer: { + [ISSUER]: { + tokens: { refresh_token: "active-refresh", token_type: "Bearer" }, + clientInformation: { client_id: "active-client" }, + }, + [OTHER]: { + tokens: { refresh_token: "other-refresh", token_type: "Bearer" }, + clientInformation: { client_id: "other-client" }, + }, + }, + serverMetadata: { + issuer: ISSUER, + token_endpoint: `${ISSUER}/token`, + }, + }, + }); + try { + const refresh = vi.fn().mockResolvedValue({ + access_token: "rotated-access", + token_type: "Bearer", + refresh_token: "rotated-refresh", + }); + const token = await refreshStoredAuthToken(SERVER, path, { + refresh, + discover: vi.fn(), + }); + expect(token).toBe("rotated-access"); + const [, opts] = refresh.mock.calls[0]!; + expect(opts.refreshToken).toBe("active-refresh"); + expect(opts.clientInformation).toEqual({ client_id: "active-client" }); + + const persisted = (await readOAuthStore(path))?.servers[SERVER]; + expect(persisted?.activeIssuer).toBe(ISSUER); + expect(persisted?.byIssuer?.[ISSUER]?.tokens?.refresh_token).toBe( + "rotated-refresh", + ); + expect(persisted?.byIssuer?.[ISSUER]?.clientInformation).toEqual({ + client_id: "active-client", + }); + // The other issuer's slot is untouched, and nothing is written to the + // legacy top-level fallback. + expect(persisted?.byIssuer?.[OTHER]?.tokens?.refresh_token).toBe( + "other-refresh", + ); + expect(persisted?.tokens).toBeUndefined(); + } finally { + rmSync(path, { force: true }); + } + }); + + it("reports no_client_information when the active slot has a refresh token but no client", async () => { + const path = writeOAuthFixture({ + [SERVER]: { + activeIssuer: ISSUER, + byIssuer: { + [ISSUER]: { + tokens: { refresh_token: "active-refresh", token_type: "Bearer" }, + }, + }, + clientInformation: { client_id: "legacy-client" }, + }, + }); + try { + await expect( + refreshStoredAuthToken(SERVER, path, { refresh: vi.fn() }), + ).rejects.toMatchObject({ + exitCode: 3, + envelope: { code: "no_client_information" }, + }); + } finally { + rmSync(path, { force: true }); + } + }); + }); + + describe("CLI flags", () => { + let server: ReturnType<typeof createTestServerHttp>; + let serverUrl: string; + + beforeAll(async () => { + server = createTestServerHttp({ + serverInfo: createTestServerInfo(), + tools: [createEchoTool()], + }); + await server.start(); + serverUrl = server.url; + }); + + afterAll(async () => { + await server.stop(); + }); + + afterEach(async () => { + await defaultSecretStore().deleteAllForServer( + oauthSecretServerId(serverUrl), + ); + }); + + it("--list-stored-auth lists a server whose token is under its active issuer", async () => { + const fixture = writeOAuthFixture({ + "https://keyed.example/mcp": issuerKeyed(ISSUER, { + access_token: "t1", + }), + "https://no-active.example/mcp": { + byIssuer: { + [ISSUER]: { tokens: { access_token: "t2", token_type: "Bearer" } }, + }, + }, + }); + try { + const result = await runCli(["--list-stored-auth"], { + env: { MCP_INSPECTOR_OAUTH_STATE_PATH: fixture }, + }); + expectCliSuccess(result); + const out = JSON.parse(result.stdout) as { storedServerUrls: string[] }; + expect(out.storedServerUrls).toEqual(["https://keyed.example/mcp"]); + } finally { + rmSync(fixture, { force: true }); + } + }); + + it("--use-stored-auth injects the active issuer's access token", async () => { + const fixture = writeOAuthFixture({ + [serverUrl]: issuerKeyed(ISSUER, { access_token: "issuer-access" }), + }); + try { + const result = await runCli( + [ + "--transport", + "http", + "--server-url", + serverUrl, + "--use-stored-auth", + "--method", + "tools/list", + ], + { env: { MCP_INSPECTOR_OAUTH_STATE_PATH: fixture } }, + ); + expectCliSuccess(result); + const last = server.getRecordedRequests().at(-1)!; + expect(last.headers?.authorization).toBe("Bearer issuer-access"); + } finally { + rmSync(fixture, { force: true }); + } + }); + + it("--use-stored-auth refreshes the active issuer's token and persists it under that issuer", async () => { + const tokenServer: Server = createServer((_req, res) => { + res.writeHead(200, { "content-type": "application/json" }); + res.end( + JSON.stringify({ + access_token: "issuer-refreshed", + token_type: "Bearer", + refresh_token: "issuer-rotated", + }), + ); + }); + await new Promise<void>((resolve) => tokenServer.listen(0, resolve)); + const addr = tokenServer.address(); + const tokenBase = + typeof addr === "object" && addr ? `http://127.0.0.1:${addr.port}` : ""; + const fixture = writeOAuthFixture({ + [serverUrl]: { + ...issuerKeyed(tokenBase, { refresh_token: "issuer-refresh" }), + serverMetadata: { + issuer: tokenBase, + token_endpoint: `${tokenBase}/token`, + response_types_supported: ["code"], + grant_types_supported: ["authorization_code", "refresh_token"], + token_endpoint_auth_methods_supported: ["client_secret_post"], + }, + }, + }); + try { + const result = await runCli( + [ + "--transport", + "http", + "--server-url", + serverUrl, + "--use-stored-auth", + "--method", + "tools/list", + ], + { env: { MCP_INSPECTOR_OAUTH_STATE_PATH: fixture } }, + ); + expectCliSuccess(result); + const last = server.getRecordedRequests().at(-1)!; + expect(last.headers?.authorization).toBe("Bearer issuer-refreshed"); + const persisted = (await readOAuthStore(fixture))?.servers[serverUrl]; + expect(persisted?.byIssuer?.[tokenBase]?.tokens?.refresh_token).toBe( + "issuer-rotated", + ); + } finally { + rmSync(fixture, { force: true }); + await new Promise<void>((resolve) => + tokenServer.close(() => resolve()), + ); + } + }); + + it("--wait-for-auth returns as soon as an issuer-keyed token lands", async () => { + const dir = mkdtempSync(join(tmpdir(), "inspector-cli-wait-issuer-")); + const file = join(dir, "oauth.json"); + setTimeout(() => { + writeFileSync( + file, + JSON.stringify({ + servers: { + [normalizeServerUrl(serverUrl)]: issuerKeyed(ISSUER, { + access_token: "waited-issuer-tok", + }), + }, + idpSessions: {}, + }), + "utf8", + ); + }, 200); + try { + const result = await runCli( + [ + "--transport", + "http", + "--server-url", + serverUrl, + "--wait-for-auth", + "5", + "--method", + "tools/list", + ], + { env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file } }, + ); + expectCliSuccess(result); + const last = server.getRecordedRequests().at(-1)!; + expect(last.headers?.authorization).toBe("Bearer waited-issuer-tok"); + } finally { + rmSync(dir, { recursive: true, force: true }); + } + }); + }); +}); diff --git a/clients/cli/__tests__/style.test.ts b/clients/cli/__tests__/style.test.ts index f9b8cfeffd..236875a8dd 100644 --- a/clients/cli/__tests__/style.test.ts +++ b/clients/cli/__tests__/style.test.ts @@ -4,7 +4,7 @@ import { resolveAnsiEnabled, styleFromOpts, PLAIN, -} from "../src/style.js"; +} from "@inspector/core/cli/style.js"; describe("resolveAnsiEnabled", () => { it("is off for --plain, json, NO_COLOR, and non-TTY", () => { diff --git a/clients/cli/src/cli.ts b/clients/cli/src/cli.ts index 293df651f1..34653b7f1e 100644 --- a/clients/cli/src/cli.ts +++ b/clients/cli/src/cli.ts @@ -1,6 +1,12 @@ import { Command } from "commander"; +import { + CATALOG_METHODS, + emitCompletionIfRequested, + isCatalogMethod, + registerCompletionOption, +} from "./completion.js"; type McpResponse = Record<string, unknown>; -import { awaitableLog } from "./utils/awaitable-log.js"; +import { awaitableLog } from "@inspector/core/cli/utils/awaitable-log.js"; import type { InspectorServerSettings, MCPServerConfig, @@ -11,9 +17,21 @@ import { eraToVersionNegotiation } from "@inspector/core/mcp/types.js"; import { DEFAULT_CONNECT_TIMEOUT_MS, withConnectTimeout, -} from "./handlers/connect-timeout.js"; -import { listServerEntries, showServerEntry } from "./handlers/servers-list.js"; -import { writeFormattedResult } from "./handlers/format-output.js"; +} from "@inspector/core/cli/handlers/connect-timeout.js"; +import { + listServerEntries, + showServerEntry, +} from "@inspector/core/cli/handlers/servers-list.js"; +import { + isCatalogWriteMethod, + runCatalogWrite, +} from "./handlers/servers-write.js"; +import { writeFormattedResult } from "@inspector/core/cli/handlers/format-output.js"; +import { + parseOutputFileFormat, + validateOutputOptions, + type OutputFileFormat, +} from "@inspector/core/cli/handlers/output-file.js"; import { clearStoredAuthForRelogin } from "./clear-stored-auth-for-relogin.js"; import { InspectorClient } from "@inspector/core/mcp/index.js"; import { cleanRoots } from "@inspector/core/mcp/serverList.js"; @@ -26,6 +44,7 @@ import { parseKeyValuePair as parseEnvPair, parseHeaderPair, parseProtocolEra, + skillCatalogLimitParser, } from "@inspector/core/mcp/node/index.js"; import type { JsonValue } from "@inspector/core/mcp/index.js"; import type { StrictJsonValue } from "@inspector/core/json/jsonUtils.js"; @@ -37,16 +56,18 @@ import { import { getStateFilePath } from "@inspector/core/auth/node/storage-node.js"; import { SecretFileLockHeldError } from "@inspector/core/auth/node/secret-store.js"; import { consumeMethodOutcome } from "./handlers/consume-outcome.js"; -import { runMethod } from "./handlers/run-method.js"; +import { runMethod } from "@inspector/core/cli/handlers/run-method.js"; import { isOneShotMethod, ONE_SHOT_METHODS, type MethodArgs, -} from "./handlers/method-types.js"; -export type { CliAppInfo } from "./handlers/method-types.js"; +} from "@inspector/core/cli/handlers/method-types.js"; +export type { CliAppInfo } from "@inspector/core/cli/handlers/method-types.js"; export { emitResult } from "./handlers/emit-result.js"; -export { collectAppInfo } from "./handlers/collect-app-info.js"; +export { collectAppInfo } from "@inspector/core/cli/handlers/collect-app-info.js"; import { type OAuthPersistSnapshot } from "@inspector/core/auth/oauth-persist.js"; +import type { ServerOAuthState } from "@inspector/core/auth/store.js"; +import { getOwnEntry } from "@inspector/core/storage/own-entry.js"; import { readOAuthStore, writeOAuthSections, @@ -64,17 +85,19 @@ import { } from "@modelcontextprotocol/client"; import type { OAuthClientInformation, - OAuthMetadata, OAuthTokens, } from "@modelcontextprotocol/client"; -import { CliExitCodeError, EXIT_CODES } from "./error-handler.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; import { MutableRedirectUrlProvider } from "@inspector/core/auth/index.js"; import { NodeOAuthStorage } from "@inspector/core/auth/node/index.js"; -import { createCliOAuthNavigation } from "./cli-oauth-navigation.js"; +import { createCliOAuthNavigation } from "@inspector/core/cli/cli-oauth-navigation.js"; import { connectInspectorWithOAuth, withCliAuthRecoveryRetry, -} from "./cliOAuth.js"; +} from "@inspector/core/cli/cliOAuth.js"; import { DEFAULT_RUNNER_OAUTH_CALLBACK_URL, formatRunnerOAuthRedirectUrl, @@ -129,7 +152,9 @@ async function callMethod( const revocation = await clearStoredAuthForRelogin(serverConfig.url, { revoke: revoke && serverSettings?.oauthRevokeOnClear !== false, }); - if (revocation?.status === "failed") { + // A warning, not the result — `--quiet` drops it (#2435). The relogin + // itself still happened; only the advisory line is suppressed. + if (revocation?.status === "failed" && !args.quiet) { process.stderr.write( `Warning: could not revoke the OAuth grant at the authorization server (${revocation.detail}); it may still be valid there.\n`, ); @@ -178,6 +203,12 @@ async function callMethod( environment, clientIdentity, initialLoggingLevel: "debug", + // A stdio server's stderr is inherited by default, so its startup banners + // and logs interleave with the CLI's own output. `--quiet` pipes it + // instead; InspectorClient drains the pipe into its `stderrLog` event, + // which nothing here subscribes to, so the lines are discarded without the + // child ever blocking on a full pipe (#2435). + pipeStderr: args.quiet === true, progress: false, sample: false, elicit: false, @@ -224,7 +255,7 @@ async function callMethod( redirectUrlProvider, callbackUrlConfig, serverSettings, - { storedAuthOnly, autoOpenControl }, + { storedAuthOnly, autoOpenControl, quiet: args.quiet }, ); const outcome = await withCliAuthRecoveryRetry( @@ -234,7 +265,7 @@ async function callMethod( callbackUrlConfig, serverSettings, () => runMethod(inspectorClient, args), - { storedAuthOnly, autoOpenControl }, + { storedAuthOnly, autoOpenControl, quiet: args.quiet }, ); await consumeMethodOutcome(outcome, args); @@ -259,15 +290,61 @@ export function normalizeServerUrl(serverUrl: string): string { } } -/** The subset of a stored server's OAuth state the CLI reads/refreshes. */ -type StoredServerState = { - tokens?: OAuthTokens; - clientInformation?: OAuthClientInformation; - serverMetadata?: OAuthMetadata; -}; +/** + * A stored server's OAuth state, in the shared store's own shape: credentials + * live under `byIssuer[issuer]` (SEP-2352, #1625), with the bare top-level + * `tokens` / `clientInformation` kept only as the legacy pre-issuer fallback. + */ +type StoredServerState = ServerOAuthState; /** The stored-server map shape the CLI reads out of the OAuth state file. */ export type StoredServers = Record<string, StoredServerState>; +/** + * The credentials that answer a stored server's ctx-less read, plus the issuer + * slot they came from (`undefined` for the legacy unkeyed fallback). + */ +export interface StoredCredentials { + issuer?: string; + tokens?: OAuthTokens; + clientInformation?: OAuthClientInformation; +} + +/** + * Resolve the credentials a stored server's state answers with, the same way + * the shared store does for a read with no `issuer` (`OAuthStorageBase`): the + * `activeIssuer` slot of `byIssuer` first, then the legacy top-level fields. It + * never picks an arbitrary issuer — with no `activeIssuer` only the legacy + * fields can answer (#2517). + * + * Tokens and client information are taken from the **same** source, so a + * refresh never pairs one AS's refresh token with another AS's client: an + * issuer slot's tokens never borrow the legacy unkeyed client, which may have + * been registered with a different AS. A preregistered (static) client is + * issuer-independent and wins over either, as it does in the auth provider. + */ +export function resolveStoredCredentials( + state: StoredServerState, +): StoredCredentials { + const issuer = state.activeIssuer; + // Own-property read: a persisted `__proto__` issuer must not resolve to + // the inherited `Object.prototype`. + const slot = + issuer !== undefined ? getOwnEntry(state.byIssuer, issuer) : undefined; + if (slot?.tokens) { + return { + issuer, + tokens: slot.tokens, + clientInformation: + state.preregisteredClientInformation ?? slot.clientInformation, + }; + } + return { + tokens: state.tokens, + clientInformation: + state.preregisteredClientInformation ?? state.clientInformation, + }; +} + /** * Read the shared OAuth state ({@link OAuthPersistSnapshot}) fresh on every * call — required for `--wait-for-auth` polling. Returns the full snapshot, @@ -320,7 +397,10 @@ function findStoredToken( servers: StoredServers, serverUrl: string, ): string | undefined { - return findStoredServerState(servers, serverUrl)?.state.tokens?.access_token; + const found = findStoredServerState(servers, serverUrl); + return found + ? resolveStoredCredentials(found.state).tokens?.access_token + : undefined; } /** @@ -383,8 +463,11 @@ export async function refreshStoredAuthToken( const snapshot = await readOAuthSnapshot(statePath); const servers = snapshot.servers as StoredServers; const found = findStoredServerState(servers, serverUrl); - const refreshToken = found?.state.tokens?.refresh_token; - const clientInformation = found?.state.clientInformation; + const credentials: StoredCredentials = found + ? resolveStoredCredentials(found.state) + : {}; + const refreshToken = credentials.tokens?.refresh_token; + const clientInformation = credentials.clientInformation; if (!found || !refreshToken) { throw new CliExitCodeError( EXIT_CODES.AUTH_REQUIRED, @@ -447,7 +530,22 @@ export async function refreshStoredAuthToken( // time (not the snapshot read before the network round-trip), under the // same cross-process lock every other writer uses, and keeps the file's // owner-only `0o600` mode + `mkdir -p` via the shared store IO. - servers[found.key] = { ...found.state, tokens }; + // + // The rotated tokens go back to the slot they were read from: the active + // issuer's `byIssuer` entry (#2517), or the legacy top-level fields for a + // pre-issuer entry — never a slot the old tokens did not come from. + const { issuer } = credentials; + servers[found.key] = + issuer !== undefined + ? { + ...found.state, + byIssuer: { + ...found.state.byIssuer, + // A computed key defines an own property, so `__proto__` is safe. + [issuer]: { ...getOwnEntry(found.state.byIssuer, issuer), tokens }, + }, + } + : { ...found.state, tokens }; await writeOAuthSections(statePath, snapshot, { servers: [found.key] }); return tokens.access_token; @@ -697,12 +795,32 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { // commander tear down the whole test worker. For --help / --version // (exitCode 0) we return without throwing, so commander falls through to its // normal clean process.exit(0) after printing — preserving that UX. See #1484. + // Commander's own `error: …` line for a usage error is held back here and + // written from `exitOverride` below — unless `--quiet` was already parsed, + // in which case it is dropped: the error is still thrown and reaches the + // envelope with the same message, so stderr stays the one envelope line + // (#2435). Deciding on Commander's parsed state rather than scanning argv + // means a `-q` that is another option's value (`--client-secret -q`) is not + // mistaken for the flag, and `-qe KEY=V` is. Commander parses left to right, + // so a `-q` placed *after* the bad option is not seen yet and the line + // prints. `--help` / `--version` write through `writeOut`, untouched. + let commanderError = ""; + program.configureOutput({ + writeErr: (text) => { + commanderError += text; + }, + }); program.exitOverride((err) => { /* v8 ignore next -- the `exitCode === 0` arm only fires for --help/--version, which cannot run through the in-process test runner (it would call the real process.exit(0) and tear down the vitest worker). That UX is covered out-of-process in e2e.test.ts; here only the throwing arm is exercised. */ - if (err.exitCode !== 0) throw err; + if (err.exitCode !== 0) { + if (program.opts().quiet !== true) { + process.stderr.write(commanderError); + } + throw err; + } }); const rawArgs = argv ?? process.argv; const scriptArgs = rawArgs.slice(2); @@ -742,6 +860,7 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { "Read-only session config file (served as-is, never written or seeded; errors if absent)", ) .option("--server <name>", "Server name from config/catalog file") + .option("--rename <name>", "New name for the entry (servers/edit only)") .option( "-e <env>", "Environment variables for the server (KEY=VALUE)", @@ -855,6 +974,16 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { "Protocol era to negotiate: legacy, auto, or modern. Overrides the file-level protocolEra for --catalog/--config runs; ad-hoc --server-url / target runs otherwise default to legacy.", parseProtocolEra, ) + .option( + "--skill-catalog-max-skills <n>", + "Skills catalog budget for --verify: the most skills one run reads (positive integer; default 256). Overrides the server's skillCatalogMaxSkills in the catalog/config file.", + skillCatalogLimitParser("--skill-catalog-max-skills"), + ) + .option( + "--skill-catalog-max-bytes <n>", + "Skills catalog budget for --verify: the most bytes one run reads across all skills (positive integer; default 67108864, 64 MiB). Overrides the server's skillCatalogMaxBytes in the catalog/config file.", + skillCatalogLimitParser("--skill-catalog-max-bytes"), + ) .option( "--format <format>", "Output format: text (default; pretty-printed) or json (one JSON object on stdout, no banners).", @@ -865,6 +994,19 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { return v; }, ) + .option( + "-q, --quiet", + "Suppress everything except the result payload on stdout (or the error envelope on stderr): status lines, warnings, advisory summaries, and a stdio server's own stderr. Interactive OAuth prompts still appear when a login is needed.", + ) + .option( + "--output <path>", + "Write the result to this file instead of stdout (the parent directory must exist; an existing file is replaced).", + ) + .option( + "--output-format <format>", + "Encoding of the --output file: json (default; the whole result, pretty-printed) or raw (the result's text, or the decoded bytes of a single image/audio/blob block). raw needs --method tools/call or resources/read.", + parseOutputFileFormat, + ) .option( "--tool-args-json <json>", 'Tool arguments as a single JSON object (e.g. \'{"zip":"10001"}\'). Values are passed verbatim — no key=value coercion. Mutually exclusive with --tool-arg.', @@ -927,12 +1069,25 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { "Print a JSON handoff block (deepLink, portForwardCmd, oauthStatePath, apiToken) for --server-url and exit. No server connection is made.", ); + registerCompletionOption(program); program.parse(preArgs); + // `--completion <shell>` (#2434): print the script and exit, no connect. + if (await emitCompletionIfRequested(program)) return { shortCircuit: true }; + + // `--quiet` mutes `console.warn` for the rest of the run; `runCli` restores + // it (#2435). Core reports advisories that way from many places the CLI + // reaches (the secret-store notice, `roots` / OAuth-endpoint settings it + // ignores, lock and persistence trouble), so muting the channel is the only + // complete answer. What `--quiet` keeps — the result, the envelope, the + // `--strict` report, the OAuth URL and step-up prompt — is written to the + // streams directly, never through `console.warn`. + if (program.opts().quiet === true) console.warn = discardWarning; const options = program.opts() as { catalog?: string; config?: string; server?: string; + rename?: string; e?: Record<string, string>; method?: string; toolName?: string; @@ -955,7 +1110,12 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { cursor?: string; connectTimeout?: number; protocolEra?: ServerProtocolEra; + skillCatalogMaxSkills?: number; + skillCatalogMaxBytes?: number; format?: OutputFormat; + quiet?: boolean; + output?: string; + outputFormat?: OutputFileFormat; toolArgsJson?: string; clientConfig?: string; clientId?: string; @@ -987,10 +1147,11 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { } if ( options.method === "servers/list" || - options.method === "servers/show" + options.method === "servers/show" || + isCatalogWriteMethod(options.method) ) { throw new Error( - "--relogin cannot be combined with --method servers/list or servers/show (no OAuth connect)", + "--relogin cannot be combined with a --method servers/* catalog command (no OAuth connect)", ); } } @@ -1004,6 +1165,10 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { throw new Error("--no-revoke requires --relogin (it has no other effect)."); } + // `--output` / `--output-format` (#2431), ahead of the short-circuit returns + // for the same reason as `--strict` below. + validateOutputOptions(options); + // `--strict` is checked HERE, ahead of every short-circuit return below // (`--list-stored-auth`, `--print-handoff`, `servers/list`, `servers/show`), // rather than beside the other method-shaped validations further down. Those @@ -1041,6 +1206,15 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { if (options.requireDigests && !options.verify) { throw new Error("--require-digests requires --verify."); } + // The skills catalog budget is read only by `verifySkills`, so without + // `--verify` it bounds nothing — and a job that set it would look bounded + // when it is not (#2420). + if (options.skillCatalogMaxSkills !== undefined && !options.verify) { + throw new Error("--skill-catalog-max-skills requires --verify."); + } + if (options.skillCatalogMaxBytes !== undefined && !options.verify) { + throw new Error("--skill-catalog-max-bytes requires --verify."); + } // `--advertise-apps` is checked here for the same reason: it shapes the // `initialize` handshake, and the short-circuit paths below never open an @@ -1050,10 +1224,11 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { (options.listStoredAuth || options.printHandoff || options.method === "servers/list" || - options.method === "servers/show") + options.method === "servers/show" || + isCatalogWriteMethod(options.method)) ) { throw new Error( - "--advertise-apps requires a command that connects to a server; it has no effect with --list-stored-auth, --print-handoff, or --method servers/list / servers/show.", + "--advertise-apps requires a command that connects to a server; it has no effect with --list-stored-auth, --print-handoff, or a --method servers/* catalog command.", ); } @@ -1066,7 +1241,9 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { if (options.listStoredAuth) { const servers = await readOAuthServers(oauthStatePath); const withToken = Object.entries(servers) - .filter(([, v]) => Boolean(v.tokens?.access_token)) + .filter(([, v]) => + Boolean(resolveStoredCredentials(v).tokens?.access_token), + ) .map(([k]) => k); await awaitableLog( JSON.stringify({ oauthStatePath, storedServerUrls: withToken }) + "\n", @@ -1092,13 +1269,38 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { "Method is required. Use --method to specify the method to invoke.", ); } - const isCatalogMethod = - options.method === "servers/list" || options.method === "servers/show"; - if (!isCatalogMethod && !isOneShotMethod(options.method)) { + if (!isCatalogMethod(options.method) && !isOneShotMethod(options.method)) { throw new Error( - `Unsupported method: ${options.method}. Supported --cli methods: ${ONE_SHOT_METHODS.join(", ")}, servers/list, servers/show.`, + `Unsupported method: ${options.method}. Supported --cli methods: ${[...ONE_SHOT_METHODS, ...CATALOG_METHODS].join(", ")}.`, ); } + if (options.rename !== undefined && options.method !== "servers/edit") { + throw new Error("--rename is only valid with --method servers/edit."); + } + + // Catalog add / edit / remove (#2433) — no MCP connection. Resolves its own + // writable catalog path, since the positional target / --server-url here + // describe the entry being written, not an ad-hoc server to connect to. + if (isCatalogWriteMethod(options.method)) { + const written = await runCatalogWrite(options.method, { + server: options.server, + rename: options.rename, + catalog: options.catalog, + config: options.config, + target: targetArgs, + transport: options.transport, + serverUrl: options.serverUrl, + cwd: options.cwd, + env: options.e, + headers: options.header as Record<string, string> | undefined, + protocolEra: options.protocolEra, + }); + await writeFormattedResult( + written, + options.format === "json" ? "json" : "text", + ); + return { shortCircuit: true }; + } // Honour MCP_CATALOG_PATH only when no ad-hoc target is given. Applying it // unconditionally meant a homespace that exports the env var could never run @@ -1125,6 +1327,10 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { // `--protocol-era` feeds `settings.protocolEra` the same way, so an ad-hoc // launch can pick a non-legacy era without an mcp.json entry (#2208). protocolEra: options.protocolEra, + // `--skill-catalog-max-*` feed the per-server skills catalog budget the + // web sets in Server Settings, read by `verifySkills` (#2420). + skillCatalogMaxSkills: options.skillCatalogMaxSkills, + skillCatalogMaxBytes: options.skillCatalogMaxBytes, }; // Catalog list / show — no MCP connection. Run before stored-auth refresh so @@ -1176,8 +1382,11 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { } else { const servers = await readOAuthServers(oauthStatePath); const stored = findStoredServerState(servers, options.serverUrl); - if (stored?.state.tokens?.refresh_token) { - const storedAccess = stored.state.tokens.access_token; + const storedTokens = stored + ? resolveStoredCredentials(stored.state).tokens + : undefined; + if (storedTokens?.refresh_token) { + const storedAccess = storedTokens.access_token; try { token = await refreshStoredAuthToken( options.serverUrl, @@ -1293,6 +1502,9 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { requireDigests: options.requireDigests === true, cursor: options.cursor, format: options.format, + quiet: options.quiet === true, + output: options.output, + outputFormat: options.outputFormat, }; return { @@ -1312,7 +1524,22 @@ async function parseArgs(argv?: string[]): Promise<ParseResult> { }; } +/** Stands in for `console.warn` during a `--quiet` run (see `parseArgs`). */ +function discardWarning(): void {} + export async function runCli(argv?: string[]): Promise<void> { + // Restored in `finally`, so a `--quiet` run cannot leave `console.warn` + // muted for whatever else shares the process — the launcher imports + // `runCli` as a module, and so does the in-process test runner. + const warn = console.warn; + try { + await runParsedCli(argv); + } finally { + if (console.warn === discardWarning) console.warn = warn; + } +} + +async function runParsedCli(argv?: string[]): Promise<void> { const parsed = await parseArgs(argv ?? process.argv); // `--list-stored-auth` / `--print-handoff` already wrote their output. if (parsed.shortCircuit) return; diff --git a/clients/cli/src/completion.ts b/clients/cli/src/completion.ts new file mode 100644 index 0000000000..90cfdf4b7f --- /dev/null +++ b/clients/cli/src/completion.ts @@ -0,0 +1,365 @@ +/** + * Shell completion for `mcp-inspector --cli` (#2434). + * + * `mcp-inspector --cli --completion <bash|zsh|fish>` prints a completion + * script on stdout and exits without connecting to anything. The flag list in + * that script is read from the live commander `program` that `parseArgs` + * builds, not from a second hand-maintained list, so a flag added to (or + * removed from) the CLI shows up in (or drops out of) the completions with no + * edit here. The same goes for which flags take a value and which take a path: + * both come from each option's commander `flags` string (`<value>`, `<path>`). + * + * The one thing commander cannot tell us is the finite set of values a flag + * accepts, because the CLI validates those in custom argument parsers rather + * than with commander `choices`. `VALUE_CHOICES` below states those sets; the + * method list reuses `ONE_SHOT_METHODS`, the same list `parseArgs` validates + * `--method` against, and the tests assert every other choice is accepted by + * the option's own parser so a stale value fails loudly. + * + * The scripts complete the published `mcp-inspector` bin: the first word + * offers the launcher's mode flags, and CLI flags are offered only once that + * first word is `--cli` (web and TUI flags are out of scope and fall back to + * the shell's default file completion). + */ +import type { Command, Option } from "commander"; +import { LoggingLevelSchema } from "@modelcontextprotocol/core"; +import { ONE_SHOT_METHODS } from "@inspector/core/cli/handlers/method-types.js"; +import { awaitableLog } from "@inspector/core/cli/utils/awaitable-log.js"; +import { OUTPUT_FILE_FORMATS } from "@inspector/core/cli/handlers/output-file.js"; +import { CATALOG_WRITE_METHODS } from "./handlers/servers-write.js"; + +export const COMPLETION_SHELLS = ["bash", "zsh", "fish"] as const; +export type CompletionShell = (typeof COMPLETION_SHELLS)[number]; + +/** The bin the scripts register for (the root package's only `bin`). */ +export const COMPLETION_COMMAND = "mcp-inspector"; + +/** Launcher mode flags, offered as the first word. */ +const MODE_FLAGS: readonly CompletionFlag[] = [ + { long: "--cli", takesValue: false, description: "Run the CLI" }, + { long: "--web", takesValue: false, description: "Run the web UI" }, + { long: "--tui", takesValue: false, description: "Run the terminal UI" }, +]; + +/** + * Catalog-only methods `parseArgs` accepts alongside `ONE_SHOT_METHODS` — + * the reads plus the writes `servers-write` implements. `parseArgs` validates + * `--method` against this same list, so the two cannot drift (#2629). + */ +export const CATALOG_METHODS = [ + "servers/list", + "servers/show", + ...CATALOG_WRITE_METHODS, +] as const; + +export function isCatalogMethod(method: string): boolean { + return (CATALOG_METHODS as readonly string[]).includes(method); +} + +/** + * Finite value sets for flags whose values the CLI validates in a custom + * parser. Keyed by the option's long name. + */ +export const VALUE_CHOICES: Readonly<Record<string, readonly string[]>> = { + "--method": [...ONE_SHOT_METHODS, ...CATALOG_METHODS], + "--transport": ["stdio", "sse", "http"], + "--log-level": Object.values(LoggingLevelSchema.enum), + "--format": ["text", "json"], + "--output-format": OUTPUT_FILE_FORMATS, + "--protocol-era": ["legacy", "auto", "modern"], + "--completion": COMPLETION_SHELLS, +}; + +export interface CompletionFlag { + /** `--name`, when the option has a long form. */ + long?: string; + /** `-e`, when the option has a short form. */ + short?: string; + takesValue: boolean; + /** Value completes as a file/directory path. */ + path?: boolean; + /** Finite set of accepted values. */ + choices?: readonly string[]; + description: string; +} + +export function isCompletionShell(value: string): value is CompletionShell { + return (COMPLETION_SHELLS as readonly string[]).includes(value); +} + +/** commander argument parser for `--completion <shell>`. */ +export function parseCompletionShell(value: string): CompletionShell { + if (!isCompletionShell(value)) { + throw new Error( + `Invalid shell: ${value}. Supported shells are: ${COMPLETION_SHELLS.join(", ")}`, + ); + } + return value; +} + +/** Shortest sentence-boundary prefix that reads as a summary. */ +const MIN_SUMMARY_LENGTH = 20; + +/** + * Summarize a help string for a completion menu: the first sentence, or the + * first few when the first alone is too terse to say what the flag does + * (`--no-revoke`'s starts "Requires --relogin."), with no trailing period. + */ +function summarize(description: string): string { + const sentences = description.split(/(?<=\.)\s+(?=[A-Z])/); + let summary = ""; + for (const sentence of sentences) { + summary = summary ? `${summary} ${sentence}` : sentence; + if (summary.length >= MIN_SUMMARY_LENGTH) break; + } + return summary.trim().replace(/\.$/, ""); +} + +function toCompletionFlag(option: Option): CompletionFlag { + const takesValue = option.required || option.optional; + const flag: CompletionFlag = { + takesValue, + description: summarize(option.description), + }; + if (option.long) flag.long = option.long; + if (option.short) flag.short = option.short; + if (takesValue && /<path>|\[path\]/.test(option.flags)) flag.path = true; + const choices = + option.argChoices ?? (option.long && VALUE_CHOICES[option.long]); + if (takesValue && choices) flag.choices = choices; + return flag; +} + +/** Every visible option on `program`, plus commander's implicit `--help`. */ +export function collectCompletionFlags(program: Command): CompletionFlag[] { + const flags = program.options + .filter((option) => !option.hidden) + .map(toCompletionFlag); + flags.push({ + long: "--help", + short: "-h", + takesValue: false, + description: "Display help for command", + }); + return flags; +} + +function names(flag: CompletionFlag): string[] { + return [flag.long, flag.short].filter((n): n is string => Boolean(n)); +} + +/** Quote for a POSIX/zsh single-quoted string. */ +function shQuote(value: string): string { + return `'${value.replace(/'/g, `'\\''`)}'`; +} + +/** Quote for a fish single-quoted string. */ +function fishQuote(value: string): string { + return `'${value.replace(/\\/g, "\\\\").replace(/'/g, "\\'")}'`; +} + +const FUNCTION_NAME = "_mcp_inspector"; + +function header(shell: CompletionShell): string { + return `# ${shell} completion for ${COMPLETION_COMMAND} --cli (generated by \`${COMPLETION_COMMAND} --cli --completion ${shell}\`).`; +} + +export function renderBash(flags: readonly CompletionFlag[]): string { + const allNames = flags.flatMap(names).join(" "); + const modeNames = MODE_FLAGS.flatMap(names).join(" "); + const choiceCases = flags + .filter((f) => f.choices) + .map( + (f) => + ` ${names(f).join("|")})\n COMPREPLY=( $(compgen -W ${shQuote(f.choices!.join(" "))} -- "$cur") )\n return 0 ;;`, + ); + // A path or a free-form value: complete nothing, so `-o default` falls back + // to file names rather than offering flags as the value. + const freeValueNames = flags + .filter((f) => f.takesValue && !f.choices) + .flatMap(names); + const cases = [ + ...choiceCases, + ` ${freeValueNames.join("|")})\n return 0 ;;`, + ].join("\n"); + return `${header("bash")} +# Load it with: source <(${COMPLETION_COMMAND} --cli --completion bash) +${FUNCTION_NAME}() { + local cur="\${COMP_WORDS[COMP_CWORD]}" + local prev="\${COMP_WORDS[COMP_CWORD-1]}" + COMPREPLY=() + if [ "$COMP_CWORD" -eq 1 ]; then + COMPREPLY=( $(compgen -W ${shQuote(modeNames)} -- "$cur") ) + return 0 + fi + [ "\${COMP_WORDS[1]}" = "--cli" ] || return 0 + # \`--opt=value\`: "=" is in COMP_WORDBREAKS, so bash splits it into + # \`--opt\`, \`=\`, \`value\`. Readline only replaces the text after the "=", + # so complete the bare value against the option before it. + if [ "$cur" = "=" ]; then + cur="" + elif [ "$prev" = "=" ] && [ "$COMP_CWORD" -ge 3 ]; then + prev="\${COMP_WORDS[COMP_CWORD-2]}" + fi + case "$prev" in +${cases} + esac + if [[ "$cur" == -* ]]; then + COMPREPLY=( $(compgen -W ${shQuote(allNames)} -- "$cur") ) + fi + return 0 +} +complete -o default -F ${FUNCTION_NAME} ${COMPLETION_COMMAND} +`; +} + +/** + * `name:description` for zsh `_describe`. `_describe` reads `\` as an escape + * and the first unescaped `:` as the separator, so backslashes in the name are + * escaped first, then colons (CodeQL #78). + */ +function zshDescribeEntry(name: string, description: string): string { + const escaped = name.replace(/\\/g, "\\\\").replace(/:/g, "\\:"); + return shQuote(`${escaped}:${description}`); +} + +export function renderZsh(flags: readonly CompletionFlag[]): string { + const optionEntries = flags + .flatMap((f) => names(f).map((n) => zshDescribeEntry(n, f.description))) + .map((entry) => ` ${entry}`) + .join("\n"); + const modeEntries = MODE_FLAGS.flatMap((f) => + names(f).map((n) => zshDescribeEntry(n, f.description)), + ).join(" "); + const cases = flags + .filter((f) => f.takesValue) + .map((f) => { + const pattern = names(f).join("|"); + if (f.choices) { + const values = f.choices.map(shQuote).join(" "); + return ` ${pattern})\n compadd -- ${values}\n return ;;`; + } + if (f.path) return ` ${pattern})\n _files\n return ;;`; + return ` ${pattern})\n _message 'value'\n return ;;`; + }) + .join("\n"); + return `#compdef ${COMPLETION_COMMAND} +${header("zsh")} +# Load it with: source <(${COMPLETION_COMMAND} --cli --completion zsh) +# or save it as _${COMPLETION_COMMAND} in a directory on your $fpath. +${FUNCTION_NAME}() { + local cur=\${words[CURRENT]} prev=\${words[CURRENT-1]} + if (( CURRENT == 2 )); then + local -a modes + modes=(${modeEntries}) + _describe -t modes 'mode' modes + return + fi + if [[ \${words[2]} != --cli ]]; then + _files + return + fi + # \`--opt=value\` stays one word: split it, and move \`--opt=\` into IPREFIX + # so only the value is replaced. + local inline=0 + if [[ $cur == --*=* ]]; then + prev=\${cur%%=*} + compset -P '*=' + cur=\${cur#*=} + inline=1 + fi + case $prev in +${cases} + esac + (( inline )) && return + if [[ $cur == -* ]]; then + local -a opts + opts=( +${optionEntries} + ) + _describe -t options 'option' opts + else + _files + fi +} +if [[ \${funcstack[1]} == _${COMPLETION_COMMAND} ]]; then + ${FUNCTION_NAME} "$@" +else + compdef ${FUNCTION_NAME} ${COMPLETION_COMMAND} +fi +`; +} + +export function renderFish(flags: readonly CompletionFlag[]): string { + const cmd = COMPLETION_COMMAND; + const lines = flags.map((f) => { + const parts = [`complete -c ${cmd} -n __mcp_inspector_cli_mode`]; + if (f.long) parts.push(`-l ${f.long.slice(2)}`); + if (f.short) parts.push(`-s ${f.short.slice(1)}`); + if (f.choices) { + parts.push(`-x -a ${fishQuote(f.choices.join(" "))}`); + } else if (f.path) { + parts.push("-r -F"); + } else if (f.takesValue) { + parts.push("-x"); + } + parts.push(`-d ${fishQuote(f.description)}`); + return parts.join(" "); + }); + const modeLines = MODE_FLAGS.map( + (f) => + `complete -c ${cmd} -n __mcp_inspector_first_arg -l ${f.long!.slice(2)} -d ${fishQuote(f.description)}`, + ); + return `${header("fish")} +# Load it with: ${cmd} --cli --completion fish | source +# or save it as ~/.config/fish/completions/${cmd}.fish +function __mcp_inspector_first_arg + test (count (commandline -opc)) -eq 1 +end +function __mcp_inspector_cli_mode + set -l tokens (commandline -opc) + test (count $tokens) -ge 2; and test "$tokens[2]" = --cli +end +${modeLines.join("\n")} +${lines.join("\n")} +`; +} + +export function renderCompletion( + shell: CompletionShell, + flags: readonly CompletionFlag[], +): string { + switch (shell) { + case "bash": + return renderBash(flags); + case "zsh": + return renderZsh(flags); + case "fish": + return renderFish(flags); + } +} + +/** + * Register `--completion <shell>` on `program`. Called before + * `program.parse` so the flag is accepted and listed in `--help`. + */ +export function registerCompletionOption(program: Command): void { + program.option( + "--completion <shell>", + `Print a shell completion script (${COMPLETION_SHELLS.join(", ")}) on stdout and exit. No server connection is made.`, + parseCompletionShell, + ); +} + +/** + * After `program.parse`: when `--completion` was given, write the script for + * that shell and return true so the caller can short-circuit. + */ +export async function emitCompletionIfRequested( + program: Command, +): Promise<boolean> { + const shell = program.opts<{ completion?: CompletionShell }>().completion; + if (!shell) return false; + await awaitableLog(renderCompletion(shell, collectCompletionFlags(program))); + return true; +} diff --git a/clients/cli/src/handlers/consume-outcome.ts b/clients/cli/src/handlers/consume-outcome.ts index 1ce45968a8..45583ae002 100644 --- a/clients/cli/src/handlers/consume-outcome.ts +++ b/clients/cli/src/handlers/consume-outcome.ts @@ -1,7 +1,16 @@ -import { awaitableError, awaitableLog } from "../utils/awaitable-log.js"; -import { CliExitCodeError, EXIT_CODES } from "../error-handler.js"; +import { + awaitableError, + awaitableLog, +} from "@inspector/core/cli/utils/awaitable-log.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; import { emitResult } from "./emit-result.js"; -import type { MethodArgs, MethodOutcome } from "./method-types.js"; +import type { + MethodArgs, + MethodOutcome, +} from "@inspector/core/cli/handlers/method-types.js"; /** * True for the error a write to a closed pipe raises — the reader went away @@ -33,8 +42,12 @@ export async function consumeMethodOutcome( await awaitableLog(JSON.stringify(line) + "\n"); } // Summary on **stderr**, after the report, so it cannot contaminate the - // NDJSON a consumer is parsing on stdout. - if (outcome.summary) await awaitableError(`${outcome.summary}\n`); + // NDJSON a consumer is parsing on stdout. `--quiet` drops it: a failing + // report still reaches the error envelope below, which carries the same + // summary as its message (#2435). + if (outcome.summary && !args.quiet) { + await awaitableError(`${outcome.summary}\n`); + } // Thrown rather than returned so it routes through the CLI's single exit // path — the report has already been written, which is why this is the // last thing that happens. diff --git a/clients/cli/src/handlers/emit-result.ts b/clients/cli/src/handlers/emit-result.ts index 459a6e7e42..e8488ad715 100644 --- a/clients/cli/src/handlers/emit-result.ts +++ b/clients/cli/src/handlers/emit-result.ts @@ -1,8 +1,19 @@ -import { awaitableLog } from "../utils/awaitable-log.js"; -import { CliExitCodeError, EXIT_CODES } from "../error-handler.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; import { lintListResult, writeSchemaLintReport } from "./schema-lint-report.js"; import { countFindings } from "@inspector/core/json/schemaLint.js"; -import type { CliAppInfo, McpResponse, MethodArgs } from "./method-types.js"; +import type { + CliAppInfo, + McpResponse, + MethodArgs, +} from "@inspector/core/cli/handlers/method-types.js"; +import { writeResultFile } from "@inspector/core/cli/handlers/output-file.js"; +import { + awaitableError, + awaitableLog, +} from "@inspector/core/cli/utils/awaitable-log.js"; /** * Write the method result (and any app-info) to stdout, honouring `--format` @@ -37,11 +48,31 @@ export async function emitResult( const lint = args.method === "tools/list" ? lintListResult(result) : undefined; + // `--output` (#2431): the result goes to the file INSTEAD of stdout, as with + // `curl -o`. Under `--format json` stdout still carries exactly one envelope — + // `{ output }` in place of `{ result }` — so a caller parsing it learns where + // the result went; in text mode the confirmation goes to stderr, leaving + // stdout empty. + const written = args.output + ? await writeResultFile( + result, + args.method ?? "", + args.output, + args.outputFormat, + ) + : undefined; + if (json) { - const envelope: Record<string, unknown> = { result }; + const envelope: Record<string, unknown> = written + ? { output: written } + : { result }; if (appInfo?.hasApp) envelope.appInfo = appInfo; if (args.strict && lint && lint.length > 0) envelope.schemaFindings = lint; await awaitableLog(JSON.stringify(envelope) + "\n"); + } else if (written) { + await awaitableError( + `Wrote ${written.bytes} bytes (${written.format}) to ${written.path}\n`, + ); } else { await awaitableLog(JSON.stringify(result, null, 2) + "\n"); } @@ -49,7 +80,13 @@ export async function emitResult( // Awaited: the throw below (and the CLI's own exit path) reaches // `process.exit()` immediately, which discards anything still buffered on a // piped stderr. - if (lint) await writeSchemaLintReport(lint, args.strict === true); + if (lint) { + await writeSchemaLintReport( + lint, + args.strict === true, + args.quiet === true, + ); + } if ((result as { isError?: unknown }).isError === true) { throw new CliExitCodeError( diff --git a/clients/cli/src/handlers/schema-lint-report.ts b/clients/cli/src/handlers/schema-lint-report.ts index a38aaa9337..2ec2d50998 100644 --- a/clients/cli/src/handlers/schema-lint-report.ts +++ b/clients/cli/src/handlers/schema-lint-report.ts @@ -5,8 +5,8 @@ import { summarizeFindings, type ToolSchemaFindings, } from "@inspector/core/json/schemaLint.js"; -import { awaitableError } from "../utils/awaitable-log.js"; -import type { McpResponse } from "./method-types.js"; +import { awaitableError } from "@inspector/core/cli/utils/awaitable-log.js"; +import type { McpResponse } from "@inspector/core/cli/handlers/method-types.js"; /** * Read the `tools` array out of a `tools/list` result. The result is typed as @@ -47,13 +47,19 @@ export function lintListResult(result: McpResponse): ToolSchemaFindings[] { * count and how to see the detail: a server author who has not asked for the * lint should still learn it found something, but a multi-page report nobody * requested would be worse than silence. + * + * `quiet` (`--quiet`) drops that one-line hint, since it is advisory. It does + * not drop the `--strict` report: that was asked for explicitly, and it is the + * detail behind the exit-6 failure the error envelope only counts (#2435). */ export async function writeSchemaLintReport( results: readonly ToolSchemaFindings[], strict: boolean, + quiet = false, ): Promise<void> { if (results.length === 0) return; if (!strict) { + if (quiet) return; await awaitableError( `Schema portability: ${summarizeFindings(results)}. Re-run with --strict for details.\n`, ); diff --git a/clients/cli/src/handlers/servers-write.ts b/clients/cli/src/handlers/servers-write.ts new file mode 100644 index 0000000000..e3722ab2e5 --- /dev/null +++ b/clients/cli/src/handlers/servers-write.ts @@ -0,0 +1,513 @@ +/** + * The CLI's write path for the server catalog (#2433): `servers/add`, + * `servers/edit` and `servers/remove`. + * + * ## Why this drives the web backend's routes in-process + * + * The catalog is not just a JSON file. The web backend's `/api/servers` + * routes (`core/mcp/remote/node/server.ts`) own every invariant a write has + * to keep: secret values (stdio `env`, the OAuth client secret) are split out + * of `mcp.json` into the selected secret store with a compensated + * keychain → disk → cleanup ordering; ids are validated + * (`validateStoreId`); duplicate and missing ids are checked with + * own-property lookups; Inspector-extension fields are smuggle-guarded and + * normalized; and the file is written atomically through `writeStoreFile`. + * A second writer here would have to re-implement all of that and would + * drift from it the first time either side changed. + * + * So this module builds the same Hono app (`createRemoteApp`) against the + * catalog file and calls its routes with `app.request()` — no socket, no + * listener, no auth token (`dangerouslyOmitAuth` is safe because nothing is + * exposed: the requests never leave this process). The CLI therefore writes + * byte-for-byte what the web UI would, and a running web backend picks the + * change up through its file watcher like any other external edit. + * + * ## Read-only sources are refused + * + * Only the writable catalog (`--catalog`, `MCP_CATALOG_PATH`, or the default + * `~/.mcp-inspector/mcp.json`) can be written. `--config` names a read-only + * session file and is rejected outright, matching the web backend's 403. + * + * ## A session-scoped secret store is refused for new secrets + * + * The web backend, on a non-durable (in-memory) secret store, keeps newly + * entered secrets in memory for the life of the process and never writes + * them to disk. For a long-running web server that is a session; for a CLI + * that exits immediately it is silent data loss. So when the store is not + * durable and this invocation supplies `-e` values, the write is refused + * with a pointer to `MCP_INSPECTOR_SECRET_STORE`. + */ +import { resolve } from "node:path"; +import { createRemoteApp } from "@inspector/core/mcp/remote/node/server.js"; +import { + defaultSecretStore, + SECRET_STORE_ENV, +} from "@inspector/core/auth/node/secret-store-selection.js"; +import { + secretStoreIsDurable, + type SecretStore, +} from "@inspector/core/auth/node/secret-store.js"; +import { + getDefaultMcpConfigPath, + parseStore, + readStoreFile, +} from "@inspector/core/storage/store-io.js"; +import { + resolveServerConfigs, + type ServerConfigOptions, +} from "@inspector/core/mcp/node/config.js"; +import { headersToServerSettings } from "@inspector/core/mcp/node/servers.js"; +import { + storedFieldsToInspectorSettings, + stripInspectorFields, +} from "@inspector/core/mcp/serverList.js"; +import type { + InspectorServerSettings, + MCPConfig, + MCPServerConfig, + ServerProtocolEra, + StoredMCPServer, +} from "@inspector/core/mcp/types.js"; +import { + DEFAULT_CONNECTION_TIMEOUT_MS, + DEFAULT_MAX_FETCH_REQUESTS, + DEFAULT_TASK_TTL_MS, +} from "@inspector/core/mcp/types.js"; + +/** The catalog-mutating `--method` values this module implements. */ +export const CATALOG_WRITE_METHODS = [ + "servers/add", + "servers/edit", + "servers/remove", +] as const; + +export type CatalogWriteMethod = (typeof CATALOG_WRITE_METHODS)[number]; + +export function isCatalogWriteMethod( + method: string | undefined, +): method is CatalogWriteMethod { + return (CATALOG_WRITE_METHODS as readonly string[]).includes(method ?? ""); +} + +/** The flags a catalog write reads, already parsed by commander. */ +export interface CatalogWriteOptions { + /** `--server`: the entry to add, edit, or remove. */ + server?: string; + /** `--rename`: new name for `servers/edit`. */ + rename?: string; + /** `--catalog`. */ + catalog?: string; + /** `--config` — present only so it can be rejected as read-only. */ + config?: string; + /** Positional command + args, or a URL. */ + target?: string[]; + transport?: "sse" | "http" | "stdio"; + serverUrl?: string; + cwd?: string; + env?: Record<string, string>; + headers?: Record<string, string>; + protocolEra?: ServerProtocolEra; + /** Test injection; defaults to the host-selected store. */ + secretStore?: SecretStore; +} + +/** What a successful write reports on stdout. */ +export interface CatalogWriteResult { + ok: true; + action: "added" | "updated" | "removed"; + server: string; + /** Present on a rename. */ + previousName?: string; + catalog: string; +} + +/** + * Resolve the writable catalog path, refusing a read-only `--config`. + * Precedence matches the read side: `--catalog` → `MCP_CATALOG_PATH` → the + * default `~/.mcp-inspector/mcp.json`. Relative paths resolve against the + * current directory, as `readServerListFile` resolves them. + */ +export function resolveWritableCatalogPath( + options: Pick<CatalogWriteOptions, "catalog" | "config">, + env: NodeJS.ProcessEnv = process.env, +): string { + // Presence, not a non-blank value: `--config ""` must not fall through and + // quietly write the default catalog instead. + if (options.config !== undefined) { + throw new Error( + "--config names a read-only session file; catalog writes need the writable catalog (--catalog <path>, MCP_CATALOG_PATH, or the default ~/.mcp-inspector/mcp.json).", + ); + } + const path = + options.catalog?.trim() || + env.MCP_CATALOG_PATH?.trim() || + getDefaultMcpConfigPath(); + return resolve(process.cwd(), path); +} + +function hasTransportTarget(options: CatalogWriteOptions): boolean { + return ( + (options.target?.length ?? 0) > 0 || Boolean(options.serverUrl?.trim()) + ); +} + +/** + * Build the SDK transport config from the same ad-hoc flags a one-shot run + * takes (positional command/URL, `--transport`, `--server-url`, `-e`, + * `--cwd`), through the shared `resolveServerConfigs` so `servers/add` and an + * ad-hoc connect can never interpret the flags differently. + */ +function buildTransportConfig(options: CatalogWriteOptions): MCPServerConfig { + const adHoc: ServerConfigOptions = { + target: options.target, + transport: options.transport, + serverUrl: options.serverUrl, + cwd: options.cwd, + env: options.env, + }; + // "single" with no catalog/config source always resolves the ad-hoc target. + const config = resolveServerConfigs(adHoc, "single")[0]!; + // The shared builder drops env/cwd for a URL target; reject them instead so + // a write never reports success for flags it did not store. + if (config.type === "sse" || config.type === "streamable-http") { + const stdioOnly = [ + Object.keys(options.env ?? {}).length > 0 && "-e", + Boolean(options.cwd?.trim()) && "--cwd", + ].filter((flag): flag is string => typeof flag === "string"); + if (stdioOnly.length > 0) { + throw new Error( + `${stdioOnly.join(" and ")} apply to stdio servers; this target is ${config.type}.`, + ); + } + } + return config; +} + +/** A settings node at product defaults, for an entry that had none. */ +function defaultSettings(): InspectorServerSettings { + return { + headers: [], + env: [], + metadata: {}, + connectionTimeout: DEFAULT_CONNECTION_TIMEOUT_MS, + requestTimeout: 0, + taskTtl: DEFAULT_TASK_TTL_MS, + maxFetchRequests: DEFAULT_MAX_FETCH_REQUESTS, + autoRefreshOnListChanged: false, + paginatedLists: false, + roots: [], + }; +} + +/** + * Overlay `--header` / `--protocol-era` onto a settings node. Returns + * undefined when neither flag was given, so the caller sends no `settings` + * and the route preserves (PUT) or omits (POST) the node. + * + * `env` / `cwd` are removed: those are config fields the settings node only + * mirrors, and every write here sends the config, which owns them. + */ +function settingsWithOverrides( + base: InspectorServerSettings | undefined, + options: CatalogWriteOptions, +): InspectorServerSettings | undefined { + const fromHeaders = headersToServerSettings(options.headers); + if (!fromHeaders && !options.protocolEra) return undefined; + const next: InspectorServerSettings = { ...(base ?? defaultSettings()) }; + if (fromHeaders) next.headers = fromHeaders.headers; + if (options.protocolEra) next.protocolEra = options.protocolEra; + next.env = []; + delete next.cwd; + return next; +} + +/** + * Refuse to hand new secret values to a store that dies with this process. + * See the module comment. + */ +async function assertSecretsPersist( + store: SecretStore, + options: CatalogWriteOptions, +): Promise<void> { + if (!options.env || Object.keys(options.env).length === 0) return; + if (await secretStoreIsDurable(store)) return; + throw new Error( + `The selected secret store is in-memory only, so the -e values would be lost when the CLI exits. Set ${SECRET_STORE_ENV}=file (or keyring) to persist them.`, + ); +} + +/** + * Read the catalog's entry map straight off disk; a missing file is empty. + * Keeps only the entries the route layer's normalizer recognizes (object + * values, never `__proto__`), so a key the routes would skip reads as "not + * found" here instead of passing the precheck and then hitting an idempotent + * no-op DELETE that reports success. + */ +async function readCatalogEntries( + catalogPath: string, +): Promise<Record<string, unknown>> { + const raw = await readStoreFile(catalogPath); + if (raw === null) return {}; + const parsed = parseStore(raw) as { mcpServers?: unknown } | null; + const servers = parsed?.mcpServers; + if (servers === null || typeof servers !== "object") return {}; + return Object.fromEntries( + Object.entries(servers as Record<string, unknown>).filter( + ([id, val]) => + id !== "__proto__" && val !== null && typeof val === "object", + ), + ); +} + +type RouteCall = ( + method: "GET" | "POST" | "PUT" | "DELETE", + path: string, + body?: unknown, +) => Promise<unknown>; + +/** + * Run `fn` against an in-process instance of the web backend's routes for + * `catalogPath`, tearing the app down afterwards. A non-2xx response throws + * the route's own `error` message. + */ +async function withCatalogRoutes<T>( + catalogPath: string, + secretStore: SecretStore, + fn: (call: RouteCall) => Promise<T>, +): Promise<T> { + const { app, close } = createRemoteApp({ + dangerouslyOmitAuth: true, + mcpConfigPath: catalogPath, + writable: true, + secretStore, + initialConfig: { defaultEnvironment: {} }, + }); + const call: RouteCall = async (method, path, body) => { + const res = await app.request(path, { + method, + ...(body !== undefined + ? { + headers: { "content-type": "application/json" }, + body: JSON.stringify(body), + } + : {}), + }); + const payload = (await res.json()) as unknown; + if (!res.ok) { + /* v8 ignore next 6 -- every /api/servers error response carries a + string `error`; the status fallback only guards a future route that + answers otherwise. */ + const message = + payload !== null && + typeof payload === "object" && + typeof (payload as { error?: unknown }).error === "string" + ? (payload as { error: string }).error + : `HTTP ${res.status}`; + throw new Error(message); + } + return payload; + }; + try { + return await fn(call); + } finally { + await close(); + } +} + +function requireServerName( + method: CatalogWriteMethod, + server: string | undefined, +): string { + const name = server?.trim(); + if (!name) { + throw new Error(`${method} requires --server <name>.`); + } + return name; +} + +/** `servers/add`: create a new catalog entry. */ +export async function addCatalogServer( + options: CatalogWriteOptions, +): Promise<CatalogWriteResult> { + const name = requireServerName("servers/add", options.server); + if (options.rename !== undefined) { + throw new Error("--rename is only valid with servers/edit."); + } + const catalog = resolveWritableCatalogPath(options); + if (!hasTransportTarget(options)) { + throw new Error( + "servers/add requires a command or URL (positional target, or --server-url).", + ); + } + const config = buildTransportConfig(options); + const settings = settingsWithOverrides(undefined, options); + const store = options.secretStore ?? defaultSecretStore(); + await assertSecretsPersist(store, options); + await withCatalogRoutes(catalog, store, (call) => + call("POST", "/api/servers", { + id: name, + config, + ...(settings ? { settings } : {}), + }), + ); + return { ok: true, action: "added", server: name, catalog }; +} + +/** + * `servers/edit`: change an existing entry. A new target (positional or + * `--server-url`) replaces the transport config; otherwise `-e` merges into + * and `--cwd` replaces the stdio config's fields. `--header` replaces the + * headers, `--protocol-era` sets the era, `--rename` renames. Everything not + * named — OAuth, timeouts, metadata, roots — is preserved. + */ +export async function editCatalogServer( + options: CatalogWriteOptions, +): Promise<CatalogWriteResult> { + const name = requireServerName("servers/edit", options.server); + const catalog = resolveWritableCatalogPath(options); + if (options.rename !== undefined && !options.rename.trim()) { + throw new Error("--rename requires a non-empty name."); + } + const newName = options.rename?.trim() || name; + const replacesTarget = hasTransportTarget(options); + const patchesStdio = + Object.keys(options.env ?? {}).length > 0 || Boolean(options.cwd?.trim()); + const changesSettings = + Object.keys(options.headers ?? {}).length > 0 || + Boolean(options.protocolEra); + if ( + !replacesTarget && + !patchesStdio && + !changesSettings && + newName === name + ) { + throw new Error( + "servers/edit needs something to change: a new command/URL, -e, --cwd, --header, --protocol-era, or --rename.", + ); + } + if (options.transport && !replacesTarget) { + throw new Error( + "--transport on servers/edit requires the new command or URL it applies to.", + ); + } + // Checked against the raw file before the routes run: GET on a missing + // catalog seeds the web UI's sample servers, which an edit must not do. + const onDisk = await readCatalogEntries(catalog); + if (!Object.hasOwn(onDisk, name)) { + throw new Error(`Server '${name}' not found in catalog ${catalog}.`); + } + const store = options.secretStore ?? defaultSecretStore(); + await assertSecretsPersist(store, options); + // A rename moves the entry's secrets from the old name to the new one. A + // session store in this process is empty — whatever it would hold lives in + // the memory of the process that wrote it (a running web backend) — so the + // move would find nothing and the entry would come back under its new name + // without them. + if (newName !== name && !(await secretStoreIsDurable(store))) { + throw new Error( + `--rename needs a durable secret store to carry the entry's secrets to the new name; the selected store is in-memory only. Set ${SECRET_STORE_ENV}=file (or keyring), or rename it from the web UI.`, + ); + } + await withCatalogRoutes(catalog, store, async (call) => { + // GET returns the entry with its secrets rehydrated, so the config sent + // back below still carries the env values the user did not touch. + const current = (await call("GET", "/api/servers")) as MCPConfig; + if (!Object.hasOwn(current.mcpServers, name)) { + throw new Error(`Server '${name}' not found in catalog ${catalog}.`); + } + const existing = current.mcpServers[name] as StoredMCPServer; + let config: MCPServerConfig; + if (replacesTarget) { + config = buildTransportConfig(options); + } else { + config = stripInspectorFields(existing); + if (patchesStdio) { + if (config.type === "sse" || config.type === "streamable-http") { + throw new Error( + `-e and --cwd apply to stdio servers; '${name}' is ${config.type}.`, + ); + } + config = { + ...config, + ...(options.env && Object.keys(options.env).length > 0 + ? { env: { ...(config.env ?? {}), ...options.env } } + : {}), + ...(options.cwd?.trim() ? { cwd: options.cwd.trim() } : {}), + }; + } + } + const settings = settingsWithOverrides( + storedFieldsToInspectorSettings(existing), + options, + ); + await call("PUT", `/api/servers/${encodeURIComponent(name)}`, { + id: newName, + config, + ...(settings ? { settings } : {}), + }); + }); + return { + ok: true, + action: "updated", + server: newName, + ...(newName !== name ? { previousName: name } : {}), + catalog, + }; +} + +/** `servers/remove`: delete an entry and its stored secrets. */ +export async function removeCatalogServer( + options: CatalogWriteOptions, +): Promise<CatalogWriteResult> { + const name = requireServerName("servers/remove", options.server); + assertNoRemoveExtras(options); + const catalog = resolveWritableCatalogPath(options); + // The route's DELETE is idempotent; a CLI user who mistypes a name should + // hear about it rather than see "removed". + const onDisk = await readCatalogEntries(catalog); + if (!Object.hasOwn(onDisk, name)) { + throw new Error(`Server '${name}' not found in catalog ${catalog}.`); + } + const store = options.secretStore ?? defaultSecretStore(); + await withCatalogRoutes(catalog, store, (call) => + call("DELETE", `/api/servers/${encodeURIComponent(name)}`), + ); + return { ok: true, action: "removed", server: name, catalog }; +} + +/** + * Flags `servers/remove` would accept and then ignore. Rejected rather than + * dropped, so a caller never reads success as "my flag took effect". + */ +function assertNoRemoveExtras(options: CatalogWriteOptions): void { + const checks: [boolean, string][] = [ + [hasTransportTarget(options), "a command/URL"], + [Boolean(options.transport), "--transport"], + [Object.keys(options.env ?? {}).length > 0, "-e"], + [Boolean(options.cwd?.trim()), "--cwd"], + [Object.keys(options.headers ?? {}).length > 0, "--header"], + [Boolean(options.protocolEra), "--protocol-era"], + [options.rename !== undefined, "--rename"], + ]; + const extras = checks.filter(([present]) => present).map(([, flag]) => flag); + if (extras.length > 0) { + throw new Error( + `servers/remove takes only --server <name>; remove ${extras.join(", ")}.`, + ); + } +} + +/** Dispatch one catalog-write method. */ +export function runCatalogWrite( + method: CatalogWriteMethod, + options: CatalogWriteOptions, +): Promise<CatalogWriteResult> { + switch (method) { + case "servers/add": + return addCatalogServer(options); + case "servers/edit": + return editCatalogServer(options); + case "servers/remove": + return removeCatalogServer(options); + } +} diff --git a/clients/cli/src/index.ts b/clients/cli/src/index.ts index a03023939b..50decb5e35 100644 --- a/clients/cli/src/index.ts +++ b/clients/cli/src/index.ts @@ -3,7 +3,7 @@ import { resolve } from "path"; import { fileURLToPath } from "url"; import { runCli, validLogLevels } from "./cli.js"; -import { handleError } from "./error-handler.js"; +import { handleError } from "@inspector/core/cli/error-handler.js"; // `handleError` is exported so the launcher (which imports `runCli` as a module // and owns the rejection) can route a `mcp-inspector --cli` failure through the diff --git a/clients/cli/tsup.config.ts b/clients/cli/tsup.config.ts index 724317cc5d..2d2cf61fe5 100644 --- a/clients/cli/tsup.config.ts +++ b/clients/cli/tsup.config.ts @@ -44,6 +44,7 @@ export default defineConfig({ // zod-to-json-schema in with it — which the #2067 guard surfaced. ESM, so it // was not failing the way `undici` did; the rule is what it violated. "@modelcontextprotocol/ext-apps", + "@modelcontextprotocol/ext-tasks", "commander", "pino", // Consolidated to the ROOT manifest by #2195, along with every other diff --git a/clients/cli/vitest.config.ts b/clients/cli/vitest.config.ts index f77d61d588..a71af73522 100644 --- a/clients/cli/vitest.config.ts +++ b/clients/cli/vitest.config.ts @@ -9,6 +9,7 @@ import { const dirname = path.dirname(fileURLToPath(import.meta.url)); const { projectResolve } = vitestSharedPaths(dirname); +const repoRoot = path.resolve(dirname, "../.."); export default defineConfig({ resolve: projectResolve, @@ -35,7 +36,13 @@ export default defineConfig({ coverage: { provider: "v8", reporter: ["text", "html", "json-summary"], - include: ["src/**/*.ts"], + // `core/cli/` is the node-only surface this client shares with mcpdo + // (#2461). It moved out of `src/` but these suites are still what + // exercise it, so it stays gated here; web's `core/*` whitelist + // deliberately omits it, since no web test reaches it. It sits outside + // this project's root, hence `allowExternal`. + include: ["src/**/*.ts", path.join(repoRoot, "core/cli/**/*.ts")], + allowExternal: true, exclude: [ // Binary bootstrap: shebang + `isMain` guard + `runCli()`/`process.exit` // wiring that only runs when launched as the real binary. Exercised by diff --git a/clients/mcpdo/README.md b/clients/mcpdo/README.md new file mode 100644 index 0000000000..3b97479277 --- /dev/null +++ b/clients/mcpdo/README.md @@ -0,0 +1,232 @@ +# MCP Inspector connection CLI (`mcpdo`) + +**Experimental** separate client — **bundled into the published `@modelcontextprotocol/inspector` package** as the `mcpdo` bin. Connect once, then run many MCP commands against a named connection via an implicit local daemon (ssh-agent style). + +> **Layout note:** Source lives in `clients/mcpdo/`. The method handlers, exit codes / `CliExitCodeError`, interactive OAuth flow and output styling it shares with the one-shot CLI live in [`core/cli/`](../../core/cli), consumed through the ordinary `@inspector/core` alias and bundled like the rest of `core/` (#2461). + +## Install + +`mcpdo` ships with the published package: + +```bash +npm install -g @modelcontextprotocol/inspector +mcpdo --help +``` + +## Install / run (from this repo) + +Build, then put `mcpdo` on your PATH with `npm link` (points at this package’s `build/mcp-bin.js`): + +```bash +# from the repo root — install deps once if needed +npm install + +cd clients/mcpdo +npm run build +npm link + +mcpdo --help +``` + +Rebuild after pulling source changes (`npm run build` in `clients/mcpdo`). You usually do **not** need to re-link unless the package `bin` entry changes. + +### Development loop + +`mcpdo` itself is a short-lived process re-executed on every invocation, so a +plain rebuild is enough for its changes to take effect on the next command. +The **connection daemon** (`build/mcpdod.js`) is different: `ensureDaemon` (see +`src/daemon/ensure.ts`) reuses an already-running daemon without checking its +code version, so a daemon started before your rebuild keeps running stale +code indefinitely. It sets `process.title = "mcpdod"`, so a stray one is +visible as `mcpdod` in `ps`/`pgrep` and killable with `pkill mcpdod`. + +Use `npm run build:dev` instead of `npm run build` while iterating: it runs +`mcpdo daemon stop` first (harmless/no-op if no daemon is running — it treats +"daemon not running" as success) and then `tsup`, so the next daemon-backed +command (`connect`, `tools/list`, …) spawns a fresh daemon from the code you +just built. Commands that never touch the daemon (`servers/list`, +`servers/show`, `--help`) don't need this — a plain `npm run build` is enough +for those. + +Without linking, run the built file directly: + +```bash +node clients/mcpdo/build/mcp-bin.js --help +``` + +Remove the link when you’re done: + +```bash +npm unlink -g @modelcontextprotocol/mcpdo +``` + +## Usage + +```bash +mcpdo servers/list --config path/to/mcp.json +mcpdo servers/show test-stdio --config path/to/mcp.json +mcpdo connect test-stdio --config path/to/mcp.json +mcpdo connect my-http --config path/to/mcp.json --relogin # (-r) ignore stored OAuth; login only if auth required +mcpdo auth/list +mcpdo auth/clear https://example.com/mcp +mcpdo auth/clear hosted-everything # or a catalog/connection name (from auth/list "known as") +mcpdo auth/clear --all --yes +mcpdo tools/list +mcpdo tools/call echo message:=hi +mcpdo tools/call echo '{"message":"hi"}' +mcpdo @test-stdio resources/list +mcpdo logging/tail # long-lived; Ctrl-C to stop +mcpdo connections/list +mcpdo disconnect --connection test-stdio +mcpdo disconnect my-http --clear-auth # (-c) also clear stored OAuth so the next connect re-triggers sign-in +mcpdo daemon status +mcpdo daemon stop + +# Optional: private daemon for this shell only +eval "$(mcpdo private)" +mcpdo connect test-stdio --config path/to/mcp.json +mcpdo tools/list +``` + +**Private mode:** `eval "$(mcpdo private)"` gives the shell its own daemon and +bearer token, separating its connections and daemon state from other mcpdo +daemons. It is not a security boundary against other processes running as your +user: anything with the same UID that learns the daemon directory can read the +token. For a hard boundary, use OS-level isolation (separate user, container). + +**Globals (before subcommand):** `--format text|json`, `--plain`, `--connection <name>` (shorthand: `--conn`), `--catalog` / `--config`, `--stored-auth-only`. + +**Output:** `--format text` (default) is human-readable (TTY ANSI unless `--plain` / `NO_COLOR`). `--format json` is pretty-printed payload with **no** `{ result }` envelope. + +**Auth:** shared `oauth.json` with other Inspector clients. Connect-time OAuth only on this CLI; mid-connection step-up remains on one-shot `mcp-inspector --cli`. `auth/list` annotates each stored URL with the catalog/connection names it is "known as" and marks `● live` when a connection currently holds it, so the raw store URLs correlate to the servers you actually use. EMA IdP login records (keyed `ema-idp:<issuer>` in the store) lead with their bare issuer URL and carry a trailing `enterprise IdP login` marker (parallel to `known as:`) rather than showing as unnamed servers. `auth/clear` accepts either a store URL **or** one of those friendly names (a name that maps to no URL — e.g. a stdio server — or to more than one URL is rejected with guidance); an EMA IdP login clears by its bare issuer URL (or its raw `ema-idp:` key). `--relogin` (`-r`) clears any URL-keyed store entry before connect (no-op for stdio). To force a clean logged-out state from an already-open connection, `disconnect <name> --clear-auth` (`-c`) tears the connection down **and** clears its stored tokens in one step, so the next plain `connect` re-triggers sign-in — no-op for stdio / servers with no stored entry. Non-TTY `connect` exits 0 with `pendingAuth: true` and an `authUrl` to relay; after the user signs in, any real command completes the connection, `connections/show` completes it too, and `connections/list` marks the entry `pendingAuthSignedIn` ("signed in — completing on next use") without dialing. + +See [`specification/v2_cli_v2.md`](../../specification/v2_cli_v2.md) for the as-built design and to-do list. + +## Isolating untrusted stdio servers + +The daemon's token controls **who can command the daemon**, not **what a +spawned server can do**: a stdio MCP server runs with your full user +privileges, like in any MCP host. To isolate a server you don't fully trust, +wrap the stdio command in a container — this works today with no mcpdo +support: + +```bash +mcpdo connect -- docker run -i --rm --network none -v "$PWD:/work:ro" <server-image> +``` + +Tighten or loosen the flags per server (drop `--network none` if it needs +egress; adjust the mount to what it should see). HTTP/SSE targets run no +local code, so they need no process isolation. + +## Protocol era support + +mcpdo shares `core`'s `InspectorClient`, so it negotiates whichever era +(`legacy` 2025-03-26-style vs. `modern`/2026-era, e.g. task-augmented calls, +`server/discover`) the target actually speaks — no extra flags needed for +that to work. Two things are mcpdo-specific: + +- **`--era <era>` on `connect`**: `legacy` (default), `auto` (probe via + `server/discover` before connecting), or `modern`. Overrides whatever a + catalog/config entry's `protocolEra` says, and is the only way to set it + for an ad-hoc target (no config entry to read one from). + + ```bash + mcpdo connect my-modern-server --config path/to/mcp.json --era modern + mcpdo connect https://example.com/mcp --era auto + ``` + +- **Era visibility in connection output**: `connections/list`, `connections/use`, and + `connect` all show the negotiated era inline (`@name (MRU) — server +[modern]`). `connections/show <name>` gives the full picture — era, negotiated + protocol version, server info, capabilities, and (when the connect probed + `server/discover`) the server's supported-versions list: + + ``` + $ mcpdo connections/show my-modern-server + Connection: my-modern-server + Server: https://example.com/mcp + Era: modern (2026-06-18) + Supported versions: 2025-03-26, 2026-06-18 + ... + ``` + +A modern (SEP-2663) task that reaches `status: "input_required"` is **not** +resumed with `tasks/update` from mcpdo. Core polls the task to a terminal +state, so the input round surfaces as an **elicitation** instead — prompted +inline on an interactive TTY, or **parked** (the RPC returns an +`elicitationPending` payload) otherwise. Answer it with `elicitation/respond`, +exactly like any other parked elicitation (see [Elicitation +support](#elicitation-support) below): + +```bash +mcpdo elicitation/respond <elicitationId> approved:=true +``` + +The `elicitationPending` payload carries no `taskId`, and `tasks/list` is +refused while a call is parked, so there is no task id to pass to +`tasks/update` — `elicitation/respond` is the only path that works. + +## Elicitation support + +mcpdo can prompt interactively for every elicitation delivery mechanism — +legacy server→client `elicitation/create` requests, modern non-task MRTR +(multi-round tool response) rounds, and modern SEP-2663 **task** input rounds +(a task that reaches `status: "input_required"`) — and both modes a server may +ask for: + +- **URL mode**: mcpdo prints the URL and waits for you to confirm you've + finished out-of-band (there's no "decline", only accept-that-you-finished + or cancel — the actual completion can't be observed locally). +- **Form mode**: mcpdo renders one prompt per field from the schema, with a + review step (edit any field again, or submit) before answering. + +Interactive callers (`--format text` on a TTY) get these prompts inline. +Non-interactive callers — `--format json`, or no TTY at all — don't get a +prompt: the daemon **parks** the elicitation and the RPC returns an +`elicitationPending` payload naming the pending id. Answer it (from any +shell) with `elicitation/respond <id>` — form answers as `key:=value` pairs +or JSON, `--done` for URL mode, or `--decline` / `--cancel` — after which the +original call completes. An unanswered parked elicitation is auto-cancelled +after 10 minutes. + +> **Decision — who answers a prompt.** Non-interactive callers never get an +> automatic decline: the elicitation is parked so whoever drives mcpdo (a +> script, an agent relaying to a human) can answer deliberately via +> `elicitation/respond`, on its own schedule. That is deliberate for an +> inspector tool. URL-mode is different: there is never an auto-accept — +> completion is only ever confirmed by an explicit answer, because the +> out-of-band action (typically an auth or consent step in a browser) is the +> user's to perform. Use `--elicit off` on `connect` to keep any elicitation +> from being asked at all. + +By default mcpdo advertises **both** modes to the server (`elicit: {url, +form}`), matching pre-#1783 behavior. Override this per connection with +`--elicit <mode>` on `connect`: + +- `off` — advertise no elicitation capability at all. Useful when whatever is + driving mcpdo (a script, an agent) can't handle an interactive prompt itself + — omitting the capability lets a well-behaved server fall back to its own + alternative (e.g. proceeding with defaults) instead of the request sitting + parked until someone answers it. +- `url` — URL mode only. +- `form` — form mode only. +- `both` — the default; both modes. + +Like `--era`, this overrides whatever a catalog/config entry's +`elicitCapability` says, and is the only way to set it for an ad-hoc target +(no config entry to read one from): + +```bash +mcpdo connect my-server --config path/to/mcp.json --elicit off +mcpdo connect https://example.com/mcp --elicit url +``` + +## Relation to one-shot CLI + +| | One-shot | Connection (`mcpdo`) | +| ------------- | ------------------------------------- | ------------------------------- | +| Entrypoint | `mcp-inspector --cli` | `mcpdo` | +| Package (dev) | `clients/cli` | `clients/mcpdo` | +| Lifecycle | Connect → one `--method` → disconnect | Connect once → many subcommands | + +One-shot docs: [`clients/cli/README.md`](../cli/README.md). diff --git a/clients/mcpdo/__tests__/agent-help.test.ts b/clients/mcpdo/__tests__/agent-help.test.ts new file mode 100644 index 0000000000..86aaf605ab --- /dev/null +++ b/clients/mcpdo/__tests__/agent-help.test.ts @@ -0,0 +1,47 @@ +import { describe, it, expect } from "vitest"; +import { existsSync } from "node:fs"; +import { runMcp } from "./helpers/mcp-runner.js"; + +describe("mcpdo agent-help", () => { + it("prints the SKILL.md body with the frontmatter stripped", async () => { + const result = await runMcp(["agent-help"]); + expect(result.exitCode).toBe(0); + expect(result.stdout).toContain("mcpdo connect"); + expect(result.stdout).not.toContain("name: mcpdo"); + expect(result.stdout.startsWith("---")).toBe(false); + }); + + it("--skill explicitly selects the default guide output", async () => { + const [bare, explicit] = [ + await runMcp(["agent-help"]), + await runMcp(["agent-help", "--skill"]), + ]; + expect(explicit.exitCode).toBe(0); + expect(explicit.stdout).toBe(bare.stdout); + }); + + it("--skill-path prints the resolved SKILL.md file path", async () => { + const result = await runMcp(["agent-help", "--skill-path"]); + expect(result.exitCode).toBe(0); + const printedPath = result.stdout.trim(); + expect(printedPath.endsWith("skills/mcpdo/SKILL.md")).toBe(true); + expect(existsSync(printedPath)).toBe(true); + }); + + it("--instructions prints the always-on CLAUDE.md/AGENTS.md snippet", async () => { + const result = await runMcp(["agent-help", "--instructions"]); + expect(result.exitCode).toBe(0); + expect(result.stdout).toContain("part of your available toolset"); + expect(result.stdout).toContain("mcpdo agent-help"); + }); + + it("rejects --instructions combined with --skill-path", async () => { + const result = await runMcp([ + "agent-help", + "--instructions", + "--skill-path", + ]); + expect(result.exitCode).not.toBe(0); + expect(result.stderr).toContain("mutually exclusive"); + }); +}); diff --git a/clients/mcpdo/__tests__/auth-helper.test.ts b/clients/mcpdo/__tests__/auth-helper.test.ts new file mode 100644 index 0000000000..e4368c5d0b --- /dev/null +++ b/clients/mcpdo/__tests__/auth-helper.test.ts @@ -0,0 +1,592 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { PassThrough } from "node:stream"; +import type { CallbackNavigation } from "@inspector/core/auth/index.js"; + +const authorizeInFrontend = vi.fn(); + +vi.mock("../src/connection/authorize.js", () => ({ + authorizeInFrontend: (...args: unknown[]) => authorizeInFrontend(...args), +})); + +import { + AUTH_HELPER_COMMAND, + obtainPendingAuthUrl, + pendingAuthMarkerPath, + readLivePendingAuthMarker, + removeOwnPendingAuthMarker, + runAuthHelper, + type PendingAuthMarker, +} from "../src/connection/auth-helper.js"; + +const SERVER_URL = "https://mcp.example.com/mcp"; + +describe("auth-helper", () => { + let dir: string; + let prevDaemonDir: string | undefined; + + beforeEach(() => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-auth-helper-")); + prevDaemonDir = process.env.MCP_INSPECTOR_DAEMON_DIR; + process.env.MCP_INSPECTOR_DAEMON_DIR = dir; + authorizeInFrontend.mockReset(); + }); + + afterEach(() => { + if (prevDaemonDir === undefined) + delete process.env.MCP_INSPECTOR_DAEMON_DIR; + else process.env.MCP_INSPECTOR_DAEMON_DIR = prevDaemonDir; + fs.rmSync(dir, { recursive: true, force: true }); + }); + + function writeMarker(marker: PendingAuthMarker): string { + const markerPath = pendingAuthMarkerPath(SERVER_URL); + fs.writeFileSync(markerPath, JSON.stringify(marker), { mode: 0o600 }); + return markerPath; + } + + describe("readLivePendingAuthMarker", () => { + it("returns undefined when no marker exists", () => { + expect(readLivePendingAuthMarker(SERVER_URL)).toBeUndefined(); + }); + + it("returns a live marker (unexpired, helper pid running)", () => { + writeMarker({ + url: "https://as.example/authorize?state=s1", + pid: process.pid, + expiresAt: Date.now() + 60_000, + }); + expect(readLivePendingAuthMarker(SERVER_URL)).toMatchObject({ + url: "https://as.example/authorize?state=s1", + }); + }); + + it("ignores an expired marker without deleting it (writers replace it)", () => { + const markerPath = writeMarker({ + url: "https://as.example/authorize", + pid: process.pid, + expiresAt: Date.now() - 1, + }); + expect(readLivePendingAuthMarker(SERVER_URL)).toBeUndefined(); + // Read-time deletion would race a just-spawned helper's fresh marker. + expect(fs.existsSync(markerPath)).toBe(true); + }); + + it("cleanup removes only this process's own marker", () => { + const markerPath = writeMarker({ + url: "https://as.example/authorize", + pid: process.pid, + expiresAt: Date.now() + 60_000, + }); + removeOwnPendingAuthMarker(markerPath); + expect(fs.existsSync(markerPath)).toBe(false); + // A replacement flow's marker (different pid) must survive the old + // helper's exit cleanup. + writeMarker({ + url: "https://as.example/authorize-2", + pid: process.pid + 1, + expiresAt: Date.now() + 60_000, + }); + removeOwnPendingAuthMarker(markerPath); + expect(fs.existsSync(markerPath)).toBe(true); + // Missing or malformed markers are a no-op, not an error. + fs.writeFileSync(markerPath, "not json\n"); + expect(() => removeOwnPendingAuthMarker(markerPath)).not.toThrow(); + fs.rmSync(markerPath, { force: true }); + expect(() => removeOwnPendingAuthMarker(markerPath)).not.toThrow(); + }); + + it("ignores a marker whose helper process is gone", () => { + writeMarker({ + url: "https://as.example/authorize", + // Out-of-range / nonexistent pid: process.kill(pid, 0) throws. + pid: 0x7fffffff, + expiresAt: Date.now() + 60_000, + }); + expect(readLivePendingAuthMarker(SERVER_URL)).toBeUndefined(); + }); + + it("ignores malformed marker files", () => { + fs.writeFileSync(pendingAuthMarkerPath(SERVER_URL), "not-json"); + expect(readLivePendingAuthMarker(SERVER_URL)).toBeUndefined(); + fs.writeFileSync(pendingAuthMarkerPath(SERVER_URL), '{"url":42}'); + expect(readLivePendingAuthMarker(SERVER_URL)).toBeUndefined(); + }); + }); + + describe("obtainPendingAuthUrl", () => { + function writeHelperScript(body: string): string { + const script = path.join(dir, "fake-helper.mjs"); + fs.writeFileSync(script, body); + return script; + } + + it("reuses a live marker's URL without spawning a second helper", async () => { + writeMarker({ + url: "https://as.example/authorize?state=reuse", + pid: process.pid, + expiresAt: Date.now() + 60_000, + }); + const url = await obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + // Would fail loudly if a spawn were attempted. + { helperArgv1: path.join(dir, "does-not-exist.mjs") }, + ); + expect(url).toBe("https://as.example/authorize?state=reuse"); + }); + + it("spawns the helper, passes params over stdin, and returns the reported URL", async () => { + const script = writeHelperScript(` + let body = ""; + process.stdin.on("data", (c) => (body += c)); + process.stdin.on("end", () => { + const params = JSON.parse(body); + if (process.argv[2] !== ${JSON.stringify(AUTH_HELPER_COMMAND)}) { + process.exit(9); + } + process.stdout.write( + JSON.stringify({ + event: "auth_url", + url: "https://as.example/authorize?server=" + + encodeURIComponent(params.serverConfig.url), + }) + "\\n", + ); + }); + `); + const url = await obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { helperArgv1: script }, + ); + expect(url).toBe( + `https://as.example/authorize?server=${encodeURIComponent(SERVER_URL)}`, + ); + }); + + it("maps a helper error event to auth_required", async () => { + const script = writeHelperScript(` + process.stdin.resume(); + process.stdin.on("end", () => { + process.stdout.write( + JSON.stringify({ event: "error", message: "no AS metadata" }) + "\\n", + ); + }); + `); + await expect( + obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { helperArgv1: script }, + ), + ).rejects.toMatchObject({ + envelope: { code: "auth_required" }, + message: expect.stringContaining("no AS metadata"), + }); + }); + + it("skips marker reuse for stdio configs and still spawns the helper", async () => { + const script = writeHelperScript(` + process.stdin.resume(); + process.stdin.on("end", () => { + process.stdout.write( + JSON.stringify({ event: "auth_url", url: "https://as.example/stdio" }) + "\\n", + ); + }); + `); + const url = await obtainPendingAuthUrl( + { type: "stdio", command: "srv" }, + undefined, + { helperArgv1: script }, + ); + expect(url).toBe("https://as.example/stdio"); + }); + + it("skips blank, malformed, and unknown-event lines before the URL", async () => { + const script = writeHelperScript(` + process.stdin.resume(); + process.stdin.on("end", () => { + process.stdout.write( + "\\n" + + "not json\\n" + + JSON.stringify({ event: "progress" }) + "\\n" + + JSON.stringify({ event: "auth_url", url: "https://as.example/after-noise" }) + "\\n", + ); + }); + `); + const url = await obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { helperArgv1: script }, + ); + expect(url).toBe("https://as.example/after-noise"); + }); + + it("waits for a concurrent reserver's marker instead of spawning", async () => { + const lockPath = `${pendingAuthMarkerPath(SERVER_URL)}.lock`; + fs.writeFileSync(lockPath, "1234\n", { mode: 0o600 }); + // Publish the marker shortly after, as the winner's helper would. + const t = setTimeout(() => { + writeMarker({ + url: "https://as.example/authorize?state=winner", + pid: process.pid, + expiresAt: Date.now() + 60_000, + }); + }, 50); + try { + const url = await obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + // Would fail loudly if a spawn were attempted. + { + helperArgv1: path.join(dir, "does-not-exist.mjs"), + waitMs: 2_000, + pollMs: 10, + }, + ); + expect(url).toBe("https://as.example/authorize?state=winner"); + } finally { + clearTimeout(t); + } + }); + + it("times out with auth_required when the reserved flow never publishes", async () => { + const lockPath = `${pendingAuthMarkerPath(SERVER_URL)}.lock`; + fs.writeFileSync(lockPath, "1234\n", { mode: 0o600 }); + await expect( + obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { + helperArgv1: path.join(dir, "does-not-exist.mjs"), + waitMs: 60, + pollMs: 10, + }, + ), + ).rejects.toMatchObject({ + envelope: { code: "auth_required" }, + message: expect.stringContaining("in-progress sign-in"), + }); + }); + + it("steals a stale reservation and releases its own after the flow", async () => { + const lockPath = `${pendingAuthMarkerPath(SERVER_URL)}.lock`; + fs.writeFileSync(lockPath, "1234\n", { mode: 0o600 }); + // Backdate past the steal threshold (wait window + slack). + const old = new Date(Date.now() - 120_000); + fs.utimesSync(lockPath, old, old); + const script = writeHelperScript(` + process.stdin.resume(); + process.stdin.on("end", () => { + process.stdout.write( + JSON.stringify({ event: "auth_url", url: "https://as.example/stolen" }) + "\\n", + ); + }); + `); + const url = await obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { helperArgv1: script }, + ); + expect(url).toBe("https://as.example/stolen"); + // The reservation is released once the URL is obtained. + expect(fs.existsSync(lockPath)).toBe(false); + }); + + it("backs off when another stealer claims the stale lock first", async () => { + const lockPath = `${pendingAuthMarkerPath(SERVER_URL)}.lock`; + fs.writeFileSync(lockPath, "1234\n", { mode: 0o600 }); + const old = new Date(Date.now() - 120_000); + fs.utimesSync(lockPath, old, old); + // Occupying this process's claim path makes the rename throw — the + // same observable outcome as losing the claim race: reserve fails, + // the caller waits, and (with no marker forthcoming) times out. + fs.mkdirSync(`${lockPath}.claim-${process.pid}`); + await expect( + obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { + helperArgv1: path.join(dir, "does-not-exist.mjs"), + waitMs: 60, + pollMs: 10, + }, + ), + ).rejects.toMatchObject({ envelope: { code: "auth_required" } }); + // The stale lock was claimed away by the rename attempt or left in + // place — either way this process never spawned a helper. + }); + + it("maps a helper spawn failure to auth_required", async () => { + const originalExecPath = process.execPath; + // A nonexistent interpreter makes spawn emit `error` instead of `exit`. + process.execPath = path.join(dir, "no-such-node"); + try { + await expect( + obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { helperArgv1: path.join(dir, "unused.mjs") }, + ), + ).rejects.toMatchObject({ + envelope: { code: "auth_required" }, + message: expect.stringContaining("Failed to spawn"), + }); + } finally { + process.execPath = originalExecPath; + } + }); + + it("fails when the helper exits before producing a URL", async () => { + const script = writeHelperScript(`process.exit(2);`); + await expect( + obtainPendingAuthUrl( + { type: "streamable-http", url: SERVER_URL }, + undefined, + { helperArgv1: script }, + ), + ).rejects.toMatchObject({ + envelope: { code: "auth_required" }, + message: expect.stringContaining("exited"), + }); + }); + }); + + describe("runAuthHelper", () => { + function stubStdin(body: string): () => void { + const stream = new PassThrough(); + const descriptor = Object.getOwnPropertyDescriptor(process, "stdin"); + Object.defineProperty(process, "stdin", { + value: stream, + configurable: true, + }); + stream.end(body); + return () => { + if (descriptor) Object.defineProperty(process, "stdin", descriptor); + }; + } + + function captureStdout(): { lines: () => string[]; restore: () => void } { + let out = ""; + const original = process.stdout.write; + process.stdout.write = ((chunk: unknown) => { + out += typeof chunk === "string" ? chunk : String(chunk); + return true; + }) as typeof process.stdout.write; + return { + lines: () => + out + .split("\n") + .filter((l) => l.trim()) + .map((l) => l), + restore: () => { + process.stdout.write = original; + }, + }; + } + + it("writes the marker while the flow runs, emits auth_url and done, and removes the marker on exit", async () => { + let markerDuringFlow: PendingAuthMarker | undefined; + authorizeInFrontend.mockImplementation( + async ( + _config: unknown, + _settings: unknown, + options: { + makeNavigation: (control: { armed: boolean }) => CallbackNavigation; + }, + ) => { + const navigation = options.makeNavigation({ armed: true }); + navigation.navigateToAuthorization( + new URL("https://as.example/authorize?state=s2"), + ); + markerDuringFlow = readLivePendingAuthMarker(SERVER_URL); + }, + ); + const restoreStdin = stubStdin( + JSON.stringify({ + serverConfig: { type: "streamable-http", url: SERVER_URL }, + }), + ); + const stdout = captureStdout(); + try { + await runAuthHelper(); + } finally { + stdout.restore(); + restoreStdin(); + } + expect(markerDuringFlow).toMatchObject({ + url: "https://as.example/authorize?state=s2", + pid: process.pid, + }); + const events = stdout + .lines() + .map((l) => JSON.parse(l) as { event: string }); + expect(events.map((e) => e.event)).toEqual(["auth_url", "done"]); + // Marker removed once the flow completed. + expect(fs.existsSync(pendingAuthMarkerPath(SERVER_URL))).toBe(false); + }); + + it("stays silent while the navigation is disarmed (SDK auth during plain connect)", async () => { + authorizeInFrontend.mockImplementation( + async ( + _config: unknown, + _settings: unknown, + options: { + makeNavigation: (control: { armed: boolean }) => CallbackNavigation; + }, + ) => { + const navigation = options.makeNavigation({ armed: false }); + navigation.navigateToAuthorization( + new URL("https://as.example/authorize?leaked=1"), + ); + }, + ); + const restoreStdin = stubStdin( + JSON.stringify({ + serverConfig: { type: "streamable-http", url: SERVER_URL }, + }), + ); + const stdout = captureStdout(); + try { + await runAuthHelper(); + } finally { + stdout.restore(); + restoreStdin(); + } + const events = stdout + .lines() + .map((l) => JSON.parse(l) as { event: string }); + expect(events.map((e) => e.event)).toEqual(["done"]); + expect(fs.existsSync(pendingAuthMarkerPath(SERVER_URL))).toBe(false); + }); + + it("emits an error event and rethrows when the flow fails", async () => { + authorizeInFrontend.mockRejectedValueOnce(new Error("flow exploded")); + const restoreStdin = stubStdin( + JSON.stringify({ + serverConfig: { type: "streamable-http", url: SERVER_URL }, + }), + ); + const stdout = captureStdout(); + try { + await expect(runAuthHelper()).rejects.toThrow("flow exploded"); + } finally { + stdout.restore(); + restoreStdin(); + } + const events = stdout + .lines() + .map((l) => JSON.parse(l) as { event: string; message?: string }); + expect(events).toEqual([{ event: "error", message: "flow exploded" }]); + }); + + it("emits the URL without writing a marker for stdio configs", async () => { + authorizeInFrontend.mockImplementation( + async ( + _config: unknown, + _settings: unknown, + options: { + makeNavigation: (control: { armed: boolean }) => CallbackNavigation; + }, + ) => { + const navigation = options.makeNavigation({ armed: true }); + navigation.navigateToAuthorization( + new URL("https://as.example/authorize?stdio=1"), + ); + }, + ); + const restoreStdin = stubStdin( + JSON.stringify({ serverConfig: { type: "stdio", command: "srv" } }), + ); + const stdout = captureStdout(); + try { + await runAuthHelper(); + // The EPIPE guard on stdout must swallow late write errors. + process.stdout.emit("error", new Error("EPIPE")); + } finally { + stdout.restore(); + restoreStdin(); + } + const events = stdout + .lines() + .map((l) => JSON.parse(l) as { event: string }); + expect(events.map((e) => e.event)).toEqual(["auth_url", "done"]); + // No url on the config → no marker file anywhere in the daemon dir. + expect(fs.readdirSync(dir).filter((f) => f.includes("auth"))).toEqual([]); + }); + + it("stringifies a non-Error flow failure in the error event", async () => { + authorizeInFrontend.mockRejectedValueOnce("string boom"); + const restoreStdin = stubStdin( + JSON.stringify({ + serverConfig: { type: "streamable-http", url: SERVER_URL }, + }), + ); + const stdout = captureStdout(); + try { + await expect(runAuthHelper()).rejects.toBe("string boom"); + } finally { + stdout.restore(); + restoreStdin(); + } + const events = stdout + .lines() + .map((l) => JSON.parse(l) as { event: string; message?: string }); + expect(events).toEqual([{ event: "error", message: "string boom" }]); + }); + + it("rejects when stdin errors before EOF", async () => { + const stream = new PassThrough(); + const descriptor = Object.getOwnPropertyDescriptor(process, "stdin"); + Object.defineProperty(process, "stdin", { + value: stream, + configurable: true, + }); + const stdout = captureStdout(); + try { + const pending = runAuthHelper(); + stream.emit("error", new Error("broken pipe")); + await expect(pending).rejects.toThrow("broken pipe"); + } finally { + stdout.restore(); + if (descriptor) Object.defineProperty(process, "stdin", descriptor); + } + }); + + it("times out when the parent never sends params", async () => { + vi.useFakeTimers(); + const stream = new PassThrough(); + const descriptor = Object.getOwnPropertyDescriptor(process, "stdin"); + Object.defineProperty(process, "stdin", { + value: stream, + configurable: true, + }); + const stdout = captureStdout(); + try { + const pending = runAuthHelper(); + const expectation = expect(pending).rejects.toThrow( + /timed out waiting for params/, + ); + await vi.advanceTimersByTimeAsync(30_000); + await expectation; + } finally { + vi.useRealTimers(); + stdout.restore(); + if (descriptor) Object.defineProperty(process, "stdin", descriptor); + } + }); + + it("rejects params without a serverConfig", async () => { + const restoreStdin = stubStdin(JSON.stringify({})); + const stdout = captureStdout(); + try { + await expect(runAuthHelper()).rejects.toThrow(/serverConfig/); + } finally { + stdout.restore(); + restoreStdin(); + } + }); + }); +}); diff --git a/clients/mcpdo/__tests__/auth-names-commands.test.ts b/clients/mcpdo/__tests__/auth-names-commands.test.ts new file mode 100644 index 0000000000..50df39a47d --- /dev/null +++ b/clients/mcpdo/__tests__/auth-names-commands.test.ts @@ -0,0 +1,238 @@ +import { describe, it, expect, vi, beforeEach } from "vitest"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; + +const callDaemon = vi.fn(); +const listServerEntries = vi.fn(); +const listStoredAuth = vi.fn(); +const clearStoredAuth = vi.fn(); + +// Stub only the three reaches of the auth commands; importOriginal keeps the +// rest (including the real normalizeServerUrl / buildAuthNameIndex) intact, so +// no real daemon socket, catalog read, or OAuth-store write happens in-process. +vi.mock("../src/daemon/index.js", async (importOriginal) => ({ + ...(await importOriginal<typeof import("../src/daemon/index.js")>()), + callDaemon: (...args: unknown[]) => callDaemon(...args), +})); + +vi.mock( + "@inspector/core/cli/handlers/servers-list.js", + async (importOriginal) => ({ + ...(await importOriginal< + typeof import("@inspector/core/cli/handlers/servers-list.js") + >()), + listServerEntries: (...args: unknown[]) => listServerEntries(...args), + }), +); + +vi.mock("../src/connection/stored-auth.js", async (importOriginal) => ({ + ...(await importOriginal< + typeof import("../src/connection/stored-auth.js") + >()), + listStoredAuth: (...args: unknown[]) => listStoredAuth(...args), + clearStoredAuth: (...args: unknown[]) => clearStoredAuth(...args), +})); + +const { runMcp } = await import("./helpers/mcp-runner.js"); + +const URL_A = "https://api.example.com/mcp"; +const URL_B = "https://other.example.com/mcp"; + +function daemonDown(): CliExitCodeError { + return new CliExitCodeError(EXIT_CODES.UNREACHABLE, "no daemon", { + code: "daemon_unreachable", + }); +} + +beforeEach(() => { + callDaemon.mockReset(); + listServerEntries.mockReset(); + listStoredAuth.mockReset(); + clearStoredAuth.mockReset(); + listServerEntries.mockResolvedValue([]); + callDaemon.mockResolvedValue({ connections: [] }); + listStoredAuth.mockResolvedValue({ + oauthStatePath: "/state/oauth.json", + servers: [], + }); + clearStoredAuth.mockImplementation((url: string) => Promise.resolve({ url })); +}); + +describe("auth/list friendly-name annotation", () => { + it("annotates a stored URL with its catalog name and live flag", async () => { + listStoredAuth.mockResolvedValue({ + oauthStatePath: "/state/oauth.json", + servers: [{ url: URL_A, hasTokens: true, hasRefreshToken: false }], + }); + listServerEntries.mockResolvedValue([ + { name: "hosted", type: "streamable-http", detail: URL_A }, + ]); + callDaemon.mockResolvedValue({ + connections: [ + { + name: "hosted", + serverIdentity: URL_A, + connectedAt: 0, + lastAccessedAt: 0, + isMru: false, + }, + ], + }); + + const result = await runMcp(["auth/list", "--format", "json"]); + expect(result.exitCode).toBe(0); + const parsed = JSON.parse(result.stdout.trim()); + expect(parsed.servers[0]).toMatchObject({ + url: URL_A, + knownAs: ["hosted"], + live: true, + }); + }); + + it("omits knownAs/live for a URL with no local name", async () => { + listStoredAuth.mockResolvedValue({ + oauthStatePath: "/state/oauth.json", + servers: [{ url: URL_A, hasTokens: false, hasRefreshToken: false }], + }); + + const result = await runMcp(["auth/list", "--format", "json"]); + const parsed = JSON.parse(result.stdout.trim()); + expect(parsed.servers[0].knownAs).toBeUndefined(); + expect(parsed.servers[0].live).toBeUndefined(); + }); + + it("labels an EMA IdP record and omits name annotation for it", async () => { + listStoredAuth.mockResolvedValue({ + oauthStatePath: "/state/oauth.json", + servers: [ + { + url: "ema-idp:https://idp.example.com", + hasTokens: false, + hasRefreshToken: false, + }, + ], + }); + + const result = await runMcp(["auth/list", "--format", "json"]); + const parsed = JSON.parse(result.stdout.trim()); + expect(parsed.servers[0]).toMatchObject({ + url: "ema-idp:https://idp.example.com", + idp: true, + issuer: "https://idp.example.com", + }); + expect(parsed.servers[0].knownAs).toBeUndefined(); + }); + + it("falls back to catalog names only when the daemon is down", async () => { + listStoredAuth.mockResolvedValue({ + oauthStatePath: "/state/oauth.json", + servers: [{ url: URL_A, hasTokens: true, hasRefreshToken: true }], + }); + listServerEntries.mockResolvedValue([ + { name: "hosted", type: "streamable-http", detail: URL_A }, + ]); + callDaemon.mockRejectedValue(daemonDown()); + + const result = await runMcp(["auth/list", "--format", "json"]); + const parsed = JSON.parse(result.stdout.trim()); + expect(parsed.servers[0].knownAs).toEqual(["hosted"]); + expect(parsed.servers[0].live).toBeUndefined(); + }); +}); + +describe("auth/clear by friendly name", () => { + it("resolves a catalog name to its URL and clears it", async () => { + listServerEntries.mockResolvedValue([ + { name: "hosted", type: "streamable-http", detail: URL_A }, + ]); + + const result = await runMcp(["auth/clear", "hosted", "--format", "json"]); + expect(clearStoredAuth).toHaveBeenCalledWith(URL_A); + expect(JSON.parse(result.stdout.trim())).toEqual({ + url: URL_A, + clearedByName: "hosted", + }); + }); + + it("keeps the URL argument path unchanged", async () => { + const result = await runMcp(["auth/clear", URL_A, "--format", "json"]); + expect(clearStoredAuth).toHaveBeenCalledWith(URL_A); + expect(JSON.parse(result.stdout.trim())).toEqual({ url: URL_A }); + }); + + it("clears a non-http store key (EMA IdP) by its exact key", async () => { + // The `ema-idp:<issuer>` key auth/list shows verbatim does not start with + // https://, so it must still route to the direct store-key path rather than + // the friendly-name resolver (which would reject it as an unknown name). + const idpKey = "ema-idp:https://idp.example.com"; + listStoredAuth.mockResolvedValue({ + oauthStatePath: "/state/oauth.json", + servers: [{ url: idpKey, hasTokens: false, hasRefreshToken: false }], + }); + + const result = await runMcp(["auth/clear", idpKey, "--format", "json"]); + expect(clearStoredAuth).toHaveBeenCalledWith(idpKey); + expect(JSON.parse(result.stdout.trim())).toEqual({ url: idpKey }); + }); + + it("clears an EMA IdP entry by its bare issuer URL", async () => { + // auth/list renders the IdP login by its issuer URL (prefix hidden), so + // auth/clear must map that issuer back to the prefixed store key. + const idpKey = "ema-idp:https://idp.example.com"; + listStoredAuth.mockResolvedValue({ + oauthStatePath: "/state/oauth.json", + servers: [{ url: idpKey, hasTokens: false, hasRefreshToken: false }], + }); + + const result = await runMcp([ + "auth/clear", + "https://idp.example.com", + "--format", + "json", + ]); + expect(clearStoredAuth).toHaveBeenCalledWith(idpKey); + expect(JSON.parse(result.stdout.trim())).toEqual({ url: idpKey }); + }); + + it("errors for a known stdio name (no stored auth)", async () => { + listServerEntries.mockResolvedValue([ + { name: "local", type: "stdio", detail: "node server.js" }, + ]); + + const result = await runMcp(["auth/clear", "local"]); + expect(result.exitCode).not.toBe(0); + expect(result.stderr).toContain("stdio server"); + expect(clearStoredAuth).not.toHaveBeenCalled(); + }); + + it("errors when a name maps to two different URLs", async () => { + listServerEntries.mockResolvedValue([ + { name: "shared", type: "streamable-http", detail: URL_A }, + ]); + callDaemon.mockResolvedValue({ + connections: [ + { + name: "shared", + serverIdentity: URL_B, + connectedAt: 0, + lastAccessedAt: 0, + isMru: false, + }, + ], + }); + + const result = await runMcp(["auth/clear", "shared"]); + expect(result.exitCode).not.toBe(0); + expect(result.stderr).toContain("multiple server URLs"); + expect(clearStoredAuth).not.toHaveBeenCalled(); + }); + + it("errors for an unknown name", async () => { + const result = await runMcp(["auth/clear", "nope"]); + expect(result.exitCode).not.toBe(0); + expect(result.stderr).toContain("No server named 'nope'"); + expect(clearStoredAuth).not.toHaveBeenCalled(); + }); +}); diff --git a/clients/mcpdo/__tests__/auth-names.test.ts b/clients/mcpdo/__tests__/auth-names.test.ts new file mode 100644 index 0000000000..514c558e2f --- /dev/null +++ b/clients/mcpdo/__tests__/auth-names.test.ts @@ -0,0 +1,146 @@ +import { describe, it, expect } from "vitest"; +import type { ServerListEntry } from "@inspector/core/cli/handlers/servers-list.js"; +import type { ConnectionInfo } from "../src/daemon/protocol.js"; +import { + buildAuthNameIndex, + resolveFriendlyName, +} from "../src/connection/auth-names.js"; + +function entry(name: string, type: string, detail: string): ServerListEntry { + return { name, type, detail }; +} + +function conn(name: string, serverIdentity: string): ConnectionInfo { + return { + name, + serverIdentity, + connectedAt: 0, + lastAccessedAt: 0, + isMru: false, + }; +} + +describe("buildAuthNameIndex", () => { + it("maps an http catalog entry both ways under the normalised URL", () => { + const index = buildAuthNameIndex( + [entry("hosted", "streamable-http", "https://api.example.com/mcp")], + [], + ); + // new URL(...).href normalises the key the store uses. + expect([...index.urlToNames.keys()]).toEqual([ + "https://api.example.com/mcp", + ]); + expect(index.urlToNames.get("https://api.example.com/mcp")).toEqual([ + { name: "hosted", source: "catalog", isLive: false }, + ]); + expect([...index.nameToUrls.get("hosted")!]).toEqual([ + "https://api.example.com/mcp", + ]); + expect(index.knownNames.has("hosted")).toBe(true); + }); + + it("records a stdio entry as a known name with no URL", () => { + const index = buildAuthNameIndex( + [entry("local", "stdio", "node server.js")], + [], + ); + expect(index.knownNames.has("local")).toBe(true); + expect(index.nameToUrls.has("local")).toBe(false); + expect(index.urlToNames.size).toBe(0); + }); + + it("marks a live connection and sorts live names first", () => { + const url = "https://api.example.com/mcp"; + const index = buildAuthNameIndex( + [entry("catalog-name", "streamable-http", url)], + [conn("live-name", url)], + ); + const refs = index.urlToNames.get(url)!; + expect(refs[0]).toEqual({ + name: "live-name", + source: "connection", + isLive: true, + }); + expect(refs[1]).toEqual({ + name: "catalog-name", + source: "catalog", + isLive: false, + }); + }); + + it("collapses a catalog+connection pair that shares a name and URL", () => { + const url = "https://api.example.com/mcp"; + const index = buildAuthNameIndex( + [entry("hosted", "streamable-http", url)], + [conn("hosted", url)], + ); + expect(index.urlToNames.get(url)).toEqual([ + { name: "hosted", source: "catalog", isLive: true }, + ]); + }); + + it("ignores a connection whose identity is not an http URL (stdio)", () => { + const index = buildAuthNameIndex([], [conn("local", "node server.js")]); + expect(index.knownNames.has("local")).toBe(true); + expect(index.nameToUrls.has("local")).toBe(false); + }); +}); + +describe("resolveFriendlyName", () => { + const url = "https://api.example.com/mcp"; + + it("resolves a name that maps to exactly one URL", () => { + const index = buildAuthNameIndex( + [entry("hosted", "streamable-http", url)], + [], + ); + expect(resolveFriendlyName("hosted", index)).toEqual({ kind: "url", url }); + }); + + it("treats two names for one URL as unambiguous (both resolve to it)", () => { + const index = buildAuthNameIndex( + [ + entry("hosted", "streamable-http", url), + entry("hosted-alias", "streamable-http", url), + ], + [], + ); + expect(resolveFriendlyName("hosted", index)).toEqual({ kind: "url", url }); + expect(resolveFriendlyName("hosted-alias", index)).toEqual({ + kind: "url", + url, + }); + }); + + it("flags one name that maps to two different URLs as ambiguous", () => { + const other = "https://other.example.com/mcp"; + const index = buildAuthNameIndex( + [entry("shared", "streamable-http", url)], + [conn("shared", other)], + ); + expect(resolveFriendlyName("shared", index)).toEqual({ + kind: "ambiguous", + name: "shared", + urls: [url, other].sort(), + }); + }); + + it("reports a known stdio name as no-url", () => { + const index = buildAuthNameIndex( + [entry("local", "stdio", "node server.js")], + [], + ); + expect(resolveFriendlyName("local", index)).toEqual({ + kind: "no-url", + name: "local", + }); + }); + + it("reports an unseen name as unknown", () => { + const index = buildAuthNameIndex([], []); + expect(resolveFriendlyName("nope", index)).toEqual({ + kind: "unknown", + name: "nope", + }); + }); +}); diff --git a/clients/mcpdo/__tests__/authorize.test.ts b/clients/mcpdo/__tests__/authorize.test.ts new file mode 100644 index 0000000000..ea73822cb0 --- /dev/null +++ b/clients/mcpdo/__tests__/authorize.test.ts @@ -0,0 +1,152 @@ +import { describe, it, expect, vi, afterEach } from "vitest"; +import type { MCPServerConfig } from "@inspector/core/mcp/types.js"; + +const connectSpy = vi.fn(); +const disconnectSpy = vi.fn().mockResolvedValue(undefined); +const navigationSpy = vi.fn(); + +vi.mock("@inspector/core/cli/cliOAuth.js", () => ({ + connectInspectorWithOAuth: (...args: unknown[]) => connectSpy(...args), +})); + +vi.mock("@inspector/core/cli/cli-oauth-navigation.js", () => ({ + createCliOAuthNavigation: (...args: unknown[]) => { + navigationSpy(...args); + return { navigate: vi.fn() }; + }, +})); + +vi.mock("@inspector/core/mcp/index.js", () => ({ + InspectorClient: class { + connect = vi.fn(); + disconnect = disconnectSpy; + }, +})); + +vi.mock("@inspector/core/client/runner.js", async (importOriginal) => { + const actual = + await importOriginal<typeof import("@inspector/core/client/runner.js")>(); + return { + ...actual, + loadRunnerClientConfig: vi.fn().mockResolvedValue({}), + buildRunnerClientAuthOptions: vi.fn().mockReturnValue({}), + }; +}); + +describe("authorizeInFrontend", () => { + afterEach(() => { + connectSpy.mockReset(); + disconnectSpy.mockClear(); + navigationSpy.mockClear(); + }); + + it("no-ops for non-OAuth-capable (stdio) configs", async () => { + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await authorizeInFrontend( + { type: "stdio", command: "x" } as MCPServerConfig, + undefined, + ); + expect(connectSpy).not.toHaveBeenCalled(); + }); + + it("runs connectInspectorWithOAuth for HTTP configs", async () => { + connectSpy.mockResolvedValue(undefined); + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await authorizeInFrontend( + { type: "streamable-http", url: "https://example.com/mcp" }, + { protocolEra: "2025-11-25" } as never, + { storedAuthOnly: true }, + ); + expect(connectSpy).toHaveBeenCalled(); + expect(disconnectSpy).toHaveBeenCalled(); + }); + + it("swallows disconnect failures in finally", async () => { + connectSpy.mockResolvedValue(undefined); + disconnectSpy.mockRejectedValueOnce(new Error("bye")); + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await expect( + authorizeInFrontend( + { type: "streamable-http", url: "https://example.com/mcp" }, + undefined, + ), + ).resolves.toBeUndefined(); + }); + + it("always admits interactive OAuth (isTTY: true), regardless of the real TTY state", async () => { + connectSpy.mockResolvedValue(undefined); + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await authorizeInFrontend( + { type: "streamable-http", url: "https://example.com/mcp" }, + undefined, + ); + const options = connectSpy.mock.calls[0]?.[5] as { isTTY?: boolean }; + expect(options.isTTY).toBe(true); + }); + + it("addresses the printed authorization line to whoever must relay it — a human directly, or an agent on behalf of one", async () => { + connectSpy.mockResolvedValue(undefined); + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await authorizeInFrontend( + { type: "streamable-http", url: "https://example.com/mcp" }, + undefined, + ); + const navOptions = navigationSpy.mock.calls[0]?.[0] as { + promptMessage: (hrefDisplay: string, tty: boolean) => string; + }; + expect(navOptions.promptMessage("https://example.com/auth", true)).toBe( + "Please navigate to: https://example.com/auth", + ); + expect(navOptions.promptMessage("https://example.com/auth", false)).toBe( + "The user needs to navigate to this link to authenticate: https://example.com/auth", + ); + }); + + it("rethrows non-EMA connect failures unchanged", async () => { + connectSpy.mockRejectedValue(new Error("network down")); + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await expect( + authorizeInFrontend( + { type: "streamable-http", url: "https://example.com/mcp" }, + undefined, + ), + ).rejects.toThrow("network down"); + expect(disconnectSpy).toHaveBeenCalled(); + }); + + it("uses a caller-provided navigation instead of the CLI default", async () => { + connectSpy.mockResolvedValue(undefined); + const makeNavigation = vi.fn().mockReturnValue({ navigate: vi.fn() }); + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await authorizeInFrontend( + { type: "streamable-http", url: "https://example.com/mcp" }, + undefined, + { makeNavigation }, + ); + expect(makeNavigation).toHaveBeenCalledTimes(1); + expect(navigationSpy).not.toHaveBeenCalled(); + }); + + it("maps EmaClientNotConfiguredError to actionable mcpdo guidance", async () => { + const { EmaClientNotConfiguredError } = + await import("@inspector/core/auth/ema/clientConfigError.js"); + connectSpy.mockRejectedValue(new EmaClientNotConfiguredError("disabled")); + const { authorizeInFrontend } = + await import("../src/connection/authorize.js"); + await expect( + authorizeInFrontend( + { type: "streamable-http", url: "https://example.com/mcp" }, + undefined, + ), + ).rejects.toThrow(/EMA.*disabled/i); + // Still tears the probe client down on the error path. + expect(disconnectSpy).toHaveBeenCalled(); + }); +}); diff --git a/clients/mcpdo/__tests__/connect-secret-storage.test.ts b/clients/mcpdo/__tests__/connect-secret-storage.test.ts new file mode 100644 index 0000000000..debc1d1f35 --- /dev/null +++ b/clients/mcpdo/__tests__/connect-secret-storage.test.ts @@ -0,0 +1,80 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import type { SecretStorageInfo } from "@inspector/core/auth/secret-storage-info.js"; + +const getSecretStorageInfo = vi.fn<() => Promise<SecretStorageInfo>>(); +const warnAboutSecretStorage = + vi.fn<(info: SecretStorageInfo, opts?: { force?: boolean }) => void>(); + +// Only these two names are consumed from the core module by mcp.ts; the rest +// of the real module is preserved so its other importers are unaffected. +vi.mock( + "@inspector/core/auth/node/secret-store-selection.js", + async (importOriginal) => { + const actual = + await importOriginal< + typeof import("@inspector/core/auth/node/secret-store-selection.js") + >(); + return { + ...actual, + getSecretStorageInfo: (...args: unknown[]) => + (getSecretStorageInfo as (...a: unknown[]) => unknown)(...args), + warnAboutSecretStorage: (...args: unknown[]) => + (warnAboutSecretStorage as (...a: unknown[]) => unknown)(...args), + }; + }, +); + +const { surfaceSecretStorageAtConnect } = + await import("../src/connection/mcp.js"); + +describe("surfaceSecretStorageAtConnect", () => { + let stderr: ReturnType<typeof vi.spyOn>; + + beforeEach(() => { + getSecretStorageInfo.mockReset(); + warnAboutSecretStorage.mockReset(); + stderr = vi.spyOn(process.stderr, "write").mockImplementation(() => true); + }); + + afterEach(() => { + stderr.mockRestore(); + }); + + it("re-surfaces the store warning once, forced (R5)", async () => { + const info: SecretStorageInfo = { + kind: "file", + reason: "fallback", + durable: true, + path: "/home/u/.mcp-inspector/secrets.json", + plaintext: true, + }; + getSecretStorageInfo.mockResolvedValue(info); + + await surfaceSecretStorageAtConnect(); + + expect(warnAboutSecretStorage).toHaveBeenCalledWith(info, { force: true }); + // Non-memory stores get no extra mcpdo-specific line. + expect(stderr).not.toHaveBeenCalled(); + }); + + it("warns that a memory store cannot persist across mcpdo's processes (R4)", async () => { + getSecretStorageInfo.mockResolvedValue({ + kind: "memory", + reason: "configured", + durable: false, + }); + + await surfaceSecretStorageAtConnect(); + + expect(warnAboutSecretStorage).toHaveBeenCalledWith( + expect.objectContaining({ kind: "memory" }), + { force: true }, + ); + const printed = stderr.mock.calls + .map((c: unknown[]) => String(c[0])) + .join(""); + expect(printed).toContain("MCP_INSPECTOR_SECRET_STORE=memory"); + expect(printed).toContain("separate processes"); + expect(printed).toContain('Use "file" or "keyring"'); + }); +}); diff --git a/clients/mcpdo/__tests__/connection-stored-auth.test.ts b/clients/mcpdo/__tests__/connection-stored-auth.test.ts new file mode 100644 index 0000000000..4f1596ef5d --- /dev/null +++ b/clients/mcpdo/__tests__/connection-stored-auth.test.ts @@ -0,0 +1,329 @@ +import { afterEach, describe, expect, it } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { resetNodeOAuthStorageCache } from "@inspector/core/auth/node/storage-node.js"; +import { writeOAuthSections } from "@inspector/core/auth/node/oauth-persist-file.js"; +import { + clearAllStoredAuth, + clearStoredAuth, + clearStoredAuthForRelogin, + listStoredAuth, + resolveStoredAuthKey, +} from "../src/connection/stored-auth.js"; +import { CliExitCodeError } from "@inspector/core/cli/error-handler.js"; +import { defaultSecretStore } from "@inspector/core/auth/node/secret-store-selection.js"; +import { oauthSecretServerId } from "@inspector/core/auth/node/oauth-secrets.js"; +import { runMcp } from "./helpers/mcp-runner.js"; +import { + expectCliSuccess, + expectCliFailure, +} from "../../cli/__tests__/helpers/assertions.js"; + +function writeOAuthFixture(dir: string): string { + const file = path.join(dir, "oauth.json"); + fs.writeFileSync( + file, + JSON.stringify({ + servers: { + "https://example.com/mcp": { + byIssuer: { + "https://as.example/": { + tokens: { + access_token: "a", + token_type: "Bearer", + refresh_token: "r", + }, + }, + }, + activeIssuer: "https://as.example/", + }, + "https://other.example/mcp": { + tokens: { access_token: "x", token_type: "Bearer" }, + }, + "https://empty.example/mcp": { + codeVerifier: "cv", + }, + // Note: entries that are not objects (null, strings) now make the + // whole file unrecognized upstream (proto-safe parsing) — the file + // indexes secret-store entries, so core refuses rather than treats + // it as empty. Junk entries therefore no longer belong in a fixture. + "https://issuer-empty.example/mcp": { + byIssuer: { + "https://as.example/": {}, + }, + }, + "https://multi.example/mcp": { + // First issuer: access only. Second: access + refresh. The summary + // must aggregate across slots, not stop at the first access token. + byIssuer: { + "https://as-a.example/": { + tokens: { access_token: "a1", token_type: "Bearer" }, + }, + "https://as-b.example/": { + tokens: { + access_token: "a2", + token_type: "Bearer", + refresh_token: "r2", + }, + }, + }, + }, + }, + idpSessions: {}, + }), + "utf8", + ); + return file; +} + +const FIXTURE_URLS = [ + "https://example.com/mcp", + "https://other.example/mcp", + "https://empty.example/mcp", + "https://issuer-empty.example/mcp", + "https://multi.example/mcp", +]; + +/** + * The pinned in-memory secret store (vitest.config.ts) is process-wide and + * keyed by server URL, not by state-file path — reads migrate fixture + * plaintext into it and joined reads prefer it over the file, so one test's + * migrated tokens would leak into the next test's fresh fixture. + */ +async function purgeFixtureSecrets(): Promise<void> { + const store = defaultSecretStore(); + for (const url of FIXTURE_URLS) { + await store.deleteAllForServer(oauthSecretServerId(url)); + } +} + +describe("connection stored-auth helpers", () => { + let dir: string | undefined; + let prevPath: string | undefined; + + afterEach(async () => { + if (prevPath === undefined) + delete process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + else process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = prevPath; + resetNodeOAuthStorageCache(); + await purgeFixtureSecrets(); + if (dir) { + fs.rmSync(dir, { recursive: true, force: true }); + dir = undefined; + } + }); + + function useFixture(): string { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-stored-auth-")); + const file = writeOAuthFixture(dir); + prevPath = process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = file; + resetNodeOAuthStorageCache(); + return file; + } + + it("lists byIssuer and legacy tokens", async () => { + const file = useFixture(); + const list = await listStoredAuth(); + expect(list.oauthStatePath).toBe(file); + expect(list.servers.map((s) => s.url)).toEqual([ + "https://empty.example/mcp", + "https://example.com/mcp", + "https://issuer-empty.example/mcp", + "https://multi.example/mcp", + "https://other.example/mcp", + ]); + expect( + list.servers.find((s) => s.url.includes("issuer-empty")), + ).toMatchObject({ hasTokens: false, hasRefreshToken: false }); + // Aggregated across issuer slots: the refresh token lives in the second + // slot, so first-match-wins would have reported hasRefreshToken: false. + expect(list.servers.find((s) => s.url.includes("multi"))).toMatchObject({ + hasTokens: true, + hasRefreshToken: true, + }); + expect( + list.servers.find((s) => s.url === "https://example.com/mcp"), + ).toMatchObject({ hasTokens: true, hasRefreshToken: true }); + expect(list.servers.find((s) => s.url.includes("other"))).toMatchObject({ + hasTokens: true, + hasRefreshToken: false, + }); + expect(list.servers.find((s) => s.url.includes("empty"))).toMatchObject({ + hasTokens: false, + hasRefreshToken: false, + }); + }); + + it("still reports tokens after they are split into the secret store", async () => { + // Regression for the #2482 adaptation: writes split tokens out of + // oauth.json into the secret store, so a raw-file parse would report + // every entry as token-less. listStoredAuth must use the joined read. + // (Read-side plaintext migration is skipped for the non-durable memory + // store pinned in vitest.config.ts, so exercise the write-side split.) + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-stored-auth-")); + const file = path.join(dir, "oauth.json"); + prevPath = process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = file; + resetNodeOAuthStorageCache(); + + await writeOAuthSections( + file, + { + servers: { + "https://example.com/mcp": { + tokens: { + access_token: "a", + token_type: "Bearer", + refresh_token: "r", + }, + }, + }, + idpSessions: {}, + }, + { servers: ["https://example.com/mcp"] }, + ); + + const raw = fs.readFileSync(file, "utf8"); + expect(raw).not.toContain("access_token"); + + const list = await listStoredAuth(); + expect( + list.servers.find((s) => s.url === "https://example.com/mcp"), + ).toMatchObject({ hasTokens: true, hasRefreshToken: true }); + }); + + it("clears one key and all keys", async () => { + useFixture(); + const cleared = await clearStoredAuth("https://example.com/mcp"); + expect(cleared.url).toBe("https://example.com/mcp"); + let list = await listStoredAuth(); + expect(list.servers.map((s) => s.url)).not.toContain( + "https://example.com/mcp", + ); + + const all = await clearAllStoredAuth(); + expect(all.cleared).toBe(4); + list = await listStoredAuth(); + expect(list.servers).toEqual([]); + }); + + it("resolveStoredAuthKey rejects unknown non-URL keys", async () => { + useFixture(); + await expect(resolveStoredAuthKey("nope")).rejects.toBeInstanceOf( + CliExitCodeError, + ); + }); + + it("clearStoredAuthForRelogin clears by URL", async () => { + useFixture(); + await clearStoredAuthForRelogin("https://other.example/mcp"); + const list = await listStoredAuth(); + expect(list.servers.map((s) => s.url)).not.toContain( + "https://other.example/mcp", + ); + await clearStoredAuthForRelogin(undefined); + await clearStoredAuthForRelogin(" "); + }); + + it("lists an empty store when the file is missing", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-stored-auth-")); + const missing = path.join(dir, "missing-oauth.json"); + prevPath = process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = missing; + resetNodeOAuthStorageCache(); + expect(await listStoredAuth()).toMatchObject({ + oauthStatePath: missing, + servers: [], + }); + }); + + it("resolves keys by normalisation and rejects blanks", async () => { + useFixture(); + await expect(resolveStoredAuthKey(" ")).rejects.toBeInstanceOf( + CliExitCodeError, + ); + await expect(resolveStoredAuthKey("https://Example.COM/mcp")).resolves.toBe( + "https://example.com/mcp", + ); + await expect( + resolveStoredAuthKey("https://brand-new.example/mcp"), + ).resolves.toBe("https://brand-new.example/mcp"); + }); +}); + +describe("mcp auth/list and auth/clear", () => { + let dir: string | undefined; + + afterEach(async () => { + resetNodeOAuthStorageCache(); + await purgeFixtureSecrets(); + if (dir) { + fs.rmSync(dir, { recursive: true, force: true }); + dir = undefined; + } + }); + + it("lists and clears via connection commands", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-auth-cmd-")); + const file = writeOAuthFixture(dir); + resetNodeOAuthStorageCache(); + + const listed = await runMcp(["auth/list", "--format", "json"], { + env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file }, + }); + expectCliSuccess(listed); + const body = JSON.parse(listed.stdout) as { + servers: { url: string }[]; + }; + expect(body.servers.length).toBe(5); + + const cleared = await runMcp( + ["auth/clear", "https://example.com/mcp", "--format", "json"], + { env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file } }, + ); + expectCliSuccess(cleared); + expect(JSON.parse(cleared.stdout)).toEqual({ + url: "https://example.com/mcp", + }); + + const all = await runMcp( + ["auth/clear", "--all", "--yes", "--format", "json"], + { env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file } }, + ); + expectCliSuccess(all); + expect(JSON.parse(all.stdout)).toMatchObject({ all: true, cleared: 4 }); + }); + + it("rejects --all without --yes when non-interactive", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-auth-cmd-")); + const file = writeOAuthFixture(dir); + const result = await runMcp(["auth/clear", "--all"], { + env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file }, + }); + expectCliFailure(result); + expect(result.stderr).toMatch(/--yes/); + }); + + it("rejects auth/clear usage errors", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-auth-cmd-")); + const file = writeOAuthFixture(dir); + const none = await runMcp(["auth/clear"], { + env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file }, + }); + expectCliFailure(none); + + const both = await runMcp( + ["auth/clear", "https://example.com/mcp", "--all", "--yes"], + { env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file } }, + ); + expectCliFailure(both); + + const human = await runMcp(["auth/list"], { + env: { MCP_INSPECTOR_OAUTH_STATE_PATH: file }, + }); + expectCliSuccess(human); + expect(human.stdout).toMatch(/Stored auth/); + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-connections.test.ts b/clients/mcpdo/__tests__/daemon-connections.test.ts new file mode 100644 index 0000000000..e21361412c --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-connections.test.ts @@ -0,0 +1,1967 @@ +import { describe, it, expect, afterEach, beforeEach, vi } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { getTestMcpServerCommand } from "@modelcontextprotocol/inspector-test-server"; +import { DaemonServer } from "../src/daemon/server.js"; +import { callDaemon } from "../src/daemon/client.js"; +import { parseRequestLine, encodeResponse } from "../src/daemon/framing.js"; +import { + DEFAULT_IDLE_MS, + elicitCapabilityToClientOption, + getLiveConnectionAuthInfo, + getConnectionAuthInfo, + isConnectionAuthRequiredError, + ConnectionRegistry, +} from "../src/daemon/connections.js"; +import { CliExitCodeError } from "@inspector/core/cli/error-handler.js"; +import { AuthRecoveryRequiredError } from "@inspector/core/auth/challenge.js"; + +describe("daemon framing", () => { + it("parses and rejects invalid request lines", () => { + expect(parseRequestLine("")).toBeNull(); + expect(parseRequestLine(" ")).toBeNull(); + expect(parseRequestLine('{"id":"1","op":"ping"}')).toEqual({ + id: "1", + op: "ping", + }); + expect(() => parseRequestLine("not-json")).toThrow(); + expect(() => parseRequestLine('{"op":"ping"}')).toThrow(/Invalid daemon/); + expect(encodeResponse({ id: "1", ok: true, result: { pong: true } })).toBe( + '{"id":"1","ok":true,"result":{"pong":true}}\n', + ); + }); +}); + +describe("elicitCapabilityToClientOption", () => { + it("maps each elicitCapability mode to the InspectorClient elicit shape", () => { + expect(elicitCapabilityToClientOption("off")).toBe(false); + expect(elicitCapabilityToClientOption("url")).toEqual({ url: true }); + expect(elicitCapabilityToClientOption("form")).toEqual({ form: true }); + expect(elicitCapabilityToClientOption("both")).toEqual({ + url: true, + form: true, + }); + }); + + it("defaults to both (url+form) when unset, matching the pre-#1783 hardcoded default", () => { + expect(elicitCapabilityToClientOption(undefined)).toEqual({ + url: true, + form: true, + }); + }); +}); + +describe("isConnectionAuthRequiredError", () => { + it("treats EMA client misconfiguration as auth_required (front-end maps it to guidance)", async () => { + const { EmaClientNotConfiguredError } = + await import("@inspector/core/auth/ema/clientConfigError.js"); + expect( + isConnectionAuthRequiredError( + new EmaClientNotConfiguredError("not_configured"), + ), + ).toBe(true); + }); + + it("recognizes unauthorized, recovery, and SDK token-exchange failures", () => { + expect(isConnectionAuthRequiredError(new Error("nope"))).toBe(false); + expect( + isConnectionAuthRequiredError( + new AuthRecoveryRequiredError(new URL("https://as.example/a"), { + reason: "unauthorized", + }), + ), + ).toBe(true); + const unauthorized = Object.assign(new Error("boom"), { status: 401 }); + expect(isConnectionAuthRequiredError(unauthorized)).toBe(true); + expect( + isConnectionAuthRequiredError( + new Error( + "Either provider.prepareTokenRequest() or authorizationCode is required", + ), + ), + ).toBe(true); + expect( + isConnectionAuthRequiredError( + new Error("redirectUrl is required for authorization_code flow"), + ), + ).toBe(true); + expect( + isConnectionAuthRequiredError( + new Error("No code verifier saved for connection"), + ), + ).toBe(true); + }); +}); + +describe("getConnectionAuthInfo", () => { + const clientWith = ( + getOAuthState: () => Promise<unknown>, + ): Parameters<typeof getConnectionAuthInfo>[0] => + ({ getOAuthState }) as unknown as Parameters< + typeof getConnectionAuthInfo + >[0]; + + it("is undefined for no-auth connections and when the state read fails", async () => { + expect( + await getConnectionAuthInfo(clientWith(async () => undefined)), + ).toBeUndefined(); + expect( + await getConnectionAuthInfo( + clientWith(async () => { + throw new Error("storage unavailable"); + }), + ), + ).toBeUndefined(); + }); + + it("projects standard OAuth state (scope + clientId when present)", async () => { + expect( + await getConnectionAuthInfo( + clientWith(async () => ({ + authorized: true, + protocol: "standard", + serverUrl: "https://mcp.example", + grantedScope: "mcp:tools", + client: { clientId: "client-123", hasClientSecret: false }, + })), + ), + ).toEqual({ + method: "oauth", + authorized: true, + scope: "mcp:tools", + clientId: "client-123", + }); + }); + + it("projects EMA state with IdP session and omits absent optionals", async () => { + expect( + await getConnectionAuthInfo( + clientWith(async () => ({ + authorized: false, + protocol: "ema", + serverUrl: "https://mcp.example", + ema: { + idpIssuer: "https://idp.example", + idpClientId: "idp-client", + idpSession: "logged_in", + }, + })), + ), + ).toEqual({ method: "ema", authorized: false, idpSession: "logged_in" }); + }); +}); + +describe("getLiveConnectionAuthInfo", () => { + it("is undefined for stdio, malformed http configs, and unengaged OAuth", async () => { + const { resetNodeOAuthStorageCache } = + await import("@inspector/core/auth/node/storage-node.js"); + const dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-conn-live-auth-")); + const saved = process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + const savedClient = process.env.MCP_CLIENT_CONFIG_PATH; + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = path.join(dir, "oauth.json"); + process.env.MCP_CLIENT_CONFIG_PATH = path.join(dir, "client.json"); + resetNodeOAuthStorageCache(); + try { + expect( + await getLiveConnectionAuthInfo({ + serverConfig: { type: "stdio", command: "x" }, + }), + ).toBeUndefined(); + // Defensive: OAuth-capable type without a usable url. + expect( + await getLiveConnectionAuthInfo({ + serverConfig: { type: "streamable-http" } as never, + }), + ).toBeUndefined(); + // http server, no oauth config anywhere, empty storage: no snapshot. + expect( + await getLiveConnectionAuthInfo({ + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + }), + ).toBeUndefined(); + // Corrupt oauth.json: the disk read fails, and the best-effort catch + // yields undefined rather than failing connections/show. + fs.writeFileSync(process.env.MCP_INSPECTOR_OAUTH_STATE_PATH!, "{nope"); + resetNodeOAuthStorageCache(); + expect( + await getLiveConnectionAuthInfo({ + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + }), + ).toBeUndefined(); + } finally { + if (saved === undefined) + delete process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + else process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = saved; + if (savedClient === undefined) delete process.env.MCP_CLIENT_CONFIG_PATH; + else process.env.MCP_CLIENT_CONFIG_PATH = savedClient; + resetNodeOAuthStorageCache(); + fs.rmSync(dir, { recursive: true, force: true }); + } + }); +}); + +describe("ConnectionRegistry", () => { + it("requires an explicit connection when asked", () => { + const registry = new ConnectionRegistry(0); + expect(() => registry.resolve(undefined, true)).toThrow(CliExitCodeError); + expect(() => registry.resolve(undefined, false)).toThrow( + /No open connections/, + ); + }); + + it("tracks MRU across connect/disconnect", async () => { + const { command, args } = getTestMcpServerCommand(); + const registry = new ConnectionRegistry(0); + const a = await registry.connect({ + name: "a", + serverConfig: { type: "stdio", command, args }, + serverIdentity: `${command} ${args.join(" ")}`, + }); + expect(a.isMru).toBe(true); + // stdio transport: no OAuth, so no auth snapshot is reported. + expect(a.auth).toBeUndefined(); + const b = await registry.connect({ + name: "b", + serverConfig: { type: "stdio", command, args }, + serverIdentity: `${command} ${args.join(" ")}`, + }); + expect(b.isMru).toBe(true); + expect(registry.getMruName()).toBe("b"); + registry.use("a"); + expect(registry.getMruName()).toBe("a"); + await registry.disconnect("b", false); + expect(registry.list().map((s) => s.name)).toEqual(["a"]); + await registry.disconnect(undefined, false); + expect(registry.connectionCount()).toBe(0); + expect(DEFAULT_IDLE_MS).toBe(60_000); + }); + + it("disconnect returns the server URL for a URL-keyed connection (for --clear-auth)", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValue(undefined); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + try { + await registry.connect({ + name: "http", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }); + // The URL comes back so the front-end can clear this server's stored + // OAuth entry; it is captured before the connection is deleted. + await expect(registry.disconnect("http", false)).resolves.toEqual({ + name: "http", + serverUrl: "https://mcp.example.com/mcp", + }); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + } + }); + + it("disconnect omits serverUrl for a stdio connection (no URL key to clear)", async () => { + const { command, args } = getTestMcpServerCommand(); + const registry = new ConnectionRegistry(0); + await registry.connect({ + name: "local", + serverConfig: { type: "stdio", command, args }, + serverIdentity: `${command} ${args.join(" ")}`, + }); + const result = await registry.disconnect("local", false); + expect(result).toEqual({ name: "local" }); + expect("serverUrl" in result).toBe(false); + }); + + it("passes the server's configured roots to the client (cleaned), so roots are advertised at initialize", async () => { + const { command, args } = getTestMcpServerCommand(); + const registry = new ConnectionRegistry(0); + try { + await registry.connect({ + name: "a", + serverConfig: { type: "stdio", command, args }, + serverIdentity: `${command} ${args.join(" ")}`, + serverSettings: { + headers: [], + metadata: {}, + env: [], + connectionTimeout: 30_000, + requestTimeout: 0, + taskTtl: 0, + maxFetchRequests: 0, + autoRefreshOnListChanged: false, + paginatedLists: false, + // The blank-uri entry exercises cleanRoots: hand-edited mcp.json + // can hold shapes the types promise away (#1797). + roots: [ + { uri: "file:///tmp/project", name: "project" }, + { uri: " " }, + ], + }, + }); + expect(registry.clientFor("a", false)?.getRoots()).toEqual([ + { uri: "file:///tmp/project", name: "project" }, + ]); + } finally { + await registry.disconnectAll(); + } + }); + + it("disconnectAll tears down the remaining connections when one disconnect fails", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValue(undefined); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + try { + const params = (name: string) => + ({ + name, + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }) as const; + await registry.connect(params("a")); + await registry.connect(params("b")); + // First teardown loses a race (connection_not_found); shutdown must + // still settle and tear down the rest instead of leaking "b". + const spy = vi.spyOn(registry, "disconnect"); + spy.mockRejectedValueOnce( + new CliExitCodeError(1, "No connection named 'a'.", { + code: "connection_not_found", + }), + ); + await expect(registry.disconnectAll()).resolves.toBeUndefined(); + expect(spy).toHaveBeenCalledTimes(2); + // "b" was genuinely disconnected, not abandoned mid-loop. + expect(registry.list().map((c) => c.name)).toEqual(["a"]); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + } + }); + + it("serializes concurrent connects for the same name so the replaced client is torn down, not leaked", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + // Slow connect widens the check→set window that raced pre-lock. + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockImplementation( + () => new Promise<void>((resolve) => setTimeout(resolve, 25)), + ); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + try { + const params = { + name: "dup", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + } as const; + const [a, b] = await Promise.all([ + registry.connect(params), + registry.connect(params), + ]); + expect(a.name).toBe("dup"); + expect(b.name).toBe("dup"); + // Exactly one tracked connection; the loser of the race was + // disconnected by the serialized reconnect path, not orphaned. + expect(registry.connectionCount()).toBe(1); + expect(connectSpy).toHaveBeenCalledTimes(2); + expect(disconnectSpy).toHaveBeenCalledTimes(1); + await registry.disconnect("dup", false); + expect(disconnectSpy).toHaveBeenCalledTimes(2); + expect(registry.connectionCount()).toBe(0); + // A queued duplicate disconnect fails cleanly rather than tearing + // down a successor's connection. + await expect(registry.disconnect("dup", false)).rejects.toThrow( + /isn't connected/, + ); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + } + }); + + it("liveClientFor revives a connection whose transport settled into a terminal state", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValue(undefined); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + let deadClient: unknown; + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockImplementation(function (this: unknown) { + // Only the original client is dead; the revived one is live. + return this === deadClient ? "error" : "connected"; + }); + try { + await registry.connect({ + name: "r", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }); + const original = registry.clientFor("r", false); + + // Live client: no re-dial. + expect(await registry.liveClientFor("r", false)).toBe(original); + expect(connectSpy).toHaveBeenCalledTimes(1); + + // Simulate the overnight drop (server expired the session). + deadClient = original; + const revived = await registry.liveClientFor("r", false); + expect(revived).not.toBe(original); + expect(connectSpy).toHaveBeenCalledTimes(2); + // The dead client's resources were released. + expect(disconnectSpy).toHaveBeenCalledTimes(1); + // The registry entry was swapped in place — still one connection. + expect(registry.connectionCount()).toBe(1); + expect(registry.clientFor("r", false)).toBe(revived); + // Live now: no further re-dial. + expect(await registry.liveClientFor("r", false)).toBe(revived); + expect(connectSpy).toHaveBeenCalledTimes(2); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + statusSpy.mockRestore(); + } + }); + + it("revive maps a credentials failure to auth_required and keeps the entry on any failure", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValueOnce(undefined) // initial connect + .mockRejectedValueOnce( + // Revive 1: SDK token-exchange failure meaning "needs full re-auth". + new Error("prepareTokenRequest() or authorizationCode is required"), + ) + .mockRejectedValueOnce(new Error("connect ECONNREFUSED 127.0.0.1:443")) // revive 2: server down + .mockResolvedValueOnce(undefined); // revive 3: server back + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + let deadClient: unknown; + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockImplementation(function (this: unknown) { + return this === deadClient ? "error" : "connected"; + }); + try { + await registry.connect({ + name: "r", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }); + deadClient = registry.clientFor("r", false); + + // Refresh credentials are gone: the ONLY case that involves the user. + await expect(registry.liveClientFor("r", false)).rejects.toMatchObject({ + envelope: { code: "auth_required" }, + }); + await expect(registry.liveClientFor("r", false)).rejects.toThrow( + /ECONNREFUSED/, + ); + // Both failures kept the entry — the user's intent persists… + expect(registry.connectionCount()).toBe(1); + // …so a later op simply revives once the server is reachable again. + const revived = await registry.liveClientFor("r", false); + expect(revived).not.toBe(deadClient); + expect(registry.clientFor("r", false)).toBe(revived); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + statusSpy.mockRestore(); + } + }); + + describe("pending auth URL relay (surfaces 1 & 2)", () => { + const RELAY_URL = "https://as.example/authorize?client_id=abc&state=xyz"; + let markerDir: string; + let prevDaemonDir: string | undefined; + + async function writeLiveMarker(serverUrl: string): Promise<void> { + const { pendingAuthMarkerPath } = + await import("../src/connection/auth-helper.js"); + fs.writeFileSync( + pendingAuthMarkerPath(serverUrl), + JSON.stringify({ + url: RELAY_URL, + pid: process.pid, + expiresAt: Date.now() + 60_000, + }), + { mode: 0o600 }, + ); + } + + beforeEach(() => { + markerDir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-relay-marker-")); + prevDaemonDir = process.env.MCP_INSPECTOR_DAEMON_DIR; + process.env.MCP_INSPECTOR_DAEMON_DIR = markerDir; + }); + + afterEach(() => { + if (prevDaemonDir === undefined) + delete process.env.MCP_INSPECTOR_DAEMON_DIR; + else process.env.MCP_INSPECTOR_DAEMON_DIR = prevDaemonDir; + fs.rmSync(markerDir, { recursive: true, force: true }); + }); + + it("pendingAuthUrlFor returns a live marker's URL, keyed by server URL", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const serverUrl = "https://mcp.example.com/mcp"; + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockRejectedValue(Object.assign(new Error("boom"), { status: 401 })); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockReturnValue("disconnected"); + const registry = new ConnectionRegistry(0); + try { + await registry.connect({ + name: "p", + serverConfig: { type: "streamable-http", url: serverUrl }, + serverIdentity: serverUrl, + pendingOnAuthRequired: true, + }); + const connection = registry.connectionFor("p", false); + // No marker yet → undefined (credential lapse, not a live sign-in). + expect(registry.pendingAuthUrlFor(connection)).toBeUndefined(); + // Live marker keyed by the server URL → the relay URL. + await writeLiveMarker(serverUrl); + expect(registry.pendingAuthUrlFor(connection)).toBe(RELAY_URL); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + statusSpy.mockRestore(); + } + }); + + it("pendingAuthUrlFor returns undefined for a stdio config (no URL to key on)", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValue(undefined); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockReturnValue("connected"); + const registry = new ConnectionRegistry(0); + try { + await registry.connect({ + name: "s", + serverConfig: { + type: "stdio", + command: "x", + }, + serverIdentity: "stdio:test", + }); + const connection = registry.connectionFor("s", false); + expect(registry.pendingAuthUrlFor(connection)).toBeUndefined(); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + statusSpy.mockRestore(); + } + }); + + it.each([ + { + who: "agent (non-interactive)", + interactive: false, + expected: /display the sign-in link again/, + embedsUrl: false, + }, + { + who: "human (interactive)", + interactive: true, + expected: /sign-in is still pending\. To see details/, + embedsUrl: false, + }, + ])( + "a failed revive with a live sign-in gives $who the right message", + async ({ interactive, expected, embedsUrl }) => { + const { InspectorClient } = + await import("@inspector/core/mcp/index.js"); + const serverUrl = "https://mcp.example.com/mcp"; + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + // Initial dial → auth_required (registers pending); revive → same. + .mockRejectedValue(Object.assign(new Error("boom"), { status: 401 })); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockReturnValue("disconnected"); + const registry = new ConnectionRegistry(0); + try { + await registry.connect({ + name: "p", + serverConfig: { type: "streamable-http", url: serverUrl }, + serverIdentity: serverUrl, + pendingOnAuthRequired: true, + }); + await writeLiveMarker(serverUrl); + await expect( + registry.liveClientFor("p", false, interactive), + ).rejects.toMatchObject({ + envelope: { code: "auth_required" }, + }); + await expect( + registry.liveClientFor("p", false, interactive), + ).rejects.toThrow(expected); + // Neither message embeds the live sign-in URL anymore: the agent is + // pointed at `connections/show @name` (which relays it with the + // right turn discipline) and the human reads it from `connect` / + // `connections/show`. So neither should carry the client_id. + if (embedsUrl) { + await expect( + registry.liveClientFor("p", false, interactive), + ).rejects.toThrow(/client_id=abc/); + } else { + await expect( + registry.liveClientFor("p", false, interactive), + ).rejects.not.toThrow(/client_id=abc/); + } + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + statusSpy.mockRestore(); + } + }, + ); + + it("a failed revive with no live sign-in keeps the 'needs re-authentication' message", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const serverUrl = "https://mcp.example.com/mcp"; + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockRejectedValue(Object.assign(new Error("boom"), { status: 401 })); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockReturnValue("disconnected"); + const registry = new ConnectionRegistry(0); + try { + await registry.connect({ + name: "p", + serverConfig: { type: "streamable-http", url: serverUrl }, + serverIdentity: serverUrl, + pendingOnAuthRequired: true, + }); + // pendingAuth is set, but NO live marker: this is a credential lapse, + // not a sign-in in flight — so the message tells the user to connect. + await expect(registry.liveClientFor("p", false, false)).rejects.toThrow( + /needs re-authentication/, + ); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + statusSpy.mockRestore(); + } + }); + }); + + it("pendingOnAuthRequired registers a dormant intent entry that completes via revive on first use", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + // Dial: no stored tokens yet — auth_required. + .mockRejectedValueOnce(Object.assign(new Error("boom"), { status: 401 })) + // Revive after the helper stored tokens: succeeds. + .mockResolvedValueOnce(undefined); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + let pendingClient: unknown; + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockImplementation(function (this: unknown) { + // The pending entry's client never connected (terminal status → + // revivable); the revived client is live. + return this === pendingClient ? "disconnected" : "connected"; + }); + try { + const info = await registry.connect({ + name: "p", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + pendingOnAuthRequired: true, + }); + // Registered as pending intent instead of throwing. + expect(info.pendingAuth).toBe(true); + expect(info.auth).toEqual({ method: "oauth", authorized: false }); + expect(registry.connectionCount()).toBe(1); + expect(registry.list()[0]).toMatchObject({ + name: "p", + pendingAuth: true, + }); + pendingClient = registry.clientFor("p", false); + + // First op after tokens land: revive dials and clears the flag. + const revived = await registry.liveClientFor("p", false); + expect(revived).not.toBe(pendingClient); + expect(registry.list()[0]?.pendingAuth).toBeUndefined(); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + statusSpy.mockRestore(); + } + }); + + it("pendingOnAuthRequired only swallows auth_required — other dial failures still throw with no entry", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockRejectedValueOnce(new Error("connect ECONNREFUSED 127.0.0.1:443")); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const registry = new ConnectionRegistry(0); + try { + await expect( + registry.connect({ + name: "p", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + pendingOnAuthRequired: true, + }), + ).rejects.toThrow(/ECONNREFUSED/); + expect(registry.connectionCount()).toBe(0); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + } + }); + + it("without pendingOnAuthRequired, an auth_required dial still throws with no entry", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockRejectedValueOnce(Object.assign(new Error("boom"), { status: 401 })); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const registry = new ConnectionRegistry(0); + try { + await expect( + registry.connect({ + name: "p", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }), + ).rejects.toMatchObject({ envelope: { code: "auth_required" } }); + expect(registry.connectionCount()).toBe(0); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + } + }); + + it("connections/show completes a pending-auth entry once signed-in tokens are on disk", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const { NodeOAuthStorage, resetNodeOAuthStorageCache } = + await import("@inspector/core/auth/node/storage-node.js"); + const serverUrl = "https://mcp.example.com/mcp"; + // Isolated client.json + oauth.json so the show handler's disk reads are + // deterministic (same pattern as the show-recomputes-from-disk test). + const stateDir = fs.mkdtempSync( + path.join(os.tmpdir(), "mcp-show-pending-"), + ); + const savedEnv = { + MCP_CLIENT_CONFIG_PATH: process.env.MCP_CLIENT_CONFIG_PATH, + MCP_INSPECTOR_OAUTH_STATE_PATH: + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH, + }; + process.env.MCP_CLIENT_CONFIG_PATH = path.join(stateDir, "client.json"); + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = path.join( + stateDir, + "oauth.json", + ); + resetNodeOAuthStorageCache(); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + // Dial at connect time: no stored tokens yet — auth_required. + .mockRejectedValueOnce(Object.assign(new Error("boom"), { status: 401 })) + // Revive triggered by connections/show after tokens land: succeeds. + .mockResolvedValueOnce(undefined); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + let pendingClient: unknown; + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockImplementation(function (this: unknown) { + return this === pendingClient ? "disconnected" : "connected"; + }); + const server = new DaemonServer({ + dir: fs.mkdtempSync(path.join(os.tmpdir(), "mcp-show-pending-daemon-")), + idleMs: 0, + }); + try { + await server.registry.connect({ + name: "p", + serverConfig: { type: "streamable-http", url: serverUrl }, + serverIdentity: serverUrl, + pendingOnAuthRequired: true, + }); + pendingClient = server.registry.clientFor("p", false); + + // Before sign-in completes, show reports the pending snapshot (no + // usable tokens on disk → no revive attempt). + const before = await server.handle({ + id: "s1", + op: "connections/show", + params: { name: "p" }, + }); + expect(before.ok).toBe(true); + if (!before.ok) throw new Error("unreachable"); + expect(before.result).toMatchObject({ + pendingAuth: true, + transport: "dormant", + }); + + // The detached helper finishes sign-in: usable tokens land on disk. + await new NodeOAuthStorage().saveTokens(serverUrl, { + access_token: "opaque-access-token", + token_type: "Bearer", + }); + resetNodeOAuthStorageCache(); + + // Now show itself completes the connection via revive. + const after = await server.handle({ + id: "s2", + op: "connections/show", + params: { name: "p" }, + }); + expect(after.ok).toBe(true); + if (!after.ok) throw new Error("unreachable"); + expect( + (after.result as { pendingAuth?: boolean }).pendingAuth, + ).toBeUndefined(); + expect(after.result).toMatchObject({ + transport: "live", + auth: { method: "oauth", authorized: true }, + }); + expect(server.registry.clientFor("p", false)).not.toBe(pendingClient); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + statusSpy.mockRestore(); + process.env.MCP_CLIENT_CONFIG_PATH = savedEnv.MCP_CLIENT_CONFIG_PATH; + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = + savedEnv.MCP_INSPECTOR_OAUTH_STATE_PATH; + resetNodeOAuthStorageCache(); + await server.stop().catch(() => {}); + fs.rmSync(stateDir, { recursive: true, force: true }); + } + }); + + it("connections/show keeps the pending snapshot when the revive attempt fails", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const { NodeOAuthStorage, resetNodeOAuthStorageCache } = + await import("@inspector/core/auth/node/storage-node.js"); + const serverUrl = "https://mcp.example.com/mcp"; + const stateDir = fs.mkdtempSync( + path.join(os.tmpdir(), "mcp-show-pending-fail-"), + ); + const savedEnv = { + MCP_CLIENT_CONFIG_PATH: process.env.MCP_CLIENT_CONFIG_PATH, + MCP_INSPECTOR_OAUTH_STATE_PATH: + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH, + }; + process.env.MCP_CLIENT_CONFIG_PATH = path.join(stateDir, "client.json"); + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = path.join( + stateDir, + "oauth.json", + ); + resetNodeOAuthStorageCache(); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + // Both the original dial and the show-triggered revive fail. + .mockRejectedValue(Object.assign(new Error("boom"), { status: 401 })); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockReturnValue("disconnected"); + const server = new DaemonServer({ + dir: fs.mkdtempSync( + path.join(os.tmpdir(), "mcp-show-pending-fail-daemon-"), + ), + idleMs: 0, + }); + try { + await server.registry.connect({ + name: "p", + serverConfig: { type: "streamable-http", url: serverUrl }, + serverIdentity: serverUrl, + pendingOnAuthRequired: true, + }); + await new NodeOAuthStorage().saveTokens(serverUrl, { + access_token: "opaque-access-token", + token_type: "Bearer", + }); + resetNodeOAuthStorageCache(); + + // Revive fails (tokens rejected on dial): show still answers with the + // honest pending snapshot instead of erroring, and the entry survives + // so a later op retries. + const shown = await server.handle({ + id: "s1", + op: "connections/show", + params: { name: "p" }, + }); + expect(shown.ok).toBe(true); + if (!shown.ok) throw new Error("unreachable"); + expect(shown.result).toMatchObject({ + pendingAuth: true, + transport: "dormant", + }); + expect(server.registry.connectionCount()).toBe(1); + + // The read-only echoes annotate instead of dialing: tokens are on + // disk, so list/use report signed-in progress while the entry stays + // pending. + const listed = await server.handle({ + id: "l1", + op: "connections/list", + params: {}, + }); + expect(listed.ok).toBe(true); + if (!listed.ok) throw new Error("unreachable"); + expect( + (listed.result as { connections: unknown[] }).connections[0], + ).toMatchObject({ + pendingAuth: true, + pendingAuthSignedIn: true, + auth: { method: "oauth", authorized: true }, + }); + const used = await server.handle({ + id: "u1", + op: "connections/use", + params: { name: "p" }, + }); + expect(used.ok).toBe(true); + if (!used.ok) throw new Error("unreachable"); + expect(used.result).toMatchObject({ + pendingAuth: true, + pendingAuthSignedIn: true, + }); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + statusSpy.mockRestore(); + process.env.MCP_CLIENT_CONFIG_PATH = savedEnv.MCP_CLIENT_CONFIG_PATH; + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = + savedEnv.MCP_INSPECTOR_OAUTH_STATE_PATH; + resetNodeOAuthStorageCache(); + await server.stop().catch(() => {}); + fs.rmSync(stateDir, { recursive: true, force: true }); + } + }); + + it("connections/show carries the sign-in URL while the auth helper is live, and omits it otherwise", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const { pendingAuthMarkerPath } = + await import("../src/connection/auth-helper.js"); + const serverUrl = "https://mcp.example.com/mcp"; + const relayUrl = "https://as.example/authorize?client_id=abc&state=xyz"; + const markerDir = fs.mkdtempSync( + path.join(os.tmpdir(), "mcp-show-authurl-marker-"), + ); + const prevDaemonDir = process.env.MCP_INSPECTOR_DAEMON_DIR; + process.env.MCP_INSPECTOR_DAEMON_DIR = markerDir; + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockRejectedValue(Object.assign(new Error("boom"), { status: 401 })); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const statusSpy = vi + .spyOn(InspectorClient.prototype, "getStatus") + .mockReturnValue("disconnected"); + const server = new DaemonServer({ + dir: fs.mkdtempSync(path.join(os.tmpdir(), "mcp-show-authurl-daemon-")), + idleMs: 0, + }); + try { + await server.registry.connect({ + name: "p", + serverConfig: { type: "streamable-http", url: serverUrl }, + serverIdentity: serverUrl, + pendingOnAuthRequired: true, + }); + + // No marker yet: show reports pending but with no URL to relay. + const before = await server.handle({ + id: "s0", + op: "connections/show", + params: { name: "p" }, + }); + expect(before.ok).toBe(true); + if (!before.ok) throw new Error("unreachable"); + expect(before.result).toMatchObject({ pendingAuth: true }); + expect((before.result as { authUrl?: string }).authUrl).toBeUndefined(); + + // Helper publishes a live marker: show now carries the URL verbatim + // (query intact — it rides the result, which is not redacted). + fs.writeFileSync( + pendingAuthMarkerPath(serverUrl), + JSON.stringify({ + url: relayUrl, + pid: process.pid, + expiresAt: Date.now() + 60_000, + }), + { mode: 0o600 }, + ); + const after = await server.handle({ + id: "s1", + op: "connections/show", + params: { name: "p" }, + }); + expect(after.ok).toBe(true); + if (!after.ok) throw new Error("unreachable"); + expect((after.result as { authUrl?: string }).authUrl).toBe(relayUrl); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + statusSpy.mockRestore(); + if (prevDaemonDir === undefined) + delete process.env.MCP_INSPECTOR_DAEMON_DIR; + else process.env.MCP_INSPECTOR_DAEMON_DIR = prevDaemonDir; + await server.stop().catch(() => {}); + fs.rmSync(markerDir, { recursive: true, force: true }); + } + }); + + it("a connect that outlives shutdown's quiesce grace tears its client down instead of leaking it", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + let releaseConnect!: () => void; + const gate = new Promise<void>((resolve) => (releaseConnect = resolve)); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockImplementation(() => gate); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + try { + const params = { + name: "late", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + } as const; + const pending = registry.connect(params); + // Ensure the client is actually dialing before the shutdown snapshot + // runs — the bounded-quiesce-expired case (otherwise the entry check + // rejects it before a client exists). + await vi.waitFor(() => expect(connectSpy).toHaveBeenCalled()); + await registry.disconnectAll(); + releaseConnect(); + await expect(pending).rejects.toMatchObject({ + envelope: { code: "daemon_stopping" }, + }); + // The freshly connected client was disconnected, not registered. + expect(disconnectSpy).toHaveBeenCalledTimes(1); + expect(registry.connectionCount()).toBe(0); + // And a connect arriving after close fails fast, before dialing. + await expect(registry.connect(params)).rejects.toMatchObject({ + envelope: { code: "daemon_stopping" }, + }); + expect(connectSpy).toHaveBeenCalledTimes(1); // no second dial + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + } + }); + + it("idle timer does not fire while another connect is still in flight", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + let releaseConnect!: () => void; + const gate = new Promise<void>((resolve) => (releaseConnect = resolve)); + let dials = 0; + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockImplementation(() => { + dials++; + // First dial (the slow, valid connect) blocks on the gate; the + // second (a concurrent connect for a different name) fails. + return dials === 1 ? gate : Promise.reject(new Error("dial failed")); + }); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(25); + const onIdle = vi.fn(); + registry.setIdleHandler(onIdle); + try { + const config = { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + } as const; + const identity = "https://mcp.example.com/mcp"; + const slow = registry.connect({ + name: "slow", + serverConfig: config, + serverIdentity: identity, + }); + await vi.waitFor(() => expect(connectSpy).toHaveBeenCalled()); + await expect( + registry.connect({ + name: "fail", + serverConfig: config, + serverIdentity: identity, + }), + ).rejects.toThrow("dial failed"); + // The failed connect must not arm the idle timer while the valid + // connect is still in flight — the daemon would otherwise stop under + // it and fail it with daemon_stopping. + expect(registry.idleRemainingMs()).toBeNull(); + await new Promise((resolve) => setTimeout(resolve, 60)); + expect(onIdle).not.toHaveBeenCalled(); + releaseConnect(); + await expect(slow).resolves.toMatchObject({ name: "slow" }); + expect(registry.connectionCount()).toBe(1); + // Self-reaping still works once the registry actually empties. + await registry.disconnect("slow", false); + await vi.waitFor(() => expect(onIdle).toHaveBeenCalled()); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + } + }); + + it("rejects a connect whose caller is already gone (pre-aborted signal)", async () => { + const registry = new ConnectionRegistry(0); + const ac = new AbortController(); + ac.abort(); + const { command, args } = getTestMcpServerCommand(); + await expect( + registry.connect( + { + name: "gone", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }, + ac.signal, + ), + ).rejects.toThrow(/Connect cancelled/); + expect(registry.connectionCount()).toBe(0); + }); + + it("cancels an in-flight connect when the caller disconnects, tearing the client down", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + let releaseConnect!: () => void; + const gate = new Promise<void>((resolve) => (releaseConnect = resolve)); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockImplementation(() => gate); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + const registry = new ConnectionRegistry(0); + try { + const ac = new AbortController(); + const pending = registry.connect( + { + name: "slow", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }, + ac.signal, + ); + // Let the connect get in flight before hanging up. + const deadline = Date.now() + 3000; + while (connectSpy.mock.calls.length === 0 && Date.now() < deadline) { + await new Promise((resolve) => setImmediate(resolve)); + } + expect(connectSpy).toHaveBeenCalledTimes(1); + ac.abort(); + await expect(pending).rejects.toThrow(/Connect cancelled/); + // The abandoned client was torn down, not left dialing. + expect(disconnectSpy).toHaveBeenCalledTimes(1); + expect(registry.connectionCount()).toBe(0); + // The late settlement of the abandoned connect is observed by the + // cancellation race, so it never surfaces as an unhandled rejection. + releaseConnect(); + await new Promise((resolve) => setTimeout(resolve, 5)); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + } + }); + + it("does not start dialing when the abort lands during reconnect teardown", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const ac = new AbortController(); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValue(undefined); + // Abort while the reconnect path is tearing down the previous client — + // after the entry abort check, before the dial. AbortSignal does not + // replay, so only withAbort's synchronous pre-start check catches this. + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockImplementation(async () => { + ac.abort(); + }); + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue(undefined as never); + const registry = new ConnectionRegistry(0); + const params = { + name: "re", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + } as const; + try { + await registry.connect(params); + expect(connectSpy).toHaveBeenCalledTimes(1); + await expect(registry.connect(params, ac.signal)).rejects.toThrow( + /Connect cancelled/, + ); + // The second dial never started: the signal was checked synchronously + // before invoking connect, with no listener-install gap to hang in. + expect(connectSpy).toHaveBeenCalledTimes(1); + // Reconnect teardown plus the cancelled attempt's cleanup. + expect(disconnectSpy).toHaveBeenCalledTimes(2); + expect(registry.connectionCount()).toBe(0); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + } + }); + + it("discards a connect that completes only after the caller hung up", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const ac = new AbortController(); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValue(undefined); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + // Abort between connect settling and registration (the auth snapshot + // read sits exactly there), hitting the post-connect abort check. + const authSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockImplementation(async () => { + ac.abort(); + return undefined as never; + }); + const registry = new ConnectionRegistry(0); + try { + await expect( + registry.connect( + { + name: "late", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }, + ac.signal, + ), + ).rejects.toThrow(/Connect cancelled/); + // Connected fine — but registering would leak a connection nobody + // asked to keep, so it was disconnected instead. + expect(disconnectSpy).toHaveBeenCalledTimes(1); + expect(registry.connectionCount()).toBe(0); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + authSpy.mockRestore(); + } + }); + + it("reports the connect-time auth snapshot, and connections/show recomputes from disk", async () => { + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const { NodeOAuthStorage, resetNodeOAuthStorageCache } = + await import("@inspector/core/auth/node/storage-node.js"); + // Isolated client.json (EMA IdP config) + oauth.json so the show + // handler's disk read is deterministic. + const stateDir = fs.mkdtempSync( + path.join(os.tmpdir(), "mcp-conn-auth-info-"), + ); + const savedEnv = { + MCP_CLIENT_CONFIG_PATH: process.env.MCP_CLIENT_CONFIG_PATH, + MCP_INSPECTOR_OAUTH_STATE_PATH: + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH, + }; + process.env.MCP_CLIENT_CONFIG_PATH = path.join(stateDir, "client.json"); + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = path.join( + stateDir, + "oauth.json", + ); + const issuer = "https://idp.example.com"; + fs.writeFileSync( + process.env.MCP_CLIENT_CONFIG_PATH, + JSON.stringify({ + enterpriseManagedAuth: { + enabled: true, + idp: { issuer, clientId: "idp-client" }, + }, + }), + ); + resetNodeOAuthStorageCache(); + // Unexpired unsigned JWT so the seeded IdP session reads as logged_in. + const b64 = (obj: object) => + Buffer.from(JSON.stringify(obj)).toString("base64url"); + const idToken = `${b64({ alg: "none" })}.${b64({ + exp: Math.floor(Date.now() / 1000) + 3600, + })}.sig`; + + // Force an auth snapshot onto the connect result without a live OAuth + // server, so the auth-present reporting paths (connect result, list, + // use) are exercised. + const stateSpy = vi + .spyOn(InspectorClient.prototype, "getOAuthState") + .mockResolvedValue({ + authorized: true, + protocol: "ema", + serverUrl: "https://mcp.example.com/mcp", + ema: { + idpIssuer: issuer, + idpClientId: "idp-client", + idpSession: "logged_in", + }, + }); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockResolvedValue(undefined); + const server = new DaemonServer({ + dir: fs.mkdtempSync(path.join(os.tmpdir(), "mcp-conn-auth-daemon-")), + idleMs: 0, + }); + const registry = server.registry; + try { + const info = await registry.connect({ + name: "a", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverSettings: { + headers: [], + metadata: {}, + env: [], + connectionTimeout: 30_000, + requestTimeout: 0, + taskTtl: 0, + maxFetchRequests: 0, + autoRefreshOnListChanged: false, + paginatedLists: false, + roots: [], + enterpriseManaged: true, + }, + serverIdentity: "https://mcp.example.com/mcp", + }); + const expected = { + method: "ema", + authorized: true, + idpSession: "logged_in", + }; + expect(info.auth).toEqual(expected); + expect(registry.list()[0]?.auth).toEqual(expected); + expect(registry.use("a").auth).toEqual(expected); + + // connections/show reads *disk*, not the client's memory-cached storage: + // seed an IdP session on disk and expect logged_in (no tokens were + // persisted, so authorized is false — matching auth/ema-status). + await new NodeOAuthStorage().saveIdpSession(issuer, { + idToken, + idTokenExpiresAt: Date.now() + 3600_000, + }); + const shown = await server.handle({ + id: "show", + op: "connections/show", + params: { name: "a" }, + }); + expect(shown.ok).toBe(true); + if (!shown.ok) throw new Error("unreachable"); + expect((shown.result as { auth?: unknown }).auth).toEqual({ + method: "ema", + authorized: false, + idpSession: "logged_in", + }); + + // Simulate a cross-process logout (e.g. auth/ema-logout): clear the + // IdP session on disk. list keeps the connect-time value; show + // reflects the new disk state. + resetNodeOAuthStorageCache(); + await new NodeOAuthStorage().clearIdpSession(issuer); + expect(registry.list()[0]?.auth).toEqual(expected); + const loggedOut = await server.handle({ + id: "show2", + op: "connections/show", + params: { name: "a" }, + }); + expect(loggedOut.ok).toBe(true); + if (!loggedOut.ok) throw new Error("unreachable"); + expect((loggedOut.result as { auth?: unknown }).auth).toEqual({ + method: "ema", + authorized: false, + idpSession: "none", + }); + } finally { + await registry.disconnectAll(); + stateSpy.mockRestore(); + connectSpy.mockRestore(); + for (const [key, value] of Object.entries(savedEnv)) { + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + resetNodeOAuthStorageCache(); + fs.rmSync(stateDir, { recursive: true, force: true }); + } + }); +}); + +describe("DaemonServer IPC", () => { + let server: DaemonServer | undefined; + let dir: string | undefined; + + afterEach(async () => { + if (server) { + await server.stop("stop"); + server = undefined; + } + if (dir) { + fs.rmSync(dir, { recursive: true, force: true }); + dir = undefined; + } + }); + + it("serves ping / connect / connections/list / disconnect over the socket", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-")); + server = new DaemonServer({ dir, idleMs: 0 }); + await server.start(); + + const pong = await callDaemon<{ pong: boolean }>( + "ping", + {}, + { socketPath: server.socketPath }, + ); + expect(pong.pong).toBe(true); + + const { command, args } = getTestMcpServerCommand(); + const connected = await callDaemon<{ name: string; isMru: boolean }>( + "connect", + { + name: "stdio", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + expect(connected.name).toBe("stdio"); + expect(connected.isMru).toBe(true); + + const listed = await callDaemon<{ connections: { name: string }[] }>( + "connections/list", + {}, + { socketPath: server.socketPath }, + ); + expect(listed.connections.map((s) => s.name)).toEqual(["stdio"]); + + const status = await callDaemon<{ pid: number; socketPath: string }>( + "daemon/status", + {}, + { socketPath: server.socketPath }, + ); + expect(status.pid).toBe(process.pid); + expect(status.socketPath).toBe(server.socketPath); + + const disc = await callDaemon<{ name: string }>( + "disconnect", + { name: "stdio" }, + { socketPath: server.socketPath }, + ); + expect(disc.name).toBe("stdio"); + }); + + it("runs rpc tools/list and initialize against a live connection", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-rpc-")); + server = new DaemonServer({ dir, idleMs: 0 }); + await server.start(); + + const { command, args } = getTestMcpServerCommand(); + await callDaemon( + "connect", + { + name: "stdio", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + + const listed = await callDaemon<{ + kind: string; + result: { tools: unknown[] }; + }>( + "rpc", + { method: "tools/list", name: "stdio" }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + expect(listed.kind).toBe("result"); + expect(listed.result.tools.length).toBeGreaterThan(0); + + const init = await callDaemon<{ + kind: string; + result: { protocolVersion?: string }; + }>( + "rpc", + { method: "initialize", name: "stdio" }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + expect(init.kind).toBe("result"); + expect(init.result.protocolVersion).toBeTruthy(); + + await callDaemon( + "disconnect", + { name: "stdio" }, + { socketPath: server.socketPath }, + ); + }); + + it("rpc transparently revives a connection whose transport died (end-to-end)", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-revive-")); + server = new DaemonServer({ dir, idleMs: 0 }); + await server.start(); + + const { command, args } = getTestMcpServerCommand(); + await callDaemon( + "connect", + { + name: "stdio", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + + // Kill the daemon-held client's transport out from under the registry — + // the in-process equivalent of the server expiring the session (or a + // stdio child dying) overnight. + const registry = (server as unknown as { registry: ConnectionRegistry }) + .registry; + await registry.clientFor("stdio", false).disconnect(); + + // connections/show is passive for a non-pending entry: it reports the + // drop, no revive (out-of-band sign-in completion is the one case where + // show itself revives — covered separately). + const shown = await callDaemon<{ transport?: string }>( + "connections/show", + { name: "stdio" }, + { socketPath: server.socketPath }, + ); + expect(shown.transport).toBe("dormant"); + + // An actual op self-heals: fresh dial, real result — never "Tools (0)". + const listed = await callDaemon<{ + kind: string; + result: { tools: unknown[] }; + }>( + "rpc", + { method: "tools/list", name: "stdio" }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + expect(listed.kind).toBe("result"); + expect(listed.result.tools.length).toBeGreaterThan(0); + + const after = await callDaemon<{ transport?: string }>( + "connections/show", + { name: "stdio" }, + { socketPath: server.socketPath }, + ); + expect(after.transport).toBe("live"); + }); + + it("rejects stream methods on rpc and rpc methods on stream", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-ops-")); + server = new DaemonServer({ dir, idleMs: 0 }); + await server.start(); + + const { command, args } = getTestMcpServerCommand(); + await callDaemon( + "connect", + { + name: "stdio", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + + await expect( + callDaemon( + "rpc", + { method: "logging/tail", name: "stdio" }, + { socketPath: server.socketPath, timeoutMs: 5000 }, + ), + ).rejects.toMatchObject({ envelope: { code: "use_stream_op" } }); + + const badStream = await server.handleOutcome({ + id: "s1", + op: "stream", + params: { method: "tools/list", name: "stdio" }, + }); + expect(badStream.response.ok).toBe(false); + + const noMethod = await server.handle({ + id: "s2", + op: "rpc", + params: { name: "stdio" } as never, + }); + expect(noMethod.ok).toBe(false); + + await callDaemon( + "disconnect", + { name: "stdio" }, + { socketPath: server.socketPath }, + ); + }); + + it("ends an open stream when its connection disconnects", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-streamlife-")); + server = new DaemonServer({ dir, idleMs: 0 }); + const { command, args } = getTestMcpServerCommand(); + await server.registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }); + + const outcome = await server.handleOutcome({ + id: "st", + op: "stream", + params: { method: "logging/tail", name: "s" } as never, + }); + expect(outcome.response.ok).toBe(true); + expect(outcome.startStream).toBeDefined(); + let ended = 0; + const stop = outcome.startStream!( + () => {}, + () => { + ended += 1; + }, + ); + + // Disconnecting the named connection must terminate its streams instead + // of leaving the caller attached to a stale client until Ctrl-C. + await server.handle({ + id: "d", + op: "disconnect", + params: { name: "s" } as never, + }); + const deadline = Date.now() + 3000; + while (ended === 0 && Date.now() < deadline) { + await new Promise((resolve) => setImmediate(resolve)); + } + expect(ended).toBe(1); + stop(); + }); + + it("ends a stream whose connection disconnected before startStream ran", async () => { + // Regression: the statusChange listener was only installed inside + // startStream, which ipc-glue invokes after the ok frame. A disconnect + // completing in the window after runMethod() returned but before the + // listener existed lost the terminal event, leaving the stream open + // against a dead client forever. Terminal status is persistent state, + // so startStream now checks the current status after installing the + // listener. + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-streamrace-")); + server = new DaemonServer({ dir, idleMs: 0 }); + const { command, args } = getTestMcpServerCommand(); + await server.registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }); + + const outcome = await server.handleOutcome({ + id: "st", + op: "stream", + params: { method: "logging/tail", name: "s" } as never, + }); + expect(outcome.response.ok).toBe(true); + + // Disconnect in the window between runMethod() and startStream. + await server.handle({ + id: "d", + op: "disconnect", + params: { name: "s" } as never, + }); + + let ended = 0; + const stop = outcome.startStream!( + () => {}, + () => { + ended += 1; + }, + ); + expect(ended).toBe(1); + stop(); + }); + + it("serializes rpc ops per connection so elicitation routing is exact", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-rpcqueue-")); + server = new DaemonServer({ dir, idleMs: 0 }); + const { command, args } = getTestMcpServerCommand(); + await server.registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }); + + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + const order: string[] = []; + let release: () => void = () => {}; + const gate = new Promise<void>((resolve) => { + release = resolve; + }); + const spy = vi + .spyOn(InspectorClient.prototype, "readResource") + .mockImplementation(async (uri: string) => { + order.push(`start:${uri}`); + if (uri === "test://a") await gate; + order.push(`end:${uri}`); + return { result: { contents: [] } } as never; + }); + try { + const first = server.handle({ + id: "1", + op: "rpc", + params: { method: "resources/read", uri: "test://a", name: "s" }, + }); + const second = server.handle({ + id: "2", + op: "rpc", + params: { method: "resources/read", uri: "test://b", name: "s" }, + }); + const deadline = Date.now() + 3000; + while (!order.includes("start:test://a") && Date.now() < deadline) { + await new Promise((resolve) => setImmediate(resolve)); + } + await new Promise((resolve) => setTimeout(resolve, 30)); + // The second rpc must not have started while the first is in flight. + expect(order).toEqual(["start:test://a"]); + release(); + const [r1, r2] = await Promise.all([first, second]); + expect(r1.ok).toBe(true); + expect(r2.ok).toBe(true); + expect(order).toEqual([ + "start:test://a", + "end:test://a", + "start:test://b", + "end:test://b", + ]); + } finally { + spy.mockRestore(); + } + }); + + it("cancels a daemon-side connect end-to-end when the caller's socket closes", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-cancel-")); + server = new DaemonServer({ dir, idleMs: 0 }); + await server.start(); + + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + let releaseConnect!: () => void; + const gate = new Promise<void>((resolve) => (releaseConnect = resolve)); + const connectSpy = vi + .spyOn(InspectorClient.prototype, "connect") + .mockImplementation(() => gate); + const disconnectSpy = vi + .spyOn(InspectorClient.prototype, "disconnect") + .mockResolvedValue(undefined); + try { + // Full real-socket path: this is the seam test that unit tests on + // either side of the IPC adapter cannot cover (a dropped signal + // argument in the adapter would pass both and fail here). + const ac = new AbortController(); + const pending = callDaemon( + "connect", + { + name: "hung", + serverConfig: { + type: "streamable-http", + url: "https://mcp.example.com/mcp", + }, + serverIdentity: "https://mcp.example.com/mcp", + }, + { socketPath: server.socketPath, timeoutMs: 0, signal: ac.signal }, + ); + let deadline = Date.now() + 3000; + while (connectSpy.mock.calls.length === 0 && Date.now() < deadline) { + await new Promise((resolve) => setImmediate(resolve)); + } + expect(connectSpy).toHaveBeenCalledTimes(1); + // Frontend hangs up (Ctrl-C): its socket is destroyed… + ac.abort(); + await expect(pending).rejects.toThrow(/cancelled/); + // …and the daemon-side dial is torn down without ever completing. + deadline = Date.now() + 3000; + while (disconnectSpy.mock.calls.length === 0 && Date.now() < deadline) { + await new Promise((resolve) => setImmediate(resolve)); + } + expect(disconnectSpy).toHaveBeenCalledTimes(1); + expect(server.registry.connectionCount()).toBe(0); + // A late success is discarded, never registered. + releaseConnect(); + await new Promise((resolve) => setTimeout(resolve, 10)); + expect(server.registry.connectionCount()).toBe(0); + } finally { + connectSpy.mockRestore(); + disconnectSpy.mockRestore(); + } + }); + + it("strips format from rpc method args so JSON tool calls skip the app-info probe", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-daemon-fmt-")); + server = new DaemonServer({ dir, idleMs: 0 }); + const { command, args } = getTestMcpServerCommand(); + await server.registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "test-stdio", + }); + + const res = await server.handle({ + id: "1", + op: "rpc", + params: { + method: "tools/call", + name: "s", + toolName: "echo", + toolArg: { message: "hi" }, + format: "json", + }, + }); + expect(res.ok).toBe(true); + const rpc = (res as { result: { kind: string; appInfo?: unknown } }).result; + expect(rpc.kind).toBe("result"); + // format is a frontend-only output concern: forwarding it used to make + // runMethod collect app info (a hidden extra resources/read) whose + // result the frontend discards. + expect(rpc.appInfo).toBeUndefined(); + + // Explicit --app-info still probes. + const withApp = await server.handle({ + id: "2", + op: "rpc", + params: { + method: "tools/call", + name: "s", + toolName: "echo", + appInfo: true, + }, + }); + expect(withApp.ok).toBe(true); + const appRes = ( + withApp as { result: { result: { hasApp?: boolean; toolName?: string } } } + ).result; + expect(appRes.result).toMatchObject({ hasApp: false, toolName: "echo" }); + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-coverage.test.ts b/clients/mcpdo/__tests__/daemon-coverage.test.ts new file mode 100644 index 0000000000..6517986567 --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-coverage.test.ts @@ -0,0 +1,1292 @@ +import { describe, it, expect, afterEach, vi } from "vitest"; +import * as fs from "node:fs"; + +// Wrap renameSync in a pass-through vi.fn so the lock-reclaim race test can +// inject a concurrent contender in the read→rename window (ESM namespaces +// cannot be spied on directly). +vi.mock("node:fs", async (importOriginal) => { + const actual = await importOriginal<typeof import("node:fs")>(); + return { ...actual, renameSync: vi.fn(actual.renameSync) }; +}); +import * as net from "node:net"; +import * as os from "node:os"; +import * as path from "node:path"; +import { getTestMcpServerCommand } from "@modelcontextprotocol/inspector-test-server"; +import { DaemonServer } from "../src/daemon/server.js"; +import { callDaemon, daemonTokenDir } from "../src/daemon/client.js"; +import { + ensureDaemon, + readLogTail, + resolveDaemonScriptPath, + waitForDaemonExit, +} from "../src/daemon/ensure.js"; +import { spawn, spawnSync } from "node:child_process"; +import { ConnectionRegistry } from "../src/daemon/connections.js"; +import { CliExitCodeError } from "@inspector/core/cli/error-handler.js"; +import { runMcp } from "./helpers/mcp-runner.js"; +import { + createSampleTestConfig, + deleteConfigFile, +} from "../../cli/__tests__/helpers/fixtures.js"; +import { + expectCliSuccess, + expectCliFailure, +} from "../../cli/__tests__/helpers/assertions.js"; + +describe("daemon coverage", () => { + let server: DaemonServer | undefined; + let dir: string | undefined; + let configPath: string | undefined; + + afterEach(async () => { + if (server) { + await server.stop("stop"); + server = undefined; + } + if (dir) { + fs.rmSync(dir, { recursive: true, force: true }); + dir = undefined; + } + if (configPath) { + deleteConfigFile(configPath); + configPath = undefined; + } + }); + + function freshDir(): string { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-cov-")); + return dir; + } + + it("handle() covers invalid connect / connections/use / unknown op", async () => { + server = new DaemonServer({ dir: freshDir(), idleMs: 0 }); + const badConnect = await server.handle({ + id: "1", + op: "connect", + params: { name: "" } as never, + }); + expect(badConnect.ok).toBe(false); + if (!badConnect.ok) expect(badConnect.error.code).toBe("invalid_params"); + + const badUse = await server.handle({ + id: "2", + op: "connections/use", + params: {}, + }); + expect(badUse.ok).toBe(false); + + // connections/show with no `params` at all exercises the `request.params ?? + // {}` fallback; with no active connection it still fails, same shape as + // connections/use above. + const badShow = await server.handle({ id: "2b", op: "connections/show" }); + expect(badShow.ok).toBe(false); + + const unknown = await server.handle({ + id: "3", + op: "nope" as never, + }); + expect(unknown.ok).toBe(false); + if (!unknown.ok) expect(unknown.error.code).toBe("unknown_op"); + + // CliExitCodeError without an envelope → default code "cli_error". + const bare = new CliExitCodeError(1, "bare"); + vi.spyOn(server.registry, "list").mockImplementationOnce(() => { + throw bare; + }); + const listed = await server.handle({ id: "4", op: "connections/list" }); + expect(listed.ok).toBe(false); + if (!listed.ok) expect(listed.error.code).toBe("cli_error"); + + vi.spyOn(server.registry, "list").mockImplementationOnce(() => { + throw new Error("boom"); + }); + const boom = await server.handle({ id: "5", op: "connections/list" }); + expect(boom.ok).toBe(false); + // Non-CliExitCodeError failures go through classifyError (code "error"). + if (!boom.ok) expect(boom.error.code).toBe("error"); + + vi.spyOn(server.registry, "list").mockImplementationOnce(() => { + throw "string-throw"; + }); + const strErr = await server.handle({ id: "6", op: "connections/list" }); + expect(strErr.ok).toBe(false); + + const disc = await server.handle({ + id: "7", + op: "disconnect", + params: undefined, + }); + expect(disc.ok).toBe(false); + + // Defaults constructor + stop without onShutdown + re-entrant stop. + const plain = new DaemonServer({ dir: freshDir(), idleMs: 0 }); + await plain.start(); + await plain.stop("stop"); + await plain.stop("stop"); + + // Constructor default dir/idle/onShutdown branches (isolated storage dir). + const prev = process.env.MCP_INSPECTOR_DAEMON_DIR; + process.env.MCP_INSPECTOR_DAEMON_DIR = freshDir(); + try { + const defs = new DaemonServer(); + expect(defs.socketPath).toContain("mcpdod.sock"); + } finally { + if (prev === undefined) delete process.env.MCP_INSPECTOR_DAEMON_DIR; + else process.env.MCP_INSPECTOR_DAEMON_DIR = prev; + } + }); + + it("rejects a second daemon while a live one holds the lock", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + const other = new DaemonServer({ dir: d, idleMs: 0 }); + await expect(other.start()).rejects.toThrow(/held by running pid/); + }); + + it("reclaims a lock left by a dead pid", async () => { + const d = freshDir(); + // No live process can have this pid-space value in practice; write a + // plausible-but-dead pid by spawning nothing and using an exited child. + fs.writeFileSync(path.join(d, "mcpdod.lock"), "999999999\n"); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + expect(fs.readFileSync(path.join(d, "mcpdod.lock"), "utf8").trim()).toBe( + String(process.pid), + ); + }); + + it("does not delete a lock it no longer owns on stop", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + const lockPath = path.join(d, "mcpdod.lock"); + // Simulate a successor's lock at the same path (reclaim race / manual + // operator cleanup): release must be ownership-checked. + fs.writeFileSync(lockPath, "424242\n"); + await server.stop("stop"); + server = undefined; + expect(fs.readFileSync(lockPath, "utf8").trim()).toBe("424242"); + fs.unlinkSync(lockPath); + }); + + it("restores a live lock created between the dead-pid read and the rename", async () => { + const d = freshDir(); + const lockPath = path.join(d, "mcpdod.lock"); + fs.writeFileSync(lockPath, "999999999\n"); // dead pid + const actualFs = await vi.importActual<typeof import("node:fs")>("node:fs"); + vi.mocked(fs.renameSync).mockImplementationOnce((( + ...args: Parameters<typeof fs.renameSync> + ) => { + // Simulate a concurrent starter finishing its own reclaim + O_EXCL + // create in the window between readLockPid() and renameSync(). + fs.writeFileSync(lockPath, `${process.pid}\n`); + return actualFs.renameSync(...args); + }) as typeof fs.renameSync); + const contender = new DaemonServer({ dir: d, idleMs: 0 }); + await expect(contender.start()).rejects.toThrow(/held by running pid/); + // The stolen live lock was restored at the canonical path. + expect(fs.readFileSync(lockPath, "utf8").trim()).toBe(String(process.pid)); + expect(fs.existsSync(`${lockPath}.reclaim.${process.pid}`)).toBe(false); + }); + + it("treats a young pidless lock as held instead of stealing it", async () => { + const d = freshDir(); + const lockPath = path.join(d, "mcpdod.lock"); + // A concurrent starter between its O_EXCL create and its pid write. + fs.writeFileSync(lockPath, ""); + const contender = new DaemonServer({ dir: d, idleMs: 0 }); + await expect(contender.start()).rejects.toThrow(/Could not acquire/); + // The other starter's lock survived untouched. + expect(fs.readFileSync(lockPath, "utf8")).toBe(""); + }); + + it("reclaims a pidless lock older than the write grace period", async () => { + const d = freshDir(); + const lockPath = path.join(d, "mcpdod.lock"); + // A starter that died between create and pid write, long ago. + fs.writeFileSync(lockPath, ""); + const past = (Date.now() - 60_000) / 1000; + fs.utimesSync(lockPath, past, past); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + expect(fs.readFileSync(lockPath, "utf8").trim()).toBe(String(process.pid)); + }); + + it("restores a young pidless lock renamed aside mid-reclaim", async () => { + const d = freshDir(); + const lockPath = path.join(d, "mcpdod.lock"); + fs.writeFileSync(lockPath, "999999999\n"); // dead pid triggers the reclaim + const actualFs = await vi.importActual<typeof import("node:fs")>("node:fs"); + vi.mocked(fs.renameSync).mockImplementationOnce((( + ...args: Parameters<typeof fs.renameSync> + ) => { + actualFs.renameSync(...args); + // Simulate the renamed-aside file really belonging to a concurrent + // starter that created it but has not written its pid yet: empty, + // freshly touched. + actualFs.truncateSync(args[1] as string); + }) as typeof fs.renameSync); + const contender = new DaemonServer({ dir: d, idleMs: 0 }); + await expect(contender.start()).rejects.toThrow(/Could not acquire/); + // The lock was restored at the canonical path, not stolen. + expect(fs.existsSync(lockPath)).toBe(true); + expect(fs.existsSync(`${lockPath}.reclaim.${process.pid}`)).toBe(false); + }); + + it("stop() force-destroys sockets whose shutdown flush never drains", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0, flushTimeoutMs: 100 }); + await server.start(); + const client = net.connect(server.socketPath); + await new Promise<void>((resolve) => client.on("connect", () => resolve())); + client.pause(); + const ipcSockets = (server as unknown as { ipcSockets: Set<net.Socket> }) + .ipcSockets; + await vi.waitFor(() => expect(ipcSockets.size).toBe(1)); + // Backpressure: buffer a payload the paused client never reads, so + // destroySoon()'s drain never completes and server.close() would wait + // forever without the bounded force-destroy. + const [serverSocket] = ipcSockets; + serverSocket!.write("x".repeat(4 * 1024 * 1024)); + const stopped = await Promise.race([ + server.stop("stop").then(() => true), + new Promise<boolean>((r) => setTimeout(() => r(false), 5000)), + ]); + expect(stopped).toBe(true); + server = undefined; + client.destroy(); + }); + + it("removes a stale socket before binding", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + fs.writeFileSync(sock, ""); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + expect(fs.existsSync(sock)).toBe(true); + }); + + it("daemon/stop responds then shuts down", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + const result = await callDaemon<{ stopping: boolean }>( + "daemon/stop", + {}, + { socketPath: server.socketPath }, + ); + expect(result.stopping).toBe(true); + // Allow async stop to finish. + await new Promise((r) => setTimeout(r, 100)); + server = undefined; + }); + + it("accepts malformed NDJSON lines without crashing", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + await new Promise<void>((resolve, reject) => { + const socket = net.createConnection(server!.socketPath); + let data = ""; + socket.on("data", (chunk) => { + data += String(chunk); + if (data.includes("invalid_request")) { + socket.on("error", () => {}); + socket.end(); + resolve(); + } + }); + socket.on("error", reject); + socket.write("not-json\n"); + }); + }); + + it("resolves the mcpdod.token directory without deriving it from pipe paths", () => { + // Explicit dir always wins. + expect(daemonTokenDir({ dir: "/x", socketPath: "/y/mcpdod.sock" })).toBe( + "/x", + ); + // A Unix socket path implies its directory. + expect(daemonTokenDir({ socketPath: "/y/mcpdod.sock" })).toBe("/y"); + // A Windows named pipe has no meaningful dirname: fall back to the + // configured daemon directory, where the token is actually published. + const prev = process.env.MCP_INSPECTOR_DAEMON_DIR; + process.env.MCP_INSPECTOR_DAEMON_DIR = "/daemon/dir"; + try { + expect(daemonTokenDir({ socketPath: "\\\\.\\pipe\\mcp-conn-abc" })).toBe( + path.resolve("/daemon/dir"), + ); + expect(daemonTokenDir({})).toBe(path.resolve("/daemon/dir")); + } finally { + if (prev === undefined) delete process.env.MCP_INSPECTOR_DAEMON_DIR; + else process.env.MCP_INSPECTOR_DAEMON_DIR = prev; + } + }); + + it("callDaemon rejects immediately on a pre-aborted signal instead of hanging", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + const ac = new AbortController(); + ac.abort(); + // timeoutMs 0 = no timer: without the pre-aborted check (AbortSignal + // does not replay) this request would hang forever. + await expect( + callDaemon( + "ping", + {}, + { socketPath: server.socketPath, timeoutMs: 0, signal: ac.signal }, + ), + ).rejects.toThrow(/cancelled/); + }); + + it("callDaemon maps error responses and unreachable sockets", async () => { + await expect( + callDaemon( + "ping", + {}, + { socketPath: path.join(freshDir(), "missing.sock") }, + ), + ).rejects.toThrow(CliExitCodeError); + + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + await expect( + callDaemon("connections/use", {}, { socketPath: server.socketPath }), + ).rejects.toThrow(/requires a connection name/); + }); + + it("callDaemon reassembles a multi-byte UTF-8 character split across chunks", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const value = "héllo 👋 wörld"; + const splitter = net.createServer((socket) => { + socket.on("error", () => {}); + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + const payload = Buffer.from( + JSON.stringify({ id: req.id, ok: true, result: { value } }) + "\n", + "utf8", + ); + // Split mid-emoji (0xf0 opens the 4-byte sequence) so the two TCP + // chunks each carry half of one UTF-8 character. + const mid = payload.indexOf(0xf0) + 2; + socket.write(payload.subarray(0, mid)); + setTimeout(() => socket.write(payload.subarray(mid)), 20); + }); + }); + await new Promise<void>((resolve) => splitter.listen(sock, resolve)); + try { + const result = await callDaemon<{ value: string }>( + "ping", + {}, + { socketPath: sock, timeoutMs: 2000 }, + ); + expect(result.value).toBe(value); + } finally { + splitter.close(); + try { + fs.unlinkSync(sock); + } catch { + // ignore + } + } + }); + + it("callDaemon with an already-aborted signal rejects without dialing", async () => { + const d = freshDir(); + const ac = new AbortController(); + ac.abort(); + // A nonexistent socket path proves no connect is attempted: dialing it + // would fail with daemon_unreachable, not cancelled. + await expect( + callDaemon( + "ping", + {}, + { + socketPath: path.join(d, "absent.sock"), + signal: ac.signal, + timeoutMs: 2000, + }, + ), + ).rejects.toMatchObject({ envelope: { code: "cancelled" } }); + }); + + it("callDaemon rejects malformed response JSON", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const bad = net.createServer((socket) => { + socket.on("error", () => {}); + socket.write("not-json\n"); + }); + await new Promise<void>((resolve) => bad.listen(sock, resolve)); + try { + await expect( + callDaemon("ping", {}, { socketPath: sock, timeoutMs: 2000 }), + ).rejects.toThrow(); + } finally { + bad.close(); + try { + fs.unlinkSync(sock); + } catch { + // ignore + } + } + }); + + it("callDaemon ignores mismatched response ids then accepts a match", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const echo = net.createServer((socket) => { + socket.on("error", () => {}); + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ id: "other", ok: true, result: {} }) + "\n", + ); + socket.write( + JSON.stringify({ id: req.id, ok: true, result: { ok: true } }) + "\n", + ); + }); + }); + await new Promise<void>((resolve) => echo.listen(sock, resolve)); + try { + const result = await callDaemon<{ ok: boolean }>( + "ping", + {}, + { socketPath: sock, timeoutMs: 2000 }, + ); + expect(result.ok).toBe(true); + } finally { + echo.close(); + try { + fs.unlinkSync(sock); + } catch { + // ignore + } + } + }); + + it("callDaemon skips blank lines and defaults missing exitCode", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const echo = net.createServer((socket) => { + socket.on("error", () => {}); + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write("\n"); + socket.write( + JSON.stringify({ + id: req.id, + ok: false, + error: { code: "usage", message: "no exit" }, + }) + "\n", + ); + }); + }); + await new Promise<void>((resolve) => echo.listen(sock, resolve)); + try { + await expect( + callDaemon("ping", {}, { socketPath: sock, timeoutMs: 2000 }), + ).rejects.toMatchObject({ exitCode: 1 }); + } finally { + echo.close(); + try { + fs.unlinkSync(sock); + } catch { + // ignore + } + } + }); + + it("stop() without start and with missing lock files is safe", async () => { + const d = freshDir(); + const orphan = new DaemonServer({ dir: d, idleMs: 0 }); + await orphan.stop("stop"); + + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + fs.unlinkSync(server.socketPath); + fs.unlinkSync(path.join(d, "mcpdod.lock")); + await server.stop("stop"); + server = undefined; + }); + + it("repeated stop() returns the same in-flight cleanup promise", async () => { + // A second SIGINT used to see `stopping` and resolve immediately, + // letting its caller process.exit() mid-teardown and strand the + // socket/token/lock. Both calls must await the same cleanup. + server = new DaemonServer({ dir: freshDir(), idleMs: 0 }); + await server.start(); + const first = server.stop("signal"); + const second = server.stop("signal"); + expect(second).toBe(first); + await first; + expect(fs.existsSync(server.socketPath)).toBe(false); + server = undefined; + }); + + it("rejects new ops while stopping but keeps status ops answerable", async () => { + server = new DaemonServer({ dir: freshDir(), idleMs: 0 }); + let release!: () => void; + const gate = new Promise<void>((resolve) => (release = resolve)); + vi.spyOn(server.registry, "disconnectAll").mockImplementation(() => gate); + const stopP = server.stop("stop"); + + const rejected = await server.handle({ id: "q1", op: "connections/list" }); + expect(rejected.ok).toBe(false); + if (!rejected.ok) expect(rejected.error.code).toBe("daemon_stopping"); + + const status = await server.handle({ id: "q2", op: "daemon/status" }); + expect(status.ok).toBe(true); + const pong = await server.handle({ id: "q3", op: "ping" }); + expect(pong.ok).toBe(true); + + release(); + await stopP; + server = undefined; + }); + + it("stop() quiesces in-flight ops before disconnecting connections", async () => { + // A connect racing shutdown used to register its client *after* + // disconnectAll's snapshot, leaking a live child process. Shutdown now + // waits for in-flight ops so the late registration is included. + server = new DaemonServer({ dir: freshDir(), idleMs: 0 }); + const order: string[] = []; + let releaseConnect!: () => void; + const gate = new Promise<void>((resolve) => (releaseConnect = resolve)); + vi.spyOn(server.registry, "connect").mockImplementation(async () => { + order.push("connect:start"); + await gate; + order.push("connect:end"); + return { name: "a" } as never; + }); + vi.spyOn(server.registry, "disconnectAll").mockImplementation(async () => { + order.push("disconnectAll"); + }); + + const opP = server.handle({ + id: "c1", + op: "connect", + params: { + name: "a", + serverConfig: { type: "stdio", command: "x" }, + serverIdentity: "x", + } as never, + }); + await vi.waitFor(() => expect(order).toContain("connect:start")); + const stopP = server.stop("stop"); + await new Promise((resolve) => setTimeout(resolve, 20)); + expect(order).toEqual(["connect:start"]); // stop is waiting, not tearing down + + releaseConnect(); + await opP; + await stopP; + expect(order).toEqual(["connect:start", "connect:end", "disconnectAll"]); + server = undefined; + }); + + it("callDaemon times out a hung server", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const hung = net.createServer((socket) => { + socket.on("error", () => {}); + }); + await new Promise<void>((resolve) => hung.listen(sock, resolve)); + try { + await expect( + callDaemon("ping", {}, { socketPath: sock, timeoutMs: 100 }), + ).rejects.toThrow(/timed out/); + } finally { + hung.close(); + try { + fs.unlinkSync(sock); + } catch { + // ignore + } + } + }, 5000); + + it("callDaemon with timeoutMs 0 arms no local deadline; daemon close still fails it", async () => { + // rpc/connect callers pass 0 because the daemon enforces the configured + // MCP timeouts; a fixed 60s local timer falsely failed long tool calls. + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const sockets: net.Socket[] = []; + const silent = net.createServer((socket) => { + sockets.push(socket); + socket.on("error", () => {}); + }); + await new Promise<void>((resolve) => silent.listen(sock, resolve)); + try { + const pending = callDaemon( + "ping", + {}, + { socketPath: sock, timeoutMs: 0 }, + ); + let settled = false; + // void: observer only; the promise itself is asserted on below. + void pending.catch(() => (settled = true)).then(() => (settled = true)); + // Longer than the "times out a hung server" test's deadline: nothing + // fires locally. + await new Promise((resolve) => setTimeout(resolve, 200)); + expect(settled).toBe(false); + for (const socket of sockets) socket.destroy(); + await expect(pending).rejects.toThrow(/closed the connection/); + } finally { + silent.close(); + try { + fs.unlinkSync(sock); + } catch { + // ignore + } + } + }, 5000); + + it("connections/use and reconnect replace an existing connection", async () => { + const { command, args } = getTestMcpServerCommand(); + const registry = new ConnectionRegistry(0); + await registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "s", + }); + await registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "s-again", + }); + expect(registry.use("s").serverIdentity).toBe("s-again"); + expect(() => registry.resolve("missing", false)).toThrow(/isn't connected/); + await registry.disconnectAll(); + }); + + it("idle handler fires after last disconnect when idleMs > 0", async () => { + const registry = new ConnectionRegistry(20); + let idle = false; + registry.setIdleHandler(() => { + idle = true; + }); + const { command, args } = getTestMcpServerCommand(); + await registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "s", + }); + await registry.disconnect("s", false); + await new Promise((r) => setTimeout(r, 60)); + expect(idle).toBe(true); + expect(registry.idleRemainingMs()).toBeNull(); + }); + + it("covers touch/auth/oauth-setup/disconnect-swallow/reconnect-before-idle", async () => { + const { command, args } = getTestMcpServerCommand(); + const registry = new ConnectionRegistry(0); + registry.touch("missing"); + + await registry.connect({ + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "s", + }); + const connection = registry.resolve("s", false); + vi.spyOn(connection.client, "disconnect").mockRejectedValueOnce( + new Error("teardown boom"), + ); + await expect(registry.disconnect("s", false)).resolves.toEqual({ + name: "s", + }); + expect(registry.getMruName()).toBeNull(); + + const { AuthRecoveryRequiredError } = + await import("@inspector/core/auth/challenge.js"); + const { InspectorClient } = await import("@inspector/core/mcp/index.js"); + vi.spyOn(InspectorClient.prototype, "connect").mockRejectedValueOnce( + new AuthRecoveryRequiredError(new URL("https://as.example/authorize"), { + reason: "unauthorized", + }), + ); + await expect( + registry.connect({ + name: "auth", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "auth", + }), + ).rejects.toMatchObject({ exitCode: 3 }); + + // SDK token-exchange failure (empty redirectUrl / stale store) must surface + // as auth_required so the front-end can re-prompt — not a hard ErrorEnvelope. + vi.spyOn(InspectorClient.prototype, "connect").mockRejectedValueOnce( + new Error( + "Either provider.prepareTokenRequest() or authorizationCode is required", + ), + ); + await expect( + registry.connect({ + name: "reauth", + serverConfig: { + type: "streamable-http", + url: "https://example.com/mcp", + }, + serverIdentity: "reauth", + }), + ).rejects.toMatchObject({ + exitCode: 3, + envelope: { code: "auth_required" }, + }); + + await expect( + registry.connect({ + name: "http", + serverConfig: { + type: "streamable-http", + url: "http://127.0.0.1:1/mcp", + }, + serverIdentity: "http", + }), + ).rejects.toThrow(); + + const idleReg = new ConnectionRegistry(80); + const onIdle = vi.fn(); + idleReg.setIdleHandler(onIdle); + await idleReg.connect({ + name: "a", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "a", + }); + await idleReg.disconnect("a", false); + await idleReg.connect({ + name: "b", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "b", + }); + await new Promise((r) => setTimeout(r, 100)); + expect(onIdle).not.toHaveBeenCalled(); + await idleReg.disconnectAll(); + }, 20000); + + it("ensureDaemon reuses a running daemon and resolveDaemonScriptPath finds build", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0 }); + await server.start(); + const ensured = await ensureDaemon({ + dir: d, + daemonScript: resolveDaemonScriptPath(), + }); + expect(ensured.spawned).toBe(false); + expect(ensured.socketPath).toBe(server.socketPath); + }); + + it("ensureDaemon auto-spawns when no daemon is present", async () => { + const d = freshDir(); + const ensured = await ensureDaemon({ + dir: d, + daemonScript: resolveDaemonScriptPath(), + }); + expect(ensured.spawned).toBe(true); + await callDaemon("daemon/stop", {}, { socketPath: ensured.socketPath }); + await new Promise((r) => setTimeout(r, 150)); + }); + + it("ping and daemon/status report stopping during shutdown", async () => { + const d = freshDir(); + const srv = new DaemonServer({ dir: d, idleMs: 0 }); + await srv.start(); + const before = await srv.handle({ id: "p1", op: "ping" }); + expect(before).toMatchObject({ ok: true, result: { stopping: false } }); + await srv.stop("stop"); + // Ping never fails — a stopping (or just-stopped in-process) daemon + // still answers, flagged so ensureDaemon knows to wait it out. + const after = await srv.handle({ id: "p2", op: "ping" }); + expect(after).toMatchObject({ + ok: true, + result: { pong: true, stopping: true }, + }); + const status = await srv.handle({ id: "s1", op: "daemon/status" }); + expect(status).toMatchObject({ ok: true, result: { stopping: true } }); + }); + + it("waitForDaemonExit resolves for a dead pid and times out on a live one", async () => { + const dead = spawnSync(process.execPath, ["-e", ""]); + expect(dead.pid).toBeGreaterThan(0); + await waitForDaemonExit(dead.pid, "/nonexistent.sock", 2000, 10); + await expect( + waitForDaemonExit(process.pid, "/nonexistent.sock", 150, 25), + ).rejects.toMatchObject({ envelope: { code: "daemon_stopping" } }); + }); + + it("waitForDaemonExit falls back to socket reachability without a pid", async () => { + const d = freshDir(); + await waitForDaemonExit(undefined, path.join(d, "absent.sock"), 500, 10); + }); + + it("ensureDaemon waits out a stopping daemon and spawns a fresh one", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + // The "old daemon": a process that takes a moment to exit, and a socket + // that answers ping with stopping:true (as the real dispatch does). + const oldDaemon = spawn( + process.execPath, + ["-e", "setTimeout(()=>{},800)"], + { + stdio: "ignore", + }, + ); + const stoppingSocket = net.createServer((socket) => { + socket.on("error", () => {}); + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ + id: req.id, + ok: true, + result: { pong: true, pid: oldDaemon.pid, stopping: true }, + }) + "\n", + ); + }); + }); + await new Promise<void>((r) => stoppingSocket.listen(sock, r)); + // Mid-shutdown the socket goes away before the process does — the gap + // where spawning too early would die on the still-held lock. + setTimeout(() => { + stoppingSocket.close(); + try { + fs.unlinkSync(sock); + } catch { + // already gone + } + }, 200); + try { + const ensured = await ensureDaemon({ + dir: d, + daemonScript: resolveDaemonScriptPath(), + }); + expect(ensured.spawned).toBe(true); + await callDaemon("daemon/stop", {}, { socketPath: ensured.socketPath }); + await new Promise((r) => setTimeout(r, 150)); + } finally { + oldDaemon.kill(); + } + }, 15000); + + it("start-timeout error quotes the daemon's stderr log", async () => { + const d = freshDir(); + // A "daemon" that logs a failure and dies without ever binding a socket + // — the silent-death case the 0600 log exists to explain. + const script = path.join(d, "dying-daemon.js"); + fs.writeFileSync( + script, + 'console.error("boom: could not start"); setTimeout(() => {}, 3000);\n', + ); + await expect( + ensureDaemon({ dir: d, daemonScript: script, readyTimeoutMs: 700 }), + ).rejects.toMatchObject({ + envelope: { code: "daemon_start_timeout" }, + message: expect.stringContaining("boom: could not start"), + }); + }, 15000); + + it("start-timeout error stays clean when the daemon logged nothing", async () => { + const d = freshDir(); + const script = path.join(d, "silent-daemon.js"); + fs.writeFileSync(script, "setTimeout(() => {}, 3000);\n"); + await expect( + ensureDaemon({ dir: d, daemonScript: script, readyTimeoutMs: 700 }), + ).rejects.toMatchObject({ + envelope: { code: "daemon_start_timeout" }, + message: expect.not.stringContaining("Daemon log"), + }); + }, 15000); + + it("readLogTail returns the last lines and empty string when unreadable", () => { + const d = freshDir(); + const logPath = path.join(d, "mcpdod.log"); + const lines = Array.from({ length: 15 }, (_, i) => `line-${i}`); + fs.writeFileSync(logPath, lines.join("\n") + "\n"); + const tail = readLogTail(logPath); + expect(tail.split("\n")).toHaveLength(10); + expect(tail).toContain("line-14"); + expect(tail).not.toContain("line-4\n"); + expect(readLogTail(path.join(d, "missing.log"))).toBe(""); + }); + + it("ensureDaemon fails loudly when a live listener rejects ping (no takeover)", async () => { + // Regression test for the daemon-takeover hole: a socket that ACCEPTS + // connections is owned by a live process. ensureDaemon must never unlink + // it and install a replacement daemon — it must surface the ping failure. + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const occupant = net.createServer((socket) => { + socket.on("error", () => {}); + socket.end(); + }); + await new Promise<void>((resolve) => occupant.listen(sock, resolve)); + try { + await expect( + ensureDaemon({ dir: d, daemonScript: resolveDaemonScriptPath() }), + ).rejects.toThrow(/closed the connection during 'ping'/); + // The occupant's socket must still be in place, untouched. + expect(fs.existsSync(sock)).toBe(true); + } finally { + occupant.close(); + } + }); + + it("ensureDaemon fails loudly on daemon_auth_failed (wrong token is not a stale socket)", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 0, requiredToken: "good" }); + await server.start(); + await expect( + ensureDaemon({ + dir: d, + daemonScript: resolveDaemonScriptPath(), + token: "wrong", + }), + ).rejects.toMatchObject({ envelope: { code: "daemon_auth_failed" } }); + // The live daemon keeps its socket and still serves the right token. + const pong = await callDaemon( + "ping", + {}, + { + socketPath: server.socketPath, + token: "good", + }, + ); + expect(pong).toBeDefined(); + }); + + it("connection-less start arms idle and self-reaps", async () => { + const d = freshDir(); + let shut = false; + server = new DaemonServer({ + dir: d, + idleMs: 40, + onShutdown: () => { + shut = true; + }, + }); + await server.start(); + // ensureDaemon from tools/list with no connections must not leak forever. + expect(server.registry.idleRemainingMs()).not.toBeNull(); + await new Promise((r) => setTimeout(r, 100)); + expect(shut).toBe(true); + server = undefined; + }); + + it("connect failure for a dead stdio command is surfaced and re-arms idle", async () => { + const registry = new ConnectionRegistry(5_000); + let idle = false; + registry.setIdleHandler(() => { + idle = true; + }); + await expect( + registry.connect({ + name: "dead", + serverConfig: { + type: "stdio", + command: path.join(os.tmpdir(), "no-such-mcp-server-binary"), + args: [], + }, + serverIdentity: "dead", + }), + ).rejects.toThrow(); + expect(registry.idleRemainingMs()).not.toBeNull(); + expect(idle).toBe(false); + }); + + it("re-arms idle when createConnectionClient fails before client.connect", async () => { + const registry = new ConnectionRegistry(5_000); + registry.setIdleHandler(() => {}); + const prev = process.env.MCP_OAUTH_CALLBACK_URL; + process.env.MCP_OAUTH_CALLBACK_URL = "https://example.com/oauth/callback"; + try { + await expect( + registry.connect({ + name: "http", + serverConfig: { + type: "streamable-http", + url: "http://127.0.0.1:1/mcp", + }, + serverIdentity: "http", + }), + ).rejects.toThrow(/http scheme|callback URL/i); + expect(registry.idleRemainingMs()).not.toBeNull(); + } finally { + if (prev === undefined) delete process.env.MCP_OAUTH_CALLBACK_URL; + else process.env.MCP_OAUTH_CALLBACK_URL = prev; + } + }); + + it("callDaemon fails immediately when the peer closes without a response", async () => { + const d = freshDir(); + const sock = path.join(d, "mcpdod.sock"); + const peer = net.createServer((socket) => { + socket.on("error", () => {}); + // Accept then FIN with no NDJSON reply. + socket.end(); + }); + await new Promise<void>((resolve) => peer.listen(sock, resolve)); + try { + await expect( + callDaemon("ping", {}, { socketPath: sock, timeoutMs: 60_000 }), + ).rejects.toMatchObject({ + envelope: { code: "daemon_unreachable" }, + }); + } finally { + peer.close(); + try { + fs.unlinkSync(sock); + } catch { + // ignore + } + } + }); + + it("connections/use via handle and blank IPC lines", async () => { + const d = freshDir(); + server = new DaemonServer({ dir: d, idleMs: 60_000 }); + await server.start(); + const { command, args } = getTestMcpServerCommand(); + await callDaemon( + "connect", + { + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "s", + }, + { socketPath: server.socketPath, timeoutMs: 15000 }, + ); + const used = await server.handle({ + id: "u", + op: "connections/use", + params: { name: "s" }, + }); + expect(used.ok).toBe(true); + expect(server.registry.idleRemainingMs()).toBeNull(); + + // connections/show over the same live connection — exercises the full case + // body (serverInfo/protocolVersion/protocolEra/capabilities lookups) + // in-process, where coverage instrumentation can see it. + const shown = await server.handle({ + id: "s2", + op: "connections/show", + params: { name: "s" }, + }); + expect(shown.ok).toBe(true); + if (shown.ok) { + const result = shown.result as { protocolVersion?: string }; + expect(result.protocolVersion).toBeTruthy(); + } + + await new Promise<void>((resolve, reject) => { + const socket = new net.Socket(); + socket.on("error", reject); + socket.connect(server!.socketPath, () => { + socket.write("\n\n"); + socket.end(); + resolve(); + }); + }); + + await callDaemon( + "disconnect", + { name: "s" }, + { socketPath: server.socketPath }, + ); + // Idle timer armed — remaining countdown is positive and ≤ configured idleMs. + const remaining = server.registry.idleRemainingMs(); + expect(remaining).not.toBeNull(); + expect(remaining!).toBeLessThanOrEqual(60_000); + expect(remaining!).toBeGreaterThan(0); + }); +}); + +describe("mcp connection coverage", () => { + let configPath: string | undefined; + let storageDir: string | undefined; + + afterEach(async () => { + if (storageDir) { + const socketPath = path.join(storageDir, "mcpdod.sock"); + if (fs.existsSync(socketPath)) { + try { + await callDaemon("daemon/stop", {}, { socketPath, timeoutMs: 2000 }); + } catch { + // ignore + } + const deadline = Date.now() + 2000; + while (fs.existsSync(socketPath) && Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 50)); + } + } + fs.rmSync(storageDir, { recursive: true, force: true }); + storageDir = undefined; + } + if (configPath) { + deleteConfigFile(configPath); + configPath = undefined; + } + }); + + function env(): Record<string, string> { + storageDir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-sess-cov-")); + return { + MCP_STORAGE_DIR: storageDir, + MCP_INSPECTOR_DAEMON_DIR: storageDir, + MCP_ALLOW_DEFAULT_CONNECTION: "1", + }; + } + + it("covers connections/use, daemon status, @connection connect, and stop no-op", async () => { + configPath = createSampleTestConfig(); + const e = env(); + + const stopIdle = await runMcp(["daemon", "stop", "--format", "json"], { + env: e, + }); + expectCliSuccess(stopIdle); + expect(stopIdle.stdout).toContain("not running"); + + const connected = await runMcp( + [ + "connect", + "@alpha", + "test-stdio", + "--config", + configPath, + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(connected); + expect(JSON.parse(connected.stdout).name).toBe("alpha"); + + const used = await runMcp( + ["connections/use", "@alpha", "--format", "text"], + { + env: e, + }, + ); + expectCliSuccess(used); + expect(used.stdout).toContain("alpha"); + + const status = await runMcp(["daemon", "status"], { env: e }); + expectCliSuccess(status); + + const listed = await runMcp(["connections/list"], { env: e }); + expectCliSuccess(listed); + + const viaServer = await runMcp( + [ + "connect", + "--server", + "test-stdio", + "--config", + configPath, + "--connection", + "via-flag", + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(viaServer); + + const stopped = await runMcp(["daemon", "stop", "--format", "json"], { + env: e, + }); + expectCliSuccess(stopped); + expect(stopped.stdout).toContain("stopping"); + }); + + it("rejects connect with no target and invalid --format", async () => { + const e = env(); + const missing = await runMcp(["connect"], { env: e }); + expectCliFailure(missing); + + const badFormat = await runMcp(["servers/list", "--format", "xml"], { + env: e, + }); + expectCliFailure(badFormat); + + const badTransport = await runMcp(["connect", "x", "--transport", "ftp"], { + env: e, + }); + expectCliFailure(badTransport); + + const badTimeout = await runMcp( + ["connect", "x", "--connect-timeout", "-1"], + { env: e }, + ); + expectCliFailure(badTimeout); + + const emptyUse = await runMcp(["connections/use", ""], { env: e }); + expectCliFailure(emptyUse); + }); + + it("connects an ad-hoc stdio target", async () => { + const { command, args } = getTestMcpServerCommand(); + const e = env(); + // Multi-token positional target → ad-hoc (not a catalog entry name). + const result = await runMcp( + [ + "connect", + "--connection", + "adhoc", + "--transport", + "stdio", + "--format", + "json", + command, + ...args, + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(result); + expect(JSON.parse(result.stdout).name).toBe("adhoc"); + }); + + it("treats a URL positional as ad-hoc", async () => { + const e = env(); + const result = await runMcp( + [ + "connect", + "http://127.0.0.1:9/mcp", + "--connection", + "url", + "--connect-timeout", + "100", + "--format", + "json", + ], + { env: e, timeout: 10000 }, + ); + // Connection should fail (nothing listening) but the ad-hoc URL path ran. + expectCliFailure(result); + }); + + it("requires explicit connection in non-interactive mode without opt-in", async () => { + configPath = createSampleTestConfig(); + storageDir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-sess-ci-")); + const e = { + MCP_STORAGE_DIR: storageDir, + MCP_INSPECTOR_DAEMON_DIR: storageDir, + // no MCP_ALLOW_DEFAULT_CONNECTION + }; + const connected = await runMcp( + ["connect", "test-stdio", "--config", configPath, "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(connected); + + // Force requireExplicit by stubbing isTTY false is default in vitest forks. + const disc = await runMcp(["disconnect", "--format", "json"], { env: e }); + expectCliFailure(disc); + expect(disc.stderr).toMatch(/Explicit|--connection|non-interactive/i); + + await runMcp(["disconnect", "--connection", "test-stdio"], { env: e }); + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-elicitation-park.test.ts b/clients/mcpdo/__tests__/daemon-elicitation-park.test.ts new file mode 100644 index 0000000000..52add1bf37 --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-elicitation-park.test.ts @@ -0,0 +1,697 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { DaemonServer } from "../src/daemon/server.js"; +import { + ElicitationParkRegistry, + ParkingElicitationChannel, +} from "../src/daemon/elicitation-park.js"; +import type { + ElicitationPendingInfo, + ElicitationRespondResult, + RpcResult, +} from "../src/daemon/protocol.js"; +import type { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; + +/** + * Covers daemon-side elicitation parking (dual-era support, phase 2): + * `rpc` from a non-interactive caller (`interactive: false`) returning + * `elicitation-pending` instead of relaying an inline prompt, `elicitation/respond` resuming the parked call + * (final result, error, or the next round), expiry, the + * one-parked-call-per-connection guard, and the registry/channel primitives. + */ + +const runMethodMock = vi.hoisted(() => ({ + impl: undefined as unknown as (...args: unknown[]) => Promise<unknown>, +})); +vi.mock("@inspector/core/cli/handlers/run-method.js", () => ({ + runMethod: (...args: unknown[]) => runMethodMock.impl(...args), +})); + +type FakeElicitationMessage = { + id: string; + origin: string; + request: { method: string; params: Record<string, unknown> }; + respond: ReturnType<typeof vi.fn>; + cancel: ReturnType<typeof vi.fn>; +}; + +function deferred<T>() { + let resolve!: (value: T) => void; + let reject!: (error: unknown) => void; + const promise = new Promise<T>((res, rej) => { + resolve = res; + reject = rej; + }); + return { promise, resolve, reject }; +} + +const FORM_SCHEMA = { + type: "object", + properties: { color: { type: "string" } }, + required: ["color"], +}; + +function makeFormMessage(id: string): { + message: FakeElicitationMessage; + answered: Promise<{ action: string; content?: Record<string, unknown> }>; + cancelled: Promise<void>; +} { + const answer = deferred<{ + action: string; + content?: Record<string, unknown>; + }>(); + const cancel = deferred<void>(); + const message: FakeElicitationMessage = { + id, + origin: "server-request", + request: { + method: "elicitation/create", + params: { message: "Pick a color", requestedSchema: FORM_SCHEMA }, + }, + respond: vi.fn(async (response) => { + answer.resolve(response as never); + }), + cancel: vi.fn(() => { + cancel.resolve(); + answer.resolve({ action: "cancel" }); + }), + }; + return { message, answered: answer.promise, cancelled: cancel.promise }; +} + +function makeUrlMessage(id: string): { + message: FakeElicitationMessage; + answered: Promise<{ action: string; content?: Record<string, unknown> }>; +} { + const answer = deferred<{ + action: string; + content?: Record<string, unknown>; + }>(); + const message: FakeElicitationMessage = { + id, + origin: "server-request", + request: { + method: "elicitation/create", + params: { + message: "Finish signup", + url: "https://example.com/signup?flow=abc", + }, + }, + respond: vi.fn(async (response) => { + answer.resolve(response as never); + }), + cancel: vi.fn(() => answer.resolve({ action: "cancel" })), + }; + return { message, answered: answer.promise }; +} + +function fakeClient(): { client: InspectorClient; emit: (m: unknown) => void } { + const target = new EventTarget(); + const client = { + addEventListener: (type: string, listener: EventListener) => + target.addEventListener(type, listener), + removeEventListener: (type: string, listener: EventListener) => + target.removeEventListener(type, listener), + getStatus: () => "connected", + cancelToolCall: () => false, + } as unknown as InspectorClient; + return { + client, + emit: (detail) => + target.dispatchEvent( + new CustomEvent("newPendingElicitation", { detail }), + ), + }; +} + +describe("daemon elicitation parking", () => { + let dir: string; + let server: DaemonServer; + let client: InspectorClient; + let emit: (m: unknown) => void; + + beforeEach(() => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-elicit-park-")); + server = new DaemonServer({ dir, idleMs: 0 }); + const fake = fakeClient(); + client = fake.client; + emit = fake.emit; + const registry = server.registry as unknown as Record<string, unknown>; + registry.connectionFor = () => ({ name: "srv", client }); + registry.liveClientFor = async () => client; + }); + + afterEach(() => { + fs.rmSync(dir, { recursive: true, force: true }); + vi.restoreAllMocks(); + }); + + function rpcCallTool(id: string) { + return server.handle({ + id, + op: "rpc", + params: { + method: "tools/call", + toolName: "collect", + name: "srv", + interactive: false, + }, + }); + } + + function respond( + id: string, + params: Record<string, unknown>, + ): ReturnType<DaemonServer["handle"]> { + return server.handle({ id, op: "elicitation/respond", params }); + } + + it("parks a form elicitation, then respond accept resumes to the final result", async () => { + const { message, answered } = makeFormMessage("elicit-1"); + runMethodMock.impl = async () => { + emit(message); + const answer = await answered; + return { + kind: "result", + result: { echoed: answer.content, action: answer.action }, + }; + }; + + const first = await rpcCallTool("r1"); + expect(first.ok).toBe(true); + const pending = (first as { result: RpcResult }).result; + expect(pending.kind).toBe("elicitation-pending"); + const info = (pending as { elicitation: ElicitationPendingInfo }) + .elicitation; + expect(info).toMatchObject({ + elicitationId: "elicit-1", + connection: "srv", + method: "tools/call", + toolName: "collect", + mode: "form", + message: "Pick a color", + requestedSchema: FORM_SCHEMA, + origin: "server-request", + }); + expect(info.expiresAt).toBeGreaterThan(Date.now()); + + const second = await respond("r2", { + elicitationId: "elicit-1", + action: "accept", + content: { color: "teal" }, + }); + expect(second.ok).toBe(true); + const result = (second as { result: ElicitationRespondResult }).result; + expect(result.method).toBe("tools/call"); + expect(result.toolName).toBe("collect"); + expect(result.outcome).toEqual({ + kind: "result", + result: { echoed: { color: "teal" }, action: "accept" }, + appInfo: undefined, + }); + expect(message.respond).toHaveBeenCalledWith({ + action: "accept", + content: { color: "teal" }, + }); + }); + + it("chains rounds: respond returns the next pending elicitation, then the result", async () => { + const round1 = makeFormMessage("elicit-a"); + const round2 = makeFormMessage("elicit-b"); + runMethodMock.impl = async () => { + emit(round1.message); + await round1.answered; + emit(round2.message); + const answer = await round2.answered; + return { kind: "result", result: { final: answer.content } }; + }; + + const first = await rpcCallTool("r1"); + expect((first as { result: RpcResult }).result.kind).toBe( + "elicitation-pending", + ); + + const mid = await respond("r2", { + elicitationId: "elicit-a", + action: "accept", + content: { color: "red" }, + }); + expect(mid.ok).toBe(true); + const midOutcome = (mid as { result: ElicitationRespondResult }).result + .outcome; + expect(midOutcome.kind).toBe("elicitation-pending"); + const nextId = (midOutcome as { elicitation: ElicitationPendingInfo }) + .elicitation.elicitationId; + expect(nextId).toBe("elicit-b"); + // The answered round's id is no longer respondable. + const stale = await respond("r3", { + elicitationId: "elicit-a", + action: "cancel", + }); + expect(stale.ok).toBe(false); + expect((stale as { error: { code: string } }).error.code).toBe( + "elicitation_not_found", + ); + + const done = await respond("r4", { + elicitationId: "elicit-b", + action: "accept", + content: { color: "blue" }, + }); + expect(done.ok).toBe(true); + expect( + (done as { result: ElicitationRespondResult }).result.outcome, + ).toMatchObject({ kind: "result", result: { final: { color: "blue" } } }); + }); + + it("relays decline and cancel; url mode accepts --done and rejects decline/content", async () => { + // decline (form) + const declineRound = makeFormMessage("elicit-d"); + runMethodMock.impl = async () => { + emit(declineRound.message); + const answer = await declineRound.answered; + return { kind: "result", result: { action: answer.action } }; + }; + await rpcCallTool("r1"); + const declined = await respond("r2", { + elicitationId: "elicit-d", + action: "decline", + }); + expect( + (declined as { result: ElicitationRespondResult }).result.outcome, + ).toMatchObject({ kind: "result", result: { action: "decline" } }); + expect(declineRound.message.respond).toHaveBeenCalledWith({ + action: "decline", + content: undefined, + }); + + // url mode + const urlRound = makeUrlMessage("elicit-u"); + runMethodMock.impl = async () => { + emit(urlRound.message); + const answer = await urlRound.answered; + return { kind: "result", result: { action: answer.action } }; + }; + const parked = await rpcCallTool("r3"); + const info = ( + (parked as { result: RpcResult }).result as { + elicitation: ElicitationPendingInfo; + } + ).elicitation; + expect(info.mode).toBe("url"); + expect(info.url).toBe("https://example.com/signup?flow=abc"); + + const badDecline = await respond("r4", { + elicitationId: "elicit-u", + action: "decline", + }); + expect(badDecline.ok).toBe(false); + expect((badDecline as { error: { code: string } }).error.code).toBe( + "invalid_params", + ); + const badContent = await respond("r5", { + elicitationId: "elicit-u", + action: "accept", + content: { nope: 1 }, + }); + expect(badContent.ok).toBe(false); + + // Validation failures put the entry back — a corrected accept still works. + const done = await respond("r6", { + elicitationId: "elicit-u", + action: "accept", + }); + expect(done.ok).toBe(true); + expect( + (done as { result: ElicitationRespondResult }).result.outcome, + ).toMatchObject({ kind: "result", result: { action: "accept" } }); + expect(urlRound.message.respond).toHaveBeenCalledWith({ + action: "accept", + content: undefined, + }); + }); + + it("returns the plain result when a parked-mode call never elicits, and propagates failures", async () => { + runMethodMock.impl = async () => ({ kind: "result", result: { n: 1 } }); + const plain = await rpcCallTool("r1"); + expect((plain as { result: RpcResult }).result).toMatchObject({ + kind: "result", + result: { n: 1 }, + }); + + runMethodMock.impl = async () => { + throw new Error("server exploded"); + }; + const failed = await rpcCallTool("r2"); + expect(failed.ok).toBe(false); + expect((failed as { error: { message: string } }).error.message).toContain( + "server exploded", + ); + }); + + it("propagates a failure that lands after the elicitation was answered", async () => { + const round = makeFormMessage("elicit-f"); + runMethodMock.impl = async () => { + emit(round.message); + await round.answered; + throw new Error("tool failed after input"); + }; + await rpcCallTool("r1"); + const failed = await respond("r2", { + elicitationId: "elicit-f", + action: "accept", + content: { color: "red" }, + }); + expect(failed.ok).toBe(false); + expect((failed as { error: { message: string } }).error.message).toContain( + "tool failed after input", + ); + }); + + it("rejects new rpcs on a connection with a parked call", async () => { + const round = makeFormMessage("elicit-g"); + runMethodMock.impl = async () => { + emit(round.message); + await round.answered; + return { kind: "result", result: {} }; + }; + await rpcCallTool("r1"); + const blocked = await server.handle({ + id: "r2", + op: "rpc", + params: { method: "tools/list", name: "srv" }, + }); + expect(blocked.ok).toBe(false); + expect((blocked as { error: { code: string } }).error.code).toBe( + "elicitation_pending", + ); + expect((blocked as { error: { message: string } }).error.message).toContain( + "elicitation/respond elicit-g", + ); + // Unblock: cancel it. + const cancelled = await respond("r3", { + elicitationId: "elicit-g", + action: "cancel", + }); + expect(cancelled.ok).toBe(true); + }); + + it("expires an unanswered parked elicitation and cancels the message", async () => { + server = new DaemonServer({ dir, idleMs: 0, elicitationTtlMs: 40 }); + const registry = server.registry as unknown as Record<string, unknown>; + registry.connectionFor = () => ({ name: "srv", client }); + registry.liveClientFor = async () => client; + + const round = makeFormMessage("elicit-x"); + runMethodMock.impl = async () => { + emit(round.message); + await round.answered; + return { kind: "result", result: {} }; + }; + const parked = await rpcCallTool("r1"); + expect((parked as { result: RpcResult }).result.kind).toBe( + "elicitation-pending", + ); + await round.cancelled; + expect(round.message.cancel).toHaveBeenCalled(); + const late = await respond("r2", { + elicitationId: "elicit-x", + action: "accept", + content: { color: "red" }, + }); + expect(late.ok).toBe(false); + expect((late as { error: { code: string } }).error.code).toBe( + "elicitation_not_found", + ); + }); + + it("expired park cannot swallow a new call's elicitation (stale subscriber unwired)", async () => { + server = new DaemonServer({ dir, idleMs: 0, elicitationTtlMs: 40 }); + const registry = server.registry as unknown as Record<string, unknown>; + registry.connectionFor = () => ({ name: "srv", client }); + registry.liveClientFor = async () => client; + + // First call parks, then its park expires while the server-side call + // keeps running (the server never gives up) — so its outcome `finally` + // (the normal unwire point) has not fired. + const round1 = makeFormMessage("elicit-stale"); + const neverSettles = deferred<never>(); + runMethodMock.impl = async () => { + emit(round1.message); + await neverSettles.promise; + return { kind: "result", result: {} }; + }; + const parked1 = await rpcCallTool("r1"); + expect((parked1 as { result: RpcResult }).result.kind).toBe( + "elicitation-pending", + ); + await round1.cancelled; // expiry fired + + // A new call in that window: its elicitation must reach ITS subscriber, + // not the expired park's closed channel (which would auto-cancel it). + const round2 = makeFormMessage("elicit-fresh"); + runMethodMock.impl = async () => { + emit(round2.message); + await round2.answered; + return { kind: "result", result: { ok: true } }; + }; + const parked2 = await rpcCallTool("r2"); + expect((parked2 as { result: RpcResult }).result.kind).toBe( + "elicitation-pending", + ); + expect(round2.message.cancel).not.toHaveBeenCalled(); + const done = await respond("r3", { + elicitationId: "elicit-fresh", + action: "accept", + content: { color: "red" }, + }); + expect(done.ok).toBe(true); + }); + + it("expiry cancels the abandoned server call so it can't emit a second elicitation", async () => { + server = new DaemonServer({ dir, idleMs: 0, elicitationTtlMs: 40 }); + const registry = server.registry as unknown as Record<string, unknown>; + registry.connectionFor = () => ({ name: "srv", client }); + registry.liveClientFor = async () => client; + const cancelToolCall = vi.spyOn(client, "cancelToolCall"); + + // The server-side call never settles on its own, so without an explicit + // cancel it would keep running past expiry and could emit another + // elicitation — which, with this park's subscriber unwired, the bridge + // would misroute to whatever rpc started next on this connection. + const round = makeFormMessage("elicit-cancel"); + const neverSettles = deferred<never>(); + runMethodMock.impl = async () => { + emit(round.message); + await neverSettles.promise; + return { kind: "result", result: {} }; + }; + const parked = await rpcCallTool("r1"); + expect((parked as { result: RpcResult }).result.kind).toBe( + "elicitation-pending", + ); + await round.cancelled; // expiry fired + expect(cancelToolCall).toHaveBeenCalled(); + }); + + it("disconnect cancels the parked call; respond then reports not found", async () => { + const registry = server.registry as unknown as Record<string, unknown>; + registry.disconnect = async () => ({ name: "srv" }); + + const round = makeFormMessage("elicit-z"); + runMethodMock.impl = async () => { + emit(round.message); + await round.answered; + return { kind: "result", result: {} }; + }; + await rpcCallTool("r1"); + const gone = await server.handle({ + id: "r2", + op: "disconnect", + params: { name: "srv" }, + }); + expect(gone.ok).toBe(true); + expect(round.message.cancel).toHaveBeenCalled(); + const late = await respond("r3", { + elicitationId: "elicit-z", + action: "cancel", + }); + expect(late.ok).toBe(false); + expect((late as { error: { code: string } }).error.code).toBe( + "elicitation_not_found", + ); + }); + + it("picks up a call that settled on its own while parked (server gave up waiting)", async () => { + const round = makeFormMessage("elicit-s"); + runMethodMock.impl = async () => { + emit(round.message); + // Server-side timeout: the call completes without our answer. + return { kind: "result", result: { timedOut: true } }; + }; + const parked = await rpcCallTool("r1"); + expect((parked as { result: RpcResult }).result.kind).toBe( + "elicitation-pending", + ); + const done = await respond("r2", { + elicitationId: "elicit-s", + action: "accept", + content: { color: "red" }, + }); + expect(done.ok).toBe(true); + expect( + (done as { result: ElicitationRespondResult }).result.outcome, + ).toMatchObject({ kind: "result", result: { timedOut: true } }); + }); + + it("validates respond params", async () => { + const missing = await respond("r1", { action: "accept" }); + expect((missing as { error: { code: string } }).error.code).toBe( + "invalid_params", + ); + const badAction = await respond("r2", { + elicitationId: "x", + action: "shrug", + }); + expect((badAction as { error: { code: string } }).error.code).toBe( + "invalid_params", + ); + const unknown = await respond("r3", { + elicitationId: "nope", + action: "cancel", + }); + expect((unknown as { error: { code: string } }).error.code).toBe( + "elicitation_not_found", + ); + }); +}); + +describe("ParkingElicitationChannel / ElicitationParkRegistry primitives", () => { + const frame = (elicitationId: string) => + ({ + id: "req-1", + kind: "elicitation-request", + elicitationId, + mode: "form", + message: "hi", + origin: "server-request", + }) as const; + + it("answer() is a no-op with nothing pending; request after close rejects", async () => { + const channel = new ParkingElicitationChannel(); + channel.answer({ + id: "req-1", + kind: "elicitation-response", + elicitationId: "none", + action: "cancel", + }); + channel.close(new Error("gone")); + await expect(channel.request(frame("later"))).rejects.toThrow("gone"); + }); + + it("close rejects a pending request and clears the waiter", async () => { + const channel = new ParkingElicitationChannel(); + const pending = channel.request(frame("e1")); + expect(channel.pendingFrame()?.elicitationId).toBe("e1"); + channel.close(new Error("teardown")); + await expect(pending).rejects.toThrow("teardown"); + expect(channel.pendingFrame()).toBeNull(); + }); + + it("waitForElicitation resolves immediately when a request is already pending", async () => { + const channel = new ParkingElicitationChannel(); + const pending = channel.request(frame("e1")); + const seen = await channel.waitForElicitation(); + expect(seen.elicitationId).toBe("e1"); + channel.close(new Error("teardown")); + await expect(pending).rejects.toThrow("teardown"); + }); + + it("forClient and cancelForConnection ignore non-matching entries", async () => { + const registry = new ElicitationParkRegistry(0); + const channel = new ParkingElicitationChannel(); + const pending = channel.request(frame("e1")); + const client = {} as InspectorClient; + registry.add({ + info: { + elicitationId: "e1", + connection: "srv", + method: "tools/call", + mode: "form", + message: "hi", + origin: "server-request", + }, + client, + channel, + outcome: new Promise<never>(() => {}), + unwire: () => {}, + cancelCall: () => {}, + }); + expect(registry.forClient({} as InspectorClient)).toBeUndefined(); + expect(registry.forClient(client)).toBeDefined(); + // A different connection's teardown must not cancel this parked call. + registry.cancelForConnection("other"); + expect(registry.forClient(client)).toBeDefined(); + registry.cancelAll(); + await expect(pending).rejects.toThrow(/going away/); + }); + + it("cancelAll settles every parked entry", async () => { + const registry = new ElicitationParkRegistry(0); + const channel = new ParkingElicitationChannel(); + const pending = channel.request(frame("e1")); + const unwire = vi.fn(); + const cancelCall = vi.fn(); + registry.add({ + info: { + elicitationId: "e1", + connection: "srv", + method: "tools/call", + mode: "form", + message: "hi", + origin: "server-request", + }, + client: {} as InspectorClient, + channel, + outcome: new Promise<never>(() => {}), + unwire, + cancelCall, + }); + registry.cancelAll(); + await expect(pending).rejects.toThrow(/going away/); + expect(() => registry.take("e1")).toThrow(/No pending elicitation/); + // Cancel must also unwire the bridge subscriber of the abandoned call. + expect(unwire).toHaveBeenCalled(); + // ...and cancel the still-running server call so it can't emit a second + // elicitation that misroutes to a later rpc on this connection. + expect(cancelCall).toHaveBeenCalled(); + }); + + it("expiry unwires and cancels the abandoned call", async () => { + const registry = new ElicitationParkRegistry(20); + const channel = new ParkingElicitationChannel(); + const pending = channel.request(frame("e2")); + const unwire = vi.fn(); + const cancelCall = vi.fn(); + registry.add({ + info: { + elicitationId: "e2", + connection: "srv", + method: "tools/call", + mode: "form", + message: "hi", + origin: "server-request", + }, + client: {} as InspectorClient, + channel, + outcome: new Promise<never>(() => {}), + unwire, + cancelCall, + }); + await expect(pending).rejects.toThrow(/expired/); + expect(unwire).toHaveBeenCalled(); + expect(cancelCall).toHaveBeenCalled(); + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-ipc-glue.test.ts b/clients/mcpdo/__tests__/daemon-ipc-glue.test.ts new file mode 100644 index 0000000000..3bc30bc45b --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-ipc-glue.test.ts @@ -0,0 +1,343 @@ +/** + * Unit tests for `acceptDaemonConnection`'s per-connection wiring: the + * elicitation channel, destroyed-socket guards, and stream cleanup. A fake + * in-memory Duplex stands in for the net.Socket so every path is exercised + * deterministically (no accept/connect races). + */ +import { describe, it, expect } from "vitest"; +import { Duplex } from "node:stream"; +import type * as net from "node:net"; +import { + acceptDaemonConnection, + MAX_REQUEST_LINE_BYTES, + type ElicitationChannel, +} from "../src/daemon/ipc-glue.js"; +import type { + DaemonRequest, + ElicitationRequestFrame, + ElicitationResponseFrame, +} from "../src/daemon/protocol.js"; + +class FakeSocket extends Duplex { + written: string[] = []; + override _read(): void {} + override _write( + chunk: unknown, + _encoding: BufferEncoding, + callback: (error?: Error | null) => void, + ): void { + this.written.push(String(chunk)); + callback(); + } + pushLine(line: string): void { + this.push(line + "\n"); + } + get all(): string { + return this.written.join(""); + } +} + +function accept( + handle: ( + request: DaemonRequest, + elicitation: ElicitationChannel, + signal?: AbortSignal, + ) => Promise<{ + response: { id: string; ok: true; result: unknown }; + startStream?: ( + writeData: (data: unknown) => void, + endStream: () => void, + ) => () => void; + }>, +): FakeSocket { + const socket = new FakeSocket(); + acceptDaemonConnection(socket as unknown as net.Socket, handle); + return socket; +} + +/** Await an event-driven condition (no fixed sleeps). */ +async function until(condition: () => boolean): Promise<void> { + while (!condition()) { + await new Promise((resolve) => setImmediate(resolve)); + } +} + +const REQUEST = JSON.stringify({ id: "r1", op: "rpc", params: {} }); + +function elicitationRequest(id: string): ElicitationRequestFrame { + return { + id, + kind: "elicitation-request", + elicitationId: `elicit-${id}`, + mode: "form", + message: "pick one", + origin: "server-request", + }; +} + +function elicitationResponse(id: string): ElicitationResponseFrame { + return { + id, + kind: "elicitation-response", + elicitationId: `elicit-${id}`, + action: "accept", + content: {}, + }; +} + +describe("acceptDaemonConnection elicitation channel", () => { + it("pauses a call for an elicitation exchange and resumes on the answer", async () => { + const socket = accept(async (request, elicitation) => { + const answer = await elicitation.request(elicitationRequest("e1")); + return { + response: { id: request.id, ok: true, result: { action: answer } }, + }; + }); + + socket.pushLine(REQUEST); + await until(() => socket.all.includes('"elicitation-request"')); + + socket.pushLine(JSON.stringify(elicitationResponse("e1"))); + await until(() => socket.all.includes('"ok":true')); + expect(socket.all).toContain('"action"'); + }); + + it("ignores non-answer lines while an exchange is pending", async () => { + const socket = accept(async (request, elicitation) => { + const answer = await elicitation.request(elicitationRequest("e2")); + return { response: { id: request.id, ok: true, result: answer } }; + }); + + socket.pushLine(REQUEST); + await until(() => socket.all.includes('"elicitation-request"')); + + // None of these are elicitation answers; each falls through to the + // request parser and earns an invalid_request response. + socket.pushLine("not-json"); + socket.pushLine("null"); + socket.pushLine(JSON.stringify({ kind: "other" })); + await until(() => socket.all.split('"invalid_request"').length - 1 === 3); + + socket.pushLine(JSON.stringify(elicitationResponse("e2"))); + await until(() => socket.all.includes('"ok":true')); + }); + + it("rejects a second exchange while one is already pending", async () => { + let secondError: Error | undefined; + const socket = accept(async (request, elicitation) => { + const first = elicitation.request(elicitationRequest("e3")); + await elicitation + .request(elicitationRequest("e4")) + .catch((error: Error) => { + secondError = error; + }); + socket.pushLine(JSON.stringify(elicitationResponse("e3"))); + await first; + return { response: { id: request.id, ok: true, result: {} } }; + }); + + socket.pushLine(REQUEST); + await until(() => socket.all.includes('"ok":true')); + expect(secondError?.message).toMatch(/already pending/); + }); + + it("rejects a pending exchange when the connection drops", async () => { + let rejection: Error | undefined; + const settled = { done: false }; + const socket = accept(async (request, elicitation) => { + elicitation.request(elicitationRequest("e5")).catch((error: Error) => { + rejection = error; + settled.done = true; + }); + return { response: { id: request.id, ok: true, result: {} } }; + }); + + socket.pushLine(REQUEST); + await until(() => socket.all.includes('"elicitation-request"')); + socket.destroy(); + await until(() => settled.done); + expect(rejection?.message).toMatch(/Connection closed/); + }); + + it("rejects immediately when the socket is already destroyed", async () => { + let rejection: Error | undefined; + const settled = { done: false }; + let channel: ElicitationChannel | undefined; + const socket = accept(async (request, elicitation) => { + channel = elicitation; + return { response: { id: request.id, ok: true, result: {} } }; + }); + + socket.pushLine(REQUEST); + await until(() => socket.all.includes('"ok":true')); + socket.destroy(); + await until(() => socket.destroyed); + await channel!.request(elicitationRequest("e6")).catch((error: Error) => { + rejection = error; + settled.done = true; + }); + expect(settled.done).toBe(true); + expect(rejection?.message).toMatch(/Connection closed/); + }); +}); + +describe("acceptDaemonConnection guards", () => { + it("aborts the per-request signal when the caller's socket closes", async () => { + let seen: AbortSignal | undefined; + let release: () => void = () => {}; + const gate = new Promise<void>((resolve) => { + release = resolve; + }); + const socket = accept(async (request, _elicitation, signal) => { + seen = signal; + await gate; + return { response: { id: request.id, ok: true, result: {} } }; + }); + + socket.pushLine(REQUEST); + await until(() => seen !== undefined); + // Caller still attached: nothing aborted. + expect(seen!.aborted).toBe(false); + socket.destroy(); + await until(() => seen!.aborted === true); + release(); + }); + + it("drops the response when the socket dies mid-handle", async () => { + let release: () => void = () => {}; + const gate = new Promise<void>((resolve) => { + release = resolve; + }); + const handled = { done: false }; + const socket = accept(async (request) => { + await gate; + handled.done = true; + return { response: { id: request.id, ok: true, result: {} } }; + }); + + socket.pushLine(REQUEST); + socket.destroy(); + await until(() => socket.destroyed); + release(); + await until(() => handled.done); + // One more tick for the post-await destroyed guard. + await new Promise((resolve) => setImmediate(resolve)); + expect(socket.all).toBe(""); + }); + + it("disposes an opened stream when the socket died mid-handle", async () => { + let release: () => void = () => {}; + const gate = new Promise<void>((resolve) => { + release = resolve; + }); + let started = 0; + let stops = 0; + const socket = accept(async (request) => { + await gate; + return { + response: { id: request.id, ok: true, result: {} }, + // e.g. resources/subscribe: producer-side state exists before the + // starter runs; the glue must start it inert and stop it so the + // daemon doesn't keep a hidden subscription with no consumer. + startStream: (writeData, endStream) => { + started += 1; + // The inert writer/end are safe to call: nothing reaches the wire. + writeData({ n: 1 }); + endStream(); + return () => { + stops += 1; + }; + }, + }; + }); + + socket.pushLine(REQUEST); + socket.destroy(); + await until(() => socket.destroyed); + release(); + await until(() => stops === 1); + expect(started).toBe(1); + expect(socket.all).toBe(""); + }); + + it("ends the stream when the producer invokes endStream", async () => { + let end: () => void = () => {}; + let stops = 0; + const socket = accept(async (request) => ({ + response: { id: request.id, ok: true, result: {} }, + startStream: (writeData, endStream) => { + end = endStream; + writeData({ n: 1 }); + return () => { + stops += 1; + }; + }, + })); + + socket.pushLine(REQUEST); + await until(() => socket.all.includes('"stream":"data"')); + + end(); + await until(() => socket.all.includes('"stream":"end"')); + expect(stops).toBe(1); + + // A duplicate end (or a later close event) does not double-stop. + end(); + socket.emit("close"); + await new Promise((resolve) => setImmediate(resolve)); + expect(stops).toBe(1); + }); + + it("cleans up a stream once on socket error and ignores late writes", async () => { + let lateWrite: (data: unknown) => void = () => {}; + let stops = 0; + const socket = accept(async (request) => ({ + response: { id: request.id, ok: true, result: {} }, + startStream: (writeData) => { + lateWrite = writeData; + writeData({ n: 1 }); + return () => { + stops += 1; + }; + }, + })); + + socket.pushLine(REQUEST); + await until(() => socket.all.includes('"stream":"data"')); + + // Error on a still-writable socket: cleanup must emit the end frame, + // half-close, and the readline teardown must not throw. + socket.emit("error", new Error("peer reset")); + await until(() => socket.all.includes('"stream":"end"')); + expect(stops).toBe(1); + + // Late writes after cleanup are no-ops, and a duplicate cleanup + // (close after error) does not double-stop. + lateWrite({ n: 2 }); + socket.emit("close"); + await new Promise((resolve) => setImmediate(resolve)); + expect(stops).toBe(1); + expect(socket.all.split('"stream":"data"').length - 1).toBe(1); + }); +}); + +describe("request line cap", () => { + it("never hands a terminated oversized line to the handler", async () => { + // Deterministic cross-chunk variant of the e2e cap tests: a valid JSON + // request padded past the cap, split so the chunk that crosses the limit + // also carries the terminating newline. Both the byte accounting and the + // post-reject line guard must hold, or the handler sees the request. + let handled = 0; + const socket = accept(async (request) => { + handled += 1; + return { response: { id: request.id, ok: true, result: {} } }; + }); + const padded = + REQUEST + " ".repeat(MAX_REQUEST_LINE_BYTES + 1024 - REQUEST.length); + socket.push(padded.slice(0, 600 * 1024)); + socket.push(padded.slice(600 * 1024) + "\n"); + await until(() => socket.destroyed); + await new Promise((resolve) => setImmediate(resolve)); + expect(handled).toBe(0); + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-paths.test.ts b/clients/mcpdo/__tests__/daemon-paths.test.ts new file mode 100644 index 0000000000..e8f51e85be --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-paths.test.ts @@ -0,0 +1,183 @@ +import { describe, it, expect, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { + assertSocketPathWithinLimit, + assertTrustedPrivateRoot, + createPrivateDaemonDir, + ensureDaemonDir, + getDaemonDir, + getDaemonLockPath, + getDaemonSocketPath, +} from "../src/daemon/paths.js"; +import { writeFormattedResult } from "@inspector/core/cli/handlers/format-output.js"; + +describe("daemon paths", () => { + const backup: Record<string, string | undefined> = {}; + + afterEach(() => { + for (const key of [ + "MCP_INSPECTOR_DAEMON_DIR", + "MCP_STORAGE_DIR", + "HOME", + "TMPDIR", + ]) { + if (key in backup) { + if (backup[key] === undefined) delete process.env[key]; + else process.env[key] = backup[key]; + delete backup[key]; + } + } + }); + + function setEnv(key: string, value: string | undefined) { + backup[key] = process.env[key]; + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + + it("prefers MCP_INSPECTOR_DAEMON_DIR over MCP_STORAGE_DIR", () => { + const a = path.join(os.tmpdir(), "daemon-a"); + const b = path.join(os.tmpdir(), "daemon-b"); + setEnv("MCP_STORAGE_DIR", b); + setEnv("MCP_INSPECTOR_DAEMON_DIR", a); + expect(getDaemonDir()).toBe(path.resolve(a)); + expect(getDaemonSocketPath()).toBe( + path.join(path.resolve(a), "mcpdod.sock"), + ); + expect(getDaemonLockPath()).toBe(path.join(path.resolve(a), "mcpdod.lock")); + }); + + it("falls back to MCP_STORAGE_DIR then ~/.mcp-inspector", () => { + const storage = path.join(os.tmpdir(), "daemon-storage"); + setEnv("MCP_INSPECTOR_DAEMON_DIR", undefined); + setEnv("MCP_STORAGE_DIR", storage); + expect(getDaemonDir()).toBe(path.resolve(storage)); + setEnv("MCP_STORAGE_DIR", undefined); + expect(getDaemonDir()).toContain(".mcp-inspector"); + }); + + it("creates the daemon directory", () => { + const dir = path.join(os.tmpdir(), `daemon-mkdir-${Date.now()}`); + ensureDaemonDir(dir); + expect(fs.statSync(dir).isDirectory()).toBe(true); + fs.rmSync(dir, { recursive: true, force: true }); + }); + + it("tightens a pre-existing loose daemon directory to 0700 and rejects symlinks", () => { + // mkdirSync never re-modes an existing dir; ~/.mcp-inspector commonly + // pre-exists at 0755, so ensureDaemonDir must tighten it itself. + const base = fs.mkdtempSync(path.join(os.tmpdir(), "daemon-tighten-")); + const loose = path.join(base, "loose"); + fs.mkdirSync(loose, { mode: 0o755 }); + fs.chmodSync(loose, 0o755); + ensureDaemonDir(loose); + expect(fs.statSync(loose).mode & 0o077).toBe(0); + + const target = path.join(base, "target"); + fs.mkdirSync(target, { mode: 0o700 }); + const link = path.join(base, "link"); + fs.symlinkSync(target, link); + expect(() => ensureDaemonDir(link)).toThrow(/not a directory/); + fs.rmSync(base, { recursive: true, force: true }); + }); + + it("createPrivateDaemonDir nests under a short 0700 tmpdir layout", () => { + const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-conn-t-")); + setEnv("TMPDIR", tmp + path.sep); + const dir = createPrivateDaemonDir(); + // $TMPDIR/mcp-conn-<uid>/<8-hex>; short enough that mcpdod.sock stays inside + // the platform sun_path limit even for macOS /var/folders tmpdirs. + expect(dir.startsWith(tmp)).toBe(true); + expect(path.basename(dir)).toMatch(/^[0-9a-f]{8}$/); + expect(path.basename(path.dirname(dir))).toMatch(/^mcp-conn-/); + expect(fs.statSync(dir).isDirectory()).toBe(true); + if (process.platform !== "win32") { + expect(fs.statSync(dir).mode & 0o777).toBe(0o700); + expect(fs.statSync(path.dirname(dir)).mode & 0o777).toBe(0o700); + } + fs.rmSync(tmp, { recursive: true, force: true }); + }); + + it("createPrivateDaemonDir refuses a symlinked mcp-conn root", () => { + if (process.platform === "win32") return; + const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-conn-sym-")); + setEnv("TMPDIR", tmp + path.sep); + // Another user pre-planting the predictable root as a symlink to a dir + // they control must fail closed, not be adopted by recursive mkdir. + const target = path.join(tmp, "attacker-controlled"); + fs.mkdirSync(target, { mode: 0o700 }); + fs.symlinkSync(target, path.join(tmp, `mcp-conn-${process.getuid!()}`)); + expect(() => createPrivateDaemonDir()).toThrow(/not a directory/); + fs.rmSync(tmp, { recursive: true, force: true }); + }); + + it("assertTrustedPrivateRoot tightens a loose pre-existing root", () => { + if (process.platform === "win32") return; + const tmp = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-conn-loose-")); + const root = path.join(tmp, "root"); + fs.mkdirSync(root, { mode: 0o755 }); + assertTrustedPrivateRoot(root); + expect(fs.statSync(root).mode & 0o777).toBe(0o700); + // A file in the root's place fails closed too. + const file = path.join(tmp, "not-a-dir"); + fs.writeFileSync(file, ""); + expect(() => assertTrustedPrivateRoot(file)).toThrow(/not a directory/); + fs.rmSync(tmp, { recursive: true, force: true }); + }); + + it("assertSocketPathWithinLimit rejects paths over the sun_path limit", () => { + expect(() => + assertSocketPathWithinLimit("/tmp/short/mcpdod.sock"), + ).not.toThrow(); + const long = "/" + "x".repeat(150) + "/mcpdod.sock"; + expect(() => assertSocketPathWithinLimit(long)).toThrow( + /too long for this platform/, + ); + }); + + it("uses a deterministic named-pipe path on Windows with no sun_path limit", () => { + const realPlatform = Object.getOwnPropertyDescriptor( + process, + "platform", + ) as PropertyDescriptor; + Object.defineProperty(process, "platform", { value: "win32" }); + try { + const pipe = getDaemonSocketPath("/some/daemon/dir"); + expect(pipe).toMatch(/^\\\\\.\\pipe\\mcp-conn-[0-9a-f]{16}$/); + // Same dir (any casing) -> same pipe; different dir -> different pipe. + expect(getDaemonSocketPath("/SOME/DAEMON/DIR")).toBe(pipe); + expect(getDaemonSocketPath("/other/daemon/dir")).not.toBe(pipe); + // Pipe names are not sun_path-constrained. + const long = "\\\\.\\pipe\\" + "x".repeat(300); + expect(() => assertSocketPathWithinLimit(long)).not.toThrow(); + } finally { + Object.defineProperty(process, "platform", realPlatform); + } + }); +}); + +describe("writeFormattedResult", () => { + it("writes text and json envelopes", async () => { + let out = ""; + const original = process.stdout.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]) => { + out += String(chunk); + const cb = rest.find((x) => typeof x === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + try { + await writeFormattedResult({ ok: 1 }, "text"); + expect(out).toContain('"ok": 1'); + out = ""; + await writeFormattedResult({ ok: 2 }, "json"); + expect(JSON.parse(out)).toEqual({ result: { ok: 2 } }); + } finally { + process.stdout.write = original; + } + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-private.test.ts b/clients/mcpdo/__tests__/daemon-private.test.ts new file mode 100644 index 0000000000..0c67b63cf7 --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-private.test.ts @@ -0,0 +1,345 @@ +import { describe, it, expect, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { getTestMcpServerCommand } from "@modelcontextprotocol/inspector-test-server"; +import { assertDaemonToken, tokensEqual } from "../src/daemon/auth.js"; +import { callDaemon } from "../src/daemon/client.js"; +import { ensureDaemon } from "../src/daemon/ensure.js"; +import { MAX_REQUEST_LINE_BYTES } from "../src/daemon/ipc-glue.js"; +import { + createPrivateDaemonDir, + DAEMON_DIR_ENV, + DAEMON_TOKEN_ENV, + getDaemonTokenPath, +} from "../src/daemon/paths.js"; +import { DaemonServer } from "../src/daemon/server.js"; +import { CliExitCodeError } from "@inspector/core/cli/error-handler.js"; +import { runMcp } from "./helpers/mcp-runner.js"; +import { + expectCliSuccess, + expectCliFailure, +} from "../../cli/__tests__/helpers/assertions.js"; +import { + createSampleTestConfig, + deleteConfigFile, +} from "../../cli/__tests__/helpers/fixtures.js"; +import { + createPrivateBinding, + formatPrivateEnvExports, +} from "../src/connection/private-env.js"; + +describe("daemon IPC token", () => { + it("compares tokens in constant time", () => { + expect(tokensEqual("abc", "abc")).toBe(true); + expect(tokensEqual("abc", "abd")).toBe(false); + expect(tokensEqual("abc", "ab")).toBe(false); + expect(tokensEqual(undefined, "x")).toBe(false); + }); + + it("assertDaemonToken allows shared mode and rejects bad private tokens", () => { + expect(() => assertDaemonToken(undefined, undefined)).not.toThrow(); + expect(() => assertDaemonToken(undefined, "x")).not.toThrow(); + expect(() => assertDaemonToken("secret", "secret")).not.toThrow(); + expect(() => assertDaemonToken("secret", "nope")).toThrow(CliExitCodeError); + expect(() => assertDaemonToken("secret", undefined)).toThrow( + CliExitCodeError, + ); + }); +}); + +describe("mcpdo private", () => { + let home: string | undefined; + let prevHome: string | undefined; + + afterEach(() => { + if (prevHome === undefined) delete process.env.HOME; + else process.env.HOME = prevHome; + prevHome = undefined; + if (home) { + fs.rmSync(home, { recursive: true, force: true }); + home = undefined; + } + }); + + function useTempHome() { + prevHome = process.env.HOME; + home = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-home-")); + process.env.HOME = home; + } + + it("prints shell exports for a new private binding", async () => { + useTempHome(); + const result = await runMcp(["private"], { + env: { HOME: home! }, + }); + expectCliSuccess(result); + expect(result.stdout).toMatch( + new RegExp( + `export ${DAEMON_DIR_ENV}='[^']+/mcp-conn-[^/']+/[0-9a-f]{8}'`, + ), + ); + expect(result.stdout).toMatch( + new RegExp(`export ${DAEMON_TOKEN_ENV}='[^']+'`), + ); + const dirMatch = result.stdout.match( + new RegExp(`${DAEMON_DIR_ENV}='([^']+)'`), + ); + expect(dirMatch?.[1]).toBeTruthy(); + expect(fs.statSync(dirMatch![1]!).isDirectory()).toBe(true); + }); + + it("formatPrivateEnvExports escapes single quotes", () => { + const text = formatPrivateEnvExports({ + dir: "/tmp/o'brian", + token: "t'ok", + }); + expect(text).toContain(`'/tmp/o'\\''brian'`); + expect(text).toContain(`'t'\\''ok'`); + }); + + it("createPrivateBinding allocates a short 0700 dir under the tmpdir", () => { + useTempHome(); + const binding = createPrivateBinding(); + expect(path.basename(binding.dir)).toMatch(/^[0-9a-f]{8}$/); + expect(path.basename(path.dirname(binding.dir))).toMatch(/^mcp-conn-/); + expect(binding.dir.startsWith(os.tmpdir())).toBe(true); + expect(binding.token.length).toBeGreaterThan(20); + }); +}); + +describe("private daemon end-to-end", () => { + let server: DaemonServer | undefined; + let dir: string | undefined; + + afterEach(async () => { + if (server) { + await server.stop("stop"); + server = undefined; + } + if (dir) { + fs.rmSync(dir, { recursive: true, force: true }); + dir = undefined; + } + }); + + it("rejects IPC with a wrong token and accepts with the right one", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-priv-")); + const token = "test-token-value"; + server = new DaemonServer({ dir, idleMs: 0, requiredToken: token }); + await server.start(); + + // The daemon publishes mcpdod.token (0600) for same-user clients, so a + // tokenless call auto-discovers it; only a wrong token must fail. + await expect( + callDaemon( + "ping", + {}, + { socketPath: server.socketPath, timeoutMs: 2000, token: "wrong" }, + ), + ).rejects.toMatchObject({ envelope: { code: "daemon_auth_failed" } }); + + // Tokenless call discovers the published token file next to the socket. + const discovered = await callDaemon<{ pong: boolean }>( + "ping", + {}, + { socketPath: server.socketPath, timeoutMs: 2000 }, + ); + expect(discovered.pong).toBe(true); + + const pong = await callDaemon<{ pong: boolean }>( + "ping", + {}, + { socketPath: server.socketPath, timeoutMs: 2000, token }, + ); + expect(pong.pong).toBe(true); + }); + + it("publishes mcpdod.token (0600) on start and removes it on stop", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-priv-tok-")); + const token = "published-token"; + server = new DaemonServer({ dir, idleMs: 0, requiredToken: token }); + await server.start(); + + const tokenPath = getDaemonTokenPath(dir); + expect(fs.readFileSync(tokenPath, "utf8").trim()).toBe(token); + if (process.platform !== "win32") { + expect(fs.statSync(tokenPath).mode & 0o777).toBe(0o600); + } + + await server.stop("stop"); + server = undefined; + expect(fs.existsSync(tokenPath)).toBe(false); + }); + + it("drops a connection whose request line exceeds the cap", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-priv-cap-")); + server = new DaemonServer({ dir, idleMs: 0 }); + await server.start(); + + const net = await import("node:net"); + const closed = await new Promise<boolean>((resolve) => { + const socket = net.connect(server!.socketPath, () => { + // One oversized line, never newline-terminated. + socket.write(Buffer.alloc(MAX_REQUEST_LINE_BYTES + 64 * 1024, 0x61)); + }); + const done = () => resolve(true); + socket.once("close", done); + socket.once("error", done); + setTimeout(() => { + socket.destroy(); + resolve(false); + }, 5000).unref(); + }); + expect(closed).toBe(true); + }); + + it("rejects an oversized line even when its terminator arrives with it", async () => { + // Regression: the old cap only counted bytes after a chunk's last + // newline, so an oversized line whose terminating "\n" arrived in the + // crossing chunk reset the counter and reached readline. + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-priv-cap2-")); + server = new DaemonServer({ dir, idleMs: 0 }); + await server.start(); + + const net = await import("node:net"); + const closed = await new Promise<boolean>((resolve) => { + const socket = net.connect(server!.socketPath, () => { + socket.write(Buffer.alloc(600 * 1024, 0x61)); + socket.write( + Buffer.concat([Buffer.alloc(600 * 1024, 0x61), Buffer.from("\n")]), + ); + }); + const done = () => resolve(true); + socket.once("close", done); + socket.once("error", done); + setTimeout(() => { + socket.destroy(); + resolve(false); + }, 5000).unref(); + }); + expect(closed).toBe(true); + }); + + it("adopts the winner's published token when a concurrent starter wins the lock", async () => { + // Two concurrent first invocations each generate a token and spawn; the + // pid lock lets one daemon survive. The loser must finish against the + // winner's daemon by re-reading its published mcpdod.token, not poll + // with its own dead token until daemon_start_timeout. + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-priv-race-")); + const prevTok = process.env[DAEMON_TOKEN_ENV]; + delete process.env[DAEMON_TOKEN_ENV]; + try { + // Our child lost the O_EXCL pid lock: it exits without binding. + const stub = path.join(dir, "losing-daemon.js"); + fs.writeFileSync(stub, "process.exit(0);\n"); + const ensured = ensureDaemon({ dir, daemonScript: stub }); + // The concurrent winner, holding a different (published) token. + server = new DaemonServer({ + dir, + idleMs: 0, + requiredToken: "winner-token", + }); + await server.start(); + + const { socketPath, spawned } = await ensured; + expect(spawned).toBe(true); + const pong = await callDaemon<{ pong: boolean }>( + "ping", + {}, + { socketPath, timeoutMs: 2000, token: "winner-token" }, + ); + expect(pong.pong).toBe(true); + } finally { + if (prevTok === undefined) delete process.env[DAEMON_TOKEN_ENV]; + else process.env[DAEMON_TOKEN_ENV] = prevTok; + } + }); + + it("connection front-end rethrows non-unreachable daemon errors", async () => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-priv-rethrow-")); + const token = "good-token"; + server = new DaemonServer({ dir, idleMs: 0, requiredToken: token }); + await server.start(); + + const env = { + MCP_STORAGE_DIR: dir, + [DAEMON_DIR_ENV]: dir, + [DAEMON_TOKEN_ENV]: "wrong-token", + }; + + const listed = await runMcp(["connections/list"], { env }); + expectCliFailure(listed); + expect(listed.stderr).toMatch(/authentication failed|daemon_auth_failed/i); + + const status = await runMcp(["daemon", "status"], { env }); + expectCliFailure(status); + + const configPath = createSampleTestConfig(); + try { + const servers = await runMcp(["servers/list", "--config", configPath], { + env, + }); + // Optional daemon probe must not swallow auth failures as empty connections. + expectCliFailure(servers); + } finally { + deleteConfigFile(configPath); + } + }); + + it("ensureDaemon spawns a token-gated daemon from env", async () => { + const home = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-home-spawn-")); + const prevHome = process.env.HOME; + process.env.HOME = home; + try { + dir = createPrivateDaemonDir(); + const token = "spawn-token-xyz"; + const prevDir = process.env[DAEMON_DIR_ENV]; + const prevTok = process.env[DAEMON_TOKEN_ENV]; + process.env[DAEMON_DIR_ENV] = dir; + process.env[DAEMON_TOKEN_ENV] = token; + try { + const { socketPath, spawned } = await ensureDaemon({ dir, token }); + expect(spawned).toBe(true); + + // Explicit wrong token — do not rely on clearing env (callDaemon + // falls back to MCP_INSPECTOR_DAEMON_TOKEN when options.token omitted). + await expect( + callDaemon( + "ping", + {}, + { socketPath, timeoutMs: 2000, token: "wrong" }, + ), + ).rejects.toMatchObject({ envelope: { code: "daemon_auth_failed" } }); + + const pong = await callDaemon<{ pong: boolean }>( + "ping", + {}, + { socketPath, timeoutMs: 2000, token }, + ); + expect(pong.pong).toBe(true); + + const { command, args } = getTestMcpServerCommand(); + await callDaemon( + "connect", + { + name: "s", + serverConfig: { type: "stdio", command, args }, + serverIdentity: "s", + }, + { socketPath, timeoutMs: 15000, token }, + ); + await callDaemon("daemon/stop", {}, { socketPath, token }); + } finally { + if (prevDir === undefined) delete process.env[DAEMON_DIR_ENV]; + else process.env[DAEMON_DIR_ENV] = prevDir; + if (prevTok === undefined) delete process.env[DAEMON_TOKEN_ENV]; + else process.env[DAEMON_TOKEN_ENV] = prevTok; + } + } finally { + if (prevHome === undefined) delete process.env.HOME; + else process.env.HOME = prevHome; + fs.rmSync(home, { recursive: true, force: true }); + } + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-rpc-abort.test.ts b/clients/mcpdo/__tests__/daemon-rpc-abort.test.ts new file mode 100644 index 0000000000..70b77d9d29 --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-rpc-abort.test.ts @@ -0,0 +1,202 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { DaemonServer } from "../src/daemon/server.js"; +import type { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; + +/** + * Covers `rpc` cancellation when the caller's IPC socket closes: an abort + * mid-call cancels the in-flight tool call (so the per-client rpc queue + * isn't wedged behind work nobody awaits), and a queued rpc whose caller + * already hung up fails fast instead of running on a dead socket's behalf. + */ + +const runMethodMock = vi.hoisted(() => ({ + impl: undefined as unknown as (...args: unknown[]) => Promise<unknown>, +})); +vi.mock("@inspector/core/cli/handlers/run-method.js", () => ({ + runMethod: (...args: unknown[]) => runMethodMock.impl(...args), +})); + +function deferred<T>() { + let resolve!: (value: T) => void; + let reject!: (error: unknown) => void; + const promise = new Promise<T>((res, rej) => { + resolve = res; + reject = rej; + }); + return { promise, resolve, reject }; +} + +function fakeClient(): { + client: InspectorClient; + cancelToolCall: ReturnType<typeof vi.fn>; + getAmbientSignal: () => AbortSignal | undefined; +} { + const target = new EventTarget(); + const cancelToolCall = vi.fn().mockReturnValue(true); + let ambientSignal: AbortSignal | undefined; + const client = { + addEventListener: (type: string, listener: EventListener) => + target.addEventListener(type, listener), + removeEventListener: (type: string, listener: EventListener) => + target.removeEventListener(type, listener), + getStatus: () => "connected", + cancelToolCall, + setAmbientRequestSignal: (signal: AbortSignal | undefined) => { + ambientSignal = signal; + return () => { + ambientSignal = undefined; + }; + }, + } as unknown as InspectorClient; + return { client, cancelToolCall, getAmbientSignal: () => ambientSignal }; +} + +describe("daemon rpc abort on caller disconnect", () => { + let dir: string; + let server: DaemonServer; + let client: InspectorClient; + let cancelToolCall: ReturnType<typeof vi.fn>; + let getAmbientSignal: () => AbortSignal | undefined; + + beforeEach(() => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-rpc-abort-")); + server = new DaemonServer({ dir, idleMs: 0 }); + const fake = fakeClient(); + client = fake.client; + cancelToolCall = fake.cancelToolCall; + getAmbientSignal = fake.getAmbientSignal; + const registry = server.registry as unknown as Record<string, unknown>; + registry.connectionFor = () => ({ name: "srv", client }); + registry.liveClientFor = async () => client; + }); + + afterEach(() => { + fs.rmSync(dir, { recursive: true, force: true }); + vi.restoreAllMocks(); + }); + + function rpc(id: string, signal?: AbortSignal, method = "tools/call") { + return server.handle( + { id, op: "rpc", params: { method, name: "srv" } }, + undefined, + signal, + ); + } + + it("cancels the in-flight tool call when the caller aborts mid-call", async () => { + const running = deferred<void>(); + const result = deferred<unknown>(); + runMethodMock.impl = async () => { + running.resolve(); + return result.promise; + }; + const abort = new AbortController(); + const call = rpc("r1", abort.signal); + await running.promise; + expect(cancelToolCall).not.toHaveBeenCalled(); + + abort.abort(); + expect(cancelToolCall).toHaveBeenCalledTimes(1); + + // The real cancel rejects the SDK request; simulate that settle. + result.reject(new Error("Pending request aborted")); + const response = await call; + expect(response.ok).toBe(false); + }); + + it("does not cancel when the call settles normally", async () => { + runMethodMock.impl = async () => ({ kind: "result", result: { n: 1 } }); + const abort = new AbortController(); + const response = await rpc("r2", abort.signal); + expect(response.ok).toBe(true); + // Abort after settle must not reach into a later call. + abort.abort(); + expect(cancelToolCall).not.toHaveBeenCalled(); + }); + + it("fails a queued rpc fast when its caller already hung up", async () => { + const firstRunning = deferred<void>(); + const firstResult = deferred<unknown>(); + runMethodMock.impl = async () => { + firstRunning.resolve(); + return firstResult.promise; + }; + const first = rpc("r3"); + await firstRunning.promise; + + // Queue a second call behind the hung first, then hang up its caller. + const abort = new AbortController(); + const runMethodCalls: number[] = []; + const prior = runMethodMock.impl; + runMethodMock.impl = async (...args) => { + runMethodCalls.push(1); + return prior(...args); + }; + const second = rpc("r4", abort.signal); + abort.abort(); + + firstResult.resolve({ kind: "result", result: {} }); + expect((await first).ok).toBe(true); + const response = await second; + expect(response.ok).toBe(false); + if (!response.ok) { + expect(response.error?.code).toBe("caller_gone"); + } + expect(runMethodCalls).toHaveLength(0); + }); + + it("frees the connection when a non-tool request's caller aborts (R2)", async () => { + // Reproduces the wedge: the daemon serializes every method on a connection, + // so a non-tool request the server never answers (e.g. `resources/read`) + // held the queue slot forever once the caller hung up — `cancelToolCall()` + // is a no-op for it. The fix makes the caller's signal the client's ambient + // request signal, so core aborts the in-flight request and the slot frees. + // This stub stands in for core honoring that ambient signal: it settles + // only when the signal aborts, and never otherwise. If the daemon failed to + // wire the ambient signal, `getAmbientSignal()` is undefined and the call + // hangs forever — the wedge this test guards against. + const firstRunning = deferred<void>(); + runMethodMock.impl = async (_clientArg, args) => { + const method = (args as { method?: string } | undefined)?.method; + if (method === "resources/read") { + firstRunning.resolve(); + const ambient = getAmbientSignal(); + return new Promise((_resolve, reject) => { + ambient?.addEventListener( + "abort", + () => reject(new Error("Pending request aborted")), + { once: true }, + ); + }); + } + // The follow-up command settles normally once it gets to run. + return { kind: "result", result: { ok: true } }; + }; + + const abort = new AbortController(); + const first = rpc("r5", abort.signal, "resources/read"); + await firstRunning.promise; + + // The caller's signal must be the client's ambient request signal while the + // call runs — that is what threads cancellation into the in-flight request. + expect(getAmbientSignal()).toBe(abort.signal); + + // A second command queues behind the hung first; it must not be wedged. + const second = rpc("r6", undefined, "tools/list"); + + // The caller hangs up → the ambient signal aborts the in-flight request. + abort.abort(); + const firstResponse = await first; + expect(firstResponse.ok).toBe(false); + + // The queue slot freed, so the follow-up runs and settles normally. + const secondResponse = await second; + expect(secondResponse.ok).toBe(true); + + // The ambient signal is cleared once the call settles (disposer ran). + expect(getAmbientSignal()).toBeUndefined(); + }); +}); diff --git a/clients/mcpdo/__tests__/daemon-stream.test.ts b/clients/mcpdo/__tests__/daemon-stream.test.ts new file mode 100644 index 0000000000..09988c1fd0 --- /dev/null +++ b/clients/mcpdo/__tests__/daemon-stream.test.ts @@ -0,0 +1,585 @@ +import { describe, it, expect, afterEach, vi } from "vitest"; +import * as fs from "node:fs"; +import * as net from "node:net"; +import * as os from "node:os"; +import * as path from "node:path"; +import { streamDaemon } from "../src/daemon/stream-client.js"; +import { + acceptDaemonConnection, + removeStaleDaemonSocket, +} from "../src/daemon/ipc-glue.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; + +describe("streamDaemon + ipc-glue", () => { + let dir: string | undefined; + let server: net.Server | undefined; + const sockets = new Set<net.Socket>(); + + afterEach(async () => { + for (const s of sockets) { + s.destroy(); + } + sockets.clear(); + if (server) { + await new Promise<void>((resolve) => { + server!.close(() => resolve()); + }); + server = undefined; + } + if (dir) { + fs.rmSync(dir, { recursive: true, force: true }); + dir = undefined; + } + }); + + function freshSock(): string { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-stream-")); + return path.join(dir, "mcpdod.sock"); + } + + async function listen( + sock: string, + onSocket: (socket: net.Socket) => void, + ): Promise<void> { + server = net.createServer((socket) => { + sockets.add(socket); + socket.on("error", () => {}); + socket.on("close", () => sockets.delete(socket)); + onSocket(socket); + }); + await new Promise<void>((resolve) => server!.listen(sock, resolve)); + } + + it("resolves immediately on an already-aborted signal without dialing", async () => { + const sock = freshSock(); + const ac = new AbortController(); + ac.abort(); + // No server listens at `sock`: a dial would reject with + // daemon_unreachable, so resolving proves connect() was never called. + await expect( + streamDaemon( + { method: "logging/tail" }, + { + socketPath: sock, + timeoutMs: 2000, + signal: ac.signal, + onData: () => {}, + }, + ), + ).resolves.toBeUndefined(); + }); + + it("delivers data frames then end (skips blank/mismatched ids)", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + // Mismatched first response id is ignored; matching ok opens the stream. + socket.write( + JSON.stringify({ id: "wrong", ok: true, result: {} }) + "\n", + ); + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + socket.write("\n"); + socket.write( + JSON.stringify({ id: "other", stream: "data", data: { skip: 1 } }) + + "\n", + ); + socket.write( + JSON.stringify({ id: req.id, stream: "noop", data: 0 }) + "\n", + ); + socket.write( + JSON.stringify({ id: req.id, stream: "data", data: { n: 1 } }) + "\n", + ); + socket.write(JSON.stringify({ id: req.id, stream: "end" }) + "\n"); + }); + }); + + const data: unknown[] = []; + await streamDaemon( + { method: "logging/tail" }, + { + socketPath: sock, + timeoutMs: 5000, + onData: (d) => { + data.push(d); + }, + }, + ); + expect(data).toEqual([{ n: 1 }]); + }); + + it("pauses socket reads while an async onData callback is pending", async () => { + const sock = freshSock(); + let serverSocket: net.Socket | undefined; + let requestId: string | undefined; + await listen(sock, (socket) => { + serverSocket = socket; + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + requestId = req.id; + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + socket.write( + JSON.stringify({ id: req.id, stream: "data", data: { n: 1 } }) + "\n", + ); + }); + }); + + const until = async (cond: () => boolean) => { + const deadline = Date.now() + 3000; + while (!cond()) { + if (Date.now() > deadline) throw new Error("condition timed out"); + await new Promise((r) => setTimeout(r, 5)); + } + }; + + const seen: unknown[] = []; + let release!: () => void; + const gate = new Promise<void>((resolve) => { + release = resolve; + }); + const done = streamDaemon( + { method: "logging/tail" }, + { + socketPath: sock, + timeoutMs: 5000, + // First callback stalls on the gate; the client must stop reading + // instead of queueing further frames behind an unbounded chain. + onData: (d) => { + seen.push(d); + return seen.length === 1 ? gate : undefined; + }, + }, + ); + + await until(() => seen.length === 1); + // Send a second frame + end while the first callback is still pending. + serverSocket!.write( + JSON.stringify({ id: requestId, stream: "data", data: { n: 2 } }) + "\n", + ); + serverSocket!.write( + JSON.stringify({ id: requestId, stream: "end" }) + "\n", + ); + await new Promise((r) => setTimeout(r, 100)); + // Reads are paused, so the second frame must not have been dispatched. + expect(seen.length).toBe(1); + + release(); + await done; + expect(seen).toEqual([{ n: 1 }, { n: 2 }]); + }); + + it("continues the stream when an async onData callback rejects", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + socket.write( + JSON.stringify({ id: req.id, stream: "data", data: { n: 1 } }) + "\n", + ); + socket.write( + JSON.stringify({ id: req.id, stream: "data", data: { n: 2 } }) + "\n", + ); + socket.write(JSON.stringify({ id: req.id, stream: "end" }) + "\n"); + }); + }); + + const seen: unknown[] = []; + // Write errors are non-fatal: a rejected callback promise must not kill + // the stream or surface as an unhandled rejection. + await streamDaemon( + { method: "logging/tail" }, + { + socketPath: sock, + timeoutMs: 5000, + onData: (d) => { + seen.push(d); + return Promise.reject(new Error("write failed")); + }, + }, + ); + expect(seen).toEqual([{ n: 1 }, { n: 2 }]); + }); + + it("rejects on socket error after the stream has opened", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + setTimeout(() => socket.destroy(), 20); + }); + }); + await expect( + streamDaemon({}, { socketPath: sock, timeoutMs: 2000, onData: () => {} }), + ).rejects.toMatchObject({ + envelope: { code: "daemon_unreachable" }, + }); + }); + + it("rejects malformed stream frames after open", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + socket.write("not-a-frame\n"); + }); + }); + await expect( + streamDaemon({}, { socketPath: sock, timeoutMs: 2000, onData: () => {} }), + ).rejects.toThrow(); + }); + + it("rejects error responses without exitCode (defaults USAGE)", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ + id: req.id, + ok: false, + error: { code: "usage", message: "nope" }, + }) + "\n", + ); + }); + }); + + await expect( + streamDaemon({}, { socketPath: sock, timeoutMs: 2000, onData: () => {} }), + ).rejects.toMatchObject({ exitCode: EXIT_CODES.USAGE }); + }); + + it("rejects malformed first-frame JSON", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", () => { + socket.write("not-json\n"); + }); + }); + await expect( + streamDaemon({}, { socketPath: sock, timeoutMs: 2000, onData: () => {} }), + ).rejects.toThrow(); + }); + + it("finishes immediately on a pre-aborted signal instead of hanging", async () => { + const sock = freshSock(); + await listen(sock, () => { + // Never respond: only the pre-aborted check can settle this promptly. + }); + const ac = new AbortController(); + ac.abort(); + await streamDaemon( + {}, + { + socketPath: sock, + timeoutMs: 5000, + signal: ac.signal, + onData: () => {}, + }, + ); + }); + + it("aborts via signal after the stream opens", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + }); + }); + + const ac = new AbortController(); + const pending = streamDaemon( + {}, + { + socketPath: sock, + timeoutMs: 5000, + signal: ac.signal, + onData: () => {}, + }, + ); + await new Promise((r) => setTimeout(r, 50)); + ac.abort(); + await pending; + }); + + it("times out a hung stream open", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", () => { + // never respond with an ok frame + }); + }); + await expect( + streamDaemon({}, { socketPath: sock, timeoutMs: 80, onData: () => {} }), + ).rejects.toThrow(/timed out/); + }, 5000); + + it("disables the open deadline entirely with timeoutMs 0", async () => { + // Regression: an unconditional setTimeout(..., 0) fired on the next + // tick, so timeoutMs 0 (documented as "no deadline") failed every + // stream immediately instead of waiting indefinitely. + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + setTimeout(() => { + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + socket.write(JSON.stringify({ id: req.id, stream: "end" }) + "\n"); + }, 120); + }); + }); + await expect( + streamDaemon({}, { socketPath: sock, timeoutMs: 0, onData: () => {} }), + ).resolves.toBeUndefined(); + }, 5000); + + it("fails when the peer FINs before the stream ok frame", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.on("error", () => {}); + socket.once("data", () => { + socket.end(); + }); + }); + await expect( + streamDaemon( + {}, + { socketPath: sock, timeoutMs: 60_000, onData: () => {} }, + ), + ).rejects.toMatchObject({ + envelope: { code: "daemon_unreachable" }, + }); + }); + + it("rejects when the peer closes mid-stream without an end frame", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + socket.end(); + }); + }); + await expect( + streamDaemon({}, { socketPath: sock, timeoutMs: 2000, onData: () => {} }), + ).rejects.toMatchObject({ + envelope: { code: "daemon_unreachable" }, + message: expect.stringMatching(/closed the stream before it ended/), + }); + }); + + it("uses env-derived defaults and sends an explicit token", async () => { + const sock = freshSock(); + const prevDir = process.env.MCP_INSPECTOR_DAEMON_DIR; + const prevToken = process.env.MCP_INSPECTOR_DAEMON_TOKEN; + process.env.MCP_INSPECTOR_DAEMON_DIR = path.dirname(sock); + delete process.env.MCP_INSPECTOR_DAEMON_TOKEN; + try { + let seenToken: string | undefined; + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { + id: string; + token?: string; + }; + seenToken = req.token; + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + "\n", + ); + socket.write(JSON.stringify({ id: req.id, stream: "end" }) + "\n"); + }); + }); + // No socketPath / timeoutMs: both fall back to defaults (the daemon + // dir env pins the socket path; the 60s default timer is cleared by + // the ok frame). + await streamDaemon({}, { token: "tok-1", onData: () => {} }); + expect(seenToken).toBe("tok-1"); + } finally { + if (prevDir === undefined) delete process.env.MCP_INSPECTOR_DAEMON_DIR; + else process.env.MCP_INSPECTOR_DAEMON_DIR = prevDir; + if (prevToken !== undefined) + process.env.MCP_INSPECTOR_DAEMON_TOKEN = prevToken; + } + }); + + it("ignores frames after the stream has already ended", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + socket.once("data", (buf) => { + const req = JSON.parse(String(buf).trim()) as { id: string }; + // ok + two end frames in one chunk: the second is handled by the + // same buffered-line loop after the promise has settled. + socket.write( + JSON.stringify({ id: req.id, ok: true, result: {} }) + + "\n" + + JSON.stringify({ id: req.id, stream: "end" }) + + "\n" + + JSON.stringify({ id: req.id, stream: "end" }) + + "\n", + ); + }); + }); + await streamDaemon( + {}, + { socketPath: sock, timeoutMs: 2000, onData: () => {} }, + ); + }); + + it("removeStaleDaemonSocket handles absent, dead, and live sockets", async () => { + const sock = freshSock(); + await removeStaleDaemonSocket(sock); + + fs.writeFileSync(sock, ""); + await removeStaleDaemonSocket(sock); + expect(fs.existsSync(sock)).toBe(false); + + await listen(sock, () => {}); + await expect(removeStaleDaemonSocket(sock)).rejects.toThrow( + /already running/, + ); + }); + + it("acceptDaemonConnection rejects invalid request lines", async () => { + const sock = freshSock(); + const chunks: string[] = []; + await listen(sock, (socket) => { + acceptDaemonConnection(socket, async () => ({ + response: { id: "x", ok: true, result: {} }, + })); + }); + + await new Promise<void>((resolve, reject) => { + const client = net.connect(sock, () => { + sockets.add(client); + client.on("data", (c) => chunks.push(String(c))); + client.write('{"op":"ping"}\n'); + setTimeout(() => { + client.destroy(); + resolve(); + }, 50); + }); + client.on("error", reject); + }); + expect(chunks.join("")).toContain("invalid_request"); + }); + + it("acceptDaemonConnection streams via startStream until socket closes", async () => { + const sock = freshSock(); + let stopCalled = false; + await listen(sock, (socket) => { + acceptDaemonConnection(socket, async (req) => ({ + response: { id: req.id, ok: true, result: {} }, + startStream: (writeData) => { + writeData({ a: 1 }); + return () => { + stopCalled = true; + throw new Error("unsubscribe boom"); + }; + }, + })); + }); + + const frames: string[] = []; + await new Promise<void>((resolve) => { + const client = net.connect(sock, () => { + sockets.add(client); + client.on("data", (c) => frames.push(String(c))); + client.on("close", () => resolve()); + client.on("error", () => {}); + client.write( + JSON.stringify({ id: "s1", op: "stream", params: {} }) + "\n", + ); + // Half-close so the server cleanup can still write the end frame. + setTimeout(() => client.end(), 80); + }); + client.on("error", () => {}); + }); + const joined = frames.join(""); + expect(joined).toContain('"stream":"data"'); + expect(stopCalled).toBe(true); + }); + + it("terminates a stream once the socket write buffer exceeds the cap", async () => { + const sock = freshSock(); + let stopCalled = false; + let writeFn: ((data: unknown) => void) | undefined; + await listen(sock, (socket) => { + acceptDaemonConnection(socket, async (req) => ({ + response: { id: req.id, ok: true, result: {} }, + startStream: (writeData) => { + writeFn = writeData; + return () => { + stopCalled = true; + }; + }, + })); + }); + + let sawEnd = false; + let client!: net.Socket; + const closed = new Promise<void>((resolve) => { + client = net.connect(sock, () => { + sockets.add(client); + // Never read: the daemon-side write buffer must hit the cap instead + // of growing without bound. + client.pause(); + client.write( + JSON.stringify({ id: "s1", op: "stream", params: {} }) + "\n", + ); + }); + client.on("data", (c) => { + if (String(c).includes('"stream":"end"')) sawEnd = true; + }); + client.on("close", () => resolve()); + client.on("error", () => {}); + }); + + await vi.waitFor(() => expect(writeFn).toBeDefined()); + const chunk = "x".repeat(64 * 1024); + for (let i = 0; i < 200 && !stopCalled; i++) { + writeFn!({ chunk }); + } + // Producer unsubscribed and socket destroyed — no clean end frame. + expect(stopCalled).toBe(true); + // The paused client never drains, so its "close" only fires once the + // test tears the socket down. + client.destroy(); + await closed; + expect(sawEnd).toBe(false); + }); + + it("unreachable socket path fails before streaming", async () => { + await expect( + streamDaemon( + {}, + { + socketPath: path.join(os.tmpdir(), "no-such-mcp-mcpdod.sock"), + timeoutMs: 500, + onData: () => {}, + }, + ), + ).rejects.toBeInstanceOf(CliExitCodeError); + }); +}); diff --git a/clients/mcpdo/__tests__/disconnect-clear-auth.test.ts b/clients/mcpdo/__tests__/disconnect-clear-auth.test.ts new file mode 100644 index 0000000000..a11d9e8fde --- /dev/null +++ b/clients/mcpdo/__tests__/disconnect-clear-auth.test.ts @@ -0,0 +1,123 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; + +const callDaemon = vi.fn(); +const ensureDaemon = vi.fn(); +const clearStoredAuthForRelogin = vi.fn(); + +// importOriginal keeps every other daemon/stored-auth export intact; only the +// two functions the disconnect path reaches are stubbed so no real daemon +// socket or OAuth-store write happens in-process. +vi.mock("../src/daemon/index.js", async (importOriginal) => ({ + ...(await importOriginal<typeof import("../src/daemon/index.js")>()), + callDaemon: (...args: unknown[]) => callDaemon(...args), + ensureDaemon: (...args: unknown[]) => ensureDaemon(...args), +})); + +vi.mock("../src/connection/stored-auth.js", async (importOriginal) => ({ + ...(await importOriginal< + typeof import("../src/connection/stored-auth.js") + >()), + clearStoredAuthForRelogin: (...args: unknown[]) => + clearStoredAuthForRelogin(...args), +})); + +describe("disconnect --clear-auth", () => { + let stdout: string; + let originalStdoutWrite: typeof process.stdout.write; + + beforeEach(() => { + stdout = ""; + originalStdoutWrite = process.stdout.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]) => { + stdout += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + callDaemon.mockReset(); + ensureDaemon.mockReset(); + clearStoredAuthForRelogin.mockReset(); + ensureDaemon.mockResolvedValue({ socketPath: "/tmp/mcpdod.sock" }); + clearStoredAuthForRelogin.mockResolvedValue(undefined); + }); + + afterEach(() => { + process.stdout.write = originalStdoutWrite; + }); + + it("clears stored auth for the server URL the daemon returns", async () => { + callDaemon.mockResolvedValue({ + name: "foo", + serverUrl: "https://mcp.example.com/mcp", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "disconnect", + "--clear-auth", + "foo", + "--format", + "json", + ]); + expect(clearStoredAuthForRelogin).toHaveBeenCalledWith( + "https://mcp.example.com/mcp", + ); + expect(JSON.parse(stdout.trim())).toEqual({ + name: "foo", + clearedAuthUrl: "https://mcp.example.com/mcp", + }); + }); + + it("accepts the -c short alias", async () => { + callDaemon.mockResolvedValue({ + name: "foo", + serverUrl: "https://mcp.example.com/mcp", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "disconnect", + "-c", + "foo", + "--format", + "json", + ]); + expect(clearStoredAuthForRelogin).toHaveBeenCalledWith( + "https://mcp.example.com/mcp", + ); + expect(JSON.parse(stdout.trim()).clearedAuthUrl).toBe( + "https://mcp.example.com/mcp", + ); + }); + + it("does not clear auth without the flag", async () => { + callDaemon.mockResolvedValue({ + name: "foo", + serverUrl: "https://mcp.example.com/mcp", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "disconnect", "foo", "--format", "json"]); + expect(clearStoredAuthForRelogin).not.toHaveBeenCalled(); + expect(JSON.parse(stdout.trim())).toEqual({ name: "foo" }); + }); + + it("is a no-op when the daemon returns no server URL (stdio)", async () => { + callDaemon.mockResolvedValue({ name: "local" }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "disconnect", + "--clear-auth", + "local", + "--format", + "json", + ]); + expect(clearStoredAuthForRelogin).not.toHaveBeenCalled(); + expect(JSON.parse(stdout.trim())).toEqual({ name: "local" }); + }); +}); diff --git a/clients/mcpdo/__tests__/dispatch.test.ts b/clients/mcpdo/__tests__/dispatch.test.ts new file mode 100644 index 0000000000..69e17814b4 --- /dev/null +++ b/clients/mcpdo/__tests__/dispatch.test.ts @@ -0,0 +1,500 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; + +const callDaemon = vi.fn(); +const ensureDaemon = vi.fn(); +const streamDaemon = vi.fn(); +const promptElicitation = vi.fn(); + +vi.mock("../src/daemon/index.js", () => ({ + callDaemon: (...args: unknown[]) => callDaemon(...args), + ensureDaemon: (...args: unknown[]) => ensureDaemon(...args), + streamDaemon: (...args: unknown[]) => streamDaemon(...args), +})); + +vi.mock("../src/connection/elicitation-prompt.js", () => ({ + promptElicitation: (...args: unknown[]) => promptElicitation(...args), +})); + +// Pass-through wrapper so tests can delay writes and observe completion +// order (the stream path must flush queued writes before returning). +const writeDelayMs = { value: 0 }; +const writeReject = { value: false }; +const writeCompletions: unknown[] = []; +vi.mock("../src/connection/format-connection.js", async (importOriginal) => { + const actual = + await importOriginal< + typeof import("../src/connection/format-connection.js") + >(); + return { + ...actual, + writeConnectionOutput: async (...args: unknown[]) => { + if (writeReject.value) throw new Error("stdout write failed"); + if (writeDelayMs.value > 0) { + await new Promise((r) => setTimeout(r, writeDelayMs.value)); + } + await ( + actual.writeConnectionOutput as (...a: unknown[]) => Promise<void> + )(...args); + writeCompletions.push(args[1]); + }, + }; +}); + +describe("dispatchConnectionRpc", () => { + let stdout: string; + let originalWrite: typeof process.stdout.write; + + beforeEach(() => { + stdout = ""; + originalWrite = process.stdout.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]) => { + stdout += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + ensureDaemon.mockResolvedValue({ socketPath: "/tmp/t.sock" }); + callDaemon.mockReset(); + streamDaemon.mockReset(); + promptElicitation.mockReset(); + writeDelayMs.value = 0; + writeReject.value = false; + writeCompletions.length = 0; + }); + + afterEach(() => { + process.stdout.write = originalWrite; + }); + + it("writes pretty JSON for --format json", async () => { + callDaemon.mockResolvedValue({ + kind: "result", + result: { tools: [] }, + }); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/list", + {}, + { format: "json", requireExplicit: false }, + ); + expect(JSON.parse(stdout.trim())).toEqual({ tools: [] }); + expect(stdout).toContain("\n"); + }); + + it("omits format from the daemon rpc params (frontend-only concern)", async () => { + callDaemon.mockResolvedValue({ kind: "result", result: {} }); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + { toolName: "echo" }, + { format: "json", requireExplicit: false }, + ); + const [op, params] = callDaemon.mock.calls[0] as [ + string, + Record<string, unknown>, + ]; + expect(op).toBe("rpc"); + // Forwarding format would make the daemon's runMethod issue a hidden + // app-info resources/read for JSON tool calls. + expect("format" in params).toBe(false); + expect(params).toMatchObject({ method: "tools/call", toolName: "echo" }); + }); + + it("writes human text for tools/list by default", async () => { + callDaemon.mockResolvedValue({ + kind: "result", + result: { + tools: [{ name: "echo", description: "Echo", inputSchema: {} }], + }, + }); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc("tools/list", {}, { requireExplicit: false }); + expect(stdout).toContain("Tools (1):"); + expect(stdout).toContain("`echo"); + }); + + it("writes human app-info list for ndjson outcomes", async () => { + callDaemon.mockResolvedValue({ + kind: "ndjson", + lines: [{ hasApp: false, toolName: "a" }], + }); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/list", + { appInfo: true }, + { requireExplicit: false }, + ); + expect(stdout).toContain("App info"); + expect(stdout).toContain("`a`"); + }); + + it("opens a stream for logging/tail and wires SIGINT abort", async () => { + // Invoke only the SIGINT listener dispatch registers, rather than + // broadcasting `process.emit("SIGINT")` process-wide: `atomically` + // (loaded via core's secret-store persistence) pulls in `when-exit`, + // whose module-level SIGINT handler re-raises the signal and kills the + // vitest worker fork mid-run (#1941). + const listenersBefore = new Set(process.listeners("SIGINT")); + streamDaemon.mockImplementation( + async ( + _params: unknown, + opts: { onData: (d: unknown) => void; signal?: AbortSignal }, + ) => { + opts.onData({ + type: "subscribed", + uri: "test://x", + }); + const added = process + .listeners("SIGINT") + .filter((listener) => !listenersBefore.has(listener)); + expect(added).toHaveLength(1); + for (const listener of added) listener("SIGINT"); + expect(opts.signal?.aborted).toBe(true); + }, + ); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "logging/tail", + {}, + { requireExplicit: false, connection: "@s" }, + ); + expect(stdout).toContain("Subscribed:"); + expect(streamDaemon).toHaveBeenCalled(); + // Core enforces the configured MCP request timeout daemon-side; the + // stream path must disable the fixed local deadline like the rpc path. + expect(streamDaemon).toHaveBeenCalledWith( + expect.anything(), + expect.objectContaining({ timeoutMs: 0 }), + ); + }); + + it("flushes queued stream writes before returning", async () => { + // Regression: stream writes were fire-and-forget, so mcp-bin's + // process.exit() right after dispatch resolved could truncate the final + // event when stdout is piped or backpressured. + writeDelayMs.value = 10; + streamDaemon.mockImplementation( + async (_params: unknown, opts: { onData: (d: unknown) => void }) => { + opts.onData({ type: "subscribed", uri: "test://one" }); + opts.onData({ type: "subscribed", uri: "test://two" }); + }, + ); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "logging/tail", + {}, + { requireExplicit: false, connection: "@s" }, + ); + expect(writeCompletions.length).toBe(2); + expect(stdout).toContain("test://two"); + }); + + it("recovers the write chain after a failed write and keeps streaming", async () => { + // Regression: one rejected write left the chain permanently rejected, so + // every later frame's `.then` was skipped and the stream went silent. + writeReject.value = true; + streamDaemon.mockImplementation( + async ( + _params: unknown, + opts: { onData: (d: unknown) => void | Promise<void> }, + ) => { + await opts.onData({ type: "subscribed", uri: "test://failed" }); + writeReject.value = false; + await opts.onData({ type: "subscribed", uri: "test://recovered" }); + }, + ); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "logging/tail", + {}, + { requireExplicit: false, connection: "@s" }, + ); + expect(stdout).not.toContain("test://failed"); + expect(stdout).toContain("test://recovered"); + }); + + it("keeps stream write failures non-fatal, as when they were fire-and-forget", async () => { + writeReject.value = true; + streamDaemon.mockImplementation( + async (_params: unknown, opts: { onData: (d: unknown) => void }) => { + opts.onData({ type: "subscribed", uri: "test://x" }); + opts.onData({ type: "subscribed", uri: "test://y" }); + }, + ); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await expect( + dispatchConnectionRpc( + "logging/tail", + {}, + { requireExplicit: false, connection: "@s" }, + ), + ).resolves.toBeUndefined(); + }); + + it("wires SIGINT/SIGTERM abort for the general rpc path (not just streams)", async () => { + // Same when-exit hazard as the stream test above: invoke only the + // SIGTERM listener dispatch registered, never a process-wide emit. + const listenersBefore = new Set(process.listeners("SIGTERM")); + callDaemon.mockImplementation( + async (_op: string, _params: unknown, opts: { signal?: AbortSignal }) => { + const added = process + .listeners("SIGTERM") + .filter((listener) => !listenersBefore.has(listener)); + expect(added).toHaveLength(1); + for (const listener of added) listener("SIGTERM"); + expect(opts.signal?.aborted).toBe(true); + return { kind: "result", result: {} }; + }, + ); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + {}, + { format: "json", requireExplicit: false }, + ); + expect(callDaemon).toHaveBeenCalled(); + }); + + it("removes the SIGINT/SIGTERM listeners after the rpc call settles", async () => { + callDaemon.mockResolvedValue({ kind: "result", result: {} }); + const before = process.listenerCount("SIGINT"); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + {}, + { format: "json", requireExplicit: false }, + ); + expect(process.listenerCount("SIGINT")).toBe(before); + }); + + it("wires onElicitation as interactive when text format + TTY stdin/stdout", async () => { + callDaemon.mockResolvedValue({ kind: "result", result: {} }); + const stdinDesc = Object.getOwnPropertyDescriptor(process.stdin, "isTTY"); + const stdoutDesc = Object.getOwnPropertyDescriptor(process.stdout, "isTTY"); + Object.defineProperty(process.stdin, "isTTY", { + configurable: true, + value: true, + }); + Object.defineProperty(process.stdout, "isTTY", { + configurable: true, + value: true, + }); + try { + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + {}, + { format: "text", requireExplicit: false }, + ); + const opts = callDaemon.mock.calls[0][2] as { + onElicitation: (frame: unknown) => unknown; + }; + expect(opts.onElicitation).toBeInstanceOf(Function); + promptElicitation.mockResolvedValue({ action: "cancel" }); + await opts.onElicitation({ id: "x" }); + expect(promptElicitation).toHaveBeenCalledWith( + { id: "x" }, + expect.objectContaining({ interactive: true }), + ); + } finally { + if (stdinDesc) Object.defineProperty(process.stdin, "isTTY", stdinDesc); + if (stdoutDesc) + Object.defineProperty(process.stdout, "isTTY", stdoutDesc); + } + }); + + it("wires onElicitation as non-interactive for --format json", async () => { + callDaemon.mockResolvedValue({ kind: "result", result: {} }); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + {}, + { format: "json", requireExplicit: false }, + ); + const opts = callDaemon.mock.calls[0][2] as { + onElicitation: (frame: unknown) => unknown; + }; + promptElicitation.mockResolvedValue({ action: "cancel" }); + await opts.onElicitation({ id: "x" }); + expect(promptElicitation).toHaveBeenCalledWith( + { id: "x" }, + expect.objectContaining({ interactive: false }), + ); + }); + + it("marks the daemon call non-interactive for --format json and for non-TTY text", async () => { + callDaemon.mockResolvedValue({ kind: "result", result: {} }); + const stdinDesc = Object.getOwnPropertyDescriptor(process.stdin, "isTTY"); + const stderrDesc = Object.getOwnPropertyDescriptor(process.stderr, "isTTY"); + Object.defineProperty(process.stdin, "isTTY", { + configurable: true, + value: undefined, + }); + Object.defineProperty(process.stderr, "isTTY", { + configurable: true, + value: undefined, + }); + try { + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + {}, + { format: "text", requireExplicit: false }, + ); + await dispatchConnectionRpc( + "tools/call", + {}, + { format: "json", requireExplicit: false }, + ); + expect(callDaemon.mock.calls[0][1]).toMatchObject({ + interactive: false, + }); + expect(callDaemon.mock.calls[1][1]).toMatchObject({ + interactive: false, + }); + } finally { + if (stdinDesc) Object.defineProperty(process.stdin, "isTTY", stdinDesc); + if (stderrDesc) + Object.defineProperty(process.stderr, "isTTY", stderrDesc); + } + }); + + it("marks the daemon call interactive for interactive text (TTY)", async () => { + callDaemon.mockResolvedValue({ kind: "result", result: {} }); + const stderrDesc = Object.getOwnPropertyDescriptor(process.stderr, "isTTY"); + Object.defineProperty(process.stderr, "isTTY", { + configurable: true, + value: true, + }); + try { + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + {}, + { format: "text", requireExplicit: false }, + ); + expect( + (callDaemon.mock.calls[0][1] as Record<string, unknown>).interactive, + ).toBe(true); + } finally { + if (stderrDesc) + Object.defineProperty(process.stderr, "isTTY", stderrDesc); + } + }); + + it("renders an elicitation-pending outcome (json and human)", async () => { + const elicitation = { + elicitationId: "e-1", + connection: "srv", + method: "tools/call", + toolName: "collect", + mode: "form", + message: "Pick a color", + requestedSchema: { + type: "object", + properties: { color: { type: "string" } }, + required: ["color"], + }, + origin: "server-request", + expiresAt: Date.now() + 600_000, + }; + callDaemon.mockResolvedValue({ kind: "elicitation-pending", elicitation }); + const { dispatchConnectionRpc } = + await import("../src/connection/dispatch.js"); + await dispatchConnectionRpc( + "tools/call", + { toolName: "collect" }, + { format: "json", requireExplicit: false }, + ); + const parsed = JSON.parse(stdout) as { + elicitationPending: { elicitationId: string }; + }; + expect(parsed.elicitationPending.elicitationId).toBe("e-1"); + + stdout = ""; + await dispatchConnectionRpc( + "tools/call", + { toolName: "collect" }, + // --plain: human rendering must stay assertable when this test runs + // under a stderr TTY (styled output would interleave ANSI codes). + { format: "text", plain: true, requireExplicit: false }, + ); + expect(stdout).toContain("Input required"); + expect(stdout).toContain("Pick a color"); + expect(stdout).toContain("elicitation/respond e-1"); + expect(stdout).toContain("color (string, required)"); + }); +}); + +describe("hoistAtConnection / stripAt / requireExplicitConnection", () => { + it("stripAt removes leading @", async () => { + const { stripAt, requireExplicitConnection } = + await import("../src/connection/dispatch.js"); + expect(stripAt("@x")).toBe("x"); + expect(stripAt(undefined)).toBeUndefined(); + const prev = process.env.MCP_ALLOW_DEFAULT_CONNECTION; + process.env.MCP_ALLOW_DEFAULT_CONNECTION = "1"; + expect(requireExplicitConnection()).toBe(false); + if (prev === undefined) delete process.env.MCP_ALLOW_DEFAULT_CONNECTION; + else process.env.MCP_ALLOW_DEFAULT_CONNECTION = prev; + }); + + it("requireExplicitConnection keys off stdin TTY (piping stdout still OK)", async () => { + const { requireExplicitConnection } = + await import("../src/connection/dispatch.js"); + const prevEnv = process.env.MCP_ALLOW_DEFAULT_CONNECTION; + delete process.env.MCP_ALLOW_DEFAULT_CONNECTION; + const stdinDesc = Object.getOwnPropertyDescriptor(process.stdin, "isTTY"); + const stdoutDesc = Object.getOwnPropertyDescriptor(process.stdout, "isTTY"); + try { + Object.defineProperty(process.stdin, "isTTY", { + configurable: true, + value: true, + }); + Object.defineProperty(process.stdout, "isTTY", { + configurable: true, + value: false, + }); + expect(requireExplicitConnection()).toBe(false); + + Object.defineProperty(process.stdin, "isTTY", { + configurable: true, + value: false, + }); + expect(requireExplicitConnection()).toBe(true); + } finally { + if (stdinDesc) Object.defineProperty(process.stdin, "isTTY", stdinDesc); + else + Object.defineProperty(process.stdin, "isTTY", { + configurable: true, + value: undefined, + }); + if (stdoutDesc) + Object.defineProperty(process.stdout, "isTTY", stdoutDesc); + else + Object.defineProperty(process.stdout, "isTTY", { + configurable: true, + value: undefined, + }); + if (prevEnv === undefined) + delete process.env.MCP_ALLOW_DEFAULT_CONNECTION; + else process.env.MCP_ALLOW_DEFAULT_CONNECTION = prevEnv; + } + }); +}); diff --git a/clients/mcpdo/__tests__/elicitation-bridge.test.ts b/clients/mcpdo/__tests__/elicitation-bridge.test.ts new file mode 100644 index 0000000000..b220ab4f2c --- /dev/null +++ b/clients/mcpdo/__tests__/elicitation-bridge.test.ts @@ -0,0 +1,265 @@ +import { describe, it, expect, vi } from "vitest"; +import { wireElicitationBridge } from "../src/daemon/elicitation-bridge.js"; +import type { ElicitationChannel } from "../src/daemon/ipc-glue.js"; +import type { ElicitationResponseFrame } from "../src/daemon/protocol.js"; +import type { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; + +/** + * Covers `wireElicitationBridge`'s event routing: task-input-required + * delivery (to an awaiting caller like any other origin; left pending — not + * cancelled — when no caller is awaiting), URL vs form mode frame shaping, + * and the channel-failure fallback to `cancel()` (since some construction + * sites, notably legacy URL-mode, never wire a reject callback). + */ +function fakeClient(): { + client: InspectorClient; + emit: (detail: unknown) => void; +} { + const target = new EventTarget(); + const client = { + addEventListener: (type: string, listener: EventListener) => + target.addEventListener(type, listener), + removeEventListener: (type: string, listener: EventListener) => + target.removeEventListener(type, listener), + } as unknown as InspectorClient; + return { + client, + emit: (detail: unknown) => + target.dispatchEvent( + new CustomEvent("newPendingElicitation", { detail }), + ), + }; +} + +function fakeMessage(overrides: Partial<Record<string, unknown>> = {}) { + return { + id: "elicitation-x", + origin: "server-request", + request: { method: "elicitation/create", params: { message: "hi" } }, + respond: vi.fn().mockResolvedValue(undefined), + cancel: vi.fn(), + reject: vi.fn(), + ...overrides, + }; +} + +describe("wireElicitationBridge", () => { + it("delivers a task-input-required elicitation to an awaiting caller", async () => { + const { client, emit } = fakeClient(); + const answer: ElicitationResponseFrame = { + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-x", + action: "accept", + }; + const request = vi.fn().mockResolvedValue(answer); + const unwire = wireElicitationBridge(client, { request }, "req-1"); + const message = fakeMessage({ origin: "task-input-required" }); + emit(message); + await vi.waitFor(() => expect(message.respond).toHaveBeenCalled()); + expect(request).toHaveBeenCalledWith( + expect.objectContaining({ origin: "task-input-required" }), + ); + unwire(); + }); + + it("leaves a task-input-required elicitation pending when no caller is awaiting it", async () => { + const { client, emit } = fakeClient(); + const request = vi.fn(); + const unwire = wireElicitationBridge(client, { request }, "req-1"); + const message = fakeMessage({ origin: "task-input-required" }); + emit(message); + unwire(); // settle before the queued microtask dispatches + // Flush the dispatch queue: a non-task origin would have been cancelled + // by now (see the cancel test below); task-input-required must stay + // pending for a later tasks/-driven answer. + await new Promise((r) => setTimeout(r, 0)); + expect(message.cancel).not.toHaveBeenCalled(); + expect(message.respond).not.toHaveBeenCalled(); + expect(request).not.toHaveBeenCalled(); + }); + + it("builds a url-mode frame and responds with the channel's answer", async () => { + const { client, emit } = fakeClient(); + const answer: ElicitationResponseFrame = { + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-x", + action: "accept", + }; + const request = vi.fn().mockResolvedValue(answer); + const channel: ElicitationChannel = { request }; + const unwire = wireElicitationBridge(client, channel, "req-1"); + const message = fakeMessage({ + request: { + method: "elicitation/create", + params: { message: "Please visit", url: "https://example.com" }, + }, + }); + emit(message); + await Promise.resolve(); + await Promise.resolve(); + await Promise.resolve(); + + expect(request).toHaveBeenCalledWith( + expect.objectContaining({ + kind: "elicitation-request", + mode: "url", + url: "https://example.com", + elicitationId: "elicitation-x", + origin: "server-request", + }), + ); + expect(message.respond).toHaveBeenCalledWith({ + action: "accept", + content: undefined, + }); + unwire(); + }); + + it("builds a form-mode frame with requestedSchema", async () => { + const { client, emit } = fakeClient(); + const answer: ElicitationResponseFrame = { + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-x", + action: "decline", + }; + const request = vi.fn().mockResolvedValue(answer); + const channel: ElicitationChannel = { request }; + const unwire = wireElicitationBridge(client, channel, "req-1"); + const message = fakeMessage({ + request: { + method: "elicitation/create", + params: { + message: "Confirm?", + requestedSchema: { type: "object", properties: {} }, + }, + }, + }); + emit(message); + await Promise.resolve(); + await Promise.resolve(); + await Promise.resolve(); + + expect(request).toHaveBeenCalledWith( + expect.objectContaining({ + mode: "form", + requestedSchema: { type: "object", properties: {} }, + url: undefined, + }), + ); + expect(message.respond).toHaveBeenCalledWith({ + action: "decline", + content: undefined, + }); + unwire(); + }); + + it("cancels the pending elicitation when the channel rejects", async () => { + const { client, emit } = fakeClient(); + const request = vi.fn().mockRejectedValue(new Error("disconnected")); + const channel: ElicitationChannel = { request }; + const unwire = wireElicitationBridge(client, channel, "req-1"); + const message = fakeMessage(); + emit(message); + await Promise.resolve(); + await Promise.resolve(); + await Promise.resolve(); + + expect(message.cancel).toHaveBeenCalled(); + expect(message.respond).not.toHaveBeenCalled(); + unwire(); + }); + + it("processes multiple elicitations in arrival order (serialized)", async () => { + const { client, emit } = fakeClient(); + const order: string[] = []; + const request = vi.fn().mockImplementation(async (frame) => { + order.push(`start:${frame.elicitationId}`); + await Promise.resolve(); + order.push(`end:${frame.elicitationId}`); + return { + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "cancel", + } satisfies ElicitationResponseFrame; + }); + const channel: ElicitationChannel = { request }; + const unwire = wireElicitationBridge(client, channel, "req-1"); + emit(fakeMessage({ id: "e1" })); + emit(fakeMessage({ id: "e2" })); + await vi.waitFor(() => { + expect(order).toEqual(["start:e1", "end:e1", "start:e2", "end:e2"]); + }); + + unwire(); + }); + + it("delivers each elicitation to exactly one of two concurrent callers (oldest first)", async () => { + const { client, emit } = fakeClient(); + const answerFor = (frame: { + id: string; + elicitationId: string; + }): ElicitationResponseFrame => ({ + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "cancel", + }); + const requestA = vi + .fn() + .mockImplementation(async (frame) => answerFor(frame)); + const requestB = vi + .fn() + .mockImplementation(async (frame) => answerFor(frame)); + const unwireA = wireElicitationBridge( + client, + { request: requestA }, + "req-a", + ); + const unwireB = wireElicitationBridge( + client, + { request: requestB }, + "req-b", + ); + + const first = fakeMessage({ id: "e1" }); + emit(first); + await vi.waitFor(() => expect(first.respond).toHaveBeenCalled()); + expect(requestA).toHaveBeenCalledTimes(1); + expect(requestB).not.toHaveBeenCalled(); + expect(first.respond).toHaveBeenCalledTimes(1); + + // Once the oldest caller settles, the next event goes to the survivor. + unwireA(); + const second = fakeMessage({ id: "e2" }); + emit(second); + await vi.waitFor(() => expect(second.respond).toHaveBeenCalled()); + expect(requestA).toHaveBeenCalledTimes(1); + expect(requestB).toHaveBeenCalledTimes(1); + unwireB(); + }); + + it("cancels an event already queued when every caller settled before dispatch", async () => { + const { client, emit } = fakeClient(); + const request = vi.fn(); + const unwire = wireElicitationBridge(client, { request }, "req-1"); + const message = fakeMessage(); + emit(message); + unwire(); // settle before the queued microtask dispatches + await vi.waitFor(() => expect(message.cancel).toHaveBeenCalled()); + expect(request).not.toHaveBeenCalled(); + expect(message.respond).not.toHaveBeenCalled(); + }); + + it("unwire stops the listener from reacting to further events", () => { + const { client, emit } = fakeClient(); + const channel: ElicitationChannel = { request: vi.fn() }; + const unwire = wireElicitationBridge(client, channel, "req-1"); + unwire(); + emit(fakeMessage()); + expect(channel.request).not.toHaveBeenCalled(); + }); +}); diff --git a/clients/mcpdo/__tests__/elicitation-client.test.ts b/clients/mcpdo/__tests__/elicitation-client.test.ts new file mode 100644 index 0000000000..a1b332bd5a --- /dev/null +++ b/clients/mcpdo/__tests__/elicitation-client.test.ts @@ -0,0 +1,305 @@ +import { describe, it, expect, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as net from "node:net"; +import * as os from "node:os"; +import * as path from "node:path"; +import { callDaemon } from "../src/daemon/client.js"; +import type { + ElicitationRequestFrame, + ElicitationResponseFrame, +} from "../src/daemon/protocol.js"; + +/** + * Covers `callDaemon`'s duplex elicitation handling (dual-era support, phase + * 1): a mid-`rpc` `elicitation-request` frame arriving before the final + * response, answered via `onElicitation` (or auto-cancelled without one), + * with the connect timeout cleared once the exchange starts. + */ +describe("callDaemon elicitation duplex", () => { + let dir: string | undefined; + let server: net.Server | undefined; + const sockets = new Set<net.Socket>(); + + afterEach(async () => { + for (const s of sockets) s.destroy(); + sockets.clear(); + if (server) { + await new Promise<void>((resolve) => server!.close(() => resolve())); + server = undefined; + } + if (dir) { + fs.rmSync(dir, { recursive: true, force: true }); + dir = undefined; + } + }); + + function freshSock(): string { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-elicit-client-")); + return path.join(dir, "mcpdod.sock"); + } + + async function listen( + sock: string, + onSocket: (socket: net.Socket) => void, + ): Promise<void> { + server = net.createServer((socket) => { + sockets.add(socket); + socket.on("error", () => {}); + socket.on("close", () => sockets.delete(socket)); + onSocket(socket); + }); + await new Promise<void>((resolve) => server!.listen(sock, resolve)); + } + + it("routes an elicitation-request frame to onElicitation and writes its answer", async () => { + const sock = freshSock(); + let receivedAnswer: ElicitationResponseFrame | undefined; + await listen(sock, (socket) => { + let buffer = ""; + socket.on("data", (chunk) => { + buffer += String(chunk); + let idx: number; + while ((idx = buffer.indexOf("\n")) >= 0) { + const line = buffer.slice(0, idx); + buffer = buffer.slice(idx + 1); + if (!line.trim()) continue; + const msg = JSON.parse(line) as { id: string; kind?: string }; + if (msg.kind === "elicitation-response") { + receivedAnswer = msg as ElicitationResponseFrame; + socket.write( + JSON.stringify({ id: msg.id, ok: true, result: { done: true } }) + + "\n", + ); + continue; + } + const frame: ElicitationRequestFrame = { + id: msg.id, + kind: "elicitation-request", + elicitationId: "elicitation-1", + mode: "url", + message: "Please confirm", + url: "https://example.com/confirm", + origin: "server-request", + }; + socket.write(JSON.stringify(frame) + "\n"); + } + }); + }); + + const seenFrames: ElicitationRequestFrame[] = []; + const result = await callDaemon<{ done: boolean }>( + "rpc", + { method: "tools/call" }, + { + socketPath: sock, + timeoutMs: 5000, + onElicitation: async (frame) => { + seenFrames.push(frame); + return { + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "accept", + }; + }, + }, + ); + + expect(result).toEqual({ done: true }); + expect(seenFrames).toHaveLength(1); + expect(seenFrames[0].mode).toBe("url"); + expect(seenFrames[0].url).toBe("https://example.com/confirm"); + expect(receivedAnswer?.action).toBe("accept"); + expect(receivedAnswer?.elicitationId).toBe("elicitation-1"); + }); + + it("auto-cancels when no onElicitation callback is provided", async () => { + const sock = freshSock(); + let receivedAnswer: ElicitationResponseFrame | undefined; + await listen(sock, (socket) => { + let buffer = ""; + socket.on("data", (chunk) => { + buffer += String(chunk); + let idx: number; + while ((idx = buffer.indexOf("\n")) >= 0) { + const line = buffer.slice(0, idx); + buffer = buffer.slice(idx + 1); + if (!line.trim()) continue; + const msg = JSON.parse(line) as { id: string; kind?: string }; + if (msg.kind === "elicitation-response") { + receivedAnswer = msg as ElicitationResponseFrame; + socket.write( + JSON.stringify({ id: msg.id, ok: true, result: { done: true } }) + + "\n", + ); + continue; + } + const frame: ElicitationRequestFrame = { + id: msg.id, + kind: "elicitation-request", + elicitationId: "elicitation-2", + mode: "url", + message: "Please confirm", + url: "https://example.com/confirm", + origin: "input-required", + }; + socket.write(JSON.stringify(frame) + "\n"); + } + }); + }); + + const result = await callDaemon<{ done: boolean }>( + "rpc", + { method: "tools/call" }, + { socketPath: sock, timeoutMs: 5000 }, + ); + + expect(result).toEqual({ done: true }); + expect(receivedAnswer?.action).toBe("cancel"); + expect(receivedAnswer?.elicitationId).toBe("elicitation-2"); + }); + + it("ignores an elicitation-request frame whose id doesn't match this call", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + let buffer = ""; + let answered = false; + socket.on("data", (chunk) => { + buffer += String(chunk); + let idx: number; + while ((idx = buffer.indexOf("\n")) >= 0) { + const line = buffer.slice(0, idx); + buffer = buffer.slice(idx + 1); + if (!line.trim()) continue; + const msg = JSON.parse(line) as { id: string }; + if (!answered) { + answered = true; + const frame: ElicitationRequestFrame = { + id: "not-this-call", + kind: "elicitation-request", + elicitationId: "elicitation-3", + mode: "url", + message: "stray frame", + url: "https://example.com", + origin: "server-request", + }; + socket.write(JSON.stringify(frame) + "\n"); + socket.write( + JSON.stringify({ id: msg.id, ok: true, result: { done: true } }) + + "\n", + ); + } + } + }); + }); + + const onElicitation = async () => + ({ + id: "n/a", + kind: "elicitation-response", + elicitationId: "n/a", + action: "cancel", + }) satisfies ElicitationResponseFrame; + + const result = await callDaemon<{ done: boolean }>( + "rpc", + { method: "tools/call" }, + { socketPath: sock, timeoutMs: 5000, onElicitation }, + ); + expect(result).toEqual({ done: true }); + }); + + it("fails the call if onElicitation itself throws", async () => { + const sock = freshSock(); + await listen(sock, (socket) => { + let buffer = ""; + socket.on("data", (chunk) => { + buffer += String(chunk); + let idx: number; + while ((idx = buffer.indexOf("\n")) >= 0) { + const line = buffer.slice(0, idx); + buffer = buffer.slice(idx + 1); + if (!line.trim()) continue; + const msg = JSON.parse(line) as { id: string; kind?: string }; + if (msg.kind === "elicitation-response") continue; + const frame: ElicitationRequestFrame = { + id: msg.id, + kind: "elicitation-request", + elicitationId: "elicitation-4", + mode: "url", + message: "boom", + url: "https://example.com", + origin: "server-request", + }; + socket.write(JSON.stringify(frame) + "\n"); + } + }); + }); + + await expect( + callDaemon<{ done: boolean }>( + "rpc", + { method: "tools/call" }, + { + socketPath: sock, + timeoutMs: 5000, + onElicitation: async () => { + throw new Error("prompt blew up"); + }, + }, + ), + ).rejects.toThrow("prompt blew up"); + }); + + it("fails with a clear cancellation error when the abort signal fires mid-call", async () => { + const sock = freshSock(); + await listen(sock, () => { + // Never respond — the call should hang until aborted, not until + // timeoutMs, proving the signal (not the timeout) ended it. + }); + + const ac = new AbortController(); + const promise = callDaemon( + "rpc", + { method: "tools/call" }, + { socketPath: sock, timeoutMs: 60_000, signal: ac.signal }, + ); + ac.abort(); + await expect(promise).rejects.toThrow("cancelled"); + }); + + it("silently swallows a post-settle socket error (e.g. late ECONNRESET)", async () => { + const sock = freshSock(); + let serverSocket: net.Socket | undefined; + await listen(sock, (socket) => { + serverSocket = socket; + let buffer = ""; + socket.on("data", (chunk) => { + buffer += String(chunk); + let idx: number; + while ((idx = buffer.indexOf("\n")) >= 0) { + const line = buffer.slice(0, idx); + buffer = buffer.slice(idx + 1); + if (!line.trim()) continue; + const msg = JSON.parse(line) as { id: string }; + socket.write( + JSON.stringify({ id: msg.id, ok: true, result: { done: true } }) + + "\n", + ); + } + }); + }); + + const result = await callDaemon<{ done: boolean }>( + "rpc", + { method: "tools/call" }, + { socketPath: sock, timeoutMs: 5000 }, + ); + expect(result).toEqual({ done: true }); + // Force a client-side 'error' after the call already settled; the + // no-op listener installed by settle() must swallow it without + // rethrowing or crashing the test. + serverSocket?.destroy(new Error("late reset")); + await new Promise((resolve) => setTimeout(resolve, 50)); + }); +}); diff --git a/clients/mcpdo/__tests__/elicitation-prompt.test.ts b/clients/mcpdo/__tests__/elicitation-prompt.test.ts new file mode 100644 index 0000000000..25391e598c --- /dev/null +++ b/clients/mcpdo/__tests__/elicitation-prompt.test.ts @@ -0,0 +1,274 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import { createStyle } from "@inspector/core/cli/style.js"; +import type { ElicitationRequestFrame } from "../src/daemon/protocol.js"; + +const question = vi.fn(); +const promptFormMock = vi.fn(); + +const once = vi.fn(); + +vi.mock("../src/connection/prompt-reader.js", () => ({ + getSharedPromptReader: () => ({ question, once }), +})); + +vi.mock("../src/connection/form-prompt.js", async () => { + const actual = await vi.importActual< + typeof import("../src/connection/form-prompt.js") + >("../src/connection/form-prompt.js"); + return { + promptForm: (...args: unknown[]) => promptFormMock(...args), + watchForClose: actual.watchForClose, + }; +}); + +/** + * Covers `promptElicitation`'s terminal UI: form mode always declines + * (rendering isn't built yet), non-interactive callers auto-cancel with a + * clear message instead of hanging, and interactive URL mode reads the + * user's accept/cancel choice. + */ +describe("promptElicitation", () => { + let stderr: string; + let originalWrite: typeof process.stderr.write; + + beforeEach(() => { + stderr = ""; + originalWrite = process.stderr.write; + process.stderr.write = ((chunk: unknown, ...rest: unknown[]) => { + stderr += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stderr.write; + question.mockReset(); + once.mockReset(); + promptFormMock.mockReset(); + }); + + afterEach(() => { + process.stderr.write = originalWrite; + }); + + const style = createStyle(false); + + function urlFrame( + overrides: Partial<ElicitationRequestFrame> = {}, + ): ElicitationRequestFrame { + return { + id: "req-1", + kind: "elicitation-request", + elicitationId: "elicitation-1", + mode: "url", + message: "Please confirm", + url: "https://example.com/confirm", + origin: "server-request", + ...overrides, + }; + } + + function formFrame( + overrides: Partial<ElicitationRequestFrame> = {}, + ): ElicitationRequestFrame { + return { + id: "req-1", + kind: "elicitation-request", + elicitationId: "elicitation-1", + mode: "form", + message: "Please provide your name", + requestedSchema: { + type: "object", + properties: { name: { type: "string" } }, + required: ["name"], + }, + origin: "server-request", + ...overrides, + }; + } + + it("declines form-mode elicitations whose schema isn't the restricted primitive shape", async () => { + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame({ mode: "form", url: undefined }); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer).toEqual({ + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-1", + action: "decline", + }); + expect(question).not.toHaveBeenCalled(); + expect(stderr).toContain("doesn't support"); + }); + + it("declines form-mode elicitations non-interactively without prompting", async () => { + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame({ + mode: "form", + url: undefined, + requestedSchema: { + type: "object", + properties: { name: { type: "string" } }, + }, + }); + const answer = await promptElicitation(frame, { + interactive: false, + style, + }); + expect(answer).toEqual({ + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-1", + action: "decline", + }); + expect(question).not.toHaveBeenCalled(); + expect(stderr).toContain("--format json"); + }); + + it("cancels when the caller isn't interactive (e.g. --format json) without prompting", async () => { + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame(); + const answer = await promptElicitation(frame, { + interactive: false, + style, + }); + expect(answer).toEqual({ + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-1", + action: "cancel", + }); + expect(question).not.toHaveBeenCalled(); + expect(stderr).toContain("--format json"); + }); + + it("cancels non-interactively without a url line when the frame has none", async () => { + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame({ url: undefined }); + const answer = await promptElicitation(frame, { + interactive: false, + style, + }); + expect(answer.action).toBe("cancel"); + expect(stderr).not.toContain("undefined"); + }); + + it("accepts when the interactive user confirms completion", async () => { + question.mockResolvedValue(""); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer).toEqual({ + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-1", + action: "accept", + }); + expect(stderr).toContain("Please confirm"); + expect(stderr).toContain("https://example.com/confirm"); + }); + + it("renders only allowlisted schemes as OSC 8 links in URL mode", async () => { + question.mockResolvedValue(""); + const ansi = createStyle(true); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + + await promptElicitation(urlFrame(), { interactive: true, style: ansi }); + expect(stderr).toContain("\u001b]8;;https://example.com/confirm"); + + stderr = ""; + const answer = await promptElicitation( + urlFrame({ url: "file:///etc/passwd" }), + { interactive: true, style: ansi }, + ); + // A server-supplied file:/custom-handler URL is shown as plain text — + // never as a clickable link inviting the local protocol handler. + expect(answer.action).toBe("accept"); + expect(stderr).not.toContain("]8;;"); + expect(stderr).toContain("file:///etc/passwd"); + }); + + it("cancels when the interactive user types 'c'", async () => { + question.mockResolvedValue("c"); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer.action).toBe("cancel"); + }); + + it("falls back to cancel if reading input throws", async () => { + question.mockRejectedValue(new Error("stdin closed")); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer.action).toBe("cancel"); + }); + + it("cancels URL mode if stdin closes before the user answers", async () => { + // Simulates a non-TTY stdin (e.g. an agent-driven pipe) hitting EOF + // before an answer arrives: question() hangs, but the "close" listener + // registered via watchForClose() fires and wins the race. + question.mockImplementation(() => new Promise(() => {})); + once.mockImplementation((event: string, cb: () => void) => { + if (event === "close") cb(); + }); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = urlFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer.action).toBe("cancel"); + }); + + it("accepts an interactive form submission and returns its content", async () => { + promptFormMock.mockResolvedValue({ + action: "accept", + content: { name: "octocat" }, + }); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = formFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer).toEqual({ + id: "req-1", + kind: "elicitation-response", + elicitationId: "elicitation-1", + action: "accept", + content: { name: "octocat" }, + }); + }); + + it("declines an interactive form when promptForm reports decline", async () => { + promptFormMock.mockResolvedValue({ action: "decline" }); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = formFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer.action).toBe("decline"); + }); + + it("cancels an interactive form when promptForm reports cancel", async () => { + promptFormMock.mockResolvedValue({ action: "cancel" }); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = formFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer.action).toBe("cancel"); + }); + + it("falls back to cancel if promptForm throws", async () => { + promptFormMock.mockRejectedValue(new Error("stdin closed")); + const { promptElicitation } = + await import("../src/connection/elicitation-prompt.js"); + const frame = formFrame(); + const answer = await promptElicitation(frame, { interactive: true, style }); + expect(answer.action).toBe("cancel"); + }); +}); diff --git a/clients/mcpdo/__tests__/ema-commands.test.ts b/clients/mcpdo/__tests__/ema-commands.test.ts new file mode 100644 index 0000000000..078ec9b371 --- /dev/null +++ b/clients/mcpdo/__tests__/ema-commands.test.ts @@ -0,0 +1,245 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import { PLAIN } from "@inspector/core/cli/style.js"; +import { formatEmaStatusHuman } from "../src/connection/format-human.js"; + +const getEmaStatus = vi.fn(); +const emaLogin = vi.fn(); +const emaLogout = vi.fn(); +const startPendingEmaLogin = vi.fn(); + +vi.mock("../src/connection/ema.js", () => ({ + getEmaStatus: (...args: unknown[]) => getEmaStatus(...args), + emaLogin: (...args: unknown[]) => emaLogin(...args), + emaLogout: (...args: unknown[]) => emaLogout(...args), +})); + +// Mocked wholesale: the real module imports ema.js (mocked above, missing the +// loadEmaIdpConfig/requireIdp/runEmaIdpInteractiveFlow exports it needs) and +// auth-helper.js. EMA_LOGIN_HELPER_COMMAND must be the real string — mcp.ts +// registers the hidden command under it. +vi.mock("../src/connection/ema-login-helper.js", () => ({ + EMA_LOGIN_HELPER_COMMAND: "auth/complete-ema-login", + runEmaLoginHelper: vi.fn(), + startPendingEmaLogin: (...args: unknown[]) => startPendingEmaLogin(...args), +})); + +describe("auth/ema-* commands", () => { + let stdout: string; + let originalStdoutWrite: typeof process.stdout.write; + const originalStderrIsTTY = process.stderr.isTTY; + const originalStdinIsTTY = process.stdin.isTTY; + + beforeEach(() => { + stdout = ""; + // Interactive path by default; individual tests flip to the non-TTY park. + process.stderr.isTTY = true; + originalStdoutWrite = process.stdout.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]) => { + stdout += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + getEmaStatus.mockReset(); + emaLogin.mockReset(); + emaLogout.mockReset(); + startPendingEmaLogin.mockReset(); + }); + + afterEach(() => { + process.stderr.isTTY = originalStderrIsTTY; + process.stdin.isTTY = originalStdinIsTTY; + process.stdout.write = originalStdoutWrite; + }); + + it("auth/ema-status prints the status as JSON", async () => { + getEmaStatus.mockResolvedValue({ + clientConfigPath: "/tmp/client.json", + configured: true, + enabled: true, + issuer: "https://idp.example.com", + clientId: "idp-client", + loginState: "logged_in", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-status", "--format", "json"]); + const parsed = JSON.parse(stdout.trim()); + expect(parsed.issuer).toBe("https://idp.example.com"); + expect(parsed.loginState).toBe("logged_in"); + }); + + it("auth/ema-status prints a human summary in text mode", async () => { + getEmaStatus.mockResolvedValue({ + clientConfigPath: "/tmp/client.json", + configured: true, + enabled: true, + issuer: "https://idp.example.com", + clientId: "idp-client", + loginState: "none", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-status"]); + expect(stdout).toContain("EMA (enterprise-managed auth):"); + expect(stdout).toContain("https://idp.example.com"); + expect(stdout).toContain("IdP session: none"); + }); + + it("auth/ema-login forwards --relogin and prints the outcome", async () => { + emaLogin.mockResolvedValue({ + issuer: "https://idp.example.com", + loginState: "logged_in", + alreadyLoggedIn: false, + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-login", "--relogin"]); + expect(emaLogin).toHaveBeenCalledWith({ relogin: true }); + expect(stdout).toContain("Signed in"); + expect(stdout).toContain("https://idp.example.com"); + }); + + it("auth/ema-login reports an already-active connection", async () => { + emaLogin.mockResolvedValue({ + issuer: "https://idp.example.com", + loginState: "logged_in", + alreadyLoggedIn: true, + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-login"]); + expect(emaLogin).toHaveBeenCalledWith({ relogin: false }); + expect(stdout).toContain("Already signed in"); + }); + + it("auth/ema-logout prints the signed-out issuer", async () => { + emaLogout.mockResolvedValue({ issuer: "https://idp.example.com" }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-logout"]); + expect(stdout).toContain("Signed out"); + expect(stdout).toContain("https://idp.example.com"); + expect(stdout).not.toContain("IdP browser session"); + }); + + it("auth/ema-logout relays the IdP end-session URL when present", async () => { + emaLogout.mockResolvedValue({ + issuer: "https://idp.example.com", + endSessionUrl: "https://idp.example.com/session/end?id_token_hint=a.b.c", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-logout"]); + expect(stdout).toContain("Signed out"); + expect(stdout).toContain("To end your IdP browser session, navigate to:"); + expect(stdout).toContain( + "https://idp.example.com/session/end?id_token_hint=a.b.c", + ); + }); + + it("auth/ema-login parks on a detached helper when no TTY is present", async () => { + process.stderr.isTTY = undefined as unknown as boolean; + process.stdin.isTTY = undefined as unknown as boolean; + startPendingEmaLogin.mockResolvedValue({ + issuer: "https://idp.example.com", + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl: "https://idp.example.com/authorize?state=abc", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-login"]); + expect(startPendingEmaLogin).toHaveBeenCalledWith({ relogin: false }); + expect(emaLogin).not.toHaveBeenCalled(); + expect(stdout).toContain("Sign-in required"); + expect(stdout).toContain("https://idp.example.com/authorize?state=abc"); + expect(stdout).toContain("auth/ema-status"); + // Non-TTY: agent relay framing, same split as the connect surface. + expect(stdout).toContain("not usable yet until the user signs in"); + expect(stdout).toContain("wait for them to confirm"); + }); + + it("auth/ema-login non-TTY forwards --relogin and emits JSON with the authUrl", async () => { + process.stderr.isTTY = undefined as unknown as boolean; + process.stdin.isTTY = undefined as unknown as boolean; + startPendingEmaLogin.mockResolvedValue({ + issuer: "https://idp.example.com", + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl: "https://idp.example.com/authorize?state=abc", + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "auth/ema-login", + "--relogin", + "--format", + "json", + ]); + expect(startPendingEmaLogin).toHaveBeenCalledWith({ relogin: true }); + const parsed = JSON.parse(stdout.trim()); + expect(parsed.pendingLogin).toBe(true); + expect(parsed.authUrl).toBe("https://idp.example.com/authorize?state=abc"); + }); + + it("auth/ema-login non-TTY short-circuit renders as already signed in", async () => { + process.stderr.isTTY = undefined as unknown as boolean; + process.stdin.isTTY = undefined as unknown as boolean; + startPendingEmaLogin.mockResolvedValue({ + issuer: "https://idp.example.com", + loginState: "logged_in", + alreadyLoggedIn: true, + }); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp(["node", "mcpdo", "auth/ema-login"]); + expect(stdout).toContain("Already signed in"); + }); +}); + +describe("formatEmaStatusHuman", () => { + it("renders the unconfigured state with configuration pointers", () => { + const text = formatEmaStatusHuman( + { clientConfigPath: "/tmp/client.json", configured: false }, + PLAIN, + ); + expect(text).toContain("not configured"); + expect(text).toContain("/tmp/client.json"); + }); + + it("renders a configured, disabled IdP without a clientId", () => { + const text = formatEmaStatusHuman( + { + clientConfigPath: "/tmp/client.json", + configured: true, + enabled: false, + issuer: "https://idp.example.com", + loginState: "expired", + }, + PLAIN, + ); + expect(text).toContain("https://idp.example.com"); + expect(text).toContain("Enabled: no"); + expect(text).toContain("IdP session: expired"); + expect(text).not.toContain("client:"); + }); + + it("highlights a live IdP session and defaults missing fields", () => { + const loggedIn = formatEmaStatusHuman( + { + clientConfigPath: "/tmp/client.json", + configured: true, + enabled: true, + issuer: "https://idp.example.com", + clientId: "idp-client", + loginState: "logged_in", + }, + PLAIN, + ); + expect(loggedIn).toContain("IdP session: logged_in"); + expect(loggedIn).toContain("(client: idp-client)"); + + // Defensive fallbacks when a JSON payload omits optional fields. + const sparse = formatEmaStatusHuman({ configured: true }, PLAIN); + expect(sparse).toContain("IdP: `?`"); + expect(sparse).toContain("IdP session: none"); + }); +}); diff --git a/clients/mcpdo/__tests__/ema-login-helper.test.ts b/clients/mcpdo/__tests__/ema-login-helper.test.ts new file mode 100644 index 0000000000..c1f358c644 --- /dev/null +++ b/clients/mcpdo/__tests__/ema-login-helper.test.ts @@ -0,0 +1,311 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; + +const runRunnerInteractiveOAuth = vi.fn(); +const startIdpOidcAuthorization = vi.fn(); +const completeIdpOidcAuthorization = vi.fn(); + +vi.mock("@inspector/core/auth/node/index.js", async (importOriginal) => { + const actual = + await importOriginal<typeof import("@inspector/core/auth/node/index.js")>(); + return { + ...actual, + runRunnerInteractiveOAuth: (...args: unknown[]) => + runRunnerInteractiveOAuth(...args), + }; +}); + +vi.mock("@inspector/core/auth/ema/idpOidc.js", () => ({ + startIdpOidcAuthorization: (...args: unknown[]) => + startIdpOidcAuthorization(...args), + completeIdpOidcAuthorization: (...args: unknown[]) => + completeIdpOidcAuthorization(...args), +})); + +import { + NodeOAuthStorage, + resetNodeOAuthStorageCache, +} from "@inspector/core/auth/node/storage-node.js"; +import { pendingAuthMarkerPath } from "../src/connection/auth-helper.js"; +import { + EMA_LOGIN_HELPER_COMMAND, + emaLoginMarkerKey, + runEmaLoginHelper, + startPendingEmaLogin, +} from "../src/connection/ema-login-helper.js"; + +const ISSUER = "https://idp.example.com"; + +/** Unexpired unsigned JWT ({ exp } one hour out). */ +function fakeIdToken(): string { + const b64 = (obj: object) => + Buffer.from(JSON.stringify(obj)).toString("base64url"); + return `${b64({ alg: "none" })}.${b64({ + exp: Math.floor(Date.now() / 1000) + 3600, + })}.sig`; +} + +describe("ema-login-helper", () => { + let dir: string; + let savedEnv: Record<string, string | undefined>; + + beforeEach(() => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-ema-login-helper-")); + savedEnv = { + MCP_CLIENT_CONFIG_PATH: process.env.MCP_CLIENT_CONFIG_PATH, + MCP_INSPECTOR_OAUTH_STATE_PATH: + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH, + MCP_INSPECTOR_DAEMON_DIR: process.env.MCP_INSPECTOR_DAEMON_DIR, + }; + process.env.MCP_CLIENT_CONFIG_PATH = path.join(dir, "client.json"); + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = path.join(dir, "oauth.json"); + process.env.MCP_INSPECTOR_DAEMON_DIR = dir; + resetNodeOAuthStorageCache(); + runRunnerInteractiveOAuth.mockReset(); + startIdpOidcAuthorization.mockReset(); + completeIdpOidcAuthorization.mockReset(); + }); + + afterEach(() => { + for (const [key, value] of Object.entries(savedEnv)) { + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + resetNodeOAuthStorageCache(); + fs.rmSync(dir, { recursive: true, force: true }); + }); + + function writeClientConfig(): void { + fs.writeFileSync( + process.env.MCP_CLIENT_CONFIG_PATH!, + JSON.stringify({ + enterpriseManagedAuth: { + idp: { + issuer: ISSUER, + clientId: "idp-client", + clientSecret: "idp-secret", + }, + }, + }), + ); + } + + async function seedIdpSession(): Promise<void> { + const storage = new NodeOAuthStorage(); + await storage.saveIdpSession(ISSUER, { + idToken: fakeIdToken(), + idTokenExpiresAt: Date.now() + 3600_000, + }); + } + + function writeHelperScript(body: string): string { + const script = path.join(dir, "fake-ema-helper.mjs"); + fs.writeFileSync(script, body); + return script; + } + + /** Fake helper emitting one auth_url event after confirming its argv. */ + function urlEmittingScript(url: string): string { + return writeHelperScript(` + process.stdin.resume(); + process.stdin.on("end", () => { + if (process.argv[2] !== ${JSON.stringify(EMA_LOGIN_HELPER_COMMAND)}) { + process.exit(9); + } + process.stdout.write( + JSON.stringify({ event: "auth_url", url: ${JSON.stringify(url)} }) + "\\n", + ); + }); + `); + } + + it("emaLoginMarkerKey namespaces the issuer", () => { + expect(emaLoginMarkerKey(ISSUER)).toBe(`ema-idp:${ISSUER}`); + }); + + describe("startPendingEmaLogin", () => { + it("fails with mcpdo guidance when EMA is not configured", async () => { + await expect(startPendingEmaLogin()).rejects.toThrow( + /not configured.*client settings/is, + ); + }); + + it("short-circuits when already signed in without spawning a helper", async () => { + writeClientConfig(); + await seedIdpSession(); + const result = await startPendingEmaLogin({ + // Would fail loudly if a spawn were attempted. + helperArgv1: path.join(dir, "does-not-exist.mjs"), + }); + expect(result).toEqual({ + issuer: ISSUER, + loginState: "logged_in", + alreadyLoggedIn: true, + }); + }); + + it("parks the login on a detached helper and returns its URL", async () => { + writeClientConfig(); + const result = await startPendingEmaLogin({ + helperArgv1: urlEmittingScript("https://idp.example.com/authorize?s=1"), + }); + expect(result).toMatchObject({ + issuer: ISSUER, + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl: "https://idp.example.com/authorize?s=1", + }); + }); + + it("--relogin clears the existing IdP session before parking", async () => { + writeClientConfig(); + await seedIdpSession(); + const result = await startPendingEmaLogin({ + relogin: true, + helperArgv1: urlEmittingScript("https://idp.example.com/authorize?s=2"), + }); + expect(result).toMatchObject({ + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl: "https://idp.example.com/authorize?s=2", + }); + }); + + it("reuses a live pending-login marker instead of spawning", async () => { + writeClientConfig(); + fs.writeFileSync( + pendingAuthMarkerPath(emaLoginMarkerKey(ISSUER)), + JSON.stringify({ + url: "https://idp.example.com/authorize?s=reuse", + pid: process.pid, + expiresAt: Date.now() + 60_000, + }), + { mode: 0o600 }, + ); + const result = await startPendingEmaLogin({ + // Would fail loudly if a spawn were attempted. + helperArgv1: path.join(dir, "does-not-exist.mjs"), + }); + expect(result.authUrl).toBe("https://idp.example.com/authorize?s=reuse"); + }); + }); + + describe("runEmaLoginHelper", () => { + let stdout: string; + let originalStdoutWrite: typeof process.stdout.write; + + beforeEach(() => { + stdout = ""; + originalStdoutWrite = process.stdout.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]) => { + stdout += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + }); + + afterEach(() => { + process.stdout.write = originalStdoutWrite; + }); + + function events(): Array<Record<string, unknown>> { + return stdout + .split("\n") + .filter((line) => line.trim() !== "") + .map((line) => JSON.parse(line) as Record<string, unknown>); + } + + it("publishes the marker, emits auth_url + done, and cleans up", async () => { + writeClientConfig(); + const markerPath = pendingAuthMarkerPath(emaLoginMarkerKey(ISSUER)); + startIdpOidcAuthorization.mockResolvedValue({ + authorizationUrl: new URL("https://idp.example.com/authorize?x=1"), + }); + completeIdpOidcAuthorization.mockImplementation(async () => { + // Mid-flow: the marker must already be on disk for repeat callers. + expect(fs.existsSync(markerPath)).toBe(true); + await seedIdpSession(); + return { idToken: fakeIdToken() }; + }); + runRunnerInteractiveOAuth.mockImplementation( + async (options: { + client: { + authenticate: () => Promise<URL | undefined>; + completeOAuthFlow: (code: string, iss?: string) => Promise<void>; + }; + redirectUrlProvider: { redirectUrl: string }; + }) => { + options.redirectUrlProvider.redirectUrl = + "http://127.0.0.1:45678/oauth/callback"; + const url = await options.client.authenticate(); + expect(url?.href).toContain("idp.example.com/authorize"); + await options.client.completeOAuthFlow("code-1", ISSUER); + return { kind: "success" }; + }, + ); + + await runEmaLoginHelper(); + // The EPIPE guard on stdout must swallow late write errors. + process.stdout.emit("error", new Error("EPIPE")); + + expect(events()).toEqual([ + { event: "auth_url", url: "https://idp.example.com/authorize?x=1" }, + { event: "done" }, + ]); + // Own marker removed on exit. + expect(fs.existsSync(markerPath)).toBe(false); + expect(startIdpOidcAuthorization).toHaveBeenCalledWith( + expect.objectContaining({ + redirectUrl: "http://127.0.0.1:45678/oauth/callback", + }), + ); + }); + + it("emits an error event and rethrows when EMA is not configured", async () => { + await expect(runEmaLoginHelper()).rejects.toThrow(/not configured/i); + expect(events()).toEqual([ + { + event: "error", + message: expect.stringContaining("not configured") as string, + }, + ]); + }); + + it("stringifies a non-Error flow failure in the error event", async () => { + writeClientConfig(); + runRunnerInteractiveOAuth.mockRejectedValueOnce("string boom"); + await expect(runEmaLoginHelper()).rejects.toBe("string boom"); + expect(events()).toEqual([{ event: "error", message: "string boom" }]); + }); + + it("survives a gone parent (stdout writes that throw)", async () => { + writeClientConfig(); + process.stdout.write = (() => { + throw new Error("EPIPE"); + }) as unknown as typeof process.stdout.write; + startIdpOidcAuthorization.mockResolvedValue({ + authorizationUrl: new URL("https://idp.example.com/authorize?x=2"), + }); + completeIdpOidcAuthorization.mockImplementation(async () => { + await seedIdpSession(); + return { idToken: fakeIdToken() }; + }); + runRunnerInteractiveOAuth.mockImplementation( + async (options: { + client: { authenticate: () => Promise<URL | undefined> }; + }) => { + await options.client.authenticate(); + return { kind: "success" }; + }, + ); + await expect(runEmaLoginHelper()).resolves.toBeUndefined(); + }); + }); +}); diff --git a/clients/mcpdo/__tests__/ema.test.ts b/clients/mcpdo/__tests__/ema.test.ts new file mode 100644 index 0000000000..6a68aa5641 --- /dev/null +++ b/clients/mcpdo/__tests__/ema.test.ts @@ -0,0 +1,302 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { CliExitCodeError } from "@inspector/core/cli/error-handler.js"; +import { + NodeOAuthStorage, + resetNodeOAuthStorageCache, +} from "@inspector/core/auth/node/storage-node.js"; + +const runRunnerInteractiveOAuth = vi.fn(); +const startIdpOidcAuthorization = vi.fn(); +const completeIdpOidcAuthorization = vi.fn(); + +vi.mock("@inspector/core/auth/node/index.js", async (importOriginal) => { + const actual = + await importOriginal<typeof import("@inspector/core/auth/node/index.js")>(); + return { + ...actual, + runRunnerInteractiveOAuth: (...args: unknown[]) => + runRunnerInteractiveOAuth(...args), + }; +}); + +vi.mock("@inspector/core/auth/ema/idpOidc.js", () => ({ + startIdpOidcAuthorization: (...args: unknown[]) => + startIdpOidcAuthorization(...args), + completeIdpOidcAuthorization: (...args: unknown[]) => + completeIdpOidcAuthorization(...args), +})); + +const ISSUER = "https://idp.example.com"; + +/** Unexpired unsigned JWT ({ exp } one hour out). */ +function fakeIdToken(): string { + const b64 = (obj: object) => + Buffer.from(JSON.stringify(obj)).toString("base64url"); + return `${b64({ alg: "none" })}.${b64({ + exp: Math.floor(Date.now() / 1000) + 3600, + })}.sig`; +} + +describe("mcpdo ema helpers", () => { + let dir: string; + let savedEnv: Record<string, string | undefined>; + + beforeEach(() => { + dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-conn-ema-")); + savedEnv = { + MCP_CLIENT_CONFIG_PATH: process.env.MCP_CLIENT_CONFIG_PATH, + MCP_INSPECTOR_OAUTH_STATE_PATH: + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH, + }; + process.env.MCP_CLIENT_CONFIG_PATH = path.join(dir, "client.json"); + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = path.join(dir, "oauth.json"); + resetNodeOAuthStorageCache(); + runRunnerInteractiveOAuth.mockReset(); + startIdpOidcAuthorization.mockReset(); + completeIdpOidcAuthorization.mockReset(); + }); + + afterEach(() => { + for (const [key, value] of Object.entries(savedEnv)) { + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + resetNodeOAuthStorageCache(); + fs.rmSync(dir, { recursive: true, force: true }); + }); + + function writeClientConfig(config: unknown): void { + fs.writeFileSync( + process.env.MCP_CLIENT_CONFIG_PATH!, + JSON.stringify(config), + ); + } + + function emaClientConfig(enabled?: boolean): unknown { + return { + enterpriseManagedAuth: { + ...(enabled !== undefined && { enabled }), + idp: { + issuer: ISSUER, + clientId: "idp-client", + clientSecret: "idp-secret", + }, + }, + }; + } + + async function seedIdpSession(): Promise<void> { + const storage = new NodeOAuthStorage(); + await storage.saveIdpSession(ISSUER, { + idToken: fakeIdToken(), + idTokenExpiresAt: Date.now() + 3600_000, + }); + } + + it("getEmaStatus reports unconfigured when client.json has no EMA block", async () => { + const { getEmaStatus } = await import("../src/connection/ema.js"); + const status = await getEmaStatus(); + expect(status.configured).toBe(false); + expect(status.enabled).toBe(false); + expect(status.loginState).toBe("unconfigured"); + expect(status.clientConfigPath).toBe(process.env.MCP_CLIENT_CONFIG_PATH); + }); + + it("getEmaStatus reports configured+enabled with no IdP session as 'none'", async () => { + writeClientConfig(emaClientConfig()); + const { getEmaStatus } = await import("../src/connection/ema.js"); + const status = await getEmaStatus(); + expect(status.configured).toBe(true); + expect(status.enabled).toBe(true); + expect(status.issuer).toBe(ISSUER); + expect(status.clientId).toBe("idp-client"); + expect(status.loginState).toBe("none"); + }); + + it("getEmaStatus reports a disabled config (still shows issuer + connection state)", async () => { + writeClientConfig(emaClientConfig(false)); + await seedIdpSession(); + const { getEmaStatus } = await import("../src/connection/ema.js"); + const status = await getEmaStatus(); + expect(status.configured).toBe(true); + expect(status.enabled).toBe(false); + expect(status.loginState).toBe("logged_in"); + }); + + it("emaLogin fails with actionable guidance when EMA is not configured", async () => { + const { emaLogin } = await import("../src/connection/ema.js"); + await expect(emaLogin()).rejects.toThrow( + /not configured.*client settings/is, + ); + await expect(emaLogin()).rejects.toThrow( + process.env.MCP_CLIENT_CONFIG_PATH!, + ); + }); + + it("emaLogin fails with actionable guidance when EMA is disabled", async () => { + writeClientConfig(emaClientConfig(false)); + const { emaLogin } = await import("../src/connection/ema.js"); + await expect(emaLogin()).rejects.toThrow(/disabled/i); + }); + + it("emaLogout fails when EMA is not configured", async () => { + const { emaLogout } = await import("../src/connection/ema.js"); + await expect(emaLogout()).rejects.toThrow(CliExitCodeError); + }); + + it("emaLogout works even when EMA is disabled, and clears the IdP session", async () => { + writeClientConfig(emaClientConfig(false)); + await seedIdpSession(); + const { emaLogout, getEmaStatus } = + await import("../src/connection/ema.js"); + const result = await emaLogout(); + expect(result.issuer).toBe(ISSUER); + expect(result.endSessionUrl).toBeUndefined(); + expect((await getEmaStatus()).loginState).toBe("none"); + }); + + it("emaLogout returns the IdP end-session URL when the IdP advertises one", async () => { + writeClientConfig(emaClientConfig()); + await seedIdpSession(); + // Login-time discovery caches the IdP metadata under the leg-1 key; the + // end-session URL is built from that cache, never a fresh network fetch. + await new NodeOAuthStorage().saveServerMetadata(`ema-idp:${ISSUER}`, { + issuer: ISSUER, + authorization_endpoint: `${ISSUER}/authorize`, + token_endpoint: `${ISSUER}/token`, + response_types_supported: ["code"], + end_session_endpoint: `${ISSUER}/session/end`, + }); + const { emaLogout, getEmaStatus } = + await import("../src/connection/ema.js"); + const result = await emaLogout(); + expect(result.issuer).toBe(ISSUER); + const url = new URL(result.endSessionUrl!); + expect(url.origin + url.pathname).toBe(`${ISSUER}/session/end`); + expect(url.searchParams.get("id_token_hint")).toMatch(/\./); + expect((await getEmaStatus()).loginState).toBe("none"); + }); + + it("emaLogin short-circuits when already signed in", async () => { + writeClientConfig(emaClientConfig()); + await seedIdpSession(); + const { emaLogin } = await import("../src/connection/ema.js"); + const result = await emaLogin(); + expect(result).toEqual({ + issuer: ISSUER, + loginState: "logged_in", + alreadyLoggedIn: true, + }); + expect(runRunnerInteractiveOAuth).not.toHaveBeenCalled(); + }); + + it("emaLogin runs the IdP flow via the runner adapter and reports the new connection", async () => { + writeClientConfig(emaClientConfig()); + let stderr = ""; + const originalWrite = process.stderr.write; + process.stderr.write = ((chunk: unknown, ...rest: unknown[]) => { + stderr += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stderr.write; + + startIdpOidcAuthorization.mockResolvedValue({ + authorizationUrl: new URL("https://idp.example.com/authorize?x=1"), + }); + completeIdpOidcAuthorization.mockImplementation(async () => { + await seedIdpSession(); + return { idToken: fakeIdToken() }; + }); + runRunnerInteractiveOAuth.mockImplementation( + async (options: { + client: { + authenticate: () => Promise<URL | undefined>; + completeOAuthFlow: (code: string, iss?: string) => Promise<void>; + }; + redirectUrlProvider: { redirectUrl: string }; + }) => { + // Mirror the real runner: bind the loopback redirect before leg 1. + options.redirectUrlProvider.redirectUrl = + "http://127.0.0.1:45678/oauth/callback"; + const url = await options.client.authenticate(); + expect(url?.href).toContain("idp.example.com/authorize"); + await options.client.completeOAuthFlow("code-1", ISSUER); + return { kind: "success" }; + }, + ); + + try { + const { emaLogin } = await import("../src/connection/ema.js"); + const result = await emaLogin(); + expect(result).toEqual({ + issuer: ISSUER, + loginState: "logged_in", + alreadyLoggedIn: false, + }); + } finally { + process.stderr.write = originalWrite; + } + + expect(startIdpOidcAuthorization).toHaveBeenCalledWith( + expect.objectContaining({ + redirectUrl: "http://127.0.0.1:45678/oauth/callback", + }), + ); + expect(completeIdpOidcAuthorization).toHaveBeenCalledWith( + expect.objectContaining({ authorizationCode: "code-1", iss: ISSUER }), + ); + // Agent-attended wording: vitest's stderr is not a TTY, so the printed + // line must direct an agent to relay the IdP link to the human user. + expect(stderr).toContain( + "The user needs to sign in to the enterprise identity provider", + ); + }); + + it("emaLogin --relogin clears the existing connection and re-runs the flow", async () => { + writeClientConfig(emaClientConfig()); + await seedIdpSession(); + startIdpOidcAuthorization.mockResolvedValue({ + authorizationUrl: new URL("https://idp.example.com/authorize"), + }); + completeIdpOidcAuthorization.mockImplementation(async () => { + await seedIdpSession(); + return { idToken: fakeIdToken() }; + }); + runRunnerInteractiveOAuth.mockImplementation( + async (options: { + client: { + authenticate: () => Promise<URL | undefined>; + completeOAuthFlow: (code: string) => Promise<void>; + }; + redirectUrlProvider: { redirectUrl: string }; + }) => { + // The pre-existing connection must already be gone before leg 1 runs. + const storage = new NodeOAuthStorage(); + expect(await storage.getIdpSession(ISSUER)).toBeUndefined(); + await options.client.authenticate(); + await options.client.completeOAuthFlow("code-2"); + return { kind: "success" }; + }, + ); + + const { emaLogin } = await import("../src/connection/ema.js"); + const result = await emaLogin({ relogin: true }); + expect(result.alreadyLoggedIn).toBe(false); + expect(result.loginState).toBe("logged_in"); + expect(runRunnerInteractiveOAuth).toHaveBeenCalledOnce(); + }); + + it("mcpdoEmaGuidance names both configuration routes", async () => { + const { mcpdoEmaGuidance } = await import("../src/connection/ema.js"); + expect(mcpdoEmaGuidance("not_configured")).toMatch( + /Client Settings.*enterpriseManagedAuth/is, + ); + expect(mcpdoEmaGuidance("disabled")).toContain("enabled"); + }); +}); diff --git a/clients/mcpdo/__tests__/form-prompt.test.ts b/clients/mcpdo/__tests__/form-prompt.test.ts new file mode 100644 index 0000000000..db6197cab2 --- /dev/null +++ b/clients/mcpdo/__tests__/form-prompt.test.ts @@ -0,0 +1,521 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import { createStyle } from "@inspector/core/cli/style.js"; +import { promptForm } from "../src/connection/form-prompt.js"; +import type { FormField } from "../src/connection/form-schema.js"; + +/** + * Covers `promptForm`'s field-by-field prompting (one branch per + * `FormField.kind`, including validation retry loops and defaults) and the + * review step (submit / edit-by-name / cancel). + */ +describe("promptForm", () => { + let stderr: string; + let originalWrite: typeof process.stderr.write; + const style = createStyle(false); + + beforeEach(() => { + stderr = ""; + originalWrite = process.stderr.write; + process.stderr.write = ((chunk: unknown, ...rest: unknown[]) => { + stderr += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stderr.write; + }); + + afterEach(() => { + process.stderr.write = originalWrite; + }); + + function fakeRl(answers: string[]) { + let i = 0; + const closeHandlers: Array<() => void> = []; + return { + question: vi.fn(async () => { + const answer = answers[i]; + i += 1; + if (answer === undefined) { + throw new Error("no more scripted answers"); + } + return answer; + }), + once: vi.fn((event: string, cb: () => void) => { + if (event === "close") closeHandlers.push(cb); + }), + // Test-only hook: simulates the underlying stdin closing (e.g. a + // redirected/piped input hitting EOF) so we can exercise the + // watchForClose() race without a real stream. + __triggerClose: () => closeHandlers.forEach((cb) => cb()), + } as unknown as Parameters<typeof promptForm>[0] & { + __triggerClose: () => void; + }; + } + + const stringField: FormField = { + name: "name", + required: true, + title: "Name", + kind: "string", + }; + + it("collects a required string field and submits on blank review answer", async () => { + const rl = fakeRl(["octocat", ""]); + const outcome = await promptForm( + rl, + "Enter your name", + [stringField], + style, + ); + expect(outcome).toEqual({ action: "accept", content: { name: "octocat" } }); + expect(stderr).toContain("Enter your name"); + }); + + it("re-prompts a required string field left blank, then accepts a default", async () => { + const field: FormField = { + ...stringField, + required: false, + default: "anon", + }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { name: "anon" } }); + }); + + it("omits an optional string field left blank with no default", async () => { + const field: FormField = { ...stringField, required: false }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: {} }); + }); + + it("accepts a blank answer for a required string field as an empty string", async () => { + // JSON Schema `required` means present, not non-empty. + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [stringField], style); + expect(outcome).toEqual({ action: "accept", content: { name: "" } }); + expect(stderr).not.toContain("This field is required"); + }); + + it("preserves a schema property named __proto__ as an own property", async () => { + // On a plain object, `content["__proto__"] = v` hits the prototype + // setter instead of creating an own property, silently dropping the + // answer; the accepted payload is built with a null prototype. + const field: FormField = { + name: "__proto__", + required: true, + title: "Proto", + kind: "string", + }; + const rl = fakeRl(["value", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome.action).toBe("accept"); + const content = (outcome as { content: Record<string, unknown> }).content; + expect(Object.prototype.hasOwnProperty.call(content, "__proto__")).toBe( + true, + ); + expect(content["__proto__"]).toBe("value"); + }); + + it("lets minLength reject a blank required answer", async () => { + const field: FormField = { ...stringField, minLength: 3 }; + const rl = fakeRl(["", "abc", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { name: "abc" } }); + expect(stderr).toContain("at least 3"); + }); + + it("keeps an empty-string default instead of dropping the field", async () => { + const field: FormField = { ...stringField, required: false, default: "" }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { name: "" } }); + }); + + it("enforces minLength/maxLength on a string field", async () => { + const field: FormField = { ...stringField, minLength: 3, maxLength: 5 }; + const rl = fakeRl(["ab", "toolong", "oka", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { name: "oka" } }); + expect(stderr).toContain("at least 3"); + expect(stderr).toContain("at most 5"); + }); + + it("measures length bounds in code points, not UTF-16 units", async () => { + // "😀😀😀" is 3 code points (6 UTF-16 units): valid for min 3 / max 5. + const field: FormField = { ...stringField, minLength: 3, maxLength: 5 }; + const rl = fakeRl(["😀😀😀", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { name: "😀😀😀" } }); + }); + + it("collects a required number field with range validation", async () => { + const field: FormField = { + name: "age", + required: true, + title: "Age", + kind: "number", + integer: false, + minimum: 18, + maximum: 100, + }; + const rl = fakeRl(["notanumber", "5", "30", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { age: 30 } }); + expect(stderr).toContain("Enter a valid number"); + }); + + it("rejects a non-integer value for an integer field", async () => { + const field: FormField = { + name: "count", + required: true, + title: "Count", + kind: "number", + integer: true, + }; + const rl = fakeRl(["1.5", "3", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { count: 3 } }); + expect(stderr).toContain("Enter a valid integer"); + }); + + it("uses a number field's default on blank, or omits when optional with none", async () => { + const withDefault: FormField = { + name: "age", + required: false, + title: "Age", + kind: "number", + integer: false, + default: 21, + }; + const rl1 = fakeRl(["", ""]); + expect(await promptForm(rl1, "msg", [withDefault], style)).toEqual({ + action: "accept", + content: { age: 21 }, + }); + + const noDefault: FormField = { + name: "age", + required: false, + title: "Age", + kind: "number", + integer: false, + }; + const rl2 = fakeRl(["", ""]); + expect(await promptForm(rl2, "msg", [noDefault], style)).toEqual({ + action: "accept", + content: {}, + }); + }); + + it("re-prompts a required number field left blank", async () => { + const field: FormField = { + name: "age", + required: true, + title: "Age", + kind: "number", + integer: false, + }; + const rl = fakeRl(["", "42", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { age: 42 } }); + }); + + it("collects a boolean field via y/n, defaulting on blank", async () => { + const field: FormField = { + name: "confirm", + required: false, + title: "Confirm", + kind: "boolean", + default: true, + }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { confirm: true } }); + }); + + it("re-prompts on an invalid boolean answer and accepts yes/no variants", async () => { + const field: FormField = { + name: "confirm", + required: true, + title: "Confirm", + kind: "boolean", + }; + const rl = fakeRl(["maybe", "yes", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { confirm: true } }); + expect(stderr).toContain("Please answer y or n"); + + const rl2 = fakeRl(["no", ""]); + expect( + await promptForm(rl2, "msg", [{ ...field, required: false }], style), + ).toEqual({ action: "accept", content: { confirm: false } }); + }); + + it("omits an optional boolean field left blank with no default", async () => { + const field: FormField = { + name: "confirm", + required: false, + title: "Confirm", + kind: "boolean", + }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: {} }); + }); + + it("collects a single-select enum by number, and accepts a default on blank", async () => { + const field: FormField = { + name: "color", + required: true, + title: "Color", + kind: "enum", + choices: [ + { value: "red", label: "Red" }, + { value: "blue", label: "Blue" }, + ], + }; + const rl = fakeRl(["2", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { color: "blue" } }); + + const withDefault: FormField = { ...field, default: "red" }; + const rl2 = fakeRl(["", ""]); + expect(await promptForm(rl2, "msg", [withDefault], style)).toEqual({ + action: "accept", + content: { color: "red" }, + }); + }); + + it("re-prompts on an out-of-range enum choice and a required blank", async () => { + const field: FormField = { + name: "color", + required: true, + title: "Color", + kind: "enum", + choices: [{ value: "red", label: "Red" }], + }; + const rl = fakeRl(["", "9", "1", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { color: "red" } }); + expect(stderr).toContain("This field is required"); + expect(stderr).toContain("Enter a number between 1 and 1"); + }); + + it("rejects malformed and multi-token single-select answers", async () => { + const field: FormField = { + name: "color", + required: true, + title: "Color", + kind: "enum", + choices: [ + { value: "red", label: "Red" }, + { value: "blue", label: "Blue" }, + ], + }; + // "1abc" must not be silently accepted as choice 1 (parseInt prefix), + // and "1,2" on a single-select must not silently submit only "red". + const rl = fakeRl(["1abc", "1,2", "2", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { color: "blue" } }); + expect(stderr).toContain("Enter a number between 1 and 2"); + expect(stderr).toContain("Enter exactly one number"); + }); + + it("omits an optional enum field left blank with no default", async () => { + const field: FormField = { + name: "color", + required: false, + title: "Color", + kind: "enum", + choices: [{ value: "red", label: "Red" }], + }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: {} }); + }); + + it("collects a multi-select enum via comma-separated numbers, enforcing minItems/maxItems", async () => { + const field: FormField = { + name: "colors", + required: true, + title: "Colors", + kind: "multiselect", + choices: [ + { value: "red", label: "Red" }, + { value: "green", label: "Green" }, + { value: "blue", label: "Blue" }, + ], + minItems: 1, + maxItems: 2, + }; + const rl = fakeRl(["1,2,3", "1,2", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ + action: "accept", + content: { colors: ["red", "green"] }, + }); + expect(stderr).toContain("Select at most 2"); + }); + + it("enforces minItems on a multi-select enum", async () => { + const field: FormField = { + name: "colors", + required: true, + title: "Colors", + kind: "multiselect", + choices: [ + { value: "red", label: "Red" }, + { value: "green", label: "Green" }, + ], + minItems: 2, + }; + const rl = fakeRl(["1", "1,2", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ + action: "accept", + content: { colors: ["red", "green"] }, + }); + expect(stderr).toContain("Select at least 2"); + }); + + it("accepts blank on a required multi-select as an empty array when minItems permits", async () => { + // JSON Schema `required` means the key must be present; [] is a valid + // value unless minItems forbids it. Previously this looped forever. + const field: FormField = { + name: "colors", + required: true, + title: "Colors", + kind: "multiselect", + choices: [ + { value: "red", label: "Red" }, + { value: "green", label: "Green" }, + ], + }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { colors: [] } }); + expect(stderr).not.toContain("This field is required"); + }); + + it("re-prompts blank on a required multi-select when minItems demands entries", async () => { + const field: FormField = { + name: "colors", + required: true, + title: "Colors", + kind: "multiselect", + choices: [ + { value: "red", label: "Red" }, + { value: "green", label: "Green" }, + ], + minItems: 1, + }; + const rl = fakeRl(["", "1", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ + action: "accept", + content: { colors: ["red"] }, + }); + expect(stderr).toContain("Select at least 1"); + }); + + it("uses a multi-select default on blank, formatted in the field description", async () => { + const field: FormField = { + name: "colors", + required: false, + title: "Colors", + kind: "multiselect", + choices: [{ value: "red", label: "Red" }], + default: ["red"], + }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: { colors: ["red"] } }); + expect( + (rl.question as ReturnType<typeof vi.fn>).mock.calls[0][0], + ).toContain("[default: red]"); + }); + + it("omits an optional multi-select field left blank with no default", async () => { + const field: FormField = { + name: "colors", + required: false, + title: "Colors", + kind: "multiselect", + choices: [{ value: "red", label: "Red" }], + }; + const rl = fakeRl(["", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ action: "accept", content: {} }); + }); + + it("shows a description when the field has one", async () => { + const field: FormField = { + ...stringField, + description: "Your full display name", + }; + const rl = fakeRl(["octocat", ""]); + await promptForm(rl, "msg", [field], style); + expect( + (rl.question as ReturnType<typeof vi.fn>).mock.calls[0][0], + ).toContain("Your full display name"); + }); + + it("cancels from the review step", async () => { + const rl = fakeRl(["octocat", "c"]); + const outcome = await promptForm(rl, "msg", [stringField], style); + expect(outcome).toEqual({ action: "cancel" }); + }); + + it("re-prompts a review answer that doesn't name a known field", async () => { + const rl = fakeRl(["octocat", "bogus", "c"]); + const outcome = await promptForm(rl, "msg", [stringField], style); + expect(outcome).toEqual({ action: "cancel" }); + expect(stderr).toContain('Unknown field "bogus"'); + }); + + it("lets the review step re-edit a named field before submitting", async () => { + const rl = fakeRl(["octocat", "name", "edited", ""]); + const outcome = await promptForm(rl, "msg", [stringField], style); + expect(outcome).toEqual({ action: "accept", content: { name: "edited" } }); + }); + + it("shows the property name in the review when it differs from the title, and edits by title", async () => { + const field: FormField = { + name: "emailAddress", + required: true, + title: "Email address", + kind: "string", + }; + // Initial value, edit via the display title, new value, submit. + const rl = fakeRl(["a@example.com", "Email address", "b@example.com", ""]); + const outcome = await promptForm(rl, "msg", [field], style); + expect(outcome).toEqual({ + action: "accept", + content: { emailAddress: "b@example.com" }, + }); + // The review label must reveal the editable property name. + expect(stderr).toContain("Email address (emailAddress):"); + }); + + it("shows '(none)' in the review for a field with no value", async () => { + const field: FormField = { ...stringField, required: false }; + const rl = fakeRl(["", ""]); + await promptForm(rl, "msg", [field], style); + expect(stderr).toContain("(none)"); + }); + + it("rejects instead of hanging when stdin closes before an answer arrives", async () => { + const rl = fakeRl([]); + (rl.question as ReturnType<typeof vi.fn>).mockImplementation( + () => new Promise(() => {}), // never resolves on its own + ); + const outcome = promptForm(rl, "msg", [stringField], style); + (rl as unknown as { __triggerClose: () => void }).__triggerClose(); + await expect(outcome).rejects.toThrow( + "stdin closed before an answer was given", + ); + }); +}); diff --git a/clients/mcpdo/__tests__/form-schema.test.ts b/clients/mcpdo/__tests__/form-schema.test.ts new file mode 100644 index 0000000000..4a554b7af6 --- /dev/null +++ b/clients/mcpdo/__tests__/form-schema.test.ts @@ -0,0 +1,443 @@ +import { describe, it, expect } from "vitest"; +import { parseFormSchema } from "../src/connection/form-schema.js"; + +describe("parseFormSchema", () => { + it("returns null for a non-object schema", () => { + expect(parseFormSchema(undefined)).toBeNull(); + expect(parseFormSchema({ type: "string" })).toBeNull(); + }); + + it("returns null when properties is missing or not an object", () => { + expect(parseFormSchema({ type: "object" })).toBeNull(); + expect(parseFormSchema({ type: "object", properties: "nope" })).toBeNull(); + }); + + it("parses a string field with title/description/length/format/default", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + name: { + type: "string", + title: "Display Name", + description: "Your name", + minLength: 2, + maxLength: 20, + format: "email", + default: "octocat", + }, + }, + required: ["name"], + }); + expect(fields).toEqual([ + { + name: "name", + required: true, + title: "Display Name", + description: "Your name", + kind: "string", + minLength: 2, + maxLength: 20, + format: "email", + default: "octocat", + }, + ]); + }); + + it("parses a number field, distinguishing integer from number", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + age: { type: "number", minimum: 18, maximum: 100, default: 30 }, + count: { type: "integer" }, + }, + properties2: undefined, + } as Record<string, unknown>); + expect(fields).toEqual([ + { + name: "age", + required: false, + title: "age", + description: undefined, + kind: "number", + integer: false, + minimum: 18, + maximum: 100, + default: 30, + }, + { + name: "count", + required: false, + title: "count", + description: undefined, + kind: "number", + integer: true, + minimum: undefined, + maximum: undefined, + default: undefined, + }, + ]); + }); + + it("parses a boolean field with a default", () => { + const fields = parseFormSchema({ + type: "object", + properties: { confirm: { type: "boolean", default: false } }, + }); + expect(fields).toEqual([ + { + name: "confirm", + required: false, + title: "confirm", + description: undefined, + kind: "boolean", + default: false, + }, + ]); + }); + + it("parses a single-select enum without titles", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + color: { + type: "string", + title: "Color", + enum: ["Red", "Green", "Blue"], + default: "Red", + }, + }, + }); + expect(fields).toEqual([ + { + name: "color", + required: false, + title: "Color", + description: undefined, + kind: "enum", + choices: [ + { value: "Red", label: "Red" }, + { value: "Green", label: "Green" }, + { value: "Blue", label: "Blue" }, + ], + default: "Red", + }, + ]); + }); + + it("parses a single-select enum with titled oneOf", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + color: { + type: "string", + oneOf: [{ const: "#FF0000", title: "Red" }, { const: "#00FF00" }], + }, + }, + }); + expect(fields).toEqual([ + { + name: "color", + required: false, + title: "color", + description: undefined, + kind: "enum", + choices: [ + { value: "#FF0000", label: "Red" }, + { value: "#00FF00", label: "#00FF00" }, + ], + default: undefined, + }, + ]); + }); + + it("returns null when oneOf entries are malformed", () => { + expect( + parseFormSchema({ + type: "object", + properties: { + color: { type: "string", oneOf: [{ notConst: true }] }, + }, + }), + ).toBeNull(); + expect( + parseFormSchema({ + type: "object", + properties: { color: { type: "string", oneOf: "nope" } }, + }), + ).toBeNull(); + }); + + it("returns null for empty choice arrays (unwinnable required prompt otherwise)", () => { + // A required field with zero options renders no choices and rejects + // every answer (1..0 range) — treat the schema as malformed instead. + expect( + parseFormSchema({ + type: "object", + properties: { color: { type: "string", enum: [] } }, + required: ["color"], + }), + ).toBeNull(); + expect( + parseFormSchema({ + type: "object", + properties: { color: { type: "string", oneOf: [] } }, + }), + ).toBeNull(); + expect( + parseFormSchema({ + type: "object", + properties: { + colors: { type: "array", items: { type: "string", enum: [] } }, + }, + }), + ).toBeNull(); + expect( + parseFormSchema({ + type: "object", + properties: { + colors: { type: "array", items: { type: "string", anyOf: [] } }, + }, + }), + ).toBeNull(); + // Non-string enum entries stay malformed too (not a freeform string). + expect( + parseFormSchema({ + type: "object", + properties: { color: { type: "string", enum: [1, 2] } }, + }), + ).toBeNull(); + }); + + it("parses a multi-select enum without titles, with min/maxItems and default", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + colors: { + type: "array", + title: "Colors", + minItems: 1, + maxItems: 2, + items: { type: "string", enum: ["Red", "Green", "Blue"] }, + default: ["Red", "Green"], + }, + }, + }); + expect(fields).toEqual([ + { + name: "colors", + required: false, + title: "Colors", + description: undefined, + kind: "multiselect", + choices: [ + { value: "Red", label: "Red" }, + { value: "Green", label: "Green" }, + { value: "Blue", label: "Blue" }, + ], + minItems: 1, + maxItems: 2, + default: ["Red", "Green"], + }, + ]); + }); + + it("parses a multi-select enum with titled anyOf", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + colors: { + type: "array", + items: { + anyOf: [ + { const: "#FF0000", title: "Red" }, + { const: "#00FF00", title: "Green" }, + ], + }, + }, + }, + }); + expect(fields?.[0]).toMatchObject({ + kind: "multiselect", + choices: [ + { value: "#FF0000", label: "Red" }, + { value: "#00FF00", label: "Green" }, + ], + }); + }); + + it("returns null for an array field without items or without enum/anyOf", () => { + expect( + parseFormSchema({ + type: "object", + properties: { colors: { type: "array" } }, + }), + ).toBeNull(); + expect( + parseFormSchema({ + type: "object", + properties: { + colors: { type: "array", items: { type: "string" } }, + }, + }), + ).toBeNull(); + }); + + it("ignores a non-string-array default on a multiselect field", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + colors: { + type: "array", + items: { type: "string", enum: ["Red"] }, + default: [1, 2], + }, + }, + }); + expect(fields?.[0]).toMatchObject({ default: undefined }); + }); + + it("returns null for an unsupported/unknown property type", () => { + expect( + parseFormSchema({ + type: "object", + properties: { nested: { type: "object", properties: {} } }, + }), + ).toBeNull(); + }); + + it("returns null when a property isn't an object", () => { + expect( + parseFormSchema({ + type: "object", + properties: { name: "not-a-schema" }, + }), + ).toBeNull(); + }); + + it("treats non-array/malformed required as no required fields", () => { + const fields = parseFormSchema({ + type: "object", + properties: { name: { type: "string" } }, + required: "name", + }); + expect(fields?.[0].required).toBe(false); + }); + + // Internally inconsistent fields are rejected like any other malformed + // schema: unsatisfiable constraints or a default violating its own + // constraints would render unwinnable / instantly-invalid prompts. + it("returns null for unsatisfiable constraints", () => { + const cases: Record<string, unknown>[] = [ + { n: { type: "number", minimum: 10, maximum: 5 } }, + { s: { type: "string", minLength: 5, maxLength: 2 } }, + { m: { type: "array", items: { enum: ["a"] }, minItems: 2 } }, + { + m: { + type: "array", + items: { enum: ["a", "b"] }, + minItems: 2, + maxItems: 1, + }, + }, + ]; + for (const properties of cases) { + expect(parseFormSchema({ type: "object", properties })).toBeNull(); + } + }); + + it("returns null for negative or non-integer length/count keywords", () => { + const cases: Record<string, unknown>[] = [ + { s: { type: "string", maxLength: -1 } }, + { s: { type: "string", minLength: -3 } }, + { s: { type: "string", minLength: 1.5 } }, + { m: { type: "array", items: { enum: ["a", "b"] }, maxItems: -1 } }, + { m: { type: "array", items: { enum: ["a", "b"] }, minItems: -2 } }, + { m: { type: "array", items: { enum: ["a", "b"] }, minItems: 0.5 } }, + ]; + for (const properties of cases) { + expect(parseFormSchema({ type: "object", properties })).toBeNull(); + } + // Zero is a valid bound. + expect( + parseFormSchema({ + type: "object", + properties: { s: { type: "string", minLength: 0 } }, + }), + ).toHaveLength(1); + }); + + it("measures default length bounds in code points, not UTF-16 units", () => { + // "😀" is 1 code point (2 UTF-16 units): a valid default for maxLength 1. + expect( + parseFormSchema({ + type: "object", + properties: { s: { type: "string", maxLength: 1, default: "😀" } }, + }), + ).toHaveLength(1); + // ...and 1 code point still violates minLength 2. + expect( + parseFormSchema({ + type: "object", + properties: { s: { type: "string", minLength: 2, default: "😀" } }, + }), + ).toBeNull(); + }); + + it("returns null for defaults that violate the field's own constraints", () => { + const cases: Record<string, unknown>[] = [ + { n: { type: "number", minimum: 1, maximum: 10, default: 11 } }, + { n: { type: "number", minimum: 1, default: 0 } }, + { i: { type: "integer", default: 1.5 } }, + { s: { type: "string", minLength: 3, default: "ab" } }, + { s: { type: "string", maxLength: 2, default: "abc" } }, + { e: { type: "string", enum: ["a", "b"], default: "c" } }, + { + e: { + type: "string", + oneOf: [{ const: "a", title: "A" }], + default: "b", + }, + }, + { m: { type: "array", items: { enum: ["a", "b"] }, default: ["c"] } }, + { + m: { + type: "array", + items: { enum: ["a", "b"] }, + minItems: 2, + default: ["a"], + }, + }, + { + m: { + type: "array", + items: { enum: ["a", "b"] }, + maxItems: 1, + default: ["a", "b"], + }, + }, + ]; + for (const properties of cases) { + expect(parseFormSchema({ type: "object", properties })).toBeNull(); + } + }); + + it("accepts consistent constraints with in-range defaults", () => { + const fields = parseFormSchema({ + type: "object", + properties: { + n: { type: "number", minimum: 1, maximum: 10, default: 5 }, + i: { type: "integer", minimum: 0, default: 0 }, + s: { type: "string", minLength: 1, maxLength: 3, default: "ab" }, + e: { type: "string", enum: ["a", "b"], default: "b" }, + m: { + type: "array", + items: { enum: ["a", "b"] }, + minItems: 1, + maxItems: 2, + default: ["a", "b"], + }, + }, + }); + expect(fields).toHaveLength(5); + }); +}); diff --git a/clients/mcpdo/__tests__/format-connection.test.ts b/clients/mcpdo/__tests__/format-connection.test.ts new file mode 100644 index 0000000000..2ab5c713e9 --- /dev/null +++ b/clients/mcpdo/__tests__/format-connection.test.ts @@ -0,0 +1,1489 @@ +import { describe, it, expect, beforeEach, afterEach } from "vitest"; +import { + formatCallToolResultHuman, + formatToolsHuman, + formatResourcesHuman, + formatResourceTemplatesHuman, + formatPromptsHuman, + formatResourceReadHuman, + formatPromptResultHuman, + formatCompletionsHuman, + formatTasksHuman, + formatTaskHuman, + formatInitializeHuman, + formatRootsHuman, + formatAuthListHuman, + formatServersListHuman, + formatServerShowHuman, + formatConnectionsListHuman, + formatConnectionInfoHuman, + formatAppInfoListHuman, + formatAppInfoHuman, + formatSkillVerifyListHuman, + formatSkillsHuman, + formatSkillGetHuman, + formatStreamEventHuman, + formatRpcResultHuman, + formatElicitationPendingHuman, +} from "../src/connection/format-human.js"; +import { writeConnectionOutput } from "../src/connection/format-connection.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { createStyle, PLAIN } from "@inspector/core/cli/style.js"; + +describe("format-human", () => { + it("formats tools with schema variants and empty list", () => { + expect(formatToolsHuman([])).toContain("(none)"); + const text = formatToolsHuman([ + { + name: "echo", + description: "Echo back\nmore", + inputSchema: { + type: "object", + properties: { + message: { type: "string" }, + n: { type: "number" }, + tags: { type: "array", items: { type: "string" } }, + extra: { type: "boolean" }, + }, + required: ["message"], + }, + annotations: { + readOnlyHint: true, + destructiveHint: true, + idempotentHint: true, + openWorldHint: true, + }, + }, + { + name: "types", + inputSchema: { + type: "object", + properties: { + emptyArr: { type: "array" }, + multi: { type: ["string", "number"] }, + bare: {}, + }, + }, + }, + { + name: "more", + inputSchema: { + type: "object", + properties: { + flag: { type: ["boolean", "null"] }, + choice: { enum: ["a", "b"] }, + obj: { type: "object" }, + }, + }, + }, + { + name: "ints", + inputSchema: { + type: "object", + properties: { + i: { type: "integer" }, + unknownType: { type: "custom" }, + nonObjProp: "x", + }, + }, + annotations: {}, + }, + { name: "plain", inputSchema: null }, + { name: "emptyProps", inputSchema: { type: "object", properties: {} } }, + { name: "noProps", inputSchema: { type: "object" } }, + {}, + ]); + expect(text).toContain("Tools (8):"); + expect(text).toContain("`echo(message:str, n?:num, tags?:[str], …)`"); + expect(text).toContain("[read-only, destructive, idempotent, open-world]"); + expect(text).toContain("emptyArr?:[any]"); + expect(text).toContain("multi?:str | num"); + expect(text).toContain("bare?:any"); + expect(text).toContain("flag?:bool"); + expect(text).toContain("choice?:enum"); + expect(text).toContain("`plain()`"); + expect(text).toContain("`?()`"); + }); + + it("formats list helpers for resources, templates, prompts, roots, tasks", () => { + expect( + formatResourcesHuman([ + { name: "r", uri: "u://x", description: "d\n2" }, + { uri: "u://only" }, + { name: "n", uri: 1, description: " " }, + { name: "no-uri" }, + ]), + ).toContain("`r` (u://x)"); + expect(formatResourcesHuman([])).toContain("(none)"); + + expect( + formatResourceTemplatesHuman([ + { name: "t", uriTemplate: "u://{id}", description: "tpl" }, + { description: " " }, + { name: "x", uriTemplate: 1 }, + ]), + ).toContain("u://{id}"); + expect(formatResourceTemplatesHuman([])).toContain("(none)"); + + expect( + formatPromptsHuman([ + { name: "p", description: "hi\nmore" }, + { description: " " }, + {}, + ]), + ).toContain("`p`"); + expect(formatPromptsHuman([])).toContain("(none)"); + + expect( + formatRootsHuman([{ uri: "file:///a", name: "a" }, { uri: "file:///b" }]), + ).toContain("file:///a (a)"); + expect(formatRootsHuman([])).toContain("(none)"); + + expect( + formatTasksHuman([ + { taskId: "1", status: "running", statusMessage: "go" }, + { id: "2", status: "done" }, + {}, + ]), + ).toContain("`1` running"); + expect(formatTasksHuman([])).toContain("(none)"); + + expect( + formatTaskHuman({ + taskId: "t1", + status: "ok", + statusMessage: "fine", + createdAt: "c", + lastUpdatedAt: "u", + }), + ).toContain("Created: c"); + expect(formatTaskHuman(null)).toContain("Task: `?`"); + expect(formatTaskHuman({})).toContain("Status: ?"); + }); + + it("formats call tool results across content block types", () => { + const structured = { ok: true }; + const withDupe = formatCallToolResultHuman({ + isError: true, + content: [ + { type: "text", text: JSON.stringify(structured) }, + { type: "text", text: "hello" }, + { type: "text", text: "{not-json" }, + { + type: "resource_link", + uri: "u://r", + name: "n", + description: "d", + mimeType: "text/plain", + }, + { type: "image", mimeType: "image/png", data: "abc" }, + { type: "audio", mimeType: "audio/wav" }, + { + type: "resource", + resource: { uri: "u://e", mimeType: "text/plain", text: "body" }, + }, + { type: "custom", x: 1 }, + ], + structuredContent: structured, + _meta: { a: 1 }, + }); + expect(withDupe).toContain("Tool error:"); + expect(withDupe).toContain("hello"); + expect(withDupe).toContain("Resource link"); + expect(withDupe).toContain("[Image:"); + expect(withDupe).toContain("[Audio:"); + expect(withDupe).toContain("Embedded resource"); + expect(withDupe).toContain('"x": 1'); + // The JSON duplicate of structuredContent is filtered from the content + // blocks, but the structured payload itself must still be rendered once. + expect(withDupe).toContain("Structured content:"); + expect(withDupe).toContain('"ok": true'); + + expect( + formatCallToolResultHuman({ + isError: true, + structuredContent: { only: true }, + content: [], + }), + ).toContain("Structured content:"); + + expect( + formatCallToolResultHuman({ + content: [{ type: "image" }, { type: "audio", data: "x" }], + }), + ).toContain("[Image: unknown"); + + expect( + formatCallToolResultHuman({ + content: [{ type: "resource" }], + }), + ).toContain("Embedded resource"); + + expect( + formatCallToolResultHuman({ + content: [ + { + type: "resource", + resource: { uri: "u://e" }, + }, + ], + }), + ).toContain("URI: u://e"); + + expect( + formatCallToolResultHuman({ + content: [{ type: "resource_link", uri: "u" }], + }), + ).toContain("Resource link"); + + expect( + formatCallToolResultHuman({ + content: [{ type: "text" }], + structuredContent: {}, + _meta: {}, + }), + ).toContain("Content:"); + + expect(formatCallToolResultHuman({})).toBe("(no content)"); + }); + + it("formats resource read, prompt get, completions, initialize", () => { + expect(formatResourceReadHuman({ contents: [] })).toBe("(empty resource)"); + expect( + formatResourceReadHuman({ + contents: [ + { uri: "u://a", mimeType: "text/plain", text: "hi" }, + { uri: "u://b", blob: "zzzz" }, + ], + }), + ).toContain("[Blob:"); + + expect(formatPromptResultHuman({})).toBe("(empty prompt)"); + expect( + formatPromptResultHuman({ + description: "desc", + messages: [ + { role: "user", content: "plain" }, + { + role: "assistant", + content: [{ type: "text", text: "block" }], + }, + { role: "user", content: { type: "text", text: "obj" } }, + ], + }), + ).toContain("[assistant]"); + + expect(formatCompletionsHuman({ values: ["a"], hasMore: true })).toContain( + "(more available)", + ); + expect(formatCompletionsHuman({ values: [] })).toContain("(none)"); + + expect( + formatInitializeHuman({ + serverInfo: { name: "s", version: "1" }, + protocolVersion: "2025-01-01", + instructions: " use me ", + capabilities: { tools: {} }, + }), + ).toContain("Capabilities: tools"); + expect(formatInitializeHuman({})).toContain("(unknown)"); + expect( + formatInitializeHuman({ + serverInfo: { name: "s" }, + instructions: " ", + capabilities: {}, + }), + ).toContain("Server: s"); + }); + + it("annotates auth/list entries with knownAs, live, and IdP labelling", () => { + const out = formatAuthListHuman({ + oauthStatePath: "/tmp/oauth.json", + servers: [ + { + url: "https://mcp.example.com/mcp", + hasTokens: true, + hasRefreshToken: false, + knownAs: ["hosted"], + live: true, + }, + { + url: "ema-idp:https://idp.example.com", + hasTokens: false, + hasRefreshToken: false, + idp: true, + issuer: "https://idp.example.com", + }, + ], + }); + expect(out).toContain("known as: hosted"); + expect(out).toContain("● live"); + // The IdP row shows its issuer and label, not the raw key or "(no local name)". + expect(out).toContain("enterprise IdP login"); + expect(out).toContain("https://idp.example.com"); + expect(out).not.toContain("ema-idp:"); + expect(out).not.toContain("(no local name)"); + }); + + it("formats admin and app-info helpers", () => { + expect( + formatAuthListHuman({ + oauthStatePath: "/tmp/oauth.json", + servers: [ + { + url: "https://example.com/mcp", + hasTokens: true, + hasRefreshToken: true, + }, + { url: "https://empty.example/mcp" }, + ], + }), + ).toMatch(/Stored auth[\s\S]*example\.com[\s\S]*tokens[\s\S]*no tokens/); + expect( + formatAuthListHuman({ oauthStatePath: "/tmp/x", servers: [] }), + ).toContain("(none)"); + expect(formatServersListHuman([])).toContain("(none)"); + expect( + formatServersListHuman([], PLAIN, { + kind: "catalog", + path: "/home/u/.mcp-inspector/mcp.json", + }), + ).toContain("Source: catalog /home/u/.mcp-inspector/mcp.json"); + expect( + formatServersListHuman([], PLAIN, { + kind: "config", + path: "./mcp.json", + }), + ).toContain("Source: config ./mcp.json"); + expect( + formatServersListHuman([{ name: "s", type: "stdio", detail: "x" }]), + ).toContain("`s`"); + expect( + formatServersListHuman([ + { + name: "s", + type: "stdio", + detail: "x", + connection: "s", + isMru: true, + }, + ]), + ).toMatch(/@s \(MRU\)/); + expect( + formatServerShowHuman( + { name: "s", type: "stdio", detail: "x", config: {} }, + PLAIN, + { kind: "catalog", path: "/tmp/cat.json" }, + ), + ).toContain("Source: catalog /tmp/cat.json"); + expect( + formatServerShowHuman({ + name: "s", + type: "stdio", + detail: "node x", + config: { type: "stdio", command: "node" }, + }), + ).toMatch(/Server[\s\S]*`s`[\s\S]*node x/); + + expect(formatConnectionsListHuman([])).toContain("connect first"); + expect( + formatConnectionsListHuman([ + { name: "a", isMru: true, serverIdentity: "id" }, + ]), + ).toContain("(MRU)"); + // protocolEra is on every ConnectionInfo now (#2298 follow-up), not just + // connections/show — connections/list renders it inline; its absence (an older + // daemon reply, hypothetically) must not print a bare "[undefined]". + expect( + formatConnectionsListHuman([ + { name: "a", isMru: true, serverIdentity: "id", protocolEra: "modern" }, + ]), + ).toContain("— id [modern]"); + expect( + formatConnectionsListHuman([ + { name: "a", isMru: false, serverIdentity: "id" }, + ]), + ).not.toContain("["); + expect( + formatConnectionInfoHuman({ + name: "a", + isMru: true, + serverIdentity: "id", + }), + ).toContain("Connection `@a`"); + // connections/show enrichment: era without a protocolVersion, serverInfo + // without a version, empty capabilities, an empty/non-array + // supportedVersions, and blank instructions each take the "nothing to + // append" branch rather than the populated one exercised elsewhere. + expect( + formatConnectionInfoHuman({ + name: "a", + protocolEra: "legacy", + serverInfo: { name: "demo" }, + capabilities: {}, + supportedVersions: [], + instructions: "", + }), + ).toMatch(/Era: legacy\nServer info: demo\nCapabilities: \(none\)/); + expect( + formatConnectionInfoHuman({ + name: "a", + protocolEra: undefined, + protocolVersion: "2025-11-25", + serverInfo: { name: "demo", version: "1.2.3" }, + supportedVersions: ["2025-11-25", "2025-06-18"], + instructions: "Say hi.", + }), + ).toMatch( + /Era: unknown \(2025-11-25\)[\s\S]*demo v1\.2\.3[\s\S]*Supported versions: 2025-11-25, 2025-06-18[\s\S]*Instructions: Say hi\./, + ); + + // Auth snapshot line: OAuth with full detail, EMA with IdP session state, + // and a bare not-authorized snapshot (no scope/clientId branches). + expect( + formatConnectionInfoHuman({ + name: "a", + auth: { + method: "oauth", + authorized: true, + scope: "mcp:tools", + clientId: "client-123", + }, + }), + ).toContain( + "Auth: OAuth (authorized; scope: mcp:tools; client: client-123)", + ); + expect( + formatConnectionInfoHuman({ + name: "a", + auth: { method: "ema", authorized: true, idpSession: "logged_in" }, + }), + ).toContain("Auth: EMA (authorized; IdP session: logged_in)"); + expect( + formatConnectionInfoHuman({ + name: "a", + auth: { method: "oauth", authorized: false }, + }), + ).toContain("Auth: OAuth (not authorized)"); + + expect( + formatAppInfoListHuman([ + { toolName: "with", hasApp: true, resourceUri: "ui://x" }, + { toolName: "err", hasApp: false, resourceError: "boom" }, + { toolName: "no", hasApp: false }, + ]), + ).toContain("no app"); + + const verifyText = formatSkillVerifyListHuman([ + { name: "ok-skill", uri: "skill://ok/SKILL.md", outcome: "verified" }, + { + name: "bad-skill", + uri: "skill://bad/SKILL.md", + outcome: "failed", + conformance: [{ severity: "error" }], + files: [{ status: "mismatch" }], + }, + { + name: "cut-short", + uri: "skill://cut/SKILL.md", + outcome: "incomplete", + incomplete: "read bounds hit", + }, + ]); + expect(verifyText).toContain("Skill verification (3):"); + expect(verifyText).toContain("`ok-skill`"); + expect(verifyText).toContain("verified"); + expect(verifyText).toContain( + "`bad-skill` (skill://bad/SKILL.md) — failed — 1 issue(s), 1 file mismatch(es)", + ); + expect(verifyText).toContain("`cut-short`"); + expect(verifyText).toContain("read bounds hit"); + + // Unverifiable is its own verdict — "failed — 0 issue(s)" misreads as a + // pass narrowly missed when nothing was checked at all. + const unverifiableText = formatSkillVerifyListHuman([ + { + name: "dyn-skill", + uri: "skill://dyn/SKILL.md", + outcome: "unverifiable", + }, + ]); + expect(unverifiableText).toContain("`dyn-skill`"); + expect(unverifiableText).toContain("unverifiable"); + expect(unverifiableText).toContain("advertised no digests"); + expect(unverifiableText).not.toContain("failed"); + expect(unverifiableText).not.toContain("0 issue(s)"); + + const skillsText = formatSkillsHuman([ + { + uri: "skill://ok/SKILL.md", + frontmatter: { name: "ok-skill", description: "Does things" }, + resources: [{ uri: "skill://ok/a.md" }, { uri: "skill://ok/b.md" }], + }, + { + uri: "skill://dyn/SKILL.md", + frontmatter: { name: "dyn-skill" }, + resources: "dynamic", + }, + ]); + expect(skillsText).toContain("Skills (2):"); + expect(skillsText).toContain("`ok-skill` (skill://ok/SKILL.md)"); + expect(skillsText).toContain("Does things"); + expect(skillsText).toContain("2 file(s)"); + expect(skillsText).toContain("dynamic resources"); + expect(formatSkillsHuman([])).toContain("(none)"); + + const skillGetText = formatSkillGetHuman({ + skill: { + uri: "skill://one/SKILL.md", + frontmatter: { name: "one" }, + resources: [], + }, + ttlMs: 60000, + cacheScope: "session", + }); + expect(skillGetText).toContain("Skill:"); + expect(skillGetText).toContain("`one`"); + expect(skillGetText).toContain("ttlMs: 60000"); + expect(skillGetText).toContain("cacheScope: session"); + + // Degenerate shapes: a bare entry (no frontmatter/uri/resources) and a + // bare envelope (no skill/ttlMs/cacheScope) must still render. + const bareEntry = formatSkillsHuman([{}]); + expect(bareEntry).toContain("`?`"); + expect(bareEntry).not.toContain("file(s)"); + const bareGet = formatSkillGetHuman({}); + expect(bareGet).toContain("Skill:"); + expect(bareGet).not.toContain("ttlMs"); + expect(bareGet).not.toContain("cacheScope"); + + expect( + formatAppInfoHuman({ + toolName: "t", + hasApp: true, + resourceUri: "ui://x", + csp: { a: 1 }, + }), + ).toContain("CSP:"); + expect( + formatAppInfoHuman({ toolName: "t", hasApp: false, resourceError: "e" }), + ).toContain("e"); + expect(formatAppInfoHuman({ toolName: "t", hasApp: false })).toContain( + "No MCP App", + ); + }); + + it("formats stream events and rpc dispatch", () => { + expect(formatStreamEventHuman(null)).toBe("null"); + // Empty URI: formatUri must pass it through without linkifying. + expect(formatStreamEventHuman({ type: "subscribed" })).toBe("Subscribed: "); + // colorLevel groups: red, yellow, dim, and the default (cyan) bucket. + const s = createStyle(true); + for (const [level, colored] of [ + ["error", s.red("error")], + ["warning", s.yellow("warning")], + ["debug", s.dim("debug")], + ["notice", s.dim("notice")], + ["info", s.cyan("info")], + ] as const) { + expect( + formatStreamEventHuman( + { + direction: "notification", + message: { params: { level, data: "x" } }, + }, + s, + ), + ).toContain(`[${colored}]`); + } + expect(formatStreamEventHuman({ type: "subscribed", uri: "u" })).toBe( + "Subscribed: u", + ); + expect( + formatStreamEventHuman({ type: "resources/updated", uri: "u" }), + ).toBe("Resource updated: u"); + expect( + formatStreamEventHuman({ + direction: "notification", + message: { + method: "notifications/message", + params: { level: "warn", logger: "L", data: "hi" }, + }, + }), + ).toBe("[warn] L: hi"); + // Object `data` renders as JSON, not "[object Object]". + expect( + formatStreamEventHuman({ + direction: "notification", + message: { + method: "notifications/message", + params: { level: "info", data: { job: "sync", ok: true } }, + }, + }), + ).toBe('[info] {"job":"sync","ok":true}'); + expect( + formatStreamEventHuman({ + direction: "notification", + message: { params: { message: "m" } }, + }), + ).toBe("[info] m"); + expect( + formatStreamEventHuman({ + direction: "notification", + message: { params: { nested: true } }, + }), + ).toContain("nested"); + expect( + formatStreamEventHuman({ + direction: "notification", + message: {}, + }), + ).toContain("[info]"); + expect(formatStreamEventHuman({ other: 1 })).toContain('"other": 1'); + expect(formatStreamEventHuman("raw")).toBe("raw"); + + expect(formatRpcResultHuman("tools/list", { tools: [] })).toContain( + "Tools", + ); + expect(formatRpcResultHuman("tools/call", { content: [] })).toBe( + "(no content)", + ); + expect(formatRpcResultHuman("resources/list", { resources: [] })).toContain( + "Resources", + ); + expect(formatRpcResultHuman("resources/read", { contents: [] })).toBe( + "(empty resource)", + ); + expect( + formatRpcResultHuman("resources/templates/list", { + resourceTemplates: [], + }), + ).toContain("templates"); + expect(formatRpcResultHuman("resources/unsubscribe", { uri: "u" })).toBe( + "Unsubscribed: u", + ); + expect(formatRpcResultHuman("prompts/list", { prompts: [] })).toContain( + "Prompts", + ); + expect(formatRpcResultHuman("prompts/get", {})).toBe("(empty prompt)"); + expect(formatRpcResultHuman("skills/list", { skills: [] })).toContain( + "Skills (0):", + ); + expect( + formatRpcResultHuman("skills/get", { skill: { uri: "skill://x" } }), + ).toContain("Skill:"); + expect(formatRpcResultHuman("prompts/complete", { values: [] })).toContain( + "Completions", + ); + expect( + formatRpcResultHuman("initialize", { serverInfo: { name: "s" } }), + ).toContain("Server: s"); + expect(formatRpcResultHuman("logging/setLevel", {})).toBe( + "Logging level updated.", + ); + expect(formatRpcResultHuman("tasks/list", { tasks: [] })).toContain( + "Tasks", + ); + expect( + formatRpcResultHuman("tasks/get", { task: { taskId: "1", status: "x" } }), + ).toContain("Task: `1`"); + expect(formatRpcResultHuman("tasks/cancel", { taskId: "1" })).toBe( + "Cancelled task: 1", + ); + expect(formatRpcResultHuman("tasks/result", { content: [] })).toBe( + "(no content)", + ); + expect(formatRpcResultHuman("roots/list", { roots: [] })).toContain( + "Roots", + ); + expect(formatRpcResultHuman("roots/set", { roots: [] })).toContain("Roots"); + expect(formatRpcResultHuman("unknown/op", { x: 1 })).toBeNull(); + }); +}); + +describe("formatElicitationPendingHuman", () => { + it("formatElicitationPendingHuman renders form fields and respond guidance", () => { + const text = formatElicitationPendingHuman({ + elicitationId: "e-9", + connection: "srv", + method: "tools/call", + toolName: "collect", + mode: "form", + message: "Pick a color", + requestedSchema: { + type: "object", + properties: { + color: { type: "string", description: "Favourite color" }, + size: { type: "string", enum: ["s", "m", "l"] }, + count: { type: "integer" }, + }, + required: ["color"], + }, + origin: "server-request", + expiresAt: Date.now() + 600_000, + }); + expect(text).toContain("Input required"); + expect(text).toContain("@srv"); + expect(text).toContain("Pick a color"); + expect(text).toContain("color (string, required)"); + expect(text).toContain("Favourite color"); + expect(text).toContain("size (enum) [s, m, l]"); + expect(text).toContain("count (integer)"); + expect(text).toContain("elicitation/respond e-9 field:=value"); + expect(text).toContain("--decline | --cancel"); + expect(text).toContain("expires"); + }); + + it("formatElicitationPendingHuman renders url mode with --done guidance; unsafe schemes stay plain", () => { + const info = { + elicitationId: "e-u", + connection: "srv", + method: "tools/call", + mode: "url", + message: "Finish signup", + url: "https://example.com/signup?flow=abc", + origin: "server-request", + expiresAt: 0, + } satisfies Parameters<typeof formatElicitationPendingHuman>[0]; + const styled = formatElicitationPendingHuman(info, createStyle(true)); + expect(styled).toContain("https://example.com/signup?flow=abc"); + expect(styled).toContain("\u001b]8;;https://example.com/signup?flow=abc"); + expect(styled).toContain("elicitation/respond e-u --done"); + expect(styled).toContain("elicitation/respond e-u --cancel"); + + const unsafe = formatElicitationPendingHuman( + { ...info, url: "file:///etc/passwd" }, + createStyle(true), + ); + expect(unsafe).toContain("file:///etc/passwd"); + expect(unsafe).not.toContain("\u001b]8"); + }); +}); + +describe("writeConnectionOutput", () => { + let stdout: string; + let stderr: string; + let original: typeof process.stdout.write; + let originalErr: typeof process.stderr.write; + + beforeEach(() => { + stdout = ""; + stderr = ""; + original = process.stdout.write; + originalErr = process.stderr.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]) => { + stdout += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + process.stderr.write = ((chunk: unknown, ...rest: unknown[]) => { + stderr += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stderr.write; + }); + + afterEach(() => { + process.stdout.write = original; + process.stderr.write = originalErr; + }); + + it("connection with authUrl: json carries the URL verbatim (query intact), human prints relay guidance", async () => { + const authUrl = "https://as.example/authorize?client_id=abc&state=xyz"; + await writeConnectionOutput( + { format: "json" }, + { + kind: "connection", + connection: { + name: "api", + serverIdentity: "https://mcp.example.com/mcp", + pendingAuth: true, + auth: { method: "oauth", authorized: false }, + }, + authUrl, + }, + ); + const parsed = JSON.parse(stdout) as Record<string, unknown>; + expect(parsed.pendingAuth).toBe(true); + expect(parsed.authUrl).toBe(authUrl); + + stdout = ""; + await writeConnectionOutput( + { format: "text" }, + { + kind: "connection", + connection: { + name: "api", + serverIdentity: "https://mcp.example.com/mcp", + pendingAuth: true, + auth: { method: "oauth", authorized: false }, + }, + authUrl, + }, + ); + expect(stdout).toContain("Sign-in required"); + expect(stdout).toContain(authUrl); + expect(stdout).toContain("Sign-in: pending"); + // Non-interactive (agent) framing: a factual statement to relay the URL + // and wait, with the resume command named for after confirmation. + expect(stdout).toContain("not usable yet until the user signs in"); + expect(stdout).toContain("wait for them to confirm"); + expect(stdout).toContain("mcpdo connections/show @api"); + expect(stdout).not.toContain("Open this link in a browser"); + + stdout = ""; + await writeConnectionOutput( + { format: "text", interactive: true }, + { + kind: "connection", + connection: { + name: "api", + serverIdentity: "https://mcp.example.com/mcp", + pendingAuth: true, + auth: { method: "oauth", authorized: false }, + }, + authUrl, + }, + ); + // Interactive (human) framing: address the reader directly, and it is the + // human path that gets the "check with connections/show" nudge. + expect(stdout).toContain("Open this link in a browser to authenticate"); + expect(stdout).toContain("connections/show @api"); + expect(stdout).toContain(authUrl); + expect(stdout).not.toContain("not usable yet until the user signs in"); + expect(stdout).not.toContain("wait for them to confirm"); + }); + + it("pendingAuthSignedIn: human output flips to completed / completing-on-next-use", async () => { + stdout = ""; + await writeConnectionOutput( + { format: "text" }, + { + kind: "connection", + connection: { + name: "api", + serverIdentity: "https://mcp.example.com/mcp", + pendingAuth: true, + pendingAuthSignedIn: true, + auth: { method: "oauth", authorized: true }, + }, + }, + ); + expect(stdout).toContain("Sign-in: completed"); + expect(stdout).toContain("finishes on next use"); + expect(stdout).not.toContain("Sign-in: pending"); + + stdout = ""; + await writeConnectionOutput( + { format: "text" }, + { + kind: "connections/list", + connections: [ + { + name: "api", + serverIdentity: "https://mcp.example.com/mcp", + connectedAt: 1, + lastAccessedAt: 1, + isMru: true, + pendingAuth: true, + pendingAuthSignedIn: true, + }, + { + name: "other", + serverIdentity: "https://mcp2.example.com/mcp", + connectedAt: 1, + lastAccessedAt: 1, + isMru: false, + pendingAuth: true, + }, + ], + }, + ); + expect(stdout).toContain("signed in — completing on next use"); + expect(stdout).toContain("(sign-in pending)"); + }); + + it("connection authUrl: only allowlisted schemes become OSC 8 links", async () => { + const connection = { + name: "api", + serverIdentity: "https://mcp.example.com/mcp", + pendingAuth: true, + auth: { method: "oauth", authorized: false }, + }; + const style = createStyle(true); + stdout = ""; + await writeConnectionOutput( + { format: "text", style, interactive: true }, + { + kind: "connection", + connection, + authUrl: "https://as.example/authorize?state=ok", + }, + ); + expect(stdout).toContain("\u001b]8;;https://as.example/authorize?state=ok"); + + // Server-controlled OAuth metadata: an unsafe scheme renders as plain + // text, never a clickable link. + stdout = ""; + await writeConnectionOutput( + { format: "text", style, interactive: true }, + { + kind: "connection", + connection, + authUrl: "file:///etc/passwd", + }, + ); + expect(stdout).toContain("file:///etc/passwd"); + expect(stdout).not.toContain("\u001b]8"); + }); + + it("pending ema-login renders relay guidance with the same link gate", async () => { + const authUrl = "https://idp.example.com/authorize?state=p1"; + await writeConnectionOutput( + { format: "text" }, + { + kind: "auth/ema-login", + result: { + issuer: "https://idp.example.com", + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl, + }, + }, + ); + expect(stdout).toContain("Sign-in required"); + expect(stdout).toContain(authUrl); + expect(stdout).toContain("auth/ema-status"); + + // JSON passthrough carries the pending fields verbatim. + stdout = ""; + await writeConnectionOutput( + { format: "json" }, + { + kind: "auth/ema-login", + result: { + issuer: "https://idp.example.com", + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl, + }, + }, + ); + const parsed = JSON.parse(stdout) as Record<string, unknown>; + expect(parsed.pendingLogin).toBe(true); + expect(parsed.authUrl).toBe(authUrl); + + // Unsafe scheme renders as plain text, never a clickable OSC 8 link. + stdout = ""; + await writeConnectionOutput( + { format: "text", style: createStyle(true) }, + { + kind: "auth/ema-login", + result: { + issuer: "https://idp.example.com", + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl: "file:///etc/passwd", + }, + }, + ); + expect(stdout).toContain("file:///etc/passwd"); + expect(stdout).not.toContain("\u001b]8"); + }); + + it("pending ema-login in interactive mode renders the human sign-in block with the link gate", async () => { + // Safe https URL becomes a clickable OSC 8 link in the human block. + await writeConnectionOutput( + { format: "text", style: createStyle(true), interactive: true }, + { + kind: "auth/ema-login", + result: { + issuer: "https://idp.example.com", + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl: "https://idp.example.com/authorize?state=h1", + }, + }, + ); + expect(stdout).toContain("Sign-in required"); + expect(stdout).toContain("\u001b]8"); + + // Unsafe scheme stays plain text even in the human block. + stdout = ""; + await writeConnectionOutput( + { format: "text", style: createStyle(true), interactive: true }, + { + kind: "auth/ema-login", + result: { + issuer: "https://idp.example.com", + loginState: "none", + alreadyLoggedIn: false, + pendingLogin: true, + authUrl: "file:///etc/passwd", + }, + }, + ); + expect(stdout).toContain("file:///etc/passwd"); + expect(stdout).not.toContain("\u001b]8"); + }); + + it("ema-logout routes the end-session URL through the same OSC 8 link gate", async () => { + // Safe https URL renders as a clickable OSC 8 link, like every other + // server-supplied URL in this formatter. + await writeConnectionOutput( + { format: "text", style: createStyle(true) }, + { + kind: "auth/ema-logout", + result: { + issuer: "https://idp.example.com", + endSessionUrl: "https://idp.example.com/session/end?id_token_hint=x", + }, + }, + ); + expect(stdout).toContain("To end your IdP browser session, navigate to:"); + expect(stdout).toContain("\u001b]8"); + + // Unsafe scheme stays plain text, never a clickable link. + stdout = ""; + await writeConnectionOutput( + { format: "text", style: createStyle(true) }, + { + kind: "auth/ema-logout", + result: { + issuer: "https://idp.example.com", + endSessionUrl: "file:///etc/passwd", + }, + }, + ); + expect(stdout).toContain("file:///etc/passwd"); + expect(stdout).not.toContain("\u001b]8"); + }); + + it("connection without authUrl renders exactly as before (no sign-in block)", async () => { + await writeConnectionOutput( + { format: "text" }, + { + kind: "connection", + connection: { name: "api", serverIdentity: "id" }, + }, + ); + expect(stdout).not.toContain("Sign-in"); + }); + + it("pretty-prints json without a result envelope", async () => { + await writeConnectionOutput( + { format: "json" }, + { + kind: "rpc", + method: "tools/list", + result: { tools: [] }, + }, + ); + expect(stdout).toBe('{\n "tools": []\n}\n'); + }); + + it("escapes C1 controls in json output (JSON.stringify only escapes C0)", async () => { + await writeConnectionOutput( + { format: "json" }, + { + kind: "rpc", + method: "tools/call", + result: { + content: [{ type: "text", text: "before\u009b31mafter" }], + }, + }, + ); + // U+009B is 8-bit CSI: it must reach the terminal as a \u escape, and + // parsing the output must restore the original value byte-for-byte. + expect(stdout).not.toContain("\u009b"); + expect(stdout).toContain("\\u009b"); + const parsed = JSON.parse(stdout) as { + content: { text: string }[]; + }; + expect(parsed.content[0]!.text).toBe("before\u009b31mafter"); + }); + + it("sanitizes server-supplied terminal escapes in text mode", async () => { + await writeConnectionOutput( + { format: "text" }, + { + kind: "rpc", + method: "tools/call", + result: { + content: [{ type: "text", text: "\u001b]52;c;c3RvbGVu\u0007hi" }], + }, + }, + ); + expect(stdout).not.toContain("\u001b"); + expect(stdout).not.toContain("\u0007"); + expect(stdout).toContain("\u241b]52;c;c3RvbGVu\u2407hi"); + }); + + it("leaves json output verbatim (JSON escaping already protects it)", async () => { + await writeConnectionOutput( + { format: "json" }, + { + kind: "rpc", + method: "tools/call", + result: { content: [{ type: "text", text: "\u001bhi" }] }, + }, + ); + expect(JSON.parse(stdout)).toEqual({ + content: [{ type: "text", text: "\u001bhi" }], + }); + expect(stdout).toContain("\\u001bhi"); + }); + + it("sanitizes the ndjson stderr summary line", async () => { + await writeConnectionOutput( + { format: "text" }, + { + kind: "ndjson", + variant: "skill-verify", + lines: [], + summary: "done \u001b[2J", + }, + ); + expect(stderr).toContain("done \u241b[2J"); + expect(stderr).not.toContain("\u001b"); + }); + + it("ignores auto-collected appInfo on tools/call json", async () => { + await writeConnectionOutput( + { format: "json" }, + { + kind: "rpc", + method: "tools/call", + result: { content: [{ type: "text", text: "ok" }] }, + appInfo: { hasApp: false, toolName: "echo" }, + }, + ); + expect(JSON.parse(stdout)).toEqual({ + content: [{ type: "text", text: "ok" }], + }); + }); + + it("throws NO_APP after printing app-info text", async () => { + await expect( + writeConnectionOutput( + { format: "text" }, + { + kind: "rpc", + method: "tools/call", + result: { hasApp: false, toolName: "x" }, + }, + ), + ).rejects.toMatchObject({ exitCode: EXIT_CODES.NO_APP }); + expect(stdout).toContain("has no MCP App"); + }); + + it("allows hasApp true app-info probes", async () => { + await writeConnectionOutput( + { format: "text" }, + { + kind: "rpc", + method: "tools/call", + result: { hasApp: true, toolName: "x", resourceUri: "ui://x" }, + }, + ); + expect(stdout).toContain("has an MCP App"); + }); + + it("throws TOOL_ERROR when isError", async () => { + await expect( + writeConnectionOutput( + { format: "json" }, + { + kind: "rpc", + method: "tools/call", + result: { isError: true, content: [] }, + toolName: "echo", + }, + ), + ).rejects.toBeInstanceOf(CliExitCodeError); + await expect( + writeConnectionOutput( + { format: "json" }, + { + kind: "rpc", + method: "tools/call", + result: { isError: true, content: [] }, + }, + ), + ).rejects.toMatchObject({ message: expect.stringContaining("tool") }); + // Server-influenced tool names are sanitized before reaching the + // terminal-bound error message (C0/C1 → visible stand-ins). + await expect( + writeConnectionOutput( + { format: "json" }, + { + kind: "rpc", + method: "tools/call", + result: { isError: true, content: [] }, + toolName: "evil\u001b]0;pwned\u0007", + }, + ), + ).rejects.toMatchObject({ + message: expect.not.stringContaining("\u001b"), + }); + }); + + it("falls back to pretty JSON for unknown rpc methods in text mode", async () => { + await writeConnectionOutput( + { format: "text" }, + { + kind: "rpc", + method: "custom/x", + result: { ok: 1 }, + }, + ); + expect(stdout).toContain('"ok": 1'); + }); + + it("renders skill-verify NDJSON with its own formatter, not app-info's", async () => { + await writeConnectionOutput( + { format: "text" }, + { + kind: "ndjson", + variant: "skill-verify", + lines: [ + { name: "ok-skill", uri: "skill://ok/SKILL.md", outcome: "verified" }, + ], + summary: "Verified 1 skill and 0 files: no conformance errors.", + }, + ); + expect(stdout).toContain("Skill verification (1):"); + expect(stdout).not.toContain("App info"); + expect(stderr).toBe( + "Verified 1 skill and 0 files: no conformance errors.\n", + ); + }); + + it("throws with the verify exit code after printing the report and summary", async () => { + await expect( + writeConnectionOutput( + { format: "json" }, + { + kind: "ndjson", + variant: "skill-verify", + lines: [ + { + name: "bad-skill", + uri: "skill://bad/SKILL.md", + outcome: "failed", + }, + ], + summary: "1 of 1 skill failed verification.", + exitCode: EXIT_CODES.SKILL_NONCONFORMANT, + }, + ), + ).rejects.toMatchObject({ + exitCode: EXIT_CODES.SKILL_NONCONFORMANT, + envelope: { code: "skills_nonconformant" }, + }); + // Report already on stdout, summary on stderr — both happen before the throw. + expect(stdout).toContain("bad-skill"); + expect(stderr).toBe("1 of 1 skill failed verification.\n"); + }); + + it("formats every admin/stream payload kind", async () => { + const kinds = [ + { + kind: "ndjson" as const, + lines: [{ toolName: "a", hasApp: false }], + }, + { kind: "stream-event" as const, data: { type: "subscribed", uri: "u" } }, + { + kind: "servers/list" as const, + servers: [{ name: "s", type: "stdio", detail: "d" }], + }, + { + kind: "servers/show" as const, + server: { + name: "s", + type: "stdio", + detail: "d", + config: { type: "stdio", command: "n" }, + }, + }, + { + kind: "connections/list" as const, + connections: [{ name: "a", serverIdentity: "id" }], + }, + { + kind: "connection" as const, + connection: { name: "a", serverIdentity: "id" }, + }, + { kind: "disconnect" as const, name: "a" }, + { + kind: "daemon/status" as const, + status: { pid: 1, socketPath: "/tmp/s", connections: [] }, + }, + { + kind: "daemon/status" as const, + status: { pid: 2, connections: "bad" }, + }, + { + kind: "daemon/stop" as const, + result: { stopping: false }, + }, + { + kind: "daemon/stop" as const, + result: { stopping: true }, + }, + { + kind: "daemon/stop" as const, + result: { stopping: false, message: "was idle" }, + }, + { + kind: "auth/list" as const, + list: { + oauthStatePath: "/tmp/oauth.json", + servers: [ + { + url: "https://example.com/mcp", + hasTokens: true, + hasRefreshToken: false, + }, + ], + }, + }, + { + kind: "auth/clear" as const, + result: { all: true, cleared: 1 }, + }, + { + kind: "auth/clear" as const, + result: { all: true, cleared: 2 }, + }, + { + kind: "auth/clear" as const, + result: { url: "https://example.com/mcp" }, + }, + { kind: "generic" as const, data: { x: 1 }, title: "Title" }, + { kind: "generic" as const, data: { y: 2 } }, + ]; + + for (const payload of kinds) { + stdout = ""; + await writeConnectionOutput({ format: "text" }, payload); + expect(stdout.length).toBeGreaterThan(0); + stdout = ""; + await writeConnectionOutput({ format: "json" }, payload); + expect(() => JSON.parse(stdout)).not.toThrow(); + } + }); + + it("defaults undefined format to text", async () => { + await writeConnectionOutput( + {}, + { + kind: "disconnect", + name: "z", + }, + ); + expect(stdout).toContain("Disconnected `@z`"); + }); + + it("notes cleared stored auth on disconnect --clear-auth (text + json)", async () => { + await writeConnectionOutput( + { format: "text" }, + { + kind: "disconnect", + name: "z", + clearedAuthUrl: "https://example.com/mcp", + }, + ); + expect(stdout).toContain("Disconnected `@z`"); + expect(stdout).toContain("Cleared stored auth for https://example.com/mcp"); + expect(stdout).toContain("re-trigger sign-in"); + + stdout = ""; + await writeConnectionOutput( + { format: "json" }, + { + kind: "disconnect", + name: "z", + clearedAuthUrl: "https://example.com/mcp", + }, + ); + expect(JSON.parse(stdout)).toEqual({ + name: "z", + clearedAuthUrl: "https://example.com/mcp", + }); + }); + + it("omits the clearedAuthUrl field when auth was not cleared (json)", async () => { + await writeConnectionOutput( + { format: "json" }, + { kind: "disconnect", name: "z" }, + ); + expect(JSON.parse(stdout)).toEqual({ name: "z" }); + }); +}); + +describe("format-human ANSI styling", () => { + it("styles human tool lists and log levels when enabled", () => { + const s = createStyle(true); + const tools = formatToolsHuman( + [ + { + name: "echo", + description: "hi", + inputSchema: { + type: "object", + properties: { message: { type: "string" } }, + required: ["message"], + }, + }, + ], + s, + ); + expect(tools).toContain("\u001b[1m"); // bold name + expect(tools).toContain("\u001b[36m"); // cyan params + expect(tools).toContain("\u001b[2m"); // dim description + expect(tools).toContain("echo"); + + const log = formatStreamEventHuman( + { + direction: "notification", + message: { params: { level: "error", data: "boom" } }, + }, + s, + ); + expect(log).toContain("\u001b[31m"); + expect(log).toContain("boom"); + }); + + it("hyperlinks only allowlisted schemes as OSC 8", () => { + const s = createStyle(true); + const out = formatResourcesHuman( + [ + { uri: "https://example.com/r", name: "web" }, + { uri: "file:///etc/passwd", name: "local" }, + { uri: "vscode://malicious/payload", name: "custom" }, + ], + s, + ); + // https renders as a clickable link; file:/custom-handler URIs must not + // invite the terminal to invoke a local protocol handler. + expect(out).toContain("\u001b]8;;https://example.com/r"); + expect(out).not.toContain("]8;;file://"); + expect(out).not.toContain("]8;;vscode://"); + expect(out).toContain("file:///etc/passwd"); + expect(out).toContain("vscode://malicious/payload"); + }); +}); diff --git a/clients/mcpdo/__tests__/helpers/mcp-runner.ts b/clients/mcpdo/__tests__/helpers/mcp-runner.ts new file mode 100644 index 0000000000..f8a43f4d34 --- /dev/null +++ b/clients/mcpdo/__tests__/helpers/mcp-runner.ts @@ -0,0 +1,88 @@ +import { runMcp as invokeMcp } from "../../src/connection/mcp.js"; +import { formatErrorOutput } from "@inspector/core/cli/error-handler.js"; + +export interface McpResult { + exitCode: number | null; + stdout: string; + stderr: string; + output: string; +} + +export interface McpOptions { + timeout?: number; + env?: Record<string, string>; +} + +type WriteArgs = [ + chunk: unknown, + encoding?: unknown, + callback?: (() => void) | undefined, +]; + +function captureWrite(append: (text: string) => void) { + return (...args: WriteArgs): boolean => { + const [chunk, encoding, callback] = args; + append(typeof chunk === "string" ? chunk : String(chunk)); + const cb = typeof encoding === "function" ? encoding : callback; + if (typeof cb === "function") cb(); + return true; + }; +} + +/** + * In-process runner for `runMcp` (connection CLI), mirroring {@link runCli}. + */ +export async function runMcp( + args: string[], + options: McpOptions = {}, +): Promise<McpResult> { + let stdout = ""; + let stderr = ""; + + const originalStdoutWrite = process.stdout.write; + const originalStderrWrite = process.stderr.write; + + const envBackup: Record<string, string | undefined> = {}; + if (options.env) { + for (const [key, value] of Object.entries(options.env)) { + envBackup[key] = process.env[key]; + process.env[key] = value; + } + } + + process.stdout.write = captureWrite((text) => { + stdout += text; + }) as typeof process.stdout.write; + process.stderr.write = captureWrite((text) => { + stderr += text; + }) as typeof process.stderr.write; + + const argv = ["node", "mcpdo", ...args]; + const timeoutMs = options.timeout ?? 15000; + let timer: ReturnType<typeof setTimeout> | undefined; + const timeout = new Promise<never>((_, reject) => { + timer = setTimeout( + () => reject(new Error(`mcpdo command timed out after ${timeoutMs}ms`)), + timeoutMs, + ); + }); + + let exitCode = 0; + try { + await Promise.race([invokeMcp(argv), timeout]); + } catch (error) { + const out = formatErrorOutput(error); + exitCode = out.exitCode; + stderr += out.stderr; + } finally { + if (timer) clearTimeout(timer); + process.stdout.write = originalStdoutWrite; + process.stderr.write = originalStderrWrite; + for (const [key, value] of Object.entries(envBackup)) { + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + } + + return { exitCode, stdout, stderr, output: stdout + stderr }; +} diff --git a/clients/mcpdo/__tests__/hoist-connection.test.ts b/clients/mcpdo/__tests__/hoist-connection.test.ts new file mode 100644 index 0000000000..613ae20018 --- /dev/null +++ b/clients/mcpdo/__tests__/hoist-connection.test.ts @@ -0,0 +1,128 @@ +import { describe, it, expect } from "vitest"; +import { hoistAtConnection } from "../src/connection/dispatch.js"; +import { expandConnAlias } from "../src/connection/mcp.js"; + +describe("hoistAtConnection", () => { + it("lifts a leading @name into connectionFromAt", () => { + const { argv, connectionFromAt } = hoistAtConnection([ + "node", + "mcpdo", + "@alpha", + "tools/list", + "--format", + "json", + ]); + expect(connectionFromAt).toBe("alpha"); + expect(argv).toEqual(["node", "mcpdo", "tools/list", "--format", "json"]); + }); + + it("leaves argv unchanged when there is no @name", () => { + const input = ["node", "mcpdo", "tools/list"]; + expect(hoistAtConnection(input)).toEqual({ argv: input }); + }); + + it("lifts @name appearing after global options (F4)", () => { + const { argv, connectionFromAt } = hoistAtConnection([ + "node", + "mcpdo", + "--format", + "json", + "@srv", + "tools/list", + ]); + expect(connectionFromAt).toBe("srv"); + expect(argv).toEqual(["node", "mcpdo", "--format", "json", "tools/list"]); + }); + + it("lifts @name after inline-value and boolean globals", () => { + const { argv, connectionFromAt } = hoistAtConnection([ + "node", + "mcpdo", + "--format=json", + "--plain", + "@srv", + "logging/tail", + ]); + expect(connectionFromAt).toBe("srv"); + expect(argv).toEqual([ + "node", + "mcpdo", + "--format=json", + "--plain", + "logging/tail", + ]); + }); + + it("does not claim an option value that looks like @name", () => { + const input = ["node", "mcpdo", "--connection", "@alpha", "tools/list"]; + // `--connection`'s value is consumed as a value, not hoisted (it is + // stripped of its @ by the consumption site instead). + expect(hoistAtConnection(input)).toEqual({ argv: input }); + }); + + it("does not claim an @-positional after the subcommand", () => { + const input = ["node", "mcpdo", "tools/call", "@notaconn"]; + expect(hoistAtConnection(input)).toEqual({ argv: input }); + }); + + it("stops at -- (child-process args)", () => { + const input = ["node", "mcpdo", "--", "@child-arg"]; + expect(hoistAtConnection(input)).toEqual({ argv: input }); + }); +}); + +describe("expandConnAlias", () => { + it("expands --conn and --conn=<name> to --connection forms", () => { + expect( + expandConnAlias(["node", "mcpdo", "--conn", "alpha", "tools/list"]), + ).toEqual(["node", "mcpdo", "--connection", "alpha", "tools/list"]); + expect(expandConnAlias(["node", "mcpdo", "--conn=alpha"])).toEqual([ + "node", + "mcpdo", + "--connection=alpha", + ]); + }); + + it("leaves --connection, --config, and other args unchanged", () => { + const input = [ + "node", + "mcpdo", + "--connection", + "alpha", + "--config", + "x.json", + "--connect-timeout", + "5", + ]; + expect(expandConnAlias(input)).toEqual(input); + }); + + it("passes tokens after -- through verbatim (child-process args)", () => { + expect( + expandConnAlias([ + "node", + "mcpdo", + "--conn", + "alpha", + "connect", + "srv", + "--", + "--conn=value", + "--conn", + ]), + ).toEqual([ + "node", + "mcpdo", + "--connection", + "alpha", + "connect", + "srv", + "--", + "--conn=value", + "--conn", + ]); + // Only the first separator ends expansion; later ones are child args too. + const onlyAfter = ["node", "mcpdo", "--", "--conn", "--", "--conn=x"]; + expect(expandConnAlias(onlyAfter)).toEqual(onlyAfter); + }); +}); diff --git a/clients/mcpdo/__tests__/mcp-auth-coverage.test.ts b/clients/mcpdo/__tests__/mcp-auth-coverage.test.ts new file mode 100644 index 0000000000..3e0980c9d0 --- /dev/null +++ b/clients/mcpdo/__tests__/mcp-auth-coverage.test.ts @@ -0,0 +1,495 @@ +import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; +import { + createSampleTestConfig, + deleteConfigFile, +} from "../../cli/__tests__/helpers/fixtures.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; + +const callDaemon = vi.fn(); +const ensureDaemon = vi.fn(); +const authorizeInFrontend = vi.fn(); +const obtainPendingAuthUrl = vi.fn(); + +vi.mock("../src/daemon/index.js", () => ({ + callDaemon: (...args: unknown[]) => callDaemon(...args), + ensureDaemon: (...args: unknown[]) => ensureDaemon(...args), + streamDaemon: vi.fn(), +})); + +vi.mock("../src/connection/authorize.js", () => ({ + authorizeInFrontend: (...args: unknown[]) => authorizeInFrontend(...args), +})); + +vi.mock("../src/connection/auth-helper.js", () => ({ + AUTH_HELPER_COMMAND: "auth/complete-signin", + PENDING_AUTH_TTL_MS: 15 * 60 * 1000, + runAuthHelper: vi.fn(), + obtainPendingAuthUrl: (...args: unknown[]) => obtainPendingAuthUrl(...args), + // ema-login-helper.js (imported real by mcp.ts) also pulls these from the + // mocked module; stubs keep its import resolvable. + obtainPendingUrlForKey: vi.fn(), + pendingAuthMarkerPath: vi.fn(() => "/tmp/pending-auth-stub.json"), + readLivePendingAuthMarker: vi.fn(), + removeOwnPendingAuthMarker: vi.fn(), + writePendingAuthMarker: vi.fn(), +})); + +describe("mcp.ts auth / daemon error paths", () => { + let configPath: string | undefined; + let stdout: string; + let originalStdoutWrite: typeof process.stdout.write; + let originalStderrWrite: typeof process.stderr.write; + + beforeEach(() => { + stdout = ""; + originalStdoutWrite = process.stdout.write; + originalStderrWrite = process.stderr.write; + process.stdout.write = ((chunk: unknown, ...rest: unknown[]) => { + stdout += typeof chunk === "string" ? chunk : String(chunk); + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stdout.write; + process.stderr.write = ((chunk: unknown, ...rest: unknown[]) => { + const cb = rest.find((r) => typeof r === "function") as + | (() => void) + | undefined; + cb?.(); + return true; + }) as typeof process.stderr.write; + + ensureDaemon.mockReset(); + ensureDaemon.mockResolvedValue({ socketPath: "/tmp/mcp-auth-cov.sock" }); + callDaemon.mockReset(); + authorizeInFrontend.mockReset(); + authorizeInFrontend.mockResolvedValue(undefined); + obtainPendingAuthUrl.mockReset(); + }); + + const originalStderrIsTTY = process.stderr.isTTY; + const originalStdinIsTTY = process.stdin.isTTY; + + afterEach(() => { + process.stderr.isTTY = originalStderrIsTTY; + process.stdin.isTTY = originalStdinIsTTY; + process.stdout.write = originalStdoutWrite; + process.stderr.write = originalStderrWrite; + if (configPath) { + deleteConfigFile(configPath); + configPath = undefined; + } + }); + + it("connect --ema overlays enterpriseManaged onto the resolved settings", async () => { + configPath = createSampleTestConfig(); + callDaemon.mockResolvedValueOnce({ + name: "test-stdio", + isMru: true, + serverIdentity: "stdio", + }); + + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connect", + "test-stdio", + "--config", + configPath, + "--ema", + "--format", + "json", + ]); + + const connectCall = callDaemon.mock.calls.find((c) => c[0] === "connect"); + const params = connectCall?.[1] as { + serverSettings?: { enterpriseManaged?: boolean }; + }; + expect(params.serverSettings?.enterpriseManaged).toBe(true); + }); + + it("retries connect after auth_required via authorizeInFrontend", async () => { + process.stderr.isTTY = true; // human path: blocking interactive OAuth + configPath = createSampleTestConfig(); + const connection = { + name: "test-stdio", + isMru: true, + serverIdentity: "stdio", + }; + callDaemon + .mockRejectedValueOnce( + new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, "need auth", { + code: "auth_required", + }), + ) + .mockResolvedValueOnce(connection); + + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connect", + "test-stdio", + "--config", + configPath, + "--format", + "json", + ]); + + expect(authorizeInFrontend).toHaveBeenCalledOnce(); + expect(callDaemon).toHaveBeenCalledTimes(2); + expect(JSON.parse(stdout.trim()).name).toBe("test-stdio"); + }); + + it("re-ensures the daemon after authorizeInFrontend, in case interactive OAuth outlasted its idle timeout", async () => { + process.stderr.isTTY = true; // human path: blocking interactive OAuth + configPath = createSampleTestConfig(); + const connection = { + name: "test-stdio", + isMru: true, + serverIdentity: "stdio", + }; + callDaemon + .mockRejectedValueOnce( + new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, "need auth", { + code: "auth_required", + }), + ) + .mockResolvedValueOnce(connection); + // Simulate the pre-auth daemon having idled out while OAuth ran: the + // retry's ensureDaemon() call returns a different (freshly respawned) + // socket than the one used for the first attempt. + ensureDaemon + .mockResolvedValueOnce({ socketPath: "/tmp/mcp-auth-cov-stale.sock" }) + .mockResolvedValueOnce({ socketPath: "/tmp/mcp-auth-cov-fresh.sock" }); + + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connect", + "test-stdio", + "--config", + configPath, + "--format", + "json", + ]); + + expect(ensureDaemon).toHaveBeenCalledTimes(2); + expect(callDaemon).toHaveBeenCalledTimes(2); + expect(callDaemon.mock.calls[0][2]).toMatchObject({ + socketPath: "/tmp/mcp-auth-cov-stale.sock", + }); + expect(callDaemon.mock.calls[1][2]).toMatchObject({ + socketPath: "/tmp/mcp-auth-cov-fresh.sock", + }); + }); + + it("non-TTY connect on auth_required: hands off to the helper, registers a pending entry, and prints the auth URL", async () => { + // Agent path: no TTY on stdin or stderr. + process.stdin.isTTY = undefined as unknown as boolean; + process.stderr.isTTY = undefined as unknown as boolean; + configPath = createSampleTestConfig(); + callDaemon + .mockRejectedValueOnce( + new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, "need auth", { + code: "auth_required", + }), + ) + .mockResolvedValueOnce({ + name: "test-stdio", + isMru: true, + serverIdentity: "stdio", + pendingAuth: true, + auth: { method: "oauth", authorized: false }, + }); + obtainPendingAuthUrl.mockResolvedValueOnce( + "https://as.example/authorize?client_id=abc&state=xyz", + ); + + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connect", + "test-stdio", + "--config", + configPath, + "--format", + "json", + ]); + + // Never the blocking interactive flow on the agent path. + expect(authorizeInFrontend).not.toHaveBeenCalled(); + expect(obtainPendingAuthUrl).toHaveBeenCalledOnce(); + // The re-dial carries the pending-intent flag. + const second = callDaemon.mock.calls[1]; + expect(second[0]).toBe("connect"); + expect(second[1]).toMatchObject({ pendingOnAuthRequired: true }); + // The auth URL rides the normal JSON payload, query string intact + // (the error envelope would redact it). + const out = JSON.parse(stdout.trim()) as Record<string, unknown>; + expect(out.pendingAuth).toBe(true); + expect(out.authUrl).toBe( + "https://as.example/authorize?client_id=abc&state=xyz", + ); + }); + + it("non-TTY connect omits authUrl when the pending re-dial actually connected (sign-in already finished)", async () => { + process.stdin.isTTY = undefined as unknown as boolean; + process.stderr.isTTY = undefined as unknown as boolean; + configPath = createSampleTestConfig(); + callDaemon + .mockRejectedValueOnce( + new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, "need auth", { + code: "auth_required", + }), + ) + .mockResolvedValueOnce({ + name: "test-stdio", + isMru: true, + serverIdentity: "stdio", + auth: { method: "oauth", authorized: true }, + }); + obtainPendingAuthUrl.mockResolvedValueOnce("https://as.example/authorize"); + + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connect", + "test-stdio", + "--config", + configPath, + "--format", + "json", + ]); + + const out = JSON.parse(stdout.trim()) as Record<string, unknown>; + expect(out.pendingAuth).toBeUndefined(); + expect(out.authUrl).toBeUndefined(); + }); + + it("MCP_AUTO_OPEN_ENABLED=true keeps the blocking interactive flow even without a TTY", async () => { + process.stdin.isTTY = undefined as unknown as boolean; + process.stderr.isTTY = undefined as unknown as boolean; + const prev = process.env.MCP_AUTO_OPEN_ENABLED; + process.env.MCP_AUTO_OPEN_ENABLED = "true"; + configPath = createSampleTestConfig(); + callDaemon + .mockRejectedValueOnce( + new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, "need auth", { + code: "auth_required", + }), + ) + .mockResolvedValueOnce({ + name: "test-stdio", + isMru: true, + serverIdentity: "stdio", + }); + + try { + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connect", + "test-stdio", + "--config", + configPath, + "--format", + "json", + ]); + expect(authorizeInFrontend).toHaveBeenCalledOnce(); + expect(obtainPendingAuthUrl).not.toHaveBeenCalled(); + } finally { + if (prev === undefined) delete process.env.MCP_AUTO_OPEN_ENABLED; + else process.env.MCP_AUTO_OPEN_ENABLED = prev; + } + }); + + it("rejects --relogin with --stored-auth-only", async () => { + configPath = createSampleTestConfig(); + const { runMcp } = await import("../src/connection/mcp.js"); + await expect( + runMcp([ + "node", + "mcpdo", + "--stored-auth-only", + "connect", + "test-stdio", + "--config", + configPath, + "--relogin", + ]), + ).rejects.toMatchObject({ exitCode: 1 }); + expect(callDaemon).not.toHaveBeenCalled(); + }); + + it("clears stored auth on connect --relogin for HTTP targets", async () => { + const fs = await import("node:fs"); + const os = await import("node:os"); + const path = await import("node:path"); + const { resetNodeOAuthStorageCache } = + await import("@inspector/core/auth/node/storage-node.js"); + const dir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-relogin-")); + const oauthFile = path.join(dir, "oauth.json"); + fs.writeFileSync( + oauthFile, + JSON.stringify({ + servers: { + "http://example.com/mcp": { + tokens: { access_token: "x", token_type: "Bearer" }, + }, + }, + idpSessions: {}, + }), + "utf8", + ); + const prev = process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = oauthFile; + resetNodeOAuthStorageCache(); + + callDaemon.mockResolvedValueOnce({ + name: "http", + isMru: true, + serverIdentity: "http://example.com/mcp", + }); + + try { + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connect", + "--connection", + "relogin-http", + "--server-url", + "http://example.com/mcp", + "--transport", + "http", + "--relogin", + "--format", + "json", + ]); + expect(callDaemon).toHaveBeenCalledOnce(); + const after = JSON.parse(fs.readFileSync(oauthFile, "utf8")) as { + servers?: Record<string, unknown>; + }; + expect(after.servers?.["http://example.com/mcp"]).toBeUndefined(); + } finally { + if (prev === undefined) delete process.env.MCP_INSPECTOR_OAUTH_STATE_PATH; + else process.env.MCP_INSPECTOR_OAUTH_STATE_PATH = prev; + resetNodeOAuthStorageCache(); + fs.rmSync(dir, { recursive: true, force: true }); + } + }); + + it("rethrows auth_required when --stored-auth-only is set", async () => { + configPath = createSampleTestConfig(); + callDaemon.mockRejectedValueOnce( + new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, "need auth", { + code: "auth_required", + }), + ); + + const { runMcp } = await import("../src/connection/mcp.js"); + await expect( + runMcp([ + "node", + "mcpdo", + "connect", + "test-stdio", + "--config", + configPath, + "--stored-auth-only", + "--format", + "json", + ]), + ).rejects.toMatchObject({ + exitCode: EXIT_CODES.AUTH_REQUIRED, + envelope: { code: "auth_required" }, + }); + expect(authorizeInFrontend).not.toHaveBeenCalled(); + }); + + it("rethrows unexpected daemon/stop errors", async () => { + callDaemon.mockRejectedValueOnce( + new CliExitCodeError(EXIT_CODES.USAGE, "boom", { code: "usage" }), + ); + + const { runMcp } = await import("../src/connection/mcp.js"); + await expect( + runMcp(["node", "mcpdo", "daemon", "stop", "--format", "json"]), + ).rejects.toMatchObject({ + exitCode: EXIT_CODES.USAGE, + envelope: { code: "usage" }, + }); + }); + + describe("connections/show pending-auth URL relay (surface 1 front-end)", () => { + const RELAY_URL = "https://as.example/authorize?client_id=abc&state=xyz"; + const showResult = { + name: "api", + serverIdentity: "https://mcp.example.com/mcp", + pendingAuth: true, + authUrl: RELAY_URL, + }; + + it("--format json always carries authUrl (machine-readable), query intact", async () => { + process.stderr.isTTY = true; // TTY must not suppress it for JSON. + callDaemon.mockResolvedValueOnce(showResult); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connections/show", + "api", + "--format", + "json", + ]); + const parsed = JSON.parse(stdout) as { authUrl?: string }; + expect(parsed.authUrl).toBe(RELAY_URL); + }); + + it("no human present (non-TTY): human text prints the relay block + URL", async () => { + process.stderr.isTTY = false; + process.stdin.isTTY = false; + callDaemon.mockResolvedValueOnce(showResult); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connections/show", + "api", + "--format", + "text", + ]); + expect(stdout).toContain("Sign-in required"); + expect(stdout).toContain(RELAY_URL); + }); + + it("human present (TTY): human text omits the relay block entirely", async () => { + process.stderr.isTTY = true; + callDaemon.mockResolvedValueOnce(showResult); + const { runMcp } = await import("../src/connection/mcp.js"); + await runMcp([ + "node", + "mcpdo", + "connections/show", + "api", + "--format", + "text", + ]); + // The connection still renders (pending status), but the agent-relay + // URL block — and the URL — are suppressed for a human at a TTY. + expect(stdout).not.toContain("Sign-in required"); + expect(stdout).not.toContain(RELAY_URL); + }); + }); +}); diff --git a/clients/mcpdo/__tests__/mcp-connection.test.ts b/clients/mcpdo/__tests__/mcp-connection.test.ts new file mode 100644 index 0000000000..478453bbfb --- /dev/null +++ b/clients/mcpdo/__tests__/mcp-connection.test.ts @@ -0,0 +1,224 @@ +import { describe, it, expect, afterEach, beforeAll } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { runMcp } from "./helpers/mcp-runner.js"; +import { runCli } from "../../cli/__tests__/helpers/cli-runner.js"; +import { + createSampleTestConfig, + deleteConfigFile, +} from "../../cli/__tests__/helpers/fixtures.js"; +import { expectCliSuccess } from "../../cli/__tests__/helpers/assertions.js"; +import { resolveDaemonScriptPath } from "../src/daemon/ensure.js"; +import { callDaemon } from "../src/daemon/client.js"; + +describe("mcp connection CLI", () => { + let configPath: string | undefined; + let storageDir: string | undefined; + + beforeAll(() => { + // Auto-spawn needs the built daemon bundle. + expect(fs.existsSync(resolveDaemonScriptPath())).toBe(true); + }); + + afterEach(async () => { + if (storageDir) { + const socketPath = path.join(storageDir, "mcpdod.sock"); + if (fs.existsSync(socketPath)) { + try { + await callDaemon("daemon/stop", {}, { socketPath, timeoutMs: 2000 }); + } catch { + // already stopped + } + const deadline = Date.now() + 2000; + while (fs.existsSync(socketPath) && Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 50)); + } + } + fs.rmSync(storageDir, { recursive: true, force: true }); + storageDir = undefined; + } + if (configPath) { + deleteConfigFile(configPath); + configPath = undefined; + } + }); + + function env(): Record<string, string> { + storageDir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-connection-")); + return { + MCP_STORAGE_DIR: storageDir, + MCP_INSPECTOR_DAEMON_DIR: storageDir, + MCP_ALLOW_DEFAULT_CONNECTION: "1", + }; + } + + it("lists servers without a daemon", async () => { + configPath = createSampleTestConfig(); + // No MCP_STORAGE_DIR — this path must not touch the daemon. + const result = await runMcp([ + "servers/list", + "--config", + configPath, + "--format", + "json", + ]); + expectCliSuccess(result); + const body = JSON.parse(result.stdout) as { + servers: { name: string }[]; + }; + expect(body.servers.some((s) => s.name === "test-stdio")).toBe(true); + }); + + it("connects, lists connections, disconnects via auto-spawned daemon", async () => { + configPath = createSampleTestConfig(); + const e = env(); + + const connected = await runMcp( + ["connect", "test-stdio", "--config", configPath, "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(connected); + const connection = JSON.parse(connected.stdout) as { + name: string; + isMru: boolean; + }; + expect(connection.name).toBe("test-stdio"); + expect(connection.isMru).toBe(true); + + const listed = await runMcp(["connections/list", "--format", "json"], { + env: e, + }); + expectCliSuccess(listed); + const connections = JSON.parse(listed.stdout) as { + connections: { name: string; isMru: boolean }[]; + }; + expect(connections.connections).toHaveLength(1); + expect(connections.connections[0]?.name).toBe("test-stdio"); + + const servers = await runMcp( + ["servers/list", "--config", configPath, "--format", "json"], + { env: e }, + ); + expectCliSuccess(servers); + const serverBody = JSON.parse(servers.stdout) as { + servers: { + name: string; + connection?: string; + isMru?: boolean; + }[]; + }; + const stdio = serverBody.servers.find((s) => s.name === "test-stdio"); + expect(stdio?.connection).toBe("test-stdio"); + expect(stdio?.isMru).toBe(true); + expect( + serverBody.servers.find((s) => s.name === "test-http")?.connection, + ).toBeUndefined(); + + const disc = await runMcp( + ["disconnect", "--connection", "test-stdio", "--format", "json"], + { env: e }, + ); + expectCliSuccess(disc); + + const stopped = await runMcp(["daemon", "stop", "--format", "json"], { + env: e, + }); + expectCliSuccess(stopped); + }); + + it("one-shot servers/list still works alongside connection mode", async () => { + configPath = createSampleTestConfig(); + const result = await runCli([ + "--config", + configPath, + "--method", + "servers/list", + ]); + expectCliSuccess(result); + expect(result.stdout).toContain("test-stdio"); + }); + + it("runs tools/list, tools/call, and connections/show over a live connection", async () => { + configPath = createSampleTestConfig(); + const e = env(); + + const connected = await runMcp( + ["connect", "test-stdio", "--config", configPath, "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(connected); + + const tools = await runMcp(["tools/list", "--format", "json"], { + env: e, + timeout: 20000, + }); + expectCliSuccess(tools); + const toolsBody = JSON.parse(tools.stdout) as { + tools: { name: string }[]; + }; + expect(toolsBody.tools.length).toBeGreaterThan(0); + + const toolsText = await runMcp(["tools/list"], { + env: e, + timeout: 20000, + }); + expectCliSuccess(toolsText); + expect(toolsText.stdout).toMatch(/Tools \(\d+\):/); + expect(toolsText.stdout).toContain("`"); + + const called = await runMcp( + ["tools/call", "echo", "message:=connection", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(called); + + const calledJson = await runMcp( + [ + "tools/call", + "echo", + '{"message":"connection-json"}', + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(calledJson); + + const resources = await runMcp(["resources/list", "--format", "json"], { + env: e, + timeout: 20000, + }); + expectCliSuccess(resources); + + const shown = await runMcp( + ["@test-stdio", "connections/show", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(shown); + const shownBody = JSON.parse(shown.stdout) as { + name?: string; + serverInfo?: { name?: string }; + protocolVersion?: string; + protocolEra?: string; + }; + expect(shownBody.protocolVersion).toBeTruthy(); + expect(shownBody.protocolEra).toBeTruthy(); + + // `connections/show <name>` (positional, no `@name`/--connection) exercises + // the opts.connection-absent fallback to the command's own argument. + const shownByArg = await runMcp( + ["connections/show", "test-stdio", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(shownByArg); + + await runMcp( + ["disconnect", "--connection", "test-stdio", "--format", "json"], + { + env: e, + }, + ); + await runMcp(["daemon", "stop", "--format", "json"], { env: e }); + }); +}); diff --git a/clients/mcpdo/__tests__/mcp-coverage.test.ts b/clients/mcpdo/__tests__/mcp-coverage.test.ts new file mode 100644 index 0000000000..1422bbf4ee --- /dev/null +++ b/clients/mcpdo/__tests__/mcp-coverage.test.ts @@ -0,0 +1,560 @@ +import { describe, it, expect, afterEach, beforeAll } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { getTestMcpServerCommand } from "@modelcontextprotocol/inspector-test-server"; +import { runMcp } from "./helpers/mcp-runner.js"; +import { + createSampleTestConfig, + createTestConfig, + deleteConfigFile, +} from "../../cli/__tests__/helpers/fixtures.js"; +import { + expectCliSuccess, + expectCliFailure, +} from "../../cli/__tests__/helpers/assertions.js"; +import { resolveDaemonScriptPath } from "../src/daemon/ensure.js"; +import { callDaemon } from "../src/daemon/client.js"; + +describe("mcp.ts coverage", () => { + let configPath: string | undefined; + let storageDir: string | undefined; + + beforeAll(() => { + expect(fs.existsSync(resolveDaemonScriptPath())).toBe(true); + }); + + afterEach(async () => { + if (storageDir) { + const socketPath = path.join(storageDir, "mcpdod.sock"); + if (fs.existsSync(socketPath)) { + try { + await callDaemon("daemon/stop", {}, { socketPath, timeoutMs: 2000 }); + } catch { + // already stopped + } + const deadline = Date.now() + 2000; + while (fs.existsSync(socketPath) && Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 50)); + } + } + fs.rmSync(storageDir, { recursive: true, force: true }); + storageDir = undefined; + } + if (configPath) { + deleteConfigFile(configPath); + configPath = undefined; + } + }); + + function env(): Record<string, string> { + storageDir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-cov-")); + return { + MCP_STORAGE_DIR: storageDir, + MCP_INSPECTOR_DAEMON_DIR: storageDir, + MCP_ALLOW_DEFAULT_CONNECTION: "1", + }; + } + + it("covers RPC registrations, metadata parse, and --plain", async () => { + configPath = createSampleTestConfig(); + const e = env(); + + const connected = await runMcp( + ["connect", "test-stdio", "--config", configPath, "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(connected); + + const withMeta = await runMcp( + [ + "tools/list", + "--metadata", + "client=connection-cov", + "--metadata", + "count=1", + // Object value must JSON.stringify (not String → "[object Object]"). + "--metadata", + 'nested={"a":1}', + "--plain", + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(withMeta); + + const badMeta = await runMcp(["tools/list", "--metadata", "novalue"], { + env: e, + }); + expectCliFailure(badMeta); + + const emptyMeta = await runMcp(["tools/list", "--metadata", "k="], { + env: e, + }); + expectCliFailure(emptyMeta); + + const read = await runMcp( + [ + "resources/read", + "demo://resource/static/document/architecture.md", + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(read); + + // Same Commander action as subscribe (uri positional / --uri); prefer + // unsubscribe so we don't open a long-lived stream in this suite. + const unsub = await runMcp( + ["resources/unsubscribe", "test://env", "--format", "json"], + { env: e, timeout: 20000 }, + ); + // Default test server does not advertise subscriptions. + expectCliFailure(unsub); + expect(unsub.stderr).toMatch(/unsubscribe|Method not found/i); + + const prompt = await runMcp( + [ + "prompts/get", + "simple_prompt", + "--prompt-args", + "unused=1", + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(prompt); + + const completeBad = await runMcp( + ["prompts/complete", "--complete-ref-type", "nope"], + { env: e }, + ); + expectCliFailure(completeBad); + + const complete = await runMcp( + [ + "prompts/complete", + "--complete-ref-type", + "ref/prompt", + "--complete-ref", + "simple_prompt", + "--complete-arg-name", + "name", + "--complete-arg-value", + "s", + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + // Completion support varies; assert the command ran (not a usage parse error). + expect(complete.stderr).not.toMatch(/complete-ref-type/); + expect([0, 1]).toContain(complete.exitCode); + + const logOk = await runMcp( + ["logging/setLevel", "debug", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(logOk); + + const logBad = await runMcp(["logging/setLevel", "--log-level", "nope"], { + env: e, + }); + expectCliFailure(logBad); + + const taskGet = await runMcp( + ["tasks/get", "missing-task", "--format", "json"], + { + env: e, + timeout: 20000, + }, + ); + expectCliFailure(taskGet); + + const taskCancel = await runMcp( + ["tasks/cancel", "--task-id", "missing-task", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliFailure(taskCancel); + + const taskResult = await runMcp( + ["tasks/result", "missing-task", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliFailure(taskResult); + + const taskUpdateNoBody = await runMcp( + ["tasks/update", "missing-task", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliFailure(taskUpdateNoBody); + expect(taskUpdateNoBody.stderr).toMatch(/--input-responses/); + + const taskUpdateBadJson = await runMcp( + [ + "tasks/update", + "missing-task", + "--input-responses", + "not-json", + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliFailure(taskUpdateBadJson); + expect(taskUpdateBadJson.stderr).toMatch(/--input-responses is invalid/); + + const taskUpdate = await runMcp( + [ + "tasks/update", + "missing-task", + "--input-responses", + '{"req-1":"answer"}', + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliFailure(taskUpdate); + + const roots = await runMcp( + ["roots/set", "--roots-json", "[]", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(roots); + + const called = await runMcp( + [ + "tools/call", + "--tool-name", + "echo", + "--tool-arg", + "message=cov", + "--tool-metadata", + "src=test", + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(called); + + const templates = await runMcp( + ["resources/templates/list", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(templates); + + const prompts = await runMcp(["prompts/list", "--format", "json"], { + env: e, + timeout: 20000, + }); + expectCliSuccess(prompts); + + const tasks = await runMcp(["tasks/list", "--format", "json"], { + env: e, + timeout: 20000, + }); + expectCliSuccess(tasks); + expect(JSON.parse(tasks.stdout)).toHaveProperty("tasks"); + + const rootsList = await runMcp(["roots/list", "--format", "json"], { + env: e, + timeout: 20000, + }); + expectCliSuccess(rootsList); + expect(JSON.parse(rootsList.stdout)).toHaveProperty("roots"); + + const show = await runMcp( + [ + "servers/show", + "test-stdio", + "--config", + configPath, + "--format", + "json", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(show); + // Entry provenance rides along, mirroring servers/list. + expect(JSON.parse(show.stdout).source).toMatchObject({ kind: "config" }); + + // No name: falls back to the MRU connection's entry (test-stdio is + // connected above and MCP_ALLOW_DEFAULT_CONNECTION opts non-TTY in). + const showMru = await runMcp( + ["servers/show", "--config", configPath, "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(showMru); + expect(JSON.parse(showMru.stdout).name).toBe("test-stdio"); + + // MRU-inferred name missing from the shell's source: the error must + // explain the name came from the MRU connection, not look like a typo. + const otherConfig = createTestConfig({ + mcpServers: { unrelated: { type: "stdio", command: "true" } }, + }); + try { + const crossCatalog = await runMcp( + ["servers/show", "--config", otherConfig], + { env: e, timeout: 20000 }, + ); + expectCliFailure(crossCatalog); + expect(crossCatalog.stderr).toMatch( + /most-recently-used connection 'test-stdio' has no entry in config/, + ); + } finally { + deleteConfigFile(otherConfig); + } + + // Skills support is optional; the default test server may not advertise + // it. Either way, the RPC action itself should run (not a usage error). + const skillsList = await runMcp(["skills/list", "--format", "json"], { + env: e, + timeout: 20000, + }); + expect([0, 1]).toContain(skillsList.exitCode); + + const skillsListVerify = await runMcp( + ["skills/list", "--verify", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expect([0, 1]).toContain(skillsListVerify.exitCode); + + const skillsGet = await runMcp( + ["skills/get", "test://skill", "--verify", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expect([0, 1]).toContain(skillsGet.exitCode); + + const skillsGetFlagUri = await runMcp( + ["skills/get", "--uri", "test://skill", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expect([0, 1]).toContain(skillsGetFlagUri.exitCode); + + await runMcp( + ["disconnect", "--connection", "test-stdio", "--format", "json"], + { + env: e, + }, + ); + await runMcp(["daemon", "stop", "--format", "json"], { env: e }); + }); + + it("covers ad-hoc connect options and servers/list catalog env", async () => { + configPath = createSampleTestConfig(); + const e = env(); + const { command, args } = getTestMcpServerCommand(); + + const adHoc = await runMcp( + [ + "connect", + "--connection", + "opts", + "--transport", + "stdio", + "--cwd", + process.cwd(), + "-e", + "COV_FLAG=1", + "--connect-timeout", + "15000", + "--era", + "auto", + "--elicit", + "url", + "--format", + "json", + command, + ...args, + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(adHoc); + + // --ema is rejected for ad-hoc targets: EMA needs per-server OAuth + // client id/secret, which only a catalog entry can supply. + const emaAdHoc = await runMcp( + ["connect", "--ema", "--format", "json", command, ...args], + { env: e, timeout: 20000 }, + ); + expectCliFailure(emaAdHoc); + expect(emaAdHoc.stderr).toMatch(/--ema cannot be used with an ad-hoc/); + + // Invalid --era is rejected before any connection is attempted. + const badEra = await runMcp( + ["connect", "--era", "bogus", "--format", "json", command, ...args], + { env: e, timeout: 20000 }, + ); + expectCliFailure(badEra); + expect(badEra.stderr).toMatch(/Invalid --era/); + + // Invalid --elicit is rejected before any connection is attempted. + const badElicit = await runMcp( + ["connect", "--elicit", "bogus", "--format", "json", command, ...args], + { env: e, timeout: 20000 }, + ); + expectCliFailure(badElicit); + expect(badElicit.stderr).toMatch(/Invalid --elicit/); + + await runMcp(["disconnect", "--connection", "opts", "--format", "json"], { + env: e, + }); + + // Ad-hoc HTTP with --server-url and no positional rest (empty-rest branch). + const urlOnly = await runMcp( + [ + "connect", + "--connection", + "urlonly", + "--transport", + "http", + "--server-url", + "http://127.0.0.1:9/mcp", + "--header", + "X-Test: 1", + "--connect-timeout", + "100", + "--format", + "json", + ], + { env: e, timeout: 10000 }, + ); + expectCliFailure(urlOnly); + // Unreachable HTTP should classify as exit 4 when the error is network-shaped. + expect([1, 4]).toContain(urlOnly.exitCode); + + const listed = await runMcp(["servers/list", "--format", "json"], { + env: { + ...e, + MCP_CATALOG_PATH: configPath, + }, + }); + expectCliSuccess(listed); + + // Whitespace --config → trim || undefined branch on servers/list. + const emptyConfig = await runMcp( + ["servers/list", "--config", " ", "--format", "json"], + { env: { ...e, MCP_CATALOG_PATH: configPath } }, + ); + expectCliSuccess(emptyConfig); + + await runMcp(["daemon", "stop", "--format", "json"], { env: e }); + }); + + it("treats a single path-like token as ad-hoc stdio, a bare word as a catalog name", async () => { + configPath = createSampleTestConfig(); + const e = { ...env(), MCP_CATALOG_PATH: configPath }; + + // Bare word: catalog/config lookup — fails as a catalog miss, before + // any daemon or spawn work. + const bareWord = await runMcp( + ["connect", "no-such-catalog-entry", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliFailure(bareWord); + expect(bareWord.stderr).toMatch(/not found/i); + + // Path-like token: ad-hoc stdio target — never touches the catalog, so + // the failure is a spawn/connect failure, not a catalog miss. + const pathToken = await runMcp( + ["connect", "./no-such-server-binary", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliFailure(pathToken); + expect(pathToken.stderr).not.toMatch(/not found\. Available servers/); + }); + + it("bare mcpdo / --help print usage without an ErrorEnvelope", async () => { + // Bare invocation: Commander writes help to stderr (help-after-error). + const bare = await runMcp([]); + expectCliSuccess(bare); + expect(bare.stderr).toMatch(/Usage:/i); + expect(bare.stderr).not.toContain('"error"'); + + const help = await runMcp(["--help"]); + expectCliSuccess(help); + expect(help.stdout).toMatch(/Usage:/i); + expect(help.stderr).not.toContain('"error"'); + }); + + it("covers exitOverride (unknown command) and default process.argv", async () => { + // Non-zero CommanderError goes through exitOverride → throw err. + const unknown = await runMcp(["not-a-command"]); + expectCliFailure(unknown); + + configPath = createSampleTestConfig(); + const originalArgv = process.argv; + process.argv = [ + "node", + "mcpdo", + "servers/list", + "--config", + configPath, + "--format", + "json", + ]; + try { + const { runMcp: invoke } = await import("../src/connection/mcp.js"); + await invoke(); + } finally { + process.argv = originalArgv; + } + }); + + it("servers/show without a name explains the MRU rules instead of a bare parse error", async () => { + configPath = createSampleTestConfig(); + const e = env(); + + // Non-interactive without the default-connection opt-in: name required. + const strict = await runMcp(["servers/show", "--config", configPath], { + env: { ...e, MCP_ALLOW_DEFAULT_CONNECTION: "" }, + timeout: 20000, + }); + expectCliFailure(strict); + expect(strict.stderr).toMatch(/requires an entry name/); + + // Opted in but nothing connected (no daemon): no MRU to infer from. + const noMru = await runMcp(["servers/show", "--config", configPath], { + env: e, + timeout: 20000, + }); + expectCliFailure(noMru); + expect(noMru.stderr).toMatch(/no most-recently-used connection/); + + // Explicit unknown name: keep the core message but say which file was + // searched, so a cross-catalog mismatch is self-explanatory. + const unknown = await runMcp( + ["servers/show", "no-such-entry", "--config", configPath], + { env: e, timeout: 20000 }, + ); + expectCliFailure(unknown); + expect(unknown.stderr).toMatch(/Server 'no-such-entry' not found/); + expect(unknown.stderr).toContain(`(config ${configPath}`); + }); + + it("connections/list and daemon status do not auto-spawn the daemon", async () => { + const e = env(); + const listed = await runMcp(["connections/list", "--format", "json"], { + env: e, + }); + expectCliSuccess(listed); + expect(JSON.parse(listed.stdout)).toEqual({ connections: [] }); + + const status = await runMcp(["daemon", "status", "--format", "json"], { + env: e, + }); + expectCliSuccess(status); + expect(JSON.parse(status.stdout)).toMatchObject({ + running: false, + message: "Daemon is not running.", + }); + + // Socket must not have been created by status/list. + expect(fs.existsSync(path.join(storageDir!, "mcpdod.sock"))).toBe(false); + }); +}); diff --git a/clients/mcpdo/__tests__/mcp-elicitation.test.ts b/clients/mcpdo/__tests__/mcp-elicitation.test.ts new file mode 100644 index 0000000000..ac4bb9c128 --- /dev/null +++ b/clients/mcpdo/__tests__/mcp-elicitation.test.ts @@ -0,0 +1,187 @@ +import { describe, it, expect, afterEach, beforeAll } from "vitest"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; +import { fileURLToPath } from "node:url"; +import { runMcp } from "./helpers/mcp-runner.js"; +import { + expectCliSuccess, + expectCliFailure, +} from "../../cli/__tests__/helpers/assertions.js"; +import { resolveDaemonScriptPath } from "../src/daemon/ensure.js"; +import { callDaemon } from "../src/daemon/client.js"; +import type { ElicitationPendingInfo } from "../src/daemon/protocol.js"; + +/** + * End-to-end non-interactive elicitation: a real daemon, a real composable + * test server (stdio) whose `collect_elicitation` tool sends a legacy + * `elicitation/create` mid-call, a non-TTY front-end that gets the exchange + * parked (`elicitationPending`), and `elicitation/respond` resuming the call + * to its final result. + */ +describe("mcp non-interactive elicitation (e2e)", () => { + let storageDir: string | undefined; + let configPath: string | undefined; + const ttyDescriptors: Array<{ + stream: NodeJS.ReadStream | NodeJS.WriteStream; + desc: PropertyDescriptor | undefined; + }> = []; + + beforeAll(() => { + expect(fs.existsSync(resolveDaemonScriptPath())).toBe(true); + }); + + afterEach(async () => { + for (const { stream, desc } of ttyDescriptors.splice(0)) { + if (desc) Object.defineProperty(stream, "isTTY", desc); + } + if (storageDir) { + const socketPath = path.join(storageDir, "mcpdod.sock"); + if (fs.existsSync(socketPath)) { + try { + await callDaemon("daemon/stop", {}, { socketPath, timeoutMs: 2000 }); + } catch { + // already stopped + } + const deadline = Date.now() + 2000; + while (fs.existsSync(socketPath) && Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 50)); + } + } + fs.rmSync(storageDir, { recursive: true, force: true }); + storageDir = undefined; + } + if (configPath) { + fs.rmSync(configPath, { force: true }); + configPath = undefined; + } + }); + + /** runMcp is in-process: force the non-TTY (parking) path regardless of + * how vitest itself was launched. */ + function stubNonTty(): void { + for (const stream of [process.stdin, process.stderr] as const) { + ttyDescriptors.push({ + stream, + desc: Object.getOwnPropertyDescriptor(stream, "isTTY"), + }); + Object.defineProperty(stream, "isTTY", { + configurable: true, + value: undefined, + }); + } + } + + function env(): Record<string, string> { + storageDir = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-elicit-e2e-")); + return { + MCP_STORAGE_DIR: storageDir, + MCP_INSPECTOR_DAEMON_DIR: storageDir, + MCP_ALLOW_DEFAULT_CONNECTION: "1", + }; + } + + function elicitServerArgs(): string[] { + const here = path.dirname(fileURLToPath(import.meta.url)); + const serverScript = path.resolve( + here, + "../../../test-servers/build/server-composable.js", + ); + expect(fs.existsSync(serverScript)).toBe(true); + configPath = path.join( + os.tmpdir(), + `elicit-server-${process.pid}-${Date.now()}.json`, + ); + fs.writeFileSync( + configPath, + JSON.stringify({ + serverInfo: { name: "elicit-e2e", version: "1.0.0" }, + tools: [{ preset: "collect_elicitation" }], + transport: { type: "stdio" }, + }), + ); + return ["node", serverScript, "--config", configPath]; + } + + it("parks a legacy form elicitation and elicitation/respond resumes to the tool result", async () => { + const e = env(); + stubNonTty(); + + const connected = await runMcp( + [ + "connect", + "--connection", + "el", + "--transport", + "stdio", + "--format", + "json", + "--", + ...elicitServerArgs(), + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(connected); + + const parked = await runMcp( + [ + "tools/call", + "collect_elicitation", + "message:=Pick a color", + 'schema:={"type":"object","properties":{"color":{"type":"string"}},"required":["color"]}', + "--format", + "json", + "--connection", + "el", + ], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(parked); + const pending = JSON.parse(parked.stdout) as { + elicitationPending: ElicitationPendingInfo; + }; + expect(pending.elicitationPending).toMatchObject({ + connection: "el", + method: "tools/call", + toolName: "collect_elicitation", + mode: "form", + message: "Pick a color", + }); + const id = pending.elicitationPending.elicitationId; + expect(id).toBeTruthy(); + + // The connection refuses new rpcs while the call is parked. + const blocked = await runMcp( + ["tools/list", "--format", "json", "--connection", "el"], + { env: e, timeout: 20000 }, + ); + expectCliFailure(blocked); + expect(blocked.output).toContain(`elicitation/respond ${id}`); + + const done = await runMcp( + ["elicitation/respond", id, "color:=teal", "--format", "json"], + { env: e, timeout: 20000 }, + ); + expectCliSuccess(done); + expect(done.stdout).toContain("accept"); + expect(done.stdout).toContain("teal"); + }, 40000); + + it("validates flag exclusivity before contacting the daemon", async () => { + const e = env(); + stubNonTty(); + const conflicting = await runMcp( + ["elicitation/respond", "e-1", "--done", "--cancel"], + { env: e, timeout: 10000 }, + ); + expectCliFailure(conflicting); + expect(conflicting.output).toContain("exactly one of"); + + const empty = await runMcp(["elicitation/respond", "e-1"], { + env: e, + timeout: 10000, + }); + expectCliFailure(empty); + expect(empty.output).toContain("key:=value"); + }); +}); diff --git a/clients/mcpdo/__tests__/parse-tool-args.test.ts b/clients/mcpdo/__tests__/parse-tool-args.test.ts new file mode 100644 index 0000000000..8885c53897 --- /dev/null +++ b/clients/mcpdo/__tests__/parse-tool-args.test.ts @@ -0,0 +1,157 @@ +import { describe, it, expect } from "vitest"; +import { + parseToolCallPositionals, + resolveToolCallArgs, +} from "../src/connection/parse-tool-args.js"; + +describe("parseToolCallPositionals", () => { + it("parses key:=value with JSON typing", () => { + expect( + parseToolCallPositionals([ + "message:=Foo", + "count:=10", + "enabled:=true", + 'cfg:={"a":1}', + 'id:="012"', + ]), + ).toEqual({ + message: "Foo", + count: 10, + enabled: true, + cfg: { a: 1 }, + id: "012", + }); + }); + + it('keeps a literal "__proto__" key as an own property, matching the inline-JSON path', () => { + // On a plain {} accumulator this key would hit the prototype setter and + // vanish while remapping the accumulator's prototype. + const out = parseToolCallPositionals(['__proto__:={"polluted":true}']); + expect(Object.getOwnPropertyNames(out)).toContain("__proto__"); + expect(Object.getOwnPropertyDescriptor(out, "__proto__")?.value).toEqual({ + polluted: true, + }); + expect(({} as Record<string, unknown>).polluted).toBeUndefined(); + }); + + it("parses a single inline JSON object", () => { + expect(parseToolCallPositionals(['{"message":"Foo","count":2}'])).toEqual({ + message: "Foo", + count: 2, + }); + }); + + it("rejects bare values, arrays, and mixed JSON+pairs", () => { + expect(() => parseToolCallPositionals(["foo"])).toThrow(/key:=value/); + expect(() => parseToolCallPositionals(["[1]"])).toThrow(/JSON object/); + expect(() => parseToolCallPositionals(["{not-json"])).toThrow( + /Invalid JSON/, + ); + expect(() => parseToolCallPositionals(['{"a":1}', "b:=2"])).toThrow( + /only one argument/, + ); + expect(() => parseToolCallPositionals([":=x"])).toThrow(/missing key/); + expect(parseToolCallPositionals([])).toEqual({}); + }); +}); + +describe("resolveToolCallArgs", () => { + it("uses positionals as the default style", () => { + expect( + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsPos: ["message:=hi"], + }), + ).toEqual({ toolName: "echo", toolArg: { message: "hi" } }); + }); + + it("treats the name slot as an arg when --tool-name is set", () => { + expect( + resolveToolCallArgs({ + toolNameFlag: "echo", + toolNamePos: "message:=hi", + }), + ).toEqual({ toolName: "echo", toolArg: { message: "hi" } }); + }); + + it("keeps --tool-arg and --tool-args-json as alternatives", () => { + expect( + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgFlag: { message: "via-flag" }, + }), + ).toEqual({ toolName: "echo", toolArg: { message: "via-flag" } }); + + expect( + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsJson: '{"message":"json"}', + }), + ).toEqual({ toolName: "echo", toolArg: { message: "json" } }); + }); + + it("rejects mixing argument styles", () => { + expect(() => + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsPos: ["message:=a"], + toolArgFlag: { message: "b" }, + }), + ).toThrow(/one style/); + expect(() => + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsPos: ["message:=a"], + toolArgsJson: '{"message":"b"}', + }), + ).toThrow(/one style/); + }); + + it("rejects non-finite numbers that JSON cannot represent", () => { + // 1e999 parses to Infinity; NDJSON serialization would send null. + expect(() => parseToolCallPositionals(["count:=1e999"])).toThrow( + /no JSON representation/, + ); + expect(() => parseToolCallPositionals(["count:=-1e999"])).toThrow( + /no JSON representation/, + ); + expect(() => parseToolCallPositionals(['{"count":1e999}'])).toThrow( + /no JSON representation/, + ); + // Nested values are validated recursively. + expect(() => parseToolCallPositionals(['{"a":{"b":[1,2,1e999]}}'])).toThrow( + /no JSON representation/, + ); + expect(() => + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsJson: '{"count":1e999}', + }), + ).toThrow(/no JSON representation/); + // Large-but-finite numbers still round-trip and are accepted. + expect(parseToolCallPositionals(["count:=1e308"])).toEqual({ + count: 1e308, + }); + }); + + it("rejects invalid --tool-args-json", () => { + expect(() => + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsJson: "{bad", + }), + ).toThrow(/not valid JSON/); + expect(() => + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsJson: "[]", + }), + ).toThrow(/must be a JSON object/); + expect(() => + resolveToolCallArgs({ + toolNamePos: "echo", + toolArgsJson: "null", + }), + ).toThrow(/must be a JSON object/); + }); +}); diff --git a/clients/mcpdo/__tests__/pin-stdio-config.test.ts b/clients/mcpdo/__tests__/pin-stdio-config.test.ts new file mode 100644 index 0000000000..b4badadce6 --- /dev/null +++ b/clients/mcpdo/__tests__/pin-stdio-config.test.ts @@ -0,0 +1,57 @@ +import { describe, it, expect } from "vitest"; +import * as path from "node:path"; +import { getDefaultEnvironment } from "@modelcontextprotocol/client/stdio"; +import type { StdioServerConfig } from "@inspector/core/mcp/types.js"; +import { pinStdioConfigToCaller } from "../src/connection/mcp.js"; + +type StdioConfig = StdioServerConfig; + +describe("pinStdioConfigToCaller", () => { + it("returns non-stdio configs unchanged", () => { + const config = { + type: "streamable-http", + url: "https://example.com/mcp", + } as const; + expect(pinStdioConfigToCaller(config)).toBe(config); + }); + + it("pins a missing cwd to the caller's cwd and resolves a relative one", () => { + const base: StdioConfig = { type: "stdio", command: process.execPath }; + expect(pinStdioConfigToCaller({ ...base }).cwd).toBe(process.cwd()); + expect(pinStdioConfigToCaller({ ...base, cwd: "sub/dir" }).cwd).toBe( + path.resolve("sub/dir"), + ); + }); + + it("treats a config without an explicit type as stdio", () => { + const pinned = pinStdioConfigToCaller<StdioConfig>({ + command: process.execPath, + }); + expect(pinned.cwd).toBe(process.cwd()); + }); + + it("snapshots the caller's default environment under the configured env", () => { + const pinned = pinStdioConfigToCaller<StdioConfig>({ + type: "stdio", + command: process.execPath, + env: { PATH: "/configured/bin", EXTRA: "1" }, + }); + const defaults: Record<string, string> = getDefaultEnvironment(); + // Configured values win over the snapshot... + expect(pinned.env).toMatchObject({ PATH: "/configured/bin", EXTRA: "1" }); + // ...and every other default-inherited var is filled from THIS process, + // so the daemon's SDK transport never falls back to its own stale env. + for (const [key, value] of Object.entries(defaults)) { + if (key === "PATH") continue; + expect(pinned.env?.[key]).toBe(value); + } + }); + + it("resolves a bare command against the caller's PATH", () => { + const pinned = pinStdioConfigToCaller<StdioConfig>({ + type: "stdio", + command: "node", + }); + expect(path.isAbsolute(pinned.command)).toBe(true); + }); +}); diff --git a/clients/mcpdo/__tests__/prompt-reader.test.ts b/clients/mcpdo/__tests__/prompt-reader.test.ts new file mode 100644 index 0000000000..48145f8f30 --- /dev/null +++ b/clients/mcpdo/__tests__/prompt-reader.test.ts @@ -0,0 +1,154 @@ +import { describe, it, expect, afterEach } from "vitest"; +import { PassThrough } from "node:stream"; +import { + PromptReader, + getSharedPromptReader, + resetSharedPromptReader, +} from "../src/connection/prompt-reader.js"; + +/** + * Covers the persistent line-queue reader that backs interactive prompts: + * piped input buffered ahead of the questions (the `printf 'a\nb\n' | + * mcpdo tools/call …` case) must answer every question, and "close" must + * mean input *exhausted* (EOF and empty queue), not merely EOF. + */ +describe("PromptReader", () => { + let reader: PromptReader | undefined; + + afterEach(() => { + reader?.dispose(); + reader = undefined; + resetSharedPromptReader(); + }); + + function make(): { + reader: PromptReader; + input: PassThrough; + output: () => string; + } { + const input = new PassThrough(); + const out = new PassThrough(); + let written = ""; + out.on("data", (chunk) => { + written += String(chunk); + }); + reader = new PromptReader( + input as unknown as NodeJS.ReadStream, + out as unknown as NodeJS.WriteStream, + ); + return { reader, input, output: () => written }; + } + + it("answers sequential questions from one up-front piped burst", async () => { + const { reader, input, output } = make(); + // All three answers arrive before any question is asked — the exact + // shape of `printf 'alice\n30\n\n' | mcpdo tools/call register`. + input.write("alice\n30\n\n"); + await expect(reader.question("Name: ")).resolves.toBe("alice"); + await expect(reader.question("Age: ")).resolves.toBe("30"); + await expect(reader.question("Color: ")).resolves.toBe(""); + expect(output()).toBe("Name: Age: Color: "); + }); + + it("resolves a pending question when its line arrives later", async () => { + const { reader, input } = make(); + const pending = reader.question("Name: "); + input.write("octocat\n"); + await expect(pending).resolves.toBe("octocat"); + }); + + it("still answers from the queue after EOF, then reports exhaustion", async () => { + const { reader, input } = make(); + let closed = 0; + reader.once("close", () => { + closed += 1; + }); + input.end("alice\n"); + // Give readline a beat to flush the final chunk and see EOF. + await expect(reader.question("Name: ")).resolves.toBe("alice"); + // EOF alone must not have fired "close" while an answer was queued. + expect(closed).toBe(0); + await expect(reader.question("Age: ")).rejects.toThrow(/stdin closed/); + expect(closed).toBe(1); + // A listener registered after exhaustion fires immediately. + reader.once("close", () => { + closed += 1; + }); + expect(closed).toBe(2); + // And later questions keep rejecting without hanging. + await expect(reader.question("More: ")).rejects.toThrow(/stdin closed/); + }); + + it("rejects a question pending at EOF and notifies close watchers", async () => { + const { reader, input } = make(); + let closed = false; + reader.once("close", () => { + closed = true; + }); + const pending = reader.question("Name: "); + input.end(); + await expect(pending).rejects.toThrow(/stdin closed/); + expect(closed).toBe(true); + }); + + it("shares one process-wide reader, and reset disposes it", () => { + const first = getSharedPromptReader(); + expect(getSharedPromptReader()).toBe(first); + resetSharedPromptReader(); + const second = getSharedPromptReader(); + expect(second).not.toBe(first); + }); + + describe("TTY cooked/raw mode", () => { + function makeTty(): { + reader: PromptReader; + input: PassThrough; + rawModes: boolean[]; + } { + const input = new PassThrough() as PassThrough & { + isTTY?: boolean; + setRawMode?: (mode: boolean) => void; + }; + const rawModes: boolean[] = []; + input.isTTY = true; + input.setRawMode = (mode: boolean) => { + rawModes.push(mode); + }; + // A non-TTY output keeps readline's own raw-mode toggling out of the + // picture, so the recorded calls are exactly the reader's park/engage. + const out = new PassThrough(); + reader = new PromptReader( + input as unknown as NodeJS.ReadStream, + out as unknown as NodeJS.WriteStream, + ); + return { reader, input, rawModes }; + } + + it("leaves a TTY in cooked mode while parked, raw only while a question is pending", async () => { + const { reader, input, rawModes } = makeTty(); + // The constructor parks immediately: the terminal must start cooked so + // Ctrl-C (SIGINT) works before the first prompt. + expect(rawModes.at(-1)).toBe(false); + + const pending = reader.question("Name: "); + // Engaging a pending question re-enters raw mode for line editing. + expect(rawModes.at(-1)).toBe(true); + + input.write("octocat\n"); + await expect(pending).resolves.toBe("octocat"); + // Answering parks again — back to cooked mode between rounds. + expect(rawModes.at(-1)).toBe(false); + }); + + it("restores cooked mode when disposed", () => { + const { reader, rawModes } = makeTty(); + // Dangling on purpose: dispose() rejects it via close — swallow so it + // doesn't surface as an unhandled rejection. + reader.question("Name: ").catch(() => {}); + expect(rawModes.at(-1)).toBe(true); + reader.dispose(); + // dispose() closes the interface; the terminal must not be left raw. + expect(rawModes.at(-1)).toBe(false); + }); + }); +}); diff --git a/clients/mcpdo/__tests__/resolve-command.test.ts b/clients/mcpdo/__tests__/resolve-command.test.ts new file mode 100644 index 0000000000..1b9b707c66 --- /dev/null +++ b/clients/mcpdo/__tests__/resolve-command.test.ts @@ -0,0 +1,81 @@ +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { afterAll, describe, expect, it } from "vitest"; +import { resolveCommandPath } from "../src/connection/resolve-command.js"; + +const tmpRoot = fs.mkdtempSync(path.join(os.tmpdir(), "mcp-conn-resolve-")); + +afterAll(() => { + fs.rmSync(tmpRoot, { recursive: true, force: true }); +}); + +function makeExecutable(dir: string, name: string): string { + fs.mkdirSync(dir, { recursive: true }); + const file = path.join(dir, name); + fs.writeFileSync(file, "#!/bin/sh\n", { mode: 0o755 }); + return file; +} + +describe("resolveCommandPath", () => { + it("resolves a bare name to the first executable on PATH", () => { + const first = path.join(tmpRoot, "first"); + const second = path.join(tmpRoot, "second"); + const expected = makeExecutable(first, "mytool"); + makeExecutable(second, "mytool"); + const env = { PATH: [first, second].join(path.delimiter) }; + expect(resolveCommandPath("mytool", env)).toBe(expected); + }); + + it("skips PATH entries where the name is missing or not a file", () => { + const missing = path.join(tmpRoot, "missing"); + const hasDir = path.join(tmpRoot, "has-dir"); + fs.mkdirSync(path.join(hasDir, "mytool2"), { recursive: true }); + const real = path.join(tmpRoot, "real"); + const expected = makeExecutable(real, "mytool2"); + const env = { PATH: [missing, hasDir, "", real].join(path.delimiter) }; + expect(resolveCommandPath("mytool2", env)).toBe(expected); + }); + + it("skips non-executable files", () => { + const dir = path.join(tmpRoot, "non-exec"); + fs.mkdirSync(dir, { recursive: true }); + fs.writeFileSync(path.join(dir, "mytool3"), "", { mode: 0o644 }); + const real = path.join(tmpRoot, "exec"); + const expected = makeExecutable(real, "mytool3"); + const env = { PATH: [dir, real].join(path.delimiter) }; + expect(resolveCommandPath("mytool3", env)).toBe(expected); + }); + + it("returns commands with a path separator unchanged", () => { + expect(resolveCommandPath("./server.js", { PATH: tmpRoot })).toBe( + "./server.js", + ); + expect(resolveCommandPath("/usr/bin/env", { PATH: tmpRoot })).toBe( + "/usr/bin/env", + ); + }); + + it("returns the name unchanged when not found on PATH", () => { + const env = { PATH: path.join(tmpRoot, "empty-dir") }; + expect(resolveCommandPath("definitely-not-a-real-tool", env)).toBe( + "definitely-not-a-real-tool", + ); + }); + + it("handles an empty command and an unset PATH", () => { + expect(resolveCommandPath("", { PATH: tmpRoot })).toBe(""); + expect(resolveCommandPath("mytool", {})).toBe("mytool"); + }); + + it("defaults to process.env", () => { + // `sh` exists on every POSIX PATH; on Windows this still exercises the + // default-env branch even if the lookup misses. + const resolved = resolveCommandPath("sh"); + if (process.platform !== "win32") { + expect(path.isAbsolute(resolved)).toBe(true); + } else { + expect(typeof resolved).toBe("string"); + } + }); +}); diff --git a/clients/mcpdo/__tests__/sanitize.test.ts b/clients/mcpdo/__tests__/sanitize.test.ts new file mode 100644 index 0000000000..0ecdb24b13 --- /dev/null +++ b/clients/mcpdo/__tests__/sanitize.test.ts @@ -0,0 +1,119 @@ +/** + * Terminal-escape sanitization (security). Server-controlled strings must + * never reach the terminal as raw control bytes — see src/connection/sanitize.ts + * for the threat catalogue (OSC 52 clipboard writes, title spoofing, CSI + * rewriting, OSC 8 hyperlink breakout). + */ +import { describe, expect, it } from "vitest"; +import { + isSafeLinkTarget, + sanitizeDeep, + sanitizeText, +} from "../src/connection/sanitize.js"; + +describe("sanitizeText", () => { + it("neutralizes an OSC 52 clipboard-write sequence", () => { + const attack = "\u001b]52;c;bWFsaWNpb3Vz\u0007done"; + const out = sanitizeText(attack); + expect(out).not.toContain("\u001b"); + expect(out).not.toContain("\u0007"); + expect(out).toBe("\u241b]52;c;bWFsaWNpb3Vz\u2407done"); + }); + + it("neutralizes OSC title spoofing and CSI cursor rewriting", () => { + expect(sanitizeText("\u001b]0;fake title\u0007")).toBe( + "\u241b]0;fake title\u2407", + ); + expect(sanitizeText("\u001b[2J\u001b[H")).toBe("\u241b[2J\u241b[H"); + }); + + it("neutralizes a BEL/ESC breakout inside a URI (OSC 8 wrapper safety)", () => { + const uri = "https://ok.test/\u0007\u001b]8;;https://evil.test\u0007"; + const out = sanitizeText(uri); + expect(out.includes("\u0007")).toBe(false); + expect(out.includes("\u001b")).toBe(false); + }); + + it("replaces C1 controls (8-bit CSI/OSC) with visible text", () => { + expect(sanitizeText("\u009b31mred")).toBe("\\u{9b}31mred"); + expect(sanitizeText("\u009d0;t\u009c")).toBe("\\u{9d}0;t\\u{9c}"); + }); + + it("replaces DEL and CR but preserves newline and tab", () => { + expect(sanitizeText("a\u007fb\rc")).toBe("a\u2421b\u240dc"); + expect(sanitizeText("line1\nline2\tend")).toBe("line1\nline2\tend"); + }); + + it("leaves ordinary text (including non-ASCII) untouched", () => { + const s = "hello — ünïcode ✅ 日本語"; + expect(sanitizeText(s)).toBe(s); + }); +}); + +describe("sanitizeDeep", () => { + it("sanitizes nested string values, array items, and object keys", () => { + const input = { + name: "tool\u001b[1m", + items: ["ok", "bad\u0007"], + nested: { "\u001bkey": { deep: "\u009btext" } }, + }; + expect(sanitizeDeep(input)).toEqual({ + name: "tool\u241b[1m", + items: ["ok", "bad\u2407"], + nested: { "\u241bkey": { deep: "\\u{9b}text" } }, + }); + }); + + it("passes non-string primitives and null through unchanged", () => { + expect(sanitizeDeep(42)).toBe(42); + expect(sanitizeDeep(true)).toBe(true); + expect(sanitizeDeep(null)).toBe(null); + expect(sanitizeDeep(undefined)).toBe(undefined); + }); + + it("does not mutate the input object", () => { + const input = { text: "esc\u001b" }; + const out = sanitizeDeep(input); + expect(input.text).toBe("esc\u001b"); + expect(out.text).toBe("esc\u241b"); + }); + + it('preserves a literal "__proto__" key instead of dropping it', () => { + // On a plain {} accumulator, assigning "__proto__" hits the prototype + // setter and silently discards the entry; the null-prototype result + // keeps it as an ordinary own property. + const input = JSON.parse( + '{"__proto__": {"polluted": "esc\\u001b"}, "a": 1}', + ); + const out = sanitizeDeep(input) as Record<string, unknown>; + expect(Object.getOwnPropertyNames(out)).toContain("__proto__"); + expect( + ( + Object.getOwnPropertyDescriptor(out, "__proto__")?.value as Record< + string, + unknown + > + ).polluted, + ).toBe("esc\u241b"); + expect(out.a).toBe(1); + // No pollution of shared prototypes either. + expect(({} as Record<string, unknown>).polluted).toBeUndefined(); + }); +}); + +describe("isSafeLinkTarget", () => { + it("allows only http(s) URLs as OSC 8 link targets", () => { + expect(isSafeLinkTarget("https://example.com/x")).toBe(true); + expect(isSafeLinkTarget("http://localhost:3001/mcp")).toBe(true); + expect(isSafeLinkTarget("file:///etc/passwd")).toBe(false); + expect(isSafeLinkTarget("javascript:alert(1)")).toBe(false); + expect(isSafeLinkTarget("vscode://malicious/payload")).toBe(false); + expect(isSafeLinkTarget("customproto://x")).toBe(false); + }); + + it("rejects strings that don't parse as URLs", () => { + expect(isSafeLinkTarget("not a url")).toBe(false); + expect(isSafeLinkTarget("")).toBe(false); + expect(isSafeLinkTarget("example.com/no-scheme")).toBe(false); + }); +}); diff --git a/clients/mcpdo/eslint.config.js b/clients/mcpdo/eslint.config.js new file mode 100644 index 0000000000..ff42ec3224 --- /dev/null +++ b/clients/mcpdo/eslint.config.js @@ -0,0 +1,34 @@ +import js from "@eslint/js"; +import globals from "globals"; +import tseslint from "typescript-eslint"; +import { defineConfig, globalIgnores } from "eslint/config"; + +export default defineConfig([ + globalIgnores(["build", "coverage"]), + { + files: ["**/*.ts"], + extends: [js.configs.recommended, tseslint.configs.recommended], + languageOptions: { + ecmaVersion: 2022, + sourceType: "module", + globals: globals.node, + }, + }, + { + // Type-aware pass for `no-floating-promises` (#1959), mirroring + // clients/cli: the rule needs type information, and the parser needs a + // project that literally contains the linted file — so both of this + // client's tsconfig projects are listed, exactly as `npm run typecheck` + // runs them (`src` is in the first, `__tests__` only in the second). + files: ["**/*.ts"], + languageOptions: { + parserOptions: { + project: ["./tsconfig.json", "./tsconfig.test.json"], + tsconfigRootDir: import.meta.dirname, + }, + }, + rules: { + "@typescript-eslint/no-floating-promises": "error", + }, + }, +]); diff --git a/clients/mcpdo/evals/evals.json b/clients/mcpdo/evals/evals.json new file mode 100644 index 0000000000..04b0e89100 --- /dev/null +++ b/clients/mcpdo/evals/evals.json @@ -0,0 +1,286 @@ +[ + { + "kind": "trigger", + "prompt": "What MCP servers am I connected to?", + "expect": "mcpdo" + }, + { + "kind": "trigger", + "prompt": "Call the echo tool on the everything MCP server and show me the result.", + "expect": "mcpdo" + }, + { + "kind": "trigger", + "prompt": "Do I have access to any MCP tools besides your built-in ones?", + "expect": "mcpdo" + }, + { + "kind": "trigger", + "prompt": "List the resources available on the MCP server I connected to earlier.", + "expect": "mcpdo" + }, + { + "kind": "trigger", + "prompt": "Use mcpdo to list the tools on the everything server.", + "expect": "mcpdo" + }, + { + "kind": "trigger", + "prompt": "What does this regex do? /^\\d{3}-\\d{4}$/", + "expect": null + }, + { + "kind": "trigger", + "prompt": "How do I write a simple HTTP server in Node.js?", + "expect": null + }, + { + "kind": "trigger", + "prompt": "Rename the variable `foo` to `bar` in this snippet: const foo = 1; console.log(foo);", + "expect": null + }, + { + "kind": "behavior", + "prompt": "Connect to the test-stdio MCP server and list its tools.", + "expectCalls": [ + { + "cmd": "connect", + "connection": "test-stdio" + }, + { + "cmd": "tools/list", + "connection": "test-stdio" + } + ] + }, + { + "kind": "behavior", + "prompt": "Use the test-stdio MCP server to add the numbers 2 and 3, and tell me the result.", + "expectCalls": [ + { + "cmd": "tools/call", + "connection": "test-stdio", + "tool": "get_sum", + "args": { + "a": 2, + "b": 3 + }, + "stdoutMatch": "5" + } + ] + }, + { + "kind": "behavior", + "prompt": "Connect to the secure-add MCP server.", + "servers": { + "secure-add": { + "serverInfo": { + "name": "secure-add", + "version": "1.0.0" + }, + "tools": [ + { + "preset": "add" + } + ], + "oauth": { + "enabled": true, + "mode": "combined", + "requireAuth": true, + "scopesSupported": ["mcp"], + "supportDCR": true + }, + "transport": { + "type": "streamable-http" + } + } + }, + "expectCalls": [ + { + "cmd": "connect", + "connection": "secure-add", + "exit": 0, + "stdoutMatch": "oauth/authorize" + } + ], + "expectReply": "https?://[^\\s\"'<>\\\\]+/oauth/authorize\\?", + "expectLastCallTurn1": "connect" + }, + { + "kind": "behavior", + "autoConsent": true, + "consentAfterReply": true, + "prompt": "List the tools available on the secure-add MCP server.", + "followUp": "I have signed in now — please continue.", + "servers": { + "secure-add": { + "serverInfo": { + "name": "secure-add", + "version": "1.0.0" + }, + "tools": [ + { + "preset": "add" + } + ], + "oauth": { + "enabled": true, + "mode": "combined", + "requireAuth": true, + "scopesSupported": ["mcp"], + "supportDCR": true + }, + "transport": { + "type": "streamable-http" + } + } + }, + "expectCalls": [ + { + "cmd": "connect", + "connection": "secure-add", + "exit": 0, + "stdoutMatch": "oauth/authorize" + }, + { + "cmd": "tools/list", + "connection": "secure-add", + "exit": 0, + "stdoutMatch": "add" + } + ], + "expectReply": "https?://[^\\s\"'<>\\\\]+/oauth/authorize\\?", + "rejectCompletedTurn1": "tools/list", + "expectReplyFollowUp": "add" + }, + { + "kind": "behavior", + "prompt": "File a ticket on the helpdesk server with the summary 'Printer on floor 3 is jammed'. If it needs contact details, I'm Ada Lovelace, ada@example.com. Tell me the ticket number.", + "servers": { + "helpdesk": { + "serverInfo": { + "name": "helpdesk", + "version": "1.0.0" + }, + "tools": [ + { + "preset": "submit_ticket" + } + ] + } + }, + "expectCalls": [ + { + "cmd": "tools/call", + "connection": "helpdesk", + "tool": "submit_ticket", + "args": { + "summary": "Printer on floor 3 is jammed" + }, + "exit": 0, + "stdoutMatch": "elicitationPending|Input required" + }, + { + "cmd": "elicitation/respond", + "args": { + "contact_name": "Ada Lovelace", + "contact_email": "ada@example.com" + }, + "exit": 0, + "stdoutMatch": "TCK-" + } + ] + }, + { + "kind": "behavior", + "prompt": "What MCP servers am I connected to right now?", + "expectCalls": [ + { + "cmd": "connections/list", + "exit": 0 + } + ] + }, + { + "kind": "behavior", + "prompt": "What MCP servers are available for me to connect to? Don't connect to anything, just tell me what's there.", + "expectCalls": [ + { + "cmd": "servers/list", + "exit": 0, + "stdoutMatch": "test-stdio" + } + ] + }, + { + "kind": "behavior", + "prompt": "My catalog has two MCP servers, alpha and beta. Use beta to compute the sum of 4 and 9.", + "servers": { + "alpha": { + "serverInfo": { + "name": "alpha", + "version": "1.0.0" + }, + "tools": [ + { + "preset": "echo" + } + ] + }, + "beta": { + "serverInfo": { + "name": "beta", + "version": "1.0.0" + }, + "tools": [ + { + "preset": "get_sum" + } + ] + } + }, + "expectCalls": [ + { + "cmd": "tools/call", + "connection": "beta", + "tool": "get_sum", + "args": { + "a": 4, + "b": 9 + }, + "stdoutMatch": "13" + } + ] + }, + { + "kind": "behavior", + "prompt": "File a ticket on the helpdesk server with the summary 'Broken keyboard'. If it asks for contact details, do NOT share any personal information — decline that request — and tell me what happened.", + "servers": { + "helpdesk": { + "serverInfo": { + "name": "helpdesk", + "version": "1.0.0" + }, + "tools": [ + { + "preset": "submit_ticket" + } + ] + } + }, + "expectCalls": [ + { + "cmd": "tools/call", + "connection": "helpdesk", + "tool": "submit_ticket", + "exit": 0, + "stdoutMatch": "elicitationPending|Input required" + }, + { + "cmd": "elicitation/respond", + "exit": 0, + "stdoutMatch": "not filed|declined" + } + ] + } +] diff --git a/clients/mcpdo/package-lock.json b/clients/mcpdo/package-lock.json new file mode 100644 index 0000000000..1d41f9f910 --- /dev/null +++ b/clients/mcpdo/package-lock.json @@ -0,0 +1,1583 @@ +{ + "name": "@modelcontextprotocol/mcpdo", + "lockfileVersion": 3, + "requires": true, + "packages": { + "": { + "name": "@modelcontextprotocol/mcpdo", + "license": "MIT", + "bin": { + "mcpdo": "build/mcp-bin.js" + }, + "devDependencies": { + "@types/express": "^5.0.6", + "tsup": "^8.5.0" + } + }, + "node_modules/@esbuild/aix-ppc64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/aix-ppc64/-/aix-ppc64-0.28.2.tgz", + "integrity": "sha512-XExcO+dvLKvVtNTibSTBej1NCAbaGhWn9Ww1ZPx80qsahhPFe/8jgWP0IchNe0F3HwkU7n8ejhH8bjonqht8mQ==", + "cpu": [ + "ppc64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "aix" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/android-arm": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/android-arm/-/android-arm-0.28.2.tgz", + "integrity": "sha512-kXXoiPVVGQcnIYGOeaovwOURpniDBpSq4A03qkQ+BMQqtGG6HYap3xne9C1O1yo4TR3qxlCX5IqqmX6fFo2Lqg==", + "cpu": [ + "arm" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "android" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/android-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/android-arm64/-/android-arm64-0.28.2.tgz", + "integrity": "sha512-5YfKeeI8qWfBZIX+u2xZC3Zlb3Os/gLS2sbEKM+I4ZOcsWmHS2WLysCcQZDAFRslDUU5Oiq44gf6PYN1vGwG5A==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "android" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/android-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/android-x64/-/android-x64-0.28.2.tgz", + "integrity": "sha512-O387ite7SzUyCcy3JQX4P4bLtEA7bLLkx+esve5JHnyYfNTxcVpXZo9jhdB0lTKN44gztELTdU7nS8Nr16Fs1Q==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "android" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/darwin-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/darwin-arm64/-/darwin-arm64-0.28.2.tgz", + "integrity": "sha512-n4KqkOQrraxHJcgjM1RvwbigfQKIKJVpM7xp+KsxiyUSrRdIXnt73VhrPAx0fV44hgfmIVKjxMN9J1t5jySVkw==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "darwin" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/darwin-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/darwin-x64/-/darwin-x64-0.28.2.tgz", + "integrity": "sha512-uq6suIWYP37qzGddBKPw5QEQPi6HiLGsO7UmkpfyaYNQ3D+rN6w6WfwH+nuqcGXWvawGwxOEroO4YGnFh95azw==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "darwin" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/freebsd-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/freebsd-arm64/-/freebsd-arm64-0.28.2.tgz", + "integrity": "sha512-n+I0BTSRIoy+d6RPKnEVwql5UwBJolytvY4mAOIEJorKlqgPII8ix6slVVrfZ5Tnj7glIZvloylbB/EJPMWEXw==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "freebsd" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/freebsd-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/freebsd-x64/-/freebsd-x64-0.28.2.tgz", + "integrity": "sha512-78XJTJkvPs0kz2w61301PJjXl4g7q3JqiYMZ/M/yVI73EHBrCRTgkhu9oqG7vPqq+a/yadEW8aD+agKlk5xrmg==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "freebsd" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-arm": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-arm/-/linux-arm-0.28.2.tgz", + "integrity": "sha512-XlDnu2q5yoqems+xay6wSAcg9DDD7K9RLKZEBOMZm3ckNpJBvOX20tSfby8KfrrhINDyv9V2YVZKY/SpoGJI8w==", + "cpu": [ + "arm" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-arm64/-/linux-arm64-0.28.2.tgz", + "integrity": "sha512-pW4AC0P3it8c7do9MVM4p51FzHzdM/TZrerurgRcHJ2WTa1VQ1CIq18xncfpBJw4ojkiZZrKW2yIBWBP92j6Ug==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-ia32": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-ia32/-/linux-ia32-0.28.2.tgz", + "integrity": "sha512-CYbnj78HsIeA+DhgUKgFCfvNsTHFhMMrinUrMZpDXJXKN8T3XViTZ/+wtHeVxEWY8ewSzTFN+nRmSwO2tZaLUQ==", + "cpu": [ + "ia32" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-loong64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-loong64/-/linux-loong64-0.28.2.tgz", + "integrity": "sha512-buwkd8nsph4R+ajRvw0qM5Hja/TXQow3ptzWO2EbG/cqcIkHloRrdlBtQlshyYGTNFvfkfJ5tpPLVkY4DtsPfQ==", + "cpu": [ + "loong64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-mips64el": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-mips64el/-/linux-mips64el-0.28.2.tgz", + "integrity": "sha512-ZVykbDyk7519VwiNb9Lcj9m8XM6v5V9uKPvrEMkkEedVewf+0itkhahp4HDpgERXhwLRpWFypsGbG/J8s0QjJA==", + "cpu": [ + "mips64el" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-ppc64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-ppc64/-/linux-ppc64-0.28.2.tgz", + "integrity": "sha512-CAXl+Dtd9UUuJd8pKKdwh6MLm3MUMiqMPmhZ3tTSXPqfyQ3vDl6R5hZdZ/kYojK4ofXtdfSv1tFq8XzWx3heNQ==", + "cpu": [ + "ppc64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-riscv64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-riscv64/-/linux-riscv64-0.28.2.tgz", + "integrity": "sha512-GeXCej4IQtU1B+QlDV8W/RRvbzI3O/Stss+/bCXv4lZls5WGRtu2a+3JkA3i4qIUlMXpcHebWpF8AkJhATowuA==", + "cpu": [ + "riscv64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-s390x": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-s390x/-/linux-s390x-0.28.2.tgz", + "integrity": "sha512-3H1weTYZPxt/WOhByszQZybS9w5lKzUn1FDMsgEChbHWQwHYQQRfBxgCcZvPhjHfKyJjIievvMmEUawJrdY9Dg==", + "cpu": [ + "s390x" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/linux-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/linux-x64/-/linux-x64-0.28.2.tgz", + "integrity": "sha512-4xTZr1FUmSoQW4XIWmit3tzQrUTZM+N3P0XV8xROKYF50XfI7xeO90+1bZvNwxIufQ9hDQVRJH5YhgPVF8A/HQ==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/netbsd-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/netbsd-arm64/-/netbsd-arm64-0.28.2.tgz", + "integrity": "sha512-sSATRjPeDBg3pdgHoQfoYBob11Kk1FGa9lui5RIHZCoCkJa9QKlvl3/vKz2usCmYYjs7ymJR/2Nnsqe+Hjt5nw==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "netbsd" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/netbsd-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/netbsd-x64/-/netbsd-x64-0.28.2.tgz", + "integrity": "sha512-lqnzCV+mM0gIADaKihiCg6ifgfU2L3h5E33rNQBN1Y4MaVGnzryzmvvf7UHxprpQdE8hpqLolJ9Rl+SkIRDpyw==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "netbsd" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/openbsd-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/openbsd-arm64/-/openbsd-arm64-0.28.2.tgz", + "integrity": "sha512-AL2qJILH7lNjrDmCQDvdxMfAUIv8KMNZOvrwAQ8i8//ntL9FflhOyMJ8OZSMBb8/AWXe3/5v5S20y3zCoZWKoQ==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "openbsd" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/openbsd-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/openbsd-x64/-/openbsd-x64-0.28.2.tgz", + "integrity": "sha512-QtiuPytchRyC4rwUKhexJdQKvDuZ6hWloi3igqPQNUJCS1/v9EiO3UTOXR6A3FoMo4fnAKbWJdqaIwhOzh8qEw==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "openbsd" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/openharmony-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/openharmony-arm64/-/openharmony-arm64-0.28.2.tgz", + "integrity": "sha512-WkhYDmpTjLvGlScA1rwjRUmhl4k8oXR3cIbtqWmELgU/dFeHHlEllxDvdWcNJV9rbzCexB5vz8gtNewWLgCT7Q==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "openharmony" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/sunos-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/sunos-x64/-/sunos-x64-0.28.2.tgz", + "integrity": "sha512-GPMSkTOtMnv2U2F8gxe4Io6qmVs+YKyp832Etqqxr0hFngmXQ3rzwytelm3GIn7T4VviRUlf3sOgBOiTdvaf7g==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "sunos" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/win32-arm64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/win32-arm64/-/win32-arm64-0.28.2.tgz", + "integrity": "sha512-PIhhEkE9uPBleRBrQEJpUn7MBnibZzbGzYWPmY3x+YoVg/95zbjB4CxPPOQ8l5tYYM4mMaCthF8/1DIfBQQyWQ==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "win32" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/win32-ia32": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/win32-ia32/-/win32-ia32-0.28.2.tgz", + "integrity": "sha512-YmJbfTlvU7Sdn9BB+4PRES4oB6pxgS37MAONj+hBr/cpXS1aBPKXxNnDbu+QCWPj0o9dgyxeq79g6c5P8KeuYA==", + "cpu": [ + "ia32" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "win32" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@esbuild/win32-x64": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/@esbuild/win32-x64/-/win32-x64-0.28.2.tgz", + "integrity": "sha512-5ebpxr3nWMzrL/rnUI755Jkuee0bHL/Gq0WTF9lvcpv73wAp5eu8MfBUgWK9bhWvZjj7yX8etf/8tI8Ney695g==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "win32" + ], + "engines": { + "node": ">=18" + } + }, + "node_modules/@jridgewell/gen-mapping": { + "version": "0.3.13", + "resolved": "https://registry.npmjs.org/@jridgewell/gen-mapping/-/gen-mapping-0.3.13.tgz", + "integrity": "sha512-2kkt/7niJ6MgEPxF0bYdQ6etZaA+fQvDcLKckhy1yIQOzaoKjBBjSj63/aLVjYE3qhRt5dvM+uUyfCg6UKCBbA==", + "dev": true, + "license": "MIT", + "dependencies": { + "@jridgewell/sourcemap-codec": "^1.5.0", + "@jridgewell/trace-mapping": "^0.3.24" + } + }, + "node_modules/@jridgewell/resolve-uri": { + "version": "3.1.2", + "resolved": "https://registry.npmjs.org/@jridgewell/resolve-uri/-/resolve-uri-3.1.2.tgz", + "integrity": "sha512-bRISgCIjP20/tbWSPWMEi54QVPRZExkuD9lJL+UIxUKtwVJA8wW1Trb1jMs1RFXo1CBTNZ/5hpC9QvmKWdopKw==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=6.0.0" + } + }, + "node_modules/@jridgewell/sourcemap-codec": { + "version": "1.6.0", + "resolved": "https://registry.npmjs.org/@jridgewell/sourcemap-codec/-/sourcemap-codec-1.6.0.tgz", + "integrity": "sha512-T7jf+5zgsZHwNJ4lvQ7/aezbyk0nNX+zJVWpmHA7VYsEx7a7qr5Rg5IbtJFqkgze5Y2sruq1RUY8Q837Od7iFw==", + "dev": true, + "license": "MIT" + }, + "node_modules/@jridgewell/trace-mapping": { + "version": "0.3.31", + "resolved": "https://registry.npmjs.org/@jridgewell/trace-mapping/-/trace-mapping-0.3.31.tgz", + "integrity": "sha512-zzNR+SdQSDJzc8joaeP8QQoCQr8NuYx2dIIytl1QeBEZHJ9uW6hebsrYgbz8hJwUQao3TWCMtmfV8Nu1twOLAw==", + "dev": true, + "license": "MIT", + "dependencies": { + "@jridgewell/resolve-uri": "^3.1.0", + "@jridgewell/sourcemap-codec": "^1.4.14" + } + }, + "node_modules/@napi-rs/lzma-linux-x64-gnu": { + "version": "1.5.1", + "resolved": "https://registry.npmjs.org/@napi-rs/lzma-linux-x64-gnu/-/lzma-linux-x64-gnu-1.5.1.tgz", + "integrity": "sha512-oTXEIha4SsuXdTA4Iyskj0kpdx2yVXdhd75c2v3xGrHFfVMsbhTPZU/nMPL4sWKo4pBHm3aucLaqGlF696dTyQ==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ], + "engines": { + "node": "^22.20 || ^24.12 || >=25" + } + }, + "node_modules/@rollup/rollup-android-arm-eabi": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-android-arm-eabi/-/rollup-android-arm-eabi-4.63.4.tgz", + "integrity": "sha512-I+BSHzTAhKN2n7ZwGZsegGcZjDpLqFOMAtJz/u6uFGe0pUFbq56dEHjqJV/ZUdRJtNXNxA+hREUatZBvMR3Oiw==", + "cpu": [ + "arm" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "android" + ] + }, + "node_modules/@rollup/rollup-android-arm64": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-android-arm64/-/rollup-android-arm64-4.63.4.tgz", + "integrity": "sha512-pu3BdjS2LtEzRu2elmGzS3fIeWSZy4BMDIaLNwjorO76+k2d0LMluijhsDx3KQyQBQ/lLUZCQA9/s6csvUfuhw==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "android" + ] + }, + "node_modules/@rollup/rollup-darwin-arm64": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-darwin-arm64/-/rollup-darwin-arm64-4.63.4.tgz", + "integrity": "sha512-xfSrj9MHnWK9GaSqT9U0ImHtH/N8WZlHLx4cZHiuLcqs640hvZ3hLPd5UR2AZS57FaE8HrRUSpltbZdWRxHiDA==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "darwin" + ] + }, + "node_modules/@rollup/rollup-darwin-x64": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-darwin-x64/-/rollup-darwin-x64-4.63.4.tgz", + "integrity": "sha512-bqU99PLJb/dqb3S0GIMdeuyAEETSUgZBoqXYd3Sd+WCsV+MmPhnN6JrotWyir31+QgH7EvvE5/mwGJlEoci8Fw==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "darwin" + ] + }, + "node_modules/@rollup/rollup-freebsd-arm64": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-freebsd-arm64/-/rollup-freebsd-arm64-4.63.4.tgz", + "integrity": "sha512-JinsFZ5G40oXQb+sUuiA5x689vhr6dDYK0H0NL+rwKdL6CqnmYN8PE4ZwfRSoIjrCxqTQG/SLfTtSvHeGxoVlw==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "freebsd" + ] + }, + "node_modules/@rollup/rollup-freebsd-x64": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-freebsd-x64/-/rollup-freebsd-x64-4.63.4.tgz", + "integrity": "sha512-GAdA4UxpiNm27cLHr2GqXBpAD0x9FqwYBY7/YSP0Ss0/PNi4k8gbviqpIpYbVSRBaS2ZcegXEzgTQMbRNCwxCw==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "freebsd" + ] + }, + "node_modules/@rollup/rollup-linux-arm-gnueabihf": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-arm-gnueabihf/-/rollup-linux-arm-gnueabihf-4.63.4.tgz", + "integrity": "sha512-qDd6NoA1znaLjp4jR5U/KWCdLAKDJNB8W9ChbbDaKbo0xA+Atln5HK6LFCZ4oJQpemtRZA288DCirFRjrspptw==", + "cpu": [ + "arm" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-arm-musleabihf": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-arm-musleabihf/-/rollup-linux-arm-musleabihf-4.63.4.tgz", + "integrity": "sha512-WtB5Tz5KTNINb8ZA+8sQ7bmjuS1JrRT7YverYIhUGdWWDlpzVWmIwuZE+jidkEXUn1l0zrEkaIMa8dHF3NGcsA==", + "cpu": [ + "arm" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-arm64-gnu": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-arm64-gnu/-/rollup-linux-arm64-gnu-4.63.4.tgz", + "integrity": "sha512-VcQ3L1tjnkKzWjryAVaFhHEWcqOfICX9uxVVoDzm2t0DpgKRHd2zOpVrJc0xsWeBZcBFyYROCIBdyR/fS174pg==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-arm64-musl": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-arm64-musl/-/rollup-linux-arm64-musl-4.63.4.tgz", + "integrity": "sha512-6+ZQX6P5s0cMDN2Ypb8Lbm2+/sZYmZjdaYny992ujUU9UKi/4CWoJWsl1pNvjWJHNHGK51m+jKGLlh1ylb2ifQ==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-loong64-gnu": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-loong64-gnu/-/rollup-linux-loong64-gnu-4.63.4.tgz", + "integrity": "sha512-D72ZnvkFkBXOfzMMQLcwfPLyGkKb7HZ9/mf97B7v6/P5Lbv4oFOtSY/uHbS8lH6uKUOxoKiuokdb50XZSzzbJw==", + "cpu": [ + "loong64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-loong64-musl": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-loong64-musl/-/rollup-linux-loong64-musl-4.63.4.tgz", + "integrity": "sha512-piU6BxeqA3O9KSu3kRCIQQtNqFFaTu21SEV4FwaRZowpnj3bLaWPZHw+xFqCs0XlJ+aOH3PTRWGoglH+mKA/OA==", + "cpu": [ + "loong64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-ppc64-gnu": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-ppc64-gnu/-/rollup-linux-ppc64-gnu-4.63.4.tgz", + "integrity": "sha512-/5PGpHwqt2EEEOUs1XwzubE/ucr0dWDQ+to3zqi4Ds7EWpwtQ79wXc4JBoxqj/OwpawTsKWzJxHfSuBOq3DrWA==", + "cpu": [ + "ppc64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-ppc64-musl": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-ppc64-musl/-/rollup-linux-ppc64-musl-4.63.4.tgz", + "integrity": "sha512-cX3beZDLWt7G2oJF+nhChiT+qtaihs+S2xi7ziGmVB+2pwPng6D0Ed0HmElQOgv2UsUmSJJLGwpBao/3TDx3VA==", + "cpu": [ + "ppc64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-riscv64-gnu": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-riscv64-gnu/-/rollup-linux-riscv64-gnu-4.63.4.tgz", + "integrity": "sha512-1uz2mGWHyptR7DgHHrlbdRAjXK7v7elGZ9lMja910/RP+ZYbX6xAmCiU9UZSX4hqmgtHMv6lr5l3kq1HIOpcag==", + "cpu": [ + "riscv64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-riscv64-musl": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-riscv64-musl/-/rollup-linux-riscv64-musl-4.63.4.tgz", + "integrity": "sha512-nLS8topojxyz7SRpKR2IODRpQ0XPZ+xaOXvT3+hqK/Uy8Lo5HFgkkIBiIrCu5tL5YqzTvgovGw55PwpahTAGig==", + "cpu": [ + "riscv64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-s390x-gnu": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-s390x-gnu/-/rollup-linux-s390x-gnu-4.63.4.tgz", + "integrity": "sha512-gs7DRKotr3l3q+jGPQBjH0ng1FjlEDm5ueQrkw5JtQvtLyEIcLASqAEaor56BhkKRzk+IcQzrcanBdb/bBQn8g==", + "cpu": [ + "s390x" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-x64-gnu": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-x64-gnu/-/rollup-linux-x64-gnu-4.63.4.tgz", + "integrity": "sha512-791ET7W17NnScOZM7h4dX5hYspxE28htPFsb1awY/NRR8+PRNkS53e475rDdxXXDrP+kwnCcNWg9CX5ztn/Aqw==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-linux-x64-musl": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-linux-x64-musl/-/rollup-linux-x64-musl-4.63.4.tgz", + "integrity": "sha512-iwZQRcmj7g88g3tzefIrQY7qvmuA/cfYwhrDtTBhsmukO4U2huVO5W+86XacUMRvdSFVAc6kZUZy21JaRwiB9w==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "linux" + ] + }, + "node_modules/@rollup/rollup-openbsd-x64": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-openbsd-x64/-/rollup-openbsd-x64-4.63.4.tgz", + "integrity": "sha512-dVHFp9gRWrdTpnqQuGfCwd7hOQDatK1VCP2iWhLY/cGrOQs/ucFzJ6A5SRqbXX12ZDI8EUuejSM5kwg+ja7Png==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "openbsd" + ] + }, + "node_modules/@rollup/rollup-openharmony-arm64": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-openharmony-arm64/-/rollup-openharmony-arm64-4.63.4.tgz", + "integrity": "sha512-t3NlauOW6gxZVVFcBEnO62Cb4wbyDFL416gTg1uFI/2tgqYQlf69FbSE115Ajre9I+c26Lk4mcmdFUsS/DGifQ==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "openharmony" + ] + }, + "node_modules/@rollup/rollup-win32-arm64-msvc": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-win32-arm64-msvc/-/rollup-win32-arm64-msvc-4.63.4.tgz", + "integrity": "sha512-xWuIaSye5FWZF8+UYtVEcHtRJDN5kN9Kfgxx3Kq8XIov9KSKbc1fiqQCm90SKrgQbUXZelbnUhnlUJmfSE7P9A==", + "cpu": [ + "arm64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "win32" + ] + }, + "node_modules/@rollup/rollup-win32-ia32-msvc": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-win32-ia32-msvc/-/rollup-win32-ia32-msvc-4.63.4.tgz", + "integrity": "sha512-9ALJJUOg/ZflMJepVo2PlgsGxSaxN7SQ4Z8GoZfVlarWr6r3rkHUNsd/zAio7p4YMtChSMXPionxej4Hkf6CXQ==", + "cpu": [ + "ia32" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "win32" + ] + }, + "node_modules/@rollup/rollup-win32-x64-gnu": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-win32-x64-gnu/-/rollup-win32-x64-gnu-4.63.4.tgz", + "integrity": "sha512-blj9z5qx/Pv4WU0W1NMFDB97e0JH5ed+aZGywW8WCvp/NhWX/4PFAq5uu6Q0AebNn+Vo6KzUYDT++JzTT5ojlQ==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "win32" + ] + }, + "node_modules/@rollup/rollup-win32-x64-msvc": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/@rollup/rollup-win32-x64-msvc/-/rollup-win32-x64-msvc-4.63.4.tgz", + "integrity": "sha512-Erx822VRBwLa124shbj+wNXe//BOgMEctDV0m1aqTQdNO1S69DgNUCFKC1RCeZfixs1J31l6igk1ziyXErbigQ==", + "cpu": [ + "x64" + ], + "dev": true, + "license": "MIT", + "optional": true, + "os": [ + "win32" + ] + }, + "node_modules/@types/body-parser": { + "version": "1.19.6", + "resolved": "https://registry.npmjs.org/@types/body-parser/-/body-parser-1.19.6.tgz", + "integrity": "sha512-HLFeCYgz89uk22N5Qg3dvGvsv46B8GLvKKo1zKG4NybA8U2DiEO3w9lqGg29t/tfLRJpJ6iQxnVw4OnB7MoM9g==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/connect": "*", + "@types/node": "*" + } + }, + "node_modules/@types/connect": { + "version": "3.4.38", + "resolved": "https://registry.npmjs.org/@types/connect/-/connect-3.4.38.tgz", + "integrity": "sha512-K6uROf1LD88uDQqJCktA4yzL1YYAK6NgfsI0v/mTgyPKWsX1CnJ0XPSDhViejru1GcRkLWb8RlzFYJRqGUbaug==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/node": "*" + } + }, + "node_modules/@types/estree": { + "version": "1.0.9", + "resolved": "https://registry.npmjs.org/@types/estree/-/estree-1.0.9.tgz", + "integrity": "sha512-GhdPgy1el4/ImP05X05Uw4cw2/M93BCUmnEvWZNStlCzEKME4Fkk+YpoA5OiHNQmoS7Cafb8Xa3Pya8m1Qrzeg==", + "dev": true, + "license": "MIT" + }, + "node_modules/@types/express": { + "version": "5.0.6", + "resolved": "https://registry.npmjs.org/@types/express/-/express-5.0.6.tgz", + "integrity": "sha512-sKYVuV7Sv9fbPIt/442koC7+IIwK5olP1KWeD88e/idgoJqDm3JV/YUiPwkoKK92ylff2MGxSz1CSjsXelx0YA==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/body-parser": "*", + "@types/express-serve-static-core": "^5.0.0", + "@types/serve-static": "^2" + } + }, + "node_modules/@types/express-serve-static-core": { + "version": "5.1.3", + "resolved": "https://registry.npmjs.org/@types/express-serve-static-core/-/express-serve-static-core-5.1.3.tgz", + "integrity": "sha512-dPfW8NFiOF4wOHc7+N/QSxlY9cfSsenewGbAz8C8U/MULPd/YZ27LvJUIlzaXie7e6Ove9YunJGgC9tbHD2cKw==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/node": "*", + "@types/qs": "*", + "@types/range-parser": "*", + "@types/send": "*" + } + }, + "node_modules/@types/http-errors": { + "version": "2.0.5", + "resolved": "https://registry.npmjs.org/@types/http-errors/-/http-errors-2.0.5.tgz", + "integrity": "sha512-r8Tayk8HJnX0FztbZN7oVqGccWgw98T/0neJphO91KkmOzug1KkofZURD4UaD5uH8AqcFLfdPErnBod0u71/qg==", + "dev": true, + "license": "MIT" + }, + "node_modules/@types/node": { + "version": "24.13.3", + "resolved": "https://registry.npmjs.org/@types/node/-/node-24.13.3.tgz", + "integrity": "sha512-Dh8vAsV36ig5wa9OX4pXvMc9D3Veibfw2wix0CUwYODLD8nkj9UsLjASr49nPg+2eKzxhBV+v7L8pXvT4e639Q==", + "dev": true, + "license": "MIT", + "dependencies": { + "undici-types": "~7.18.0" + } + }, + "node_modules/@types/qs": { + "version": "6.15.1", + "resolved": "https://registry.npmjs.org/@types/qs/-/qs-6.15.1.tgz", + "integrity": "sha512-GZHUBZR9hckSUhrxmp1nG6NwdpM9fCunJwyThLW1X3AyHgd9IlHb6VANpQQqDr2o/qQp6McZ3y/IA2rVzKzSbw==", + "dev": true, + "license": "MIT" + }, + "node_modules/@types/range-parser": { + "version": "1.2.7", + "resolved": "https://registry.npmjs.org/@types/range-parser/-/range-parser-1.2.7.tgz", + "integrity": "sha512-hKormJbkJqzQGhziax5PItDUTMAM9uE2XXQmM37dyd4hVM+5aVl7oVxMVUiVQn2oCQFN/LKCZdvSM0pFRqbSmQ==", + "dev": true, + "license": "MIT" + }, + "node_modules/@types/send": { + "version": "1.2.1", + "resolved": "https://registry.npmjs.org/@types/send/-/send-1.2.1.tgz", + "integrity": "sha512-arsCikDvlU99zl1g69TcAB3mzZPpxgw0UQnaHeC1Nwb015xp8bknZv5rIfri9xTOcMuaVgvabfIRA7PSZVuZIQ==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/node": "*" + } + }, + "node_modules/@types/serve-static": { + "version": "2.2.0", + "resolved": "https://registry.npmjs.org/@types/serve-static/-/serve-static-2.2.0.tgz", + "integrity": "sha512-8mam4H1NHLtu7nmtalF7eyBH14QyOASmcxHhSfEoRyr0nP/YdoesEtU+uSRvMe96TW/HPTtkoKqQLl53N7UXMQ==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/http-errors": "*", + "@types/node": "*" + } + }, + "node_modules/acorn": { + "version": "8.18.0", + "resolved": "https://registry.npmjs.org/acorn/-/acorn-8.18.0.tgz", + "integrity": "sha512-lGq+9yr1/GuAWaVYIHRjvvySG5/4VfKIvC8EWxStPdcDh/Ka7FG3twP6v4d5BkravUilhIAsG4Qj83t02LWUPQ==", + "dev": true, + "license": "MIT", + "bin": { + "acorn": "bin/acorn" + }, + "engines": { + "node": ">=0.4.0" + } + }, + "node_modules/any-promise": { + "version": "1.3.0", + "resolved": "https://registry.npmjs.org/any-promise/-/any-promise-1.3.0.tgz", + "integrity": "sha512-7UvmKalWRt1wgjL1RrGxoSJW/0QZFIegpeGvZG9kjp8vrRu55XTHbwnqq2GpXm9uLbcuhxm3IqX9OB4MZR1b2A==", + "dev": true, + "license": "MIT" + }, + "node_modules/bundle-require": { + "version": "5.1.0", + "resolved": "https://registry.npmjs.org/bundle-require/-/bundle-require-5.1.0.tgz", + "integrity": "sha512-3WrrOuZiyaaZPWiEt4G3+IffISVC9HYlWueJEBWED4ZH4aIAC2PnkdnuRrR94M+w6yGWn4AglWtJtBI8YqvgoA==", + "dev": true, + "license": "MIT", + "dependencies": { + "load-tsconfig": "^0.2.3" + }, + "engines": { + "node": "^12.20.0 || ^14.13.1 || >=16.0.0" + }, + "peerDependencies": { + "esbuild": ">=0.18" + } + }, + "node_modules/cac": { + "version": "6.7.14", + "resolved": "https://registry.npmjs.org/cac/-/cac-6.7.14.tgz", + "integrity": "sha512-b6Ilus+c3RrdDk+JhLKUAQfzzgLEPy6wcXqS7f/xe1EETvsDP6GORG7SFuOs6cID5YkqchW/LXZbX5bc8j7ZcQ==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=8" + } + }, + "node_modules/chokidar": { + "version": "4.0.3", + "resolved": "https://registry.npmjs.org/chokidar/-/chokidar-4.0.3.tgz", + "integrity": "sha512-Qgzu8kfBvo+cA4962jnP1KkS6Dop5NS6g7R5LFYJr4b8Ub94PPQXUksCw9PvXoeXPRRddRNC5C1JQUR2SMGtnA==", + "dev": true, + "license": "MIT", + "dependencies": { + "readdirp": "^4.0.1" + }, + "engines": { + "node": ">= 14.16.0" + }, + "funding": { + "url": "https://paulmillr.com/funding/" + } + }, + "node_modules/commander": { + "version": "13.1.0", + "resolved": "https://registry.npmjs.org/commander/-/commander-13.1.0.tgz", + "integrity": "sha512-/rFeCpNJQbhSZjGVwO9RFV3xPqbnERS8MmIQzCtD/zl6gpJuV/bMLuN92oG3F7d8oDEHHRrujSXNUr8fpjntKw==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=18" + } + }, + "node_modules/confbox": { + "version": "0.1.8", + "resolved": "https://registry.npmjs.org/confbox/-/confbox-0.1.8.tgz", + "integrity": "sha512-RMtmw0iFkeR4YV+fUOSucriAQNb9g8zFR52MWCtl+cCZOFRNL6zeB395vPzFhEjjn4fMxXudmELnl/KF/WrK6w==", + "dev": true, + "license": "MIT" + }, + "node_modules/consola": { + "version": "3.4.2", + "resolved": "https://registry.npmjs.org/consola/-/consola-3.4.2.tgz", + "integrity": "sha512-5IKcdX0nnYavi6G7TtOhwkYzyjfJlatbjMjuLSfE2kYT5pMDOilZ4OvMhi637CcDICTmz3wARPoyhqyX1Y+XvA==", + "dev": true, + "license": "MIT", + "engines": { + "node": "^14.18.0 || >=16.10.0" + } + }, + "node_modules/debug": { + "version": "4.4.3", + "resolved": "https://registry.npmjs.org/debug/-/debug-4.4.3.tgz", + "integrity": "sha512-RGwwWnwQvkVfavKVt22FGLw+xYSdzARwm0ru6DhTVA3umU5hZc28V3kO4stgYryrTlLpuvgI9GiijltAjNbcqA==", + "dev": true, + "license": "MIT", + "dependencies": { + "ms": "^2.1.3" + }, + "engines": { + "node": ">=6.0" + }, + "peerDependenciesMeta": { + "supports-color": { + "optional": true + } + } + }, + "node_modules/esbuild": { + "version": "0.28.2", + "resolved": "https://registry.npmjs.org/esbuild/-/esbuild-0.28.2.tgz", + "integrity": "sha512-HKVLS8dvII+xoKW9kmqxbRKrnWEXfJJr/FZhhJmiqIB0e053QNYFqOBouTMO/k5sID4MvCiUCvv8b9M4h32wIA==", + "dev": true, + "hasInstallScript": true, + "license": "MIT", + "bin": { + "esbuild": "bin/esbuild" + }, + "engines": { + "node": ">=18" + }, + "optionalDependencies": { + "@esbuild/aix-ppc64": "0.28.2", + "@esbuild/android-arm": "0.28.2", + "@esbuild/android-arm64": "0.28.2", + "@esbuild/android-x64": "0.28.2", + "@esbuild/darwin-arm64": "0.28.2", + "@esbuild/darwin-x64": "0.28.2", + "@esbuild/freebsd-arm64": "0.28.2", + "@esbuild/freebsd-x64": "0.28.2", + "@esbuild/linux-arm": "0.28.2", + "@esbuild/linux-arm64": "0.28.2", + "@esbuild/linux-ia32": "0.28.2", + "@esbuild/linux-loong64": "0.28.2", + "@esbuild/linux-mips64el": "0.28.2", + "@esbuild/linux-ppc64": "0.28.2", + "@esbuild/linux-riscv64": "0.28.2", + "@esbuild/linux-s390x": "0.28.2", + "@esbuild/linux-x64": "0.28.2", + "@esbuild/netbsd-arm64": "0.28.2", + "@esbuild/netbsd-x64": "0.28.2", + "@esbuild/openbsd-arm64": "0.28.2", + "@esbuild/openbsd-x64": "0.28.2", + "@esbuild/openharmony-arm64": "0.28.2", + "@esbuild/sunos-x64": "0.28.2", + "@esbuild/win32-arm64": "0.28.2", + "@esbuild/win32-ia32": "0.28.2", + "@esbuild/win32-x64": "0.28.2" + } + }, + "node_modules/fdir": { + "version": "6.5.0", + "resolved": "https://registry.npmjs.org/fdir/-/fdir-6.5.0.tgz", + "integrity": "sha512-tIbYtZbucOs0BRGqPJkshJUYdL+SDH7dVM8gjy+ERp3WAUjLEFJE+02kanyHtwjWOnwrKYBiwAmM0p4kLJAnXg==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=12.0.0" + }, + "peerDependencies": { + "picomatch": "^3 || ^4" + }, + "peerDependenciesMeta": { + "picomatch": { + "optional": true + } + } + }, + "node_modules/fix-dts-default-cjs-exports": { + "version": "1.0.1", + "resolved": "https://registry.npmjs.org/fix-dts-default-cjs-exports/-/fix-dts-default-cjs-exports-1.0.1.tgz", + "integrity": "sha512-pVIECanWFC61Hzl2+oOCtoJ3F17kglZC/6N94eRWycFgBH35hHx0Li604ZIzhseh97mf2p0cv7vVrOZGoqhlEg==", + "dev": true, + "license": "MIT", + "dependencies": { + "magic-string": "^0.30.17", + "mlly": "^1.7.4", + "rollup": "^4.34.8" + } + }, + "node_modules/fsevents": { + "version": "2.3.3", + "resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.3.tgz", + "integrity": "sha512-5xoDfX+fL7faATnagmWPpbFtwh/R77WmMMqqHGS65C3vvB0YHrgF+B1YmZ3441tMj5n63k0212XNoJwzlhffQw==", + "dev": true, + "hasInstallScript": true, + "license": "MIT", + "optional": true, + "os": [ + "darwin" + ], + "engines": { + "node": "^8.16.0 || ^10.6.0 || >=11.0.0" + } + }, + "node_modules/joycon": { + "version": "3.1.1", + "resolved": "https://registry.npmjs.org/joycon/-/joycon-3.1.1.tgz", + "integrity": "sha512-34wB/Y7MW7bzjKRjUKTa46I2Z7eV62Rkhva+KkopW7Qvv/OSWBqvkSY7vusOPrNuZcUG3tApvdVgNB8POj3SPw==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=10" + } + }, + "node_modules/lilconfig": { + "version": "3.1.3", + "resolved": "https://registry.npmjs.org/lilconfig/-/lilconfig-3.1.3.tgz", + "integrity": "sha512-/vlFKAoH5Cgt3Ie+JLhRbwOsCQePABiU3tJ1egGvyQ+33R/vcwM2Zl2QR/LzjsBeItPt3oSVXapn+m4nQDvpzw==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=14" + }, + "funding": { + "url": "https://github.com/sponsors/antonk52" + } + }, + "node_modules/lines-and-columns": { + "version": "1.2.4", + "resolved": "https://registry.npmjs.org/lines-and-columns/-/lines-and-columns-1.2.4.tgz", + "integrity": "sha512-7ylylesZQ/PV29jhEDl3Ufjo6ZX7gCqJr5F7PKrqc93v7fzSymt1BpwEU8nAUXs8qzzvqhbjhK5QZg6Mt/HkBg==", + "dev": true, + "license": "MIT" + }, + "node_modules/load-tsconfig": { + "version": "0.2.5", + "resolved": "https://registry.npmjs.org/load-tsconfig/-/load-tsconfig-0.2.5.tgz", + "integrity": "sha512-IXO6OCs9yg8tMKzfPZ1YmheJbZCiEsnBdcB03l0OcfK9prKnJb96siuHCr5Fl37/yo9DnKU+TLpxzTUspw9shg==", + "dev": true, + "license": "MIT", + "engines": { + "node": "^12.20.0 || ^14.13.1 || >=16.0.0" + } + }, + "node_modules/magic-string": { + "version": "0.30.21", + "resolved": "https://registry.npmjs.org/magic-string/-/magic-string-0.30.21.tgz", + "integrity": "sha512-vd2F4YUyEXKGcLHoq+TEyCjxueSeHnFxyyjNp80yg0XV4vUhnDer/lvvlqM/arB5bXQN5K2/3oinyCRyx8T2CQ==", + "dev": true, + "license": "MIT", + "dependencies": { + "@jridgewell/sourcemap-codec": "^1.5.5" + } + }, + "node_modules/mlly": { + "version": "1.8.2", + "resolved": "https://registry.npmjs.org/mlly/-/mlly-1.8.2.tgz", + "integrity": "sha512-d+ObxMQFmbt10sretNDytwt85VrbkhhUA/JBGm1MPaWJ65Cl4wOgLaB1NYvJSZ0Ef03MMEU/0xpPMXUIQ29UfA==", + "dev": true, + "license": "MIT", + "dependencies": { + "acorn": "^8.16.0", + "pathe": "^2.0.3", + "pkg-types": "^1.3.1", + "ufo": "^1.6.3" + } + }, + "node_modules/ms": { + "version": "2.1.3", + "resolved": "https://registry.npmjs.org/ms/-/ms-2.1.3.tgz", + "integrity": "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA==", + "dev": true, + "license": "MIT" + }, + "node_modules/mz": { + "version": "2.7.0", + "resolved": "https://registry.npmjs.org/mz/-/mz-2.7.0.tgz", + "integrity": "sha512-z81GNO7nnYMEhrGh9LeymoE4+Yr0Wn5McHIZMK5cfQCl+NDX08sCZgUc9/6MHni9IWuFLm1Z3HTCXu2z9fN62Q==", + "dev": true, + "license": "MIT", + "dependencies": { + "any-promise": "^1.0.0", + "object-assign": "^4.0.1", + "thenify-all": "^1.0.0" + } + }, + "node_modules/object-assign": { + "version": "4.1.1", + "resolved": "https://registry.npmjs.org/object-assign/-/object-assign-4.1.1.tgz", + "integrity": "sha512-rJgTQnkUnH1sFw8yT6VSU3zD3sWmu6sZhIseY8VX+GRu3P6F7Fu+JNDoXfklElbLJSnc3FUQHVe4cU5hj+BcUg==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=0.10.0" + } + }, + "node_modules/pathe": { + "version": "2.0.3", + "resolved": "https://registry.npmjs.org/pathe/-/pathe-2.0.3.tgz", + "integrity": "sha512-WUjGcAqP1gQacoQe+OBJsFA7Ld4DyXuUIjZ5cc75cLHvJ7dtNsTugphxIADwspS+AraAUePCKrSVtPLFj/F88w==", + "dev": true, + "license": "MIT" + }, + "node_modules/picocolors": { + "version": "1.1.1", + "resolved": "https://registry.npmjs.org/picocolors/-/picocolors-1.1.1.tgz", + "integrity": "sha512-xceH2snhtb5M9liqDsmEw56le376mTZkEX/jEb/RxNFyegNul7eNslCXP9FDj/Lcu0X8KEyMceP2ntpaHrDEVA==", + "dev": true, + "license": "ISC" + }, + "node_modules/picomatch": { + "version": "4.0.7", + "resolved": "https://registry.npmjs.org/picomatch/-/picomatch-4.0.7.tgz", + "integrity": "sha512-qcJu88Q2IWqJsDD529JKMdwGm/dvInW4HvQnRwiH9JtihJvzGOscDtHE3x1pBKeUOTysQ8kVmLnJ2kJu7yhcGA==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=12" + }, + "funding": { + "url": "https://github.com/sponsors/jonschlinkert" + } + }, + "node_modules/pirates": { + "version": "4.0.7", + "resolved": "https://registry.npmjs.org/pirates/-/pirates-4.0.7.tgz", + "integrity": "sha512-TfySrs/5nm8fQJDcBDuUng3VOUKsd7S+zqvbOTiGXHfxX4wK31ard+hoNuvkicM/2YFzlpDgABOevKSsB4G/FA==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">= 6" + } + }, + "node_modules/pkg-types": { + "version": "1.3.1", + "resolved": "https://registry.npmjs.org/pkg-types/-/pkg-types-1.3.1.tgz", + "integrity": "sha512-/Jm5M4RvtBFVkKWRu2BLUTNP8/M2a+UwuAX+ae4770q1qVGtfjG+WTCupoZixokjmHiry8uI+dlY8KXYV5HVVQ==", + "dev": true, + "license": "MIT", + "dependencies": { + "confbox": "^0.1.8", + "mlly": "^1.7.4", + "pathe": "^2.0.1" + } + }, + "node_modules/postcss-load-config": { + "version": "6.0.1", + "resolved": "https://registry.npmjs.org/postcss-load-config/-/postcss-load-config-6.0.1.tgz", + "integrity": "sha512-oPtTM4oerL+UXmx+93ytZVN82RrlY/wPUV8IeDxFrzIjXOLF1pN+EmKPLbubvKHT2HC20xXsCAH2Z+CKV6Oz/g==", + "dev": true, + "funding": [ + { + "type": "opencollective", + "url": "https://opencollective.com/postcss/" + }, + { + "type": "github", + "url": "https://github.com/sponsors/ai" + } + ], + "license": "MIT", + "dependencies": { + "lilconfig": "^3.1.1" + }, + "engines": { + "node": ">= 18" + }, + "peerDependencies": { + "jiti": ">=1.21.0", + "postcss": ">=8.0.9", + "tsx": "^4.8.1", + "yaml": "^2.4.2" + }, + "peerDependenciesMeta": { + "jiti": { + "optional": true + }, + "postcss": { + "optional": true + }, + "tsx": { + "optional": true + }, + "yaml": { + "optional": true + } + } + }, + "node_modules/readdirp": { + "version": "4.1.2", + "resolved": "https://registry.npmjs.org/readdirp/-/readdirp-4.1.2.tgz", + "integrity": "sha512-GDhwkLfywWL2s6vEjyhri+eXmfH6j1L7JE27WhqLeYzoh/A3DBaYGEj2H/HFZCn/kMfim73FXxEJTw06WtxQwg==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">= 14.18.0" + }, + "funding": { + "type": "individual", + "url": "https://paulmillr.com/funding/" + } + }, + "node_modules/resolve-from": { + "version": "5.0.0", + "resolved": "https://registry.npmjs.org/resolve-from/-/resolve-from-5.0.0.tgz", + "integrity": "sha512-qYg9KP24dD5qka9J47d0aVky0N+b4fTU89LN9iDnjB5waksiC49rvMB0PrUJQGoTmH50XPiqOvAjDfaijGxYZw==", + "dev": true, + "license": "MIT", + "engines": { + "node": ">=8" + } + }, + "node_modules/rollup": { + "version": "4.63.4", + "resolved": "https://registry.npmjs.org/rollup/-/rollup-4.63.4.tgz", + "integrity": "sha512-4U0liVayNIoLp3GFl1FcI8561WepLnZ1rqfraGh7S9B3Ur5F9S283y8Futii7RUU2C/97tOBmBy7nYvhoiOpbQ==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/estree": "1.0.9" + }, + "bin": { + "rollup": "dist/bin/rollup" + }, + "engines": { + "node": ">=18.0.0", + "npm": ">=8.0.0" + }, + "optionalDependencies": { + "@napi-rs/lzma-linux-x64-gnu": "1.5.1", + "@rollup/rollup-android-arm-eabi": "4.63.4", + "@rollup/rollup-android-arm64": "4.63.4", + "@rollup/rollup-darwin-arm64": "4.63.4", + "@rollup/rollup-darwin-x64": "4.63.4", + "@rollup/rollup-freebsd-arm64": "4.63.4", + "@rollup/rollup-freebsd-x64": "4.63.4", + "@rollup/rollup-linux-arm-gnueabihf": "4.63.4", + "@rollup/rollup-linux-arm-musleabihf": "4.63.4", + "@rollup/rollup-linux-arm64-gnu": "4.63.4", + "@rollup/rollup-linux-arm64-musl": "4.63.4", + "@rollup/rollup-linux-loong64-gnu": "4.63.4", + "@rollup/rollup-linux-loong64-musl": "4.63.4", + "@rollup/rollup-linux-ppc64-gnu": "4.63.4", + "@rollup/rollup-linux-ppc64-musl": "4.63.4", + "@rollup/rollup-linux-riscv64-gnu": "4.63.4", + "@rollup/rollup-linux-riscv64-musl": "4.63.4", + "@rollup/rollup-linux-s390x-gnu": "4.63.4", + "@rollup/rollup-linux-x64-gnu": "4.63.4", + "@rollup/rollup-linux-x64-musl": "4.63.4", + "@rollup/rollup-openbsd-x64": "4.63.4", + "@rollup/rollup-openharmony-arm64": "4.63.4", + "@rollup/rollup-win32-arm64-msvc": "4.63.4", + "@rollup/rollup-win32-ia32-msvc": "4.63.4", + "@rollup/rollup-win32-x64-gnu": "4.63.4", + "@rollup/rollup-win32-x64-msvc": "4.63.4", + "fsevents": "~2.3.2" + } + }, + "node_modules/source-map": { + "version": "0.7.6", + "resolved": "https://registry.npmjs.org/source-map/-/source-map-0.7.6.tgz", + "integrity": "sha512-i5uvt8C3ikiWeNZSVZNWcfZPItFQOsYTUAOkcUPGd8DqDy1uOUikjt5dG+uRlwyvR108Fb9DOd4GvXfT0N2/uQ==", + "dev": true, + "license": "BSD-3-Clause", + "engines": { + "node": ">= 12" + } + }, + "node_modules/sucrase": { + "version": "3.35.1", + "resolved": "https://registry.npmjs.org/sucrase/-/sucrase-3.35.1.tgz", + "integrity": "sha512-DhuTmvZWux4H1UOnWMB3sk0sbaCVOoQZjv8u1rDoTV0HTdGem9hkAZtl4JZy8P2z4Bg0nT+YMeOFyVr4zcG5Tw==", + "dev": true, + "license": "MIT", + "dependencies": { + "@jridgewell/gen-mapping": "^0.3.2", + "commander": "^4.0.0", + "lines-and-columns": "^1.1.6", + "mz": "^2.7.0", + "pirates": "^4.0.1", + "tinyglobby": "^0.2.11", + "ts-interface-checker": "^0.1.9" + }, + "bin": { + "sucrase": "bin/sucrase", + "sucrase-node": "bin/sucrase-node" + }, + "engines": { + "node": ">=16 || 14 >=14.17" + } + }, + "node_modules/thenify": { + "version": "3.3.1", + "resolved": "https://registry.npmjs.org/thenify/-/thenify-3.3.1.tgz", + "integrity": "sha512-RVZSIV5IG10Hk3enotrhvz0T9em6cyHBLkH/YAZuKqd8hRkKhSfCGIcP2KUY0EPxndzANBmNllzWPwak+bheSw==", + "dev": true, + "license": "MIT", + "dependencies": { + "any-promise": "^1.0.0" + } + }, + "node_modules/thenify-all": { + "version": "1.6.0", + "resolved": "https://registry.npmjs.org/thenify-all/-/thenify-all-1.6.0.tgz", + "integrity": "sha512-RNxQH/qI8/t3thXJDwcstUO4zeqo64+Uy/+sNVRBx4Xn2OX+OZ9oP+iJnNFqplFra2ZUVeKCSa2oVWi3T4uVmA==", + "dev": true, + "license": "MIT", + "dependencies": { + "thenify": ">= 3.1.0 < 4" + }, + "engines": { + "node": ">=0.8" + } + }, + "node_modules/tinyexec": { + "version": "0.3.2", + "resolved": "https://registry.npmjs.org/tinyexec/-/tinyexec-0.3.2.tgz", + "integrity": "sha512-KQQR9yN7R5+OSwaK0XQoj22pwHoTlgYqmUscPYoknOoWCWfj/5/ABTMRi69FrKU5ffPVh5QcFikpWJI/P1ocHA==", + "dev": true, + "license": "MIT" + }, + "node_modules/tinyglobby": { + "version": "0.2.17", + "resolved": "https://registry.npmjs.org/tinyglobby/-/tinyglobby-0.2.17.tgz", + "integrity": "sha512-wXR/dYpcqKmfWpEdZjiKJOwCNFndD0DMnrW/cYjVGttEkBfVgcLFHoNrlj47mjOVic9yyNu65alsgF4NQyTa2g==", + "dev": true, + "license": "MIT", + "dependencies": { + "fdir": "^6.5.0", + "picomatch": "^4.0.4" + }, + "engines": { + "node": ">=12.0.0" + }, + "funding": { + "url": "https://github.com/sponsors/SuperchupuDev" + } + }, + "node_modules/tree-kill": { + "version": "1.2.2", + "resolved": "https://registry.npmjs.org/tree-kill/-/tree-kill-1.2.2.tgz", + "integrity": "sha512-L0Orpi8qGpRG//Nd+H90vFB+3iHnue1zSSGmNOOCh1GLJ7rUKVwV2HvijphGQS2UmhUZewS9VgvxYIdgr+fG1A==", + "dev": true, + "license": "MIT", + "bin": { + "tree-kill": "cli.js" + } + }, + "node_modules/ts-interface-checker": { + "version": "0.1.13", + "resolved": "https://registry.npmjs.org/ts-interface-checker/-/ts-interface-checker-0.1.13.tgz", + "integrity": "sha512-Y/arvbn+rrz3JCKl9C4kVNfTfSm2/mEp5FSz5EsZSANGPSlQrpRI5M4PKF+mJnE52jOO90PnPSc3Ur3bTQw0gA==", + "dev": true, + "license": "Apache-2.0" + }, + "node_modules/tsup": { + "version": "8.5.1", + "resolved": "https://registry.npmjs.org/tsup/-/tsup-8.5.1.tgz", + "integrity": "sha512-xtgkqwdhpKWr3tKPmCkvYmS9xnQK3m3XgxZHwSUjvfTjp7YfXe5tT3GgWi0F2N+ZSMsOeWeZFh7ZZFg5iPhing==", + "dev": true, + "license": "MIT", + "dependencies": { + "bundle-require": "^5.1.0", + "cac": "^6.7.14", + "chokidar": "^4.0.3", + "consola": "^3.4.0", + "debug": "^4.4.0", + "esbuild": "^0.27.0", + "fix-dts-default-cjs-exports": "^1.0.0", + "joycon": "^3.1.1", + "picocolors": "^1.1.1", + "postcss-load-config": "^6.0.1", + "resolve-from": "^5.0.0", + "rollup": "^4.34.8", + "source-map": "^0.7.6", + "sucrase": "^3.35.0", + "tinyexec": "^0.3.2", + "tinyglobby": "^0.2.11", + "tree-kill": "^1.2.2" + }, + "bin": { + "tsup": "dist/cli-default.js", + "tsup-node": "dist/cli-node.js" + }, + "engines": { + "node": ">=18" + }, + "peerDependencies": { + "@microsoft/api-extractor": "^7.36.0", + "@swc/core": "^1", + "postcss": "^8.4.12", + "typescript": ">=4.5.0" + }, + "peerDependenciesMeta": { + "@microsoft/api-extractor": { + "optional": true + }, + "@swc/core": { + "optional": true + }, + "postcss": { + "optional": true + }, + "typescript": { + "optional": true + } + } + }, + "node_modules/ufo": { + "version": "1.6.4", + "resolved": "https://registry.npmjs.org/ufo/-/ufo-1.6.4.tgz", + "integrity": "sha512-JFNbkD1Svwe0KvGi8GOeLcP4kAWQ609twvCdcHxq1oSL8svv39ZuSvajcD8B+5D0eL4+s1Is2D/O6KN3qcTeRA==", + "dev": true, + "license": "MIT" + }, + "node_modules/undici-types": { + "version": "7.18.2", + "resolved": "https://registry.npmjs.org/undici-types/-/undici-types-7.18.2.tgz", + "integrity": "sha512-AsuCzffGHJybSaRrmr5eHr81mwJU3kjw6M+uprWvCXiNeN9SOGwQ3Jn8jb8m3Z6izVgknn1R0FTCEAP2QrLY/w==", + "dev": true, + "license": "MIT" + } + } +} diff --git a/clients/mcpdo/package.json b/clients/mcpdo/package.json new file mode 100644 index 0000000000..49c5bd84bc --- /dev/null +++ b/clients/mcpdo/package.json @@ -0,0 +1,41 @@ +{ + "name": "@modelcontextprotocol/mcpdo", + "private": true, + "description": "Connection-oriented MCP Inspector CLI (mcpdo) — connect once, run many commands", + "license": "MIT", + "type": "module", + "main": "build/mcp-bin.js", + "bin": { + "mcpdo": "./build/mcp-bin.js" + }, + "files": [ + "build", + "README.md" + ], + "scripts": { + "build": "tsup", + "build:dev": "node scripts/stop-dev-daemon.mjs && tsup", + "typecheck": "tsc --noEmit -p tsconfig.json && tsc --noEmit -p tsconfig.test.json", + "check": "npm run format:check && npm run lint && npm run typecheck", + "validate": "npm run check && npm run test", + "test": "vitest run", + "test:watch": "vitest", + "test:coverage": "npm run test-servers:build && npm run build && vitest run --coverage", + "test-servers:build": "tsc -p ../../test-servers --noCheck", + "pretest": "npm run test-servers:build && npm run build", + "lint": "eslint . --max-warnings 0", + "format": "prettier --write src __tests__ scripts \"*.{ts,tsx,mts,cts,js,jsx,mjs,cjs}\"", + "format:check": "prettier --check src __tests__ scripts \"*.{ts,tsx,mts,cts,js,jsx,mjs,cjs}\"" + }, + "devDependencies": { + "@types/express": "^5.0.6", + "tsup": "^8.5.0" + }, + "overrides": { + "@types/node": "^24.12.4", + "esbuild": "^0.28.2", + "sucrase": { + "commander": "^13.1.0" + } + } +} diff --git a/clients/mcpdo/scripts/stop-dev-daemon.mjs b/clients/mcpdo/scripts/stop-dev-daemon.mjs new file mode 100644 index 0000000000..e5d8d56070 --- /dev/null +++ b/clients/mcpdo/scripts/stop-dev-daemon.mjs @@ -0,0 +1,18 @@ +/** + * Pre-build step for `build:dev`: stop a daemon still running from a previous + * build so the fresh bundle isn't shadowed by a stale resident process. + * + * Best-effort by design — a missing `build/` (first build) or no running + * daemon must not fail the build. Kept as a script rather than shell syntax + * so the npm script works on Windows too (cmd.exe has no `;` sequencing or + * `/dev/null`). + */ +import { execFileSync } from "node:child_process"; + +try { + execFileSync(process.execPath, ["build/mcp-bin.js", "daemon", "stop"], { + stdio: "ignore", + }); +} catch { + // Nothing to stop (or nothing built yet) — proceed with the build. +} diff --git a/clients/mcpdo/src/connection/auth-helper.ts b/clients/mcpdo/src/connection/auth-helper.ts new file mode 100644 index 0000000000..073779a867 --- /dev/null +++ b/clients/mcpdo/src/connection/auth-helper.ts @@ -0,0 +1,468 @@ +import { spawn } from "node:child_process"; +import { createHash } from "node:crypto"; +import * as fs from "node:fs"; +import * as path from "node:path"; +import { CallbackNavigation } from "@inspector/core/auth/index.js"; +import type { + InspectorServerSettings, + MCPServerConfig, +} from "@inspector/core/mcp/types.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { getDaemonDir } from "../daemon/paths.js"; +import { authorizeInFrontend } from "./authorize.js"; +import { sanitizeText } from "./sanitize.js"; + +/** + * Detached OAuth completion helper for the non-TTY `connect` path. + * + * An agent driving mcpdo over pipes cannot sit on a blocking interactive + * OAuth flow: the auth URL stays invisible in a buffered foreground pipe, and + * killing the foreground process would tear down the loopback callback + * listener the URL points at (staling the link). Instead, `connect` spawns + * this helper detached: the helper owns the whole interactive flow + * (callback listener, authorization-code exchange, token persistence to the + * shared `oauth.json`), reports the freshly minted authorize URL back over + * its stdout pipe, and keeps running after the parent exits — bounded by the + * flow's own 15-minute callback wait. The parent relays the URL and exits; + * the daemon-side pending entry completes on first use once tokens land. + * + * Params travel over **stdin as JSON**, never argv: `serverConfig` may carry + * header secrets, and argv is world-visible in `ps`. + */ + +/** Hidden subcommand name (see registerAuthCommands in mcp.ts). */ +export const AUTH_HELPER_COMMAND = "auth/complete-signin"; + +/** Params the parent writes to the helper's stdin as one JSON document. */ +export type AuthHelperParams = { + serverConfig: MCPServerConfig; + serverSettings?: InspectorServerSettings; +}; + +/** One NDJSON line on the helper's stdout. */ +export type AuthHelperEvent = + | { event: "auth_url"; url: string } + | { event: "done" } + | { event: "error"; message: string }; + +/** + * Pending sign-in marker, one per flow key (a server URL, or an + * `ema-idp:<issuer>` key for the EMA IdP login), in the daemon dir (0700). + * A repeat `connect` while a helper is still waiting must reprint the SAME + * URL rather than mint a second flow: the fixed loopback callback port makes + * a second listener fail, and a fresh PKCE state would stale the link the + * user is already holding. + */ +export type PendingAuthMarker = { + url: string; + pid: number; + /** Epoch ms; matches the flow's own callback-wait bound. */ + expiresAt: number; +}; + +/** Matches the interactive flow's 15-minute loopback callback wait. */ +export const PENDING_AUTH_TTL_MS = 15 * 60 * 1000; + +/** Bound on the parent's wait for the helper to report the auth URL. */ +const AUTH_URL_WAIT_MS = 60 * 1000; + +/** Bound on the helper's wait for params on stdin (parent writes eagerly). */ +const HELPER_STDIN_TIMEOUT_MS = 30 * 1000; + +export function pendingAuthMarkerPath(key: string): string { + const hash = createHash("sha256").update(key).digest("hex").slice(0, 16); + return path.join(getDaemonDir(), `pending-auth-${hash}.json`); +} + +/** + * Read the marker for `key` if it is still live: unexpired AND its + * helper process is still running (a killed/crashed helper must not pin a + * dead URL for up to 15 minutes). Stale markers are removed best-effort. + */ +export function readLivePendingAuthMarker( + key: string, +): PendingAuthMarker | undefined { + const markerPath = pendingAuthMarkerPath(key); + let marker: PendingAuthMarker; + try { + const parsed = JSON.parse(fs.readFileSync(markerPath, "utf8")) as unknown; + if ( + typeof parsed !== "object" || + parsed === null || + typeof (parsed as PendingAuthMarker).url !== "string" || + typeof (parsed as PendingAuthMarker).pid !== "number" || + typeof (parsed as PendingAuthMarker).expiresAt !== "number" + ) { + throw new Error("malformed marker"); + } + marker = parsed as PendingAuthMarker; + } catch { + return undefined; + } + const live = + marker.expiresAt > Date.now() && + (() => { + try { + process.kill(marker.pid, 0); + return true; + } catch { + return false; + } + })(); + if (!live) { + // Deliberately NOT deleted here: unlinking by pathname after the read + // would race a just-spawned helper replacing the marker (TOCTOU — the + // rm could delete the fresh URL). Stale markers are inert (re-validated + // on every read) and the next flow's writePendingAuthMarker replaces + // them. + return undefined; + } + return marker; +} + +/** + * Remove the pending-auth marker only if THIS process wrote the one on disk. + * Past the 15-minute TTL a replacement flow may have published a fresh + * marker at the same shared pathname, and an unconditional rm at helper exit + * would delete the replacement's URL out from under its callers. A read→rm + * microsecond window remains (POSIX has no compare-and-delete); losing it + * costs one extra sign-in prompt, never a wrong URL. + */ +export function removeOwnPendingAuthMarker(markerPath: string) { + try { + const onDisk = JSON.parse( + fs.readFileSync(markerPath, "utf8"), + ) as PendingAuthMarker; + if (onDisk.pid === process.pid) fs.rmSync(markerPath, { force: true }); + } catch { + // Missing or unreadable marker: nothing of ours to clean up. + } +} + +export function writePendingAuthMarker( + markerPath: string, + marker: PendingAuthMarker, +) { + // Recreate exclusively (same symlink hardening as the daemon log): an + // append/overwrite open would follow a planted symlink and only apply the + // 0600 mode on create. + fs.rmSync(markerPath, { force: true }); + fs.writeFileSync(markerPath, `${JSON.stringify(marker)}\n`, { + flag: "wx", + mode: 0o600, + }); +} + +/** Read the helper's stdin to EOF and parse the params document. */ +async function readHelperParams(): Promise<AuthHelperParams> { + const chunks: Buffer[] = []; + const body = await new Promise<string>((resolve, reject) => { + const timer = setTimeout(() => { + reject(new Error("timed out waiting for params on stdin")); + }, HELPER_STDIN_TIMEOUT_MS); + timer.unref(); + process.stdin.on("data", (chunk: Buffer) => chunks.push(chunk)); + process.stdin.on("end", () => { + clearTimeout(timer); + resolve(Buffer.concat(chunks).toString("utf8")); + }); + process.stdin.on("error", (error) => { + clearTimeout(timer); + reject(error); + }); + }); + const parsed = JSON.parse(body) as AuthHelperParams; + if (typeof parsed !== "object" || parsed === null || !parsed.serverConfig) { + throw new Error("auth helper params must include serverConfig"); + } + return parsed; +} + +/** + * Entry point for the hidden helper subcommand. Runs the full interactive + * OAuth flow with a navigation that reports the authorize URL as an NDJSON + * event on stdout (instead of printing a prompt line) and never opens a + * browser — the parent (or the human it relayed the URL to) does that. + */ +export async function runAuthHelper(): Promise<void> { + // The parent unrefs and exits once it has the URL; every later stdout + // write would EPIPE without this guard. + const emit = (event: AuthHelperEvent) => { + try { + process.stdout.write(`${JSON.stringify(event)}\n`); + } catch { + // Parent is gone; the flow itself is unaffected. + } + }; + process.stdout.on("error", () => {}); + + const params = await readHelperParams(); + const serverUrl = + "url" in params.serverConfig ? params.serverConfig.url : undefined; + let markerPath: string | undefined; + try { + await authorizeInFrontend(params.serverConfig, params.serverSettings, { + makeNavigation: (autoOpenControl) => + new CallbackNavigation(async (url) => { + // Mirror createCliOAuthNavigation's arming: SDK-internal auth() + // during the plain connect() attempt must not leak a URL the + // flow isn't listening for yet. + if (!autoOpenControl.armed) return; + if (serverUrl !== undefined) { + markerPath = pendingAuthMarkerPath(serverUrl); + writePendingAuthMarker(markerPath, { + url: url.href, + pid: process.pid, + expiresAt: Date.now() + PENDING_AUTH_TTL_MS, + }); + } + emit({ event: "auth_url", url: url.href }); + }), + }); + emit({ event: "done" }); + } catch (error) { + emit({ + event: "error", + message: error instanceof Error ? error.message : String(error), + }); + throw error; + } finally { + if (markerPath !== undefined) removeOwnPendingAuthMarker(markerPath); + } +} + +/** + * Atomically reserve the right to spawn the sign-in helper for one server. + * `wx` creation is the atomicity (O_EXCL also refuses a planted symlink). + * + * A leftover lock from a crashed reserver is stolen once it is older than + * the URL wait window. The steal claims the specific stale file by an + * atomic rename to a per-pid path — concurrent stealers cannot both win, + * and a winner that renamed a lock which turned out to be fresh backs off + * (POSIX has no compare-and-delete; the rename makes the claim itself + * exclusive, which is what prevents a double spawn). + * + * @returns true if this process holds the reservation. + */ +function tryReserveAuthFlow(lockPath: string): boolean { + const isStale = (mtimeMs: number) => + Date.now() - mtimeMs > AUTH_URL_WAIT_MS + 5_000; + const create = () => + fs.writeFileSync(lockPath, `${process.pid}\n`, { flag: "wx", mode: 0o600 }); + try { + create(); + return true; + } catch { + try { + if (!isStale(fs.statSync(lockPath).mtimeMs)) return false; + // Claim the stale lock atomically: only one renamer succeeds. + const claimPath = `${lockPath}.claim-${process.pid}`; + fs.renameSync(lockPath, claimPath); + const claimedFresh = !isStale(fs.statSync(claimPath).mtimeMs); + fs.rmSync(claimPath, { force: true }); + // The claimed file was recreated fresh between stat and rename: an + // active reserver holds the flow — back off and wait for its marker. + // (Un-injectable microsecond race; the guard is what matters.) + /* v8 ignore next */ + if (claimedFresh) return false; + create(); + return true; + } catch { + // Lock vanished, was claimed by another stealer, or was recreated + // mid-steal: treat as held by another process. + return false; + } + } +} + +/** + * Another connect holds the flow reservation: wait for its helper to publish + * the marker and reuse that URL instead of spawning a competing helper. + */ +async function waitForPendingAuthUrl( + serverUrl: string, + waitMs: number, + pollMs: number, +): Promise<string> { + const deadline = Date.now() + waitMs; + for (;;) { + const marker = readLivePendingAuthMarker(serverUrl); + if (marker !== undefined) return marker.url; + if (Date.now() >= deadline) { + throw new CliExitCodeError( + EXIT_CODES.AUTH_REQUIRED, + "Timed out waiting for the in-progress sign-in flow to produce an authorization URL.", + { code: "auth_required" }, + ); + } + await new Promise<void>((resolve) => { + // NOT unref'ed: for a reservation loser this timer may be the only + // live handle, and an unref'ed one would let Node exit cleanly + // mid-wait — the connect would print no authorization URL at all. + setTimeout(resolve, pollMs); + }); + } +} + +/** + * Non-TTY connect path: return the authorize URL for `serverConfig`, either + * from a still-live pending marker (helper already waiting — reuse its URL) + * or by spawning a fresh detached helper and reading the URL off its stdout. + * + * The marker check and helper spawn are made atomic by a per-server lock + * file: concurrent connects for the same server would otherwise both pass + * the check and spawn helpers that contend for the OAuth callback port. The + * loser of the reservation waits for the winner's helper to publish the + * marker (written before the helper reports the URL) and reuses it. + * + * After this resolves the helper is unrefed and survives this process: it + * holds the loopback callback listener and completes the token exchange when + * the user finishes signing in. + */ +export async function obtainPendingAuthUrl( + serverConfig: MCPServerConfig, + serverSettings: InspectorServerSettings | undefined, + options?: { helperArgv1?: string; waitMs?: number; pollMs?: number }, +): Promise<string> { + const serverUrl = "url" in serverConfig ? serverConfig.url : undefined; + return obtainPendingUrlForKey( + serverUrl, + AUTH_HELPER_COMMAND, + { serverConfig, serverSettings }, + options, + ); +} + +/** + * Shared engine behind {@link obtainPendingAuthUrl} and the EMA login's + * non-TTY park (ema-login-helper.ts): marker reuse, flow reservation, helper + * spawn. `markerKey` is any stable string identifying the one flow callers + * must share (a server URL, an `ema-idp:<issuer>` key); `undefined` skips + * marker reuse entirely and always spawns (stdio configs have no URL to key + * on). + */ +export async function obtainPendingUrlForKey( + markerKey: string | undefined, + helperCommand: string, + helperParams: unknown, + options?: { helperArgv1?: string; waitMs?: number; pollMs?: number }, +): Promise<string> { + const waitMs = options?.waitMs ?? AUTH_URL_WAIT_MS; + let lockPath: string | undefined; + if (markerKey !== undefined) { + const marker = readLivePendingAuthMarker(markerKey); + if (marker !== undefined) return marker.url; + lockPath = `${pendingAuthMarkerPath(markerKey)}.lock`; + if (!tryReserveAuthFlow(lockPath)) { + return waitForPendingAuthUrl(markerKey, waitMs, options?.pollMs ?? 250); + } + } + + try { + return await spawnAuthHelperForUrl( + helperCommand, + helperParams, + waitMs, + options?.helperArgv1, + ); + } finally { + // Success: the helper's marker is already on disk (written before the + // URL event), so later connects reuse it. Failure: releasing lets the + // next attempt spawn a fresh helper. + if (lockPath !== undefined) fs.rmSync(lockPath, { force: true }); + } +} + +/** Spawn the detached helper and read the authorize URL off its stdout. */ +async function spawnAuthHelperForUrl( + helperCommand: string, + helperParams: unknown, + waitMs: number, + helperArgv1?: string, +): Promise<string> { + /* v8 ignore next 6 -- argv[1] is always the mcpdo bin in production. */ + const script = helperArgv1 ?? process.argv[1]; + if (!script) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "Cannot locate the mcpdo entry script to spawn the sign-in helper.", + { code: "usage" }, + ); + } + const child = spawn(process.execPath, [script, helperCommand], { + detached: true, + stdio: ["pipe", "pipe", "ignore"], + env: process.env, + }); + child.stdin.on("error", () => {}); + child.stdin.write(JSON.stringify(helperParams)); + child.stdin.end(); + + try { + return await new Promise<string>((resolve, reject) => { + let buffer = ""; + const fail = (message: string) => { + reject( + new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, message, { + code: "auth_required", + }), + ); + }; + const timer = setTimeout(() => { + fail( + "Timed out waiting for the sign-in helper to produce an authorization URL.", + ); + }, waitMs); + timer.unref(); + child.stdout.setEncoding("utf8"); + child.stdout.on("data", (chunk: string) => { + buffer += chunk; + let newline; + while ((newline = buffer.indexOf("\n")) !== -1) { + const line = buffer.slice(0, newline); + buffer = buffer.slice(newline + 1); + if (!line.trim()) continue; + let event: AuthHelperEvent; + try { + event = JSON.parse(line) as AuthHelperEvent; + } catch { + continue; + } + if (event.event === "auth_url") { + clearTimeout(timer); + resolve(event.url); + return; + } + if (event.event === "error") { + clearTimeout(timer); + // The helper relays server-derived text; strip C0/C1 controls + // before it reaches a terminal via the error envelope. + fail(`Sign-in helper failed: ${sanitizeText(event.message)}`); + return; + } + } + }); + // "close", not "exit": exit can fire while the final stdout line + // (e.g. `{"event":"error",...}`) is still buffered; close waits for + // the stdio streams to drain so that line is parsed first. + child.on("close", (code) => { + clearTimeout(timer); + fail( + `Sign-in helper exited (code ${String(code)}) before producing an authorization URL.`, + ); + }); + child.on("error", (error) => { + clearTimeout(timer); + fail(`Failed to spawn sign-in helper: ${error.message}`); + }); + }); + } finally { + // Release the helper: close our ends of its pipes and drop it from this + // process's ref graph so `connect` can exit while it keeps waiting. + child.stdout.destroy(); + child.unref(); + } +} diff --git a/clients/mcpdo/src/connection/auth-names.ts b/clients/mcpdo/src/connection/auth-names.ts new file mode 100644 index 0000000000..eaf5350816 --- /dev/null +++ b/clients/mcpdo/src/connection/auth-names.ts @@ -0,0 +1,159 @@ +/** + * Friendly-name resolution for the URL-keyed OAuth store. + * + * The OAuth store is identified by server URL (the credential *is* the URL, and + * `oauth.json` is shared by every client sharing a config dir). Catalog entries + * (`mcp.json`, per-shell) and live daemon connections (daemon-global) each carry + * a user-chosen *name* for the same URL, so this module builds the two-way map + * between them: `auth/list` annotates each stored URL with the names it is + * `knownAs`, and `auth/clear` accepts a name as an alternate to the raw URL. + * + * Both functions here are pure — the caller fetches the catalog entries + * (`listServerEntries`) and the daemon's connection list (`connections/list`) + * and passes them in, so this stays trivially unit-testable with no daemon or + * filesystem. URL canonicalisation is the *same* `normalizeServerUrl` the store + * resolves keys with, so a catalog `https://api.example.com/mcp` and the store's + * `new URL(...).href` key line up. + */ +import type { ServerListEntry } from "@inspector/core/cli/handlers/servers-list.js"; +import type { ConnectionInfo } from "../daemon/protocol.js"; +import { normalizeServerUrl } from "./stored-auth.js"; + +/** A friendly name that points at a stored server URL. */ +export type AuthNameRef = { + name: string; + /** Where the name came from — a catalog entry or a live daemon connection. */ + source: "catalog" | "connection"; + /** True when this name is a currently-open connection. */ + isLive: boolean; +}; + +export type AuthNameIndex = { + /** Normalised URL → the friendly names that resolve to it (0..n). */ + urlToNames: Map<string, AuthNameRef[]>; + /** Friendly name → the normalised URL(s) it maps to (usually exactly one). */ + nameToUrls: Map<string, Set<string>>; + /** + * Every catalog/connection name seen, *including* stdio ones that have no + * URL — lets `auth/clear <name>` tell "known server, just no stored auth" + * (stdio) from "no such name" (typo). + */ + knownNames: Set<string>; +}; + +/** Only http(s) identities have a URL-keyed OAuth entry. */ +function urlIfHttp(value: string | undefined): string | undefined { + if (value && /^https?:\/\//i.test(value)) return normalizeServerUrl(value); + return undefined; +} + +/** + * Build the two-way name↔URL index from catalog entries and live connections. + * A connection and a catalog entry that share a name and URL collapse to one + * `AuthNameRef` (with `isLive: true`) rather than listing the name twice. + */ +export function buildAuthNameIndex( + entries: ServerListEntry[], + connections: ConnectionInfo[], +): AuthNameIndex { + const urlToNames = new Map<string, AuthNameRef[]>(); + const nameToUrls = new Map<string, Set<string>>(); + const knownNames = new Set<string>(); + + const add = ( + name: string, + url: string | undefined, + source: AuthNameRef["source"], + isLive: boolean, + ): void => { + knownNames.add(name); + if (!url) return; + let urls = nameToUrls.get(name); + if (!urls) { + urls = new Set<string>(); + nameToUrls.set(name, urls); + } + urls.add(url); + + let refs = urlToNames.get(url); + if (!refs) { + refs = []; + urlToNames.set(url, refs); + } + const existing = refs.find((r) => r.name === name); + if (existing) { + // Same (name, url) from both catalog and connection: keep one ref and + // let the live flag win, since a live connection is the stronger fact. + if (isLive) existing.isLive = true; + return; + } + refs.push({ name, source, isLive }); + }; + + // Catalog entries: `detail` is the URL for sse/streamable-http, a command + // line for stdio (no URL-keyed entry). + for (const entry of entries) { + add(entry.name, urlIfHttp(entry.detail), "catalog", false); + } + + // Live connections (daemon-global) catch ad-hoc connects the per-shell + // catalog never had an entry for. + for (const conn of connections) { + add(conn.name, urlIfHttp(conn.serverIdentity), "connection", true); + } + + return index(urlToNames, nameToUrls, knownNames); +} + +/** Sort each URL's names (live first, then alphabetical) for stable output. */ +function index( + urlToNames: Map<string, AuthNameRef[]>, + nameToUrls: Map<string, Set<string>>, + knownNames: Set<string>, +): AuthNameIndex { + for (const refs of urlToNames.values()) { + refs.sort( + (a, b) => + Number(b.isLive) - Number(a.isLive) || a.name.localeCompare(b.name), + ); + } + return { urlToNames, nameToUrls, knownNames }; +} + +/** The outcome of resolving a friendly name to a store URL. */ +export type AuthNameResolution = + | { kind: "url"; url: string } + | { kind: "no-url"; name: string } + | { kind: "ambiguous"; name: string; urls: string[] } + | { kind: "unknown"; name: string }; + +/** + * Resolve a friendly (non-URL) `auth/clear` argument against the index. + * + * - `url`: the name maps to exactly one store URL — clear it. + * - `no-url`: a known catalog/connection name with no URL (stdio) — nothing to + * clear, but not an error worth a non-zero typo message. + * - `ambiguous`: the name maps to two *different* URLs (a catalog entry and a + * live ad-hoc connection that share a name but point elsewhere) — the one + * genuine collision; ask for the explicit URL. + * - `unknown`: no such name anywhere. + * + * Two names sharing one URL is *not* ambiguous: both resolve to the same + * credential and clearing is idempotent. + */ +export function resolveFriendlyName( + name: string, + index: AuthNameIndex, +): AuthNameResolution { + const urls = index.nameToUrls.get(name); + if (urls && urls.size === 1) { + return { kind: "url", url: [...urls][0] }; + } + if (urls && urls.size > 1) { + return { kind: "ambiguous", name, urls: [...urls].sort() }; + } + if (index.knownNames.has(name)) { + return { kind: "no-url", name }; + } + return { kind: "unknown", name }; +} diff --git a/clients/mcpdo/src/connection/authorize.ts b/clients/mcpdo/src/connection/authorize.ts new file mode 100644 index 0000000000..6aff3f2a38 --- /dev/null +++ b/clients/mcpdo/src/connection/authorize.ts @@ -0,0 +1,154 @@ +import { MutableRedirectUrlProvider } from "@inspector/core/auth/index.js"; +import { NodeOAuthStorage } from "@inspector/core/auth/node/index.js"; +import { + DEFAULT_RUNNER_OAUTH_CALLBACK_URL, + formatRunnerOAuthRedirectUrl, + parseRunnerOAuthCallbackUrl, +} from "@inspector/core/auth/node/runner-oauth-callback.js"; +import { + buildRunnerClientAuthOptions, + isOAuthCapableServerConfig, + loadRunnerClientConfig, +} from "@inspector/core/client/runner.js"; +import { InspectorClient } from "@inspector/core/mcp/index.js"; +import { createTransportNode } from "@inspector/core/mcp/node/index.js"; +import { + eraToVersionNegotiation, + type InspectorClientEnvironment, + type InspectorServerSettings, + type MCPServerConfig, +} from "@inspector/core/mcp/types.js"; +import { readInspectorVersion } from "@inspector/core/node/version.js"; +import { createCliOAuthNavigation } from "@inspector/core/cli/cli-oauth-navigation.js"; +import { connectInspectorWithOAuth } from "@inspector/core/cli/cliOAuth.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { isEmaClientNotConfiguredError } from "@inspector/core/auth/ema/clientConfigError.js"; +import type { CallbackNavigation } from "@inspector/core/auth/index.js"; +import type { CliOAuthAutoOpenControl } from "@inspector/core/cli/cli-oauth-navigation.js"; +import { mcpdoEmaGuidance } from "./ema.js"; + +/** + * Run interactive (or stored-auth-only) OAuth in the front-end process so tokens + * land in the shared `oauth.json` store, then the daemon can reconnect. + * + * `makeNavigation` overrides how the authorize URL is surfaced once the + * flow's interactive window arms it: the default prints the relay-worded + * prompt line; the detached auth helper injects a navigation that reports + * the raw URL over its stdout pipe instead (see auth-helper.ts). + */ +export async function authorizeInFrontend( + serverConfig: MCPServerConfig, + serverSettings: InspectorServerSettings | undefined, + options?: { + storedAuthOnly?: boolean; + makeNavigation?: ( + autoOpenControl: CliOAuthAutoOpenControl, + ) => CallbackNavigation; + }, +): Promise<void> { + if (!isOAuthCapableServerConfig(serverConfig)) { + return; + } + + const environment: InspectorClientEnvironment = { + transport: createTransportNode, + }; + const redirectUrlProvider = new MutableRedirectUrlProvider(); + const callbackUrlConfig = parseRunnerOAuthCallbackUrl( + process.env.MCP_OAUTH_CALLBACK_URL ?? DEFAULT_RUNNER_OAUTH_CALLBACK_URL, + ); + redirectUrlProvider.redirectUrl = + formatRunnerOAuthRedirectUrl(callbackUrlConfig); + // Disarmed until connectInspectorWithOAuth's own interactive-OAuth window + // runs — mirrors the one-shot CLI's autoOpenControl (clients/cli/src/cli.ts): + // SDK `auth()` during plain connect() must not print/open before that + // window (or --stored-auth-only) gates it. + const autoOpenControl = { armed: false }; + environment.oauth = { + storage: new NodeOAuthStorage(), + // mcpdo always attempts interactive OAuth (see the isTTY override below) — + // whoever is running it (human or agent) may not have a real TTY on + // stdin/stderr. Reword the printed line so an agent knows it must relay + // the link to a human rather than treating "Please navigate to" as + // addressed to itself. + navigation: options?.makeNavigation + ? options.makeNavigation(autoOpenControl) + : createCliOAuthNavigation({ + autoOpenControl, + disableAutoOpen: options?.storedAuthOnly, + promptMessage: (hrefDisplay, tty) => + tty + ? `Please navigate to: ${hrefDisplay}` + : `The user needs to navigate to this link to authenticate: ${hrefDisplay}`, + }), + redirectUrlProvider, + }; + + const clientConfig = await loadRunnerClientConfig({}); + const clientAuthOptions = buildRunnerClientAuthOptions( + clientConfig, + serverSettings, + {}, + ); + + const client = new InspectorClient(serverConfig, { + environment, + clientIdentity: { + name: "inspector-cli", + version: readInspectorVersion(import.meta.url), + }, + initialLoggingLevel: "debug", + progress: false, + sample: false, + elicit: false, + serverSettings, + ...(serverSettings?.protocolEra && { + versionNegotiation: eraToVersionNegotiation(serverSettings.protocolEra), + }), + ...clientAuthOptions, + }); + + try { + await connectInspectorWithOAuth( + client, + serverConfig, + redirectUrlProvider, + callbackUrlConfig, + serverSettings, + { + storedAuthOnly: options?.storedAuthOnly, + // mcpdo runs as a front-end for whatever invoked it (human terminal or + // agent subprocess) — always admit interactive OAuth rather than + // refusing when stdin/stderr aren't a real TTY. The CI-hang concern + // behind that gate (see clients/cli/README.md OAuth section) doesn't + // apply here: an agent without a TTY is still expected to relay the + // printed URL to an attended human, not run unattended. --stored-auth-only + // (checked above assertInteractiveOAuthAllowed, so unaffected by this) + // remains the way to opt out of interactive OAuth entirely. + isTTY: true, + autoOpenControl, + }, + ); + } catch (err) { + // An EMA server without active install-level IdP config: interactive + // OAuth cannot fix this, so replace the core error (which points at the + // web Client Settings dialog only) with mcpdo-appropriate guidance. + if (isEmaClientNotConfiguredError(err)) { + throw new CliExitCodeError( + EXIT_CODES.AUTH_REQUIRED, + mcpdoEmaGuidance(err.reason), + { code: "auth_required" }, + ); + } + throw err; + } finally { + try { + await client.disconnect(); + } catch { + // best-effort + } + } +} diff --git a/clients/mcpdo/src/connection/dispatch.ts b/clients/mcpdo/src/connection/dispatch.ts new file mode 100644 index 0000000000..1d4d92c021 --- /dev/null +++ b/clients/mcpdo/src/connection/dispatch.ts @@ -0,0 +1,252 @@ +import { callDaemon, ensureDaemon, streamDaemon } from "../daemon/index.js"; +import type { RpcParams, RpcResult } from "../daemon/protocol.js"; +import type { + CliAppInfo, + MethodArgs, +} from "@inspector/core/cli/handlers/method-types.js"; +import type { OutputFormat } from "@inspector/core/cli/handlers/format-output.js"; +import { writeConnectionOutput } from "./format-connection.js"; +import { styleFromOpts, type Style } from "@inspector/core/cli/style.js"; +import { promptElicitation } from "./elicitation-prompt.js"; + +const STREAM_METHODS = new Set(["logging/tail", "resources/subscribe"]); + +/** + * The only two methods whose NDJSON output is a `--verify` conformance report + * rather than `tools/list --app-info` probe lines. Everything else that ever + * returns `kind: "ndjson"` is the app-info shape, so this is a short + * allow-list rather than the other way round. + */ +const NDJSON_VARIANTS = new Set(["skills/list", "skills/get"]); + +export type ConnectionDispatchOpts = { + format?: OutputFormat; + plain?: boolean; + connection?: string; + requireExplicit: boolean; +}; + +/** + * Run one connection MCP method via daemon `rpc` or `stream`. + */ +export async function dispatchConnectionRpc( + method: string, + methodArgs: MethodArgs, + opts: ConnectionDispatchOpts, +): Promise<void> { + const format: OutputFormat = opts.format ?? "text"; + const style = styleFromOpts({ plain: opts.plain, format }); + const params: RpcParams = { + ...methodArgs, + // `format` stays frontend-only: forwarding it would make the daemon's + // runMethod treat `format: "json"` tool calls as app-info requests and + // issue a hidden extra resources/read whose result we discard. The + // daemon also strips it defensively (see stripConnectionFields). + method, + name: stripAt(opts.connection), + requireExplicit: opts.requireExplicit, + }; + + const { socketPath } = await ensureDaemon(); + + if (STREAM_METHODS.has(method)) { + const ac = new AbortController(); + const onSignal = () => ac.abort(); + process.on("SIGINT", onSignal); + process.on("SIGTERM", onSignal); + // Stream writes are chained and awaited before returning: mcp-bin calls + // process.exit() right after, which truncates a still-pending stdout + // write when output is piped or backpressured. + let writeChain: Promise<void> = Promise.resolve(); + try { + await streamDaemon(params, { + socketPath, + // Core enforces the configured MCP request timeout daemon-side; a + // fixed local deadline would falsely fail stream setups (e.g. a + // subscribe against a slow server) that are still valid. + timeoutMs: 0, + signal: ac.signal, + onData: (data) => { + writeChain = writeChain + .then(() => + writeConnectionOutput( + { format, style }, + { + kind: "stream-event", + data, + }, + ), + ) + // Recover the chain itself, not just observe it: a rejected + // chain would skip every later `.then`, silently dropping all + // subsequent events after one failed write. Write errors stay + // non-fatal, as they were when these writes were + // fire-and-forget. + .catch(() => {}); + // Returning the chain lets streamDaemon pause socket reads until + // the write settles, bounding memory when stdout is slow. + return writeChain; + }, + }); + } finally { + process.off("SIGINT", onSignal); + process.off("SIGTERM", onSignal); + await writeChain.catch(() => {}); + } + return; + } + + const ac = new AbortController(); + const onSignal = () => ac.abort(); + process.on("SIGINT", onSignal); + process.on("SIGTERM", onSignal); + // Interactive callers get inline prompts; everyone else — `--format json` + // (single machine-readable payload) or no TTY at all (an agent's stdin is + // not wired to the human, so a prompt would hang until auto-cancel) — has + // the daemon park the elicitation and answers via `elicitation/respond`. + const interactive = + format === "text" && + (process.stdin.isTTY === true || process.stderr.isTTY === true); + params.interactive = interactive; + let outcome: RpcResult; + try { + outcome = await callDaemon<RpcResult>("rpc", params, { + socketPath, + // Core enforces the configured MCP request timeout daemon-side; a + // fixed local deadline would falsely fail long-running tool calls. + timeoutMs: 0, + signal: ac.signal, + onElicitation: (frame) => + promptElicitation(frame, { + style, + // Prompting only needs a readable stdin and a text-based reply + // channel; parked (non-interactive) callers never receive frames, + // so this only ever runs interactively. + interactive, + }), + }); + } finally { + process.off("SIGINT", onSignal); + process.off("SIGTERM", onSignal); + } + await writeRpcOutcome( + { format, style }, + method, + methodArgs.toolName, + outcome, + ); +} + +/** + * Render one rpc outcome — final result, NDJSON report, or a parked + * elicitation round. Shared by the originating call above and by + * `elicitation/respond`, whose result is the same shape (the resumed call's + * outcome or the next round). + */ +export async function writeRpcOutcome( + out: { format?: OutputFormat; style: Style }, + method: string, + toolName: string | undefined, + outcome: RpcResult, +): Promise<void> { + const { format, style } = out; + if (outcome.kind === "elicitation-pending") { + await writeConnectionOutput( + { format, style }, + { kind: "elicitation-pending", elicitation: outcome.elicitation }, + ); + return; + } + if (outcome.kind === "ndjson") { + await writeConnectionOutput( + { format, style }, + { + kind: "ndjson", + lines: outcome.lines, + variant: NDJSON_VARIANTS.has(method) ? "skill-verify" : "app-info", + summary: outcome.summary, + exitCode: outcome.exitCode, + }, + ); + return; + } + await writeConnectionOutput( + { format, style }, + { + kind: "rpc", + method, + result: outcome.result, + appInfo: outcome.appInfo as CliAppInfo | undefined, + toolName, + }, + ); +} + +export function stripAt(name: string | undefined): string | undefined { + if (!name) return undefined; + return name.startsWith("@") ? name.slice(1) : name; +} + +/** + * Non-interactive runs must pass an explicit connection for MRU-targeting ops. + * Key off stdin (not stdout) so piping output (`mcpdo tools/list | jq`) still + * uses MRU when a human is at the keyboard. + */ +export function requireExplicitConnection(): boolean { + if (process.env.MCP_ALLOW_DEFAULT_CONNECTION === "1") return false; + return process.stdin.isTTY !== true; +} + +/** + * Global options that take a value, as registered on the program in `runMcp` + * (after `expandConnAlias` has rewritten `--conn` to `--connection`). The + * hoist below needs this to know that the token after `--format` is its value + * rather than the subcommand. + */ +const VALUE_TAKING_GLOBALS = new Set([ + "--format", + "--connection", + "--catalog", + "--config", +]); + +const AT_CONNECTION_RE = /^@[A-Za-z0-9_.-]+$/; + +/** + * Hoist an `@name` connection token from argv so `mcpdo @alpha tools/list` + * works. The token is recognised anywhere before the subcommand — the README + * documents global options as going before the subcommand, so + * `mcpdo --format json @alpha tools/list` must work too, not only a leading + * `@alpha`. Scanning stops at the subcommand (the first token that is neither + * an option, an option's value, nor the `@name` itself) and at `--`, so a + * positional argument that happens to start with `@` is never claimed. + */ +export function hoistAtConnection(argv: string[]): { + argv: string[]; + connectionFromAt?: string; +} { + const start = 2; + for (let i = start; i < argv.length; i++) { + const token = argv[i]!; + if (token === "--") break; + if (AT_CONNECTION_RE.test(token)) { + return { + argv: [...argv.slice(0, i), ...argv.slice(i + 1)], + connectionFromAt: token.slice(1), + }; + } + if (token.startsWith("-")) { + // `--opt=value` carries its value inline; a value-taking global + // consumes the next token. Any other option (boolean globals, -h) is a + // single token. An option this table doesn't know is treated as + // boolean, which at worst stops the scan early at its value — never + // claims one as a connection. + if (!token.includes("=") && VALUE_TAKING_GLOBALS.has(token)) i++; + continue; + } + // First non-option token is the subcommand: @name past this point is a + // positional argument, not a connection selector. + break; + } + return { argv }; +} diff --git a/clients/mcpdo/src/connection/elicitation-prompt.ts b/clients/mcpdo/src/connection/elicitation-prompt.ts new file mode 100644 index 0000000000..c21d9d415e --- /dev/null +++ b/clients/mcpdo/src/connection/elicitation-prompt.ts @@ -0,0 +1,177 @@ +/** + * Terminal UI for a mid-`rpc` elicitation exchange (dual-era support). Covers + * both delivery mechanisms (legacy server→client request, modern non-task + * MRTR round) and both modes: + * + * - **URL mode** mirrors the web client's convention + * (`InlineElicitationRequest`/`PendingClientRequestModal`): the actual + * out-of-band completion can't be observed here, so the user self-reports + * it by answering a confirm prompt — there is no "decline" for URL mode, + * only accept (they say they finished) or cancel. + * - **Form mode** renders one prompt per field from the schema (see + * `form-schema.ts`/`form-prompt.ts`), with a review step before submitting. + * Schemas outside the spec's restricted primitive-field shape (should + * never happen from a well-behaved server) fall back to a clear decline. + */ +import type { Style } from "@inspector/core/cli/style.js"; +import type { + ElicitationRequestFrame, + ElicitationResponseFrame, +} from "../daemon/protocol.js"; +import { parseFormSchema } from "./form-schema.js"; +import { promptForm, watchForClose } from "./form-prompt.js"; +import { getSharedPromptReader } from "./prompt-reader.js"; +import { isSafeLinkTarget, sanitizeText } from "./sanitize.js"; + +export type PromptElicitationOpts = { + /** + * False only for callers where a text prompt can't sensibly be shown + * (currently just `--format json`, whose stdout is a single + * machine-readable payload). A prompt works the same over a plain, + * non-TTY stdin/stderr as it does at a real terminal — a human at a + * keyboard and an agent relaying/answering on their behalf both just + * read a line of text and reply with one. A stdin that's already closed + * (e.g. `mcpdo ... </dev/null`) is handled by declining/cancelling once + * reading fails, not by refusing to try in the first place. + */ + interactive: boolean; + style: Style; +}; + +function cancelResponse( + frame: ElicitationRequestFrame, +): ElicitationResponseFrame { + return { + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "cancel", + }; +} + +function declineResponse( + frame: ElicitationRequestFrame, +): ElicitationResponseFrame { + return { + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "decline", + }; +} + +/** + * Prompt the user for one elicitation exchange and return their answer. + * Never throws — falls back to cancel/decline on any failure to read input, + * so a mid-prompt problem always unblocks the daemon-side call. + */ +export async function promptElicitation( + frame: ElicitationRequestFrame, + opts: PromptElicitationOpts, +): Promise<ElicitationResponseFrame> { + const { style } = opts; + // Server-controlled display strings must not reach the terminal raw + // (escape injection — see sanitize.ts). Protocol ids on `frame` stay + // untouched so responses still correlate. + const message = sanitizeText(frame.message); + const url = frame.url === undefined ? undefined : sanitizeText(frame.url); + + if (frame.mode === "form") { + // The schema is parsed raw: sanitizing it wholesale would mutate protocol + // data (property names, enum values, defaults), so the accepted response + // could carry keys/values the server never defined. Server-controlled + // strings are instead sanitized at each render point in form-prompt.ts. + const fields = parseFormSchema(frame.requestedSchema); + if (!fields) { + // Schema outside the spec's restricted primitive-field shape — + // shouldn't happen from a well-behaved server; decline clearly rather + // than silently guessing at field values. + process.stderr.write( + style.yellow( + "This server's form request uses a schema mcpdo doesn't support " + + "— declining.\n", + ) + ` ${message}\n`, + ); + return declineResponse(frame); + } + + if (!opts.interactive) { + process.stderr.write( + style.yellow( + "This server is asking for form input, which isn't supported " + + "with --format json — declining.\n", + ) + ` ${message}\n`, + ); + return declineResponse(frame); + } + + // The shared reader outlives this exchange on purpose: piped answers + // for later fields/rounds arrive before their questions are asked, and + // a per-exchange interface would drop them (see prompt-reader.ts). + const rl = getSharedPromptReader(); + try { + const outcome = await promptForm(rl, message, fields, style); + if (outcome.action === "accept") { + return { + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "accept", + content: outcome.content, + }; + } + if (outcome.action === "decline") return declineResponse(frame); + return cancelResponse(frame); + } catch { + return cancelResponse(frame); + } + } + + if (!opts.interactive) { + process.stderr.write( + style.yellow( + "This server is asking for input via a URL (elicitation), which " + + "isn't supported with --format json — cancelling.\n", + ) + + ` ${message}\n` + + (url ? ` ${url}\n` : ""), + ); + return cancelResponse(frame); + } + + process.stderr.write( + "\n" + + style.bold("Action required: ") + + message + + "\n" + + " " + + // Only allowlisted schemes render as a clickable OSC 8 link; a server + // supplying file:/custom-handler URLs gets plain text (see sanitize.ts). + (url !== undefined && isSafeLinkTarget(url) + ? style.link(url, url) + : (url ?? "")) + + "\n\n", + ); + + const rl = getSharedPromptReader(); + try { + const answer = await Promise.race([ + rl.question( + "Open the URL above, complete it, then press Enter to continue " + + "(or type 'c' to cancel): ", + ), + watchForClose(rl), + ]); + if (answer.trim().toLowerCase() === "c") { + return cancelResponse(frame); + } + return { + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "accept", + }; + } catch { + return cancelResponse(frame); + } +} diff --git a/clients/mcpdo/src/connection/ema-login-helper.ts b/clients/mcpdo/src/connection/ema-login-helper.ts new file mode 100644 index 0000000000..cc63adf901 --- /dev/null +++ b/clients/mcpdo/src/connection/ema-login-helper.ts @@ -0,0 +1,146 @@ +/** + * Non-TTY park path for `auth/ema-login`, mirroring the connect-time OAuth + * park in auth-helper.ts: instead of blocking for up to 15 minutes on the + * loopback callback (printing the IdP URL to a stderr no agent relays), the + * parent spawns a detached helper that owns the callback wait, reads the IdP + * authorization URL off its stdout, and exits 0 with that URL in the result + * payload so the agent can relay it and poll `auth/ema-status`. + * + * This lives in its own module (not ema.ts, not auth-helper.ts) because it + * imports from both and neither imports it — keeping the dependency graph + * acyclic. + */ +import { + clearEmaIdpSession, + getEmaIdpLoginState, + normalizeIdpIssuer, +} from "@inspector/core/auth/ema/index.js"; +import { CallbackNavigation } from "@inspector/core/auth/index.js"; +import { NodeOAuthStorage } from "@inspector/core/auth/node/index.js"; +import { + obtainPendingUrlForKey, + pendingAuthMarkerPath, + removeOwnPendingAuthMarker, + writePendingAuthMarker, + PENDING_AUTH_TTL_MS, + type AuthHelperEvent, +} from "./auth-helper.js"; +import { + loadEmaIdpConfig, + requireIdp, + runEmaIdpInteractiveFlow, + type EmaLoginResult, +} from "./ema.js"; + +/** Hidden subcommand the detached EMA login helper runs as. */ +export const EMA_LOGIN_HELPER_COMMAND = "auth/complete-ema-login"; + +/** + * Marker key for the pending-login marker: per IdP issuer, not per server — + * leg 1 is server-less, and every EMA server behind the same issuer shares + * the one IdP session a login establishes. + */ +export function emaLoginMarkerKey(issuer: string): string { + return `ema-idp:${issuer}`; +} + +export type PendingEmaLoginResult = EmaLoginResult & { + /** Present (true) when the login was parked on a detached helper. */ + pendingLogin?: true; + /** IdP authorization URL for the agent to relay to the human. */ + authUrl?: string; +}; + +/** + * Non-TTY `auth/ema-login`: short-circuit if already logged in (or clear the + * session on --relogin), then park the interactive IdP flow on a detached + * helper and return its authorization URL. The helper survives this process + * and completes the code exchange when the user finishes signing in; callers + * poll `auth/ema-status` for `loginState: "logged_in"`. + */ +export async function startPendingEmaLogin(options?: { + relogin?: boolean; + helperArgv1?: string; + waitMs?: number; + pollMs?: number; +}): Promise<PendingEmaLoginResult> { + const { idp, enabled } = await loadEmaIdpConfig(); + const active = requireIdp(idp, enabled); + const issuer = normalizeIdpIssuer(active.issuer); + const storage = new NodeOAuthStorage(); + + if (options?.relogin) { + await clearEmaIdpSession(storage, active.issuer); + } else if ( + (await getEmaIdpLoginState(storage, active.issuer)) === "logged_in" + ) { + return { issuer, loginState: "logged_in", alreadyLoggedIn: true }; + } + + const authUrl = await obtainPendingUrlForKey( + emaLoginMarkerKey(issuer), + EMA_LOGIN_HELPER_COMMAND, + // The helper re-reads config itself; params document is empty. + {}, + options, + ); + + return { + issuer, + loginState: await getEmaIdpLoginState(storage, active.issuer), + alreadyLoggedIn: false, + pendingLogin: true, + authUrl, + }; +} + +/** + * Entry point for the hidden helper subcommand. Runs the full IdP OIDC flow + * with a navigation that publishes the authorization URL as a pending-login + * marker plus an NDJSON event on stdout (instead of a prompt line), and + * never opens a browser — the human the URL is relayed to does that. + */ +export async function runEmaLoginHelper(): Promise<void> { + // The parent unrefs and exits once it has the URL; every later stdout + // write would EPIPE without this guard. + const emit = (event: AuthHelperEvent) => { + try { + process.stdout.write(`${JSON.stringify(event)}\n`); + } catch { + // Parent is gone; the flow itself is unaffected. + } + }; + process.stdout.on("error", () => {}); + + let markerPath: string | undefined; + try { + const { idp, enabled } = await loadEmaIdpConfig(); + const active = requireIdp(idp, enabled); + const issuer = normalizeIdpIssuer(active.issuer); + const storage = new NodeOAuthStorage(); + await runEmaIdpInteractiveFlow( + active, + storage, + new CallbackNavigation(async (url) => { + // No arming guard needed: unlike connect-time OAuth there is no + // SDK-internal auth() phase — this flow owns its one authorize URL. + markerPath = pendingAuthMarkerPath(emaLoginMarkerKey(issuer)); + writePendingAuthMarker(markerPath, { + url: url.href, + pid: process.pid, + expiresAt: Date.now() + PENDING_AUTH_TTL_MS, + }); + emit({ event: "auth_url", url: url.href }); + }), + ); + emit({ event: "done" }); + } catch (error) { + emit({ + event: "error", + message: error instanceof Error ? error.message : String(error), + }); + throw error; + } finally { + if (markerPath !== undefined) removeOwnPendingAuthMarker(markerPath); + } +} diff --git a/clients/mcpdo/src/connection/ema.ts b/clients/mcpdo/src/connection/ema.ts new file mode 100644 index 0000000000..89e11683d0 --- /dev/null +++ b/clients/mcpdo/src/connection/ema.ts @@ -0,0 +1,270 @@ +import { + clearEmaIdpSession, + getEmaIdpLoginState, + normalizeIdpIssuer, + type EmaIdpLoginState, +} from "@inspector/core/auth/ema/index.js"; +import type { EmaClientNotConfiguredReason } from "@inspector/core/auth/ema/clientConfigError.js"; +import { + completeIdpOidcAuthorization, + startIdpOidcAuthorization, +} from "@inspector/core/auth/ema/idpOidc.js"; +import { MutableRedirectUrlProvider } from "@inspector/core/auth/index.js"; +import type { CallbackNavigation } from "@inspector/core/auth/index.js"; +import { + NodeOAuthStorage, + runRunnerInteractiveOAuth, +} from "@inspector/core/auth/node/index.js"; +import { resetNodeOAuthStorageCache } from "@inspector/core/auth/node/storage-node.js"; +import { + DEFAULT_RUNNER_OAUTH_CALLBACK_URL, + formatRunnerOAuthRedirectUrl, + parseRunnerOAuthCallbackUrl, +} from "@inspector/core/auth/node/runner-oauth-callback.js"; +import { getClientConfigFilePath } from "@inspector/core/client/index.js"; +import { loadRunnerClientConfig } from "@inspector/core/client/runner.js"; +import type { EnterpriseManagedAuthIdpConfig } from "@inspector/core/client/types.js"; +import { createCliOAuthNavigation } from "@inspector/core/cli/cli-oauth-navigation.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; + +/** Where install-level EMA IdP config lives (honours MCP_CLIENT_CONFIG_PATH). */ +function clientConfigPath(): string { + return getClientConfigFilePath( + process.env.MCP_CLIENT_CONFIG_PATH?.trim() || undefined, + ); +} + +/** + * mcpdo-flavoured guidance for a missing/disabled EMA client configuration. + * The core `EmaClientNotConfiguredError` message points at the web Client + * Settings dialog; mcpdo users may equally well edit `client.json` directly, + * so name both, with the resolved path. + */ +export function mcpdoEmaGuidance(reason: EmaClientNotConfiguredReason): string { + const path = clientConfigPath(); + if (reason === "disabled") { + return ( + "Enterprise-managed auth (EMA) is configured but disabled. Enable it in " + + "the web Inspector's Client Settings, or set " + + `enterpriseManagedAuth.enabled to true in ${path}.` + ); + } + return ( + "Enterprise-managed auth (EMA) is not configured. Configure the " + + "enterprise IdP (issuer, client ID, client secret) in the web Inspector's " + + `Client Settings, or add an enterpriseManagedAuth block to ${path}.` + ); +} + +export type EmaStatus = { + /** Resolved client.json path the config was read from. */ + clientConfigPath: string; + /** An IdP block exists in client.json (even if disabled). */ + configured: boolean; + /** Configured and not explicitly disabled. */ + enabled: boolean; + issuer?: string; + clientId?: string; + /** IdP session state; "unconfigured" when no IdP block exists. */ + loginState: EmaIdpLoginState | "unconfigured"; +}; + +/** Read install-level EMA config; the raw idp block, even when disabled. */ +export async function loadEmaIdpConfig(): Promise<{ + idp: EnterpriseManagedAuthIdpConfig | undefined; + enabled: boolean; +}> { + const clientConfig = await loadRunnerClientConfig({}); + const ema = clientConfig.enterpriseManagedAuth; + return { + idp: ema?.idp, + enabled: Boolean(ema?.idp) && ema?.enabled !== false, + }; +} + +export function requireIdp( + idp: EnterpriseManagedAuthIdpConfig | undefined, + enabled: boolean, + options?: { allowDisabled?: boolean }, +): EnterpriseManagedAuthIdpConfig { + if (!idp) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + mcpdoEmaGuidance("not_configured"), + { + code: "usage", + }, + ); + } + if (!enabled && !options?.allowDisabled) { + throw new CliExitCodeError(EXIT_CODES.USAGE, mcpdoEmaGuidance("disabled"), { + code: "usage", + }); + } + return idp; +} + +/** EMA configuration + IdP session state for `auth/ema-status`. */ +export async function getEmaStatus(): Promise<EmaStatus> { + const { idp, enabled } = await loadEmaIdpConfig(); + if (!idp) { + return { + clientConfigPath: clientConfigPath(), + configured: false, + enabled: false, + loginState: "unconfigured", + }; + } + const storage = new NodeOAuthStorage(); + const loginState = await getEmaIdpLoginState(storage, idp.issuer); + return { + clientConfigPath: clientConfigPath(), + configured: true, + enabled, + issuer: normalizeIdpIssuer(idp.issuer), + clientId: idp.clientId, + loginState, + }; +} + +export type EmaLogoutResult = { + issuer: string; + /** + * OIDC RP-initiated logout URL, when the IdP advertises one. The local + * clear cannot end the IdP's browser SSO session; navigating here does. + */ + endSessionUrl?: string; +}; + +/** + * Sign out of the enterprise IdP locally: clears the cached IdP OIDC + * connection and all EMA-minted resource-server tokens. The IdP's own + * browser session is untouched — when the IdP advertises an + * `end_session_endpoint`, the returned `endSessionUrl` lets the user end it + * too. Works even when EMA is disabled (state cleanup should never be + * blocked by the enabled flag). + */ +export async function emaLogout(): Promise<EmaLogoutResult> { + const { idp, enabled } = await loadEmaIdpConfig(); + const active = requireIdp(idp, enabled, { allowDisabled: true }); + const storage = new NodeOAuthStorage(); + const { endSessionUrl } = await clearEmaIdpSession(storage, active.issuer); + resetNodeOAuthStorageCache(); + return { + issuer: normalizeIdpIssuer(active.issuer), + ...(endSessionUrl === undefined ? {} : { endSessionUrl }), + }; +} + +export type EmaLoginResult = { + issuer: string; + loginState: EmaIdpLoginState; + alreadyLoggedIn: boolean; +}; + +/** + * Sign in to the enterprise IdP (EMA leg 1 only — no server required): print + * the IdP authorization URL, wait on the loopback callback, and exchange the + * code for an IdP session. Subsequent connects to EMA servers mint resource + * tokens silently from this connection. + * + * Non-TTY (agent-attended) callers get wording that directs the agent to + * relay the link to the human user, mirroring `authorizeInFrontend`. SIGINT / + * SIGTERM and the callback timeout are handled by + * {@link runRunnerInteractiveOAuth}. + */ +export async function emaLogin(options?: { + /** Clear any existing IdP session (and EMA server tokens) first. */ + relogin?: boolean; +}): Promise<EmaLoginResult> { + const { idp, enabled } = await loadEmaIdpConfig(); + const active = requireIdp(idp, enabled); + const issuer = normalizeIdpIssuer(active.issuer); + const storage = new NodeOAuthStorage(); + + if (options?.relogin) { + await clearEmaIdpSession(storage, active.issuer); + } else if ( + (await getEmaIdpLoginState(storage, active.issuer)) === "logged_in" + ) { + return { issuer, loginState: "logged_in", alreadyLoggedIn: true }; + } + + // Armed from the start: unlike connect-time OAuth there is no SDK-internal + // auth() phase to guard against — this flow owns its one authorize URL. + const navigation = createCliOAuthNavigation({ + autoOpenControl: { armed: true }, + promptMessage: (hrefDisplay, tty) => + tty + ? `Sign in to your enterprise IdP: ${hrefDisplay}` + : "The user needs to sign in to the enterprise identity provider " + + `(IdP) at this link: ${hrefDisplay}`, + }); + await runEmaIdpInteractiveFlow(active, storage, navigation); + + return { + issuer, + loginState: await getEmaIdpLoginState(storage, active.issuer), + alreadyLoggedIn: false, + }; +} + +/** + * Run EMA leg 1 (the IdP OIDC authorization-code flow) to completion: + * loopback callback server, 15-minute wait, SIGINT/SIGTERM cancellation. + * `navigation` decides how the authorization URL reaches the user — the + * interactive path above prints a prompt line (and may auto-open a browser); + * the detached non-TTY helper reports it over its stdout pipe instead (see + * ema-login-helper.ts). + */ +export async function runEmaIdpInteractiveFlow( + active: EnterpriseManagedAuthIdpConfig, + storage: NodeOAuthStorage, + navigation: Pick<CallbackNavigation, "navigateToAuthorization">, +): Promise<void> { + const callbackUrlConfig = parseRunnerOAuthCallbackUrl( + process.env.MCP_OAUTH_CALLBACK_URL ?? DEFAULT_RUNNER_OAUTH_CALLBACK_URL, + ); + const redirectUrlProvider = new MutableRedirectUrlProvider(); + redirectUrlProvider.redirectUrl = + formatRunnerOAuthRedirectUrl(callbackUrlConfig); + + // Adapter over the server-bound runner-interactive-OAuth surface: EMA leg 1 + // is server-less, so authenticate/completeOAuthFlow map straight onto the + // IdP OIDC start/complete helpers. This reuses the loopback callback + // server, 15-minute timeout, and SIGINT/SIGTERM cancellation. + await runRunnerInteractiveOAuth({ + client: { + authenticate: async () => { + const { authorizationUrl } = await startIdpOidcAuthorization({ + idp: active, + redirectUrl: redirectUrlProvider.redirectUrl, + storage, + }); + navigation.navigateToAuthorization(authorizationUrl); + return authorizationUrl; + }, + /* v8 ignore next 2 -- only reached when options.authorizationUrl is set, which this flow never does */ + beginInteractiveAuthorization: async () => {}, + completeOAuthFlow: async (authorizationCode, iss) => { + await completeIdpOidcAuthorization({ + idp: active, + authorizationCode, + iss, + redirectUrl: redirectUrlProvider.redirectUrl, + storage, + }); + }, + /* v8 ignore next 2 -- only reached when options.authChallenge is set, which this flow never does */ + checkAuthChallengeSatisfied: async () => false, + }, + redirectUrlProvider, + callbackListen: callbackUrlConfig, + // mcpdo is a plain CLI (no Ink); own Ctrl-C during the IdP wait. + handleSignals: true, + }); + resetNodeOAuthStorageCache(); +} diff --git a/clients/mcpdo/src/connection/form-prompt.ts b/clients/mcpdo/src/connection/form-prompt.ts new file mode 100644 index 0000000000..1ece55eb80 --- /dev/null +++ b/clients/mcpdo/src/connection/form-prompt.ts @@ -0,0 +1,299 @@ +/** + * Interactive terminal renderer for a form-mode elicitation + * (dual-era support, phase 3). Prompts once per field (type-appropriate: + * text, numeric, y/n, numbered single-select, numbered multi-select), + * pre-fills defaults, does light client-side validation (required/length/ + * range), then shows a review step before submitting so the user can + * re-edit any field or cancel outright. + */ +import type { Style } from "@inspector/core/cli/style.js"; +import type { PromptInput } from "./prompt-reader.js"; +import type { FormField } from "./form-schema.js"; +import { codePointLength } from "./form-schema.js"; +import { sanitizeText } from "./sanitize.js"; + +export type FormOutcome = + | { action: "accept"; content: Record<string, unknown> } + | { action: "decline" } + | { action: "cancel" }; + +/** + * A promise that rejects the first time `rl` reports exhausted input (for a + * {@link PromptReader}: EOF on a redirected/piped stdin *with no buffered + * line left to answer from*, or the reader being disposed elsewhere). Racing + * every `rl.question()` against this means a closed-before-answered stdin + * (e.g. `mcpdo ... </dev/null`) falls through to the caller's cancel/decline + * handling instead of hanging forever waiting for a line that will never + * arrive. + */ +export function watchForClose(rl: PromptInput): Promise<never> { + return new Promise((_, reject) => { + rl.once("close", () => + reject(new Error("stdin closed before an answer was given")), + ); + }); +} + +/** `rl.question()`, but rejects instead of hanging if stdin closes first. */ +function ask( + rl: PromptInput, + closed: Promise<never>, + prompt: string, +): Promise<string> { + return Promise.race([rl.question(prompt), closed]); +} + +function formatDefault(field: FormField): string | undefined { + if (field.default === undefined) return undefined; + if (field.kind === "multiselect") { + return (field.default as string[]).join(", "); + } + return String(field.default); +} + +function describeField(field: FormField, style: Style): string { + // Titles, descriptions and defaults are server-controlled: sanitize at the + // render point only, so the raw values still travel in the response. + const req = field.required ? style.yellow(" (required)") : ""; + const desc = field.description ? ` — ${sanitizeText(field.description)}` : ""; + const def = formatDefault(field); + const defHint = + def !== undefined ? style.dim(` [default: ${sanitizeText(def)}]`) : ""; + return `${style.bold(sanitizeText(field.title))}${req}${desc}${defHint}`; +} + +/** Prompts for one field's value; loops until a valid answer or a default/blank-when-optional. */ +async function promptField( + rl: PromptInput, + closed: Promise<never>, + field: FormField, + style: Style, +): Promise<unknown> { + for (;;) { + if (field.kind === "boolean") { + const def = field.default; + const hint = def === undefined ? "y/n" : def ? "Y/n" : "y/N"; + const raw = ( + await ask(rl, closed, `${describeField(field, style)}\n [${hint}]: `) + ) + .trim() + .toLowerCase(); + if (raw === "" && def !== undefined) return def; + if (raw === "y" || raw === "yes") return true; + if (raw === "n" || raw === "no") return false; + if (raw === "" && !field.required) return undefined; + process.stderr.write(style.red(" Please answer y or n.\n")); + continue; + } + + if (field.kind === "enum" || field.kind === "multiselect") { + const lines = field.choices.map( + (choice, i) => ` ${i + 1}. ${sanitizeText(choice.label)}`, + ); + const multi = field.kind === "multiselect"; + const prompt = multi + ? "Enter one or more numbers separated by commas" + : "Enter a number"; + const raw = ( + await ask( + rl, + closed, + `${describeField(field, style)}\n${lines.join("\n")}\n ${prompt}: `, + ) + ).trim(); + if (raw === "") { + if (field.default !== undefined) return field.default; + if (!field.required) return undefined; + if (multi) { + const m = field as Extract<FormField, { kind: "multiselect" }>; + // JSON Schema "required" only means the key must be present; an + // empty array is a valid value unless minItems forbids it. Without + // this, "none selected" on a required multiselect loops forever. + if ((m.minItems ?? 0) === 0) return []; + process.stderr.write(style.red(` Select at least ${m.minItems}.\n`)); + continue; + } + process.stderr.write(style.red(" This field is required.\n")); + continue; + } + // Strict whole-token integers only: parseInt would accept "1abc" as 1, + // silently submitting a different answer than the user typed. + const tokens = raw.split(",").map((s) => s.trim()); + const indices = tokens.map((s) => + /^\d+$/.test(s) ? Number.parseInt(s, 10) : Number.NaN, + ); + if ( + indices.some( + (n) => !Number.isInteger(n) || n < 1 || n > field.choices.length, + ) + ) { + process.stderr.write( + style.red( + ` Enter a number between 1 and ${field.choices.length}.\n`, + ), + ); + continue; + } + if (!multi && indices.length !== 1) { + // "1,2" on a single-select would silently drop everything after the + // first choice — re-prompt instead. + process.stderr.write(style.red(" Enter exactly one number.\n")); + continue; + } + const values = indices.map((n) => field.choices[n - 1]!.value); + if (multi) { + const m = field as Extract<FormField, { kind: "multiselect" }>; + if (m.minItems !== undefined && values.length < m.minItems) { + process.stderr.write(style.red(` Select at least ${m.minItems}.\n`)); + continue; + } + if (m.maxItems !== undefined && values.length > m.maxItems) { + process.stderr.write(style.red(` Select at most ${m.maxItems}.\n`)); + continue; + } + return values; + } + return values[0]; + } + + if (field.kind === "number") { + const def = field.default; + const raw = ( + await ask( + rl, + closed, + `${describeField(field, style)}\n ${def !== undefined ? `[${def}]` : ""}: `, + ) + ).trim(); + if (raw === "") { + if (def !== undefined) return def; + if (!field.required) return undefined; + process.stderr.write(style.red(" This field is required.\n")); + continue; + } + const n = Number(raw); + if ( + // isFinite (not isNaN): "Infinity" is not a valid JSON number and + // would serialize as null in the response frame. + !Number.isFinite(n) || + (field.integer && !Number.isInteger(n)) || + (field.minimum !== undefined && n < field.minimum) || + (field.maximum !== undefined && n > field.maximum) + ) { + const range = + field.minimum !== undefined || field.maximum !== undefined + ? ` (${field.minimum ?? "-∞"}..${field.maximum ?? "∞"})` + : ""; + process.stderr.write( + style.red( + ` Enter a valid ${field.integer ? "integer" : "number"}${range}.\n`, + ), + ); + continue; + } + return n; + } + + // string + const def = field.default; + const raw = await ask( + rl, + closed, + `${describeField(field, style)}\n ${def !== undefined ? `[${sanitizeText(def)}]` : ""}: `, + ); + // Enter with a default selects it — even an empty-string default; the + // schema gate already rejected defaults violating their own constraints. + if (raw === "" && def !== undefined) return def; + // A blank answer with no default omits an optional field. For a required + // field "" is a value — JSON Schema `required` means present, not + // non-empty — so minLength (if any) decides below. + if (raw === "" && !field.required) return undefined; + const value = raw; + // Code points, not UTF-16 units: JSON Schema length semantics. + const length = codePointLength(value); + if (field.minLength !== undefined && length < field.minLength) { + process.stderr.write( + style.red(` Must be at least ${field.minLength} characters.\n`), + ); + continue; + } + if (field.maxLength !== undefined && length > field.maxLength) { + process.stderr.write( + style.red(` Must be at most ${field.maxLength} characters.\n`), + ); + continue; + } + return value; + } +} + +/** + * Collect one value per field, then loop on a review step (submit / edit a + * field by name / cancel) until the user submits or cancels. + */ +export async function promptForm( + rl: PromptInput, + message: string, + fields: FormField[], + style: Style, +): Promise<FormOutcome> { + process.stderr.write(`\n${style.bold("Input requested: ")}${message}\n\n`); + const closed = watchForClose(rl); + + const values = new Map<string, unknown>(); + for (const field of fields) { + values.set(field.name, await promptField(rl, closed, field, style)); + } + + for (;;) { + process.stderr.write(`\n${style.bold("Review your answers:")}\n`); + for (const field of fields) { + const v = values.get(field.name); + // The edit prompt below accepts the schema property *name*; show it + // whenever it differs from the display title so the user can discover + // what to type. + const label = + field.title === field.name + ? field.title + : `${field.title} (${field.name})`; + process.stderr.write( + ` ${sanitizeText(label)}: ${v === undefined ? style.dim("(none)") : sanitizeText(String(v))}\n`, + ); + } + const answer = ( + await ask( + rl, + closed, + "\nPress Enter to submit, type a field name to edit it, or 'c' to cancel: ", + ) + ).trim(); + if (answer === "") { + // Null prototype: a schema is entitled to a property named + // "__proto__", which on a plain object would hit the prototype + // setter instead of creating an own property, silently dropping the + // answer (mirrors sanitizeDeep's handling of untrusted keys). + const content: Record<string, unknown> = Object.create(null) as Record< + string, + unknown + >; + for (const field of fields) { + const v = values.get(field.name); + if (v !== undefined) content[field.name] = v; + } + return { action: "accept", content }; + } + if (answer.toLowerCase() === "c") { + return { action: "cancel" }; + } + const field = + fields.find((f) => f.name === answer) ?? + fields.find((f) => f.title === answer); + if (!field) { + process.stderr.write( + style.red(` Unknown field "${answer}". Try again.\n`), + ); + continue; + } + values.set(field.name, await promptField(rl, closed, field, style)); + } +} diff --git a/clients/mcpdo/src/connection/form-schema.ts b/clients/mcpdo/src/connection/form-schema.ts new file mode 100644 index 0000000000..5f807e6113 --- /dev/null +++ b/clients/mcpdo/src/connection/form-schema.ts @@ -0,0 +1,302 @@ +/** + * Parses a form-mode elicitation `requestedSchema` into a flat list of + * fields mcpdo can prompt for. Per the MRTR elicitation spec (2026-07-28), + * form-mode schemas are restricted to a flat object whose properties are + * primitive types only — string, number/integer, boolean, single-select + * enum (`enum` or titled `oneOf`), or multi-select enum (`array` of one of + * those) — so this never needs to handle nesting, arrays of objects, or + * other general JSON Schema features. + * + * Returns `null` if the schema doesn't match that shape (defensive: a + * well-behaved server never sends anything else, but this is untrusted + * wire input from an arbitrary MCP server). + */ + +export type Choice = { value: string; label: string }; + +type FieldExtra = + | { + kind: "string"; + minLength?: number; + maxLength?: number; + format?: string; + default?: string; + } + | { + kind: "number"; + integer: boolean; + minimum?: number; + maximum?: number; + default?: number; + } + | { kind: "boolean"; default?: boolean } + | { kind: "enum"; choices: Choice[]; default?: string } + | { + kind: "multiselect"; + choices: Choice[]; + minItems?: number; + maxItems?: number; + default?: string[]; + }; + +export type FormField = { + name: string; + required: boolean; + title: string; + description?: string; +} & FieldExtra; + +function isRecord(value: unknown): value is Record<string, unknown> { + return typeof value === "object" && value !== null && !Array.isArray(value); +} + +function parseChoicesFromEnum(value: unknown): Choice[] | undefined { + if ( + !Array.isArray(value) || + value.length === 0 || + value.some((v) => typeof v !== "string") + ) { + return undefined; + } + return (value as string[]).map((v) => ({ value: v, label: v })); +} + +function parseChoicesFromOneOf(value: unknown): Choice[] | undefined { + // Empty choice sets are rejected (like empty `enum`): a required field + // with zero options would render an unwinnable prompt. + if (!Array.isArray(value) || value.length === 0) return undefined; + const choices: Choice[] = []; + for (const entry of value) { + if (!isRecord(entry) || typeof entry.const !== "string") return undefined; + choices.push({ + value: entry.const, + label: typeof entry.title === "string" ? entry.title : entry.const, + }); + } + return choices; +} + +/** + * Length/count keywords (`minLength`, `maxLength`, `minItems`, `maxItems`) + * must be non-negative integers. A negative bound (e.g. `maxLength: -1` on a + * required string) would otherwise parse fine and then reject every possible + * answer — an unwinnable prompt loop. + */ +function isValidCount(value: number | undefined): boolean { + return value === undefined || (Number.isInteger(value) && value >= 0); +} + +/** + * JSON Schema `minLength`/`maxLength` count Unicode code points, not UTF-16 + * code units — `"😀"` has length 1 under the spec but `.length === 2` in + * JavaScript. Every bound check on user-visible strings must use this. + */ +export function codePointLength(value: string): number { + // String iteration yields code points, unlike .length's UTF-16 units. + return [...value].length; +} + +/** + * A structurally valid field can still be internally inconsistent — + * unsatisfiable constraints (`minimum > maximum`, `minItems` above the + * choice count) or a default that violates its own constraints. Those would + * render unwinnable or instantly-invalid prompts, so treat them like any + * other malformed schema and reject the field. + */ +function isConsistent(field: FieldExtra): boolean { + switch (field.kind) { + case "boolean": + return true; + case "number": + if ( + field.minimum !== undefined && + field.maximum !== undefined && + field.minimum > field.maximum + ) { + return false; + } + if (field.default !== undefined) { + if (field.integer && !Number.isInteger(field.default)) return false; + if (field.minimum !== undefined && field.default < field.minimum) + return false; + if (field.maximum !== undefined && field.default > field.maximum) + return false; + } + return true; + case "string": + if (!isValidCount(field.minLength) || !isValidCount(field.maxLength)) { + return false; + } + if ( + field.minLength !== undefined && + field.maxLength !== undefined && + field.minLength > field.maxLength + ) { + return false; + } + if (field.default !== undefined) { + if ( + field.minLength !== undefined && + codePointLength(field.default) < field.minLength + ) { + return false; + } + if ( + field.maxLength !== undefined && + codePointLength(field.default) > field.maxLength + ) { + return false; + } + } + return true; + case "enum": + return ( + field.default === undefined || + field.choices.some((c) => c.value === field.default) + ); + case "multiselect": { + if (!isValidCount(field.minItems) || !isValidCount(field.maxItems)) { + return false; + } + if ( + field.minItems !== undefined && + field.maxItems !== undefined && + field.minItems > field.maxItems + ) { + return false; + } + if ( + field.minItems !== undefined && + field.minItems > field.choices.length + ) { + return false; + } + const def = field.default; + if (def !== undefined) { + if (!def.every((v) => field.choices.some((c) => c.value === v))) { + return false; + } + if (field.minItems !== undefined && def.length < field.minItems) + return false; + if (field.maxItems !== undefined && def.length > field.maxItems) + return false; + } + return true; + } + } +} + +function parseField(prop: unknown): FieldExtra | null { + const parsed = parseFieldShape(prop); + if (!parsed || !isConsistent(parsed)) return null; + return parsed; +} + +function parseFieldShape(prop: unknown): FieldExtra | null { + if (!isRecord(prop)) return null; + const type = prop.type; + + if (type === "boolean") { + return { + kind: "boolean", + default: typeof prop.default === "boolean" ? prop.default : undefined, + }; + } + + if (type === "number" || type === "integer") { + return { + kind: "number", + integer: type === "integer", + minimum: typeof prop.minimum === "number" ? prop.minimum : undefined, + maximum: typeof prop.maximum === "number" ? prop.maximum : undefined, + default: typeof prop.default === "number" ? prop.default : undefined, + }; + } + + if (type === "string") { + if (prop.enum !== undefined) { + const enumChoices = parseChoicesFromEnum(prop.enum); + // Present-but-invalid (non-string entries or an empty list) is a + // malformed schema, not a freeform string field: an empty required + // choice prompt would be unwinnable. + if (!enumChoices) return null; + return { + kind: "enum", + choices: enumChoices, + default: typeof prop.default === "string" ? prop.default : undefined, + }; + } + if (prop.oneOf !== undefined) { + const oneOfChoices = parseChoicesFromOneOf(prop.oneOf); + if (!oneOfChoices) return null; + return { + kind: "enum", + choices: oneOfChoices, + default: typeof prop.default === "string" ? prop.default : undefined, + }; + } + return { + kind: "string", + minLength: + typeof prop.minLength === "number" ? prop.minLength : undefined, + maxLength: + typeof prop.maxLength === "number" ? prop.maxLength : undefined, + format: typeof prop.format === "string" ? prop.format : undefined, + default: typeof prop.default === "string" ? prop.default : undefined, + }; + } + + if (type === "array") { + const items = prop.items; + if (!isRecord(items)) return null; + const choices = + parseChoicesFromEnum(items.enum) ?? parseChoicesFromOneOf(items.anyOf); + if (!choices) return null; + const defaultValue = + Array.isArray(prop.default) && + prop.default.every((v) => typeof v === "string") + ? (prop.default as string[]) + : undefined; + return { + kind: "multiselect", + choices, + minItems: typeof prop.minItems === "number" ? prop.minItems : undefined, + maxItems: typeof prop.maxItems === "number" ? prop.maxItems : undefined, + default: defaultValue, + }; + } + + return null; +} + +/** Parse a `requestedSchema` into an ordered list of {@link FormField}s. */ +export function parseFormSchema( + schema: Record<string, unknown> | undefined, +): FormField[] | null { + if (!isRecord(schema)) return null; + const properties = schema.properties; + if (!isRecord(properties)) return null; + const required = Array.isArray(schema.required) + ? (schema.required.filter((v) => typeof v === "string") as string[]) + : []; + + const fields: FormField[] = []; + for (const [name, prop] of Object.entries(properties)) { + const parsed = parseField(prop); + if (!parsed) return null; + const title = + isRecord(prop) && typeof prop.title === "string" ? prop.title : name; + const description = + isRecord(prop) && typeof prop.description === "string" + ? prop.description + : undefined; + fields.push({ + name, + required: required.includes(name), + title, + description, + ...parsed, + } as FormField); + } + return fields; +} diff --git a/clients/mcpdo/src/connection/format-connection.ts b/clients/mcpdo/src/connection/format-connection.ts new file mode 100644 index 0000000000..ed5077d1ea --- /dev/null +++ b/clients/mcpdo/src/connection/format-connection.ts @@ -0,0 +1,527 @@ +import { + awaitableError, + awaitableLog, +} from "@inspector/core/cli/utils/awaitable-log.js"; +import type { + ConnectionInfo, + ElicitationPendingInfo, +} from "../daemon/protocol.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import type { OutputFormat } from "@inspector/core/cli/handlers/format-output.js"; +import type { CliAppInfo } from "@inspector/core/cli/handlers/method-types.js"; +import { + formatAppInfoHuman, + formatAppInfoListHuman, + formatAuthListHuman, + formatEmaStatusHuman, + formatRpcResultHuman, + formatServersListHuman, + formatServerShowHuman, + formatConnectionInfoHuman, + formatConnectionsListHuman, + formatElicitationPendingHuman, + formatSkillVerifyListHuman, + formatStreamEventHuman, +} from "./format-human.js"; +import { isSafeLinkTarget, sanitizeDeep, sanitizeText } from "./sanitize.js"; +import { PLAIN, type Style } from "@inspector/core/cli/style.js"; + +type JsonObject = Record<string, unknown>; + +/** + * Pretty-print JSON for connection `--format json`. + * Unlike one-shot, this does **not** wrap in `{ result }` — the payload is the + * MCP / admin object itself (convenient for scripting). + * + * `JSON.stringify` escapes C0 controls but emits C1 controls (U+0080–U+009F, + * including 8-bit CSI/OSC) literally, which terminals can interpret. Escape + * them as standard `\uXXXX` sequences: the serialized text is terminal-safe + * while parsed values stay byte-identical. + */ +export function formatConnectionJson(data: unknown): string { + return ( + JSON.stringify(data, null, 2).replace( + /[\u0080-\u009F]/g, + (ch) => `\\u${ch.charCodeAt(0).toString(16).padStart(4, "0")}`, + ) + "\n" + ); +} + +export type ConnectionWriteKind = + | { + kind: "rpc"; + method: string; + result: JsonObject; + /** + * Auto-collected by `runMethod` for `tools/call` + `--format json`. + * Connection output ignores this side-channel (no `{ result, appInfo }` + * envelope); only `result` is printed. `--app-info` probes put the + * info object in `result` itself. + */ + appInfo?: CliAppInfo; + /** For exit-code messages when result.isError. */ + toolName?: string; + } + | { + kind: "ndjson"; + lines: unknown[]; + /** Distinguishes `tools/list --app-info` probe lines from a `--verify` report. */ + variant?: "app-info" | "skill-verify"; + /** `--verify` one-line stderr verdict; absent for `--app-info`. */ + summary?: string; + /** Non-zero when the emitted `--verify` report is itself a failure. */ + exitCode?: number; + } + | { kind: "stream-event"; data: unknown } + | { + kind: "servers/list"; + servers: unknown[]; + /** Which file produced the entries (writable catalog vs read-only config). */ + source?: { kind: "catalog" | "config"; path: string }; + } + | { + kind: "servers/show"; + server: JsonObject; + /** Which file produced the entry (writable catalog vs read-only config). */ + source?: { kind: "catalog" | "config"; path: string }; + } + | { kind: "connections/list"; connections: unknown[] } + | { + kind: "connection"; + connection: ConnectionInfo | JsonObject; + /** + * Non-TTY pending sign-in (see auth-helper.ts): the authorize URL the + * caller must relay to a human. Rides the normal output payload — the + * error envelope redacts URL query strings, which would strip the + * client_id/PKCE/state this URL is made of. + */ + authUrl?: string; + } + | { + /** + * A parked elicitation (non-interactive caller): everything needed to + * relay the request to a human and answer it with + * `elicitation/respond`. Rides the normal output payload for the same + * redaction reason as `authUrl` (URL-mode elicitations carry a URL + * whose query is meaningful). + */ + kind: "elicitation-pending"; + elicitation: ElicitationPendingInfo; + } + | { kind: "disconnect"; name: string; clearedAuthUrl?: string } + | { kind: "daemon/status"; status: JsonObject } + | { kind: "daemon/stop"; result: JsonObject } + | { + kind: "auth/list"; + list: { oauthStatePath: string; servers: unknown[] }; + } + | { + kind: "auth/clear"; + result: { + url?: string; + cleared?: number; + all?: boolean; + clearedByName?: string; + }; + } + | { + kind: "auth/ema-status"; + status: { + clientConfigPath: string; + configured: boolean; + enabled: boolean; + issuer?: string; + clientId?: string; + loginState: string; + }; + } + | { + kind: "auth/ema-login"; + result: { + issuer: string; + loginState: string; + alreadyLoggedIn: boolean; + /** Present when the login was parked on a detached helper (non-TTY). */ + pendingLogin?: boolean; + /** IdP authorization URL to relay to the human. */ + authUrl?: string; + }; + } + | { + kind: "auth/ema-logout"; + result: { issuer: string; endSessionUrl?: string }; + } + | { kind: "generic"; data: unknown; title?: string }; + +export type ConnectionWriteOpts = { + format?: OutputFormat; + /** Human-output styling; ignored for `--format json`. Defaults to plain. */ + style?: Style; + /** + * Whether a human is reading this output live on a TTY. Sign-in prompts + * address the reader directly when `true` ("Open this link…") and instruct + * the *agent* to relay the link to its user when `false`/absent — the same + * split the `auth_required` error draws (see `connections.ts` + * `authRequiredError`), because a non-interactive agent's user never sees + * this command's output. Defaults to the agent framing. + */ + interactive?: boolean; +}; + +/** + * Write connection CLI output honouring `--format text|json`. + * One-shot output paths are unchanged (`emitResult` / `writeFormattedResult`). + */ +export async function writeConnectionOutput( + opts: ConnectionWriteOpts, + payload: ConnectionWriteKind, +): Promise<void> { + const format: OutputFormat = opts.format === "json" ? "json" : "text"; + const style = opts.style ?? PLAIN; + const interactive = opts.interactive === true; + + if (format === "json") { + await awaitableLog(formatConnectionJson(jsonPayload(payload))); + await writeNdjsonSummary(payload); + applyExitCodes(payload); + return; + } + + // Server-controlled strings must never reach the terminal raw (escape + // injection: OSC 52 clipboard writes, title spoofing, output rewriting). + // Sanitize the whole payload before human formatting; the formatter's own + // ANSI styling is applied afterwards and stays intact. JSON output above + // is made safe by formatConnectionJson (C0 via JSON.stringify, C1 via its + // own escaping). + await awaitableLog( + humanPayload(sanitizeDeep(payload), style, interactive) + "\n", + ); + await writeNdjsonSummary(payload); + applyExitCodes(payload); +} + +/** + * `skills/list --verify` / `skills/get --verify`: the one-line verdict goes to + * **stderr**, after the report, in both `--format text` and `--format json` — + * mirrors the one-shot CLI (`consumeMethodOutcome`), so a reader piping stdout + * into `jq` still sees it and a `--format json` caller isn't left without one + * just because the report itself is already structured. + */ +async function writeNdjsonSummary(payload: ConnectionWriteKind): Promise<void> { + if (payload.kind === "ndjson" && payload.summary) { + // Human-facing stderr line in both formats; may embed server-derived + // names, so sanitize (see sanitize.ts). + await awaitableError(`${sanitizeText(payload.summary)}\n`); + } +} + +function jsonPayload(payload: ConnectionWriteKind): unknown { + switch (payload.kind) { + case "rpc": + // Pretty payload only — never the one-shot `{ result[, appInfo] }` wrap. + return payload.result; + case "ndjson": + return payload.lines; + case "stream-event": + return payload.data; + case "servers/list": + return { + servers: payload.servers, + ...(payload.source && { source: payload.source }), + }; + case "servers/show": + return payload.source + ? { ...payload.server, source: payload.source } + : payload.server; + case "connections/list": + return { connections: payload.connections }; + case "connection": + return payload.authUrl !== undefined + ? { ...(payload.connection as JsonObject), authUrl: payload.authUrl } + : payload.connection; + case "elicitation-pending": + // The key doubles as the discriminator: a caller can tell "input + // required" from a final tool result by `elicitationPending` alone. + return { elicitationPending: payload.elicitation }; + case "disconnect": + return { + name: payload.name, + ...(payload.clearedAuthUrl && { + clearedAuthUrl: payload.clearedAuthUrl, + }), + }; + case "daemon/status": + return payload.status; + case "daemon/stop": + return payload.result; + case "auth/list": + return payload.list; + case "auth/clear": + return payload.result; + case "auth/ema-status": + return payload.status; + case "auth/ema-login": + return payload.result; + case "auth/ema-logout": + return payload.result; + case "generic": + return payload.data; + } +} + +/** + * Headline for a pending browser sign-in, addressed to whoever reads it. + * + * - Interactive (a human on a TTY): ask them to open the link themselves. + * - Non-interactive (an agent relaying for its user): the user never sees this + * command's output, so the link must be reproduced verbatim in the agent's + * own reply — a bare reference ("the link above") reaches no one. Mirrors the + * agent/human split the `auth_required` error already draws in + * `connections.ts` `authRequiredError`. + */ +function signInHeadline(): string { + return "Sign-in required. Open this link in a browser to authenticate:"; +} + +/** + * Agent-facing (non-TTY) sign-in block. + * + * In a captured-stdout agent harness (e.g. Claude Code), assistant text written + * *before* a tool call in the same turn is not reliably shown to the user, and + * the command's own output is collapsed behind a disclosure. So the URL is + * embedded here as literal plain text (never an OSC 8 link) and the copy is a + * factual statement: the connection is pending and not usable, show the URL to + * the user, and wait for them to sign in before continuing. The authoritative + * stop-and-wait guidance lives in the mcpdo skill; this block is a concise + * restatement for the agent reading the command's own output. + * + * `resumeCmd` is the single command the agent runs *after* the user confirms. + */ +function agentSignInBlock( + url: string, + subject: string, + resumeCmd: string, + finishPhrase: string, +): string { + return [ + `Sign-in required for ${subject}. This command exited 0, but ${subject} is not usable yet until the user signs in — this is not an error and not success.`, + "", + `Show the sign-in URL below to the user as literal plain text (make it the last thing in your reply), ask them to open it and sign in, and wait for them to confirm before continuing. Do not retry, poll, sleep, or run another command this turn — the user cannot see your reply until the turn ends.`, + "", + "Sign-in URL (show this to the user):", + "", + ` ${url}`, + "", + `Once the user confirms they have signed in, run \`mcpdo ${resumeCmd}\` to ${finishPhrase}.`, + ].join("\n"); +} + +function humanPayload( + payload: ConnectionWriteKind, + style: Style, + interactive: boolean, +): string { + switch (payload.kind) { + case "rpc": { + if (asAppInfoProbe(payload.result)) { + return formatAppInfoHuman(payload.result, style); + } + const formatted = formatRpcResultHuman( + payload.method, + payload.result, + style, + ); + return formatted ?? JSON.stringify(payload.result, null, 2); + } + case "ndjson": + return payload.variant === "skill-verify" + ? formatSkillVerifyListHuman(payload.lines, style) + : formatAppInfoListHuman(payload.lines, style); + case "stream-event": + return formatStreamEventHuman(payload.data, style); + case "servers/list": + return formatServersListHuman(payload.servers, style, payload.source); + case "servers/show": + return formatServerShowHuman(payload.server, style, payload.source); + case "connections/list": + return formatConnectionsListHuman(payload.connections, style); + case "connection": { + const info = formatConnectionInfoHuman( + payload.connection as JsonObject, + style, + ); + if (payload.authUrl === undefined) return info; + const name = String((payload.connection as JsonObject).name ?? ""); + if (!interactive) { + return [ + info, + "", + agentSignInBlock( + payload.authUrl, + `@${name}`, + `connections/show @${name}`, + "finish the connection", + ), + ].join("\n"); + } + return [ + info, + "", + signInHeadline(), + // The URL comes from server-controlled OAuth metadata: only + // allowlisted schemes become clickable OSC 8 links (same gate as + // every other server-supplied link — see sanitize.ts). + ` ${isSafeLinkTarget(payload.authUrl) ? style.link(payload.authUrl) : payload.authUrl}`, + style.dim( + `The connection completes automatically after sign-in — check with \`connections/show @${name}\`, or just run the next command.`, + ), + ].join("\n"); + } + case "elicitation-pending": + return formatElicitationPendingHuman(payload.elicitation, style); + case "disconnect": { + const line = `${style.bold("Disconnected")} ${`\`${style.bold(`@${payload.name}`)}\``}`; + if (!payload.clearedAuthUrl) return line; + return [ + line, + style.dim( + `Cleared stored auth for ${payload.clearedAuthUrl} — the next connect will re-trigger sign-in.`, + ), + ].join("\n"); + } + case "daemon/status": { + const s = payload.status; + if (s.running === false) { + return String(s.message ?? "Daemon is not running."); + } + const connections = Array.isArray(s.connections) + ? (s.connections as unknown[]) + : []; + return [ + `${style.bold("Daemon")} pid ${String(s.pid)}` + + (s.stopping === true ? ` ${style.yellow("(shutting down)")}` : ""), + style.dim(`Socket: ${String(s.socketPath ?? "")}`), + formatConnectionsListHuman(connections, style), + ].join("\n"); + } + case "daemon/stop": + if (payload.result.stopping === false) { + return String(payload.result.message ?? "Daemon was not running."); + } + return style.green("Daemon stopping."); + case "auth/list": + return formatAuthListHuman(payload.list, style); + case "auth/clear": + if (payload.result.all === true) { + return style.green( + `Cleared ${String(payload.result.cleared ?? 0)} stored auth entr${ + payload.result.cleared === 1 ? "y" : "ies" + }.`, + ); + } + return `${style.green("Cleared")} \`${style.bold(String(payload.result.url ?? ""))}\`${ + payload.result.clearedByName + ? style.dim(` (${String(payload.result.clearedByName)})`) + : "" + }`; + case "auth/ema-status": + return formatEmaStatusHuman(payload.status, style); + case "auth/ema-login": + if (payload.result.alreadyLoggedIn) { + return `${style.green("Already signed in")} to \`${style.bold(payload.result.issuer)}\` ${style.dim("(use auth/ema-login --relogin for a fresh connection)")}`; + } + if (payload.result.pendingLogin === true && payload.result.authUrl) { + if (!interactive) { + return agentSignInBlock( + payload.result.authUrl, + "enterprise IdP login", + "auth/ema-status", + "finish the login", + ); + } + return [ + signInHeadline(), + // Same OSC 8 allowlist gate as the connection authUrl above. + ` ${isSafeLinkTarget(payload.result.authUrl) ? style.link(payload.result.authUrl) : payload.result.authUrl}`, + style.dim( + "The sign-in completes in the background — check with `auth/ema-status` (loginState becomes logged_in).", + ), + ].join("\n"); + } + return `${style.green("Signed in")} to \`${style.bold(payload.result.issuer)}\``; + case "auth/ema-logout": { + const signedOut = `${style.green("Signed out")} of \`${style.bold(payload.result.issuer)}\` ${style.dim("(EMA server tokens cleared)")}`; + if (payload.result.endSessionUrl === undefined) return signedOut; + // Local clear only — the IdP's browser SSO cookie survives it. Relay + // the RP-initiated logout URL so the user can end that session too. + return [ + signedOut, + "To end your IdP browser session, navigate to:", + // Same OSC 8 allowlist gate as the sign-in URLs above, so the logout + // link renders clickable in a terminal instead of as plain text. + ` ${isSafeLinkTarget(payload.result.endSessionUrl) ? style.link(payload.result.endSessionUrl) : payload.result.endSessionUrl}`, + ].join("\n"); + } + case "generic": { + if (payload.title) { + return `${style.bold(payload.title)}\n${JSON.stringify(payload.data, null, 2)}`; + } + return JSON.stringify(payload.data, null, 2); + } + } +} + +function asAppInfoProbe(result: JsonObject): CliAppInfo | undefined { + if ( + typeof result.hasApp !== "boolean" || + typeof result.toolName !== "string" || + result.content !== undefined || + result.tools !== undefined + ) { + return undefined; + } + // Narrowed by the structural checks above; CliAppInfo adds optional fields. + // `JsonObject`'s index signature doesn't structurally overlap with + // `CliAppInfo`'s concrete shape, so `as` needs the `unknown` bridge. + return result as unknown as CliAppInfo; +} + +function applyExitCodes(payload: ConnectionWriteKind): void { + if (payload.kind === "ndjson" && payload.exitCode) { + // Report already written above; thrown last so it routes through the + // connection CLI's single exit path, same as the one-shot CLI's + // `consumeMethodOutcome` (Copilot). + throw new CliExitCodeError(payload.exitCode, payload.summary ?? "", { + code: + payload.exitCode === EXIT_CODES.SKILL_INCOMPLETE + ? "skills_incomplete" + : "skills_nonconformant", + }); + } + if (payload.kind === "rpc") { + // Only `--app-info` probes (result is the info object) map to NO_APP. + // Auto-collected `payload.appInfo` from tools/call+json must not. + const info = asAppInfoProbe(payload.result); + if (info) { + if (!info.hasApp) { + throw new CliExitCodeError( + EXIT_CODES.NO_APP, + // toolName echoes server-influenced text into a terminal-bound + // error message; sanitize like the success path does. + `Tool '${sanitizeText(info.toolName)}' has no MCP App UI resource (_meta.ui.resourceUri).`, + ); + } + return; + } + if (payload.result.isError === true) { + throw new CliExitCodeError( + EXIT_CODES.TOOL_ERROR, + `Tool '${sanitizeText(payload.toolName ?? "tool")}' returned isError:true.`, + { code: "tool_is_error" }, + ); + } + } +} diff --git a/clients/mcpdo/src/connection/format-human.ts b/clients/mcpdo/src/connection/format-human.ts new file mode 100644 index 0000000000..966c48b956 --- /dev/null +++ b/clients/mcpdo/src/connection/format-human.ts @@ -0,0 +1,1030 @@ +/** + * Human-readable (markdown-ish) formatters for the connection CLI. + * Styling (color / bold / dim / OSC 8 links) is parameterized via {@link Style}. + */ + +import { PLAIN, type Style } from "@inspector/core/cli/style.js"; +import { isSafeLinkTarget } from "./sanitize.js"; +import { parseFormSchema } from "./form-schema.js"; +import type { ElicitationPendingInfo } from "../daemon/protocol.js"; + +type JsonObject = Record<string, unknown>; + +function asArray<T>(value: unknown): T[] { + return Array.isArray(value) ? (value as T[]) : []; +} + +function shortType(schema: unknown): string { + if (!schema || typeof schema !== "object") return "any"; + const s = schema as JsonObject; + const t = s.type; + if (t === "array") { + if (s.items) return `[${shortType(s.items)}]`; + return "[any]"; + } + if (Array.isArray(t)) { + const filtered = t.filter((x) => x !== "null"); + if (filtered.length === 1) return shortTypeName(String(filtered[0])); + return filtered.map((x) => shortTypeName(String(x))).join(" | "); + } + if (Array.isArray(s.enum)) return "enum"; + if (typeof t === "string") return shortTypeName(t); + return "any"; +} + +function shortTypeName(type: string): string { + const map: Record<string, string> = { + string: "str", + number: "num", + integer: "int", + boolean: "bool", + object: "obj", + array: "[any]", + }; + return map[type] ?? type; +} + +function formatToolParamsInline(schema: unknown): string { + if (!schema || typeof schema !== "object") return "()"; + const s = schema as JsonObject; + const properties = s.properties as Record<string, unknown> | undefined; + if (!properties || Object.keys(properties).length === 0) return "()"; + const required = new Set(asArray<string>(s.required)); + const names = Object.keys(properties); + const ordered = [ + ...names.filter((n) => required.has(n)), + ...names.filter((n) => !required.has(n)), + ]; + const shown = ordered.slice(0, 3); + const hidden = ordered.length - shown.length; + const parts = shown.map((name) => { + const typeStr = shortType(properties[name]); + return required.has(name) ? `${name}:${typeStr}` : `${name}?:${typeStr}`; + }); + if (hidden > 0) parts.push("…"); + return `(${parts.join(", ")})`; +} + +function toolHints(tool: JsonObject): string | undefined { + const ann = tool.annotations as JsonObject | undefined; + if (!ann) return undefined; + const hints: string[] = []; + if (ann.readOnlyHint === true) hints.push("read-only"); + if (ann.destructiveHint === true) hints.push("destructive"); + if (ann.idempotentHint === true) hints.push("idempotent"); + if (ann.openWorldHint === true) hints.push("open-world"); + return hints.length > 0 ? hints.join(", ") : undefined; +} + +function code(style: Style, name: string): string { + return `\`${style.bold(name)}\``; +} + +function heading(style: Style, text: string): string { + return style.bold(text); +} + +function descSuffix(style: Style, description: unknown): string { + if (typeof description !== "string" || !description.trim()) return ""; + return style.dim(` — ${description.trim().split("\n")[0]}`); +} + +function formatUri(style: Style, uri: string): string { + if (!uri) return uri; + // Only allowlisted schemes become clickable OSC 8 links (see sanitize.ts); + // file:/custom-handler URIs from a server render as plain colored text. + if (isSafeLinkTarget(uri)) return style.link(uri); + return style.cyan(uri); +} + +function colorLevel(style: Style, level: string): string { + switch (level) { + case "error": + case "critical": + case "alert": + case "emergency": + return style.red(level); + case "warning": + return style.yellow(level); + case "debug": + case "notice": + return style.dim(level); + default: + return style.cyan(level); + } +} + +/** Format tools/list for human display. */ +export function formatToolsHuman( + tools: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Tools (${tools.length}):`)]; + for (const raw of tools) { + const tool = raw as JsonObject; + const name = String(tool.name ?? "?"); + const params = formatToolParamsInline(tool.inputSchema); + const hints = toolHints(tool); + const hintSuffix = hints ? style.dim(` [${hints}]`) : ""; + lines.push( + `* \`${style.bold(name)}${style.cyan(params)}\`${hintSuffix}${descSuffix(style, tool.description)}`, + ); + } + if (tools.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format resources/list. */ +export function formatResourcesHuman( + resources: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Resources (${resources.length}):`)]; + for (const raw of resources) { + const r = raw as JsonObject; + const name = typeof r.name === "string" ? r.name : String(r.uri ?? "?"); + const uri = typeof r.uri === "string" ? r.uri : ""; + const uriPart = uri ? ` (${formatUri(style, uri)})` : ""; + lines.push( + `* ${code(style, name)}${uriPart}${descSuffix(style, r.description)}`, + ); + } + if (resources.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format resources/templates/list. */ +export function formatResourceTemplatesHuman( + templates: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Resource templates (${templates.length}):`)]; + for (const raw of templates) { + const t = raw as JsonObject; + const name = String(t.name ?? "?"); + const uri = typeof t.uriTemplate === "string" ? t.uriTemplate : ""; + const uriPart = uri ? ` (${formatUri(style, uri)})` : ""; + lines.push( + `* ${code(style, name)}${uriPart}${descSuffix(style, t.description)}`, + ); + } + if (templates.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format prompts/list. */ +export function formatPromptsHuman( + prompts: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Prompts (${prompts.length}):`)]; + for (const raw of prompts) { + const p = raw as JsonObject; + const name = String(p.name ?? "?"); + lines.push(`* ${code(style, name)}${descSuffix(style, p.description)}`); + } + if (prompts.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +function formatContentBlock(block: JsonObject, style: Style): string[] { + const lines: string[] = []; + switch (block.type) { + case "text": + lines.push("````"); + lines.push(String(block.text ?? "")); + lines.push("````"); + break; + case "resource_link": + lines.push(heading(style, "Resource link")); + lines.push(`* URI: ${formatUri(style, String(block.uri ?? ""))}`); + if (block.name) lines.push(`* Name: ${String(block.name)}`); + if (block.description) + lines.push(`* Description: ${String(block.description)}`); + if (block.mimeType) lines.push(`* MIME type: ${String(block.mimeType)}`); + break; + case "image": + lines.push( + style.dim( + `[Image: ${String(block.mimeType ?? "unknown")}${ + typeof block.data === "string" + ? `, ${block.data.length} chars base64` + : "" + }]`, + ), + ); + break; + case "audio": + lines.push( + style.dim( + `[Audio: ${String(block.mimeType ?? "unknown")}${ + typeof block.data === "string" + ? `, ${block.data.length} chars base64` + : "" + }]`, + ), + ); + break; + case "resource": { + lines.push(heading(style, "Embedded resource")); + const res = block.resource as JsonObject | undefined; + if (res) { + lines.push(`* URI: ${formatUri(style, String(res.uri ?? ""))}`); + if (res.mimeType) lines.push(`* MIME type: ${String(res.mimeType)}`); + if (typeof res.text === "string") { + lines.push("````"); + lines.push(res.text); + lines.push("````"); + } + } + break; + } + default: + lines.push(JSON.stringify(block, null, 2)); + } + return lines; +} + +function findDuplicateTextBlocks( + content: JsonObject[], + structuredContent: JsonObject, +): Set<number> { + const dupes = new Set<number>(); + const canonical = JSON.stringify(structuredContent); + for (let i = 0; i < content.length; i++) { + const block = content[i]; + if (!block || block.type !== "text" || typeof block.text !== "string") + continue; + try { + const parsed: unknown = JSON.parse(block.text.trim()); + if (JSON.stringify(parsed) === canonical) dupes.add(i); + } catch { + // keep + } + } + return dupes; +} + +/** + * Format a CallToolResult (also used for tasks/result) for human display. + */ +export function formatCallToolResultHuman( + result: JsonObject, + style: Style = PLAIN, +): string { + const lines: string[] = []; + if (result.isError === true) { + lines.push(style.red(heading(style, "Tool error:"))); + } + + const sc = result.structuredContent as JsonObject | undefined; + const hasStructuredContent = !!sc && Object.keys(sc).length > 0; + const content = asArray<JsonObject>(result.content); + const skipIndices = hasStructuredContent + ? findDuplicateTextBlocks(content, sc!) + : new Set<number>(); + const visible = content.filter((_, i) => !skipIndices.has(i)); + + if (visible.length > 0) { + lines.push(heading(style, "Content:")); + for (let i = 0; i < visible.length; i++) { + if (i > 0) lines.push(""); + lines.push(...formatContentBlock(visible[i]!, style)); + } + } + + // Always render structuredContent once. The duplicate-text filter above + // may have removed its JSON copy from the content blocks, so gating this + // on `visible.length === 0` would drop the structured payload whenever + // any other content block is present alongside the duplicate. + if (hasStructuredContent) { + if (lines.length > 0) lines.push(""); + lines.push(heading(style, "Structured content:")); + lines.push(JSON.stringify(sc, null, 2)); + } + + const meta = result._meta as JsonObject | undefined; + if (meta && Object.keys(meta).length > 0) { + if (lines.length > 0) lines.push(""); + lines.push(style.dim("Metadata:")); + lines.push(style.dim(JSON.stringify(meta, null, 2))); + } + + if (lines.length === 0) return style.dim("(no content)"); + return lines.join("\n"); +} + +/** Format resources/read contents. */ +export function formatResourceReadHuman( + result: JsonObject, + style: Style = PLAIN, +): string { + const contents = asArray<JsonObject>(result.contents); + if (contents.length === 0) return style.dim("(empty resource)"); + const lines: string[] = [ + heading(style, `Resource contents (${contents.length}):`), + ]; + for (const c of contents) { + lines.push(""); + lines.push(`URI: ${formatUri(style, String(c.uri ?? ""))}`); + if (c.mimeType) lines.push(style.dim(`MIME: ${String(c.mimeType)}`)); + if (typeof c.text === "string") { + lines.push("````"); + lines.push(c.text); + lines.push("````"); + } else if (typeof c.blob === "string") { + lines.push(style.dim(`[Blob: ${c.blob.length} chars base64]`)); + } + } + return lines.join("\n"); +} + +/** Format prompts/get. */ +export function formatPromptResultHuman( + result: JsonObject, + style: Style = PLAIN, +): string { + const description = + typeof result.description === "string" ? result.description : undefined; + const messages = asArray<JsonObject>(result.messages); + const lines: string[] = []; + if (description) { + lines.push(style.dim(description)); + lines.push(""); + } + lines.push(heading(style, `Messages (${messages.length}):`)); + for (const msg of messages) { + const role = String(msg.role ?? "?"); + lines.push(""); + lines.push(style.cyan(`[${role}]`)); + const content = msg.content; + if (typeof content === "string") { + lines.push("````"); + lines.push(content); + lines.push("````"); + } else if (content && typeof content === "object") { + if (Array.isArray(content)) { + for (const block of content as JsonObject[]) { + lines.push(...formatContentBlock(block, style)); + } + } else { + lines.push(...formatContentBlock(content as JsonObject, style)); + } + } + } + if (messages.length === 0 && !description) return style.dim("(empty prompt)"); + return lines.join("\n"); +} + +/** Format prompts/complete. */ +export function formatCompletionsHuman( + result: JsonObject, + style: Style = PLAIN, +): string { + const values = asArray<string>(result.values); + const lines = [heading(style, `Completions (${values.length}):`)]; + for (const v of values) lines.push(`* ${v}`); + if (values.length === 0) lines.push(style.dim("(none)")); + if (result.hasMore === true) lines.push(style.dim("(more available)")); + return lines.join("\n"); +} + +/** Format tasks/list. */ +export function formatTasksHuman( + tasks: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Tasks (${tasks.length}):`)]; + for (const raw of tasks) { + const t = raw as JsonObject; + const id = String(t.taskId ?? t.id ?? "?"); + const status = String(t.status ?? "?"); + const msg = + typeof t.statusMessage === "string" + ? style.dim(` — ${t.statusMessage}`) + : ""; + lines.push(`* ${code(style, id)} ${status}${msg}`); + } + if (tasks.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format tasks/get. */ +export function formatTaskHuman(task: unknown, style: Style = PLAIN): string { + const t = (task ?? {}) as JsonObject; + const lines = [ + `${heading(style, "Task:")} ${code(style, String(t.taskId ?? t.id ?? "?"))}`, + `Status: ${String(t.status ?? "?")}`, + ]; + if (typeof t.statusMessage === "string") { + lines.push(`Message: ${t.statusMessage}`); + } + if (t.createdAt) lines.push(style.dim(`Created: ${String(t.createdAt)}`)); + if (t.lastUpdatedAt) + lines.push(style.dim(`Updated: ${String(t.lastUpdatedAt)}`)); + return lines.join("\n"); +} + +/** Format initialize / server probe. */ +export function formatInitializeHuman( + result: JsonObject, + style: Style = PLAIN, +): string { + const info = (result.serverInfo ?? {}) as JsonObject; + const lines = [ + `${heading(style, "Server:")} ${style.bold(String(info.name ?? "(unknown)"))}${ + info.version ? style.dim(` v${String(info.version)}`) : "" + }`, + ]; + if (result.protocolVersion) { + lines.push(`Protocol: ${String(result.protocolVersion)}`); + } + if (typeof result.instructions === "string" && result.instructions.trim()) { + lines.push(""); + lines.push(heading(style, "Instructions:")); + lines.push(result.instructions.trim()); + } + const caps = result.capabilities; + if (caps && typeof caps === "object") { + const keys = Object.keys(caps as JsonObject); + if (keys.length > 0) { + lines.push(""); + lines.push(`${heading(style, "Capabilities:")} ${keys.join(", ")}`); + } + } + return lines.join("\n"); +} + +/** Format roots/list or roots/set. */ +export function formatRootsHuman( + roots: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Roots (${roots.length}):`)]; + for (const raw of roots) { + const r = raw as JsonObject; + const name = typeof r.name === "string" ? style.dim(` (${r.name})`) : ""; + lines.push(`* ${formatUri(style, String(r.uri ?? "?"))}${name}`); + } + if (roots.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format auth/list. */ +export function formatAuthListHuman( + list: { + oauthStatePath?: string; + servers?: unknown[]; + }, + style: Style = PLAIN, +): string { + const servers = Array.isArray(list.servers) ? list.servers : []; + const lines = [ + heading(style, `Stored auth (${servers.length}):`), + style.dim(String(list.oauthStatePath ?? "")), + ]; + for (const raw of servers) { + const s = raw as JsonObject; + const flags: string[] = []; + if (s.hasTokens === true) flags.push("tokens"); + if (s.hasRefreshToken === true) flags.push("refresh"); + const flagText = + flags.length > 0 + ? style.dim(` (${flags.join(", ")})`) + : style.dim(" (no tokens)"); + // An EMA IdP login record is not a server: lead with its issuer URL (which + // auth/clear also accepts) like every other row, and mark it in the trailing + // annotation slot — parallel to "known as:" — rather than prefixing the line + // or showing the raw `ema-idp:` key with a misleading "(no local name)". + if (s.idp === true && typeof s.issuer === "string") { + lines.push( + `* ${code(style, s.issuer)}${flagText}${style.dim( + " enterprise IdP login", + )}`, + ); + continue; + } + const knownAs = Array.isArray(s.knownAs) + ? s.knownAs.map((n) => String(n)) + : []; + const nameText = + knownAs.length > 0 + ? style.dim(` known as: ${knownAs.join(", ")}`) + : style.dim(" (no local name)"); + const liveText = s.live === true ? ` ${style.green("● live")}` : ""; + lines.push( + `* ${code(style, String(s.url))}${flagText}${nameText}${liveText}`, + ); + } + if (servers.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format auth/ema-status. */ +export function formatEmaStatusHuman( + status: { + clientConfigPath?: string; + configured?: boolean; + enabled?: boolean; + issuer?: string; + clientId?: string; + loginState?: string; + }, + style: Style = PLAIN, +): string { + const lines = [heading(style, "EMA (enterprise-managed auth):")]; + lines.push(style.dim(String(status.clientConfigPath ?? ""))); + if (status.configured !== true) { + lines.push( + "IdP: " + + style.dim( + "(not configured — set enterpriseManagedAuth in client.json or the web Inspector's Client Settings)", + ), + ); + return lines.join("\n"); + } + const client = status.clientId + ? style.dim(` (client: ${status.clientId})`) + : ""; + lines.push(`IdP: ${code(style, String(status.issuer ?? "?"))}${client}`); + lines.push(`Enabled: ${status.enabled === true ? "yes" : style.dim("no")}`); + const loginState = String(status.loginState ?? "none"); + const stateText = + loginState === "logged_in" + ? style.green(loginState) + : style.dim(loginState); + lines.push(`IdP session: ${stateText}`); + return lines.join("\n"); +} + +/** Format servers/list. */ +export function formatServersListHuman( + servers: unknown[], + style: Style = PLAIN, + source?: { kind: "catalog" | "config"; path: string }, +): string { + const lines = [heading(style, `Servers (${servers.length}):`)]; + // Say where the entries came from — shells with different --catalog / + // MCP_CATALOG_PATH / --config see different lists from the same daemon. + if (source) { + lines.push(`Source: ${style.dim(`${source.kind} ${source.path}`)}`); + } + for (const raw of servers) { + const s = raw as JsonObject; + const connectionName = + typeof s.connection === "string" && s.connection.length > 0 + ? s.connection + : undefined; + const connectionMark = connectionName + ? ` ${style.green(`@${connectionName}`)}${s.isMru === true ? style.green(" (MRU)") : ""}` + : ""; + lines.push( + `* ${code(style, String(s.name))} ${style.dim(`[${String(s.type)}]`)} ${style.dim(String(s.detail ?? ""))}${connectionMark}`, + ); + } + if (servers.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format servers/show (one catalog entry). */ +export function formatServerShowHuman( + server: JsonObject, + style: Style = PLAIN, + source?: { kind: "catalog" | "config"; path: string }, +): string { + const name = String(server.name ?? "?"); + const type = String(server.type ?? "?"); + const detail = String(server.detail ?? ""); + const header = `${heading(style, "Server")} ${code(style, name)} ${style.dim(`[${type}]`)}`; + const body: Record<string, unknown> = {}; + if (server.config && typeof server.config === "object") { + body.config = server.config; + } + if (server.settings && typeof server.settings === "object") { + body.settings = server.settings; + } + return [ + header, + detail ? style.dim(detail) : style.dim("(no detail)"), + // Same provenance line as servers/list — which file the entry came from. + ...(source + ? [`Source: ${style.dim(`${source.kind} ${source.path}`)}`] + : []), + JSON.stringify(body, null, 2), + ].join("\n"); +} + +/** Format connections/list. */ +/** + * A parked elicitation (`kind: "elicitation-pending"`): show the caller — + * typically an agent relaying to a human — what the server is asking and + * exactly how to answer it. Guidance lives here in human output only; the + * JSON payload stays data-only (the mcpdo skill carries the procedure). + */ +export function formatElicitationPendingHuman( + elicitation: ElicitationPendingInfo, + style: Style = PLAIN, +): string { + const id = elicitation.elicitationId; + const mode = elicitation.mode === "url" ? "url" : "form"; + const message = elicitation.message; + const lines = [ + `${heading(style, "Input required")} — ${code(style, elicitation.method)}${ + typeof elicitation.toolName === "string" + ? ` (tool ${code(style, elicitation.toolName)})` + : "" + } on ${code(style, `@${elicitation.connection}`)} is waiting on the user:`, + ` ${message}`, + ]; + if (mode === "url") { + const url = typeof elicitation.url === "string" ? elicitation.url : ""; + lines.push( + "", + "The user needs to open this link and complete it:", + // Only allowlisted schemes render as a clickable OSC 8 link; a server + // supplying file:/custom-handler URLs gets plain text (see sanitize.ts). + ` ${isSafeLinkTarget(url) ? style.link(url, url) : url}`, + "", + `When they're done, run: ${code(style, `elicitation/respond ${id} --done`)}`, + style.dim(`To give up instead: elicitation/respond ${id} --cancel`), + ); + } else { + const fields = parseFormSchema(elicitation.requestedSchema); + if (fields && fields.length > 0) { + lines.push("", heading(style, `Fields (${fields.length}):`)); + for (const field of fields) { + const kind = + field.kind === "number" && field.integer ? "integer" : field.kind; + const flags = field.required ? `${kind}, required` : kind; + const choices = + field.kind === "enum" || field.kind === "multiselect" + ? ` [${field.choices.map((c) => c.value).join(", ")}]` + : ""; + lines.push( + ` ${style.bold(field.name)} (${flags})${choices}${descSuffix(style, field.description)}`, + ); + } + } + lines.push( + "", + `Answer with: ${code(style, `elicitation/respond ${id} field:=value ...`)}`, + style.dim(`Or: elicitation/respond ${id} --decline | --cancel`), + ); + } + const expiresAt = Number(elicitation.expiresAt ?? 0); + if (Number.isFinite(expiresAt) && expiresAt > Date.now()) { + const minutes = Math.max(1, Math.round((expiresAt - Date.now()) / 60_000)); + lines.push( + style.dim( + `Unanswered, this expires (auto-cancels) in about ${minutes} minute${minutes === 1 ? "" : "s"}.`, + ), + ); + } + return lines.join("\n"); +} + +export function formatConnectionsListHuman( + connections: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Connections (${connections.length}):`)]; + for (const raw of connections) { + const s = raw as JsonObject; + const mru = s.isMru === true ? style.green(" (MRU)") : ""; + const era = + s.protocolEra !== undefined + ? style.dim(` [${String(s.protocolEra)}]`) + : ""; + const pending = + s.pendingAuth === true + ? s.pendingAuthSignedIn === true + ? style.yellow(" (signed in — completing on next use)") + : style.yellow(" (sign-in pending)") + : ""; + lines.push( + `* ${code(style, `@${String(s.name)}`)}${mru}${style.dim(` — ${String(s.serverIdentity ?? "")}`)}${era}${pending}`, + ); + } + if (connections.length === 0) lines.push(style.dim("(none — connect first)")); + return lines.join("\n"); +} + +/** Format a single connection info (connect / connections/use / connections/show). */ +export function formatConnectionInfoHuman( + connection: JsonObject, + style: Style = PLAIN, +): string { + const mru = connection.isMru === true ? style.green(" (MRU)") : ""; + const lines = [ + `${heading(style, "Connection")} ${code(style, `@${String(connection.name)}`)}${mru}`, + `Server: ${style.dim(String(connection.serverIdentity ?? ""))}`, + ]; + + // Connection details. `protocolEra` is now on every `ConnectionInfo` (#2298 + // follow-up), so it renders for plain `connect`/`connections/use` results too; + // `protocolVersion` and everything below it are `connections/show`-only. + const era = connection.protocolEra; + const protocolVersion = connection.protocolVersion; + if (era !== undefined || protocolVersion !== undefined) { + const versionSuffix = + protocolVersion !== undefined ? ` (${String(protocolVersion)})` : ""; + lines.push( + `Era: ${style.dim(`${String(era ?? "unknown")}${versionSuffix}`)}`, + ); + } + // Authorization snapshot (connect-time; omitted for stdio / no-auth + // servers — see `ConnectionInfo.auth`). + const auth = connection.auth as JsonObject | undefined; + if (auth !== undefined) { + const method = auth.method === "ema" ? "EMA" : "OAuth"; + const parts = [auth.authorized === true ? "authorized" : "not authorized"]; + if (typeof auth.scope === "string" && auth.scope !== "") { + parts.push(`scope: ${auth.scope}`); + } + if (typeof auth.clientId === "string" && auth.clientId !== "") { + parts.push(`client: ${auth.clientId}`); + } + if (typeof auth.idpSession === "string") { + parts.push(`IdP session: ${auth.idpSession}`); + } + lines.push(`Auth: ${method} ${style.dim(`(${parts.join("; ")})`)}`); + } + // Sign-in pending (non-TTY connect handed OAuth to the detached helper): + // the connection completes automatically on first use after sign-in. Once + // the tokens are on disk the read-only echoes flag it, so a poller knows + // the user's part is done. + if (connection.pendingAuth === true) { + lines.push( + connection.pendingAuthSignedIn === true + ? `Sign-in: ${style.green("completed")} ${style.dim("(connection finishes on next use)")}` + : `Sign-in: ${style.yellow("pending")} ${style.dim("(completes automatically after the user signs in)")}`, + ); + } + // Live transport state (`connections/show` only). Dormant is informational: + // the next op transparently re-dials with stored credentials. + if (typeof connection.transport === "string") { + const detail = + connection.transport === "dormant" + ? "dormant (reconnects on next use)" + : connection.transport; + lines.push(`Transport: ${style.dim(detail)}`); + } + const serverInfo = connection.serverInfo as JsonObject | undefined; + if (serverInfo?.name !== undefined) { + const version = + serverInfo.version !== undefined ? ` v${String(serverInfo.version)}` : ""; + lines.push( + `Server info: ${style.dim(`${String(serverInfo.name)}${version}`)}`, + ); + } + const capabilities = connection.capabilities as JsonObject | undefined; + if (capabilities !== undefined) { + const keys = Object.keys(capabilities); + lines.push( + `Capabilities: ${style.dim(keys.length > 0 ? keys.join(", ") : "(none)")}`, + ); + } + const supportedVersions = connection.supportedVersions; + if (Array.isArray(supportedVersions) && supportedVersions.length > 0) { + lines.push( + `Supported versions: ${style.dim(supportedVersions.join(", "))}`, + ); + } + if ( + typeof connection.instructions === "string" && + connection.instructions !== "" + ) { + lines.push(`Instructions: ${style.dim(connection.instructions)}`); + } + + return lines.join("\n"); +} + +/** Format tools/list --app-info lines. */ +export function formatAppInfoListHuman( + lines: unknown[], + style: Style = PLAIN, +): string { + const out = [heading(style, `App info (${lines.length} tools):`)]; + for (const raw of lines) { + const info = raw as JsonObject; + const name = String(info.toolName ?? "?"); + if (info.hasApp === true) { + const uri = String(info.resourceUri ?? "ui://?"); + out.push( + `* ${code(style, name)} — ${style.green("app")} (${formatUri(style, uri)})`, + ); + } else { + const err = + typeof info.resourceError === "string" + ? style.dim(` — ${info.resourceError}`) + : style.dim(" — no app"); + out.push(`* ${code(style, name)}${err}`); + } + } + return out.join("\n"); +} + +/** + * Format a `skills/list` result: one line per skill — frontmatter name, + * URI, description, and whether its file manifest is enumerated or dynamic. + */ +export function formatSkillsHuman( + skills: unknown[], + style: Style = PLAIN, +): string { + const lines = [heading(style, `Skills (${skills.length}):`)]; + for (const raw of skills) { + lines.push(formatSkillEntryLine(raw as JsonObject, style)); + } + if (skills.length === 0) lines.push(style.dim("(none)")); + return lines.join("\n"); +} + +/** Format a `skills/get` result (the envelope: `skill` + cache fields). */ +export function formatSkillGetHuman( + result: JsonObject, + style: Style = PLAIN, +): string { + const skill = (result.skill ?? {}) as JsonObject; + const lines = [heading(style, "Skill:"), formatSkillEntryLine(skill, style)]; + if (result.ttlMs !== undefined) + lines.push(style.dim(`ttlMs: ${String(result.ttlMs)}`)); + if (result.cacheScope !== undefined) + lines.push(style.dim(`cacheScope: ${String(result.cacheScope)}`)); + return lines.join("\n"); +} + +function formatSkillEntryLine(skill: JsonObject, style: Style): string { + const frontmatter = (skill.frontmatter ?? {}) as JsonObject; + const name = String(frontmatter.name ?? "?"); + const uri = String(skill.uri ?? ""); + const resources = skill.resources; + const manifest = + resources === "dynamic" + ? style.dim(" — dynamic resources") + : Array.isArray(resources) + ? style.dim(` — ${resources.length} file(s)`) + : ""; + return `* ${code(style, name)} (${formatUri(style, uri)})${descSuffix(style, frontmatter.description)}${manifest}`; +} + +/** + * Format `skills/list --verify` / `skills/get --verify` NDJSON lines. + * Each line is a {@link SkillVerifyReport}; the caller already computed the + * one-line stderr summary (`summarizeSkillVerification`) shared with the + * one-shot CLI, so this only renders the per-skill breakdown. + */ +export function formatSkillVerifyListHuman( + lines: unknown[], + style: Style = PLAIN, +): string { + const out = [heading(style, `Skill verification (${lines.length}):`)]; + for (const raw of lines) { + const report = raw as JsonObject; + const name = String(report.name ?? "?"); + const uri = String(report.uri ?? ""); + const outcome = report.outcome as string | undefined; + const conformance = asArray<JsonObject>(report.conformance); + const frontmatter = asArray<JsonObject>(report.frontmatter); + const files = asArray<JsonObject>(report.files); + const errorCount = [...conformance, ...frontmatter].filter( + (issue) => issue.severity === "error", + ).length; + const mismatchCount = files.filter( + (file) => file.status === "mismatch" || file.status === "read-error", + ).length; + const verdict = + outcome === "verified" + ? style.green("verified") + : outcome === "incomplete" + ? style.dim("incomplete") + : outcome === "unverifiable" + ? style.dim("unverifiable") + : style.red("failed"); + const detail = + outcome === "verified" + ? "" + : outcome === "incomplete" + ? style.dim( + ` — ${String(report.incomplete ?? "read bounds cut the walk short")}`, + ) + : outcome === "unverifiable" + ? // Nothing was checked, so "0 issue(s)" would misread as a pass + // narrowly missed; name the actual condition instead (#2405). + style.dim( + ` — advertised no digests (resources: "dynamic"), so integrity was not checked`, + ) + : style.dim( + ` — ${errorCount} issue(s), ${mismatchCount} file mismatch(es)`, + ); + out.push( + `* ${code(style, name)} (${formatUri(style, uri)}) — ${verdict}${detail}`, + ); + } + return out.join("\n"); +} + +/** Format a single app-info probe. */ +export function formatAppInfoHuman( + info: JsonObject, + style: Style = PLAIN, +): string { + const name = String(info.toolName ?? "?"); + if (info.hasApp === true) { + const lines = [ + `Tool ${code(style, name)} ${style.green("has an MCP App")}`, + `Resource: ${formatUri(style, String(info.resourceUri ?? ""))}`, + ]; + if (info.csp) lines.push(style.dim(`CSP: ${JSON.stringify(info.csp)}`)); + return lines.join("\n"); + } + const err = + typeof info.resourceError === "string" + ? info.resourceError + : "No MCP App UI resource (_meta.ui.resourceUri)."; + return `Tool ${code(style, name)} ${style.red("has no MCP App")}\n${style.dim(err)}`; +} + +/** Format a stream event for human display. */ +export function formatStreamEventHuman( + data: unknown, + style: Style = PLAIN, +): string { + if (!data || typeof data !== "object") return String(data); + const ev = data as JsonObject; + if (ev.type === "subscribed") { + return `${heading(style, "Subscribed:")} ${formatUri(style, String(ev.uri ?? ""))}`; + } + if (ev.type === "resources/updated") { + return `${heading(style, "Resource updated:")} ${formatUri(style, String(ev.uri ?? ""))}`; + } + // logging/tail MessageEntry-shaped + if (ev.direction === "notification" && ev.message) { + const msg = ev.message as JsonObject; + const params = (msg.params ?? {}) as JsonObject; + const level = String(params.level ?? "info"); + const logger = params.logger ? style.dim(` ${String(params.logger)}:`) : ""; + // Servers may log structured `data`; String() would render it as + // "[object Object]", so non-strings get JSON instead. + const raw = params.data ?? params.message ?? params; + const text = typeof raw === "string" ? raw : JSON.stringify(raw); + return `[${colorLevel(style, level)}]${logger} ${text}`; + } + return JSON.stringify(ev, null, 2); +} + +/** + * Dispatch human formatting for an RPC method result. + * Returns null when the caller should fall back to pretty JSON. + */ +export function formatRpcResultHuman( + method: string, + result: JsonObject, + style: Style = PLAIN, +): string | null { + switch (method) { + case "tools/list": + return formatToolsHuman(asArray(result.tools), style); + case "tools/call": + return formatCallToolResultHuman(result, style); + case "resources/list": + return formatResourcesHuman(asArray(result.resources), style); + case "resources/read": + return formatResourceReadHuman(result, style); + case "resources/templates/list": + return formatResourceTemplatesHuman( + asArray(result.resourceTemplates), + style, + ); + case "resources/unsubscribe": + return `${heading(style, "Unsubscribed:")} ${formatUri(style, String(result.uri ?? ""))}`; + case "prompts/list": + return formatPromptsHuman(asArray(result.prompts), style); + case "prompts/get": + return formatPromptResultHuman(result, style); + case "prompts/complete": + return formatCompletionsHuman(result, style); + case "initialize": + return formatInitializeHuman(result, style); + case "logging/setLevel": + return style.green("Logging level updated."); + case "tasks/list": + return formatTasksHuman(asArray(result.tasks), style); + case "tasks/get": + return formatTaskHuman(result.task, style); + case "tasks/cancel": + return `${heading(style, "Cancelled task:")} ${String(result.taskId ?? "")}`; + case "tasks/result": + return formatCallToolResultHuman(result, style); + case "roots/list": + case "roots/set": + return formatRootsHuman(asArray(result.roots), style); + case "skills/list": + return formatSkillsHuman(asArray(result.skills), style); + case "skills/get": + return formatSkillGetHuman(result, style); + default: + return null; + } +} diff --git a/clients/mcpdo/src/connection/mcp.ts b/clients/mcpdo/src/connection/mcp.ts new file mode 100644 index 0000000000..131e374238 --- /dev/null +++ b/clients/mcpdo/src/connection/mcp.ts @@ -0,0 +1,1736 @@ +import { Command, type Command as CommandType } from "commander"; +import { existsSync, readFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; +import type { JsonValue } from "@inspector/core/mcp/index.js"; +import type { + ElicitCapabilityMode, + InspectorServerSettings, + ServerProtocolEra, +} from "@inspector/core/mcp/types.js"; +import { + DEFAULT_MAX_FETCH_REQUESTS, + DEFAULT_TASK_TTL_MS, +} from "@inspector/core/mcp/types.js"; +import { + loadServerEntries, + parseHeaderPair, + parseKeyValuePair as parseEnvPair, + selectServerEntry, +} from "@inspector/core/mcp/node/index.js"; +import { type LoggingLevel } from "@modelcontextprotocol/client"; +import { getDefaultEnvironment } from "@modelcontextprotocol/client/stdio"; +import type { MCPServerConfig } from "@inspector/core/mcp/types.js"; +import { LoggingLevelSchema } from "@modelcontextprotocol/core"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { readInspectorVersion } from "@inspector/core/node/version.js"; +import { + getSecretStorageInfo, + warnAboutSecretStorage, +} from "@inspector/core/auth/node/secret-store-selection.js"; +import { callDaemon, ensureDaemon } from "../daemon/index.js"; +import type { + ConnectionInfo, + ConnectionShowResult, + ElicitationRespondParams, + ElicitationRespondResult, +} from "../daemon/protocol.js"; +import { + annotateServerEntriesWithConnections, + listServerEntries, + resolveServerListSource, + type ServerListSource, + showServerEntry, + summarizeServerConfig, +} from "@inspector/core/cli/handlers/servers-list.js"; +import { type OutputFormat } from "@inspector/core/cli/handlers/format-output.js"; +import { + DEFAULT_CONNECT_TIMEOUT_MS, + withConnectTimeout, +} from "@inspector/core/cli/handlers/connect-timeout.js"; +import { + CONNECTION_RPC_METHODS, + type MethodArgs, +} from "@inspector/core/cli/handlers/method-types.js"; +import { authorizeInFrontend } from "./authorize.js"; +import { + AUTH_HELPER_COMMAND, + obtainPendingAuthUrl, + runAuthHelper, +} from "./auth-helper.js"; +import { isCliAutoOpenForced } from "@inspector/core/cli/cli-oauth-navigation.js"; +import { emaLogin, emaLogout, getEmaStatus } from "./ema.js"; +import { + EMA_LOGIN_HELPER_COMMAND, + runEmaLoginHelper, + startPendingEmaLogin, +} from "./ema-login-helper.js"; +import { + assertJsonRoundTrips, + parseToolCallPositionals, + resolveToolCallArgs, +} from "./parse-tool-args.js"; +import { resolveCommandPath } from "./resolve-command.js"; +import { + dispatchConnectionRpc, + hoistAtConnection, + requireExplicitConnection, + stripAt, + writeRpcOutcome, +} from "./dispatch.js"; +import { writeConnectionOutput } from "./format-connection.js"; +import { + createPrivateBinding, + formatPrivateEnvExports, +} from "./private-env.js"; +import { + clearAllStoredAuth, + clearStoredAuth, + clearStoredAuthForRelogin, + listStoredAuth, + normalizeServerUrl, +} from "./stored-auth.js"; +import { + type AuthNameIndex, + buildAuthNameIndex, + resolveFriendlyName, +} from "./auth-names.js"; +import { + idpOAuthStorageKey, + parseIdpOAuthStorageKey, +} from "@inspector/core/auth/ema/storage.js"; +import { styleFromOpts } from "@inspector/core/cli/style.js"; +import { awaitableLog } from "@inspector/core/cli/utils/awaitable-log.js"; +import { createInterface } from "node:readline/promises"; + +function isDaemonUnreachable(error: unknown): boolean { + return ( + error instanceof CliExitCodeError && + error.envelope?.code === "daemon_unreachable" + ); +} + +/** + * `servers/show` with no name falls back to the MRU connection's entry name, + * under the same non-interactive gate as MRU connection targeting: an agent + * shell must name the entry explicitly. When there is no MRU (no daemon, or + * nothing connected), say so — a bare Commander "missing required argument" + * doesn't tell the user why a name is needed. + */ +async function resolveMruEntryName(): Promise<string> { + if (requireExplicitConnection()) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "servers/show requires an entry name in non-interactive mode. Pass one (see mcpdo servers/list).", + { code: "server_name_required" }, + ); + } + let connections: ConnectionInfo[] = []; + try { + const result = await callDaemon<{ connections: ConnectionInfo[] }>( + "connections/list", + {}, + ); + connections = result.connections; + } catch (error) { + if (!isDaemonUnreachable(error)) throw error; + } + const mru = connections.find((c) => c.isMru); + if (!mru) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "No entry name given and there is no most-recently-used connection to infer one from. Pass a catalog entry name (see mcpdo servers/list).", + { code: "server_name_required" }, + ); + } + return mru.name; +} + +/** + * The core "not found" error is source-agnostic by design; here we know both + * where the list came from and whether the user typed the name. An explicit + * name gets the source appended; an MRU-inferred name gets a full explanation, + * because "Server 'X' not found" is baffling when the user never typed X — + * connections are daemon-global while the catalog is per-shell, so the MRU + * connection's entry may simply not exist in this shell's catalog. + */ +function describeServerShowNotFound( + error: unknown, + entryName: string, + inferredFromMru: boolean, + source: ServerListSource | null, +): unknown { + if ( + !(error instanceof Error) || + !error.message.startsWith(`Server '${entryName}' not found`) + ) { + return error; + } + const where = source + ? `${source.kind} ${source.path}` + : "the resolved server list"; + if (inferredFromMru) { + return new CliExitCodeError( + EXIT_CODES.USAGE, + `The most-recently-used connection '${entryName}' has no entry in ${where}. Connections and catalog entries are separate — it may have been connected ad-hoc or from a different catalog. Pass an entry name (see mcpdo servers/list).`, + { code: "server_not_found" }, + ); + } + return new CliExitCodeError(EXIT_CODES.USAGE, `${error.message} (${where})`, { + code: "server_not_found", + }); +} + +/** Commander help/version exits — text already written; not real failures. */ +function isCommanderDisplayOnly(error: unknown): boolean { + if (error == null || typeof error !== "object") return false; + const code = (error as { code?: unknown }).code; + return ( + code === "commander.help" || + code === "commander.helpDisplayed" || + code === "commander.version" + ); +} + +type GlobalOpts = { + format?: OutputFormat; + plain?: boolean; + connection?: string; + catalog?: string; + config?: string; + storedAuthOnly?: boolean; +}; + +function outOpts(opts: GlobalOpts) { + return { + format: opts.format, + style: styleFromOpts({ plain: opts.plain === true, format: opts.format }), + // Matches dispatch.ts: a human is "present" only when output is text and a + // TTY is attached. `--format json` or no TTY means an agent is reading, so + // sign-in prompts switch to relay framing (see ConnectionWriteOpts). + interactive: + opts.format !== "json" && + (process.stdin.isTTY === true || process.stderr.isTTY === true), + }; +} + +const validLogLevels: LoggingLevel[] = Object.values(LoggingLevelSchema.enum); + +/** + * `--conn` is a documented shorthand for `--connection`. Expanding it at the + * argv level keeps a single option registration (one help entry, one + * GlobalOpts field) instead of two options merged at every consumption site. + * Expansion stops at the first `--`: everything after the separator belongs + * to the child process (`connect … -- <server args>`) and must pass through + * verbatim. + */ +export function expandConnAlias(argv: string[]): string[] { + const sep = argv.indexOf("--"); + const end = sep === -1 ? argv.length : sep; + return argv.map((arg, i) => + i >= end + ? arg + : arg === "--conn" + ? "--connection" + : arg.startsWith("--conn=") + ? `--connection=${arg.slice("--conn=".length)}` + : arg, + ); +} + +/** + * Connection-first CLI entry (`mcpdo`). Talks to the implicit connection daemon over + * IPC for connect/disconnect/connections and MCP RPCs; `servers/list` and + * `servers/show` are local (no daemon). + */ +/** + * Surface the secret-storage state once, at `connect`. + * + * The automatic store-selection banner is quieted process-wide in the + * front-end (see mcp-bin.ts) so it does not print on every command; this + * re-emits it exactly here, with `force`, so an agent still learns once — + * when a connection is being established — that secrets are in a plaintext + * fallback file (R5). + * + * It also catches the mcpdo-specific memory-store trap (R4): an explicit + * `MCP_INSPECTOR_SECRET_STORE=memory` keeps secrets in this process's RAM + * only, but mcpdo runs the daemon and the OAuth sign-in helper as separate + * processes — so a token one of them saves is invisible to the others and a + * sign-in that looked successful fails later with an opaque refresh error. + * Core's generic "lost on exit" caveat does not explain that, so we add a + * line that does, up front. + */ +export async function surfaceSecretStorageAtConnect(): Promise<void> { + const info = await getSecretStorageInfo(); + warnAboutSecretStorage(info, { force: true }); + if (info.kind === "memory") { + process.stderr.write( + `[mcpdo] MCP_INSPECTOR_SECRET_STORE=memory keeps secrets only in this ` + + `process's memory. mcpdo runs the daemon and the OAuth sign-in helper ` + + `as separate processes, so a token saved by one is invisible to the ` + + `others and sign-in cannot complete (you will see a credential-refresh ` + + `failure after signing in). Use "file" or "keyring" to persist across ` + + `mcpdo's processes.\n`, + ); + } +} + +/** + * Pin a stdio config's cwd, command, and environment to the CALLER's shell + * before it crosses the socket to the daemon. + * + * The daemon is a persistent detached process: it chdir()s at startup and + * keeps the environment of whichever mcpdo invocation first spawned it. Left + * unpinned, a relative cwd or bare command name — and every default-inherited + * env var the SDK transport fills in (PATH, HOME, SHELL, ...) — would resolve + * against that stale context instead of the shell that ran `connect`, which + * is what `mcpdo connect node ./server.js` means to the user. A cwd/env + * configured in the catalog entry (or flags) still wins; this only pins the + * defaults and resolves relative values. + */ +export function pinStdioConfigToCaller<T extends MCPServerConfig>( + config: T, +): T { + // `type` is optional on stdio configs (stdio is the implicit default), so + // narrow by excluding the URL transports rather than matching "stdio". + if (config.type === "sse" || config.type === "streamable-http") return config; + const resolved = resolveCommandPath(config.command); + return { + ...config, + cwd: path.resolve(config.cwd ?? process.cwd()), + ...(resolved !== config.command ? { command: resolved } : {}), + // The SDK transport spawns with {...getDefaultEnvironment(), ...env} + // evaluated in the DAEMON process; snapshotting the same default set + // here makes those fallbacks the caller's. + env: { ...getDefaultEnvironment(), ...config.env }, + }; +} + +export async function runMcp(argv?: string[]): Promise<void> { + const raw = argv ?? process.argv; + const { argv: rewritten, connectionFromAt } = hoistAtConnection( + expandConnAlias(raw), + ); + + const program = new Command(); + // Commander would print its own usage-error text before exitOverride runs; + // the JSON error envelope handleError writes is this CLI's single error + // channel, so the duplicate human line is suppressed (help output still + // prints through writeOut/writeErr as usual). + program.configureOutput({ outputError: () => {} }); + program.exitOverride((err) => { + // Help/version already printed. Always throw so Commander does not + // process.exit (which would tear down in-process tests); runMcp treats + // these as success. Bare `mcpdo` uses code `commander.help` with exitCode 1 + // — must not reach handleError as an ErrorEnvelope. + if (isCommanderDisplayOnly(err)) throw err; + if (err.exitCode !== 0) { + // Re-shape as a usage error so the envelope carries the right code and + // exit status instead of the generic "error". + throw new CliExitCodeError( + EXIT_CODES.USAGE, + err.message.replace(/^error: /, ""), + { code: "usage" }, + ); + } + }); + + program + .name("mcpdo") + .description( + "MCP Inspector connection CLI — connect once, run many commands against a named connection.\n\n" + + "Agent skill for mcpdo: install with `npx skills add modelcontextprotocol/inspector --skill mcpdo`, or see `agent-help` below.", + ) + .helpOption("-h, --help", "Display help for command") + .helpCommand("help [command]", "Display help for command") + .version(readInspectorVersion(import.meta.url), "--version") + .option( + "--format <format>", + "Output format: text (default; human-readable) or json (pretty-printed)", + (v: string): OutputFormat => { + if (v !== "text" && v !== "json") { + throw new Error(`--format must be 'text' or 'json'.`); + } + return v; + }, + ) + .option( + "--plain", + "Disable ANSI styling (color, bold/dim, hyperlinks) in human text output", + ) + .option( + "--connection <name>", + "Connection name (without required @). Overrides MRU / positional @name. `--conn` is a supported shorthand.", + ) + .option( + "--catalog <path>", + "Writable catalog file (default: ~/.mcp-inspector/mcp.json or MCP_CATALOG_PATH)", + ) + .option( + "--config <path>", + "Read-only connection config file (never written or seeded)", + ) + .option( + "--stored-auth-only", + "Never start interactive OAuth; use the shared store if present, otherwise fail.", + ); + + if (connectionFromAt) { + program.setOptionValue("connection", connectionFromAt); + } + + program + .command("servers/list") + .description( + "List catalog/config server entries (marks live connections when the daemon is running; no MCP connection)", + ) + .action(async () => { + const opts = program.opts<GlobalOpts>(); + const envCatalog = process.env.MCP_CATALOG_PATH; + const serverOptions = { + catalogPath: opts.catalog?.trim() || envCatalog, + configPath: opts.config?.trim() || undefined, + }; + const entries = await listServerEntries(serverOptions); + const source = resolveServerListSource(serverOptions); + let connections: ConnectionInfo[] = []; + try { + const result = await callDaemon<{ connections: ConnectionInfo[] }>( + "connections/list", + {}, + ); + connections = result.connections; + } catch (error) { + if (!isDaemonUnreachable(error)) throw error; + } + await writeConnectionOutput(outOpts(opts), { + kind: "servers/list", + servers: annotateServerEntriesWithConnections(entries, connections), + ...(source && { source }), + }); + }); + + program + .command("servers/show") + .description( + "Show one catalog/config entry in detail (no MCP connection; secrets redacted)", + ) + .argument( + "[name]", + "Catalog entry name (defaults to the MRU connection's entry on an interactive TTY)", + ) + .action(async (name: string | undefined) => { + const opts = program.opts<GlobalOpts>(); + const envCatalog = process.env.MCP_CATALOG_PATH; + const serverOptions = { + catalogPath: opts.catalog?.trim() || envCatalog, + configPath: opts.config?.trim() || undefined, + }; + const explicitName = stripAt(name?.trim() || undefined); + const entryName = explicitName ?? (await resolveMruEntryName()); + const source = resolveServerListSource(serverOptions); + let entry; + try { + entry = await showServerEntry(entryName, serverOptions); + } catch (error) { + throw describeServerShowNotFound( + error, + entryName, + explicitName === undefined, + source, + ); + } + await writeConnectionOutput(outOpts(opts), { + kind: "servers/show", + server: entry, + ...(source && { source }), + }); + }); + + registerConnect(program); + registerConnectionAdmin(program); + registerAuthCommands(program); + registerRpcCommands(program); + registerElicitationCommands(program); + // Keep infra commands last in --help (just before Commander's built-in help). + registerDaemonCommands(program); + registerPrivateCommand(program); + registerAgentHelpCommand(program); + + try { + await program.parseAsync(rewritten); + } catch (error) { + if (isCommanderDisplayOnly(error)) return; + throw error; + } +} + +function registerConnect(program: CommandType): void { + program + .command("connect") + .description( + "Connect a catalog entry or ad-hoc target as a named connection", + ) + .argument( + "[target...]", + "Catalog entry name, or command/URL (use -- for command args). A " + + "single bare word is a catalog name; a URL, a path (contains / or " + + "starts with . or ~), multiple tokens, or --transport force an " + + "ad-hoc target.", + ) + .option("--server <name>", "Server name from catalog/config") + .option( + "-e <env>", + "Environment variables for the server (KEY=VALUE)", + parseEnvPair, + {}, + ) + .option("--cwd <path>", "Working directory for stdio server process") + .option( + "--transport <type>", + "Transport type (sse, http, or stdio)", + (value: string) => { + const valid = ["sse", "http", "stdio"]; + if (!valid.includes(value)) { + throw new Error(`Invalid transport type: ${value}`); + } + return value as "sse" | "http" | "stdio"; + }, + ) + .option("--server-url <url>", "Server URL for SSE/HTTP transport") + .option( + "--header <headers...>", + 'HTTP headers as "HeaderName: Value" pairs', + parseHeaderPair, + {}, + ) + .option( + "--connect-timeout <ms>", + `Connection timeout in ms (default ${DEFAULT_CONNECT_TIMEOUT_MS} for ad-hoc)`, + (v: string) => { + const n = Number(v); + if (!Number.isFinite(n) || n < 0) { + throw new Error(`--connect-timeout must be a non-negative number.`); + } + return n; + }, + ) + .option( + "--era <era>", + "Protocol era to negotiate: legacy (default), auto, or modern. " + + "Overrides the catalog/config entry's protocolEra; the only way to " + + "set it for an ad-hoc target, which has no config entry of its own.", + (value: string) => { + const valid: ServerProtocolEra[] = ["legacy", "auto", "modern"]; + if (!valid.includes(value as ServerProtocolEra)) { + throw new Error( + `Invalid --era: ${value}. Use legacy, auto, or modern.`, + ); + } + return value as ServerProtocolEra; + }, + ) + .option( + "-r, --relogin", + "Ignore stored OAuth for this connect (HTTP/SSE URL keys only); interactive login runs only if the server requires auth. No-op for stdio / servers with no stored entry", + ) + .option( + "--elicit <mode>", + "Elicitation capability to advertise: off, url, form, or both (default). " + + "Overrides the catalog/config entry's elicitCapability; the only way to " + + "set it for an ad-hoc target, which has no config entry of its own. Use " + + "off when the caller of mcpdo can't handle an elicitation request, so the " + + "server sees no elicitation capability and can fall back on its own.", + (value: string) => { + const valid: ElicitCapabilityMode[] = ["off", "url", "form", "both"]; + if (!valid.includes(value as ElicitCapabilityMode)) { + throw new Error( + `Invalid --elicit: ${value}. Use off, url, form, or both.`, + ); + } + return value as ElicitCapabilityMode; + }, + ) + .option( + "--ema", + "Treat the server as enterprise-managed (EMA): mint tokens from the " + + "signed-in enterprise IdP session instead of standard OAuth. " + + "Overrides the catalog/config entry's oauth.enterpriseManaged. " + + "Requires install-level IdP config (see auth/ema-status) and " + + "per-server OAuth client id/secret from the catalog entry, so it " + + "cannot be used with an ad-hoc target.", + ) + .action(async (target: string[], cmdOpts) => { + const opts = program.opts<GlobalOpts>(); + const { name: positionalConnection, rest } = + splitConnectionTarget(target); + const connectionName = + stripAt(opts.connection) ?? + positionalConnection ?? + cmdOpts.server?.trim() ?? + rest[0]; + + if (!connectionName) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "connect requires a catalog entry name, --server <name>, or an ad-hoc target.", + { code: "usage" }, + ); + } + + await surfaceSecretStorageAtConnect(); + + const relogin = cmdOpts.relogin === true; + if (relogin && opts.storedAuthOnly) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "--relogin cannot be combined with --stored-auth-only", + { code: "usage" }, + ); + } + + const adHoc = + rest.length > 1 || + Boolean(cmdOpts.transport) || + Boolean(cmdOpts.serverUrl?.trim()) || + (rest.length === 1 && + (looksLikeUrl(rest[0]!) || looksLikePath(rest[0]!))); + + // EMA needs the resource server's OAuth client id/secret, which only a + // catalog/config entry can carry (oauth.clientId / oauth.clientSecret). + // An ad-hoc target has no entry and this CLI deliberately offers no + // secret-bearing flags, so the flow would only fail later with an + // opaque error — reject up front with actionable guidance instead. + if (cmdOpts.ema === true && adHoc) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "--ema cannot be used with an ad-hoc target: EMA requires per-server " + + "OAuth client id/secret from a catalog entry. Add the server to a " + + "catalog with oauth.clientId and oauth.clientSecret (and " + + "oauth.enterpriseManaged), then connect by entry name.", + { code: "usage" }, + ); + } + + const envCatalog = adHoc ? undefined : process.env.MCP_CATALOG_PATH; + const serverOptions = { + catalogPath: opts.catalog?.trim() || envCatalog, + configPath: opts.config?.trim() || undefined, + target: adHoc ? (rest.length > 0 ? rest : undefined) : undefined, + transport: cmdOpts.transport as "sse" | "http" | "stdio" | undefined, + serverUrl: cmdOpts.serverUrl as string | undefined, + // Resolve --cwd against the CALLER's working directory. The daemon + // that spawns the stdio server inherits an unrelated cwd (see + // daemon/run.ts), so a relative --cwd must be pinned here. + cwd: cmdOpts.cwd ? path.resolve(cmdOpts.cwd as string) : undefined, + env: cmdOpts.e as Record<string, string> | undefined, + headers: cmdOpts.header as Record<string, string> | undefined, + }; + + const selectName = adHoc + ? undefined + : ((cmdOpts.server as string | undefined)?.trim() ?? rest[0]); + + const entries = await loadServerEntries(serverOptions); + const selected = selectServerEntry(entries, selectName); + // Pin cwd/command/env to THIS shell before the config crosses to the + // persistent daemon, whose own cwd and environment are stale (they + // belong to whichever invocation first spawned it). + const serverConfig = pinStdioConfigToCaller(selected.config); + const serverSettings = withEmaOverride( + withElicitOverride( + withEraOverride( + withConnectTimeout( + selected.settings, + (cmdOpts.connectTimeout as number | undefined) ?? + (adHoc ? DEFAULT_CONNECT_TIMEOUT_MS : undefined), + ), + cmdOpts.era as ServerProtocolEra | undefined, + ), + cmdOpts.elicit as ElicitCapabilityMode | undefined, + ), + cmdOpts.ema === true ? true : undefined, + ); + const { detail } = summarizeServerConfig(serverConfig); + const name = stripAt(connectionName)!; + + if (relogin && "url" in serverConfig && serverConfig.url) { + await clearStoredAuthForRelogin(serverConfig.url); + } + + const { socketPath } = await ensureDaemon(); + const connectParams = { + name, + serverConfig, + serverSettings, + serverIdentity: detail, + }; + + let result: ConnectionInfo; + try { + result = await callDaemon<ConnectionInfo>("connect", connectParams, { + socketPath, + // The daemon enforces the configured connect timeout (which may be + // 0 = unlimited or exceed 60s); no fixed local deadline. + timeoutMs: 0, + }); + } catch (error) { + if ( + !(error instanceof CliExitCodeError) || + error.envelope?.code !== "auth_required" + ) { + throw error; + } + if (opts.storedAuthOnly) { + throw error; + } + // Agent path: no TTY anywhere means the blocking interactive flow is + // hostile — the URL sits invisible in a buffered pipe and a timeout + // kill would tear down the callback listener the link points at. + // Hand the flow to a detached helper, register the connection as + // pending intent, and exit with the link so the caller can relay it. + // `MCP_AUTO_OPEN_ENABLED=true` (forced auto-open) keeps the blocking + // flow: that's an explicit unattended-automation opt-in. + const humanPresent = + process.stdin.isTTY === true || process.stderr.isTTY === true; + if (!humanPresent && !isCliAutoOpenForced()) { + const authUrl = await obtainPendingAuthUrl( + serverConfig, + serverSettings, + ); + // The dial re-attempt is cheap (it fails auth_required again) but + // makes the daemon register the pending entry, so + // `connections/show @name` polls sign-in state (completing the + // connection itself once tokens land) and any real op completes it + // via revive. + const { socketPath: pendingSocketPath } = await ensureDaemon(); + const pending = await callDaemon<ConnectionInfo>( + "connect", + { ...connectParams, pendingOnAuthRequired: true }, + { socketPath: pendingSocketPath, timeoutMs: 0 }, + ); + await writeConnectionOutput(outOpts(opts), { + kind: "connection", + connection: pending, + ...(pending.pendingAuth === true && { authUrl }), + }); + return; + } + await authorizeInFrontend(serverConfig, serverSettings, { + storedAuthOnly: false, + }); + // Interactive OAuth can run well past the daemon's idle timeout + // (60s, armed while it holds zero connections) — a slow human login + // (SSO, MFA) can leave the daemon we ensured above already exited. + // Re-ensure so the retry lands on a live daemon instead of a stale + // socket; ensureDaemon() is a no-op when the existing one still + // answers pings. + const { socketPath: freshSocketPath } = await ensureDaemon(); + result = await callDaemon<ConnectionInfo>("connect", connectParams, { + socketPath: freshSocketPath, + timeoutMs: 0, + }); + } + await writeConnectionOutput(outOpts(opts), { + kind: "connection", + connection: result, + }); + }); +} + +/** + * Build the friendly-name index for the auth commands: catalog entries + * (per-shell) joined with the daemon's live connections (daemon-global). The + * daemon call is best-effort — offline just means catalog names only and no + * live marker, so `auth/list` / `auth/clear` keep working with the daemon down. + */ +async function collectAuthNameIndex(opts: GlobalOpts): Promise<AuthNameIndex> { + const envCatalog = process.env.MCP_CATALOG_PATH; + const serverOptions = { + catalogPath: opts.catalog?.trim() || envCatalog, + configPath: opts.config?.trim() || undefined, + }; + const entries = await listServerEntries(serverOptions); + let connections: ConnectionInfo[] = []; + try { + const result = await callDaemon<{ connections: ConnectionInfo[] }>( + "connections/list", + {}, + ); + connections = result.connections; + } catch (error) { + if (!isDaemonUnreachable(error)) throw error; + } + return buildAuthNameIndex(entries, connections); +} + +function registerAuthCommands(program: CommandType): void { + // Internal detached sign-in helper for the non-TTY connect path (see + // auth-helper.ts). Hidden: params arrive as JSON on stdin, never argv. + program + .command(AUTH_HELPER_COMMAND, { hidden: true }) + .description("Internal: complete an OAuth sign-in (params JSON on stdin)") + .action(async () => { + await runAuthHelper(); + }); + + // Same park machinery for the EMA IdP login (see ema-login-helper.ts). + program + .command(EMA_LOGIN_HELPER_COMMAND, { hidden: true }) + .description("Internal: complete an EMA IdP sign-in") + .action(async () => { + await runEmaLoginHelper(); + }); + + program + .command("auth/list") + .description( + "List servers in the shared OAuth store, annotated with catalog/connection names", + ) + .action(async () => { + const opts = program.opts<GlobalOpts>(); + const list = await listStoredAuth(); + const index = await collectAuthNameIndex(opts); + const servers = list.servers.map((s) => { + // EMA IdP login records live in the same store, keyed `ema-idp:<issuer>`. + // They are not servers, so they never carry a catalog/connection name — + // surface them as their own category rather than an unnamed server. + const issuer = parseIdpOAuthStorageKey(s.url); + if (issuer !== null) { + return { ...s, idp: true, issuer }; + } + const refs = index.urlToNames.get(s.url) ?? []; + return { + ...s, + ...(refs.length > 0 && { knownAs: refs.map((r) => r.name) }), + ...(refs.some((r) => r.isLive) && { live: true }), + }; + }); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/list", + list: { oauthStatePath: list.oauthStatePath, servers }, + }); + }); + + program + .command("auth/clear") + .description( + "Clear stored OAuth state for one server (URL or catalog/connection name from auth/list) or all entries", + ) + .argument("[key]", "Server URL or catalog/connection name from auth/list") + .option("--all", "Clear every stored OAuth server entry") + .option("--yes", "Skip confirmation when using --all") + .action(async (key: string | undefined, cmdOpts) => { + const opts = program.opts<GlobalOpts>(); + const all = cmdOpts.all === true; + if (all && key?.trim()) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "auth/clear: pass a key or --all, not both", + { code: "usage" }, + ); + } + if (!all && !key?.trim()) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "auth/clear requires a server URL key (from auth/list) or --all", + { code: "usage" }, + ); + } + if (all) { + if (!cmdOpts.yes) { + if (!process.stdin.isTTY || !process.stdout.isTTY) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "auth/clear --all requires --yes in non-interactive mode", + { code: "usage" }, + ); + } + /* v8 ignore next 22 -- interactive y/N confirm needs a real TTY */ + const rl = createInterface({ + input: process.stdin, + output: process.stderr, + }); + try { + const answer = await rl.question( + "Clear ALL stored OAuth credentials? [y/N] ", + ); + const ok = + answer.trim().toLowerCase() === "y" || + answer.trim().toLowerCase() === "yes"; + if (!ok) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "auth/clear --all cancelled", + { code: "usage" }, + ); + } + } finally { + rl.close(); + } + } + const result = await clearAllStoredAuth(); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/clear", + result: { all: true, cleared: result.cleared }, + }); + return; + } + const trimmed = key!.trim(); + // A URL argument — or any string that is already an exact key in the + // store (e.g. an `ema-idp:<issuer>` EMA login key, which auth/list shows + // verbatim) — keeps the direct store-key path. Only a bare friendly name + // falls through to catalog + live-connection resolution. + const storedKeys = (await listStoredAuth()).servers.map((s) => s.url); + if (storedKeys.includes(trimmed)) { + const result = await clearStoredAuth(trimmed); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/clear", + result: { url: result.url }, + }); + return; + } + if (/^https?:\/\//i.test(trimmed)) { + // auth/list renders an EMA IdP login by its bare issuer URL (prefix + // hidden), so accept that issuer here and map it back to the prefixed + // store key — preferring a real server entry on the off chance both + // exist. Otherwise fall through to clearStoredAuth's own URL handling. + const idpKey = idpOAuthStorageKey(trimmed); + const target = + storedKeys.includes(idpKey) && + !storedKeys.includes(normalizeServerUrl(trimmed)) + ? idpKey + : trimmed; + const result = await clearStoredAuth(target); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/clear", + result: { url: result.url }, + }); + return; + } + const index = await collectAuthNameIndex(opts); + const resolution = resolveFriendlyName(trimmed, index); + switch (resolution.kind) { + case "no-url": + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `'${trimmed}' is a stdio server — it has no stored OAuth entry to clear.`, + { code: "usage" }, + ); + case "ambiguous": + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `'${trimmed}' maps to multiple server URLs (${resolution.urls.join( + ", ", + )}). Pass the explicit URL from auth/list.`, + { code: "usage" }, + ); + case "unknown": + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `No server named '${trimmed}'. Use auth/list or servers/list to see names, or pass a URL.`, + { code: "usage" }, + ); + case "url": { + const result = await clearStoredAuth(resolution.url); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/clear", + result: { url: result.url, clearedByName: trimmed }, + }); + return; + } + } + }); + + program + .command("auth/ema-status") + .description( + "Show enterprise-managed auth (EMA) configuration and IdP login state", + ) + .action(async () => { + const opts = program.opts<GlobalOpts>(); + const status = await getEmaStatus(); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/ema-status", + status, + }); + }); + + program + .command("auth/ema-login") + .description( + "Sign in to the enterprise IdP (EMA); subsequent connects to EMA servers mint tokens silently from this connection", + ) + .option( + "-r, --relogin", + "Clear the existing IdP session (and EMA server tokens) and sign in fresh", + ) + .action(async (cmdOpts) => { + const opts = program.opts<GlobalOpts>(); + const relogin = cmdOpts.relogin === true; + // Same split as connect: with no TTY anywhere the blocking flow's IdP + // URL sits invisible in a buffered pipe — park the flow on a detached + // helper and exit with the link so the caller can relay it (then poll + // auth/ema-status). Forced auto-open keeps the blocking flow. + const humanPresent = + process.stdin.isTTY === true || process.stderr.isTTY === true; + const result = + !humanPresent && !isCliAutoOpenForced() + ? await startPendingEmaLogin({ relogin }) + : await emaLogin({ relogin }); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/ema-login", + result, + }); + }); + + program + .command("auth/ema-logout") + .description( + "Sign out of the enterprise IdP locally and clear EMA-minted server tokens (prints the IdP end-session URL when available)", + ) + .action(async () => { + const opts = program.opts<GlobalOpts>(); + const result = await emaLogout(); + await writeConnectionOutput(outOpts(opts), { + kind: "auth/ema-logout", + result, + }); + }); +} + +function registerConnectionAdmin(program: CommandType): void { + program + .command("disconnect") + .description("Disconnect a connection (MRU when omitted on a TTY)") + .argument("[connection]", "Optional @name / name to disconnect") + .option( + "-c, --clear-auth", + "Also clear this server's stored OAuth tokens so the next connect re-triggers sign-in (HTTP/SSE URL keys only; no-op for stdio / no stored entry)", + ) + .action(async (connectionArg: string | undefined, cmdOpts) => { + const opts = program.opts<GlobalOpts>(); + const name = stripAt(opts.connection) ?? stripAt(connectionArg); + const { socketPath } = await ensureDaemon(); + const result = await callDaemon<{ name: string; serverUrl?: string }>( + "disconnect", + { + name, + requireExplicit: requireExplicitConnection(), + }, + { socketPath }, + ); + // Clear stored auth only after the daemon tore the connection down, so + // the live client's in-memory tokens can't linger against a cleared + // store. No-op when the server has no URL key (stdio / no stored entry). + let clearedAuthUrl: string | undefined; + if (cmdOpts.clearAuth === true && result.serverUrl) { + await clearStoredAuthForRelogin(result.serverUrl); + clearedAuthUrl = result.serverUrl; + } + await writeConnectionOutput(outOpts(opts), { + kind: "disconnect", + name: result.name, + ...(clearedAuthUrl && { clearedAuthUrl }), + }); + }); + + program + .command("connections/list") + .description("List open connections (marks MRU); does not start the daemon") + .action(async () => { + const opts = program.opts<GlobalOpts>(); + try { + const result = await callDaemon<{ connections: ConnectionInfo[] }>( + "connections/list", + {}, + ); + await writeConnectionOutput(outOpts(opts), { + kind: "connections/list", + connections: result.connections, + }); + } catch (error) { + if (isDaemonUnreachable(error)) { + await writeConnectionOutput(outOpts(opts), { + kind: "connections/list", + connections: [], + }); + return; + } + throw error; + } + }); + + program + .command("connections/use") + .description("Set the MRU connection without an MCP RPC") + .argument("<connection>", "Connection @name / name") + .action(async (connectionArg: string) => { + const opts = program.opts<GlobalOpts>(); + const name = stripAt(opts.connection) ?? stripAt(connectionArg); + if (!name) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "connections/use requires a connection name", + { code: "usage" }, + ); + } + const { socketPath } = await ensureDaemon(); + const result = await callDaemon<ConnectionInfo>( + "connections/use", + { name }, + { socketPath }, + ); + await writeConnectionOutput(outOpts(opts), { + kind: "connection", + connection: result, + }); + }); + + program + .command("connections/show") + .description( + "Show connection + connection details: server info, capabilities, negotiated protocol era (defaults to MRU)", + ) + .argument("[connection]", "Connection @name / name (defaults to MRU)") + .action(async (connectionArg: string | undefined) => { + const opts = program.opts<GlobalOpts>(); + const name = stripAt(opts.connection) ?? stripAt(connectionArg); + const { socketPath } = await ensureDaemon(); + const result = await callDaemon<ConnectionShowResult>( + "connections/show", + { name, requireExplicit: requireExplicitConnection() }, + { socketPath }, + ); + // `authUrl` rides the result, so `--format json` always carries it (the + // JSON payload spreads the result object). The human URL block, however, + // is agent-relay text — show it only when no human is present; a human + // at a TTY sees the plain pending status in the connection info. + const humanPresent = + process.stdin.isTTY === true || process.stderr.isTTY === true; + await writeConnectionOutput(outOpts(opts), { + kind: "connection", + connection: result, + ...(!humanPresent && result.authUrl && { authUrl: result.authUrl }), + }); + }); +} + +function registerDaemonCommands(program: CommandType): void { + const daemon = program.command("daemon").description("Daemon control"); + + daemon + .command("status") + .description("Show daemon pid, socket, and connections (does not start it)") + .action(async () => { + const opts = program.opts<GlobalOpts>(); + try { + const result = await callDaemon("daemon/status", {}); + await writeConnectionOutput(outOpts(opts), { + kind: "daemon/status", + status: result as Record<string, unknown>, + }); + } catch (error) { + if (isDaemonUnreachable(error)) { + await writeConnectionOutput(outOpts(opts), { + kind: "daemon/status", + status: { + running: false, + message: "Daemon is not running.", + }, + }); + return; + } + throw error; + } + }); + + daemon + .command("stop") + .description("Stop the daemon and disconnect all connections") + .action(async () => { + const opts = program.opts<GlobalOpts>(); + try { + const result = await callDaemon("daemon/stop", {}); + await writeConnectionOutput(outOpts(opts), { + kind: "daemon/stop", + result: result as Record<string, unknown>, + }); + } catch (error) { + if (isDaemonUnreachable(error)) { + await writeConnectionOutput(outOpts(opts), { + kind: "daemon/stop", + result: { + stopping: false, + message: "Daemon was not running.", + }, + }); + return; + } + throw error; + } + }); +} + +function registerPrivateCommand(program: CommandType): void { + program + .command("private") + .description( + 'Print shell exports for a private daemon (eval "$(mcpdo private)"). ' + + "Later mcpdo commands in that shell use an isolated, token-gated daemon.", + ) + .action(async () => { + const binding = createPrivateBinding(); + await awaitableLog(formatPrivateEnvExports(binding)); + }); +} + +/** + * Locates the repo-root `skills/mcpdo/SKILL.md` relative to this module. + * Tries both the built (bundled single-file, `clients/mcpdo/build/`) and + * source (`clients/mcpdo/src/connection/`) layouts, since the two sit at + * different depths from the repo root. + */ +function resolveAgentSkillPath(): string | undefined { + const here = path.dirname(fileURLToPath(import.meta.url)); + const candidates = [ + path.resolve(here, "../../../skills/mcpdo/SKILL.md"), + path.resolve(here, "../../../../skills/mcpdo/SKILL.md"), + ]; + return candidates.find((candidate) => existsSync(candidate)); +} + +/** + * The always-on awareness snippet for a project's CLAUDE.md / AGENTS.md. + * Skills are pull-only (an agent sees just the description until something + * triggers a load), so task-shaped prompts that never mention MCP won't + * activate the skill; a line in an always-in-context instructions file is + * the reliable mechanism for standing awareness. + */ +const AGENT_INSTRUCTIONS_SNIPPET = + "mcpdo manages connections to additional MCP servers. Treat the tools, " + + "resources, and prompts on its open connections (`mcpdo connections/list`) " + + "as part of your available toolset — check them before deciding a task " + + "can't be done, and include them when asked what tools or MCP servers " + + "you have. Run `mcpdo agent-help` for the usage guide.\n"; + +/** Returns SKILL.md content with the YAML frontmatter block removed. */ +function stripFrontmatter(content: string): string { + if (!content.startsWith("---\n")) return content; + const end = content.indexOf("\n---\n", 4); + if (end === -1) return content; + return content.slice(end + 5).replace(/^\n+/, ""); +} + +function registerAgentHelpCommand(program: CommandType): void { + program + .command("agent-help") + .description( + "Print an agent-oriented usage guide, comparable to the mcpdo SKILL. " + + "Agents should read this before using other mcpdo commands.", + ) + .option( + "--skill", + "Print the full usage guide — the mcpdo SKILL.md body, usable as if " + + "the skill had been loaded (default when no option is given)", + ) + .option( + "--instructions", + "Print a short snippet to append to a project's CLAUDE.md/AGENTS.md " + + "so agents treat mcpdo connections as part of their toolset on " + + "every turn (skills and this guide load only on demand). " + + "Pipeable: mcpdo agent-help --instructions >> AGENTS.md", + ) + .option( + "--skill-path", + "Print the path of the installable SKILL.md file — this guide plus " + + "its skill frontmatter — for skill runtimes " + + "(e.g. copy into ~/.claude/skills/mcpdo/)", + ) + .action( + async (o: { + skill?: boolean; + instructions?: boolean; + skillPath?: boolean; + }) => { + const picked = [o.skill, o.instructions, o.skillPath].filter( + (v) => v === true, + ).length; + if (picked > 1) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "--skill, --instructions, and --skill-path are mutually exclusive.", + { code: "agent_help_flag_conflict" }, + ); + } + if (o.instructions === true) { + await awaitableLog(AGENT_INSTRUCTIONS_SNIPPET); + return; + } + const skillPath = resolveAgentSkillPath(); + if (!skillPath) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "Could not locate skills/mcpdo/SKILL.md relative to this install.", + { code: "agent_help_not_found" }, + ); + } + if (o.skillPath === true) { + await awaitableLog(skillPath + "\n"); + return; + } + await awaitableLog(stripFrontmatter(readFileSync(skillPath, "utf8"))); + }, + ); +} + +function registerRpcCommands(program: CommandType): void { + for (const method of CONNECTION_RPC_METHODS) { + const cmd = program + .command(method) + .description(`MCP ${method} against the current connection`); + + cmd.option( + "--metadata <pairs...>", + "General metadata as key=value pairs", + parseKeyValue, + {}, + ); + + switch (method) { + case "tools/list": + cmd.option("--app-info", "Emit one NDJSON app-info line per tool"); + cmd.action(async (o) => { + await runRpc(program, method, { + appInfo: o.appInfo === true, + metadata: o.metadata, + }); + }); + break; + case "tools/call": + cmd + .argument("[toolName]", "Tool name") + .argument( + "[toolArgs...]", + "Arguments as key:=value pairs or a JSON object", + ) + .option("--tool-name <name>", "Tool name") + .option( + "--tool-arg <pairs...>", + "Tool argument as key=value pair (alternative to key:=value positionals)", + parseKeyValue, + {}, + ) + .option( + "--tool-args-json <json>", + "Tool arguments as a JSON object (alternative to inline JSON positional)", + ) + .option( + "--tool-metadata <pairs...>", + "Tool-specific metadata", + parseKeyValue, + {}, + ) + .option("--task", "Task-augmented tool call (callToolStream)") + .option("--app-info", "Probe MCP App metadata only"); + cmd.action( + async ( + toolNamePos: string | undefined, + toolArgsPos: string[] | undefined, + o, + ) => { + const { toolName, toolArg } = resolveToolCallArgs({ + toolNameFlag: o.toolName as string | undefined, + toolNamePos, + toolArgsPos, + toolArgFlag: (o.toolArg ?? {}) as Record<string, JsonValue>, + toolArgsJson: o.toolArgsJson as string | undefined, + }); + await runRpc(program, method, { + toolName, + toolArg, + toolMeta: o.toolMetadata, + metadata: o.metadata, + task: o.task === true, + appInfo: o.appInfo === true, + }); + }, + ); + break; + case "resources/read": + case "resources/subscribe": + case "resources/unsubscribe": + cmd + .argument("[uri]", "Resource URI") + .option("--uri <uri>", "Resource URI"); + cmd.action(async (uriPos: string | undefined, o) => { + await runRpc(program, method, { + uri: (o.uri as string | undefined) ?? uriPos, + metadata: o.metadata, + }); + }); + break; + case "skills/list": + cmd.option( + "--verify", + "Run the SEP-2640 conformance and digest checks over every skill returned", + ); + cmd.action(async (o) => { + await runRpc(program, method, { + verify: o.verify === true, + metadata: o.metadata, + }); + }); + break; + case "skills/get": + cmd + .argument("[uri]", "Skill URI") + .option("--uri <uri>", "Skill URI") + .option( + "--verify", + "Run the SEP-2640 conformance and digest checks over this skill", + ); + cmd.action(async (uriPos: string | undefined, o) => { + await runRpc(program, method, { + uri: (o.uri as string | undefined) ?? uriPos, + verify: o.verify === true, + metadata: o.metadata, + }); + }); + break; + case "prompts/get": + cmd + .argument("[promptName]", "Prompt name") + .option("--prompt-name <name>", "Prompt name") + .option( + "--prompt-args <pairs...>", + "Prompt arguments", + parseKeyValue, + {}, + ); + cmd.action(async (promptPos: string | undefined, o) => { + await runRpc(program, method, { + promptName: (o.promptName as string | undefined) ?? promptPos, + promptArgs: (o.promptArgs ?? {}) as Record<string, JsonValue>, + metadata: o.metadata, + }); + }); + break; + case "prompts/complete": + cmd + .option("--complete-ref-type <type>", "ref/prompt or ref/resource") + .option("--complete-ref <ref>", "Prompt name or resource URI") + .option("--complete-arg-name <name>", "Argument name") + .option("--complete-arg-value <value>", "Partial value", ""); + cmd.action(async (o) => { + const refType = o.completeRefType as string | undefined; + if (refType !== "ref/prompt" && refType !== "ref/resource") { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "prompts/complete requires --complete-ref-type ref/prompt|ref/resource", + { code: "usage" }, + ); + } + await runRpc(program, method, { + completeRefType: refType, + completeRef: o.completeRef as string | undefined, + completeArgName: o.completeArgName as string | undefined, + completeArgValue: (o.completeArgValue as string | undefined) ?? "", + metadata: o.metadata, + }); + }); + break; + case "logging/setLevel": + cmd + .argument("[level]", "Logging level") + .option("--log-level <level>", "Logging level"); + cmd.action(async (levelPos: string | undefined, o) => { + const level = (o.logLevel as string | undefined) ?? levelPos; + if (level && !validLogLevels.includes(level as LoggingLevel)) { + throw new Error( + `Invalid log level: ${level}. Valid: ${validLogLevels.join(", ")}`, + ); + } + await runRpc(program, method, { + logLevel: level as LoggingLevel | undefined, + metadata: o.metadata, + }); + }); + break; + case "tasks/get": + case "tasks/cancel": + case "tasks/result": + cmd.argument("[taskId]", "Task id").option("--task-id <id>", "Task id"); + cmd.action(async (taskPos: string | undefined, o) => { + await runRpc(program, method, { + taskId: (o.taskId as string | undefined) ?? taskPos, + metadata: o.metadata, + }); + }); + break; + case "tasks/update": + // Modern-only (SEP-2663): resumes a task paused on `input_required`. + // `--input-responses` mirrors `roots/set`'s JSON-blob convention + // rather than trying to model arbitrary per-request shapes as flags. + cmd + .argument("[taskId]", "Task id") + .option("--task-id <id>", "Task id") + .option( + "--input-responses <json>", + "JSON object keyed by the server's inputRequests id", + ); + cmd.action(async (taskPos: string | undefined, o) => { + await runRpc(program, method, { + taskId: (o.taskId as string | undefined) ?? taskPos, + inputResponsesJson: o.inputResponses as string | undefined, + metadata: o.metadata, + }); + }); + break; + case "roots/set": + cmd.option("--roots-json <json>", "JSON array of {uri, name?}"); + cmd.action(async (o) => { + await runRpc(program, method, { + rootsJson: o.rootsJson as string | undefined, + metadata: o.metadata, + }); + }); + break; + default: + cmd.action(async (o) => { + await runRpc(program, method, { + metadata: o.metadata, + }); + }); + break; + } + } +} + +/** + * `elicitation/respond` — answers an elicitation the daemon parked for a + * non-interactive caller (`elicitationPending` output). One respond per + * round: the result is either the resumed call's final output or the next + * pending round. + */ +function registerElicitationCommands(program: CommandType): void { + program + .command("elicitation/respond") + .description( + "Answer a pending server elicitation (from elicitationPending output): form answers as key:=value pairs / JSON, --done for URL mode, or --decline / --cancel", + ) + .argument("<elicitationId>", "Id from the elicitationPending payload") + .argument( + "[fields...]", + "Form answers as key:=value pairs or one JSON object (accepts)", + ) + .option( + "--done", + "URL mode: report the linked interaction as finished (accept)", + ) + .option("--decline", "Decline the request (form mode only)") + .option("--cancel", "Cancel the elicitation") + .action(async (elicitationId: string, fields: string[] | undefined, o) => { + const opts = program.opts<GlobalOpts>(); + const flags = [ + o.done === true && "--done", + o.decline === true && "--decline", + o.cancel === true && "--cancel", + ].filter(Boolean) as string[]; + const hasFields = (fields?.length ?? 0) > 0; + if (flags.length > 1 || (flags.length === 1 && hasFields)) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `Provide field values, or exactly one of --done / --decline / --cancel — not ${[...(hasFields ? ["field values"] : []), ...flags].join(" and ")}.`, + { code: "usage" }, + ); + } + if (flags.length === 0 && !hasFields) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "Provide form answers as key:=value pairs (or one JSON object), or one of --done / --decline / --cancel.", + { code: "usage" }, + ); + } + const params: ElicitationRespondParams = o.cancel + ? { elicitationId, action: "cancel" } + : o.decline + ? { elicitationId, action: "decline" } + : hasFields + ? { + elicitationId, + action: "accept", + content: parseToolCallPositionals(fields!), + } + : { elicitationId, action: "accept" }; + const { socketPath } = await ensureDaemon(); + const result = await callDaemon<ElicitationRespondResult>( + "elicitation/respond", + params, + // The resumed call's duration is governed by MCP timeouts the + // daemon enforces; a fixed local deadline would falsely fail it. + { socketPath, timeoutMs: 0 }, + ); + await writeRpcOutcome( + outOpts(opts), + result.method, + result.toolName, + result.outcome, + ); + }); +} + +async function runRpc( + program: CommandType, + method: string, + methodArgs: MethodArgs, +): Promise<void> { + const opts = program.opts<GlobalOpts>(); + await dispatchConnectionRpc(method, methodArgs, { + format: opts.format, + plain: opts.plain === true, + connection: opts.connection, + requireExplicit: requireExplicitConnection(), + }); +} + +/** + * Overlay `--era` onto the settings lifted from the file/ad-hoc target. + * Mirrors `withConnectTimeout`'s shape: only `protocolEra` is overridden, and a + * bare-defaults settings object is synthesized when the target had none (the + * common ad-hoc case, which otherwise has no way to request `auto`/`modern`). + */ +function withEraOverride( + settings: InspectorServerSettings | undefined, + era: ServerProtocolEra | undefined, +): InspectorServerSettings | undefined { + if (era === undefined) return settings; + if (settings) return { ...settings, protocolEra: era }; + return { + headers: [], + metadata: {}, + env: [], + connectionTimeout: DEFAULT_CONNECT_TIMEOUT_MS, + requestTimeout: 0, + taskTtl: DEFAULT_TASK_TTL_MS, + maxFetchRequests: DEFAULT_MAX_FETCH_REQUESTS, + autoRefreshOnListChanged: false, + paginatedLists: false, + roots: [], + protocolEra: era, + }; +} + +/** + * Overlay `--elicit` onto the settings lifted from the file/ad-hoc target. + * Mirrors `withEraOverride`: only `elicitCapability` is overridden, and a + * bare-defaults settings object is synthesized when the target had none (the + * common ad-hoc case, which otherwise has no way to request anything but the + * default `both`). + */ +function withElicitOverride( + settings: InspectorServerSettings | undefined, + elicit: ElicitCapabilityMode | undefined, +): InspectorServerSettings | undefined { + if (elicit === undefined) return settings; + if (settings) return { ...settings, elicitCapability: elicit }; + return { + headers: [], + metadata: {}, + env: [], + connectionTimeout: DEFAULT_CONNECT_TIMEOUT_MS, + requestTimeout: 0, + taskTtl: DEFAULT_TASK_TTL_MS, + maxFetchRequests: DEFAULT_MAX_FETCH_REQUESTS, + autoRefreshOnListChanged: false, + paginatedLists: false, + roots: [], + elicitCapability: elicit, + }; +} + +/** + * Overlay `--ema` onto the settings lifted from the catalog/config target. + * Mirrors `withEraOverride`: only `enterpriseManaged` is overridden, and a + * bare-defaults settings object is synthesized when the entry had none. + * Ad-hoc targets never reach here with `--ema` set — connect rejects that + * combination up front, since EMA needs per-server OAuth credentials only a + * catalog entry can supply. + */ +function withEmaOverride( + settings: InspectorServerSettings | undefined, + ema: true | undefined, +): InspectorServerSettings | undefined { + if (ema === undefined) return settings; + if (settings) return { ...settings, enterpriseManaged: true }; + return { + headers: [], + metadata: {}, + env: [], + connectionTimeout: DEFAULT_CONNECT_TIMEOUT_MS, + requestTimeout: 0, + taskTtl: DEFAULT_TASK_TTL_MS, + maxFetchRequests: DEFAULT_MAX_FETCH_REQUESTS, + autoRefreshOnListChanged: false, + paginatedLists: false, + roots: [], + enterpriseManaged: true, + }; +} + +function parseKeyValue( + value: string, + previous: Record<string, JsonValue> = {}, +): Record<string, JsonValue> { + const parts = value.split("="); + const key = parts[0]; + const val = parts.slice(1).join("="); + if (!key || val === undefined || val === "") { + throw new Error( + `Invalid parameter format: ${value}. Use key=value format.`, + ); + } + let parsedValue: JsonValue; + try { + parsedValue = JSON.parse(val) as JsonValue; + } catch { + parsedValue = val; + } + assertJsonRoundTrips(parsedValue, `parameter "${value}"`); + return { ...previous, [key]: parsedValue }; +} + +function looksLikeUrl(value: string): boolean { + return /^https?:\/\//i.test(value); +} + +/** + * A single positional token is ambiguous between a catalog entry name and a + * bare stdio command. Disambiguate deterministically: a token that looks + * like a filesystem path (contains a separator, or starts with `.` or `~`) + * is an ad-hoc stdio target; a bare word is a catalog/config name. A bare + * command name can still be run ad-hoc with an explicit `--transport stdio`. + */ +function looksLikePath(value: string): boolean { + return ( + value.includes("/") || + value.includes("\\") || + value.startsWith(".") || + value.startsWith("~") + ); +} + +function splitConnectionTarget(target: string[]): { + name: string | undefined; + rest: string[]; +} { + if (target.length > 0 && target[0]!.startsWith("@")) { + return { name: stripAt(target[0]), rest: target.slice(1) }; + } + return { name: undefined, rest: target }; +} + +export { hoistAtConnection } from "./dispatch.js"; diff --git a/clients/mcpdo/src/connection/parse-tool-args.ts b/clients/mcpdo/src/connection/parse-tool-args.ts new file mode 100644 index 0000000000..809961d56f --- /dev/null +++ b/clients/mcpdo/src/connection/parse-tool-args.ts @@ -0,0 +1,164 @@ +import type { JsonValue } from "@inspector/core/mcp/index.js"; + +/** + * `JSON.parse` accepts numbers JSON cannot represent (`1e999` → `Infinity`); + * the NDJSON serialization to the daemon would then silently send `null`, + * invoking the tool with a different value than the user typed. Reject + * anything that cannot round-trip instead. + */ +export function assertJsonRoundTrips(value: unknown, context: string): void { + if (typeof value === "number" && !Number.isFinite(value)) { + throw new Error( + `Invalid ${context}: ${value} has no JSON representation ` + + `(it would silently reach the tool as null).`, + ); + } + if (Array.isArray(value)) { + for (const item of value) assertJsonRoundTrips(item, context); + } else if (value !== null && typeof value === "object") { + for (const item of Object.values(value)) { + assertJsonRoundTrips(item, context); + } + } +} + +/** + * Parse connection `tools/call` positionals after the tool name: + * - `key:=value` pairs (JSON-typed when the value parses as JSON, else string) + * - a single inline JSON object (`{"message":"Foo"}`) + */ +export function parseToolCallPositionals( + args: string[], +): Record<string, JsonValue> { + if (args.length === 0) return {}; + + const first = args[0]!; + if (first.startsWith("{") || first.startsWith("[")) { + if (args.length > 1) { + throw new Error( + "When using inline JSON, only one argument is allowed after the tool name.", + ); + } + let parsed: unknown; + try { + parsed = JSON.parse(first); + } catch (e) { + throw new Error( + `Invalid JSON tool arguments: ${e instanceof Error ? e.message : String(e)}`, + { cause: e }, + ); + } + if ( + parsed === null || + typeof parsed !== "object" || + Array.isArray(parsed) + ) { + throw new Error("Inline JSON tool arguments must be a JSON object."); + } + assertJsonRoundTrips(parsed, "inline JSON tool arguments"); + return parsed as Record<string, JsonValue>; + } + + // Null prototype: a "__proto__" key must become an ordinary own property + // (as the inline-JSON path preserves it), not hit the {} prototype setter + // and vanish while remapping the accumulator's prototype. + const out: Record<string, JsonValue> = Object.create(null); + for (const pair of args) { + const sep = pair.indexOf(":="); + if (sep === -1) { + throw new Error( + `Invalid tool argument "${pair}". Use key:=value pairs or a JSON object.\n` + + `Examples: message:=hello count:=10 '{"message":"hello"}'`, + ); + } + const key = pair.slice(0, sep); + const rawValue = pair.slice(sep + 2); + if (!key) { + throw new Error( + `Invalid tool argument "${pair}" — missing key before :=`, + ); + } + out[key] = autoParseValue(rawValue, `tool argument "${pair}"`); + } + return out; +} + +function autoParseValue(raw: string, context: string): JsonValue { + let parsed: unknown; + try { + parsed = JSON.parse(raw) as JsonValue; + } catch { + return raw; + } + assertJsonRoundTrips(parsed, context); + return parsed as JsonValue; +} + +export type ResolveToolCallArgsInput = { + toolNameFlag?: string; + toolNamePos?: string; + /** Remaining positionals after the tool-name slot. */ + toolArgsPos?: string[]; + toolArgFlag?: Record<string, JsonValue>; + toolArgsJson?: string; +}; + +/** + * Resolve tool name + arguments from positionals and/or legacy flags. + * Styles are mutually exclusive: positionals, `--tool-arg`, or `--tool-args-json`. + */ +export function resolveToolCallArgs(input: ResolveToolCallArgsInput): { + toolName: string | undefined; + toolArg: Record<string, JsonValue>; +} { + const flagArgs = input.toolArgFlag ?? {}; + const hasFlagArgs = Object.keys(flagArgs).length > 0; + const hasJson = input.toolArgsJson !== undefined; + + let toolName = input.toolNameFlag ?? input.toolNamePos; + let positionals = [...(input.toolArgsPos ?? [])]; + + // `tools/call --tool-name echo message:=Foo` — commander puts message:=Foo + // in the toolName slot when the name came from the flag. + if (input.toolNameFlag && input.toolNamePos) { + positionals = [input.toolNamePos, ...positionals]; + toolName = input.toolNameFlag; + } + + const hasPositionals = positionals.length > 0; + const styles = [hasPositionals, hasFlagArgs, hasJson].filter(Boolean).length; + if (styles > 1) { + throw new Error( + "Tool arguments must use one style: key:=value / JSON positionals, " + + "--tool-arg, or --tool-args-json.", + ); + } + + if (hasJson) { + return { + toolName, + toolArg: parseJsonObject(input.toolArgsJson!, "--tool-args-json"), + }; + } + if (hasPositionals) { + return { toolName, toolArg: parseToolCallPositionals(positionals) }; + } + return { toolName, toolArg: flagArgs }; +} + +function parseJsonObject(raw: string, flag: string): Record<string, JsonValue> { + let parsed: unknown; + try { + parsed = JSON.parse(raw); + } catch (e) { + throw new Error( + `${flag} is not valid JSON: ${e instanceof Error ? e.message : String(e)}`, + { cause: e }, + ); + } + if (parsed === null || typeof parsed !== "object" || Array.isArray(parsed)) { + throw new Error(`${flag} must be a JSON object.`); + } + assertJsonRoundTrips(parsed, flag); + return parsed as Record<string, JsonValue>; +} diff --git a/clients/mcpdo/src/connection/private-env.ts b/clients/mcpdo/src/connection/private-env.ts new file mode 100644 index 0000000000..8f5b00ddac --- /dev/null +++ b/clients/mcpdo/src/connection/private-env.ts @@ -0,0 +1,37 @@ +import { randomBytes } from "node:crypto"; +import { + createPrivateDaemonDir, + DAEMON_DIR_ENV, + DAEMON_TOKEN_ENV, +} from "../daemon/paths.js"; + +export type PrivateEnvBinding = { + dir: string; + token: string; +}; + +/** + * Allocate a private daemon directory and mint an IPC token. + * Does not start the daemon (lazy on first `ensureDaemon`). + */ +export function createPrivateBinding(): PrivateEnvBinding { + const dir = createPrivateDaemonDir(); + const token = randomBytes(32).toString("base64url"); + return { dir, token }; +} + +/** + * Shell exports for `eval "$(mcpdo private)"` (POSIX sh / bash / zsh). + */ +export function formatPrivateEnvExports(binding: PrivateEnvBinding): string { + return [ + `export ${DAEMON_DIR_ENV}=${shellSingleQuote(binding.dir)}`, + `export ${DAEMON_TOKEN_ENV}=${shellSingleQuote(binding.token)}`, + "", + ].join("\n"); +} + +function shellSingleQuote(value: string): string { + // POSIX-safe: 'foo'\''bar' for embedded quotes. + return `'${value.replace(/'/g, `'\\''`)}'`; +} diff --git a/clients/mcpdo/src/connection/prompt-reader.ts b/clients/mcpdo/src/connection/prompt-reader.ts new file mode 100644 index 0000000000..a4a0a0d45b --- /dev/null +++ b/clients/mcpdo/src/connection/prompt-reader.ts @@ -0,0 +1,166 @@ +/** + * Process-wide line reader for interactive prompts (elicitation forms and + * URL confirmations). + * + * Why not one `readline` interface per prompt exchange: readline discards + * `line` events that fire while no `question()` is pending, and closing an + * interface discards whatever input it already buffered. With a piped stdin + * (`printf 'alice\n30\n' | mcpdo tools/call register`) every buffered line + * after the first arrives while no question is listening — and the next + * exchange's fresh interface starts from an empty (often already-EOF) + * stream. So answers beyond the first were silently dropped. + * + * This reader owns a single persistent interface: every `line` event lands + * in a queue, `question()` consumes from the queue before waiting for new + * input, and "close" means *input exhausted* — EOF **and** an empty queue — + * not merely EOF, so piped answers already received still get delivered. + * Between questions the underlying stream is paused and unref'd, so the + * reader never pins the event loop or holds a TTY hostage. + */ +import * as readline from "node:readline"; + +/** + * The structural surface prompts consume: sequential questions plus an + * exhausted-input notification. Implemented by {@link PromptReader}; + * narrow enough for tests to fake. + */ +export type PromptInput = { + question(prompt: string): Promise<string>; + once(event: "close", listener: () => void): unknown; +}; + +type Waiter = { + resolve: (line: string) => void; + reject: (error: Error) => void; +}; + +function exhaustedError(): Error { + return new Error("stdin closed before an answer was given"); +} + +export class PromptReader implements PromptInput { + private readonly rl: readline.Interface; + private readonly input: NodeJS.ReadStream; + private readonly output: NodeJS.WriteStream; + private readonly queued: string[] = []; + private waiter: Waiter | undefined; + private closeListeners: Array<() => void> = []; + private eof = false; + private exhaustedNotified = false; + + constructor( + input: NodeJS.ReadStream = process.stdin, + output: NodeJS.WriteStream = process.stderr, + ) { + this.input = input; + this.output = output; + this.rl = readline.createInterface({ input, output }); + this.rl.on("line", (line) => { + const waiter = this.waiter; + if (waiter) { + this.waiter = undefined; + this.park(); + waiter.resolve(line); + } else { + // No question pending (between fields, or input arrived up front): + // keep the line for the next question instead of dropping it. + this.queued.push(line); + } + }); + this.rl.once("close", () => { + this.eof = true; + const waiter = this.waiter; + if (waiter) { + this.waiter = undefined; + this.notifyExhausted(); + waiter.reject(exhaustedError()); + } + }); + this.park(); + } + + /** + * Write `prompt` and resolve with the next input line — a queued one + * first, else the next to arrive. Rejects once input is exhausted (EOF + * with nothing queued). + */ + question(prompt: string): Promise<string> { + this.output.write(prompt); + const queued = this.queued.shift(); + if (queued !== undefined) return Promise.resolve(queued); + if (this.eof) { + this.notifyExhausted(); + return Promise.reject(exhaustedError()); + } + this.engage(); + return new Promise<string>((resolve, reject) => { + this.waiter = { resolve, reject }; + }); + } + + /** + * `close` here means input is exhausted: EOF *and* no queued line left to + * answer with. A raw stream close while answers are still queued must not + * cancel a form those answers can complete. + */ + once(event: "close", listener: () => void): this { + if (event === "close") { + if (this.exhaustedNotified) listener(); + else this.closeListeners.push(listener); + } + return this; + } + + /** Tear down the underlying interface (tests / process cleanup). Restores + * cooked mode first so a disposed/exiting reader never leaves the terminal + * raw (where Ctrl-C would stay dead in the user's next shell). */ + dispose(): void { + if (this.input.isTTY) this.input.setRawMode?.(false); + this.rl.close(); + } + + private notifyExhausted(): void { + if (this.exhaustedNotified) return; + this.exhaustedNotified = true; + const listeners = this.closeListeners; + this.closeListeners = []; + for (const listener of listeners) listener(); + } + + /** Actively waiting for a line: let the stream flow and hold the loop. On a + * TTY, re-enter raw mode so readline can do line editing while a question is + * pending (the counterpart to {@link park} restoring cooked mode). */ + private engage(): void { + this.input.ref?.(); + if (this.input.isTTY) this.input.setRawMode?.(true); + this.rl.resume(); + } + + /** Idle between questions: stop reading and release the event loop. On a TTY, + * restore cooked mode so the terminal is not left raw between prompts — raw + * mode disables the kernel's Ctrl-C (SIGINT), and a persistent reader that + * stayed raw while parked left Ctrl-C dead in the gap between rounds. */ + private park(): void { + this.rl.pause(); + if (this.input.isTTY) this.input.setRawMode?.(false); + this.input.unref?.(); + } +} + +let shared: PromptReader | undefined; + +/** + * The stdin/stderr reader shared by every prompt in this process. Lazy: a + * run that never prompts never touches stdin. Persistent: consecutive + * elicitation exchanges in one command must share buffered piped input. + */ +export function getSharedPromptReader(): PromptReader { + shared ??= new PromptReader(); + return shared; +} + +/** Test hook: drop (and dispose) the shared reader between cases. */ +export function resetSharedPromptReader(): void { + shared?.dispose(); + shared = undefined; +} diff --git a/clients/mcpdo/src/connection/resolve-command.ts b/clients/mcpdo/src/connection/resolve-command.ts new file mode 100644 index 0000000000..2731472c21 --- /dev/null +++ b/clients/mcpdo/src/connection/resolve-command.ts @@ -0,0 +1,56 @@ +import fs from "node:fs"; +import path from "node:path"; + +/** + * Resolve a bare stdio command name to an absolute path using the CALLER's + * `PATH`, before the config crosses the IPC boundary. + * + * The daemon inherits the environment of whichever mcpdo invocation first + * spawned it, so a bare `node` would otherwise be looked up in a stale + * `PATH` (a different nvm version, a venv from another shell) — the daemon + * could run a different binary than the one the user's shell would. + * Resolving here spawns exactly the caller's binary without forwarding any + * environment across the boundary. + * + * Commands containing a path separator are returned unchanged: the daemon + * resolves those against the connection cwd, which connect already pins to the + * caller's cwd. Names not found on `PATH` are also returned unchanged so the + * daemon's spawn error remains the user-visible failure. + */ +export function resolveCommandPath( + command: string, + env: NodeJS.ProcessEnv = process.env, +): string { + if (!command || command.includes("/") || command.includes(path.sep)) { + return command; + } + const pathVar = env.PATH ?? ""; + /* v8 ignore next 7 -- platform-only branch: PATHEXT applies on win32 only */ + const extensions = + process.platform === "win32" + ? // cmd.exe-like: an already-suffixed name ("node.exe") is tried as-is + // before PATHEXT variants — otherwise only "node.exe.EXE" etc. would + // be searched and resolution would silently fall to the daemon's PATH. + [ + ...(path.extname(command) ? [""] : []), + ...(env.PATHEXT ?? ".COM;.EXE;.BAT;.CMD").split(";"), + ] + : [""]; + for (const dir of pathVar.split(path.delimiter)) { + // POSIX: an empty PATH entry means the current directory. Resolve it (and + // any relative entry) against the caller's cwd so the daemon always + // receives an absolute path. + for (const ext of extensions) { + const candidate = path.resolve(dir === "" ? "." : dir, command + ext); + try { + const stat = fs.statSync(candidate); + if (!stat.isFile()) continue; + fs.accessSync(candidate, fs.constants.X_OK); + return candidate; + } catch { + // Not there / not executable — keep looking. + } + } + } + return command; +} diff --git a/clients/mcpdo/src/connection/sanitize.ts b/clients/mcpdo/src/connection/sanitize.ts new file mode 100644 index 0000000000..fcc3a1f936 --- /dev/null +++ b/clients/mcpdo/src/connection/sanitize.ts @@ -0,0 +1,74 @@ +/** + * Terminal-output sanitization for server-controlled text (security). + * + * Every string a server sends (tool results, descriptions, resource text, + * elicitation messages, URIs) reaches the user's terminal through the human + * formatter. Raw C0/C1 control bytes in that text are attacker-controlled + * terminal commands: OSC 52 writes the clipboard, OSC 0 spoofs the window + * title, CSI moves/erases earlier output, and a BEL/ESC inside a URI breaks + * out of an OSC 8 hyperlink wrapper. `--format json` is safe (JSON escapes + * them); this module makes `--format text` safe by replacing every control + * character except `\n` and `\t` with a visible stand-in before any styling + * (so the CLI's own ANSI styling, added afterwards, is unaffected). + * + * Replacements: C0 → Unicode Control Pictures (␀…␟, e.g. ESC → ␛), + * DEL → ␡, C1 (0x80–0x9F, includes 8-bit CSI/OSC) → ␡-style `\u{9b}` text. + */ + +const CONTROL_CHARS = + // C0 minus \t (0x09) and \n (0x0A), plus DEL and the C1 range. + // eslint-disable-next-line no-control-regex + /[\u0000-\u0008\u000B-\u001F\u007F-\u009F]/g; + +function visibleControl(ch: string): string { + const code = ch.codePointAt(0)!; + if (code <= 0x1f) return String.fromCodePoint(0x2400 + code); + if (code === 0x7f) return "\u2421"; // ␡ + return `\\u{${code.toString(16)}}`; // C1: no control picture exists +} + +/** Replace terminal control characters (except `\n`/`\t`) with visible text. */ +export function sanitizeText(value: string): string { + return value.replace(CONTROL_CHARS, visibleControl); +} + +/** + * Deep-sanitize every string in a payload (values AND object keys) ahead of + * human formatting. Input is JSON-shaped data (daemon responses are + * JSON-parsed), so plain objects/arrays/primitives are the whole universe. + */ +export function sanitizeDeep<T>(value: T): T { + if (typeof value === "string") return sanitizeText(value) as T; + if (Array.isArray(value)) return value.map((v) => sanitizeDeep(v)) as T; + if (value !== null && typeof value === "object") { + // Null prototype: a JSON key named "__proto__" must become an own + // property, not invoke the inherited prototype setter (which would + // silently drop the field from formatted output). + const out: Record<string, unknown> = Object.create(null) as Record< + string, + unknown + >; + for (const [k, v] of Object.entries(value as Record<string, unknown>)) { + out[sanitizeText(k)] = sanitizeDeep(v); + } + return out as T; + } + return value; +} + +/** + * Schemes a server-supplied URI may be rendered as an OSC 8 hyperlink. + * A hyperlink is an invitation for the user to invoke the local handler for + * the scheme, so an untrusted MCP server only gets the web ones: `file:`, + * custom protocol handlers, `javascript:` and the rest render as plain text. + */ +const SAFE_LINK_SCHEMES = new Set(["https:", "http:"]); + +/** True when `uri` parses and its scheme is on the OSC 8 allowlist. */ +export function isSafeLinkTarget(uri: string): boolean { + try { + return SAFE_LINK_SCHEMES.has(new URL(uri).protocol); + } catch { + return false; + } +} diff --git a/clients/mcpdo/src/connection/stored-auth.ts b/clients/mcpdo/src/connection/stored-auth.ts new file mode 100644 index 0000000000..f87934273c --- /dev/null +++ b/clients/mcpdo/src/connection/stored-auth.ts @@ -0,0 +1,148 @@ +import { + clearAllOAuthClientState, + getStateFilePath, + NodeOAuthStorage, + resetNodeOAuthStorageCache, +} from "@inspector/core/auth/node/storage-node.js"; +import { readOAuthStore } from "@inspector/core/auth/node/oauth-persist-file.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; + +/** Same canonicalisation as one-shot `normalizeServerUrl` (avoid importing cli.ts). */ +export function normalizeServerUrl(serverUrl: string): string { + try { + return new URL(serverUrl).href; + } catch { + return serverUrl; + } +} + +export type StoredAuthEntry = { + url: string; + hasTokens: boolean; + hasRefreshToken: boolean; +}; + +export type StoredAuthList = { + oauthStatePath: string; + servers: StoredAuthEntry[]; +}; + +type TokenBlob = { + access_token?: string; + refresh_token?: string; +}; + +function tokenFlagsFromState(state: unknown): { + hasTokens: boolean; + hasRefreshToken: boolean; +} { + if (state == null || typeof state !== "object") { + return { hasTokens: false, hasRefreshToken: false }; + } + const s = state as { + tokens?: TokenBlob; + byIssuer?: Record<string, { tokens?: TokenBlob }>; + }; + // Aggregate across every token slot: with multiple issuers, returning at + // the first access-token-bearing slot would make hasRefreshToken depend on + // object insertion order. + let hasTokens = Boolean(s.tokens?.access_token); + let hasRefreshToken = hasTokens && Boolean(s.tokens?.refresh_token); + for (const slot of Object.values(s.byIssuer ?? {})) { + if (slot?.tokens?.access_token) { + hasTokens = true; + if (slot.tokens.refresh_token) hasRefreshToken = true; + } + } + return { hasTokens, hasRefreshToken }; +} + +async function readServersMap( + statePath: string, +): Promise<Record<string, unknown>> { + // readOAuthStore rejoins secrets (tokens, client secrets) from the secret + // store into the snapshot — the raw oauth.json blob no longer carries them, + // so parsing the file directly would report every entry as token-less. + const snapshot = await readOAuthStore(statePath); + if (snapshot?.servers && typeof snapshot.servers === "object") { + return snapshot.servers as Record<string, unknown>; + } + return {}; +} + +/** List every server key in the shared OAuth store (tokens optional). */ +export async function listStoredAuth(): Promise<StoredAuthList> { + const oauthStatePath = getStateFilePath(); + const servers = await readServersMap(oauthStatePath); + const entries = Object.keys(servers) + .sort((a, b) => a.localeCompare(b)) + .map((url) => ({ + url, + ...tokenFlagsFromState(servers[url]), + })); + return { oauthStatePath, servers: entries }; +} + +/** + * Resolve a user-supplied key to a stored server URL (exact, then normalised). + */ +export async function resolveStoredAuthKey(key: string): Promise<string> { + const trimmed = key.trim(); + if (!trimmed) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "auth/clear requires a server URL key (from auth/list) or --all", + { code: "usage" }, + ); + } + const { servers } = await listStoredAuth(); + const urls = servers.map((s) => s.url); + if (urls.includes(trimmed)) return trimmed; + const normalized = normalizeServerUrl(trimmed); + if (urls.includes(normalized)) return normalized; + // Allow clearing a key that is not listed (no-op clear) when it normalises + // to a URL — still useful after partial writes. + if (normalized !== trimmed || /^https?:\/\//i.test(trimmed)) { + return normalized; + } + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `No stored auth entry for '${trimmed}'. Use auth/list to see keys.`, + { code: "usage" }, + ); +} + +/** Clear one server's OAuth state from the shared store. */ +export async function clearStoredAuth(key: string): Promise<{ url: string }> { + const url = await resolveStoredAuthKey(key); + const storage = new NodeOAuthStorage(); + await storage.clear(url); + resetNodeOAuthStorageCache(); + return { url }; +} + +/** Clear every server entry in the shared OAuth store. */ +export async function clearAllStoredAuth(): Promise<{ cleared: number }> { + const before = await listStoredAuth(); + await clearAllOAuthClientState(); + resetNodeOAuthStorageCache(); + return { cleared: before.servers.length }; +} + +/** + * Drop stored OAuth state for an HTTP(S) server URL so the next connect cannot + * silently reuse tokens (`--relogin`). No-op when `serverUrl` is missing + * (stdio / no URL-keyed entry) — interactive login still only runs if auth is required. + */ +export async function clearStoredAuthForRelogin( + serverUrl: string | undefined, +): Promise<void> { + if (!serverUrl?.trim()) return; + const url = normalizeServerUrl(serverUrl.trim()); + const storage = new NodeOAuthStorage(); + await storage.clear(url); + resetNodeOAuthStorageCache(); +} diff --git a/clients/mcpdo/src/daemon/auth.ts b/clients/mcpdo/src/daemon/auth.ts new file mode 100644 index 0000000000..638dec52fa --- /dev/null +++ b/clients/mcpdo/src/daemon/auth.ts @@ -0,0 +1,69 @@ +import { randomBytes, timingSafeEqual } from "node:crypto"; +import * as fs from "node:fs"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { DAEMON_TOKEN_ENV, getDaemonTokenPath } from "./paths.js"; + +/** Fresh random IPC token for a daemon whose environment didn't supply one. */ +export function generateDaemonToken(): string { + return randomBytes(32).toString("hex"); +} + +/** + * Read the token a running daemon published to `mcpdod.token` (see + * {@link getDaemonTokenPath}). Undefined when missing/unreadable — the + * request will then fail authentication with a clear error. + */ +export function readDaemonTokenFile(dir?: string): string | undefined { + try { + const token = fs.readFileSync(getDaemonTokenPath(dir), "utf8").trim(); + return token || undefined; + } catch { + return undefined; + } +} + +/** + * Read the IPC token from the environment (parent client or daemon child). + * Empty / unset → shared mode, which is still authenticated: the daemon + * generates its own required token (see `daemon/run.ts`) and publishes it + * to `mcpdod.token` for same-user clients to read. Every daemon requires a + * token; the environment variable only selects who supplies it. + */ +export function getDaemonTokenFromEnv( + env: NodeJS.ProcessEnv = process.env, +): string | undefined { + const token = env[DAEMON_TOKEN_ENV]?.trim(); + return token || undefined; +} + +/** Constant-time compare; false if either side is missing or lengths differ. */ +export function tokensEqual( + expected: string | undefined, + provided: string | undefined, +): boolean { + if (expected === undefined || provided === undefined) return false; + const a = Buffer.from(expected, "utf8"); + const b = Buffer.from(provided, "utf8"); + if (a.length !== b.length) return false; + return timingSafeEqual(a, b); +} + +/** + * When {@link requiredToken} is set, reject requests that omit or mismatch it. + */ +export function assertDaemonToken( + requiredToken: string | undefined, + provided: string | undefined, +): void { + if (requiredToken === undefined) return; + if (!tokensEqual(requiredToken, provided)) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "Daemon IPC authentication failed (missing or invalid token).", + { code: "daemon_auth_failed" }, + ); + } +} diff --git a/clients/mcpdo/src/daemon/client.ts b/clients/mcpdo/src/daemon/client.ts new file mode 100644 index 0000000000..8002b08a29 --- /dev/null +++ b/clients/mcpdo/src/daemon/client.ts @@ -0,0 +1,285 @@ +import { randomUUID } from "node:crypto"; +import * as net from "node:net"; +import * as path from "node:path"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { getDaemonTokenFromEnv, readDaemonTokenFile } from "./auth.js"; +import { encodeRequest } from "./framing.js"; +import { getDaemonDir, getDaemonSocketPath } from "./paths.js"; +import { sanitizeText } from "../connection/sanitize.js"; +import type { + DaemonOp, + DaemonRequest, + DaemonResponse, + ElicitationRequestFrame, + ElicitationResponseFrame, +} from "./protocol.js"; + +export type DaemonClientOptions = { + socketPath?: string; + /** + * Daemon directory that owns `mcpdod.token`. On Windows `socketPath` is a + * named pipe (`\\.\pipe\...`), so the token location cannot be derived + * from the endpoint; callers using a non-default directory with an + * explicit `socketPath` should pass it. Defaults to the socket's directory + * for Unix socket paths, else the shared daemon directory. + */ + dir?: string; + /** Per-request timeout in ms. */ + /** + * Client-side deadline for the whole request; `0` disables it. Defaults to + * 60s, which suits short control ops (ping, status, list). Callers of ops + * whose duration is governed by configured MCP timeouts the daemon already + * enforces (`connect` honouring `--connect-timeout`, `rpc` honouring the + * request timeout — either may validly run past 60s or be unlimited) must + * pass `0` so the fixed local timer can't fail an op the daemon is still + * executing. Daemon death is still detected via socket error/close. + */ + timeoutMs?: number; + /** IPC token; defaults to `MCP_INSPECTOR_DAEMON_TOKEN` when set. */ + token?: string; + /** + * Called when the in-flight `rpc` call surfaces a legacy or modern + * non-task MRTR elicitation mid-call (dual-era support, phase 1). Omit to + * auto-answer `{action: "cancel"}` — appropriate for non-interactive + * callers (e.g. `--format json`, non-TTY) that shouldn't hang waiting on a + * human. The connect timeout is cleared once the first such frame arrives, + * so an interactive prompt isn't bounded by the original request timeout. + */ + onElicitation?: ( + frame: ElicitationRequestFrame, + ) => Promise<ElicitationResponseFrame>; + /** + * Abort the in-flight request (e.g. on SIGINT/SIGTERM), failing it with a + * clear cancellation error instead of leaving the caller to kill the + * process abruptly mid-call (mid-`tools/call`, mid-elicitation-wait, etc). + */ + signal?: AbortSignal; +}; + +/** + * Directory holding `mcpdod.token` for a request. An explicit `dir` wins; a + * filesystem `socketPath` implies its directory (Unix sockets live next to + * the token file); otherwise — the default endpoint, or a Windows named + * pipe, which has no meaningful dirname — the shared daemon directory. + */ +export function daemonTokenDir(options: { + dir?: string; + socketPath?: string; +}): string { + if (options.dir !== undefined) return options.dir; + if ( + options.socketPath !== undefined && + !options.socketPath.startsWith("\\\\.\\pipe\\") + ) { + return path.dirname(options.socketPath); + } + return getDaemonDir(); +} + +/** + * Short-lived NDJSON client for one request/response against the daemon. + */ +export async function callDaemon<T = unknown>( + op: DaemonOp, + params?: DaemonRequest["params"], + options: DaemonClientOptions = {}, +): Promise<T> { + const socketPath = options.socketPath ?? getDaemonSocketPath(); + const timeoutMs = options.timeoutMs ?? 60_000; + const id = randomUUID(); + // Env token wins (private mode / spawner); otherwise read the token the + // daemon published in its directory (see getDaemonTokenPath). The + // directory is only derived from `socketPath` when that is a filesystem + // path — dirname of a Windows named pipe is the pipe namespace, not the + // daemon dir. + const token = + options.token ?? + getDaemonTokenFromEnv() ?? + readDaemonTokenFile(daemonTokenDir(options)); + const request: DaemonRequest = { id, op, params }; + if (token !== undefined) request.token = token; + + return new Promise<T>((resolve, reject) => { + let settled = false; + let buffer = ""; + let queue: Promise<void> = Promise.resolve(); + // `let` so settle() can clearTimeout before the assignment if connect fails + // synchronously (prefer-const would put `timer` in the TDZ for that race). + let timer: ReturnType<typeof setTimeout> | undefined; + const socket = new net.Socket(); + // Decode at the socket: a multi-byte UTF-8 character split across TCP + // chunks must be reassembled by the stream's StringDecoder, not mangled + // into U+FFFD by a per-chunk String() conversion. + socket.setEncoding("utf8"); + + function settle(fn: () => void) { + /* v8 ignore next -- settle() no-op when already settled (connect/timeout race) */ + if (settled) return; + settled = true; + if (timer !== undefined) clearTimeout(timer); + options.signal?.removeEventListener("abort", onAbort); + socket.removeAllListeners(); + socket.on("error", () => {}); + fn(); + } + + function onAbort() { + fail( + new CliExitCodeError(EXIT_CODES.USAGE, `'${op}' cancelled.`, { + code: "cancelled", + }), + ); + } + + function fail(error: unknown) { + settle(() => { + socket.destroy(); + reject(error); + }); + } + + function succeed(value: T) { + settle(() => { + socket.end(); + resolve(value); + }); + } + + function handleLine(line: string): Promise<void> { + const trimmed = line.trim(); + if (!trimmed) return Promise.resolve(); + let parsed: DaemonResponse | ElicitationRequestFrame; + try { + parsed = JSON.parse(trimmed) as + | DaemonResponse + | ElicitationRequestFrame; + } catch (error) { + fail(error); + return Promise.resolve(); + } + if ( + parsed !== null && + typeof parsed === "object" && + "kind" in parsed && + parsed.kind === "elicitation-request" + ) { + return handleElicitationRequest(parsed as ElicitationRequestFrame); + } + handleResponse(parsed as DaemonResponse); + return Promise.resolve(); + } + + async function handleElicitationRequest( + frame: ElicitationRequestFrame, + ): Promise<void> { + if (frame.id !== id) return; + // A human (or a multi-round MRTR exchange) answering this shouldn't be + // bounded by the original fixed request timeout. + if (timer !== undefined) { + clearTimeout(timer); + timer = undefined; + } + const answer = options.onElicitation + ? await options.onElicitation(frame) + : ({ + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "cancel", + } satisfies ElicitationResponseFrame); + if (settled || socket.destroyed) return; + socket.write(JSON.stringify(answer) + "\n"); + } + + function handleResponse(response: DaemonResponse) { + if (response.id !== id && response.id !== "?") { + return; + } + if (!response.ok) { + fail( + new CliExitCodeError( + response.error.exitCode ?? EXIT_CODES.USAGE, + // Daemon error text can embed server-supplied strings; sanitize + // before it reaches a terminal via the shared error handler. + sanitizeText(response.error.message), + { code: response.error.code }, + ), + ); + return; + } + succeed(response.result as T); + } + + socket.on("error", (err) => { + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Cannot reach connection daemon at ${socketPath}: ${err.message}`, + { code: "daemon_unreachable" }, + ), + ); + }); + + // Clean FIN with no response must not sit until timeoutMs (mirrors + // streamDaemon's close guard). + socket.on("close", () => { + if (!settled) { + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Connection daemon closed the connection during '${op}'`, + { code: "daemon_unreachable" }, + ), + ); + } + }); + + timer = + timeoutMs > 0 + ? setTimeout(() => { + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Daemon request '${op}' timed out after ${timeoutMs}ms`, + { code: "daemon_timeout" }, + ), + ); + }, timeoutMs) + : undefined; + + // AbortSignal does not replay: a signal that aborted before this point + // (e.g. SIGINT during ensureDaemon) would never fire the listener, and + // with timeoutMs 0 the request would hang forever. Check first. + if (options.signal?.aborted) { + onAbort(); + } else { + options.signal?.addEventListener("abort", onAbort, { once: true }); + } + + socket.once("connect", () => { + socket.write(encodeRequest(request)); + }); + + socket.on("data", (chunk) => { + buffer += String(chunk); + let idx: number; + while ((idx = buffer.indexOf("\n")) >= 0) { + const line = buffer.slice(0, idx); + buffer = buffer.slice(idx + 1); + // Sequential so an awaited onElicitation prompt fully settles (and + // its answer is written) before the next buffered line is handled. + queue = queue + .then(() => handleLine(line)) + .catch((error) => fail(error)); + } + }); + + // A pre-aborted signal settles above via onAbort() and destroys the + // socket; connect() would silently un-destroy it and leak a live socket + // that pins the event loop. + if (!settled) socket.connect(socketPath); + }); +} diff --git a/clients/mcpdo/src/daemon/connections.ts b/clients/mcpdo/src/daemon/connections.ts new file mode 100644 index 0000000000..0f175fb8fe --- /dev/null +++ b/clients/mcpdo/src/daemon/connections.ts @@ -0,0 +1,928 @@ +import { InspectorClient } from "@inspector/core/mcp/index.js"; +import type { InspectorClientEnvironment } from "@inspector/core/mcp/types.js"; +import { + DEFAULT_ELICIT_CAPABILITY, + eraToVersionNegotiation, + isTerminalStatus, + type ElicitCapabilityMode, + type InspectorClientOptions, + type InspectorServerSettings, + type MCPServerConfig, +} from "@inspector/core/mcp/types.js"; +import { createTransportNode } from "@inspector/core/mcp/node/index.js"; +import { + buildOAuthConnectionState, + ConsoleNavigation, + hasPersistedOAuthServerState, + isServerOAuthConfigured, + MutableRedirectUrlProvider, + protocolFromOAuthConfig, +} from "@inspector/core/auth/index.js"; +import type { OAuthConnectionState } from "@inspector/core/auth/types.js"; +import { NodeOAuthStorage } from "@inspector/core/auth/node/index.js"; +import { resetNodeOAuthStorageCache } from "@inspector/core/auth/node/storage-node.js"; +import { + DEFAULT_RUNNER_OAUTH_CALLBACK_URL, + formatRunnerOAuthRedirectUrl, + parseRunnerOAuthCallbackUrl, +} from "@inspector/core/auth/node/runner-oauth-callback.js"; +import { + buildRunnerClientAuthOptions, + isOAuthCapableServerConfig, + loadRunnerClientConfig, +} from "@inspector/core/client/runner.js"; +import { readInspectorVersion } from "@inspector/core/node/version.js"; +import { cleanRoots } from "@inspector/core/mcp/serverList.js"; +import { + AuthRecoveryRequiredError, + isUnauthorizedError, +} from "@inspector/core/auth/index.js"; +import { isEmaClientNotConfiguredError } from "@inspector/core/auth/ema/clientConfigError.js"; +import { readLivePendingAuthMarker } from "../connection/auth-helper.js"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import type { + ConnectionAuthInfo, + ConnectionInfo, + ConnectParams, +} from "./protocol.js"; + +const CONNECTION_CLIENT_NAME = "inspector-cli"; + +/** Default idle timeout after the last connection disconnects (~60s). */ +export const DEFAULT_IDLE_MS = 60_000; + +type LiveConnection = { + name: string; + serverIdentity: string; + connectedAt: number; + lastAccessedAt: number; + client: InspectorClient; + /** Retained for `connections/show`'s live auth recompute. */ + serverConfig: MCPServerConfig; + serverSettings?: InspectorServerSettings; + /** Connect-time snapshot (see {@link ConnectionInfo.auth}). */ + auth?: ConnectionAuthInfo; + /** + * Auth-pending intent entry (see {@link ConnectionInfo.pendingAuth}): + * registered with a never-connected client while the detached auth helper + * completes sign-in out of band. Cleared by the first successful revive. + */ + pendingAuth?: boolean; +}; + +/** + * In-memory registry of live MCP connections owned by the daemon. + */ +export class ConnectionRegistry { + private readonly connections = new Map<string, LiveConnection>(); + private mruName: string | null = null; + private idleTimer: ReturnType<typeof setTimeout> | null = null; + /** Absolute deadline for idle shutdown while the timer is armed. */ + private idleDeadline: number | null = null; + private onIdle: (() => void) | null = null; + private readonly idleMs: number; + + constructor(idleMs: number = DEFAULT_IDLE_MS) { + this.idleMs = idleMs; + } + + /** Register a callback invoked when the idle timer fires with no connections. */ + setIdleHandler(handler: (() => void) | null): void { + this.onIdle = handler; + } + + /** + * Per-name serialization for connect/disconnect. Both hold `connections` + * state across awaits; two simultaneous connects for the same + * previously-unused name would otherwise both pass the reconnect check and + * race through `connections.set`, leaving the loser's live client + * untracked and undisconnectable. + */ + private readonly nameLocks = new Map<string, Promise<void>>(); + + /** + * Set by {@link disconnectAll} (daemon shutdown). A connect that was + * in-flight when the shutdown snapshot was taken — e.g. one that outlived + * the bounded quiesce grace — must not register a live client afterwards: + * nothing would ever disconnect it once the daemon exits. + */ + private closed = false; + + private async withNameLock<T>( + name: string, + fn: () => Promise<T>, + ): Promise<T> { + const prev = this.nameLocks.get(name) ?? Promise.resolve(); + let release!: () => void; + const current = new Promise<void>((resolve) => (release = resolve)); + this.nameLocks.set(name, current); + await prev; + try { + return await fn(); + } finally { + release(); + if (this.nameLocks.get(name) === current) this.nameLocks.delete(name); + } + } + + /** + * Connects currently in flight (past {@link connectLocked} entry, not yet + * registered or failed). The idle timer must not fire while one is + * pending: a *different* name's failed connect (or a disconnect) would + * otherwise re-arm the timer and stop the daemon under a valid connect + * still awaiting `client.connect()`. + */ + private pendingConnects = 0; + + /** + * Arm the idle shutdown timer when there are no connections and no + * connects in flight. Called at daemon start so a spawn that never + * connects still self-reaps, and after a failed connect/disconnect that + * left the registry empty. + */ + armIdleTimerIfEmpty(): void { + if (this.closed) return; + if (this.connections.size === 0 && this.pendingConnects === 0) { + this.armIdleTimer(); + } + } + + list(): ConnectionInfo[] { + return [...this.connections.values()] + .map((s) => ({ + name: s.name, + serverIdentity: s.serverIdentity, + connectedAt: s.connectedAt, + lastAccessedAt: s.lastAccessedAt, + isMru: s.name === this.mruName, + protocolEra: s.client.getProtocolEra(), + ...(s.auth && { auth: s.auth }), + ...(s.pendingAuth && { pendingAuth: true }), + })) + .sort((a, b) => b.lastAccessedAt - a.lastAccessedAt); + } + + /** + * Annotate pending-auth entries whose out-of-band sign-in has completed: + * a disk check only (no dial, no MRU touch), setting + * {@link ConnectionInfo.pendingAuthSignedIn} and refreshing `auth` to the + * live disk state so the echo isn't the contradictory "pending + + * authorized: false". Used by the read-only echoes (`connections/list`, + * `connections/use`, `daemon/status`); `connections/show` goes further and + * revives (see the show handler). + */ + async annotateAuthProgress( + infos: ConnectionInfo[], + ): Promise<ConnectionInfo[]> { + for (const info of infos) { + if (info.pendingAuth !== true) continue; + const connection = this.connections.get(String(info.name)); + if (!connection) continue; + const auth = await getLiveConnectionAuthInfo(connection); + if (auth?.authorized === true) { + info.pendingAuthSignedIn = true; + info.auth = auth; + } + } + return infos; + } + + getMruName(): string | null { + return this.mruName; + } + + connectionCount(): number { + return this.connections.size; + } + + /** + * Resolve a connection by explicit name or MRU. Throws {@link CliExitCodeError} + * when missing / ambiguous under CI rules. + */ + resolve( + name: string | undefined, + requireExplicit: boolean | undefined, + ): LiveConnection { + if (!name) { + if (requireExplicit) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "Explicit --connection / @name is required in non-interactive mode.", + { code: "connection_required" }, + ); + } + if (!this.mruName) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "No open connections. Connect first (e.g. mcpdo servers/list, mcpdo connect <entry>).", + { code: "no_connection" }, + ); + } + name = this.mruName; + } + const connection = this.connections.get(name); + if (!connection) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `Connection '${name}' isn't connected. Before using it you must connect using:\n mcpdo connect ${name}`, + { code: "connection_not_found" }, + ); + } + return connection; + } + + touch(name: string): void { + const connection = this.connections.get(name); + if (!connection) return; + connection.lastAccessedAt = Date.now(); + this.mruName = name; + this.clearIdleTimer(); + } + + /** + * Resolve a connection for an RPC/stream/show, touch MRU, and return the + * live connection (name/serverIdentity/timestamps plus the client). + */ + connectionFor( + name: string | undefined, + requireExplicit: boolean | undefined, + ): LiveConnection { + const connection = this.resolve(name, requireExplicit); + this.touch(connection.name); + return connection; + } + + /** + * Resolve a connection for an RPC/stream, touch MRU, and return its client. + */ + clientFor( + name: string | undefined, + requireExplicit: boolean | undefined, + ): InspectorClient { + return this.connectionFor(name, requireExplicit).client; + } + + /** + * Like {@link clientFor}, but guarantees the returned client's transport is + * live, transparently re-dialing when it has settled into a terminal state. + * + * The registry entry represents user intent — "I connected; it's mine until + * I disconnect" — while the `InspectorClient` inside it is a disposable + * transport artifact. A long-held connection's transport can die without any + * user action (the server expires its session, drops the SSE stream, or a + * stdio child exits), so ops resolve their client through here and the dead + * client is replaced with a freshly dialed one on demand. Stored credentials + * (refresh token → new access token) make the revive silent; the user is + * only involved when re-auth genuinely needs them (`auth_required`, e.g. a + * missing/expired refresh token). No background retry loop: a connection + * nobody is using costs nothing, matching increasingly session-less servers. + * + * On revive failure the registry entry is kept — the intent persists, and + * the next op simply tries again. + */ + async liveClientFor( + name: string | undefined, + requireExplicit: boolean | undefined, + interactive?: boolean, + ): Promise<InspectorClient> { + const connection = this.connectionFor(name, requireExplicit); + if (!isTerminalStatus(connection.client.getStatus())) { + return connection.client; + } + return this.withNameLock(connection.name, () => + this.reviveLocked(connection.name, interactive), + ); + } + + private async reviveLocked( + name: string, + interactive?: boolean, + ): Promise<InspectorClient> { + this.assertOpen(); + // Re-resolve under the lock: a disconnect or replacing connect queued + // ahead of this revive changes what the name means (or removes it). + const connection = this.resolve(name, true); + if (!isTerminalStatus(connection.client.getStatus())) { + // A queued sibling op already revived it. + return connection.client; + } + const dead = connection.client; + let client: InspectorClient; + try { + client = await this.dial(connection); + } catch (error) { + if ( + error instanceof CliExitCodeError && + error.envelope?.code === "auth_required" + ) { + throw this.authRequiredError(connection, interactive, error); + } + throw error; + } + // Free the dead client's resources (child process reaping, timers); + // best-effort, its transport is already gone. + await safeDisconnect(dead); + if (this.closed) { + // Shutdown ran while dialing; disconnectAll's snapshot already passed + // this entry, so registering the fresh client would leak it past exit. + await safeDisconnect(client); + this.assertOpen(); + } + connection.client = client; + connection.lastAccessedAt = Date.now(); + // A successful revive is the completion of any out-of-band sign-in. + delete connection.pendingAuth; + // Refresh the auth snapshot `connections/list`/`use` report — the revive + // may have rotated tokens. + const auth = await getConnectionAuthInfo(client); + if (auth) { + connection.auth = auth; + } else { + delete connection.auth; + } + return client; + } + + /** + * The live authorize URL for a pending connection, or undefined. Keyed by + * the server URL (the OAuth marker key — see auth-helper.ts + * `obtainPendingAuthUrl`); stdio configs have no URL to key on and so never + * carry one. "Live" means the detached auth helper is still running and the + * marker is unexpired (see `readLivePendingAuthMarker`): a completed or + * crashed sign-in returns undefined, which is the discriminator between a + * sign-in still in flight and a genuine credential lapse. + */ + pendingAuthUrlFor(connection: LiveConnection): string | undefined { + const cfg = connection.serverConfig; + if (!("url" in cfg) || typeof cfg.url !== "string" || cfg.url === "") { + return undefined; + } + return readLivePendingAuthMarker(cfg.url)?.url; + } + + /** + * Compose the `auth_required` error a failed silent revive raises, picking + * the message by *why* it failed and *who* is reading it: + * + * - **Sign-in still pending** (`pendingAuth` + a live auth-helper marker): + * the out-of-band flow the user was handed hasn't completed. The sign-in + * URL is NOT embedded here — `connect` is the one reliable place it + * surfaces with the right turn discipline (paste-as-last, end-turn), and + * `connections/show @name` re-displays it the same way. So this error only + * names the command to run; it never relays the URL itself, which keeps + * the relay in one place instead of smeared across every RPC. An agent + * (non-interactive) is told, as the message's last line, to run + * `connections/show @name` if it needs to show the link again; a human + * reads it from `connect` / `connections/show` and gets a terse pointer. + * - **Credentials lapsed** (no pending flow in flight): the stored tokens + * could not be refreshed — re-run `mcpdo connect`. Unchanged wording. + */ + private authRequiredError( + connection: LiveConnection, + interactive: boolean | undefined, + cause: CliExitCodeError, + ): CliExitCodeError { + const name = connection.name; + // A live sign-in marker is what distinguishes "flow in flight" from + // "credentials lapsed"; we branch on its presence but no longer relay the + // URL out of this error (connect / connections/show own that). + const signInInFlight = + connection.pendingAuth === true && + this.pendingAuthUrlFor(connection) !== undefined; + if (signInInFlight) { + const message = + interactive === true + ? `Connection '${name}' sign-in is still pending. To see details, run:\n mcpdo connections/show @${name}` + : `Connection '${name}' is not authenticated yet — sign-in is still pending; ` + + `the user must finish signing in. To display the sign-in link again ` + + `if you need it, run:\n mcpdo connections/show @${name}`; + return new CliExitCodeError(EXIT_CODES.AUTH_REQUIRED, message, { + code: "auth_required", + }); + } + // Silent revive is out of credentials with nothing in flight; only now + // does the user need to act. The front-end `connect` command runs the + // interactive flow. + return new CliExitCodeError( + EXIT_CODES.AUTH_REQUIRED, + `Connection '${name}' needs re-authentication — stored credentials could not be refreshed ` + + `(${cause.message}). To sign in again, run:\n mcpdo connect ${name}`, + { code: "auth_required" }, + ); + } + + /** + * Build a client for a server config and connect it, using whatever + * credentials are on disk (silent: interactive login runs in the front-end, + * never here). Shared by first connect and revive. Auth failures — including + * SDK token-exchange mistakes that mean "needs a full re-auth" — map to an + * `auth_required` envelope; every failure path tears the client down. + */ + private async dial( + params: { + serverConfig: MCPServerConfig; + serverSettings?: InspectorServerSettings; + }, + signal?: AbortSignal, + ): Promise<InspectorClient> { + // Front-end authorize / auth/clear write oauth.json in another process. + // Drop the daemon's cached store so this dial re-reads disk. + resetNodeOAuthStorageCache(); + + const client = await createConnectionClient( + params.serverConfig, + params.serverSettings, + ); + + try { + // Race the connect against caller hang-up: when the requesting + // socket closes mid-dial (Ctrl-C, frontend crash) the daemon must + // not keep the attempt alive — with `--connect-timeout 0` it would + // otherwise pin `pendingConnects` (blocking idle shutdown) or + // register a connection the user cancelled. On abort the shared + // catch below tears the client down, which also cancels the + // still-in-flight connect. + await withAbort(() => client.connect(), signal, connectCancelledError); + } catch (error) { + await safeDisconnect(client); + if (isConnectionAuthRequiredError(error)) { + throw new CliExitCodeError( + EXIT_CODES.AUTH_REQUIRED, + error instanceof Error ? error.message : String(error), + { code: "auth_required" }, + ); + } + throw error; + } + return client; + } + + use(name: string): ConnectionInfo { + const connection = this.resolve(name, true); + this.touch(connection.name); + return { + name: connection.name, + serverIdentity: connection.serverIdentity, + connectedAt: connection.connectedAt, + lastAccessedAt: connection.lastAccessedAt, + isMru: true, + protocolEra: connection.client.getProtocolEra(), + ...(connection.auth && { auth: connection.auth }), + ...(connection.pendingAuth && { pendingAuth: true }), + }; + } + + async connect( + params: ConnectParams, + signal?: AbortSignal, + ): Promise<ConnectionInfo> { + return this.withNameLock(params.name, () => + this.connectLocked(params, signal), + ); + } + + private async connectLocked( + params: ConnectParams, + signal?: AbortSignal, + ): Promise<ConnectionInfo> { + this.assertOpen(); + this.clearIdleTimer(); + this.pendingConnects++; + + try { + // Caller may already be gone (e.g. Ctrl-C while queued on the name + // lock); don't start dialing on behalf of nobody. + if (signal?.aborted) throw connectCancelledError(); + + if (this.connections.has(params.name)) { + // Reconnect: tear down the previous client first. + await this.disconnectLocked(params.name); + } + + let client: InspectorClient; + let pendingAuth = false; + try { + client = await this.dial(params, signal); + } catch (error) { + if ( + !params.pendingOnAuthRequired || + !(error instanceof CliExitCodeError) || + error.envelope?.code !== "auth_required" + ) { + throw error; + } + // Sign-in is completing out of band (detached auth helper). + // Register the intent anyway with a never-connected client: its + // status is "disconnected" (terminal), so the first op after tokens + // land goes through liveClientFor → revive and dials with the fresh + // credentials — the entry self-completes on use. + client = await createConnectionClient( + params.serverConfig, + params.serverSettings, + ); + pendingAuth = true; + } + + const now = Date.now(); + // Pending entries snapshot as unauthorized OAuth: the whole point is + // that tokens aren't in storage yet. + const auth: ConnectionAuthInfo | undefined = pendingAuth + ? { method: "oauth", authorized: false } + : await getConnectionAuthInfo(client); + if (this.closed) { + // Shutdown proceeded past its bounded quiesce grace while this + // connect was still in flight; the disconnectAll snapshot has already + // run, so registering now would leak a live transport/child process + // past daemon exit. Tear the fresh client down instead. + await safeDisconnect(client); + this.assertOpen(); + } + if (signal?.aborted) { + // Caller hung up after the connect completed but before + // registration; keeping the client would leak a live connection + // nobody asked to retain. + await safeDisconnect(client); + throw connectCancelledError(); + } + this.connections.set(params.name, { + name: params.name, + serverIdentity: params.serverIdentity, + connectedAt: now, + lastAccessedAt: now, + client, + serverConfig: params.serverConfig, + ...(params.serverSettings && { serverSettings: params.serverSettings }), + ...(auth && { auth }), + ...(pendingAuth && { pendingAuth: true }), + }); + this.mruName = params.name; + + return { + name: params.name, + serverIdentity: params.serverIdentity, + connectedAt: now, + lastAccessedAt: now, + isMru: true, + protocolEra: client.getProtocolEra(), + ...(auth && { auth }), + ...(pendingAuth && { pendingAuth: true }), + }; + } finally { + this.pendingConnects--; + // Re-arm on any exit. On success the registered connection makes this + // a no-op; on failure (createConnectionClient, reconnect disconnect, + // client.connect, …) it restores self-reaping — but only once no other + // connect is still in flight. + this.armIdleTimerIfEmpty(); + } + } + + async disconnect( + name: string | undefined, + requireExplicit: boolean | undefined, + ): Promise<{ name: string; serverUrl?: string }> { + const connectionName = this.resolve(name, requireExplicit).name; + return this.withNameLock(connectionName, () => + this.disconnectLocked(connectionName), + ); + } + + private async disconnectLocked( + connectionName: string, + ): Promise<{ name: string; serverUrl?: string }> { + // Re-resolve under the lock: a queued duplicate disconnect must fail + // with connection_not_found, not tear down a successor's connection. + const connection = this.resolve(connectionName, true); + // Captured before delete so the front-end can clear stored auth for this + // server's URL (disconnect --clear-auth). Undefined for stdio / no-URL + // configs, where there is no URL-keyed OAuth entry to clear. + const cfg = connection.serverConfig; + const serverUrl = + "url" in cfg && typeof cfg.url === "string" && cfg.url !== "" + ? cfg.url + : undefined; + this.connections.delete(connectionName); + if (this.mruName === connectionName) { + // Promote the next most-recently-accessed connection, if any. + const remaining = [...this.connections.values()].sort( + (a, b) => b.lastAccessedAt - a.lastAccessedAt, + ); + this.mruName = remaining[0]?.name ?? null; + } + await safeDisconnect(connection.client); + this.armIdleTimerIfEmpty(); + return { name: connectionName, ...(serverUrl && { serverUrl }) }; + } + + async disconnectAll(): Promise<void> { + this.closed = true; + const names = [...this.connections.keys()]; + for (const name of names) { + // Settle each teardown independently: one failed disconnect (e.g. a + // connection_not_found race with a concurrent explicit disconnect) + // must not abandon the remaining connections — that would leak live + // clients and stdio child processes, and make daemon shutdown reject + // before the socket server closes. + try { + await this.disconnect(name, false); + } catch { + // Best-effort teardown on shutdown; the connection is already gone + // or its transport close failed, neither of which should block the + // rest. + } + } + this.clearIdleTimer(); + } + + private assertOpen(): void { + if (this.closed) { + throw new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "Connection daemon is shutting down.", + { code: "daemon_stopping" }, + ); + } + } + + private armIdleTimer(): void { + this.clearIdleTimer(); + if (this.idleMs <= 0 || !this.onIdle) return; + this.idleDeadline = Date.now() + this.idleMs; + this.idleTimer = setTimeout(() => { + this.idleTimer = null; + this.idleDeadline = null; + if (this.connections.size === 0 && this.pendingConnects === 0) { + this.onIdle?.(); + } + }, this.idleMs); + // Don't keep the process alive solely for the idle timer when nothing else + // is pending — the socket server keeps the event loop alive. + this.idleTimer.unref?.(); + } + + private clearIdleTimer(): void { + if (this.idleTimer) { + clearTimeout(this.idleTimer); + this.idleTimer = null; + } + this.idleDeadline = null; + } + + /** Remaining ms until idle shutdown, or null if not armed. */ + idleRemainingMs(): number | null { + if (this.idleDeadline === null) return null; + return Math.max(0, this.idleDeadline - Date.now()); + } +} + +/** + * Connect failures that should trigger front-end interactive OAuth (then retry), + * not a hard ErrorEnvelope. Includes SDK token-exchange mistakes that happen when + * stored creds need a full re-auth. + */ +export function isConnectionAuthRequiredError(error: unknown): boolean { + if ( + error instanceof AuthRecoveryRequiredError || + isUnauthorizedError(error) + ) { + return true; + } + // EMA misconfiguration (no/disabled install-level IdP) must surface via the + // front-end too: authorizeInFrontend re-hits it in-process and maps it to + // actionable mcpdo guidance, instead of this daemon relaying the web-centric + // core message in an opaque error envelope. + if (isEmaClientNotConfiguredError(error)) { + return true; + } + const message = error instanceof Error ? error.message : String(error); + return ( + /prepareTokenRequest\(\) or authorizationCode is required/i.test(message) || + /redirectUrl is required for authorization_code/i.test(message) || + /No code verifier saved for connection/i.test(message) + ); +} + +/** + * Maps a persisted/overridden `elicitCapability` mode onto the `elicit` shape + * `InspectorClient` expects. Absence reads back as {@link + * DEFAULT_ELICIT_CAPABILITY} (`"both"`), matching the pre-#1783 hardcoded + * default so existing connections keep behaving the same until a caller opts + * into something narrower via `--elicit` or a catalog entry's + * `elicitCapability` field. + */ +export function elicitCapabilityToClientOption( + mode: ElicitCapabilityMode | undefined, +): InspectorClientOptions["elicit"] { + switch (mode ?? DEFAULT_ELICIT_CAPABILITY) { + case "off": + return false; + case "url": + return { url: true }; + case "form": + return { form: true }; + case "both": + return { url: true, form: true }; + } +} + +/** + * Project the core `OAuthConnectionState` down to the slim + * {@link ConnectionAuthInfo} reported on `ConnectionInfo`. + */ +function projectAuthState(state: OAuthConnectionState): ConnectionAuthInfo { + return { + method: state.protocol === "ema" ? "ema" : "oauth", + authorized: state.authorized, + ...(state.grantedScope && { scope: state.grantedScope }), + ...(state.client?.clientId && { clientId: state.client.clientId }), + ...(state.ema?.idpSession && { idpSession: state.ema.idpSession }), + }; +} + +/** + * Connect-time auth snapshot, read through the live client's own storage. + * Undefined for stdio servers and HTTP servers that never engaged OAuth + * (`getOAuthState()` returns undefined for both), so no-auth connections simply + * omit the field. Best-effort: a storage read failure must never fail the + * connect that already succeeded. + */ +export async function getConnectionAuthInfo( + client: InspectorClient, +): Promise<ConnectionAuthInfo | undefined> { + let state; + try { + state = await client.getOAuthState(); + } catch { + return undefined; + } + if (!state) return undefined; + return projectAuthState(state); +} + +/** + * Live auth snapshot for `connections/show`, read from *disk* rather than the + * client's storage. `NodeOAuthStorage` is load-once/memory-authoritative, so + * the live client never observes cross-process changes to `oauth.json` (an + * `auth/clear`, `auth/ema-logout`, or a web-client re-auth) — a fresh storage + * after a cache reset does. Mirrors `OAuthManager.getOAuthState()`'s inputs: + * the oauth config assembled from client.json + the saved server settings. + * Best-effort: any failure falls back to the connect-time snapshot's absence + * semantics (undefined). + */ +export async function getLiveConnectionAuthInfo(connection: { + serverConfig: MCPServerConfig; + serverSettings?: InspectorServerSettings; +}): Promise<ConnectionAuthInfo | undefined> { + try { + const config = connection.serverConfig; + if (!isOAuthCapableServerConfig(config)) return undefined; + const serverUrl = "url" in config ? config.url : undefined; + if (typeof serverUrl !== "string" || serverUrl === "") return undefined; + resetNodeOAuthStorageCache(); + const storage = new NodeOAuthStorage(); + const clientConfig = await loadRunnerClientConfig({}); + const authOptions = buildRunnerClientAuthOptions( + clientConfig, + connection.serverSettings, + {}, + ); + const oauthConfig = authOptions.oauth ?? {}; + if ( + !isServerOAuthConfigured(oauthConfig) && + !(await hasPersistedOAuthServerState(storage, serverUrl)) + ) { + return undefined; + } + return projectAuthState( + await buildOAuthConnectionState({ + serverUrl, + protocol: protocolFromOAuthConfig(oauthConfig), + configuredScope: oauthConfig.scope, + enterpriseManagedAuth: authOptions.enterpriseManagedAuth, + storage, + }), + ); + } catch { + return undefined; + } +} + +async function createConnectionClient( + serverConfig: MCPServerConfig, + serverSettings: InspectorServerSettings | undefined, +): Promise<InspectorClient> { + const environment: InspectorClientEnvironment = { + transport: createTransportNode, + }; + const redirectUrlProvider = new MutableRedirectUrlProvider(); + if (isOAuthCapableServerConfig(serverConfig)) { + // Must be non-empty: SDK treats a falsy redirectUrl as "non-interactive" and + // calls fetchToken() without an authorization code (breaking stored-token / + // refresh reconnect). Interactive login still runs in the front-end on + // auth_required; this value only keeps the daemon's silent path correct. + const callbackUrlConfig = parseRunnerOAuthCallbackUrl( + process.env.MCP_OAUTH_CALLBACK_URL ?? DEFAULT_RUNNER_OAUTH_CALLBACK_URL, + ); + redirectUrlProvider.redirectUrl = + formatRunnerOAuthRedirectUrl(callbackUrlConfig); + environment.oauth = { + storage: new NodeOAuthStorage(), + navigation: new ConsoleNavigation(), + redirectUrlProvider, + }; + } + + const clientConfig = await loadRunnerClientConfig({}); + const clientAuthOptions = buildRunnerClientAuthOptions( + clientConfig, + serverSettings, + {}, + ); + + return new InspectorClient(serverConfig, { + environment, + clientIdentity: { + name: CONNECTION_CLIENT_NAME, + version: readInspectorVersion(import.meta.url), + }, + initialLoggingLevel: "debug", + progress: false, + sample: false, + // Advertise the roots configured for this server in mcp.json, exactly as + // the one-shot CLI (`cli.ts`), web and TUI do. Passing the option (even + // empty) is what negotiates `capabilities.roots` at `initialize` and + // registers the `roots/list` handler — omitting it meant a server that + // asks for roots was told the client doesn't support them, and + // `roots/set` changed local state only (#1797). + roots: cleanRoots(serverSettings?.roots ?? []), + // Elicitation capability advertised to the server: derived from + // `serverSettings.elicitCapability` (settable via a catalog entry or the + // `--elicit` connect flag), defaulting to url+form when unset. A server + // that ignores our (possibly empty) capabilities and elicits anyway is + // defensively auto-declined by the daemon's elicitation prompt. + elicit: elicitCapabilityToClientOption(serverSettings?.elicitCapability), + serverSettings, + ...(serverSettings?.protocolEra && { + versionNegotiation: eraToVersionNegotiation(serverSettings.protocolEra), + }), + ...clientAuthOptions, + }); +} + +async function safeDisconnect(client: InspectorClient): Promise<void> { + try { + await client.disconnect(); + } catch { + // Best-effort teardown. + } +} + +/** + * Run a cancellable async operation. Guards the whole class of + * abort-listener races: `AbortSignal` does not replay its event, so any + * "check aborted, await something, then addEventListener" sequence can miss + * an abort that fired in the gap and hang forever. Here the aborted check is + * synchronous and the listener is installed *before* the operation starts, + * so no abort can interleave. When aborted pre-start the operation is never + * invoked; when aborted mid-flight its eventual settlement is still + * observed by the race handlers, so nothing rejects unhandled. + */ +async function withAbort<T>( + start: () => Promise<T>, + signal: AbortSignal | undefined, + makeError: () => Error, +): Promise<T> { + if (!signal) return start(); + if (signal.aborted) throw makeError(); + return await new Promise<T>((resolve, reject) => { + const onAbort = () => reject(makeError()); + signal.addEventListener("abort", onAbort, { once: true }); + start().then( + (value) => { + signal.removeEventListener("abort", onAbort); + resolve(value); + }, + (error: unknown) => { + signal.removeEventListener("abort", onAbort); + reject(error); + }, + ); + }); +} + +/** + * Error thrown when a connect is abandoned because the requesting client's + * socket closed. The response is written to a dead socket, so the exit code + * only matters for in-process callers; UNREACHABLE ("no connection was + * established") is the closest fit. + */ +function connectCancelledError(): CliExitCodeError { + return new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "Connect cancelled: the requesting client disconnected while the connection was still in progress.", + { code: "connect_cancelled" }, + ); +} diff --git a/clients/mcpdo/src/daemon/elicitation-bridge.ts b/clients/mcpdo/src/daemon/elicitation-bridge.ts new file mode 100644 index 0000000000..3255275967 --- /dev/null +++ b/clients/mcpdo/src/daemon/elicitation-bridge.ts @@ -0,0 +1,144 @@ +/** + * Bridges `InspectorClient`'s `newPendingElicitation` events to a mid-`rpc` + * duplex exchange with the CLI, for legacy and modern MRTR elicitations + * (dual-era support). Task-augmented MRTR elicitation (SEP-2663 + * `origin: "task-input-required"`) goes through the same delivery when a call + * is awaiting it — in the modern era core's `tools/call` polls the task to a + * terminal state, so the originating call IS still in flight and would hang + * forever without the prompt. Only when nothing subscribes (the legacy shape, + * where the call returned immediately) is it left pending for a later + * `tasks/`-driven answer instead of being cancelled. + */ +import type { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; +import type { ElicitationCreateMessage } from "@inspector/core/mcp/elicitationCreateMessage.js"; +import type { TypedEventGeneric } from "@inspector/core/mcp/typedEventTarget.js"; +import type { InspectorClientEventMap } from "@inspector/core/mcp/inspectorClientEventTarget.js"; +import type { ElicitationChannel } from "./ipc-glue.js"; +import type { ElicitationRequestFrame } from "./protocol.js"; + +/** + * Per-client bridge registry. Concurrent RPCs on the same connection would + * otherwise each install their own `newPendingElicitation` listener, so one + * server elicitation would be delivered to every active caller — duplicate + * prompts and multiple `respond()` calls. One listener per client dispatches + * each event to exactly one active subscriber. The daemon serializes `rpc` + * ops per client (see `DaemonServer.rpcQueues`), so at most one subscriber + * is active at a time and the dispatch is exact; the subscriber list (with + * its oldest-first pick) remains as defense in depth should that + * serialization ever change. + */ +type BridgeSubscriber = { channel: ElicitationChannel; requestId: string }; + +type BridgeRegistry = { + subscribers: BridgeSubscriber[]; + queue: Promise<void>; + listener: ( + event: TypedEventGeneric<InspectorClientEventMap, "newPendingElicitation">, + ) => void; +}; + +const bridgeRegistries = new WeakMap<InspectorClient, BridgeRegistry>(); + +/** + * Wires `client`'s pending-elicitation events to `channel` for the duration + * of one in-flight call. Returns a cleanup function that must be called + * (typically in a `finally`) once the call settles, so the subscription + * doesn't outlive the request. + * + * Core resolves elicitations sequentially — never more than one pending at a + * time (see `inspectorClient.ts`'s `fulfilInputRequests` and + * `requestWithInputRequired`'s retry loop) — but a single call can pause and + * resume through several of these in turn across MRTR rounds. The `queue` + * here is a defensive belt-and-suspenders in case that guarantee ever + * changes; each event is still handled one at a time, in arrival order. + */ +export function wireElicitationBridge( + client: InspectorClient, + channel: ElicitationChannel, + requestId: string, +): () => void { + let registry = bridgeRegistries.get(client); + if (!registry) { + const created: BridgeRegistry = { + subscribers: [], + queue: Promise.resolve(), + listener: (event) => { + const message = event.detail; + created.queue = created.queue.then(() => { + const subscriber = created.subscribers[0]; + if (!subscriber) { + if (message.origin === "task-input-required") { + // Task-augmented with no call awaiting it (the legacy shape, + // where the originating call returned a task id immediately): + // leave it pending for a later tasks/-driven command to + // answer, rather than cancelling — cancel would fail the task. + return; + } + // Every subscribing call settled before this event was + // dispatched — nothing is awaiting it, settle it like a channel + // failure would. + message.cancel(); + return; + } + return handleOne(subscriber.channel, subscriber.requestId, message); + }); + }, + }; + client.addEventListener("newPendingElicitation", created.listener); + bridgeRegistries.set(client, created); + registry = created; + } + const subscriber: BridgeSubscriber = { channel, requestId }; + registry.subscribers.push(subscriber); + + return () => { + const index = registry.subscribers.indexOf(subscriber); + if (index >= 0) registry.subscribers.splice(index, 1); + if (registry.subscribers.length === 0) { + client.removeEventListener("newPendingElicitation", registry.listener); + bridgeRegistries.delete(client); + } + }; +} + +async function handleOne( + channel: ElicitationChannel, + requestId: string, + message: ElicitationCreateMessage, +): Promise<void> { + const params = message.request.params; + const isUrlMode = params != null && "url" in params; + const frame: ElicitationRequestFrame = { + id: requestId, + kind: "elicitation-request", + elicitationId: message.id, + mode: isUrlMode ? "url" : "form", + message: params?.message ?? "", + requestedSchema: isUrlMode + ? undefined + : (params as { requestedSchema?: Record<string, unknown> }) + .requestedSchema, + url: isUrlMode ? (params as { url?: string }).url : undefined, + origin: message.origin, + }; + + try { + const answer = await channel.request(frame); + // Defensive: if the answer's elicitationId somehow doesn't match what we + // asked for, proceed with it anyway (single connection, single pending + // exchange at a time — this should never happen in practice) rather than + // hang the call. + await message.respond({ + action: answer.action, + content: answer.content as + | { [x: string]: string | number | boolean | string[] } + | undefined, + }); + } catch { + // Channel failure (e.g. CLI disconnected mid-prompt). `cancel()` settles + // the pending elicitation regardless of origin/mode — some construction + // sites (notably legacy URL-mode's `awaitUrlElicitation`) never wire a + // reject callback, so `reject()` alone would leave the call hanging. + message.cancel(); + } +} diff --git a/clients/mcpdo/src/daemon/elicitation-park.ts b/clients/mcpdo/src/daemon/elicitation-park.ts new file mode 100644 index 0000000000..3c5ddf4d1e --- /dev/null +++ b/clients/mcpdo/src/daemon/elicitation-park.ts @@ -0,0 +1,265 @@ +/** + * Daemon-side parking for elicitations from non-interactive callers + * (dual-era support, phase 2). Instead of relaying an elicitation over the + * socket for an inline prompt — which a non-TTY agent can never answer, its + * stdin isn't wired to the human — the in-flight call is parked here: the + * originating `rpc` returns immediately with `kind: "elicitation-pending"`, + * and a later `elicitation/respond` op answers the exchange and picks up + * either the final call result or the next pending round. Works identically + * for both eras because the bridge funnels legacy server→client requests and + * modern non-task MRTR rounds through the same `ElicitationChannel` seam. + */ +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import type { ElicitationChannel } from "./ipc-glue.js"; +import type { + ElicitationPendingInfo, + ElicitationRequestFrame, + ElicitationResponseFrame, + RpcResult, +} from "./protocol.js"; + +/** + * How long an unanswered parked elicitation lives before the daemon cancels + * it. Long enough for an agent to relay a form to a human and collect + * answers; bounded so a caller that vanishes can't hold the server's + * elicitation request (and the parked call) open forever. + */ +export const PARKED_ELICITATION_TTL_MS = 10 * 60_000; + +/** + * {@link ElicitationChannel} that parks instead of prompting: `request()` + * returns a promise nobody answers until `elicitation/respond` calls + * {@link answer}. `waitForElicitation()` lets the rpc/respond handlers race + * the in-flight call against the next elicitation arriving. `close()` + * cancel-settles the current exchange and every future one (expiry or + * connection teardown) — the bridge's channel-failure path then `cancel()`s + * the underlying message, so the parked call always settles. + */ +export class ParkingElicitationChannel implements ElicitationChannel { + private pending: { + frame: ElicitationRequestFrame; + resolve: (frame: ElicitationResponseFrame) => void; + reject: (error: Error) => void; + } | null = null; + private waiter: ((frame: ElicitationRequestFrame) => void) | null = null; + private closed: Error | null = null; + + request(frame: ElicitationRequestFrame): Promise<ElicitationResponseFrame> { + if (this.closed) return Promise.reject(this.closed); + return new Promise((resolve, reject) => { + this.pending = { frame, resolve, reject }; + if (this.waiter) { + const waiter = this.waiter; + this.waiter = null; + waiter(frame); + } + }); + } + + /** Resolves when the next elicitation arrives; never rejects. */ + waitForElicitation(): Promise<ElicitationRequestFrame> { + if (this.pending) return Promise.resolve(this.pending.frame); + return new Promise((resolve) => { + this.waiter = resolve; + }); + } + + /** The frame of the exchange currently awaiting an answer, if any. */ + pendingFrame(): ElicitationRequestFrame | null { + return this.pending?.frame ?? null; + } + + /** + * Answer the pending exchange. A no-op when nothing is pending (the call + * settled on its own — e.g. the server timed out its elicitation and + * completed anyway); the caller then just picks up the settled outcome. + */ + answer(response: ElicitationResponseFrame): void { + const pending = this.pending; + this.pending = null; + pending?.resolve(response); + } + + /** Cancel-settle the pending exchange and auto-cancel all future ones. */ + close(error: Error): void { + this.closed = error; + const pending = this.pending; + this.pending = null; + this.waiter = null; + pending?.reject(error); + } +} + +export type ParkedCall = { + /** Current round's payload; re-pointed by {@link ElicitationParkRegistry.rearm}. */ + info: ElicitationPendingInfo; + /** Client the call runs on — guards against new rpcs interleaving. */ + client: InspectorClient; + channel: ParkingElicitationChannel; + /** Settles when the parked daemon-side call finishes (result or error). */ + outcome: Promise<RpcResult>; + /** + * Removes the call's bridge subscriber. Normally the call's own `finally` + * unwires when the underlying call settles — but a cancelled/expired park + * abandons a call that is still running server-side, and its subscriber + * would otherwise stay first in the bridge's dispatch order until the + * server settles it, swallowing (auto-cancelling) the next elicitation of + * any new call started in that window. Park teardown owns the unwire so a + * closed channel is never a dispatch target. Idempotent. + */ + unwire: () => void; + /** + * Cancels the still-running server call (sends `notifications/cancelled`). + * `unwire` only drops the dead subscriber; the abandoned call itself keeps + * running server-side, and a tool call that survives its errored elicitation + * could emit a *second* one. With this park's subscriber gone, the bridge — + * which cannot attribute an elicitation to a specific call — would misroute + * that second prompt to whatever new rpc has since started on this + * connection. Cancelling the call on teardown/expiry closes that window. + * A no-op when the parked method isn't a tool call (no active controller), + * the case where no second elicitation realistically arises. Idempotent. + */ + cancelCall: () => void; + /** + * `awaiting` = parked, answerable; `responding` = an `elicitation/respond` + * is in flight for it (a concurrent respond must not double-answer). + */ + state: "awaiting" | "responding"; + timer: ReturnType<typeof setTimeout> | null; +}; + +/** + * All parked calls, keyed by the current round's `elicitationId`. At most + * one per connection: the daemon serializes rpcs per client and refuses new + * rpcs on a connection with a parked call (the bridge would misroute a + * second call's elicitations to the parked subscriber). + */ +export class ElicitationParkRegistry { + private readonly byId = new Map<string, ParkedCall>(); + private readonly ttlMs: number; + + constructor(ttlMs: number = PARKED_ELICITATION_TTL_MS) { + this.ttlMs = ttlMs; + } + + /** The parked call running on `client`, whatever its state, if any. */ + forClient(client: InspectorClient): ParkedCall | undefined { + for (const entry of this.byId.values()) { + if (entry.client === client) return entry; + } + return undefined; + } + + /** Park a call. `info.expiresAt` is set here from the registry's TTL. */ + add(entry: { + info: Omit<ElicitationPendingInfo, "expiresAt">; + client: InspectorClient; + channel: ParkingElicitationChannel; + outcome: Promise<RpcResult>; + unwire: () => void; + cancelCall: () => void; + }): ParkedCall { + const parked: ParkedCall = { + ...entry, + info: { ...entry.info, expiresAt: Date.now() + this.ttlMs }, + state: "awaiting", + timer: null, + }; + this.byId.set(parked.info.elicitationId, parked); + this.armTimer(parked); + return parked; + } + + /** + * Claim a parked call for one `elicitation/respond`. Removes the id from + * the awaitable state so a concurrent respond can't double-answer. + */ + take(elicitationId: string): ParkedCall { + const entry = this.byId.get(elicitationId); + if (!entry || entry.state !== "awaiting") { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `No pending elicitation '${elicitationId}' — it may have expired, been answered, or belong to a connection that closed.`, + { code: "elicitation_not_found" }, + ); + } + entry.state = "responding"; + this.clearTimer(entry); + return entry; + } + + /** Park the next round of an already-claimed call under a new id. */ + rearm(entry: ParkedCall, info: Omit<ElicitationPendingInfo, "expiresAt">) { + this.byId.delete(entry.info.elicitationId); + entry.info = { ...info, expiresAt: Date.now() + this.ttlMs }; + entry.state = "awaiting"; + this.byId.set(entry.info.elicitationId, entry); + this.armTimer(entry); + return entry.info; + } + + /** The parked call settled; forget it. */ + finish(entry: ParkedCall): void { + this.clearTimer(entry); + this.byId.delete(entry.info.elicitationId); + } + + /** Connection going away (disconnect / replacing connect): cancel its parked call. */ + cancelForConnection(connectionName: string): void { + for (const entry of this.byId.values()) { + if (entry.info.connection === connectionName) this.cancel(entry); + } + } + + /** Daemon shutdown: cancel everything so no server request is left held. */ + cancelAll(): void { + for (const entry of this.byId.values()) this.cancel(entry); + } + + private cancel(entry: ParkedCall): void { + this.finish(entry); + // Stop the abandoned call at the source before tearing down local state: + // a tool call that outlives this park could otherwise emit a second + // elicitation that misroutes to a new rpc (see ParkedCall.cancelCall). + entry.cancelCall(); + // The abandoned call may run server-side long after this park is gone; + // unwire its bridge subscriber now so it cannot shadow a new call's + // elicitations (see ParkedCall.unwire). + entry.unwire(); + entry.channel.close( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "Parked elicitation cancelled: the connection or daemon is going away.", + { code: "elicitation_cancelled" }, + ), + ); + } + + private armTimer(entry: ParkedCall): void { + if (this.ttlMs <= 0) return; + entry.timer = setTimeout(() => { + this.finish(entry); + entry.cancelCall(); + entry.unwire(); + entry.channel.close( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "Parked elicitation expired unanswered.", + { code: "elicitation_expired" }, + ), + ); + }, this.ttlMs); + entry.timer.unref?.(); + } + + private clearTimer(entry: ParkedCall): void { + if (entry.timer) { + clearTimeout(entry.timer); + entry.timer = null; + } + } +} diff --git a/clients/mcpdo/src/daemon/ensure.ts b/clients/mcpdo/src/daemon/ensure.ts new file mode 100644 index 0000000000..71493b9545 --- /dev/null +++ b/clients/mcpdo/src/daemon/ensure.ts @@ -0,0 +1,277 @@ +import { spawn } from "node:child_process"; +import * as fs from "node:fs"; +import * as net from "node:net"; +import * as path from "node:path"; +import { fileURLToPath } from "node:url"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { + generateDaemonToken, + getDaemonTokenFromEnv, + readDaemonTokenFile, +} from "./auth.js"; +import { callDaemon } from "./client.js"; +import { + DAEMON_DIR_ENV, + DAEMON_TOKEN_ENV, + assertSocketPathWithinLimit, + ensureDaemonDir, + getDaemonDir, + getDaemonLogPath, + getDaemonSocketPath, +} from "./paths.js"; + +const READY_TIMEOUT_MS = 10_000; +const READY_POLL_MS = 50; + +/** + * Resolve the built daemon entry (`build/mcpdod.js`) next to this package's + * build output. When running from source under vitest, prefer the built file + * if present; otherwise throw a clear error. + */ +export function resolveDaemonScriptPath(): string { + // ensure.ts lives at src/daemon/ensure.ts → ../../build/mcpdod.js + // In the bundle, import.meta.url is build/mcpdod-*.js or similar; tsup emits + // ensure into the daemon entry chunk. Prefer an explicit sibling mcpdod.js. + const here = path.dirname(fileURLToPath(import.meta.url)); + const candidates = [ + path.resolve(here, "mcpdod.js"), + path.resolve(here, "../mcpdod.js"), + path.resolve(here, "../../build/mcpdod.js"), + path.resolve(here, "../build/mcpdod.js"), + ]; + for (const candidate of candidates) { + if (fs.existsSync(candidate)) return candidate; + } + /* v8 ignore next 6 -- only when clients/cli/build is missing; pretest always + builds, and fs.existsSync cannot be spied in this ESM package under vitest. */ + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `Connection daemon bundle not found (looked for mcpdod.js near ${here}). Run npm run build in clients/mcpdo.`, + { code: "daemon_not_built" }, + ); +} + +async function isDaemonReachable(socketPath: string): Promise<boolean> { + return new Promise((resolve) => { + let settled = false; + const socket = new net.Socket(); + const done = (ok: boolean) => { + /* v8 ignore next -- re-entry when connect and error both fire */ + if (settled) return; + settled = true; + socket.removeAllListeners(); + socket.on("error", () => {}); + socket.destroy(); + resolve(ok); + }; + socket.on("error", () => done(false)); + socket.setTimeout(500); + socket.once("connect", () => done(true)); + /* v8 ignore next -- 500ms probe timeout; ensureDaemon usually connects faster */ + socket.once("timeout", () => done(false)); + socket.connect(socketPath); + }); +} + +async function waitForDaemon( + socketPath: string, + token: string, + logPath: string, + opts?: { + timeoutMs?: number; + /** + * Set only when `token` was self-generated (shared mode). Two concurrent + * first invocations each generate a token and spawn; the pid lock lets + * one daemon survive, and it may not be ours. Re-reading the winner's + * published `mcpdod.token` between polls lets the losing caller finish + * against the surviving daemon instead of timing out on auth failures. + * Explicitly supplied / private-mode tokens never fall back — a mismatch + * there must stay a loud failure. + */ + rereadTokenDir?: string; + }, +): Promise<void> { + const deadline = Date.now() + (opts?.timeoutMs ?? READY_TIMEOUT_MS); + while (Date.now() < deadline) { + if (await isDaemonReachable(socketPath)) { + const effectiveToken = opts?.rereadTokenDir + ? (readDaemonTokenFile(opts.rereadTokenDir) ?? token) + : token; + try { + await callDaemon( + "ping", + {}, + { socketPath, timeoutMs: 2000, token: effectiveToken }, + ); + return; + } catch { + // connected but not ready yet + } + } + await new Promise((r) => setTimeout(r, READY_POLL_MS)); + } + const logTail = readLogTail(logPath); + throw new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Timed out waiting for connection daemon at ${socketPath}` + + (logTail ? `\nDaemon log (${logPath}):\n${logTail}` : ""), + { code: "daemon_start_timeout" }, + ); +} + +/** Last few lines of the daemon's stderr log — the only trace of a spawn + * that died before binding its socket. Best-effort. Exported for tests. */ +export function readLogTail(logPath: string, maxLines = 10): string { + try { + const text = fs.readFileSync(logPath, "utf8"); + return text.trimEnd().split("\n").slice(-maxLines).join("\n"); + } catch { + return ""; + } +} + +/** How long ensureDaemon waits for a stopping daemon to finish exiting. */ +export const STOPPING_EXIT_TIMEOUT_MS = 10_000; +const STOPPING_POLL_MS = 100; + +/** Signal-0 liveness probe; EPERM means alive but not ours. */ +function pidAlive(pid: number): boolean { + try { + process.kill(pid, 0); + return true; + } catch (error) { + return (error as NodeJS.ErrnoException).code === "EPERM"; + } +} + +/** + * Wait for a shutting-down daemon to actually exit. Keyed on the process + * (which holds the daemon lock until it dies), not the socket — the socket + * closes earlier in shutdown, and spawning in that gap would die on the + * still-held lock. Falls back to socket reachability when the ping predates + * the `pid` field. Exported for tests. + */ +export async function waitForDaemonExit( + pid: number | undefined, + socketPath: string, + timeoutMs = STOPPING_EXIT_TIMEOUT_MS, + pollMs = STOPPING_POLL_MS, +): Promise<void> { + const deadline = Date.now() + timeoutMs; + for (;;) { + const gone = + pid !== undefined + ? !pidAlive(pid) + : !(await isDaemonReachable(socketPath)); + if (gone) return; + if (Date.now() >= deadline) { + throw new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "Connection daemon is shutting down but did not exit in time; retry shortly.", + { code: "daemon_stopping" }, + ); + } + await new Promise((r) => setTimeout(r, pollMs)); + } +} + +/** + * Ensure a connection daemon is running for the current {@link getDaemonDir}. + * Auto-spawns a detached Node process when the socket is not reachable. + * + * When `MCP_INSPECTOR_DAEMON_TOKEN` is set (private mode), the child inherits + * that token; otherwise a fresh token is generated for the child. Either way + * every IPC call must present it (clients that didn't spawn the daemon read + * it from the published `mcpdod.token` file). + */ +export async function ensureDaemon(options?: { + dir?: string; + daemonScript?: string; + token?: string; + /** Startup wait override (tests exercise the timeout path). */ + readyTimeoutMs?: number; +}): Promise<{ socketPath: string; spawned: boolean }> { + const dir = options?.dir ?? getDaemonDir(); + let token = options?.token ?? getDaemonTokenFromEnv(); + ensureDaemonDir(dir); + const socketPath = getDaemonSocketPath(dir); + // Fail here with an actionable error rather than letting the daemon's + // listen() die over sun_path limits with only a generic start timeout. + assertSocketPathWithinLimit(socketPath); + + if (await isDaemonReachable(socketPath)) { + // Something accepted the connection, so a live daemon owns this socket. + // Any ping failure here (daemon_auth_failed, timeout, protocol error) + // must fail loudly: unlinking and respawning would let a caller with the + // wrong token (or none) silently replace a live private daemon and + // orphan its connections. Only a socket nothing is listening on — the + // unreachable path below — is stale, and the spawned daemon itself + // removes it after a connect probe (removeStaleDaemonSocket). + token ??= readDaemonTokenFile(dir); + const pong = await callDaemon<{ + pong: boolean; + pid?: number; + stopping?: boolean; + }>("ping", {}, { socketPath, timeoutMs: 2000, token }); + if (pong?.stopping !== true) { + return { socketPath, spawned: false }; + } + // The daemon answered but is shutting down; it would reject real work + // with daemon_stopping, and a respawn now would die on its still-held + // lock. To the caller the daemon is simply "up": wait for the old + // process to finish exiting, then fall through and spawn a fresh one. + await waitForDaemonExit(pong.pid, socketPath); + } + + // Every daemon requires a token; generate one for the child when the + // caller/environment didn't supply one. The daemon republishes it to + // mcpdod.token (0600) so unrelated clients can still connect. + const tokenWasGenerated = token === undefined; + token ??= generateDaemonToken(); + const script = options?.daemonScript ?? resolveDaemonScriptPath(); + const childEnv: NodeJS.ProcessEnv = { + ...process.env, + // Pin the socket directory explicitly so parent and child agree even when + // MCP_STORAGE_DIR is unset (default ~/.mcp-inspector). + [DAEMON_DIR_ENV]: dir, + [DAEMON_TOKEN_ENV]: token, + }; + + // Detached + stdio "ignore" made every startup failure invisible. Capture + // stderr in a 0600 log the start-timeout error can quote. + const logPath = getDaemonLogPath(dir); + let stderrTarget: number | "ignore" = "ignore"; + try { + // Recreate exclusively: append-open follows symlinks and applies the mode + // only on create, so a pre-existing mcpdod.log could stay group/other- + // readable or redirect daemon stderr to a planted target. The parent dir + // was just tightened to 0700; removing the entry closes the window for + // children planted before that. + fs.rmSync(logPath, { force: true }); + stderrTarget = fs.openSync(logPath, "ax", 0o600); + /* v8 ignore next 3 -- log capture is best-effort; openSync on a freshly + ensured 0700 dir cannot be made to fail portably in tests. */ + } catch { + // The daemon still runs without a log. + } + + const child = spawn(process.execPath, [script], { + detached: true, + stdio: ["ignore", "ignore", stderrTarget], + env: childEnv, + }); + child.unref(); + /* v8 ignore next -- "ignore" only when the best-effort openSync failed */ + if (typeof stderrTarget === "number") { + fs.closeSync(stderrTarget); + } + + await waitForDaemon(socketPath, token, logPath, { + timeoutMs: options?.readyTimeoutMs, + rereadTokenDir: tokenWasGenerated ? dir : undefined, + }); + return { socketPath, spawned: true }; +} diff --git a/clients/mcpdo/src/daemon/framing.ts b/clients/mcpdo/src/daemon/framing.ts new file mode 100644 index 0000000000..fb9822f008 --- /dev/null +++ b/clients/mcpdo/src/daemon/framing.ts @@ -0,0 +1,28 @@ +import type { DaemonRequest, DaemonResponse } from "./protocol.js"; + +/** + * Parse one NDJSON line into a daemon request. Returns null for blank lines. + */ +export function parseRequestLine(line: string): DaemonRequest | null { + const trimmed = line.trim(); + if (!trimmed) return null; + const value: unknown = JSON.parse(trimmed); + if ( + value === null || + typeof value !== "object" || + Array.isArray(value) || + typeof (value as DaemonRequest).id !== "string" || + typeof (value as DaemonRequest).op !== "string" + ) { + throw new Error("Invalid daemon request: expected { id, op, params? }"); + } + return value as DaemonRequest; +} + +export function encodeResponse(response: DaemonResponse): string { + return JSON.stringify(response) + "\n"; +} + +export function encodeRequest(request: DaemonRequest): string { + return JSON.stringify(request) + "\n"; +} diff --git a/clients/mcpdo/src/daemon/index.ts b/clients/mcpdo/src/daemon/index.ts new file mode 100644 index 0000000000..3b55d12fa7 --- /dev/null +++ b/clients/mcpdo/src/daemon/index.ts @@ -0,0 +1,40 @@ +export { + assertDaemonToken, + generateDaemonToken, + getDaemonTokenFromEnv, + readDaemonTokenFile, + tokensEqual, +} from "./auth.js"; +export { callDaemon } from "./client.js"; +export { streamDaemon } from "./stream-client.js"; +export { ensureDaemon, resolveDaemonScriptPath } from "./ensure.js"; +export { encodeRequest, encodeResponse, parseRequestLine } from "./framing.js"; +export { + assertSocketPathWithinLimit, + createPrivateDaemonDir, + DAEMON_DIR_ENV, + DAEMON_TOKEN_ENV, + ensureDaemonDir, + getDaemonDir, + getDaemonLockPath, + getDaemonLogPath, + getDaemonSocketPath, + getDaemonTokenPath, +} from "./paths.js"; +export type { + ConnectParams, + DaemonOp, + DaemonRequest, + DaemonResponse, + DaemonStatus, + RpcParams, + RpcResult, + ConnectionInfo, + ConnectionNameParams, +} from "./protocol.js"; +export { DaemonServer } from "./server.js"; +export { + DEFAULT_IDLE_MS, + isConnectionAuthRequiredError, + ConnectionRegistry, +} from "./connections.js"; diff --git a/clients/mcpdo/src/daemon/ipc-glue.ts b/clients/mcpdo/src/daemon/ipc-glue.ts new file mode 100644 index 0000000000..9b78745281 --- /dev/null +++ b/clients/mcpdo/src/daemon/ipc-glue.ts @@ -0,0 +1,311 @@ +/** + * Low-level Unix-socket accept / stale-socket helpers for {@link DaemonServer}. + */ +import * as fs from "node:fs"; +import * as net from "node:net"; +import { createInterface } from "node:readline"; +import { encodeResponse, parseRequestLine } from "./framing.js"; +import type { + DaemonRequest, + DaemonResponse, + DaemonStreamFrame, + ElicitationRequestFrame, + ElicitationResponseFrame, +} from "./protocol.js"; + +/** + * Starts a stream producer. `writeData` pushes one data frame; `endStream` + * lets the producer side finish the stream cleanly (end frame + socket end), + * e.g. when the underlying connection is torn down. Returns an unsubscribe. + */ +export type StreamStarter = ( + writeData: (data: unknown) => void, + endStream: () => void, +) => () => void; + +/** Result of handling one daemon request — optional long-lived stream. */ +export type HandleOutcome = { + response: DaemonResponse; + /** When set, keep the socket open and push stream frames until closed. */ + startStream?: StreamStarter; +}; + +/** + * Bridges a single in-flight `rpc` call to its owning connection so it can + * pause mid-call for a legacy/modern-non-task elicitation, and resume once + * the CLI answers. See `ElicitationRequestFrame`'s doc comment in + * `protocol.ts` for why one exchange (repeatable) is all a single connection + * ever needs. + */ +export type ElicitationChannel = { + request(frame: ElicitationRequestFrame): Promise<ElicitationResponseFrame>; +}; + +export type HandleRequest = ( + request: DaemonRequest, + elicitation: ElicitationChannel, + /** + * Aborted when the requesting socket closes, so long-running handlers + * (notably `connect`) can stop work whose caller is gone. + */ + signal?: AbortSignal, +) => Promise<HandleOutcome>; + +/** + * Upper bound on a single NDJSON request line. A client that streams an + * unterminated line would otherwise grow readline's buffer without limit — + * a trivial local DoS on the daemon. 1 MiB is far beyond any legitimate + * request (tool args included) while staying cheap to buffer. + */ +export const MAX_REQUEST_LINE_BYTES = 1024 * 1024; + +/** + * Upper bound on unflushed stream-frame bytes buffered for one socket. + * `socket.write()` queues without limit when the peer stops reading; a + * high-rate stream (e.g. `logging/tail`) to a slow client would otherwise + * grow the daemon's heap without bound. Once exceeded, the stream is + * terminated: the producer is unsubscribed and the socket destroyed, which + * the client reports as an interrupted stream (`daemon_unreachable`). + */ +export const MAX_STREAM_BUFFER_BYTES = 1024 * 1024; + +/** + * Per-connection {@link ElicitationChannel}. Writes an elicitation-request + * frame straight onto the socket (ahead of the eventual `DaemonResponse`) and + * waits for the next line to answer it; `acceptDaemonConnection`'s line + * handler gives that next line to {@link tryConsumeLine} instead of parsing + * it as a new top-level request. Rejects any pending exchange if the socket + * disconnects, so a dropped client can't hang the daemon-side call forever. + */ +class ConnectionElicitationChannel implements ElicitationChannel { + private pending: { + resolve: (frame: ElicitationResponseFrame) => void; + reject: (error: Error) => void; + } | null = null; + + constructor(private readonly socket: net.Socket) { + const onDisconnect = () => this.rejectPending("Connection closed"); + socket.once("close", onDisconnect); + socket.once("error", onDisconnect); + } + + request(frame: ElicitationRequestFrame): Promise<ElicitationResponseFrame> { + if (this.pending) { + return Promise.reject( + new Error("Another elicitation is already pending on this connection"), + ); + } + return new Promise((resolve, reject) => { + this.pending = { resolve, reject }; + if (this.socket.destroyed) { + this.rejectPending("Connection closed"); + return; + } + this.socket.write(JSON.stringify(frame) + "\n"); + }); + } + + /** Returns true if this line was consumed as a pending elicitation answer. */ + tryConsumeLine(line: string): boolean { + if (!this.pending) return false; + let parsed: ElicitationResponseFrame; + try { + parsed = JSON.parse(line); + } catch { + return false; + } + if (!parsed || parsed.kind !== "elicitation-response") return false; + const { resolve } = this.pending; + this.pending = null; + resolve(parsed); + return true; + } + + private rejectPending(message: string): void { + if (!this.pending) return; + const { reject } = this.pending; + this.pending = null; + reject(new Error(message)); + } +} + +export function acceptDaemonConnection( + socket: net.Socket, + handle: HandleRequest, +): void { + // Enforce the line cap below readline: track bytes since the last newline + // and drop the connection once a single line exceeds the limit. Every + // newline-delimited segment is checked at its full accumulated size before + // the counter resets — checking only the tail of a chunk would let an + // oversized line slip through whenever the chunk that crosses the limit + // also contains the terminating newline. + let bytesSinceNewline = 0; + let rejected = false; + socket.on("data", (chunk: Buffer) => { + if (rejected) return; + let start = 0; + for (;;) { + const idx = chunk.indexOf(0x0a, start); + bytesSinceNewline += (idx === -1 ? chunk.length : idx) - start; + if (bytesSinceNewline > MAX_REQUEST_LINE_BYTES) { + rejected = true; + // No error argument: nothing useful can be written back on a socket + // that's mid-way through an oversized line; just drop it. + socket.destroy(); + return; + } + if (idx === -1) return; + bytesSinceNewline = 0; + start = idx + 1; + } + }); + const rl = createInterface({ input: socket, crlfDelay: Infinity }); + // readline re-emits input errors on the interface; without a listener a + // client RST would crash the daemon with an unhandled 'error' event. The + // socket's own error handler below owns the teardown. + rl.on("error", () => {}); + const elicitationChannel = new ConnectionElicitationChannel(socket); + // Cancellation for in-flight handlers: when the caller hangs up (Ctrl-C + // closes its socket) the handler should stop working on its behalf — + // e.g. abort a `connect` that would otherwise keep dialing indefinitely. + const requestAbort = new AbortController(); + socket.once("close", () => requestAbort.abort()); + rl.on("line", (line) => { + void (async () => { + // readline sees the same chunks as the cap enforcement above, so an + // oversized-but-terminated line can still surface here in the same + // tick the connection was rejected — never hand it to a handler. + if (rejected) return; + if (elicitationChannel.tryConsumeLine(line)) return; + let request: DaemonRequest; + try { + const parsed = parseRequestLine(line); + if (!parsed) return; + request = parsed; + } catch (error) { + socket.write( + encodeResponse({ + id: "?", + ok: false, + error: { + code: "invalid_request", + message: error instanceof Error ? error.message : String(error), + }, + }), + ); + return; + } + const outcome = await handle( + request, + elicitationChannel, + requestAbort.signal, + ); + if (socket.destroyed) { + // The caller vanished while the handler ran. A stream outcome may + // already hold producer-side state (resources/subscribe subscribes + // before returning its starter), so start it inert and stop it + // immediately — otherwise the daemon keeps a hidden subscription + // with no consumer. + if (outcome.response.ok && outcome.startStream) { + try { + const stop = outcome.startStream( + () => {}, + () => {}, + ); + stop(); + } catch { + // ignore cleanup errors + } + } + return; + } + socket.write(encodeResponse(outcome.response)); + + if (!outcome.response.ok || !outcome.startStream) { + return; + } + + const id = request.id; + let stopped = false; + let stop: (() => void) | undefined = undefined; + const stopProducer = () => { + try { + stop?.(); + } catch { + // ignore unsubscribe errors + } + }; + const writeData = (data: unknown) => { + if (stopped || socket.destroyed) return; + const frame: DaemonStreamFrame = { id, stream: "data", data }; + socket.write(JSON.stringify(frame) + "\n"); + if (socket.writableLength > MAX_STREAM_BUFFER_BYTES) { + // Slow/non-reading client: cap the buffered backlog instead of + // exhausting the daemon heap. Destroying (no end frame) makes the + // client report an interrupted stream rather than a clean finish. + stopped = true; + stopProducer(); + socket.destroy(); + } + }; + const cleanup = () => { + if (stopped) return; + stopped = true; + stopProducer(); + if (!socket.destroyed) { + const end: DaemonStreamFrame = { id, stream: "end" }; + socket.write(JSON.stringify(end) + "\n"); + socket.end(); + } + }; + stop = outcome.startStream(writeData, cleanup); + if (stopped) { + // Overflow or a producer-side end hit while startStream was still + // running (synchronous producer): `stop` wasn't assigned yet, + // unsubscribe it now. + stopProducer(); + return; + } + socket.once("close", cleanup); + socket.once("error", cleanup); + })(); + }); + socket.on("error", () => { + rl.close(); + }); +} + +export async function removeStaleDaemonSocket( + socketPath: string, +): Promise<void> { + if (!fs.existsSync(socketPath)) return; + const live = await canConnect(socketPath); + if (live) { + throw new Error( + `Daemon already running at ${socketPath}. Use mcpdo daemon stop first.`, + ); + } + try { + fs.unlinkSync(socketPath); + } catch { + // ignore + } +} + +async function canConnect(socketPath: string): Promise<boolean> { + return new Promise((resolve) => { + let settled = false; + const socket = new net.Socket(); + const done = (ok: boolean) => { + if (settled) return; + settled = true; + socket.removeAllListeners(); + socket.on("error", () => {}); + socket.destroy(); + resolve(ok); + }; + socket.once("connect", () => done(true)); + socket.once("error", () => done(false)); + socket.connect(socketPath); + }); +} diff --git a/clients/mcpdo/src/daemon/paths.ts b/clients/mcpdo/src/daemon/paths.ts new file mode 100644 index 0000000000..dfded04795 --- /dev/null +++ b/clients/mcpdo/src/daemon/paths.ts @@ -0,0 +1,172 @@ +import { createHash, randomBytes } from "node:crypto"; +import * as fs from "node:fs"; +import * as os from "node:os"; +import * as path from "node:path"; + +/** Env: directory that owns mcpdod.sock + mcpdod.lock. */ +export const DAEMON_DIR_ENV = "MCP_INSPECTOR_DAEMON_DIR"; + +/** + * Env: IPC bearer token. Every daemon requires one: set it explicitly for + * private mode, or leave it unset and the daemon generates one at startup + * and publishes it to `mcpdod.token` (see {@link getDaemonTokenPath}). + */ +export const DAEMON_TOKEN_ENV = "MCP_INSPECTOR_DAEMON_TOKEN"; + +/** + * Directory that owns the daemon socket + lock. + * Precedence: + * 1. `MCP_INSPECTOR_DAEMON_DIR` — explicit (private mode / auto-spawn parent) + * 2. `MCP_STORAGE_DIR` — CI / parallel isolation (same override as oauth.json) + * 3. `~/.mcp-inspector` + */ +export function getDaemonDir(): string { + const daemonDir = process.env[DAEMON_DIR_ENV]?.trim(); + if (daemonDir) return path.resolve(daemonDir); + const storage = process.env.MCP_STORAGE_DIR?.trim(); + if (storage) return path.resolve(storage); + /* v8 ignore next 2 -- USERPROFILE is the Windows fallback; CI/darwin use HOME. */ + const home = process.env.HOME || process.env.USERPROFILE || os.homedir(); + return path.join(home, ".mcp-inspector"); +} + +/** + * Create a new private daemon directory (mode `0700`). Does not start the + * daemon. + * + * Lives under `$TMPDIR/mcp-conn-<uid>/<id>/`, not `~/.mcp-inspector`: `sun_path` + * caps Unix socket paths at 104 bytes on macOS (108 on Linux), and the tmp + * dir is short on every platform (macOS's per-user `/var/folders/...` is the + * long case, and even that fits with the 8-char id). The parent + * `mcp-conn-<uid>` dir is also created 0700 so the layout never depends on the + * platform's default tmp permissions. + */ +export function createPrivateDaemonDir(): string { + /* v8 ignore next 2 -- getuid is missing only on Windows */ + const uid = typeof process.getuid === "function" ? process.getuid() : "u"; + const root = path.join(os.tmpdir(), `mcp-conn-${uid}`); + fs.mkdirSync(root, { recursive: true, mode: 0o700 }); + // The tmpdir parent is world-writable on shared machines, and mkdir with + // `recursive: true` succeeds silently over a pre-existing entry — never + // trust a root another user could have planted (dir or symlink) before + // this user's first run. + assertTrustedPrivateRoot(root); + const id = randomBytes(4).toString("hex"); + const dir = path.join(root, id); + // Non-recursive mkdir is exclusive: an existing entry (however unlikely + // under a now-verified 0700 root) throws instead of being adopted. + fs.mkdirSync(dir, { mode: 0o700 }); + try { + fs.chmodSync(dir, 0o700); + } catch { + // best-effort on platforms that ignore mode + } + return dir; +} + +/** + * Fail closed unless `dir` is a real directory (not a symlink) owned by the + * current user, and tighten its mode to 0700. On a shared `/tmp`, another + * user who pre-created the predictable `mcp-conn-<uid>` path — or planted a + * symlink there — would otherwise keep write control over where the daemon's + * socket and token land. No-op on Windows (no getuid/UNIX mode semantics). + * Exported for tests. + */ +export function assertTrustedPrivateRoot(dir: string): void { + /* v8 ignore next -- Windows-only: no getuid */ + if (typeof process.getuid !== "function") return; + const st = fs.lstatSync(dir); + if (!st.isDirectory()) { + throw new Error( + `Refusing to use ${dir}: not a directory (a file or symlink was planted in its place).`, + ); + } + /* v8 ignore next 5 -- requires a second uid to create the dir; untestable without root */ + if (st.uid !== process.getuid()) { + throw new Error( + `Refusing to use ${dir}: owned by uid ${st.uid}, not the current user (uid ${process.getuid()}).`, + ); + } + if ((st.mode & 0o077) !== 0) { + // We own it, so chmod either succeeds or throws — a failure here must + // stay fatal rather than leaving a group/other-accessible daemon dir. + fs.chmodSync(dir, 0o700); + } +} + +/** + * IPC endpoint for the daemon owning `dir`. On Unix this is a socket file + * inside the directory. On Windows, `net` requires named-pipe paths + * (`\\.\pipe\...`) — a filesystem path never binds — so derive a + * deterministic per-directory pipe name: every client of the same daemon + * dir dials the same pipe, and distinct dirs (private mode, tests) never + * collide. The dir is resolved and lowercased first, matching Windows + * path-comparison semantics. + */ +export function getDaemonSocketPath(dir: string = getDaemonDir()): string { + if (process.platform === "win32") { + const hash = createHash("sha256") + .update(path.resolve(dir).toLowerCase()) + .digest("hex") + .slice(0, 16); + return `\\\\.\\pipe\\mcp-conn-${hash}`; + } + return path.join(dir, "mcpdod.sock"); +} + +export function getDaemonLockPath(dir: string = getDaemonDir()): string { + return path.join(dir, "mcpdod.lock"); +} + +/** + * IPC token published by a running daemon (0600, inside the 0700 daemon + * dir). Written on start, removed on shutdown. Lets clients that didn't + * spawn the daemon (and so have no `MCP_INSPECTOR_DAEMON_TOKEN` in their + * environment) authenticate: filesystem permissions on the file are the + * trust boundary, which is exactly the same-user boundary the socket has — + * but requests now always carry a token, so there is no unauthenticated + * request path at all. + */ +export function getDaemonTokenPath(dir: string = getDaemonDir()): string { + return path.join(dir, "mcpdod.token"); +} + +/** Daemon stderr log (0600) — the only visibility into a detached daemon + * that died during startup. */ +export function getDaemonLogPath(dir: string = getDaemonDir()): string { + return path.join(dir, "mcpdod.log"); +} + +/** + * `sun_path` limit for Unix sockets: 104 bytes on macOS/BSD, 108 on Linux + * (both including the trailing NUL). `listen()` fails opaquely above it — + * historically the daemon then died silently and the client reported only a + * generic start timeout. Validate up front with an actionable error instead. + */ +export function assertSocketPathWithinLimit(socketPath: string): void { + // Windows named pipes are not sun_path-constrained. + if (process.platform === "win32") return; + /* v8 ignore next -- one arm per platform; CI runs each on its own OS */ + const limit = process.platform === "linux" ? 107 : 103; + const bytes = Buffer.byteLength(socketPath); + if (bytes > limit) { + throw new Error( + `Connection daemon socket path is too long for this platform ` + + `(${bytes} bytes > ${limit}): ${socketPath}. ` + + `Point MCP_INSPECTOR_DAEMON_DIR (or MCP_STORAGE_DIR) at a shorter directory.`, + ); + } +} + +/** Ensure the daemon directory exists before binding the socket. + * Created 0700: the socket lives inside, so its own mode never has to be + * the enforcement boundary (BSDs are inconsistent about socket modes). + * `mkdirSync` never changes the mode of a pre-existing directory — and the + * default `~/.mcp-inspector` commonly already exists at 0755 from other + * inspector components — so an existing directory is re-validated and + * tightened with the same symlink/ownership/mode checks as the private + * tmp root. */ +export function ensureDaemonDir(dir: string = getDaemonDir()): void { + fs.mkdirSync(dir, { recursive: true, mode: 0o700 }); + assertTrustedPrivateRoot(dir); +} diff --git a/clients/mcpdo/src/daemon/protocol.ts b/clients/mcpdo/src/daemon/protocol.ts new file mode 100644 index 0000000000..f3f90d0228 --- /dev/null +++ b/clients/mcpdo/src/daemon/protocol.ts @@ -0,0 +1,362 @@ +import type { + InspectorServerSettings, + MCPServerConfig, + PendingRequestOrigin, +} from "@inspector/core/mcp/types.js"; +import type { + CliAppInfo, + MethodArgs, +} from "@inspector/core/cli/handlers/method-types.js"; +import type { + Implementation, + ProtocolEra, + ServerCapabilities, +} from "@modelcontextprotocol/client"; + +/** Operations the connection daemon accepts over IPC. */ +export type DaemonOp = + | "ping" + | "connect" + | "disconnect" + | "connections/list" + | "connections/use" + | "connections/show" + | "daemon/status" + | "daemon/stop" + | "rpc" + | "stream" + | "elicitation/respond"; + +export type ConnectParams = { + name: string; + serverConfig: MCPServerConfig; + serverSettings?: InspectorServerSettings; + /** Human-readable server identity for `connections/list`. */ + serverIdentity: string; + /** + * When true and the dial fails with `auth_required`, register the + * connection anyway as a dormant intent entry (never-connected client, + * terminal status) and return `ConnectionInfo` with `pendingAuth: true` + * instead of throwing. The front-end sets this on the non-TTY connect path + * after handing interactive OAuth to the detached auth helper: once the + * user finishes signing in, the next op on this connection revives it with + * the freshly stored credentials — no second `connect` required. + */ + pendingOnAuthRequired?: boolean; +}; + +export type ConnectionNameParams = { + /** Omit to target the MRU connection (TTY). */ + name?: string; + /** + * When true (non-TTY / CI), omit is an error — require an explicit connection. + * Front-end sets this from `!process.stdin.isTTY` (not stdout — keying off + * stdin lets piping output, e.g. `mcpdo tools/list | jq`, still use MRU when + * a human is at the keyboard) unless opted out via + * `MCP_ALLOW_DEFAULT_CONNECTION=1`. + */ + requireExplicit?: boolean; +}; + +/** Params for `rpc` / `stream` — connection targeting plus method args. */ +export type RpcParams = ConnectionNameParams & + MethodArgs & { + method: string; + /** + * Whether a human is driving this call: set by the front-end to + * `format === "text" && (stdin.isTTY || stderr.isTTY)` (see dispatch.ts). + * One honest fact — "is there a human at the stream?" — drives two + * daemon behaviours, so the daemon reads interactivity here rather than + * inferring it from a side effect: + * + * - **Elicitation routing.** The daemon parks a legacy/modern non-task + * MRTR elicitation — instead of relaying it for an inline prompt — only + * when this field is *explicitly* `false` (`park = interactive === false`): + * it returns immediately with `kind: "elicitation-pending"` and the + * caller answers via the `elicitation/respond` op. Parking is opt-in, so + * an absent field does NOT park (a caller that never announced itself + * must not have its elicitations silently parked and hang). A + * non-interactive caller (`--format json` or non-TTY) has no human at the + * stream to answer an inline prompt. + * - **Agent-vs-human message text.** Auth-pending errors use agent-tuned + * wording (relay instructions) unless this field is *explicitly* `true`. + * See `ConnectionRegistry.reviveLocked`. + * + * The two reads differ in the absent case on purpose: parking defaults OFF + * (absent ⇒ don't park), message audience defaults to agent-facing + * (absent ⇒ non-interactive). The front-end always sends it, so absent is + * only ever a daemon-internal caller. + */ + interactive?: boolean; + }; + +export type DaemonRequest = { + id: string; + op: DaemonOp; + /** + * IPC auth token. Every daemon requires one: private mode passes it via + * `MCP_INSPECTOR_DAEMON_TOKEN`, and the shared default daemon generates + * one at startup and publishes it to `mcpdod.token` for clients to read. + * Optional only at the wire/type boundary so a request missing the token + * can still be parsed — and then rejected — rather than failing framing. + */ + token?: string; + params?: + | ConnectParams + | ConnectionNameParams + | RpcParams + | ElicitationRespondParams + | Record<string, never>; +}; + +export type DaemonErrorBody = { + code: string; + message: string; + /** Suggested CLI exit code when applicable. */ + exitCode?: number; +}; + +export type DaemonResponse = + | { id: string; ok: true; result: unknown } + | { id: string; ok: false; error: DaemonErrorBody }; + +/** Frames after the initial ok response on a `stream` connection. */ +export type DaemonStreamFrame = + | { id: string; stream: "data"; data: unknown } + | { id: string; stream: "end" }; + +/** + * One elicitation request/answer exchange, carried mid-`rpc` call when the + * in-flight tool/prompt/resource call surfaces a legacy or modern non-task + * MRTR elicitation (dual-era support, phase 1 — task-augmented MRTR + * elicitation is a separate follow-up, since that call already returns + * immediately and never blocks a `rpc` round-trip in the first place). + * + * Written by the daemon onto the SAME connection as the originating `rpc` + * request, before its `DaemonResponse`; the CLI answers on that same + * connection with an {@link ElicitationResponseFrame}, and the daemon resumes + * the (still in-flight) call. See `ipc-glue.ts`'s `acceptDaemonConnection` for + * why this needs no new channel: each `rpc` request already owns its + * connection exclusively, and core itself never has more than one elicitation + * pending at a time (sequential by design) — though a single call can + * pause/resume through several of these exchanges before its final response. + */ +export type ElicitationRequestFrame = { + id: string; + kind: "elicitation-request"; + /** `ElicitationCreateMessage.id` — echoed back so the answer can be matched. */ + elicitationId: string; + mode: "form" | "url"; + message: string; + /** Form mode only. */ + requestedSchema?: Record<string, unknown>; + /** URL mode only. */ + url?: string; + /** Legacy server→client request vs. modern non-task MRTR round. */ + origin: PendingRequestOrigin; +}; + +export type ElicitationResponseFrame = { + id: string; + kind: "elicitation-response"; + elicitationId: string; + action: "accept" | "decline" | "cancel"; + /** Form mode `action: "accept"` only. */ + content?: Record<string, unknown>; +}; + +/** + * A parked elicitation, as reported to a non-interactive caller + * ({@link RpcParams.interactive} false): everything an agent needs to relay + * the request to a human and answer it with `elicitation/respond`. Rides the + * normal success payload — like the pending-auth URL, an elicitation URL's + * query string is meaningful data the error envelope would redact. + */ +export type ElicitationPendingInfo = { + /** Key for `elicitation/respond`; changes on every round. */ + elicitationId: string; + /** Connection whose in-flight call is parked. */ + connection: string; + /** Originating rpc method (e.g. `tools/call`), for output rendering. */ + method: string; + /** Originating tool, when the method was `tools/call`. */ + toolName?: string; + mode: "form" | "url"; + message: string; + /** Form mode only. */ + requestedSchema?: Record<string, unknown>; + /** URL mode only. */ + url?: string; + /** Legacy server→client request vs. modern non-task MRTR round. */ + origin: PendingRequestOrigin; + /** Epoch ms; the exchange is auto-cancelled if unanswered by then. */ + expiresAt: number; +}; + +/** Params for the `elicitation/respond` op. */ +export type ElicitationRespondParams = { + elicitationId: string; + action: "accept" | "decline" | "cancel"; + /** Form mode `action: "accept"` only. */ + content?: Record<string, unknown>; +}; + +/** + * `elicitation/respond` result. `outcome` is either the parked call's final + * result — the response the original `rpc` would have produced — or the next + * `elicitation-pending` round; `method`/`toolName` echo the originating call + * so the front-end can render that result the same way `rpc` output is. + */ +export type ElicitationRespondResult = { + method: string; + toolName?: string; + outcome: RpcResult; +}; + +/** + * Slim connect-time snapshot of a connection's authorization, projected from the + * core `OAuthConnectionState` (see {@link ConnectionInfo.auth}). Absent entirely + * for stdio servers and HTTP servers that never engaged OAuth — cleaner than + * reporting "none" for every local server. + */ +export type ConnectionAuthInfo = { + method: "oauth" | "ema"; + /** Whether tokens for this server are present in storage. */ + authorized: boolean; + /** Granted scope (from the token response), when known. */ + scope?: string; + /** OAuth client id used with the authorization server, when known. */ + clientId?: string; + /** EMA only: IdP session state at connect time. */ + idpSession?: "none" | "logged_in" | "expired"; +}; + +export type ConnectionInfo = { + name: string; + serverIdentity: string; + connectedAt: number; + lastAccessedAt: number; + isMru: boolean; + /** + * Negotiated era for this connection's connection — legacy `initialize` vs. + * modern `server/discover` (#2298 follow-up). Present everywhere a live + * connection is reported (`connect`, `connections/list`, `connections/use`), not + * just `connections/show`, so a user with several open connections can see which + * era each negotiated without querying them one at a time. Absent only if + * the client hasn't connected (never observed in practice — every code + * path constructing a `ConnectionInfo` does so from an already-connected + * connection). + */ + protocolEra?: ProtocolEra; + /** + * Authorization snapshot. Like `protocolEra`, present everywhere a live + * connection is reported so both humans and agents can see *how* a connection is + * authenticated (OAuth vs. EMA, authorized or not) without a separate + * query. Freshness varies by op: `connect` computes it right after the + * connection succeeds; `connections/list` and `connections/use` reuse that + * connect-time value; `connections/show` recomputes it live *from disk* so it + * reflects the current persisted state (e.g. after `auth/clear` or + * `auth/ema-logout`, even from another process). Note a live connection may + * keep working on its in-memory tokens after storage was cleared — `show` + * reports the persisted state, matching `auth/ema-status`. + */ + auth?: ConnectionAuthInfo; + /** + * True when this entry was registered as auth-pending intent + * ({@link ConnectParams.pendingOnAuthRequired}): the dial hit + * `auth_required` and interactive sign-in is completing out of band in the + * detached auth helper. The entry holds a never-connected client, so the + * first op after tokens land revives (dials) it transparently. Reported by + * `connect` and echoed by `connections/list`/`connections/show` until a + * revive succeeds; `connections/show` itself revives once it sees the + * signed-in tokens on disk, so polling it observes the completion. + */ + pendingAuth?: boolean; + /** + * Only alongside `pendingAuth: true`: the out-of-band sign-in has already + * stored usable tokens on disk, so the connection completes on its next + * use (or on the next `connections/show`, which revives it). Set by the + * read-only echoes (`connections/list`, `connections/use`, `daemon/status`) + * from a disk check — no dial. A poller seeing this can stop waiting on the + * user and proceed to its next command. + */ + pendingAuthSignedIn?: boolean; +}; + +/** + * `connections/show` result: daemon bookkeeping ({@link ConnectionInfo}, which as of + * #2298 already carries `protocolEra`) plus the live MCP connection state — + * era-agnostic (`serverInfo`/`capabilities`/`instructions`/`protocolVersion` + * are populated the same way whether they came from a legacy `initialize` + * response or a modern `server/discover`) and era-specific (`supportedVersions`, + * only set when the connect actually probed `server/discover`, i.e. + * `auto`/`modern`). + */ +export type ConnectionShowResult = ConnectionInfo & { + serverInfo?: Implementation; + protocolVersion?: string; + capabilities?: ServerCapabilities; + instructions?: string; + supportedVersions?: string[]; + /** + * Live transport state, `connections/show` only. The connection itself is + * user intent ("connected until I disconnect"); the transport under it is + * disposable and self-healing. `"live"` = the client session is up; + * `"connecting"` = mid-dial; `"dormant"` = the transport dropped (server + * expired the session, SSE stream closed, stdio child exited) and the next + * op will transparently re-dial with stored credentials. Debug detail, not + * something the user must act on. + */ + transport?: "live" | "connecting" | "dormant"; + /** + * Only alongside `pendingAuth: true`: the authorize URL the human must open + * to complete an out-of-band sign-in, present while the detached auth + * helper is still live (unexpired, PID alive — see + * `readLivePendingAuthMarker`). Rides the result payload, not the error + * envelope, because the envelope redacts URL query strings and this URL IS + * its query (client_id/PKCE/state). A poller reads it here and relays it; + * the `auth_required` error on other ops points here rather than carrying a + * (redacted, useless) copy. + */ + authUrl?: string; +}; + +export type DaemonStatus = { + pid: number; + socketPath: string; + connections: ConnectionInfo[]; + idleMs: number | null; + /** True once shutdown has begun (status stays answerable while stopping). */ + stopping: boolean; +}; + +/** Serializable RPC outcome (no live stream callbacks). */ +export type RpcResult = + | { + kind: "result"; + result: Record<string, unknown>; + appInfo?: CliAppInfo; + } + | { + kind: "ndjson"; + lines: unknown[]; + /** + * `skills/list --verify` / `skills/get --verify` one-line stderr + * verdict (#2248). Carried across the daemon socket so the connection CLI + * can report the same summary the one-shot CLI does, rather than + * silently dropping it the way an earlier pass through this file did. + */ + summary?: string; + /** Non-zero when the emitted report is itself a failure (`--verify`). */ + exitCode?: number; + } + | { + /** + * The call surfaced an elicitation while + * {@link RpcParams.interactive} was false: the call is parked + * daemon-side awaiting `elicitation/respond`, and this is everything + * the caller needs to answer it. + */ + kind: "elicitation-pending"; + elicitation: ElicitationPendingInfo; + }; diff --git a/clients/mcpdo/src/daemon/run.ts b/clients/mcpdo/src/daemon/run.ts new file mode 100644 index 0000000000..c34273d2a2 --- /dev/null +++ b/clients/mcpdo/src/daemon/run.ts @@ -0,0 +1,54 @@ +#!/usr/bin/env node +/** + * Connection daemon entrypoint. Spawned detached by {@link ensureDaemon}. + * Optional foreground `mcpdo daemon run` is not shipped yet (see v2_cli_v2.md). + */ +import { DaemonServer } from "./server.js"; +import { generateDaemonToken, getDaemonTokenFromEnv } from "./auth.js"; +import { ensureDaemonDir } from "./paths.js"; +import { disallowMemorySecretStoreFallback } from "@inspector/core/auth/node/secret-store-selection.js"; +import { awaitableError } from "@inspector/core/cli/utils/awaitable-log.js"; + +// Name the process `mcpdod` (Unix d-suffix convention) so `ps`/`pgrep`/`pkill` +// see the daemon under a greppable name instead of a bare `node .../mcpdod.js`. +process.title = "mcpdod"; + +// Same policy as mcp-bin.ts: mcpdo is multi-process, so the keychain-less +// automatic fallback must be the (shared) secrets file, never memory. +disallowMemorySecretStoreFallback(); + +async function main(): Promise<void> { + const server = new DaemonServer({ + // No tokenless daemons: when the spawner didn't hand one down via + // MCP_INSPECTOR_DAEMON_TOKEN, generate one. start() publishes it to + // mcpdod.token (0600) for clients to read. + requiredToken: getDaemonTokenFromEnv() ?? generateDaemonToken(), + onShutdown: () => { + // Allow natural exit once the server closes and idle work finishes. + process.exitCode = 0; + }, + }); + + // Never keep the cwd of whichever mcpdo invocation happened to spawn this + // daemon: connects would resolve relative stdio paths against it (and pin + // the directory against unmounting). The front end always sends an + // explicit cwd for stdio servers, so the daemon's own cwd is inert. + ensureDaemonDir(server.dir); + process.chdir(server.dir); + + const shutdown = () => { + void server.stop("signal").then(() => process.exit(0)); + }; + process.on("SIGINT", shutdown); + process.on("SIGTERM", shutdown); + + await server.start(); +} + +main().catch(async (error: unknown) => { + const message = error instanceof Error ? error.message : String(error); + // Exit only once the write has been performed: on a pipe or file stderr is + // asynchronous, and process.exit() would discard the diagnostic (#2638). + await awaitableError(`mcpdo daemon: ${message}\n`); + process.exit(1); +}); diff --git a/clients/mcpdo/src/daemon/server.ts b/clients/mcpdo/src/daemon/server.ts new file mode 100644 index 0000000000..7b45e83709 --- /dev/null +++ b/clients/mcpdo/src/daemon/server.ts @@ -0,0 +1,1149 @@ +import * as fs from "node:fs"; +import * as net from "node:net"; +import { + classifyError, + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { runMethod } from "@inspector/core/cli/handlers/run-method.js"; +import type { MethodArgs } from "@inspector/core/cli/handlers/method-types.js"; +import { + acceptDaemonConnection, + removeStaleDaemonSocket, + type ElicitationChannel, + type HandleOutcome, +} from "./ipc-glue.js"; +import { wireElicitationBridge } from "./elicitation-bridge.js"; +import { + ElicitationParkRegistry, + ParkingElicitationChannel, +} from "./elicitation-park.js"; +import { assertDaemonToken, getDaemonTokenFromEnv } from "./auth.js"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import { isTerminalStatus } from "@inspector/core/mcp/types.js"; +import type { InspectorClientEventMap } from "@inspector/core/mcp/inspectorClientEventTarget.js"; +import type { TypedEventGeneric } from "@inspector/core/mcp/typedEventTarget.js"; +import { + assertSocketPathWithinLimit, + ensureDaemonDir, + getDaemonDir, + getDaemonLockPath, + getDaemonSocketPath, + getDaemonTokenPath, +} from "./paths.js"; +import type { + ConnectParams, + DaemonRequest, + DaemonResponse, + DaemonStatus, + ElicitationPendingInfo, + ElicitationRequestFrame, + ElicitationRespondParams, + ElicitationRespondResult, + RpcParams, + RpcResult, + ConnectionNameParams, + ConnectionShowResult, +} from "./protocol.js"; +import { + DEFAULT_IDLE_MS, + getLiveConnectionAuthInfo, + ConnectionRegistry, +} from "./connections.js"; + +/** + * Default channel used when a caller doesn't wire a real one (in-process + * `handle`/`handleOutcome` test call sites that predate elicitation support). + * Immediately cancels any elicitation, matching `elicit: false` behavior — + * these callers never advertise elicitation support to the server anyway. + */ +const autoCancelElicitationChannel: ElicitationChannel = { + request(frame) { + return Promise.resolve({ + id: frame.id, + kind: "elicitation-response", + elicitationId: frame.elicitationId, + action: "cancel", + }); + }, +}; + +export type DaemonServerOptions = { + dir?: string; + idleMs?: number; + /** + * When set, every IPC request must present this token. Defaults to + * `MCP_INSPECTOR_DAEMON_TOKEN` from the environment (private mode). + */ + requiredToken?: string; + /** Called when the daemon should exit (idle timeout or daemon/stop). */ + onShutdown?: () => void; + /** + * Grace period at shutdown for flushing buffered response bytes before + * still-open sockets are force-destroyed (a client that stopped reading + * must not hang `daemon stop`). Tests use a short value. + */ + flushTimeoutMs?: number; + /** + * TTL for parked elicitations (parked when `RpcParams.interactive` is + * false); defaults to {@link PARKED_ELICITATION_TTL_MS}. Tests use a short + * value. + */ + elicitationTtlMs?: number; +}; + +/** + * Unix-socket NDJSON daemon that owns {@link ConnectionRegistry}. + */ +export class DaemonServer { + readonly registry: ConnectionRegistry; + readonly socketPath: string; + readonly lockPath: string; + readonly dir: string; + private readonly requiredToken: string | undefined; + private server: net.Server | null = null; + private readonly onShutdown: (() => void) | null; + private stopping = false; + /** In-flight stop, memoized so a repeated stop (e.g. a second SIGINT) + * awaits the original cleanup instead of resolving immediately and letting + * its caller `process.exit()` mid-teardown, stranding the socket, token, + * and lock on disk. */ + private stopPromise: Promise<void> | null = null; + /** In-flight handleOutcome calls; shutdown quiesces these before the + * registry snapshot so a concurrent connect cannot register a live client + * after disconnectAll and leak it. */ + private activeOps = 0; + private opsIdleResolvers: (() => void)[] = []; + /** Accepted IPC sockets. Long-lived stream sockets never end on their own, + * so shutdown flushes and destroys them — otherwise server.close() would + * wait forever. */ + private readonly ipcSockets = new Set<net.Socket>(); + /** Serializes `rpc` ops per client. Core cannot attribute an elicitation + * to a specific in-flight call, so with concurrent RPCs on one connection + * the bridge would route a prompt to the wrong caller's terminal; running + * at most one rpc per connection at a time makes the routing exact. */ + private readonly rpcQueues = new WeakMap<InspectorClient, Promise<void>>(); + /** Parked elicitations for non-interactive callers (see elicitation-park.ts). */ + private readonly parks: ElicitationParkRegistry; + + constructor(options: DaemonServerOptions = {}) { + this.dir = options.dir ?? getDaemonDir(); + this.socketPath = getDaemonSocketPath(this.dir); + this.lockPath = getDaemonLockPath(this.dir); + this.requiredToken = options.requiredToken ?? getDaemonTokenFromEnv(); + this.flushTimeoutMs = + options.flushTimeoutMs ?? DaemonServer.FLUSH_TIMEOUT_MS; + this.registry = new ConnectionRegistry(options.idleMs ?? DEFAULT_IDLE_MS); + this.parks = new ElicitationParkRegistry(options.elicitationTtlMs); + this.onShutdown = options.onShutdown ?? null; + this.registry.setIdleHandler(() => { + void this.stop("idle"); + }); + } + + async start(): Promise<void> { + ensureDaemonDir(this.dir); + assertSocketPathWithinLimit(this.socketPath); + this.acquireLock(); + try { + await removeStaleDaemonSocket(this.socketPath); + + // Publish the IPC token (0600, inside the 0700 daemon dir) before the + // socket exists, so a client can never connect without being able to + // read the token it needs. See getDaemonTokenPath. + if (this.requiredToken !== undefined) { + const tokenPath = getDaemonTokenPath(this.dir); + // Exclusive create after removing any existing entry: writeFileSync + // follows symlinks, and ensureDaemonDir's tightening of the parent + // does not remove children planted while the dir was writable — a + // planted symlink would leak the token into an attacker-readable + // file. rmSync removes a symlink itself, never its target. + fs.rmSync(tokenPath, { force: true }); + fs.writeFileSync(tokenPath, this.requiredToken + "\n", { + mode: 0o600, + flag: "wx", + }); + try { + fs.chmodSync(tokenPath, 0o600); + } catch { + // Unsupported on some platforms. + } + } + + this.server = net.createServer((socket) => { + this.ipcSockets.add(socket); + socket.once("close", () => this.ipcSockets.delete(socket)); + acceptDaemonConnection(socket, (req, elicitation, signal) => + this.handleOutcome(req, elicitation, signal), + ); + }); + + await new Promise<void>((resolve, reject) => { + this.server!.once("error", reject); + this.server!.listen(this.socketPath, () => { + this.server!.off("error", reject); + resolve(); + }); + }); + + // Restrict socket + lock to the creating user. Private mode also requires + // an IPC token (see specification/v2_cli_v2.md §5.3). + try { + fs.chmodSync(this.socketPath, 0o600); + fs.chmodSync(this.lockPath, 0o600); + } catch { + // Unsupported on some platforms (e.g. Windows named pipes). + } + + // Connection-less spawn (e.g. ensureDaemon from tools/list with no connections) + // must still self-reap — idle was previously only armed after disconnect. + this.registry.armIdleTimerIfEmpty(); + } catch (error) { + // Never leave a lock we own but no daemon behind it. The socket is only + // unlinked by removeStaleDaemonSocket after a dead connect probe, so a + // live daemon's socket is never touched here. + this.releaseLock(); + throw error; + } + } + + stop(reason: "idle" | "stop" | "signal" = "stop"): Promise<void> { + this.stopPromise ??= this.doStop(reason); + return this.stopPromise; + } + + private async doStop(reason: "idle" | "stop" | "signal"): Promise<void> { + void reason; + this.stopping = true; + // Settle parked elicitations first: their held server requests must be + // cancelled before the connections under them are torn down. + this.parks.cancelAll(); + // Quiesce: new ops are rejected above; wait (bounded — an rpc blocked on + // an interactive elicitation prompt must not hang shutdown forever) for + // in-flight ops so a concurrent connect lands in the registry before the + // disconnect snapshot below. + await this.waitForActiveOps(DaemonServer.QUIESCE_TIMEOUT_MS); + await this.registry.disconnectAll(); + // Flush pending response writes (e.g. daemon/stop's own {stopping:true}) + // then drop the sockets: long-lived stream sockets never end on their + // own and would keep server.close() waiting forever. + for (const socket of [...this.ipcSockets]) { + socket.destroySoon(); + } + await new Promise<void>((resolve) => { + if (!this.server) { + resolve(); + return; + } + // destroySoon() only destroys once queued bytes drain — a client that + // stopped reading with a buffered response would keep server.close() + // waiting forever. Bounded grace for the flush, then force-destroy + // whatever is left. + const force = setTimeout(() => { + for (const socket of [...this.ipcSockets]) { + socket.destroy(); + } + }, this.flushTimeoutMs); + force.unref(); + this.server.close(() => { + clearTimeout(force); + resolve(); + }); + }); + this.server = null; + this.removeLockAndSocket(); + this.onShutdown?.(); + } + + /** Grace period for in-flight ops during shutdown before teardown proceeds + * anyway. Exported for tests. */ + static readonly QUIESCE_TIMEOUT_MS = 3_000; + + /** Default shutdown flush grace before force-destroying sockets. */ + static readonly FLUSH_TIMEOUT_MS = 2_000; + private readonly flushTimeoutMs: number; + + private waitForActiveOps(timeoutMs: number): Promise<void> { + if (this.activeOps === 0) return Promise.resolve(); + return new Promise((resolve) => { + const timer = setTimeout(resolve, timeoutMs); + timer.unref?.(); + this.opsIdleResolvers.push(() => { + clearTimeout(timer); + resolve(); + }); + }); + } + + status(): DaemonStatus { + return { + pid: process.pid, + socketPath: this.socketPath, + connections: this.registry.list(), + idleMs: this.registry.idleRemainingMs(), + stopping: this.stopping, + }; + } + + /** Handle one request; returns the response body (used by in-process tests). */ + async handle( + request: DaemonRequest, + elicitation: ElicitationChannel = autoCancelElicitationChannel, + signal?: AbortSignal, + ): Promise<DaemonResponse> { + return (await this.handleOutcome(request, elicitation, signal)).response; + } + + /** Full handle including optional stream starter (socket accept path). */ + async handleOutcome( + request: DaemonRequest, + elicitation: ElicitationChannel = autoCancelElicitationChannel, + signal?: AbortSignal, + ): Promise<HandleOutcome> { + try { + assertDaemonToken(this.requiredToken, request.token); + this.activeOps++; + try { + return await this.dispatch(request, elicitation, signal); + } finally { + this.activeOps--; + if (this.activeOps === 0) { + for (const resolve of this.opsIdleResolvers.splice(0)) resolve(); + } + } + } catch (error) { + if (error instanceof CliExitCodeError) { + return { + response: { + id: request.id, + ok: false, + error: { + code: error.envelope?.code ?? "cli_error", + message: error.message, + exitCode: error.exitCode, + }, + }, + }; + } + // Match one-shot CLI exit codes (e.g. unreachable → 4, not always 1). + const { exitCode, envelope } = classifyError(error); + return { + response: { + id: request.id, + ok: false, + error: { + code: envelope.code, + message: envelope.message, + exitCode, + }, + }, + }; + } + } + + private async dispatch( + request: DaemonRequest, + elicitation: ElicitationChannel, + signal?: AbortSignal, + ): Promise<HandleOutcome> { + // Once shutdown starts, new work is rejected: an op accepted here could + // otherwise register a live client after disconnectAll's snapshot. + // Status-style ops stay answerable; a repeated daemon/stop joins the + // in-flight stop via the memoized promise. + if ( + this.stopping && + request.op !== "ping" && + request.op !== "daemon/status" && + request.op !== "daemon/stop" + ) { + throw new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "Connection daemon is shutting down.", + { code: "daemon_stopping" }, + ); + } + switch (request.op) { + case "ping": + return { + response: { + id: request.id, + ok: true, + // `stopping` lets ensureDaemon treat a shutting-down daemon as + // "about to be gone" (wait for exit, respawn) instead of alive — + // ping itself always succeeds so status checks never fail. + result: { pong: true, pid: process.pid, stopping: this.stopping }, + }, + }; + case "connect": { + const params = request.params as ConnectParams; + if (!params?.name || !params.serverConfig || !params.serverIdentity) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "connect requires name, serverConfig, and serverIdentity", + { code: "invalid_params" }, + ); + } + // A replacing connect tears down any previous connection under this + // name; a call parked on it can never be answered — settle it now. + this.parks.cancelForConnection(params.name); + return { + response: { + id: request.id, + ok: true, + result: await this.registry.connect(params, signal), + }, + }; + } + case "disconnect": { + const params = (request.params ?? {}) as ConnectionNameParams; + const result = await this.registry.disconnect( + params.name, + params.requireExplicit, + ); + // The connection is gone; a call parked on it can never be answered. + this.parks.cancelForConnection(result.name); + return { + response: { id: request.id, ok: true, result }, + }; + } + case "connections/list": + return { + response: { + id: request.id, + ok: true, + result: { + connections: await this.registry.annotateAuthProgress( + this.registry.list(), + ), + }, + }, + }; + case "connections/use": { + const params = (request.params ?? {}) as ConnectionNameParams; + if (!params.name) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "connections/use requires a connection name", + { code: "invalid_params" }, + ); + } + return { + response: { + id: request.id, + ok: true, + result: ( + await this.registry.annotateAuthProgress([ + this.registry.use(params.name), + ]) + )[0], + }, + }; + } + case "connections/show": { + const params = (request.params ?? {}) as ConnectionNameParams; + let connection = this.registry.connectionFor( + params.name, + params.requireExplicit, + ); + // Recomputed live from disk (not the connect-time cache and not the + // client's memory-cached storage): `show` reports the *current* + // persisted auth state, so an auth/clear, auth/ema-logout, or a + // web-client re-auth since connect is reflected here. + let auth = await getLiveConnectionAuthInfo(connection); + if (connection.pendingAuth === true && auth?.authorized === true) { + // The out-of-band sign-in completed (tokens are on disk) but no op + // has revived the entry yet. `connect` tells callers to poll here + // ("completes automatically after sign-in — check with + // `connections/show`"), so make that true: run the same revive the + // first op would, instead of reporting "pending" forever. + try { + await this.registry.liveClientFor( + params.name, + params.requireExplicit, + ); + } catch { + // Revive failed (server unreachable, tokens rejected mid-flight, + // raced disconnect). Keep show read-only-honest: fall through to + // the snapshot — the entry stays pending and a later op retries. + } + // Re-resolve: a successful revive replaced the client and cleared + // pendingAuth; a raced disconnect removed the entry (thrown here + // as the usual unknown-connection error). + connection = this.registry.connectionFor( + params.name, + params.requireExplicit, + ); + auth = await getLiveConnectionAuthInfo(connection); + } + const client = connection.client; + const result: ConnectionShowResult = { + name: connection.name, + serverIdentity: connection.serverIdentity, + connectedAt: connection.connectedAt, + lastAccessedAt: connection.lastAccessedAt, + isMru: true, + serverInfo: client.getServerInfo(), + protocolVersion: client.getProtocolVersion(), + protocolEra: client.getProtocolEra(), + ...(auth && { auth }), + ...(connection.pendingAuth && { pendingAuth: true }), + // While sign-in is still in flight, surface the authorize URL so a + // poller can relay it. Lives on the result (un-redacted), not the + // error envelope (which would strip its query string). Absent once + // the helper completes or dies — `pendingAuthUrlFor` checks liveness. + ...(connection.pendingAuth && + (() => { + const authUrl = this.registry.pendingAuthUrlFor(connection); + return authUrl ? { authUrl } : {}; + })()), + capabilities: client.getCapabilities(), + instructions: client.getInstructions(), + supportedVersions: client.getDiscoverResult()?.supportedVersions, + transport: isTerminalStatus(client.getStatus()) + ? "dormant" + : client.getStatus() === "connecting" + ? "connecting" + : "live", + }; + return { + response: { id: request.id, ok: true, result }, + }; + } + case "daemon/status": { + const status = this.status(); + status.connections = await this.registry.annotateAuthProgress( + status.connections, + ); + return { + response: { id: request.id, ok: true, result: status }, + }; + } + case "daemon/stop": + queueMicrotask(() => { + void this.stop("stop"); + }); + return { + response: { id: request.id, ok: true, result: { stopping: true } }, + }; + case "rpc": + return { + response: { + id: request.id, + ok: true, + result: await this.runRpc( + request.id, + request.params as RpcParams, + elicitation, + signal, + ), + }, + }; + case "stream": + return this.openStream(request.id, request.params as RpcParams); + case "elicitation/respond": + return { + response: { + id: request.id, + ok: true, + result: await this.respondElicitation( + request.params as ElicitationRespondParams, + ), + }, + }; + default: + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `Unknown daemon op: ${(request as DaemonRequest).op}`, + { code: "unknown_op" }, + ); + } + } + + private async runRpc( + requestId: string, + params: RpcParams, + elicitation: ElicitationChannel, + signal?: AbortSignal, + ): Promise<RpcResult> { + if (!params?.method) { + throw new CliExitCodeError(EXIT_CODES.USAGE, "rpc requires a method", { + code: "invalid_params", + }); + } + // Two independent reads of `interactive`, and they differ in the ABSENT + // case on purpose: + // - Parking is opt-in. Park only when the caller *explicitly* says it is + // non-interactive (`interactive === false`); an absent field must NOT + // park (matching the pre-rename `parkElicitations`-absent default), or + // a caller that never announced itself would have its elicitations + // silently parked and hang. + // - Message audience defaults the other way: absent ⇒ agent-facing. The + // raw `params.interactive` (boolean | undefined) is threaded to + // `liveClientFor`, and `authRequiredError` treats only `=== true` as a + // human — so absent falls to the agent-tuned wording. + const park = params.interactive === false; + // Parking needs the connection *name* for the pending payload and for + // teardown-keyed cancellation; resolve it before reviving the client. + const connectionName = park + ? this.registry.connectionFor(params.name, params.requireExplicit).name + : undefined; + const client = await this.registry.liveClientFor( + params.name, + params.requireExplicit, + params.interactive, + ); + const previous = this.rpcQueues.get(client) ?? Promise.resolve(); + const run = previous.then(() => { + // The caller hung up while this call sat queued behind another op — + // don't run work on behalf of a socket that can't receive the result. + if (signal?.aborted) { + throw new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "The caller disconnected before this command could run.", + { code: "caller_gone" }, + ); + } + return park + ? this.runRpcParked(client, connectionName!, requestId, params) + : this.runRpcOnClient(client, requestId, params, elicitation, signal); + }); + // Keep the queue alive past failures; each caller still sees its own + // error through `run`. + this.rpcQueues.set( + client, + run.then( + () => undefined, + () => undefined, + ), + ); + return run; + } + + private async runRpcOnClient( + client: InspectorClient, + requestId: string, + params: RpcParams, + elicitation: ElicitationChannel, + signal?: AbortSignal, + ): Promise<RpcResult> { + const methodArgs = stripConnectionFields(params); + this.assertNoParkedCall(client); + // Backstop against the silent-empty class: `runMethod`'s list states + // return `[]` without error when the client isn't connected, which would + // render as "Tools (0)" for a connection that actually dropped. The + // resolve above revived a dead client, but a drop can still land between + // that and here (e.g. while queued behind a long op) — fail honestly and + // let a retry revive it. + if (isTerminalStatus(client.getStatus())) { + throw new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "The connection dropped before this command could run; re-run the command to reconnect.", + { code: "connection_stale" }, + ); + } + const unwire = wireElicitationBridge(client, elicitation, requestId); + // When the caller's socket closes mid-call, cancel the in-flight request so + // the per-client rpc queue isn't wedged behind work nobody is waiting for. + // Two mechanisms, because a tool call and a plain read cancel differently: + // - `cancelToolCall()` runs the MCP cancellation flow for a tool call (it + // sends `notifications/cancelled` with a reason) — a no-op otherwise. + // - the ambient request signal (#1783) covers *every other* method. A + // `resources/read`, a `prompts/get`, any `*/list` the server never + // answers would otherwise hold the queue slot forever, wedging the whole + // connection for all later commands until a reconnect. Making the + // caller's signal the client's ambient request signal threads it into + // that request's options, so a caller disconnect aborts it and the slot + // frees. + const onAbort = () => { + client.cancelToolCall(); + }; + signal?.addEventListener("abort", onAbort, { once: true }); + const clearAmbientSignal = client.setAmbientRequestSignal(signal); + let outcome; + try { + outcome = await runMethod(client, methodArgs); + } finally { + clearAmbientSignal(); + signal?.removeEventListener("abort", onAbort); + unwire(); + } + return toRpcResult(outcome, params.method); + } + + /** + * A connection with a parked call must not accept new rpcs: the bridge + * routes elicitations to its oldest subscriber, so a second in-flight call + * would have its elicitations misdelivered to the parked exchange. + */ + private assertNoParkedCall(client: InspectorClient): void { + const parked = this.parks.forClient(client); + if (parked) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `A server elicitation is pending on this connection; answer it first: elicitation/respond ${parked.info.elicitationId} (or --cancel).`, + { code: "elicitation_pending" }, + ); + } + } + + /** + * `rpc` from a non-interactive caller (`RpcParams.interactive` false): run + * the call racing its completion against the first elicitation. Completion + * first → ordinary result. Elicitation first → park the still-running call + * and return `elicitation-pending`; `elicitation/respond` picks it up from + * there. + */ + private async runRpcParked( + client: InspectorClient, + connectionName: string, + requestId: string, + params: RpcParams, + ): Promise<RpcResult> { + const methodArgs = stripConnectionFields(params); + this.assertNoParkedCall(client); + if (isTerminalStatus(client.getStatus())) { + throw new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + "The connection dropped before this command could run; re-run the command to reconnect.", + { code: "connection_stale" }, + ); + } + const channel = new ParkingElicitationChannel(); + const unwire = wireElicitationBridge(client, channel, requestId); + const outcome: Promise<RpcResult> = (async () => { + try { + return toRpcResult(await runMethod(client, methodArgs), params.method); + } finally { + unwire(); + } + })(); + const first = await raceCallOrElicitation(outcome, channel); + if (first.kind === "settled") return first.result; + if (first.kind === "failed") throw first.error; + // Parked: the call keeps running with nothing here awaiting it — the + // eventual settle is picked up by elicitation/respond, or discarded on + // expiry/teardown. The swallow keeps a discarded failure from becoming + // an unhandled rejection. + outcome.catch(() => {}); + const entry = this.parks.add({ + client, + channel, + outcome, + unwire, + // Cancel the whole call on expiry/teardown, not just its elicitation, so + // an abandoned tool call can't emit a second prompt that misroutes to a + // later rpc on this connection (see ParkedCall.cancelCall). Same lever + // runRpcOnClient uses on caller disconnect. + cancelCall: () => { + client.cancelToolCall(); + }, + info: pendingInfo(first.frame, connectionName, { + method: params.method, + toolName: params.toolName, + }), + }); + return { kind: "elicitation-pending", elicitation: entry.info }; + } + + /** + * Answer a parked elicitation and pick up what the resumed call does + * next: its final result, its failure, or another elicitation round + * (re-parked under a fresh id). + */ + private async respondElicitation( + params: ElicitationRespondParams, + ): Promise<ElicitationRespondResult> { + if (!params?.elicitationId || typeof params.elicitationId !== "string") { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "elicitation/respond requires an elicitationId", + { code: "invalid_params" }, + ); + } + const action = params.action; + if (action !== "accept" && action !== "decline" && action !== "cancel") { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "elicitation/respond action must be accept, decline, or cancel", + { code: "invalid_params" }, + ); + } + const entry = this.parks.take(params.elicitationId); + try { + if (entry.info.mode === "url") { + if (action === "decline") { + // Mirrors the interactive prompt: URL mode has no decline — the + // user either reports completion (--done) or cancels. + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "A URL elicitation can't be declined — use --done once the linked interaction is finished, or --cancel.", + { code: "invalid_params" }, + ); + } + if (params.content !== undefined) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "A URL elicitation takes no field values — use --done once the linked interaction is finished.", + { code: "invalid_params" }, + ); + } + } + if (params.content !== undefined && action !== "accept") { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + "Field values are only valid when accepting (omit --decline/--cancel).", + { code: "invalid_params" }, + ); + } + } catch (error) { + // Validation failed after the claim — put the entry back so a + // corrected respond can still answer it. + this.parks.rearm(entry, entry.info); + throw error; + } + entry.channel.answer({ + id: entry.channel.pendingFrame()?.id ?? "", + kind: "elicitation-response", + elicitationId: entry.info.elicitationId, + action, + ...(action === "accept" && entry.info.mode === "form" + ? { content: params.content ?? {} } + : {}), + }); + const { method, toolName, connection } = entry.info; + const next = await raceCallOrElicitation(entry.outcome, entry.channel); + if (next.kind === "elicited") { + const info = this.parks.rearm( + entry, + pendingInfo(next.frame, connection, { method, toolName }), + ); + return { + method, + ...(toolName !== undefined && { toolName }), + outcome: { kind: "elicitation-pending", elicitation: info }, + }; + } + this.parks.finish(entry); + if (next.kind === "failed") throw next.error; + return { + method, + ...(toolName !== undefined && { toolName }), + outcome: next.result, + }; + } + + private async openStream( + id: string, + params: RpcParams, + ): Promise<HandleOutcome> { + if (!params?.method) { + throw new CliExitCodeError(EXIT_CODES.USAGE, "stream requires a method", { + code: "invalid_params", + }); + } + const client = await this.registry.liveClientFor( + params.name, + params.requireExplicit, + params.interactive, + ); + const methodArgs = stripConnectionFields(params); + const outcome = await runMethod(client, methodArgs); + if (outcome.kind !== "stream") { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `Method '${params.method}' is not a stream; use the rpc op.`, + { code: "use_rpc_op" }, + ); + } + return { + response: { + id, + ok: true, + result: { streaming: true, label: outcome.label }, + }, + startStream: (write, end) => { + // Tie the stream to its connection's lifecycle: when the named + // connection reaches a terminal state (mcpdo disconnect, a + // connections/use replacement, or a transport failure), end the + // stream instead of leaving the caller attached to a stale client + // until Ctrl-C or daemon idle shutdown. + const onStatus = ( + event: TypedEventGeneric<InspectorClientEventMap, "statusChange">, + ) => { + if (isTerminalStatus(event.detail)) end(); + }; + client.addEventListener("statusChange", onStatus); + const stop = outcome.start(write); + // Terminal status is persistent state, not just an event: a + // disconnect completing between runMethod() and the listener + // install above would never fire statusChange again, leaving the + // stream open against a dead client. Checking the current status + // after installing the listener closes both sides of that race. + if (isTerminalStatus(client.getStatus())) end(); + return () => { + client.removeEventListener("statusChange", onStatus); + stop(); + }; + }, + }; + } + + /** + * `mcpdod.lock` is a real lock, not bookkeeping: `O_EXCL`-create it with + * our pid, and refuse to start while another *live* daemon holds it. A + * lock left by a dead pid is reclaimed atomically: the stale file is + * `rename`d aside first, so exactly one contender wins the reclaim and a + * concurrent starter's freshly-created lock can never be deleted by the + * read-pid → unlink window of another. If the renamed-aside file turns out + * to hold a *live* pid (created between our read and the rename), it is + * restored with a create-only `link` — ownership-preserving, same inode. + * A lock with *no* pid is only stale once it outlives + * {@link LOCK_WRITE_GRACE_MS}: younger than that, it belongs to a starter + * that has created the file but not yet written its pid, so it is treated + * as held (and restored if already renamed aside) rather than stolen. + */ + private acquireLock(): void { + for (let attempt = 0; attempt < 3; attempt++) { + try { + const fd = fs.openSync(this.lockPath, "wx", 0o600); + fs.writeSync(fd, `${process.pid}\n`); + fs.closeSync(fd); + return; + } catch (error) { + if ((error as NodeJS.ErrnoException).code !== "EEXIST") throw error; + const holder = this.readLockPid(); + if (holder !== undefined && isPidAlive(holder)) { + throw new Error( + `Connection daemon lock ${this.lockPath} is held by running pid ${holder}. ` + + `Use \`mcpdo daemon stop\`, or remove the file if that pid is not an mcpdo daemon.`, + { cause: error }, + ); + } + if ( + holder === undefined && + lockFileAgeMs(this.lockPath) < LOCK_WRITE_GRACE_MS + ) { + // Empty (or unparsable) but young: a concurrent starter is between + // its O_EXCL create and its pid write. Stealing it here would let + // both daemons win — treat it as held and retry after a beat. Only + // a lock still empty past the grace period (a starter that died + // mid-create) is stale. + sleepSync(LOCK_RETRY_DELAY_MS); + continue; + } + const claimed = `${this.lockPath}.reclaim.${process.pid}`; + try { + fs.renameSync(this.lockPath, claimed); + } catch { + // Another contender renamed it first; retry the O_EXCL create. + continue; + } + const claimedPid = this.readPidFile(claimed); + if ( + (claimedPid !== undefined && isPidAlive(claimedPid)) || + (claimedPid === undefined && + lockFileAgeMs(claimed) < LOCK_WRITE_GRACE_MS) + ) { + // We renamed away a lock that a concurrent starter created between + // our dead-pid read and the rename — either it already holds a + // live pid, or it is still empty inside the pid-write grace + // period (the starter's fd targets this same inode, so its write + // still lands after the restore). Put it back without breaking + // that starter's ownership: link() re-creates the path for the + // same inode and fails (EEXIST) rather than overwriting. + try { + fs.linkSync(claimed, this.lockPath); + } catch { + // A third contender created a new lock meanwhile; the retry's + // O_EXCL create / live-pid check decides. + } + try { + fs.unlinkSync(claimed); + } catch { + // best-effort temp cleanup + } + if (claimedPid === undefined) { + sleepSync(LOCK_RETRY_DELAY_MS); + continue; + } + throw new Error( + `Connection daemon lock ${this.lockPath} is held by running pid ${claimedPid}. ` + + `Use \`mcpdo daemon stop\`, or remove the file if that pid is not an mcpdo daemon.`, + { cause: error }, + ); + } + try { + fs.unlinkSync(claimed); + } catch { + // best-effort temp cleanup + } + } + } + throw new Error( + `Could not acquire connection daemon lock ${this.lockPath}`, + ); + } + + private readLockPid(): number | undefined { + return this.readPidFile(this.lockPath); + } + + private readPidFile(filePath: string): number | undefined { + try { + const pid = Number.parseInt(fs.readFileSync(filePath, "utf8").trim(), 10); + return Number.isInteger(pid) && pid > 0 ? pid : undefined; + } catch { + return undefined; + } + } + + private releaseLock(): void { + // Only release a lock this process still owns: after a reclaim race or + // an operator's manual cleanup, the path may hold a successor's lock. + if (this.readLockPid() !== process.pid) return; + try { + fs.unlinkSync(this.lockPath); + } catch { + // absent is fine + } + } + + private removeLockAndSocket(): void { + try { + fs.unlinkSync(this.socketPath); + } catch { + // absent is fine + } + try { + fs.unlinkSync(getDaemonTokenPath(this.dir)); + } catch { + // absent is fine + } + this.releaseLock(); + } +} + +/** `kill(pid, 0)` liveness probe; EPERM means alive but not ours. */ +function isPidAlive(pid: number): boolean { + try { + process.kill(pid, 0); + return true; + } catch (error) { + return (error as NodeJS.ErrnoException).code === "EPERM"; + } +} + +/** + * How long a pidless `mcpdod.lock` is presumed to belong to a concurrent + * starter that is between its O_EXCL create and its pid write, rather than + * to a starter that died mid-create. Generous against a stalled writer while + * still reclaiming a genuinely abandoned empty lock promptly. + */ +const LOCK_WRITE_GRACE_MS = 2000; + +/** Backoff between lock-acquisition retries while inside the grace period. */ +const LOCK_RETRY_DELAY_MS = 100; + +/** Age of `filePath` since last write; missing/unstattable counts as stale. */ +function lockFileAgeMs(filePath: string): number { + try { + return Date.now() - fs.statSync(filePath).mtimeMs; + } catch { + return Number.POSITIVE_INFINITY; + } +} + +/** Synchronous sleep — acquireLock() runs in the sync startup path. */ +function sleepSync(ms: number): void { + Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), 0, 0, ms); +} + +function stripConnectionFields( + params: RpcParams, +): MethodArgs & { method: string } { + // `format` is a frontend-only output concern; forwarding it would make + // runMethod's `format === "json"` branch collect app info (an extra + // resources/read) whose result the frontend discards. `interactive` is + // daemon routing (elicitation parking, message audience), not a method + // argument. + const { name, requireExplicit, format, interactive, method, ...rest } = + params; + void name; + void requireExplicit; + void format; + void interactive; + return { method, ...rest }; +} + +/** Convert a `runMethod` outcome into the serializable `rpc` result. */ +function toRpcResult( + outcome: Awaited<ReturnType<typeof runMethod>>, + method: string, +): RpcResult { + if (outcome.kind === "stream") { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `Method '${method}' is a stream; use the stream op.`, + { code: "use_stream_op" }, + ); + } + if (outcome.kind === "ndjson") { + return { + kind: "ndjson", + lines: outcome.lines, + summary: outcome.summary, + exitCode: outcome.exitCode, + }; + } + return { + kind: "result", + result: outcome.result, + appInfo: outcome.appInfo, + }; +} + +/** Project one elicitation frame into the caller-facing pending payload. */ +function pendingInfo( + frame: ElicitationRequestFrame, + connectionName: string, + call: { method: string; toolName?: string }, +): Omit<ElicitationPendingInfo, "expiresAt"> { + return { + elicitationId: frame.elicitationId, + connection: connectionName, + method: call.method, + ...(call.toolName !== undefined && { toolName: call.toolName }), + mode: frame.mode, + message: frame.message, + ...(frame.requestedSchema !== undefined && { + requestedSchema: frame.requestedSchema, + }), + ...(frame.url !== undefined && { url: frame.url }), + origin: frame.origin, + }; +} + +type CallOrElicitation = + | { kind: "settled"; result: RpcResult } + | { kind: "failed"; error: unknown } + | { kind: "elicited"; frame: ElicitationRequestFrame }; + +/** + * Race a (possibly parked) call's completion against its next elicitation. + * Both parking sites — the original `rpc` and each `elicitation/respond` + * round — end in exactly this decision. + */ +function raceCallOrElicitation( + outcome: Promise<RpcResult>, + channel: ParkingElicitationChannel, +): Promise<CallOrElicitation> { + return Promise.race([ + outcome.then( + (result): CallOrElicitation => ({ kind: "settled", result }), + (error): CallOrElicitation => ({ kind: "failed", error }), + ), + channel + .waitForElicitation() + .then((frame): CallOrElicitation => ({ kind: "elicited", frame })), + ]); +} diff --git a/clients/mcpdo/src/daemon/stream-client.ts b/clients/mcpdo/src/daemon/stream-client.ts new file mode 100644 index 0000000000..4526a03015 --- /dev/null +++ b/clients/mcpdo/src/daemon/stream-client.ts @@ -0,0 +1,247 @@ +/** + * Long-lived daemon stream client. + */ +import { randomUUID } from "node:crypto"; +import * as net from "node:net"; +import { + CliExitCodeError, + EXIT_CODES, +} from "@inspector/core/cli/error-handler.js"; +import { getDaemonTokenFromEnv, readDaemonTokenFile } from "./auth.js"; +import { encodeRequest } from "./framing.js"; +import { getDaemonSocketPath } from "./paths.js"; +import type { + DaemonRequest, + DaemonResponse, + DaemonStreamFrame, +} from "./protocol.js"; +import type { DaemonClientOptions } from "./client.js"; +import { daemonTokenDir } from "./client.js"; +import { sanitizeText } from "../connection/sanitize.js"; + +export type StreamDaemonOptions = DaemonClientOptions & { + /** + * Called for every data frame. A returned promise applies backpressure: + * socket reads pause until it settles, so a fast daemon stream cannot + * queue unbounded output ahead of a slow consumer. + */ + onData: (data: unknown) => void | Promise<void>; + /** Abort / cancel the stream (closes the socket). */ + signal?: AbortSignal; +}; + +/** + * Long-lived `stream` op: first frame is a DaemonResponse; subsequent frames + * are {@link DaemonStreamFrame} until `end` or the socket closes. + */ +export async function streamDaemon( + params: DaemonRequest["params"], + options: StreamDaemonOptions, +): Promise<void> { + const socketPath = options.socketPath ?? getDaemonSocketPath(); + const timeoutMs = options.timeoutMs ?? 60_000; + const id = randomUUID(); + const token = + options.token ?? + getDaemonTokenFromEnv() ?? + readDaemonTokenFile(daemonTokenDir(options)); + const request: DaemonRequest = { id, op: "stream", params }; + if (token !== undefined) request.token = token; + + return new Promise<void>((resolve, reject) => { + let settled = false; + let buffer = ""; + let streaming = false; + let pendingCallbacks = 0; + let timer: ReturnType<typeof setTimeout> | undefined; + const socket = new net.Socket(); + // Decode at the socket: a multi-byte UTF-8 character split across TCP + // chunks must be reassembled by the stream's StringDecoder, not mangled + // into U+FFFD by a per-chunk String() conversion. + socket.setEncoding("utf8"); + + function settle(fn: () => void) { + if (settled) return; + settled = true; + if (timer !== undefined) clearTimeout(timer); + options.signal?.removeEventListener("abort", onAbort); + socket.removeAllListeners(); + socket.on("error", () => {}); + fn(); + } + + function fail(error: unknown) { + settle(() => { + socket.destroy(); + reject(error); + }); + } + + function succeed() { + settle(() => { + socket.destroy(); + resolve(); + }); + } + + function onAbort() { + succeed(); + } + + function handleLine(line: string) { + const trimmed = line.trim(); + if (!trimmed) return; + + if (!streaming) { + let response: DaemonResponse; + try { + response = JSON.parse(trimmed) as DaemonResponse; + } catch (error) { + fail(error); + return; + } + if (response.id !== id && response.id !== "?") return; + if (!response.ok) { + fail( + new CliExitCodeError( + response.error.exitCode ?? EXIT_CODES.USAGE, + sanitizeText(response.error.message), + { code: response.error.code }, + ), + ); + return; + } + streaming = true; + if (timer !== undefined) { + clearTimeout(timer); + timer = undefined; + } + return; + } + + let frame: DaemonStreamFrame; + try { + frame = JSON.parse(trimmed) as DaemonStreamFrame; + } catch (error) { + fail(error); + return; + } + if (frame.id !== id) return; + if (frame.stream === "data") { + const result = options.onData(frame.data); + if ( + result !== undefined && + typeof (result as Promise<void>).then === "function" + ) { + // Backpressure: stop reading until the consumer's write settles. + // Frames already split from the current chunk still dispatch + // synchronously (bounded by one socket read), but no further + // chunks are read while any callback is pending. Callback errors + // stay non-fatal, matching the previous fire-and-forget behavior. + pendingCallbacks++; + socket.pause(); + void Promise.resolve(result) + .catch(() => {}) + .finally(() => { + pendingCallbacks--; + if (pendingCallbacks === 0 && !settled) socket.resume(); + }); + } + return; + } + if (frame.stream === "end") { + succeed(); + } + } + + socket.on("error", (err) => { + if (streaming) { + // A socket error mid-stream means the daemon crashed or the + // transport broke — not a clean finish. A deliberate cancel settles + // first via onAbort, so only unsolicited errors reach here. + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Connection daemon stream failed: ${err.message}`, + { code: "daemon_unreachable" }, + ), + ); + return; + } + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Cannot reach connection daemon at ${socketPath}: ${err.message}`, + { code: "daemon_unreachable" }, + ), + ); + }); + + socket.on("close", () => { + if (settled) return; + // Only an explicit `end` frame (or a caller abort, which settles via + // onAbort before destroying) is a clean finish. EOF without `end` + // means the daemon exited or dropped the socket mid-stream. + if (streaming) { + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Connection daemon closed the stream before it ended`, + { code: "daemon_unreachable" }, + ), + ); + return; + } + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Connection daemon closed the connection before the stream opened`, + { code: "daemon_unreachable" }, + ), + ); + }); + + // timeoutMs 0 disables the deadline (mirrors callDaemon): an + // unconditional setTimeout(..., 0) would fire immediately, failing + // every stream on the next tick instead of never. + timer = + timeoutMs > 0 + ? setTimeout(() => { + fail( + new CliExitCodeError( + EXIT_CODES.UNREACHABLE, + `Daemon stream open timed out after ${timeoutMs}ms`, + { code: "daemon_timeout" }, + ), + ); + }, timeoutMs) + : undefined; + + // AbortSignal does not replay: a pre-aborted signal would never fire + // the listener, leaving the stream open until the timeout. Check first. + if (options.signal?.aborted) { + onAbort(); + } else { + options.signal?.addEventListener("abort", onAbort, { once: true }); + } + + socket.once("connect", () => { + socket.write(encodeRequest(request)); + }); + + socket.on("data", (chunk) => { + buffer += String(chunk); + let idx: number; + while ((idx = buffer.indexOf("\n")) >= 0) { + const line = buffer.slice(0, idx); + buffer = buffer.slice(idx + 1); + handleLine(line); + } + }); + + // A pre-aborted signal settles above via onAbort() and destroys the + // socket; connect() would silently un-destroy it and leak a live socket + // that pins the event loop. + if (!settled) socket.connect(socketPath); + }); +} diff --git a/clients/mcpdo/src/mcp-bin.ts b/clients/mcpdo/src/mcp-bin.ts new file mode 100644 index 0000000000..2109fbf4aa --- /dev/null +++ b/clients/mcpdo/src/mcp-bin.ts @@ -0,0 +1,42 @@ +#!/usr/bin/env node + +import { realpathSync } from "fs"; +import { resolve } from "path"; +import { fileURLToPath } from "url"; +import { handleError } from "@inspector/core/cli/error-handler.js"; +import { + disallowMemorySecretStoreFallback, + setSecretStorageWarningsQuiet, +} from "@inspector/core/auth/node/secret-store-selection.js"; +import { runMcp } from "./connection/mcp.js"; + +export { runMcp }; + +// mcpdo splits one logical session across processes (front-end, OAuth +// helper, daemon) — a per-process memory store can never serve it, so the +// keychain-less automatic fallback is always the secrets file here. +disallowMemorySecretStoreFallback(); + +// A fresh front-end process runs for every command and most touch a stored +// secret, so the automatic store-selection banner would print on every one. +// Quiet it here; `connect` re-surfaces it once (see connection/mcp.ts). +setSecretStorageWarningsQuiet(true); + +const __filename = fileURLToPath(import.meta.url); + +/** True when this file is the process entry (works through npm-link symlinks). */ +function isMainModule(): boolean { + const entry = process.argv[1]; + if (entry === undefined) return false; + try { + return realpathSync(resolve(entry)) === realpathSync(resolve(__filename)); + } catch { + return resolve(entry) === resolve(__filename); + } +} + +if (isMainModule()) { + runMcp(process.argv) + .then(() => process.exit(0)) + .catch(handleError); +} diff --git a/clients/mcpdo/tsconfig.json b/clients/mcpdo/tsconfig.json new file mode 100644 index 0000000000..b11826efd5 --- /dev/null +++ b/clients/mcpdo/tsconfig.json @@ -0,0 +1,23 @@ +{ + "extends": "../../tsconfig.base.json", + "compilerOptions": { + "noEmit": true, + // Match clients/cli/tsconfig.json's module/lib *resolution* options (mcpdo + // reaches into @inspector/core/* — including the core/cli/ surface it + // shares with the one-shot CLI — the same way cli does) so core/ is + // validated the same way its own gates validate it, rather than under + // base's stricter noUncheckedIndexedAccess, which core/ was never written + // against. + "lib": ["ES2023", "DOM", "DOM.Iterable"], + "types": ["node"], + "moduleResolution": "bundler", + "allowImportingTsExtensions": true, + "module": "ESNext", + "noUncheckedIndexedAccess": false, + "paths": { + "@inspector/core/*": ["../../core/*"] + } + }, + "include": ["src/**/*", "vitest.config.ts", "tsup.config.ts"], + "exclude": ["node_modules", "**/*.test.ts", "build"] +} diff --git a/clients/mcpdo/tsconfig.test.json b/clients/mcpdo/tsconfig.test.json new file mode 100644 index 0000000000..5c694c0343 --- /dev/null +++ b/clients/mcpdo/tsconfig.test.json @@ -0,0 +1,28 @@ +{ + // Typecheck the __tests__ dir (the src-only tsconfig.json excludes tests). + // Mirrors clients/cli/tsconfig.test.json (and clients/web's): the test-server + // barrel is aliased to its source and the module paths below resolve what + // vitest resolves via vitest.shared.mts's projectResolve, so tsc validates + // the tests against the same graph the runner executes. See #1791. + "extends": "./tsconfig.json", + "compilerOptions": { + "types": ["node", "express"], + "paths": { + "@inspector/core/*": ["../../core/*"], + "@modelcontextprotocol/inspector-test-server": [ + "../../test-servers/src/index.ts" + ], + "express": ["./node_modules/@types/express"], + "vitest": ["./node_modules/vitest"] + } + }, + // Some tests import cli's own test helpers by relative path + // (../../cli/__tests__/helpers/*) — tsc follows those transitively, no + // separate include entry needed. + // + // Only the tests root the project; tsc pulls in the `src` they import. The + // src-only tsconfig.json already validates all of `src` (without the + // test-only aliases), so listing it here too would just check it twice. + "include": ["__tests__/**/*"], + "exclude": ["node_modules", "build"] +} diff --git a/clients/mcpdo/tsup.config.ts b/clients/mcpdo/tsup.config.ts new file mode 100644 index 0000000000..67091dcaf2 --- /dev/null +++ b/clients/mcpdo/tsup.config.ts @@ -0,0 +1,55 @@ +import { defineConfig } from "tsup"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const dirname = path.dirname(fileURLToPath(import.meta.url)); +const repoRoot = path.resolve(dirname, "../.."); + +export default defineConfig({ + entry: { + "mcp-bin": "src/mcp-bin.ts", + mcpdod: "src/daemon/run.ts", + }, + format: ["esm"], + outDir: "build", + clean: true, + // No source maps in the published bundle — they roughly double the on-disk + // size and aren't needed at runtime (debug via `npm run dev` on the source). + sourcemap: false, + target: "node22", + platform: "node", + // Bundle core, including the CLI-client surface it shares with the one-shot + // CLI (`core/cli/` — handlers, error-handler, OAuth helpers; #2461). + noExternal: [/^@inspector\/core/], + // Mirrors clients/cli/tsup.config.ts (which documents each entry's story): + // this client declares NO runtime dependencies (AGENTS.md dependency- + // placement rule), so tsup's nearest-manifest auto-externalization sees + // nothing — every root-declared runtime package `core/` imports must be named here or esbuild inlines it, + // and inlining a CJS module into this ESM bundle leaves esbuild's + // `Dynamic require of "..." is not supported` shim (#2067). + // `npm run verify:bundle-externals` enforces this against the built output. + external: [ + "undici", + "@napi-rs/keyring", + "proper-lockfile", + "@modelcontextprotocol/client", + "@modelcontextprotocol/core", + "@modelcontextprotocol/ext-apps", + "@modelcontextprotocol/ext-tasks", + "commander", + "pino", + "ajv", + "atomically", + "open", + "zod", + "yaml", + "chokidar", + "hono", + "react", + ], + esbuildOptions(options) { + options.alias = { + "@inspector/core": path.join(repoRoot, "core"), + }; + }, +}); diff --git a/clients/mcpdo/vitest.config.ts b/clients/mcpdo/vitest.config.ts new file mode 100644 index 0000000000..dede3f7a13 --- /dev/null +++ b/clients/mcpdo/vitest.config.ts @@ -0,0 +1,45 @@ +import path from "node:path"; +import { fileURLToPath } from "node:url"; +import { defineConfig } from "vitest/config"; +import { + NO_RETRY_SETUP, + TIMEOUTS, + vitestSharedPaths, +} from "../../vitest.shared.mts"; + +const dirname = path.dirname(fileURLToPath(import.meta.url)); +const { projectResolve } = vitestSharedPaths(dirname); + +export default defineConfig({ + resolve: projectResolve, + test: { + globals: false, + environment: "node", + include: ["__tests__/**/*.test.ts"], + setupFiles: [NO_RETRY_SETUP], + // OAuth tokens/client secrets are split into the selected secret store by + // the shared file persistence backend. Pin the in-memory store so + // stored-auth tests (and spawned daemons, which inherit process.env) + // never probe or write the real OS keychain on a dev machine. + env: { MCP_INSPECTOR_SECRET_STORE: "memory" }, + // Shared budgets (#2323). + ...TIMEOUTS, + pool: "forks", + coverage: { + provider: "v8", + reporter: ["text", "html", "json-summary"], + include: ["src/**/*.ts"], + // Process entry points only: argv/env wiring plus a top-level call into + // covered modules. Exercised by spawning real processes (daemon spawn in + // tests, smoke), which v8 coverage can't observe from the parent. + exclude: ["src/mcp-bin.ts", "src/daemon/run.ts"], + thresholds: { + perFile: true, + lines: 90, + statements: 90, + functions: 90, + branches: 90, + }, + }, + }, +}); diff --git a/clients/tui/README.md b/clients/tui/README.md index ba2c3c083d..876ba5a773 100644 --- a/clients/tui/README.md +++ b/clients/tui/README.md @@ -31,6 +31,8 @@ npx @modelcontextprotocol/inspector --tui --config mcp.json # read-only sessi Options that specify the MCP server(s) (catalog/config file, ad-hoc command/URL, env vars, headers) are shared by the Web, CLI, and TUI and are documented in [MCP server configuration](../../docs/mcp-server-configuration.md): `--catalog` (writable catalog, seeded **empty** if missing; default `~/.mcp-inspector/mcp.json` or `MCP_CATALOG_PATH`), `--config` (read-only session, errors if absent), `-e`, `--cwd`, `--header`, `--protocol-era` (`legacy`/`auto`/`modern`; sets the era an ad-hoc server negotiates, or overrides a file's `protocolEra`), `--transport`, `--server-url`, and the positional `[target...]`. `--catalog` and `--config` are mutually exclusive, and neither combines with an ad-hoc target. +`--skill-catalog-max-skills <n>` and `--skill-catalog-max-bytes <n>` set the per-server skills catalog budget (defaults `256` skills and `67108864` bytes, 64 MiB), the same setting the web client exposes in Server Settings. Each takes a positive integer (anything else is rejected at startup) and overrides the file's [`skillCatalogMaxSkills` / `skillCatalogMaxBytes`](../../docs/mcp-server-configuration.md#inspector-specific-per-server-fields) for every server loaded, like `--header`. The budget bounds a verification run that covers **several** skills: in the **Skills** pane that is **`v`**, which verifies the whole listing in one run the way the CLI's `--verify` does, and skills past the budget come back incomplete and unread. **Enter** verifies one skill, which always reads it in full ([#2590](https://github.com/modelcontextprotocol/inspector/issues/2590)). + ### TUI-specific (OAuth for HTTP servers) The TUI supports OAuth for **SSE** and **Streamable HTTP** servers. Per-server OAuth fields in `mcp.json` (static client id/secret, scopes, enterprise-managed flag) are applied automatically when loaded from `--catalog` or `--config`. Install-wide settings (CIMD, enterprise IdP) come from **`~/.mcp-inspector/storage/client.json`** — the same file the web **Client Settings** dialog writes. You can point at a different file with `--client-config` or `MCP_CLIENT_CONFIG_PATH`. @@ -75,10 +77,14 @@ See also [EMA / enterprise-managed auth](../../specification/v2_auth_ema.md) and The TUI provides terminal-native tabs and panes for interacting with your MCP server: +- **Info**: The server's configuration and identity, plus the **Roots** the client advertises to it. With the pane focused, `e` opens the roots editor (`n` add, `x` remove), which applies changes through `setRoots` and so sends `notifications/roots/list_changed`. Edits last for the session only; the TUI does not write them back to the catalog ([#2432](https://github.com/modelcontextprotocol/inspector/issues/2432)). - **Resources**: Browse and read resources exposed by the server. +- **Subscriptions** (`u`): Shown only when the server advertises `resources.subscribe`. **Enter** subscribes to or unsubscribes from the selected resource (legacy `resources/subscribe`, or the modern `subscriptions/listen` filter, chosen by era), and the detail pane shows each subscription's last update, the modern listen-stream status, and a live feed of `notifications/resources/updated` read from the protocol log ([#2432](https://github.com/modelcontextprotocol/inspector/issues/2432)). - **Prompts**: List and test prompts. - **Tools**: View available tools and execute them with form-like inputs. A tool whose advertised schema carries a portability problem is flagged in the list — red `!` for a construct a shipping MCP client refuses, yellow `?` for one handled unevenly — and the detail pane lists each finding under **Schema Portability** with the path, the problem, and a concrete fix. The verdict comes from [`core/json/schemaLint.ts`](../../core/json/schemaLint.ts), shared with the web Tools tab and the CLI's `--strict` report, so the three cannot disagree ([#1005](https://github.com/modelcontextprotocol/inspector/issues/1005)). -- **Skills**: Shown only when the connected server declares the SEP-2640 Skills extension (`io.modelcontextprotocol/skills`), since it is a *server* declaration and so only knowable after connecting. The list marks each skill with its structural verdict — `✓` conforms, `!` warnings only, `✗` an error — using a glyph as well as a colour, because this pane is read over ssh, in tmux and through `script(1)`. The detail pane shows the entry's URI, description, conformance findings and manifest. **Enter** verifies the selected skill: one `resources/read` per manifest file, each hashed against its advertised digest, plus the frontmatter cross-check that compares the served `SKILL.md`'s own frontmatter against the one the listing advertised. Verification is a gesture rather than a page load because SEP-2640 says hosts MUST NOT retrieve a skill's files ahead of need. The checks are the same ones the web Skills tab and the CLI's `--verify` run ([#2234](https://github.com/modelcontextprotocol/inspector/issues/2234), [#2248](https://github.com/modelcontextprotocol/inspector/issues/2248)). + - **Saving a result**: on the tool result view, press **`w`** to save the result to a file. The prompt opens on `<tool-name>-result.json` (relative paths resolve against the directory the TUI was launched from); **Tab** switches between `json` — the whole result, pretty-printed — and `raw` — the text of its text blocks, or the decoded bytes of its single binary block — rendered by the shared [`core/mcp/resultFile.ts`](../../core/mcp/resultFile.ts), whose encodings match the CLI's planned `--output` ([#2431](https://github.com/modelcontextprotocol/inspector/issues/2431)). **Enter** writes it and confirms the path; **Escape** cancels. A failed write is reported on the result view ([#2571](https://github.com/modelcontextprotocol/inspector/issues/2571)). +- **Skills**: Shown only when the connected server declares the SEP-2640 Skills extension (`io.modelcontextprotocol/skills`), since it is a *server* declaration and so only knowable after connecting. The list marks each skill with its structural verdict — `✓` conforms, `!` warnings only, `✗` an error — using a glyph as well as a colour, because this pane is read over ssh, in tmux and through `script(1)`. The detail pane shows the entry's URI, description, conformance findings and manifest. **Enter** verifies the selected skill: one `resources/read` per manifest file, each hashed against its advertised digest, plus the frontmatter cross-check that compares the served `SKILL.md`'s own frontmatter against the one the listing advertised. **`v`** verifies every skill in the listing in one run — the whole catalog, even while a `/` filter narrows the list — under the server's [catalog budget](../../docs/mcp-server-configuration.md#inspector-specific-per-server-fields); a line under the list tallies the verdicts held for the current listing (`✓` verified, `✗` failed, `…` stopped by the budget, `?` unverifiable, `·` not yet checked) ([#2590](https://github.com/modelcontextprotocol/inspector/issues/2590)). Verification is a gesture rather than a page load because SEP-2640 says hosts MUST NOT retrieve a skill's files ahead of need. The checks are the same ones the web Skills tab and the CLI's `--verify` run ([#2234](https://github.com/modelcontextprotocol/inspector/issues/2234), [#2248](https://github.com/modelcontextprotocol/inspector/issues/2248)). +- **Tasks** (`s`): Shown only when the server supports Tasks (legacy `capabilities.tasks` or the negotiated SEP-2663 extension). Lists the tasks this client created, with live status. **Enter** fetches a completed or failed task's result, `x` cancels a running one, `f` refreshes and `l` clears finished tasks. `s` does not switch here while the Auth pane is focused, where it clears OAuth state ([#2432](https://github.com/modelcontextprotocol/inspector/issues/2432)). - **Protocol**: View JSON-RPC request/response/notification history (matches the web Protocol monitor). - **Network**: View HTTP fetch traffic for SSE / Streamable HTTP servers (matches the web Network monitor). - **Console**: View stdio stderr from the connected server process (matches the web Console monitor). @@ -88,7 +94,18 @@ The TUI provides terminal-native tabs and panes for interacting with your MCP se - Use the **Arrow Keys** (Left/Right) or **Tab** to switch between the main tabs (Resources, Tools, Prompts, Skills, etc.). - Use the **Arrow Keys** (Up/Down) to scroll through lists of items. - Press **Enter** to select an item, execute a tool, or fetch a resource. +- Press **`/`** on a focused Resources, Prompts, Skills or Tools list to filter it as you type (case-insensitive, by name or title — and URI, for resources). **Enter** keeps the filter and returns the keys to the list; **Escape** clears it. While the filter is being typed, the tab accelerators and Escape-to-exit are paused, so a query can contain any letter. To clear a kept filter, press `/` then **Escape**. - Press **Escape** or `Ctrl+C` to exit the application. +- Press **`?`** for an in-app list of the keybindings for the current tab; press `?` or **Escape** to close it. The list is the `src/utils/keybindings.ts` table, so a new keybinding is added there. + +## Copying values + +Any details view (opened from Resources, Prompts, Tools, Protocol or Network) and the **Auth** tab's access token can be copied out of the TUI ([#2421](https://github.com/modelcontextprotocol/inspector/issues/2421)): + +- **Y** copies the value to your clipboard with an [OSC 52](https://invisible-island.net/xterm/ctlseqs/ctlseqs.html#h3-Operating-System-Commands) escape sequence. The sequence asks the _terminal emulator_ to set its clipboard, so it reaches the machine you are sitting at even when the TUI runs over SSH. A details view copies the entry's raw JSON; the Auth tab copies the full token, not the truncated prefix it displays. +- **W** saves the value to a private temp file (`mcp-inspector-copy-*/value.txt`, mode `0600`) and shows its path — the fallback when OSC 52 does not work. + +OSC 52 is fire-and-forget: a terminal that does not support it (or has it disabled) fails silently, which is why the status line after **Y** points at **W**. Most modern terminals support it — iTerm2 (enable _Applications in terminal may access clipboard_), kitty, WezTerm, Alacritty, Windows Terminal, foot, and xterm with `allowWindowOps`. Inside **tmux**, set `set -g set-clipboard on` so tmux forwards the sequence. Some terminals cap the payload (often around 100 KB); a copy above that is flagged in the status line, and **W** has no such limit. ## Development diff --git a/clients/tui/__tests__/App.test.tsx b/clients/tui/__tests__/App.test.tsx index d259ccc507..6d70a59252 100644 --- a/clients/tui/__tests__/App.test.tsx +++ b/clients/tui/__tests__/App.test.tsx @@ -1,12 +1,22 @@ import React from "react"; import { describe, it, expect, vi, beforeEach, afterEach } from "vitest"; import { render } from "./helpers/renderTui"; +import { BodyLines } from "../src/components/BodyLines.js"; type RenderResult = ReturnType<typeof render>; vi.mock("ink-scroll-view", () => import("./helpers/inkScrollViewMock.js")); vi.mock("ink-form", () => import("./helpers/inkFormMock.js")); +// Passthrough spy on the shared, capped body renderer, so a test can prove a +// view routes its body through it even where the modal is too short to show +// the cap marker (#2539). +vi.mock("../src/components/BodyLines.js", async (importOriginal) => { + const actual = + await importOriginal<typeof import("../src/components/BodyLines.js")>(); + return { ...actual, BodyLines: vi.fn(actual.BodyLines) }; +}); + // --------------------------------------------------------------------------- // Controllable mock of the entire @inspector/core surface App.tsx depends on. // `ctrl` is mutated by individual tests (reset in beforeEach) to drive what the @@ -29,6 +39,10 @@ const h = vi.hoisted(() => { messages: unknown[]; fetchRequests: unknown[]; stderrLogs: unknown[]; + tasks: unknown[]; + subscriptions: unknown[]; + roots: unknown[]; + tasksExtension: boolean; } const ctrl: Ctrl = { status: "disconnected", @@ -46,10 +60,19 @@ const h = vi.hoisted(() => { messages: [], fetchRequests: [], stderrLogs: [], + tasks: [], + subscriptions: [], + roots: [], + tasksExtension: false, }; + const refreshTasks = vi.fn(async () => []); + const clearCompletedTasks = vi.fn(); const connect = vi.fn().mockResolvedValue(undefined); const disconnect = vi.fn().mockResolvedValue(undefined); const openUrl = vi.fn().mockResolvedValue(undefined); + // The callback App hands each OAuth-capable server's CallbackNavigation, so + // a test can drive the browser-open path directly (#2533). + const navigationCallbacks: Array<(url: URL) => unknown> = []; // Shared OAuth-related spies so a test can configure resolve/reject and // assert calls regardless of which per-server FakeClient instance App built. // Each spy is typed against the real InspectorClient method signature so its @@ -168,6 +191,13 @@ const h = vi.hoisted(() => { // "not declared" — the tab is hidden unless a test opts in by pointing // `ctrl.skillsExtension` at a declaration. getSkillsExtension = vi.fn(() => ctrl.skillsExtension); + isTasksExtensionNegotiated = vi.fn(() => ctrl.tasksExtension); + getRoots = vi.fn(() => ctrl.roots); + setRoots = vi.fn(async () => {}); + subscribeToResource = vi.fn(async () => {}); + unsubscribeFromResource = vi.fn(async () => {}); + cancelRequestorTask = vi.fn(async () => {}); + getRequestorTaskResult = vi.fn(async () => ({ content: [] })); authenticate = (...a: Parameters<InspectorClient["authenticate"]>) => clientSpies.authenticate(...a); clearOAuthTokens = ( @@ -214,6 +244,7 @@ const h = vi.hoisted(() => { connect, disconnect, openUrl, + navigationCallbacks, clientSpies, cb, createOAuthCallbackServer, @@ -248,6 +279,20 @@ const h = vi.hoisted(() => { useMessageLog: vi.fn(() => ({ messages: ctrl.messages })), useFetchRequestLog: vi.fn(() => ({ fetchRequests: ctrl.fetchRequests })), useStderrLog: vi.fn(() => ({ stderrLogs: ctrl.stderrLogs })), + refreshTasks, + clearCompletedTasks, + useManagedRequestorTasks: vi.fn(() => ({ + tasks: ctrl.tasks, + refresh: refreshTasks, + clearCompleted: clearCompletedTasks, + })), + useResourceSubscriptions: vi.fn(() => ({ + subscriptions: ctrl.subscriptions, + streamState: { active: false, status: "ended", honoredUris: [] }, + })), + // Roots are read with a direct store snapshot; the fake client is not a + // real TypedEventTarget, so the snapshot is driven from `ctrl.roots`. + useStoreSnapshot: vi.fn(() => ctrl.roots), }; }); @@ -263,6 +308,8 @@ vi.mock("@inspector/core/mcp/state/index.js", () => ({ MessageLogState: h.FakeManager, FetchRequestLogState: h.FakeManager, StderrLogState: h.FakeManager, + ManagedRequestorTasksState: h.FakeManager, + ResourceSubscriptionsState: h.FakeManager, })); vi.mock("@inspector/core/mcp/node/index.js", () => ({ createTransportNode: vi.fn(), @@ -300,12 +347,25 @@ vi.mock("@inspector/core/react/useFetchRequestLog.js", () => ({ vi.mock("@inspector/core/react/useStderrLog.js", () => ({ useStderrLog: h.useStderrLog, })); +vi.mock("@inspector/core/react/useManagedRequestorTasks.js", () => ({ + useManagedRequestorTasks: h.useManagedRequestorTasks, +})); +vi.mock("@inspector/core/react/useResourceSubscriptions.js", () => ({ + useResourceSubscriptions: h.useResourceSubscriptions, +})); +vi.mock("@inspector/core/react/useStoreSnapshot.js", () => ({ + useStoreSnapshot: h.useStoreSnapshot, +})); vi.mock("@inspector/core/auth/index.js", async (importOriginal) => { const actual = await importOriginal<typeof import("@inspector/core/auth/index.js")>(); return { ...actual, - CallbackNavigation: class {}, + CallbackNavigation: class { + constructor(callback: (url: URL) => unknown) { + h.navigationCallbacks.push(callback); + } + }, MutableRedirectUrlProvider: class { redirectUrl = ""; }, @@ -686,13 +746,20 @@ beforeEach(() => { messages: [], fetchRequests: [], stderrLogs: [], + tasks: [], + subscriptions: [], + roots: [], + tasksExtension: false, }); + h.refreshTasks.mockClear(); + h.clearCompletedTasks.mockClear(); h.connect.mockClear(); h.connect.mockResolvedValue(undefined); h.disconnect.mockClear(); h.disconnect.mockResolvedValue(undefined); h.openUrl.mockClear(); h.openUrl.mockResolvedValue(undefined); + h.navigationCallbacks.length = 0; h.cb.opts = null; h.callbackStart.mockClear(); h.callbackStop.mockClear(); @@ -952,6 +1019,45 @@ describe("App (foundation)", () => { await expectFrame(r, "OAuth"); }); + it("shows the manual-open note on the Auth tab when the browser cannot be opened", async () => { + // #2533: the opener failing (e.g. missing from PATH) must surface as a + // note on the Auth tab rather than crash the TUI — switching to that tab, + // since the flow can start from anywhere. + h.ctrl.serverType = "streamable-http"; + h.openUrl.mockImplementation( + async (_url: URL, onFailure?: (message: string) => void) => { + onFailure?.("Open it by hand"); + }, + ); + const r = await mount(httpServer()); + expect(r.lastFrame() ?? "").not.toContain("Open it by hand"); + const navigate = h.navigationCallbacks.at(-1); + expect(navigate).toBeDefined(); + await navigate!(new URL("https://auth.example/start")); + expect(h.openUrl).toHaveBeenCalledWith( + new URL("https://auth.example/start"), + expect.any(Function), + ); + await expectFrame(r, "Open it by hand"); + await expectFrame(r, "OAuth"); + }); + + it("ignores a browser-open failure for a server that is no longer selected", async () => { + h.ctrl.serverType = "streamable-http"; + h.openUrl.mockImplementation( + async (_url: URL, onFailure?: (message: string) => void) => { + onFailure?.("Open it by hand"); + }, + ); + const r = await mount(twoHttp()); + // One callback per OAuth-capable server, in catalog order: web, then api. + expect(h.navigationCallbacks).toHaveLength(2); + await h.navigationCallbacks[1]!(new URL("https://auth.example/api")); + await tick(); + expect(h.openUrl).toHaveBeenCalledOnce(); + expect(r.lastFrame() ?? "").not.toContain("Open it by hand"); + }); + it("renders connected status with capabilities", async () => { h.ctrl.status = "connected"; h.ctrl.capabilities = { tools: {}, resources: {}, prompts: {} }; @@ -1009,6 +1115,41 @@ describe("App (status, layout, modals)", () => { expect(r.lastFrame() ?? "").not.toContain("MOCK_FORM"); }); + it("mutes the global accelerators while a list filter is edited (#2430)", async () => { + h.ctrl.status = "connected"; + h.ctrl.tools = [ + sampleTool, + { ...sampleTool, name: "dpt-tool", description: "D" }, + ]; + const r = await mount(oneStdio()); + await press(r, ["t", TAB, "/"]); + await expectFrame(r, "Enter keep"); + // `d` disconnects, `p` switches to Prompts, Esc quits — none may fire + // while the keystrokes belong to the query. + await press(r, ["d", "p", "t"]); + await expectFrame(r, "Tools (1/2)"); + expect(h.disconnect).not.toHaveBeenCalled(); + expect(r.lastFrame() ?? "").toContain("/dpt"); + // Esc clears the filter rather than quitting; the app still answers keys. + await press(r, [ESC]); + await expectFrame(r, "Tools (2)"); + await press(r, ["d"]); + await waitUntil(() => h.disconnect.mock.calls.length > 0); + expect(h.disconnect).toHaveBeenCalled(); + }); + + it("types '?' into a list filter instead of opening the help (#2430 + #2436)", async () => { + h.ctrl.status = "connected"; + h.ctrl.tools = [sampleTool]; + const r = await mount(oneStdio()); + await press(r, ["t", TAB, "/", "?"]); + await expectFrame(r, "/?"); + expect(r.lastFrame() ?? "").not.toContain("Keyboard shortcuts"); + // Once the filter is cleared, '?' is the help key again. + await press(r, [ESC, "?"]); + await expectFrame(r, "Keyboard shortcuts"); + }); + it("opens the tool details modal with '+' and closes it on ESC", async () => { h.ctrl.status = "connected"; h.ctrl.tools = [sampleTool]; @@ -1094,6 +1235,41 @@ describe("App (status, layout, modals)", () => { await expectFrame(r, "Response:"); }); + // The zoom modal is bounded by the terminal, so a cap marker would sit below + // the visible frame. Assert instead that the zoom renders every body through + // the shared, capped BodyLines — the one the Network zoom uses (#2539). The + // `zoom-` key prefix tells the modal's calls apart from the Protocol pane's. + const zoomBodies = () => + Object.fromEntries( + vi + .mocked(BodyLines) + .mock.calls.map(([props]) => [props.keyPrefix, props.body]) + .filter(([key]) => key.startsWith("zoom-")), + ); + + it("renders Protocol zoom request and response bodies through the body cap (#2539)", async () => { + vi.mocked(BodyLines).mockClear(); + h.ctrl.messages = [reqMessage]; + const r = await mount(oneStdio()); + await press(r, ["p", TAB, TAB, "+"]); + await expectFrame(r, "Response:"); + expect(zoomBodies()).toEqual({ + "zoom-req": JSON.stringify(reqMessage.message), + "zoom-resp": JSON.stringify(reqMessage.response), + }); + }); + + it("renders a Protocol zoom response body through the body cap (#2539)", async () => { + vi.mocked(BodyLines).mockClear(); + h.ctrl.messages = [respMessage]; + const r = await mount(oneStdio()); + await press(r, ["p", TAB, TAB, "+"]); + await expectFrame(r, "Response:"); + expect(zoomBodies()).toEqual({ + "zoom-msg": JSON.stringify(respMessage.message), + }); + }); + it("opens in-progress request details (no status, error, or bodies)", async () => { h.ctrl.status = "connected"; h.ctrl.fetchRequests = [bareRequest]; @@ -1494,6 +1670,17 @@ describe("App (mid-session auth lifecycle events)", () => { await expectFrame(r, "unreachable"); }); + it("redacts URL query secrets in a failed revocation's detail (#2490)", async () => { + h.clientSpies.clearOAuthTokens.mockResolvedValue({ + status: "failed", + detail: "at https://a.x/?token=s3cr3t", + }); + const r = await mount(oneHttp()); + await press(r, ["a", "s"]); + await expectFrame(r, "REDACTED"); + expect(r.lastFrame() ?? "").not.toContain("s3cr3t"); + }); + // The wiring is the contract here, not just `AuthTab`'s own behavior: a // `void`-ing arrow between them resolves instantly, which makes the pending // state, the repeat lock and the rejection path all inert while revocation @@ -1848,6 +2035,39 @@ describe("App (OAuth result branches)", () => { await expectFrame(r, "Authorization updated. Retry your action"); }); + // #2490: an OAuth error quoting a secret-bearing URL is redacted on screen. + it("redacts a thrown OAuth error on an interactive reauth", async () => { + h.clientSpies.checkAuthChallengeSatisfied.mockResolvedValue(false); + h.runner.override = async () => { + throw new Error("cb https://a.x/?code=s3cr3t"); + }; + const r = await mount(oneHttp()); + await press(r, ["a"]); + h.fireClientEvent("authChallengeInteractive", { + authorizationUrl: authUrl(), + challenge: { reason: "unauthorized" }, + }); + await expectFrame(r, "REDACTED"); + expect(r.lastFrame() ?? "").not.toContain("s3cr3t"); + }); + + it("redacts a thrown OAuth error on a standard step-up authorize", async () => { + h.clientSpies.checkAuthChallengeSatisfied.mockResolvedValue(false); + h.runner.override = async () => { + throw new Error("cb https://a.x/?code=s3cr3t"); + }; + const r = await mount(oneHttp()); + await press(r, ["a"]); + h.fireClientEvent("authChallengeInteractive", { + authorizationUrl: authUrl(), + challenge: stepUpChallenge, + }); + await expectFrame(r, "needs additional OAuth scopes"); + await press(r, ["a"]); + await expectFrame(r, "REDACTED"); + expect(r.lastFrame() ?? "").not.toContain("s3cr3t"); + }); + it("completes reauth when OAuth returns already_authorized", async () => { h.clientSpies.checkAuthChallengeSatisfied.mockResolvedValue(false); h.runner.override = async () => ({ kind: "already_authorized" }); @@ -1905,3 +2125,270 @@ describe("App (OAuth result branches)", () => { await expectFrame(r, "web error"); }); }); + +describe("App (keybinding help, #2436)", () => { + it("advertises the help key in the footer", async () => { + const r = await mount(stdioServer()); + await expectFrame(r, "? help"); + }); + + it("opens with '?' showing the active tab's bindings and closes with '?'", async () => { + h.ctrl.status = "connected"; + h.ctrl.tools = [sampleTool]; + const r = await mount(oneStdio()); + await press(r, ["t", "?"]); + await expectFrame(r, "Keyboard shortcuts"); + expect(r.lastFrame() ?? "").toContain("Tools tab"); + expect(r.lastFrame() ?? "").toContain("Test the tool"); + await press(r, ["?"]); + await waitUntil( + () => !(r.lastFrame() ?? "").includes("Keyboard shortcuts"), + ); + expect(r.lastFrame() ?? "").not.toContain("Keyboard shortcuts"); + }); + + it("keeps the panes underneath inert, then restores focus on close", async () => { + h.ctrl.status = "connected"; + h.ctrl.tools = [sampleTool]; + const r = await mount(oneStdio()); + await press(r, ["t", TAB, TAB, "?"]); // tool details focused, then help + await expectFrame(r, "Keyboard shortcuts"); + // '+' would open the details dialog if the details pane still had focus; + // a tab accelerator and disconnect are global keys it must swallow too. + await press(r, ["+", "i", "d"]); + expect(h.disconnect).not.toHaveBeenCalled(); + await press(r, ["?"]); // close the help + await waitUntil( + () => !(r.lastFrame() ?? "").includes("Keyboard shortcuts"), + ); + expect(r.lastFrame() ?? "").not.toContain("Full JSON:"); + expect(r.lastFrame() ?? "").toContain("Tools (1)"); + await press(r, ["+"]); // focus is back on the details pane + await expectFrame(r, "Full JSON:"); + }); + + it("closes on ESC without exiting the app", async () => { + const r = await mount(stdioServer()); + await press(r, ["?"]); + await expectFrame(r, "Keyboard shortcuts"); + await press(r, [ESC]); + await waitUntil( + () => !(r.lastFrame() ?? "").includes("Keyboard shortcuts"), + ); + // Still mounted and listening: the help opens again. + await press(r, ["?"]); + await expectFrame(r, "Keyboard shortcuts"); + }); + + it("lists only the visible tabs' accelerators", async () => { + // A stdio server shows no Auth, Network or Skills tab, so `a`, `n` and `k` + // do nothing and must not be advertised. + const r = await mount(oneStdio()); + await press(r, ["?"]); + await expectFrame(r, "Keyboard shortcuts"); + expect(r.lastFrame() ?? "").toContain("i r m t p o "); + }); + + it("keeps a focus move made while it is open from waking a hidden pane", async () => { + h.clientSpies.checkAuthChallengeSatisfied.mockResolvedValue(false); + // Connected, so the global `c` (Connect) is inert and `c` can only mean + // the Auth pane's cancel. + h.ctrl.status = "connected"; + const r = await mount(oneHttp()); + await press(r, ["?"]); + await expectFrame(r, "Keyboard shortcuts"); + // A step-up landing now moves focus to the Auth pane underneath. + h.fireClientEvent("authChallengeInteractive", { + authorizationUrl: new URL("https://as.example/authorize"), + challenge: { + reason: "insufficient_scope" as const, + requiredScopes: ["env:read"], + authorizationScopes: ["tools:read", "env:read"], + context: { toolName: "get-env" }, + }, + }); + await expectFrame(r, "Auth tab"); + await press(r, ["c"]); // would cancel the step-up if Auth were live + await press(r, ["?"]); + await expectFrame(r, "needs additional OAuth scopes"); + expect(r.lastFrame() ?? "").not.toContain("Authorization cancelled"); + // Once the help closes, the step-up's focus move takes effect. + await press(r, ["c"]); + await expectFrame(r, "Authorization cancelled"); + }); + + it("holds back a details dialog that arrives while it is open", async () => { + h.ctrl.status = "connected"; + h.ctrl.prompts = [{ name: "greet", description: "no-arg prompt" }]; + const r = await mount(oneStdio()); + let resolvePrompt: (value: { result: { messages: [] } }) => void = () => {}; + const pending = new Promise<{ result: { messages: [] } }>((resolve) => { + resolvePrompt = resolve; + }); + for (const client of h.clientInstances) { + Object.assign(client, { getPrompt: vi.fn(() => pending) }); + } + await press(r, ["m", TAB, ENTER, "?"]); // fetch starts, then help opens + await expectFrame(r, "Keyboard shortcuts"); + resolvePrompt({ result: { messages: [] } }); + await tick(); + await press(r, [ESC]); // closes only the help + await expectFrame(r, "Prompt: greet"); + await press(r, [ESC]); // and then the details dialog + await waitUntil(() => !(r.lastFrame() ?? "").includes("Prompt: greet")); + expect(r.lastFrame() ?? "").not.toContain("Prompt: greet"); + }); + + it("ignores '?' while another dialog is open", async () => { + h.ctrl.status = "connected"; + h.ctrl.tools = [sampleTool]; + const r = await mount(oneStdio()); + await press(r, ["t", TAB, TAB, "+"]); + await expectFrame(r, "Full JSON:"); + await press(r, ["?"]); + expect(r.lastFrame() ?? "").not.toContain("Keyboard shortcuts"); + }); +}); + +/** + * The FakeClient App built for the first server. `clientInstances` is typed by + * the narrow config/options shape the mount-option tests read, while the + * instance is the hoisted `FakeClient`, which `vi.hoisted` cannot export as a + * type; the double cast bridges exactly that gap, and `InstanceType` keeps the + * spies typed against the real class rather than an ad-hoc shape. + */ +const firstFakeClient = () => + h.clientInstances[0] as unknown as InstanceType<typeof h.FakeClient>; + +describe("App (tasks, subscriptions, roots — #2432)", () => { + const task = { + taskId: "task-1", + status: "completed", + createdAt: "2026-01-01T00:00:00Z", + lastUpdatedAt: "2026-01-01T00:00:01Z", + ttl: null, + }; + + it("hides the Subscriptions and Tasks tabs for a server that serves neither", async () => { + h.ctrl.status = "connected"; + h.ctrl.capabilities = { resources: {} }; + const r = await mount(oneStdio()); + await expectFrame(r, "Tools"); + const frame = r.lastFrame() ?? ""; + expect(frame).not.toContain("Subscriptions"); + expect(frame).not.toContain("Tasks"); + }); + + it("opens the Subscriptions tab with 'u' when the server can subscribe", async () => { + h.ctrl.status = "connected"; + h.ctrl.capabilities = { resources: { subscribe: true } }; + h.ctrl.resources = [{ uri: "file:///a", name: "a" }]; + h.ctrl.subscriptions = [{ resource: { uri: "file:///a", name: "a" } }]; + const r = await mount(oneStdio()); + await expectFrame(r, "Subscriptions (1)"); + await press(r, ["u", TAB, ENTER]); + await expectFrame(r, "Subscriptions (1/1)"); + const client = firstFakeClient(); + await waitUntil(() => client.unsubscribeFromResource.mock.calls.length > 0); + expect(client.unsubscribeFromResource).toHaveBeenCalledWith("file:///a"); + }); + + it("opens the Tasks tab with 's' for a legacy tasks server and zooms a task", async () => { + h.ctrl.status = "connected"; + h.ctrl.capabilities = { tasks: {} }; + h.ctrl.tasks = [task]; + const r = await mount(oneStdio()); + await expectFrame(r, "Tasks (1)"); + await press(r, ["s", TAB]); + await expectFrame(r, "task-1"); + // Fetch the result, then zoom the details pane into the modal. + await press(r, [ENTER, TAB]); + const client = firstFakeClient(); + await waitUntil(() => client.getRequestorTaskResult.mock.calls.length > 0); + await press(r, ["+"]); + await press(r, ["f", "l"]); + expect(h.refreshTasks).not.toHaveBeenCalled(); + }); + + it("zooms a task with no fetched result", async () => { + h.ctrl.status = "connected"; + h.ctrl.tasksExtension = true; + h.ctrl.tasks = [{ ...task, status: "working" }]; + const r = await mount(oneStdio()); + await press(r, ["s", TAB, TAB, "+"]); + // The details modal swallows App input, so 's' no longer switches tabs. + await press(r, [ESC]); + await expectFrame(r, "task-1"); + }); + + it("refreshes and clears tasks through the store hook", async () => { + h.ctrl.status = "connected"; + h.ctrl.tasksExtension = true; + h.ctrl.tasks = [task]; + const r = await mount(oneStdio()); + await press(r, ["s", TAB, "f", "l"]); + await waitUntil(() => h.refreshTasks.mock.calls.length > 0); + expect(h.refreshTasks).toHaveBeenCalled(); + expect(h.clearCompletedTasks).toHaveBeenCalled(); + }); + + it("walks the new tabs with the arrow keys", async () => { + h.ctrl.status = "connected"; + h.ctrl.capabilities = { resources: { subscribe: true }, tasks: {} }; + const r = await mount(oneStdio()); + // info -> resources -> subscriptions + await press(r, ["i", RIGHT, RIGHT]); + await expectFrame(r, "Subscriptions (0/0)"); + // tools -> tasks + await press(r, ["t", RIGHT]); + await expectFrame(r, "No tasks yet"); + }); + + it("does not jump to Tasks on 's' while the Auth pane is focused", async () => { + h.ctrl.status = "connected"; + h.ctrl.capabilities = { tasks: {} }; + const r = await mount(oneHttp()); + await press(r, ["a"]); + await expectFrame(r, "Auth"); + await press(r, ["s"]); + expect(r.lastFrame() ?? "").not.toContain("No tasks yet"); + // From the tab bar, the accelerator works as usual. + await press(r, [STAB, "s"]); + await expectFrame(r, "No tasks yet"); + }); + + it("leaves the Subscriptions and Tasks tabs when the server stops serving them", async () => { + h.ctrl.status = "connected"; + h.ctrl.capabilities = { resources: { subscribe: true }, tasks: {} }; + const r = await mount(oneStdio()); + await press(r, ["s"]); + await expectFrame(r, "No tasks yet"); + h.ctrl.capabilities = {}; + r.rerender( + <App + mcpServers={oneStdio()} + clientConfig={emptyClientConfig} + callbackUrlConfig={callbackUrlConfig} + />, + ); + await expectFrame(r, "Server Configuration"); + expect(r.lastFrame() ?? "").not.toContain("No tasks yet"); + }); + + it("lists roots on the Info tab and edits them in the roots modal", async () => { + h.ctrl.status = "connected"; + h.ctrl.roots = [{ uri: "file:///work", name: "work" }]; + const r = await mount(oneStdio()); + // The 24-row test terminal shows the section heading; the entries below it + // are InfoTab's to render (covered in InfoTab.test.tsx). + await expectFrame(r, "Roots (1)"); + await press(r, ["i", TAB, "e", "x"]); + const client = firstFakeClient(); + await waitUntil(() => client.setRoots.mock.calls.length > 0); + expect(client.setRoots).toHaveBeenCalledWith([]); + // Esc closes the modal rather than exiting the app, and `e` reopens it. + await press(r, [ESC, "e", "x"]); + await waitUntil(() => client.setRoots.mock.calls.length > 1); + expect(client.setRoots).toHaveBeenCalledTimes(2); + }); +}); diff --git a/clients/tui/__tests__/AuthTab.test.tsx b/clients/tui/__tests__/AuthTab.test.tsx index 4f462e48ab..af941a9670 100644 --- a/clients/tui/__tests__/AuthTab.test.tsx +++ b/clients/tui/__tests__/AuthTab.test.tsx @@ -7,6 +7,7 @@ import type { InspectorClient } from "@inspector/core/mcp/index.js"; vi.mock("ink-scroll-view", () => import("./helpers/inkScrollViewMock.js")); import { AuthTab } from "../src/components/AuthTab.js"; +import { buildOsc52Sequence } from "../src/utils/clipboard.js"; const tick = async () => { for (let i = 0; i < 8; i++) @@ -377,6 +378,28 @@ describe("AuthTab", () => { ); expect(lastFrame() ?? "").toContain("Authenticating"); + // A note raised mid-flow stays visible while authenticating (#2533). + rerender( + <AuthTab + {...baseProps} + inspectorClient={client} + oauthStatus="authenticating" + oauthMessage="Open it by hand" + oauthMessageTone="warning" + />, + ); + expect(lastFrame() ?? "").toContain("Authenticating"); + expect(lastFrame() ?? "").toContain("Open it by hand"); + rerender( + <AuthTab + {...baseProps} + inspectorClient={client} + oauthStatus="authenticating" + oauthMessage="Re-authenticating" + />, + ); + expect(lastFrame() ?? "").toContain("Re-authenticating"); + rerender( <AuthTab {...baseProps} @@ -623,4 +646,114 @@ describe("AuthTab", () => { unmount(); expect(listeners.get("oauthComplete")?.size).toBe(0); }); + + it("Y copies the full access token via OSC 52 (#2421)", async () => { + const { client } = makeClient(sampleOAuthState); + const { stdin, stdout, lastFrame } = render( + <AuthTab + {...baseProps} + inspectorClient={client} + oauthStatus="idle" + oauthMessage={null} + focused + />, + ); + await tick(); + expect(lastFrame()).toContain("Y copy token, W save token"); + stdin.write("Y"); + await tick(); + expect(stdout.frames).toContain( + buildOsc52Sequence("tok-abcdefghijklmnopqrstuvwxyz"), + ); + expect(lastFrame()).toContain("Copied access token (30 chars)"); + }); + + it("offers no token copy when there is no access token", async () => { + const { client } = makeClient(undefined); + const { stdin, stdout, lastFrame } = render( + <AuthTab + {...baseProps} + inspectorClient={client} + oauthStatus="idle" + oauthMessage={null} + focused + />, + ); + await tick(); + expect(lastFrame()).not.toContain("Y copy token"); + stdin.write("y"); + await tick(); + expect(stdout.frames.some((f) => f.includes("\u001b]52;"))).toBe(false); + }); + + describe("server switch (#2421)", () => { + const withToken = (token: string): OAuthConnectionState => ({ + ...sampleOAuthState, + tokens: { access_token: token, token_type: "Bearer" }, + }); + /** A client whose getOAuthState resolves only when the test says so. */ + function deferredClient() { + let resolve: (s: OAuthConnectionState) => void = () => {}; + const pending = new Promise<OAuthConnectionState>((r) => { + resolve = r; + }); + const client = { + getOAuthState: vi.fn(() => pending), + addEventListener: () => {}, + removeEventListener: () => {}, + // Double cast: a partial test double covering only the three members + // AuthTab calls, same as makeClient above; InspectorClient is a class + // with private state, so no structural single cast can reach it. + } as unknown as InspectorClient; + return { client, resolve }; + } + const tab = (client: InspectorClient, serverName: string) => ( + <AuthTab + {...baseProps} + serverName={serverName} + inspectorClient={client} + oauthStatus="idle" + oauthMessage={null} + focused + /> + ); + + it("does not show or copy the previous server's token while the next one loads", async () => { + const a = makeClient(withToken("token-for-server-A-xxxxxxxx")); + const b = deferredClient(); + const { stdin, stdout, lastFrame, rerender } = render(tab(a.client, "A")); + await tick(); + expect(lastFrame()).toContain("token-for-server-A"); + + rerender(tab(b.client, "B")); + await tick(); + expect(lastFrame()).not.toContain("token-for-server-A"); + stdin.write("y"); + await tick(); + expect( + stdout.frames.some((f) => + f.includes(buildOsc52Sequence("token-for-server-A-xxxxxxxx")), + ), + ).toBe(false); + + b.resolve(withToken("token-for-server-B-yyyyyyyy")); + await tick(); + expect(lastFrame()).toContain("token-for-server-B"); + }); + + it("ignores a read that settles after the selection moved on", async () => { + const a = deferredClient(); + const b = makeClient(withToken("token-for-server-B-yyyyyyyy")); + const { lastFrame, rerender } = render(tab(a.client, "A")); + await tick(); + rerender(tab(b.client, "B")); + await tick(); + expect(lastFrame()).toContain("token-for-server-B"); + + a.resolve(withToken("token-for-server-A-xxxxxxxx")); + await tick(); + expect(lastFrame()).toContain("token-for-server-B"); + expect(lastFrame()).not.toContain("token-for-server-A"); + }); + }); }); diff --git a/clients/tui/__tests__/DetailsModal.test.tsx b/clients/tui/__tests__/DetailsModal.test.tsx index 655041a4d4..6fe28444d1 100644 --- a/clients/tui/__tests__/DetailsModal.test.tsx +++ b/clients/tui/__tests__/DetailsModal.test.tsx @@ -1,13 +1,14 @@ import React from "react"; import { describe, it, expect, vi } from "vitest"; import { render } from "./helpers/renderTui"; -import { Text } from "ink"; +import { Box, Text } from "ink"; // ScrollView: passthrough so `content` mounts and the imperative ref API // (scrollBy / getViewportHeight) exists for the scroll-key handlers. vi.mock("ink-scroll-view", () => import("./helpers/inkScrollViewMock.js")); import { DetailsModal } from "../src/components/DetailsModal.js"; +import { buildOsc52Sequence } from "../src/utils/clipboard.js"; // Ink processes stdin keypresses asynchronously — await this after stdin.write. const tick = async () => { @@ -99,4 +100,52 @@ describe("DetailsModal", () => { expect(onClose).toHaveBeenCalledTimes(1); }); + + it("Y copies copyText via OSC 52 and shows the status line (#2421)", async () => { + const onClose = vi.fn(); + // Wrapped in a sized box: the modal is position="absolute", so on its own + // it contributes nothing to the frame and the status line can't be read. + const { stdin, stdout, lastFrame } = render( + <Box width={120} height={30}> + <DetailsModal + title="Details" + content={<Text>x</Text>} + width={120} + height={30} + onClose={onClose} + copyText='{"raw":true}' + /> + </Box>, + ); + + await tick(); + expect(lastFrame()).toContain("Y to copy, W to save to a file"); + stdin.write("y"); + await tick(); + + expect(stdout.frames).toContain(buildOsc52Sequence('{"raw":true}')); + expect(lastFrame()).toContain("Copied details (12 chars)"); + expect(onClose).not.toHaveBeenCalled(); + }); + + it("offers no copy affordance without copyText", async () => { + const { stdin, stdout, lastFrame } = render( + <Box width={120} height={30}> + <DetailsModal + title="Details" + content={<Text>x</Text>} + width={120} + height={30} + onClose={() => {}} + /> + </Box>, + ); + + await tick(); + stdin.write("y"); + await tick(); + + expect(lastFrame()).not.toContain("Y to copy"); + expect(stdout.frames.some((f) => f.includes("\u001b]52;"))).toBe(false); + }); }); diff --git a/clients/tui/__tests__/HelpOverlay.test.tsx b/clients/tui/__tests__/HelpOverlay.test.tsx new file mode 100644 index 0000000000..cb246bd258 --- /dev/null +++ b/clients/tui/__tests__/HelpOverlay.test.tsx @@ -0,0 +1,104 @@ +import React from "react"; +import { describe, it, expect, vi } from "vitest"; +import { render } from "./helpers/renderTui"; +import { Box } from "ink"; + +// ScrollView: a passthrough like the shared `inkScrollViewMock`, but with a +// spied `scrollBy`, so the scroll keys' deltas are asserted rather than merely +// exercised (the shared double's `scrollBy` is a no-op). +const scroll = vi.hoisted(() => ({ scrollBy: vi.fn(), viewportHeight: 7 })); +vi.mock("ink-scroll-view", async () => { + const React = await import("react"); + const { Box } = await import("ink"); + const ScrollView = React.forwardRef<unknown, { children?: React.ReactNode }>( + function ScrollView({ children }, ref) { + React.useImperativeHandle(ref, () => ({ + scrollBy: scroll.scrollBy, + scrollTo: () => {}, + getViewportHeight: () => scroll.viewportHeight, + })); + return React.createElement(Box, { flexDirection: "column" }, children); + }, + ); + return { ScrollView }; +}); + +import { HelpOverlay } from "../src/components/HelpOverlay.js"; +import { keybindingSections } from "../src/utils/keybindings.js"; + +// Ink processes stdin keypresses asynchronously — await this after stdin.write. +const tick = async () => { + for (let i = 0; i < 8; i++) + await new Promise((resolve) => setTimeout(resolve, 4)); +}; + +const ESC = String.fromCharCode(27); +const UP = `${ESC}[A`; +const DOWN = `${ESC}[B`; +const PAGE_UP = `${ESC}[5~`; +const PAGE_DOWN = `${ESC}[6~`; + +/** + * The overlay is `position="absolute"`, which a bare render does not lay out + * into the frame; a sized parent gives it a box to draw into, as `App` does. + */ +function renderOverlay(onClose: () => void) { + return render( + <Box width={100} height={60}> + <HelpOverlay + sections={keybindingSections("tools")} + width={100} + height={60} + onClose={onClose} + /> + </Box>, + ); +} + +describe("HelpOverlay", () => { + it("lists the global and active-tab bindings", async () => { + const r = renderOverlay(() => {}); + await tick(); + const frame = r.lastFrame() ?? ""; + expect(frame).toContain("Keyboard shortcuts"); + expect(frame).toContain("Global"); + expect(frame).toContain("Tools tab"); + expect(frame).toContain("Test the tool"); + expect(frame).toContain("Show or hide this help"); + r.unmount(); + }); + + it("scrolls by a line or a viewport, without closing", async () => { + scroll.scrollBy.mockClear(); + const onClose = vi.fn(); + const r = renderOverlay(onClose); + await tick(); + for (const k of [DOWN, UP, PAGE_DOWN, PAGE_UP, "x"]) { + r.stdin.write(k); + await tick(); + } + expect(scroll.scrollBy.mock.calls).toEqual([[1], [-1], [7], [-7]]); + expect(onClose).not.toHaveBeenCalled(); + r.unmount(); + }); + + it("closes on '?'", async () => { + const onClose = vi.fn(); + const r = renderOverlay(onClose); + await tick(); + r.stdin.write("?"); + await tick(); + expect(onClose).toHaveBeenCalledTimes(1); + r.unmount(); + }); + + it("closes on ESC", async () => { + const onClose = vi.fn(); + const r = renderOverlay(onClose); + await tick(); + r.stdin.write(ESC); + await tick(); + expect(onClose).toHaveBeenCalledTimes(1); + r.unmount(); + }); +}); diff --git a/clients/tui/__tests__/HistoryTab.test.tsx b/clients/tui/__tests__/HistoryTab.test.tsx index 6d872522ad..b43c42a594 100644 --- a/clients/tui/__tests__/HistoryTab.test.tsx +++ b/clients/tui/__tests__/HistoryTab.test.tsx @@ -9,6 +9,7 @@ import type { MessageEntry } from "@inspector/core/mcp/index.js"; vi.mock("ink-scroll-view", () => import("./helpers/inkScrollViewMock.js")); import { HistoryTab } from "../src/components/HistoryTab.js"; +import { MAX_BODY_LINES } from "../src/utils/bodyLines.js"; // Ink processes stdin keypresses asynchronously — await this after stdin.write // and after rerender() before asserting. @@ -202,6 +203,71 @@ describe("HistoryTab", () => { expect(frame).toContain("notifications/message"); }); + describe("caps large bodies like the Network pane (#2539)", () => { + // 600 tools pretty-print to well over MAX_BODY_LINES lines (6 per tool), + // so every body below must stop at the cap and say what it hid. + const bigTools = Array.from({ length: 600 }, (_, i) => ({ + name: `tool_${i}`, + })); + const bigReq = entry({ + id: "big-req", + direction: "request", + message: { + jsonrpc: "2.0", + id: 9, + method: "tools/call", + params: { name: "x", arguments: { tools: bigTools } }, + }, + response: { jsonrpc: "2.0", id: 9, result: { tools: bigTools } }, + }); + const bigResp = entry({ + id: "big-resp", + direction: "response", + message: { jsonrpc: "2.0", id: 10, result: { tools: bigTools } }, + }); + const bigNotif = entry({ + id: "big-notif", + direction: "notification", + message: { + jsonrpc: "2.0", + method: "notifications/message", + params: { level: "info", data: bigTools }, + }, + }); + + const frameFor = (msg: MessageEntry) => + render( + <HistoryTab + serverName="srv" + messages={[msg]} + width={120} + // Tall enough that the outer Box does not clip the cap marker. + height={4000} + />, + ).lastFrame() ?? ""; + + it("caps both the request and its response", () => { + const frame = frameFor(bigReq); + const markers = frame.match(/more lines not shown/g) ?? []; + expect(markers).toHaveLength(2); + expect(frame).not.toContain("tool_599"); + }); + + it("caps a response entry", () => { + const frame = frameFor(bigResp); + expect(frame).toMatch(/… \d+ more lines not shown \(\d+ total\)/); + expect(frame).not.toContain("tool_599"); + // The cap is the shared one, not a local constant that could drift. + expect(frame).toContain(`tool_${Math.floor(MAX_BODY_LINES / 3) - 2}`); + }); + + it("caps a notification entry", () => { + const frame = frameFor(bigNotif); + expect(frame).toContain("more lines not shown"); + expect(frame).not.toContain("tool_599"); + }); + }); + it("falls back to the Message header for a methodless notification", () => { const { lastFrame } = render( <HistoryTab diff --git a/clients/tui/__tests__/InfoTab.test.tsx b/clients/tui/__tests__/InfoTab.test.tsx index 099eee47ef..ac0144402f 100644 --- a/clients/tui/__tests__/InfoTab.test.tsx +++ b/clients/tui/__tests__/InfoTab.test.tsx @@ -396,4 +396,72 @@ describe("InfoTab", () => { ); expect(lastFrame() ?? "").not.toContain("to scroll"); }); + + describe("roots (#2432)", () => { + it("lists the advertised roots, with and without a name", () => { + const { lastFrame } = render( + <InfoTab + serverName="my-server" + serverConfig={stdioConfig} + serverState={baseState} + width={80} + height={60} + roots={[{ uri: "file:///a", name: "alpha" }, { uri: "file:///b" }]} + />, + ); + const frame = lastFrame() ?? ""; + expect(frame).toContain("Roots (2)"); + expect(frame).toContain("file:///a (alpha)"); + expect(frame).toContain("file:///b"); + }); + + it("says None when no roots are advertised", () => { + const { lastFrame } = render( + <InfoTab + serverName="my-server" + serverConfig={stdioConfig} + serverState={baseState} + width={80} + height={60} + />, + ); + expect(lastFrame() ?? "").toContain("Roots (0)"); + expect(lastFrame() ?? "").toContain("None"); + }); + + it("opens the roots editor on 'e' when focused", async () => { + const onEditRoots = vi.fn(); + const { stdin, lastFrame } = render( + <InfoTab + serverName="my-server" + serverConfig={stdioConfig} + serverState={baseState} + width={80} + height={60} + focused + onEditRoots={onEditRoots} + />, + ); + expect(lastFrame() ?? "").toContain("e to edit roots"); + stdin.write("e"); + await tick(); + expect(onEditRoots).toHaveBeenCalledTimes(1); + }); + + it("ignores 'e' when no editor is wired", async () => { + const { stdin, lastFrame } = render( + <InfoTab + serverName="my-server" + serverConfig={stdioConfig} + serverState={baseState} + width={80} + height={60} + focused + />, + ); + stdin.write("e"); + await tick(); + expect(lastFrame() ?? "").not.toContain("e to edit roots"); + }); + }); }); diff --git a/clients/tui/__tests__/ListFilterBar.test.tsx b/clients/tui/__tests__/ListFilterBar.test.tsx new file mode 100644 index 0000000000..4064dac5c0 --- /dev/null +++ b/clients/tui/__tests__/ListFilterBar.test.tsx @@ -0,0 +1,46 @@ +import React from "react"; +import { describe, it, expect } from "vitest"; +import { render } from "./helpers/renderTui"; +import { ListFilterBar, filterCount } from "../src/components/ListFilterBar.js"; + +describe("ListFilterBar", () => { + it("shows the query, a cursor and the ways out while editing", () => { + const { lastFrame } = render( + <ListFilterBar query="echo" editing focused />, + ); + expect(lastFrame()).toContain("/echo"); + expect(lastFrame()).toContain("Enter keep · Esc clear"); + }); + + it("shows a kept query and how to change it", () => { + const { lastFrame } = render( + <ListFilterBar query="echo" editing={false} focused={false} />, + ); + expect(lastFrame()).toContain("/echo"); + expect(lastFrame()).toContain("(/ to edit)"); + }); + + it("advertises '/' on a focused list with no query", () => { + const { lastFrame } = render( + <ListFilterBar query=" " editing={false} focused />, + ); + expect(lastFrame()).toContain("/ to filter"); + }); + + it("stays blank on an unfocused list with no query", () => { + const { lastFrame } = render( + <ListFilterBar query="" editing={false} focused={false} />, + ); + expect((lastFrame() ?? "").trim()).toBe(""); + }); +}); + +describe("filterCount", () => { + it("is the total alone when nothing is filtered", () => { + expect(filterCount(false, 3, 3)).toBe("3"); + }); + + it("is matches over total while a filter narrows the list", () => { + expect(filterCount(true, 1, 3)).toBe("1/3"); + }); +}); diff --git a/clients/tui/__tests__/PromptsTab.test.tsx b/clients/tui/__tests__/PromptsTab.test.tsx index 3d7b4ba0c1..8675653bcc 100644 --- a/clients/tui/__tests__/PromptsTab.test.tsx +++ b/clients/tui/__tests__/PromptsTab.test.tsx @@ -289,3 +289,39 @@ describe("PromptsTab", () => { expect(onFetchPrompt).not.toHaveBeenCalled(); }); }); + +describe("PromptsTab list filter (#2430)", () => { + it("narrows the list by name or title and shows a no-match state", async () => { + const filterable: Prompt[] = [ + makePrompt({ name: "alpha" }), + makePrompt({ name: "", title: "Gamma Title", description: "G desc" }), + makePrompt({ name: "delta" }), + ]; + const { lastFrame, stdin } = render( + <PromptsTab + prompts={filterable} + inspectorClient={null} + width={120} + height={30} + focusedPane="list" + />, + ); + await tick(); + for (const k of ["/", "g", "a", "m"]) { + stdin.write(k); + await tick(); + } + let frame = lastFrame() ?? ""; + expect(frame).toContain("Prompts (1/3)"); + expect(frame).toContain("Prompt 2"); + expect(frame).toContain("G desc"); + stdin.write("z"); + await tick(); + frame = lastFrame() ?? ""; + expect(frame).toContain("No prompts match the filter"); + expect(frame).toContain("Select a prompt to view details"); + stdin.write(ESC); + await tick(); + expect(lastFrame()).toContain("Prompts (3)"); + }); +}); diff --git a/clients/tui/__tests__/ResourcesTab.test.tsx b/clients/tui/__tests__/ResourcesTab.test.tsx index 621abd7caa..68f2ec1946 100644 --- a/clients/tui/__tests__/ResourcesTab.test.tsx +++ b/clients/tui/__tests__/ResourcesTab.test.tsx @@ -410,3 +410,75 @@ describe("ResourcesTab", () => { expect(frame).toContain("[Enter to Fetch Resource]"); }); }); + +describe("ResourcesTab list filter (#2430)", () => { + it("narrows resources and templates by name or URI", async () => { + const onCountChange = vi.fn(); + const { lastFrame, stdin } = render( + <ResourcesTab + resources={resources} + resourceTemplates={templates} + inspectorClient={null} + width={120} + height={30} + focusedPane="list" + onCountChange={onCountChange} + />, + ); + await tick(); + for (const k of ["/", "{", "i", "d"]) { + stdin.write(k); + await tick(); + } + let frame = lastFrame() ?? ""; + // Only the template's URI carries "{id}". + expect(frame).toContain("Resources (1/5)"); + expect(frame).toContain("tmpl-alpha"); + expect(frame).not.toContain("res-alpha"); + + for (const k of ["\x7f", "\x7f", "\x7f", "/", "b"]) { + stdin.write(k); + await tick(); + } + frame = lastFrame() ?? ""; + expect(frame).toContain("Resources (1/5)"); + expect(frame).toContain("file:///b"); + + stdin.write("q"); + await tick(); + expect(lastFrame()).toContain("No resources match the filter"); + // The tab count reports the unfiltered total, never the filtered view. + expect(onCountChange).not.toHaveBeenCalled(); + }); +}); + +describe("ResourcesTab list filter — template titles (#2430)", () => { + it("matches a resource template by its title", async () => { + const { lastFrame, stdin } = render( + <ResourcesTab + resources={[]} + resourceTemplates={[ + { + name: "rows", + title: "Database Table Row", + uriTemplate: "db://{id}", + }, + { name: "other", uriTemplate: "x://{id}" }, + ]} + inspectorClient={null} + width={120} + height={30} + focusedPane="list" + />, + ); + await tick(); + for (const k of ["/", "t", "a", "b", "l", "e"]) { + stdin.write(k); + await tick(); + } + const frame = lastFrame() ?? ""; + expect(frame).toContain("Resources (1/2)"); + expect(frame).toContain("rows"); + expect(frame).not.toContain("other"); + }); +}); diff --git a/clients/tui/__tests__/RootsModal.test.tsx b/clients/tui/__tests__/RootsModal.test.tsx new file mode 100644 index 0000000000..9a6249dcf3 --- /dev/null +++ b/clients/tui/__tests__/RootsModal.test.tsx @@ -0,0 +1,236 @@ +import React from "react"; +import { describe, it, expect, vi, afterEach } from "vitest"; +import { render } from "./helpers/renderTui"; +import type { Root } from "@modelcontextprotocol/client"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; + +vi.mock("ink-form", () => import("./helpers/inkFormMock.js")); +// Passthrough spy: the modal's frame is empty under ink-testing-library (see +// below), so the redaction test asserts the error reached the display boundary. +vi.mock("../src/utils/errorText.js", async (importOriginal) => { + const actual = + await importOriginal<typeof import("../src/utils/errorText.js")>(); + return { ...actual, errorMessage: vi.fn(actual.errorMessage) }; +}); +import * as errorText from "../src/utils/errorText.js"; + +import { RootsModal, rootFromForm } from "../src/components/RootsModal.js"; + +// The modal renders position="absolute", which ink-testing-library draws as an +// EMPTY frame (see ResourceTestModal.test.tsx), so these tests assert on the +// client fake and the callbacks rather than on lastFrame(). + +const tick = async () => { + for (let i = 0; i < 8; i++) + await new Promise((resolve) => setTimeout(resolve, 4)); +}; +const ESC = String.fromCharCode(27); +const UP = `${ESC}[A`; +const DOWN = `${ESC}[B`; +const DELETE = `${ESC}[3~`; + +const setSubmitValue = (value: Record<string, unknown>) => { + (globalThis as Record<string, unknown>).__INK_FORM_SUBMIT_VALUE__ = value; +}; + +afterEach(() => { + delete (globalThis as Record<string, unknown>).__INK_FORM_SUBMIT_VALUE__; +}); + +// A deliberately partial fake: RootsModal calls only `setRoots` on the client, +// and InspectorClient is a class with private members that no structural object +// literal can satisfy, so a single `as` is refused. The double cast is confined +// to this factory, and the intersection keeps `setRoots` typed as the spy. +const fakeClient = ( + setRoots: ReturnType<typeof vi.fn> = vi.fn(async () => {}), +) => + ({ setRoots }) as unknown as InspectorClient & { setRoots: typeof setRoots }; + +const roots: Root[] = [{ uri: "file:///a", name: "a" }, { uri: "file:///b" }]; + +function renderModal(props: Partial<React.ComponentProps<typeof RootsModal>>) { + const onClose = vi.fn(); + const client = fakeClient(); + const api = render( + <RootsModal + roots={roots} + inspectorClient={client} + connected + width={80} + height={24} + onClose={onClose} + {...props} + />, + ); + return { ...api, onClose, client }; +} + +describe("rootFromForm", () => { + it("trims, requires a URI and drops an empty name", () => { + expect(rootFromForm({ uri: " file:///x ", name: " x " })).toEqual({ + uri: "file:///x", + name: "x", + }); + expect(rootFromForm({ uri: "file:///x", name: " " })).toEqual({ + uri: "file:///x", + }); + expect(rootFromForm({ uri: "file:///x" })).toEqual({ uri: "file:///x" }); + expect(rootFromForm({ uri: " " })).toEqual({ error: "A root needs a URI" }); + expect(rootFromForm({})).toEqual({ error: "A root needs a URI" }); + }); +}); + +describe("RootsModal", () => { + it("removes the selected root with 'x' after moving the selection", async () => { + const { stdin, client } = renderModal({}); + await tick(); + stdin.write(DOWN); + await tick(); + stdin.write(DOWN); // at the end + await tick(); + stdin.write(UP); + await tick(); + stdin.write(UP); // at the top + await tick(); + stdin.write(DOWN); + await tick(); + stdin.write("x"); + await tick(); + expect(client.setRoots).toHaveBeenCalledWith([ + { uri: "file:///a", name: "a" }, + ]); + }); + + it("removes with the Delete key too", async () => { + const { stdin, client } = renderModal({}); + await tick(); + stdin.write(DELETE); + await tick(); + expect(client.setRoots).toHaveBeenCalledWith([{ uri: "file:///b" }]); + }); + + it("adds a root through the form", async () => { + const { stdin, client } = renderModal({}); + await tick(); + stdin.write("n"); + await tick(); + setSubmitValue({ uri: "file:///c", name: "c" }); + stdin.write("\r"); + await tick(); + expect(client.setRoots).toHaveBeenCalledWith([ + ...roots, + { uri: "file:///c", name: "c" }, + ]); + }); + + it("refuses a root without a URI and returns to the list", async () => { + const { stdin, client, onClose } = renderModal({}); + await tick(); + stdin.write("+"); + await tick(); + setSubmitValue({ uri: "" }); + stdin.write("\r"); + await tick(); + expect(client.setRoots).not.toHaveBeenCalled(); + // Back on the list, Esc closes. + stdin.write(ESC); + await tick(); + expect(onClose).toHaveBeenCalledTimes(1); + }); + + it("backs out of the form with Esc before closing", async () => { + const { stdin, onClose } = renderModal({}); + await tick(); + stdin.write("n"); + await tick(); + stdin.write("x"); // the form owns this key, so nothing is removed + await tick(); + stdin.write(ESC); + await tick(); + expect(onClose).not.toHaveBeenCalled(); + stdin.write(ESC); + await tick(); + expect(onClose).toHaveBeenCalledTimes(1); + }); + + it("keeps the form open and reports a failed save", async () => { + const setRoots = vi.fn(async () => { + throw new Error("Client is not connected"); + }); + const { stdin } = renderModal({ inspectorClient: fakeClient(setRoots) }); + await tick(); + stdin.write("x"); + await tick(); + expect(setRoots).toHaveBeenCalledTimes(1); + }); + + it("shows a failed save through the redacting display boundary (#2638)", async () => { + const failure = new Error( + "Request failed: https://auth.example/cb?code=s3cret&state=ok", + ); + const setRoots = vi.fn(async () => { + throw failure; + }); + const { stdin } = renderModal({ inspectorClient: fakeClient(setRoots) }); + await tick(); + stdin.write("x"); + await tick(); + expect(errorText.errorMessage).toHaveBeenCalledWith(failure); + expect(vi.mocked(errorText.errorMessage).mock.results.at(-1)?.value).toBe( + "Request failed: https://auth.example/cb?code=%5BREDACTED%5D&state=ok", + ); + }); + + it("reports a non-Error failure", async () => { + const setRoots = vi.fn(() => Promise.reject("nope")); + const { stdin } = renderModal({ inspectorClient: fakeClient(setRoots) }); + await tick(); + stdin.write("x"); + await tick(); + expect(setRoots).toHaveBeenCalledTimes(1); + }); + + it("ignores keys while a save is in flight", async () => { + let release: () => void = () => {}; + const setRoots = vi.fn( + () => + new Promise<void>((resolve) => { + release = resolve; + }), + ); + const { stdin } = renderModal({ inspectorClient: fakeClient(setRoots) }); + await tick(); + // Two keys in one burst, before any re-render, still save once. + stdin.write("x"); + stdin.write("x"); + await tick(); + stdin.write("x"); + await tick(); + expect(setRoots).toHaveBeenCalledTimes(1); + release(); + await tick(); + }); + + it("is read-only while disconnected", async () => { + const { stdin, client } = renderModal({ connected: false }); + await tick(); + stdin.write("n"); + await tick(); + stdin.write("x"); + await tick(); + expect(client.setRoots).not.toHaveBeenCalled(); + }); + + it("does nothing on an empty list or without a client", async () => { + const { stdin, client } = renderModal({ roots: [] }); + await tick(); + stdin.write("x"); + await tick(); + expect(client.setRoots).not.toHaveBeenCalled(); + + const noClient = renderModal({ inspectorClient: null }); + await tick(); + noClient.stdin.write("x"); + await tick(); + }); +}); diff --git a/clients/tui/__tests__/SaveResultBar.test.tsx b/clients/tui/__tests__/SaveResultBar.test.tsx new file mode 100644 index 0000000000..27c00474d7 --- /dev/null +++ b/clients/tui/__tests__/SaveResultBar.test.tsx @@ -0,0 +1,44 @@ +import React from "react"; +import { describe, it, expect } from "vitest"; +import { render } from "./helpers/renderTui"; +import { SaveResultBar } from "../src/components/SaveResultBar.js"; + +describe("SaveResultBar (#2571)", () => { + it("shows the prompt with its format, path and key hints", () => { + const { lastFrame } = render( + <SaveResultBar + prompt={{ path: "alpha-result.json", format: "json", edited: false }} + status={{ ok: true, message: "hidden while prompting" }} + />, + ); + const frame = lastFrame() ?? ""; + expect(frame).toContain("Save as json: alpha-result.json"); + expect(frame).toContain("Enter to save, Tab to switch json/raw, ESC"); + expect(frame).not.toContain("hidden while prompting"); + }); + + it("shows a confirmation once the prompt closes", () => { + const { lastFrame } = render( + <SaveResultBar + prompt={null} + status={{ ok: true, message: "Saved json result to /tmp/a.json" }} + />, + ); + expect(lastFrame()).toContain("Saved json result to /tmp/a.json"); + }); + + it("shows a failure message", () => { + const { lastFrame } = render( + <SaveResultBar + prompt={null} + status={{ ok: false, message: "Could not write /nope/a.json" }} + />, + ); + expect(lastFrame()).toContain("Could not write /nope/a.json"); + }); + + it("renders nothing when idle", () => { + const { lastFrame } = render(<SaveResultBar prompt={null} status={null} />); + expect(lastFrame()).toBe(""); + }); +}); diff --git a/clients/tui/__tests__/SkillsTab.test.tsx b/clients/tui/__tests__/SkillsTab.test.tsx index 9caade4b8a..eb95a02e3a 100644 --- a/clients/tui/__tests__/SkillsTab.test.tsx +++ b/clients/tui/__tests__/SkillsTab.test.tsx @@ -897,3 +897,315 @@ describe("SkillsTab (#2248)", () => { ); }); }); + +describe("SkillsTab list filter (#2430)", () => { + it("narrows the catalog by skill name", async () => { + const { lastFrame, stdin } = render( + <SkillsTab + skills={skills} + pageCount={1} + inspectorClient={null} + width={140} + height={30} + focusedPane="list" + />, + ); + await tick(); + for (const k of ["/", "r", "i", "g", "h", "t"]) { + stdin.write(k); + await tick(); + } + let frame = lastFrame() ?? ""; + expect(frame).toContain("Skills (1/4)"); + expect(frame).toContain("skill://wrong-folder/SKILL.md"); + stdin.write("!"); + await tick(); + frame = lastFrame() ?? ""; + expect(frame).toContain("No skills match the filter"); + expect(frame).toContain("Select a skill to view details"); + }); +}); + +describe("SkillsTab verify all (#2590)", () => { + /** A client whose server settings carry a catalog budget, as `verifySkills` reads it. */ + function budgetedClient( + readResource: ReturnType<typeof vi.fn>, + settings: { skillCatalogMaxSkills?: number; skillCatalogMaxBytes?: number }, + ): InspectorClient { + // Built on `mockClient` so no further assertion is needed: the accessor is + // the one thing `verifySkills` reads the budget through. + return Object.assign(mockClient(readResource), { + getServerSettings: () => settings, + }); + } + + // Structurally clean, so its outcome is decided by reading it — `broken` + // fails its static checks whether or not its files are read. + const other: SkillEntry = { + uri: "skill://other/SKILL.md", + frontmatter: { name: "other", description: "Another clean skill" }, + resources: [ + { uri: "skill://other/SKILL.md", digest: CLEAN_DIGEST, size: 51 }, + ], + }; + + const readUris = (readResource: ReturnType<typeof vi.fn>) => + readResource.mock.calls.map((call) => String(call[0])); + + it("hints at the gesture until something has been verified", () => { + const { lastFrame } = render( + <SkillsTab + skills={skills} + pageCount={1} + inspectorClient={null} + width={140} + height={30} + />, + ); + expect(lastFrame() ?? "").toContain("[v: verify all]"); + }); + + it("verifies the whole listing in one run, so the catalog budget bounds it", async () => { + const readResource = vi.fn().mockResolvedValue({ + result: { contents: [{ uri: "skill://clean/SKILL.md", text: SKILL_MD }] }, + }); + const { lastFrame, stdin } = render( + <SkillsTab + skills={[clean, other]} + pageCount={1} + inspectorClient={budgetedClient(readResource, { + skillCatalogMaxSkills: 1, + })} + width={140} + height={30} + focusedPane="list" + />, + ); + stdin.write("v"); + await tick(); + // The budget of one skill stopped the walk after `clean`: `other`'s file + // was never fetched. + const uris = readUris(readResource); + expect(uris.some((uri) => uri.includes("clean"))).toBe(true); + expect(uris.some((uri) => uri.includes("other"))).toBe(false); + let frame = lastFrame() ?? ""; + expect(frame).toContain("✓0 ✗1 …1"); + // The skill past the budget says why it was not read. + stdin.write(DOWN); + await tick(); + frame = lastFrame() ?? ""; + expect(frame).toContain("Incomplete:"); + expect(frame).toContain("catalog budget of 1 skills"); + expect(frame).toContain("Verification INCOMPLETE"); + }); + + it("covers the whole catalog even while a filter narrows the list", async () => { + const readResource = vi.fn().mockResolvedValue({ + result: { contents: [{ uri: "skill://clean/SKILL.md", text: SKILL_MD }] }, + }); + const { lastFrame, stdin } = render( + <SkillsTab + skills={[clean, broken]} + pageCount={1} + inspectorClient={mockClient(readResource)} + width={140} + height={30} + focusedPane="list" + />, + ); + await tick(); + for (const k of ["/", "c", "l", "e", "a", "n", ENTER]) { + stdin.write(k); + await tick(); + } + expect(lastFrame() ?? "").toContain("Skills (1/2)"); + stdin.write("v"); + await tick(); + expect( + readUris(readResource).some((uri) => uri.includes("wrong-folder")), + ).toBe(true); + expect(lastFrame() ?? "").toContain("✓0 ✗2"); + }); + + it("tallies single verifications against what is still unchecked", async () => { + const { lastFrame, stdin } = render( + <SkillsTab + skills={[clean, broken]} + pageCount={1} + inspectorClient={mockClient()} + width={140} + height={30} + focusedPane="list" + />, + ); + stdin.write(ENTER); + await tick(); + expect(lastFrame() ?? "").toContain("✓0 ✗1 ·1"); + }); + + it("counts an unverifiable skill separately", async () => { + const genMd = "---\nname: gen\ndescription: Generated\n---\n\n# G\n"; + const readResource = vi.fn().mockResolvedValue({ + result: { contents: [{ uri: "skill://gen/SKILL.md", text: genMd }] }, + }); + const { lastFrame, stdin } = render( + <SkillsTab + skills={[dynamic]} + pageCount={1} + inspectorClient={mockClient(readResource)} + width={140} + height={30} + focusedPane="list" + />, + ); + stdin.write("v"); + await tick(); + expect(lastFrame() ?? "").toContain("✓0 ✗0 ?1"); + }); + + it("does nothing without a client", async () => { + const { lastFrame, stdin } = render( + <SkillsTab + skills={skills} + pageCount={1} + inspectorClient={null} + width={140} + height={30} + focusedPane="list" + />, + ); + stdin.write("v"); + await tick(); + expect(lastFrame() ?? "").toContain("[v: verify all]"); + }); + + it("keeps a read verdict when a repeated entry is cut off by the budget", async () => { + // A malformed listing repeating `clean` verbatim: the walk reads the first + // copy and the budget stops it before the second. The second copy's unread + // `incomplete` must not overwrite the first copy's real failure. + const { lastFrame, stdin } = render( + <SkillsTab + skills={[clean, clean]} + pageCount={1} + inspectorClient={budgetedClient( + vi.fn().mockResolvedValue({ + result: { + contents: [{ uri: "skill://clean/SKILL.md", text: SKILL_MD }], + }, + }), + { skillCatalogMaxSkills: 1 }, + )} + width={140} + height={30} + focusedPane="list" + />, + ); + stdin.write("v"); + await tick(); + const frame = lastFrame() ?? ""; + expect(frame).toContain("✓0 ✗2"); + expect(frame).toContain("Verification FAILED"); + }); + + it("drops verdicts for entries a refresh removed", async () => { + const { lastFrame, stdin, rerender } = render( + <SkillsTab + skills={[clean, other]} + pageCount={1} + inspectorClient={mockClient()} + width={140} + height={30} + focusedPane="list" + />, + ); + stdin.write("v"); + await tick(); + expect(lastFrame() ?? "").toContain("✓0 ✗2"); + // `other` is gone; `clean` is re-verified, which is when pruning happens. + const tree = ( + <SkillsTab + skills={[clean]} + pageCount={1} + inspectorClient={mockClient()} + width={140} + height={30} + focusedPane="list" + /> + ); + rerender(tree); + await tick(); + stdin.write("v"); + await tick(); + rerender( + <SkillsTab + skills={[clean, other]} + pageCount={1} + inspectorClient={mockClient()} + width={140} + height={30} + focusedPane="list" + />, + ); + await tick(); + // Had `other`'s old verdict survived, it would be counted again here. + expect(lastFrame() ?? "").toContain("✓0 ✗1 ·1"); + }); + + it("fits the tally in an 80-column terminal's list pane", async () => { + // App hands the pane 56 columns at an 80-column terminal. + const { lastFrame, stdin } = render( + <SkillsTab + skills={[clean, other, dynamic]} + pageCount={1} + inspectorClient={budgetedClient( + vi.fn().mockResolvedValue({ + result: { + contents: [{ uri: "skill://clean/SKILL.md", text: SKILL_MD }], + }, + }), + { skillCatalogMaxSkills: 1 }, + )} + width={56} + height={30} + focusedPane="list" + />, + ); + stdin.write("v"); + await tick(); + expect(lastFrame() ?? "").toMatch(/✓0 ✗1 …\d/); + }); + + it("does not cache a result for an entry removed while the run was in flight", async () => { + let release: () => void = () => {}; + const gate = new Promise<void>((resolve) => { + release = resolve; + }); + const readResource = vi.fn(async (uri: string) => { + await gate; + return { result: { contents: [{ uri, text: SKILL_MD }] } }; + }); + const client = mockClient(readResource); + const pane = (list: SkillEntry[]) => ( + <SkillsTab + skills={list} + pageCount={1} + inspectorClient={client} + width={140} + height={30} + focusedPane="list" + /> + ); + const { lastFrame, stdin, rerender } = render(pane([clean, other])); + stdin.write("v"); + await tick(); + // A refresh drops `other` before the pending run resolves. + rerender(pane([clean])); + await tick(); + release(); + await tick(); + // `other` returns: had its in-flight verdict been cached, it would count. + rerender(pane([clean, other])); + await tick(); + expect(lastFrame() ?? "").toContain("✓0 ✗1 ·1"); + }); +}); diff --git a/clients/tui/__tests__/SubscriptionsTab.test.tsx b/clients/tui/__tests__/SubscriptionsTab.test.tsx new file mode 100644 index 0000000000..236551a875 --- /dev/null +++ b/clients/tui/__tests__/SubscriptionsTab.test.tsx @@ -0,0 +1,283 @@ +import React from "react"; +import { describe, it, expect, vi } from "vitest"; +import { render } from "./helpers/renderTui"; +import type { Resource } from "@modelcontextprotocol/client"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import type { + InspectorResourceSubscription, + MessageEntry, + ResourceSubscriptionStreamState, +} from "@inspector/core/mcp/types.js"; +import { INACTIVE_SUBSCRIPTION_STREAM_STATE } from "@inspector/core/mcp/types.js"; +import { AuthRecoveryRequiredError } from "@inspector/core/auth/challenge.js"; + +const CHALLENGE = { reason: "insufficient_scope" as const }; + +vi.mock("ink-scroll-view", () => import("./helpers/inkScrollViewMock.js")); + +import { SubscriptionsTab } from "../src/components/SubscriptionsTab.js"; + +const tick = async () => { + for (let i = 0; i < 8; i++) + await new Promise((resolve) => setTimeout(resolve, 4)); +}; +const ESC = String.fromCharCode(27); +const UP = `${ESC}[A`; +const DOWN = `${ESC}[B`; +const PAGE_UP = `${ESC}[5~`; +const PAGE_DOWN = `${ESC}[6~`; + +const resA: Resource = { uri: "file:///a", name: "alpha" }; +const resB: Resource = { uri: "file:///b", name: "" }; + +const updated = (id: string, uri: string): MessageEntry => ({ + id, + timestamp: new Date(Date.UTC(2026, 0, 1, 12, 0, 0)), + direction: "notification", + message: { + jsonrpc: "2.0", + method: "notifications/resources/updated", + params: { uri }, + }, +}); + +interface FakeOps { + subscribeToResource: ReturnType<typeof vi.fn>; + unsubscribeFromResource: ReturnType<typeof vi.fn>; +} + +// A deliberately partial fake: SubscriptionsTab calls only the two methods in +// `FakeOps` on the client, and InspectorClient is a class with private members +// that no structural object literal can satisfy, so a single `as` is refused. +// The double cast is confined to this factory, and `FakeOps` keeps the methods +// the tests assert on typed. +const fakeClient = (over: Partial<FakeOps> = {}): FakeOps & InspectorClient => + ({ + subscribeToResource: vi.fn(async () => {}), + unsubscribeFromResource: vi.fn(async () => {}), + ...over, + }) as unknown as FakeOps & InspectorClient; + +function renderTab( + props: Partial<React.ComponentProps<typeof SubscriptionsTab>> = {}, +) { + const client = fakeClient(); + const api = render( + <SubscriptionsTab + resources={[resA, resB]} + subscriptions={[]} + streamState={INACTIVE_SUBSCRIPTION_STREAM_STATE} + messages={[]} + inspectorClient={client} + width={100} + height={30} + focusedPane="list" + {...props} + />, + ); + return { ...api, client }; +} + +const subscribedA: InspectorResourceSubscription[] = [ + { resource: resA, lastUpdated: new Date(Date.UTC(2026, 0, 1, 12, 0, 0)) }, +]; + +describe("SubscriptionsTab", () => { + it("shows an empty state", () => { + const { lastFrame } = renderTab({ resources: [] }); + const frame = lastFrame() ?? ""; + expect(frame).toContain("Subscriptions (0/0)"); + expect(frame).toContain("No resources to subscribe to"); + expect(frame).toContain("Resource updates"); + expect(frame).toContain("No resources/updated notifications yet"); + expect(frame).not.toContain("Enter to subscribe"); + }); + + it("lists resources with subscription state and the update feed", () => { + const { lastFrame } = renderTab({ + subscriptions: subscribedA, + messages: [updated("m1", "file:///a"), updated("m2", "file:///b")], + }); + const frame = lastFrame() ?? ""; + expect(frame).toContain("Subscriptions (1/2)"); + expect(frame).toContain("● alpha"); + // A resource with an empty name falls back to its URI. + expect(frame).toContain("○ file:///b"); + expect(frame).toContain("Subscribed"); + expect(frame).toContain("last updated"); + expect(frame).toContain("Updates (2):"); + expect(frame).toContain("Enter to unsubscribe"); + }); + + it("subscribes on Enter", async () => { + const { stdin, client } = renderTab(); + stdin.write("\r"); + await tick(); + expect(client.subscribeToResource).toHaveBeenCalledWith("file:///a"); + }); + + it("unsubscribes on Enter when subscribed", async () => { + const { stdin, client } = renderTab({ subscriptions: subscribedA }); + stdin.write("\r"); + await tick(); + expect(client.unsubscribeFromResource).toHaveBeenCalledWith("file:///a"); + }); + + it("shows a pending toggle and ignores a second Enter until it settles", async () => { + let release: () => void = () => {}; + const pending = new Promise<void>((resolve) => { + release = resolve; + }); + const client = fakeClient({ subscribeToResource: vi.fn(() => pending) }); + const { stdin, lastFrame } = renderTab({ inspectorClient: client }); + // Two presses in one burst, before any re-render, start one request. + stdin.write("\r"); + stdin.write("\r"); + await tick(); + expect(lastFrame()).toContain("Updating subscription…"); + stdin.write("\r"); + await tick(); + expect(client.subscribeToResource).toHaveBeenCalledTimes(1); + release(); + await tick(); + expect(lastFrame()).not.toContain("Updating subscription…"); + }); + + it("surfaces a subscribe failure", async () => { + const client = fakeClient({ + subscribeToResource: vi.fn(async () => { + throw new Error("Server does not support resource subscriptions"); + }), + }); + const { stdin, lastFrame } = renderTab({ inspectorClient: client }); + stdin.write("\r"); + await tick(); + expect(lastFrame()).toContain("does not support resource subscriptions"); + }); + + it("redacts URL query secrets in a surfaced failure (#2638)", async () => { + const client = fakeClient({ + subscribeToResource: vi.fn(async () => { + throw new Error( + "Request failed: https://auth.example/cb?code=s3cret&state=ok", + ); + }), + }); + const { stdin, lastFrame } = renderTab({ inspectorClient: client }); + stdin.write("\r"); + await tick(); + const frame = (lastFrame() ?? "").replace(/\s+/g, ""); + expect(frame).toContain("code=%5BREDACTED%5D&state=ok"); + expect(frame).not.toContain("s3cret"); + }); + + it("surfaces a non-Error failure", async () => { + const client = fakeClient({ + subscribeToResource: vi.fn(() => Promise.reject("nope")), + }); + const { stdin, lastFrame } = renderTab({ inspectorClient: client }); + stdin.write("\r"); + await tick(); + expect(lastFrame()).toContain("nope"); + }); + + it("hands auth recovery to the caller, and tolerates no handler", async () => { + const recovery = new AuthRecoveryRequiredError( + new URL("https://auth.example/start"), + CHALLENGE, + ); + const client = fakeClient({ + // Thrown the way the real client throws it: wrapped, with the recovery + // error as `cause` (inspectorClient.ts subscribeToResource). + subscribeToResource: vi.fn(async () => { + throw new Error("Failed to subscribe to resource: auth", { + cause: recovery, + }); + }), + }); + const onAuthRecoveryRequired = vi.fn(); + const first = renderTab({ + inspectorClient: client, + onAuthRecoveryRequired, + }); + first.stdin.write("\r"); + await tick(); + expect(onAuthRecoveryRequired).toHaveBeenCalledWith(recovery); + first.unmount(); + + const second = renderTab({ inspectorClient: client }); + second.stdin.write("\r"); + await tick(); + expect(second.lastFrame()).toContain("Not subscribed"); + }); + + it("does nothing without a client", async () => { + const { stdin, lastFrame } = renderTab({ inspectorClient: null }); + stdin.write("\r"); + await tick(); + expect(lastFrame()).toContain("Not subscribed"); + }); + + it("navigates the list and highlights the selection's updates", async () => { + const { stdin, lastFrame } = renderTab({ + messages: [updated("m1", "file:///b")], + }); + stdin.write(DOWN); + await tick(); + expect(lastFrame()).toContain("URI: file:///b"); + stdin.write(DOWN); // at the end + await tick(); + stdin.write(UP); + await tick(); + expect(lastFrame()).toContain("URI: file:///a"); + stdin.write(UP); // at the top + await tick(); + expect(lastFrame()).toContain("URI: file:///a"); + }); + + it("scrolls the details pane", async () => { + const { stdin, lastFrame } = renderTab({ focusedPane: "details" }); + for (const k of [UP, DOWN, PAGE_UP, PAGE_DOWN, "z"]) { + stdin.write(k); + await tick(); + } + expect(lastFrame()).toContain("URI: file:///a"); + }); + + it("shows the modern listen stream and an unhonored URI", () => { + const streamState: ResourceSubscriptionStreamState = { + active: true, + status: "acknowledged", + honoredUris: [], + }; + const { lastFrame } = renderTab({ + subscriptions: subscribedA, + streamState, + }); + const frame = lastFrame() ?? ""; + expect(frame).toContain("Listen stream: acknowledged"); + expect(frame).toContain("did not honor"); + }); + + it("does not flag a URI the server honored", () => { + const { lastFrame } = renderTab({ + subscriptions: [{ resource: resA }], + streamState: { + active: true, + status: "reconnecting", + honoredUris: ["file:///a"], + }, + }); + const frame = lastFrame() ?? ""; + expect(frame).toContain("Listen stream: reconnecting"); + expect(frame).not.toContain("did not honor"); + expect(frame).not.toContain("last updated"); + }); + + it("ignores keys while a modal is open", async () => { + const { stdin, client } = renderTab({ modalOpen: true }); + stdin.write("\r"); + await tick(); + expect(client.subscribeToResource).not.toHaveBeenCalled(); + }); +}); diff --git a/clients/tui/__tests__/Tabs.test.tsx b/clients/tui/__tests__/Tabs.test.tsx index 8854d12660..6c3ca965fd 100644 --- a/clients/tui/__tests__/Tabs.test.tsx +++ b/clients/tui/__tests__/Tabs.test.tsx @@ -69,6 +69,26 @@ describe("Tabs", () => { expect(shown.lastFrame() ?? "").toContain("Skills"); }); + it("hides Subscriptions and Tasks by default and shows them when supported (#2432)", () => { + const hidden = render( + <Tabs activeTab="info" onTabChange={noop} width={160} />, + ); + expect(hidden.lastFrame() ?? "").not.toContain("Subscriptions"); + expect(hidden.lastFrame() ?? "").not.toContain("Tasks"); + const shown = render( + <Tabs + activeTab="tasks" + onTabChange={noop} + width={160} + showSubscriptions={true} + showTasks={true} + counts={{ subscriptions: 2, tasks: 3 }} + />, + ); + expect(shown.lastFrame() ?? "").toContain("Subscriptions (2)"); + expect(shown.lastFrame() ?? "").toContain("Tasks (3)"); + }); + it("renders a count on the skills tab", () => { const { lastFrame } = render( <Tabs diff --git a/clients/tui/__tests__/TasksTab.test.tsx b/clients/tui/__tests__/TasksTab.test.tsx new file mode 100644 index 0000000000..43d7f2c984 --- /dev/null +++ b/clients/tui/__tests__/TasksTab.test.tsx @@ -0,0 +1,349 @@ +import React from "react"; +import { describe, it, expect, vi } from "vitest"; +import { render } from "./helpers/renderTui"; +import type { Task } from "@modelcontextprotocol/client"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import { AuthRecoveryRequiredError } from "@inspector/core/auth/challenge.js"; + +const CHALLENGE = { reason: "insufficient_scope" as const }; + +vi.mock("ink-scroll-view", () => import("./helpers/inkScrollViewMock.js")); + +import { + TasksTab, + hasTaskResult, + isTaskActive, + taskStatusStyle, +} from "../src/components/TasksTab.js"; + +const tick = async () => { + for (let i = 0; i < 8; i++) + await new Promise((resolve) => setTimeout(resolve, 4)); +}; +const ESC = String.fromCharCode(27); +const UP = `${ESC}[A`; +const DOWN = `${ESC}[B`; +const PAGE_UP = `${ESC}[5~`; +const PAGE_DOWN = `${ESC}[6~`; + +const task = (over: Partial<Task> = {}): Task => + ({ + taskId: "task-1", + status: "working", + createdAt: "2026-01-01T00:00:00Z", + lastUpdatedAt: "2026-01-01T00:00:01Z", + ttl: 60000, + ...over, + }) as Task; + +interface FakeOps { + cancelRequestorTask: ReturnType<typeof vi.fn>; + getRequestorTaskResult: ReturnType<typeof vi.fn>; +} + +// A deliberately partial fake: TasksTab calls only the two methods in +// `FakeOps` on the client, and InspectorClient is a class with private members +// that no structural object literal can satisfy, so a single `as` is refused. +// The double cast is confined to this factory, and `FakeOps` keeps the methods +// the tests assert on typed. +const fakeClient = (over: Partial<FakeOps> = {}): FakeOps & InspectorClient => + ({ + cancelRequestorTask: vi.fn(async () => {}), + getRequestorTaskResult: vi.fn(async () => ({ + content: [{ type: "text", text: "done" }], + })), + ...over, + }) as unknown as FakeOps & InspectorClient; + +function renderTab(props: Partial<React.ComponentProps<typeof TasksTab>> = {}) { + const onRefresh = vi.fn(async () => []); + const onClearCompleted = vi.fn(); + const client = fakeClient(); + const api = render( + <TasksTab + tasks={[task()]} + inspectorClient={client} + width={100} + height={30} + focusedPane="list" + onRefresh={onRefresh} + onClearCompleted={onClearCompleted} + {...props} + />, + ); + return { ...api, onRefresh, onClearCompleted, client }; +} + +describe("task status helpers", () => { + it("styles known and unknown statuses", () => { + expect(taskStatusStyle("completed").color).toBe("green"); + expect(taskStatusStyle("mystery")).toEqual({ glyph: "·", color: "gray" }); + }); + + it("classifies active and result-bearing statuses", () => { + expect(isTaskActive("working")).toBe(true); + expect(isTaskActive("input_required")).toBe(true); + expect(isTaskActive("completed")).toBe(false); + expect(hasTaskResult("completed")).toBe(true); + expect(hasTaskResult("failed")).toBe(true); + expect(hasTaskResult("cancelled")).toBe(false); + }); +}); + +describe("TasksTab", () => { + it("shows an empty state", () => { + const { lastFrame } = renderTab({ tasks: [] }); + expect(lastFrame()).toContain("Tasks (0)"); + expect(lastFrame()).toContain("No tasks yet"); + expect(lastFrame()).toContain("Select a task to view details"); + }); + + it("renders the selected task's details and footer", () => { + const { lastFrame } = renderTab({ + tasks: [task({ statusMessage: "halfway", pollInterval: 500 })], + }); + const frame = lastFrame() ?? ""; + expect(frame).toContain("Tasks (1)"); + expect(frame).toContain("Status: working"); + expect(frame).toContain("halfway"); + expect(frame).toContain("Updated: 2026-01-01T00:00:01Z"); + expect(frame).toContain("TTL: 60000 Poll: 500ms"); + expect(frame).toContain("x cancel"); + }); + + it("renders a task without optional fields", () => { + const { lastFrame } = renderTab({ + tasks: [task({ ttl: null, lastUpdatedAt: undefined })], + focusedPane: "details", + }); + const frame = lastFrame() ?? ""; + expect(frame).toContain("TTL: none"); + expect(frame).not.toContain("Updated:"); + expect(frame).toContain("+ zoom"); + }); + + it("navigates the list with the arrow keys", async () => { + const { lastFrame, stdin } = renderTab({ + tasks: [task(), task({ taskId: "task-2", status: "failed" })], + }); + stdin.write(DOWN); + await tick(); + expect(lastFrame()).toContain("Status: failed"); + stdin.write(DOWN); // already at the end + await tick(); + stdin.write(UP); + await tick(); + expect(lastFrame()).toContain("Status: working"); + stdin.write(UP); // already at the top + await tick(); + expect(lastFrame()).toContain("Status: working"); + }); + + it("refreshes on 'f' and clears completed on 'l'", async () => { + const { stdin, onRefresh, onClearCompleted } = renderTab(); + stdin.write("f"); + await tick(); + stdin.write("l"); + await tick(); + expect(onRefresh).toHaveBeenCalledTimes(1); + expect(onClearCompleted).toHaveBeenCalledTimes(1); + }); + + it("surfaces a refresh failure", async () => { + const { stdin, lastFrame } = renderTab({ + onRefresh: vi.fn(async () => { + throw new Error("list failed"); + }), + }); + stdin.write("f"); + await tick(); + expect(lastFrame()).toContain("list failed"); + }); + + it("redacts URL query secrets in a surfaced failure (#2638)", async () => { + const { stdin, lastFrame } = renderTab({ + onRefresh: vi.fn(async () => { + throw new Error( + "Request failed: https://auth.example/cb?code=s3cret&state=ok", + ); + }), + }); + stdin.write("f"); + await tick(); + const frame = (lastFrame() ?? "").replace(/\s+/g, ""); + expect(frame).toContain("code=%5BREDACTED%5D&state=ok"); + expect(frame).not.toContain("s3cret"); + }); + + it("surfaces a non-Error failure with no task selected", async () => { + const { stdin, lastFrame } = renderTab({ + tasks: [], + onRefresh: vi.fn(() => Promise.reject("plain failure")), + }); + stdin.write("f"); + await tick(); + expect(lastFrame()).toContain("plain failure"); + }); + + it("shows the busy line while an operation is pending", async () => { + let release: () => void = () => {}; + const pending = new Promise<never[]>((resolve) => { + release = () => resolve([]); + }); + const onRefresh = vi.fn(() => pending); + const { stdin, lastFrame } = renderTab({ onRefresh }); + stdin.write("f"); + await tick(); + expect(lastFrame()).toContain("Refreshing…"); + // A second press while the first is in flight starts nothing. + stdin.write("f"); + stdin.write("x"); + stdin.write("l"); + await tick(); + expect(onRefresh).toHaveBeenCalledTimes(1); + release(); + await tick(); + expect(lastFrame()).not.toContain("Refreshing…"); + }); + + it("shows the busy line in the empty state too", async () => { + let release: () => void = () => {}; + const pending = new Promise<never[]>((resolve) => { + release = () => resolve([]); + }); + const { stdin, lastFrame } = renderTab({ + tasks: [], + onRefresh: () => pending, + }); + stdin.write("f"); + await tick(); + expect(lastFrame()).toContain("Refreshing…"); + release(); + await tick(); + }); + + it("cancels an active task on 'x'", async () => { + const client = fakeClient(); + const { stdin } = renderTab({ inspectorClient: client }); + stdin.write("x"); + await tick(); + expect(client.cancelRequestorTask).toHaveBeenCalledWith("task-1"); + }); + + it("does not cancel a terminal task", async () => { + const client = fakeClient(); + const { stdin } = renderTab({ + inspectorClient: client, + tasks: [task({ status: "completed" })], + }); + stdin.write("x"); + await tick(); + expect(client.cancelRequestorTask).not.toHaveBeenCalled(); + }); + + it("fetches and shows a terminal task's result on Enter, and zooms it", async () => { + const client = fakeClient(); + const onViewDetails = vi.fn(); + const { stdin, lastFrame, rerender, onRefresh, onClearCompleted } = + renderTab({ + inspectorClient: client, + tasks: [task({ status: "completed" })], + }); + stdin.write("\r"); + await tick(); + expect(client.getRequestorTaskResult).toHaveBeenCalledWith("task-1"); + expect(lastFrame()).toContain("Result:"); + expect(lastFrame()).toContain("done"); + + rerender( + <TasksTab + tasks={[task({ status: "completed" })]} + inspectorClient={client} + width={100} + height={30} + focusedPane="details" + onRefresh={onRefresh} + onClearCompleted={onClearCompleted} + onViewDetails={onViewDetails} + />, + ); + await tick(); + stdin.write("+"); + await tick(); + expect(onViewDetails).toHaveBeenCalledWith( + expect.objectContaining({ taskId: "task-1" }), + { content: [{ type: "text", text: "done" }] }, + ); + }); + + it("does not fetch a result for a task still running", async () => { + const client = fakeClient(); + const { stdin } = renderTab({ inspectorClient: client }); + stdin.write("\r"); + await tick(); + expect(client.getRequestorTaskResult).not.toHaveBeenCalled(); + }); + + it("hands auth recovery to the caller", async () => { + const recovery = new AuthRecoveryRequiredError( + new URL("https://auth.example/start"), + CHALLENGE, + ); + const client = fakeClient({ + cancelRequestorTask: vi.fn(async () => { + throw recovery; + }), + }); + const onAuthRecoveryRequired = vi.fn(); + const { stdin, lastFrame } = renderTab({ + inspectorClient: client, + onAuthRecoveryRequired, + }); + stdin.write("x"); + await tick(); + expect(onAuthRecoveryRequired).toHaveBeenCalledWith(recovery); + expect(lastFrame()).not.toContain("Error"); + }); + + it("tolerates auth recovery with no handler", async () => { + const client = fakeClient({ + cancelRequestorTask: vi.fn(async () => { + throw new AuthRecoveryRequiredError( + new URL("https://auth.example"), + CHALLENGE, + ); + }), + }); + const { stdin, lastFrame } = renderTab({ inspectorClient: client }); + stdin.write("x"); + await tick(); + expect(lastFrame()).toContain("Status: working"); + }); + + it("scrolls the details pane and ignores zoom without a handler", async () => { + const { stdin, lastFrame } = renderTab({ focusedPane: "details" }); + for (const k of [UP, DOWN, PAGE_UP, PAGE_DOWN, "+", "z"]) { + stdin.write(k); + await tick(); + } + expect(lastFrame()).toContain("Status: working"); + }); + + it("does nothing without a client", async () => { + const { stdin, lastFrame } = renderTab({ + inspectorClient: null, + tasks: [task({ status: "completed" })], + }); + stdin.write("x"); + stdin.write("\r"); + await tick(); + expect(lastFrame()).not.toContain("Result:"); + }); + + it("ignores keys while a modal is open", async () => { + const { stdin, onRefresh } = renderTab({ modalOpen: true }); + stdin.write("f"); + await tick(); + expect(onRefresh).not.toHaveBeenCalled(); + }); +}); diff --git a/clients/tui/__tests__/ToolTestModal.saveRace.test.tsx b/clients/tui/__tests__/ToolTestModal.saveRace.test.tsx new file mode 100644 index 0000000000..477d428fd0 --- /dev/null +++ b/clients/tui/__tests__/ToolTestModal.saveRace.test.tsx @@ -0,0 +1,223 @@ +import React from "react"; +import { describe, it, expect, vi, afterEach } from "vitest"; +import { render } from "./helpers/renderTui"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import type { Tool } from "@modelcontextprotocol/client"; +import type { + SavePrompt, + SaveStatus, +} from "../src/components/SaveResultBar.js"; +import type { SavedResult } from "../src/utils/saveResult.js"; + +vi.mock("ink-scroll-view", () => import("./helpers/inkScrollViewMock.js")); +vi.mock("ink-form", () => import("./helpers/inkFormMock.js")); + +// Each save resolves only when the test says so, so a write can be held in +// flight deterministically. +const pending: Array<{ + resolve: (v: SavedResult) => void; + reject: (e: unknown) => void; +}> = []; +vi.mock("../src/utils/saveResult.js", () => ({ + defaultResultFileName: () => "alpha-result.json", + saveResultToFile: () => + new Promise<SavedResult>((resolve, reject) => { + pending.push({ resolve, reject }); + }), +})); + +// The modal's frame is empty under ink-testing-library (position="absolute"), +// so the status it hands the bar is captured from the bar's props instead. +const statuses: Array<SaveStatus | null> = []; +const prompts: Array<SavePrompt | null> = []; +vi.mock("../src/components/SaveResultBar.js", () => ({ + SaveResultBar: ({ + prompt, + status, + }: { + prompt: SavePrompt | null; + status: SaveStatus | null; + }) => { + prompts.push(prompt); + statuses.push(status); + return null; + }, +})); + +import { ToolTestModal } from "../src/components/ToolTestModal.js"; + +// Double cast: the modal only calls `callTool`, and a full InspectorClient is +// a class with private state no structural fake can satisfy. +const client = (callTool: unknown) => + ({ callTool }) as unknown as InspectorClient; + +// Condition waits rather than fixed sleeps: each step waits for the state it +// needs, bounded so a regression fails instead of hanging. +const waitUntil = async (predicate: () => boolean) => { + for (let i = 0; i < 500 && !predicate(); i++) + await new Promise((resolve) => setTimeout(resolve, 4)); + expect(predicate()).toBe(true); +}; +const lastStatus = () => statuses.at(-1); +let resultSettledAt: number | null = null; +const promptOpen = () => prompts.at(-1) != null; + +type Api = ReturnType<typeof render>; + +// Submit the form and open the save prompt. Enter is re-pressed until the call +// goes out, since the form ignores it until its input handler has subscribed. +// `w` is pressed once, only after the modal has re-rendered with the result — +// from then on its input handler sees the result view, so the key cannot be +// lost (and re-pressing it could type a stray `w` into the path). +const submitAndOpenPrompt = async (api: Api, callTool: () => unknown) => { + await waitUntil(() => { + if (vi.mocked(callTool).mock.calls.length === 0) api.stdin.write("\r"); + return vi.mocked(callTool).mock.calls.length > 0; + }); + await waitUntil( + () => resultSettledAt !== null && statuses.length > resultSettledAt, + ); + api.stdin.write("w"); + await waitUntil(promptOpen); +}; + +const renderModal = () => { + // Notes how many renders had happened when the call settled: the modal's + // own continuation (which moves it to the result view) runs after this one, + // and both are microtasks, so the next render shows the result view. + const callTool = vi.fn(() => { + const settled = Promise.resolve({ + success: true, + result: { content: [{ type: "text", text: "hello" }] }, + }); + void settled.then(() => { + resultSettledAt = statuses.length; + }); + return settled; + }); + const api = render( + <ToolTestModal + tool={{ name: "alpha", inputSchema: { type: "object" } } as Tool} + inspectorClient={client(callTool)} + width={80} + height={24} + onClose={vi.fn()} + />, + ); + return { api, callTool }; +}; + +afterEach(() => { + resultSettledAt = null; + pending.length = 0; + statuses.length = 0; + prompts.length = 0; + delete (globalThis as Record<string, unknown>).__INK_FORM_SUBMIT_VALUE__; +}); + +describe("ToolTestModal save serialization (#2571)", () => { + it("shows progress, then a confirmation with the right byte unit", async () => { + const { api, callTool } = renderModal(); + await submitAndOpenPrompt(api, callTool); + api.stdin.write("\r"); + await waitUntil(() => pending.length === 1); + await waitUntil( + () => lastStatus()?.message === "Saving to alpha-result.json…", + ); + pending[0]!.resolve({ path: "/a.json", format: "json", bytes: 1 }); + await waitUntil( + () => lastStatus()?.message === "Saved json result to /a.json (1 byte)", + ); + // A second save from the same view, reported with the plural unit. + api.stdin.write("w"); + await waitUntil(promptOpen); + api.stdin.write("\r"); + await waitUntil(() => pending.length === 2); + pending[1]!.resolve({ path: "/a.json", format: "json", bytes: 2 }); + await waitUntil( + () => lastStatus()?.message === "Saved json result to /a.json (2 bytes)", + ); + api.unmount(); + }); + + it("an older save settling does not replace the newer one's status", async () => { + const { api, callTool } = renderModal(); + await submitAndOpenPrompt(api, callTool); + api.stdin.write("\r"); + await waitUntil(() => pending.length === 1); + api.stdin.write("w"); + await waitUntil(promptOpen); + // Edit the path so the two saves are told apart in the status. + for (const len of [16, 15, 14, 13]) { + api.stdin.write("\b"); + await waitUntil(() => prompts.at(-1)?.path.length === len); + } + api.stdin.write("newer"); + await waitUntil(() => prompts.at(-1)?.path === "alpha-result.newer"); + api.stdin.write("\r"); + await waitUntil(() => pending.length === 2); + await waitUntil( + () => lastStatus()?.message === "Saving to alpha-result.newer…", + ); + pending[0]!.resolve({ path: "/old.json", format: "json", bytes: 3 }); + // A negative assertion has no condition to wait on, so the stale + // completion gets a few macrotasks to land (it would within one). + for (let i = 0; i < 10; i++) await new Promise((r) => setTimeout(r, 4)); + expect(lastStatus()?.message).toBe("Saving to alpha-result.newer…"); + pending[1]!.resolve({ path: "/new.json", format: "json", bytes: 2 }); + await waitUntil( + () => + lastStatus()?.message === "Saved json result to /new.json (2 bytes)", + ); + api.unmount(); + }); + + it("Ctrl+W and Meta+W do not open the save prompt", async () => { + const { api, callTool } = renderModal(); + await submitAndOpenPrompt(api, callTool); + api.stdin.write("\u001b"); // close the prompt; the result view stays up + await waitUntil(() => !promptOpen()); + // Negative assertions: each chord gets a few macrotasks to open the + // prompt (it would within one) before checking it did not. + const settle = async () => { + for (let i = 0; i < 10; i++) await new Promise((r) => setTimeout(r, 4)); + }; + api.stdin.write("\u0017"); // Ctrl+W + await settle(); + expect(promptOpen()).toBe(false); + api.stdin.write("\u001bw"); // Meta+W + await settle(); + expect(promptOpen()).toBe(false); + // A plain w still opens it. + api.stdin.write("w"); + await waitUntil(promptOpen); + api.unmount(); + }); + + it("applies keys typed in one burst to the prompt as it now stands", async () => { + const { api, callTool } = renderModal(); + await submitAndOpenPrompt(api, callTool); + // Written back to back, before React re-renders: each key must see the + // edit the previous one made, not the prompt from the last render. + api.stdin.write("\b"); + api.stdin.write("\b"); + api.stdin.write("\r"); + await waitUntil(() => pending.length === 1); + await waitUntil( + () => lastStatus()?.message === "Saving to alpha-result.js…", + ); + expect(pending).toHaveLength(1); + api.unmount(); + }); + + it("reports a non-Error rejection as text", async () => { + const { api, callTool } = renderModal(); + await submitAndOpenPrompt(api, callTool); + api.stdin.write("\r"); + await waitUntil(() => pending.length === 1); + pending[0]!.reject("plain string"); + await waitUntil(() => lastStatus()?.message === "plain string"); + expect(lastStatus()?.ok).toBe(false); + api.unmount(); + }); +}); diff --git a/clients/tui/__tests__/ToolTestModal.test.tsx b/clients/tui/__tests__/ToolTestModal.test.tsx index 1bf58296e5..696175c537 100644 --- a/clients/tui/__tests__/ToolTestModal.test.tsx +++ b/clients/tui/__tests__/ToolTestModal.test.tsx @@ -1,5 +1,8 @@ import React from "react"; -import { describe, it, expect, vi, afterEach } from "vitest"; +import { describe, it, expect, vi, afterEach, beforeEach } from "vitest"; +import { existsSync, mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; import { render } from "./helpers/renderTui"; import type { InspectorClient } from "@inspector/core/mcp/index.js"; import type { Tool } from "@modelcontextprotocol/client"; @@ -399,3 +402,196 @@ describe("ToolTestModal", () => { api.unmount(); }); }); + +// The frame is empty (see the note at the top), so these assert on what lands +// on disk. The prompt and status text themselves are asserted in +// SaveResultBar.test.tsx. +describe("ToolTestModal — w saves the result to a file (#2571)", () => { + const BACKSPACE = "\b"; + const DELETE = "\x7f"; + const TAB = "\t"; + let dir: string; + + beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), "tui-save-")); + // Relative save paths resolve against the launch directory. + vi.spyOn(process, "cwd").mockReturnValue(dir); + }); + + afterEach(() => { + rmSync(dir, { recursive: true, force: true }); + }); + + const okClient = (result: unknown) => + fakeClient(vi.fn().mockResolvedValue({ success: true, result })); + + const TEXT_RESULT = { content: [{ type: "text", text: "hello" }] }; + + // The write is async; wait for it to land rather than racing the fs call. + const waitForFile = async (path: string) => { + for (let i = 0; i < 50 && !existsSync(path); i++) await tick(); + await tick(); + }; + + const press = async (stdin: { write: (s: string) => void }, s: string) => { + stdin.write(s); + await tick(); + }; + + it("w then Enter writes the whole result as pretty JSON to <tool>-result.json", async () => { + const { stdin, onClose, unmount } = await renderAndSubmit( + okClient(TEXT_RESULT), + ); + await press(stdin, "w"); + await press(stdin, "\r"); + const target = join(dir, "alpha-result.json"); + await waitForFile(target); + expect(readFileSync(target, "utf8")).toBe( + JSON.stringify(TEXT_RESULT, null, 2) + "\n", + ); + expect(onClose).not.toHaveBeenCalled(); + unmount(); + }); + + it("Tab switches to raw and its default name; Tab again switches back", async () => { + const { stdin, unmount } = await renderAndSubmit(okClient(TEXT_RESULT)); + await press(stdin, "w"); + await press(stdin, TAB); + await press(stdin, TAB); + await press(stdin, TAB); + await press(stdin, "\r"); + const target = join(dir, "alpha-result.txt"); + await waitForFile(target); + expect(readFileSync(target, "utf8")).toBe("hello"); + unmount(); + }); + + it("a typed path is kept across a format switch, and backspace/delete edit it", async () => { + const { stdin, unmount } = await renderAndSubmit(okClient(TEXT_RESULT)); + await press(stdin, "w"); + // Clear "alpha-result.json" (17 chars), alternating the two erase keys. + for (let i = 0; i < 9; i++) { + await press(stdin, BACKSPACE); + await press(stdin, DELETE); + } + await press(stdin, "out.txtx"); + await press(stdin, BACKSPACE); + // A ctrl chord is not text and must not land in the path. + await press(stdin, "\x01"); + await press(stdin, TAB); + await press(stdin, "\r"); + const target = join(dir, "out.txt"); + await waitForFile(target); + expect(readFileSync(target, "utf8")).toBe("hello"); + unmount(); + }); + + it("ESC cancels the prompt without closing the modal; a second ESC closes it", async () => { + const { stdin, onClose, unmount } = await renderAndSubmit( + okClient(TEXT_RESULT), + ); + await press(stdin, "w"); + await press(stdin, ESC); + await tick(); + expect(onClose).not.toHaveBeenCalled(); + // The prompt is gone, so Enter no longer saves. + await press(stdin, "\r"); + await tick(); + expect(existsSync(join(dir, "alpha-result.json"))).toBe(false); + await press(stdin, ESC); + await tick(); + expect(onClose).toHaveBeenCalledTimes(1); + unmount(); + }); + + it("saves an isError result too — it is still the server's result", async () => { + const errResult = { + isError: true, + content: [{ type: "text", text: "oops" }], + }; + const { stdin, unmount } = await renderAndSubmit(okClient(errResult)); + await press(stdin, "w"); + await press(stdin, "\r"); + const target = join(dir, "alpha-result.json"); + await waitForFile(target); + expect(JSON.parse(readFileSync(target, "utf8"))).toEqual(errResult); + unmount(); + }); + + it("w with no result to save opens no prompt and writes nothing", async () => { + const callTool = vi + .fn() + .mockResolvedValue({ success: false, result: null, error: "boom" }); + const { stdin, onClose, unmount } = await renderAndSubmit( + fakeClient(callTool), + ); + await press(stdin, "w"); + await press(stdin, "\r"); + await tick(); + expect(existsSync(join(dir, "alpha-result.json"))).toBe(false); + expect(onClose).not.toHaveBeenCalled(); + unmount(); + }); + + it("a failed write is reported in the TUI, not thrown, and a retry still works", async () => { + const { stdin, onClose, unmount } = await renderAndSubmit( + okClient(TEXT_RESULT), + ); + await press(stdin, "w"); + // Point the default name into a directory that does not exist. + for (let i = 0; i < 17; i++) await press(stdin, BACKSPACE); + await press(stdin, "missing/dir/x.json"); + await press(stdin, "\r"); + await tick(); + await tick(); + expect(existsSync(join(dir, "missing"))).toBe(false); + expect(onClose).not.toHaveBeenCalled(); + // The modal is still alive and the default comes back on the next w. + await press(stdin, "w"); + await press(stdin, "\r"); + const target = join(dir, "alpha-result.json"); + await waitForFile(target); + expect(existsSync(target)).toBe(true); + unmount(); + }); + + it("a result with no raw form is refused when saved as raw", async () => { + const linkOnly = { + content: [{ type: "resource_link", uri: "x://a", name: "a" }], + }; + const { stdin, onClose, unmount } = await renderAndSubmit( + okClient(linkOnly), + ); + await press(stdin, "w"); + await press(stdin, TAB); + await press(stdin, "\r"); + await tick(); + await tick(); + expect(existsSync(join(dir, "alpha-result.txt"))).toBe(false); + expect(onClose).not.toHaveBeenCalled(); + unmount(); + }); + + it("falls back to a 'tool' file name when the tool has no name", async () => { + const onClose = vi.fn(); + const api = render( + <ToolTestModal + tool={makeTool({ name: "" })} + inspectorClient={okClient(TEXT_RESULT)} + width={80} + height={24} + onClose={onClose} + />, + ); + await tick(); + setSubmitValue({}); + await press(api.stdin, "\r"); + await tick(); + await press(api.stdin, "w"); + await press(api.stdin, "\r"); + const target = join(dir, "tool-result.json"); + await waitForFile(target); + expect(existsSync(target)).toBe(true); + api.unmount(); + }); +}); diff --git a/clients/tui/__tests__/ToolsTab.test.tsx b/clients/tui/__tests__/ToolsTab.test.tsx index 2871ab6369..b5a39cbc0e 100644 --- a/clients/tui/__tests__/ToolsTab.test.tsx +++ b/clients/tui/__tests__/ToolsTab.test.tsx @@ -286,3 +286,89 @@ describe("schemaMarker", () => { ).toEqual({ glyph: "!", color: "red" }); }); }); + +describe("ToolsTab list filter (#2430)", () => { + it("narrows the list by name or title, keeps fallback ordinals, and clears", async () => { + const onFilterEditingChange = vi.fn(); + const filterable: Tool[] = [ + makeTool({ name: "alpha", description: "Alpha desc" }), + makeTool({ name: "", title: "Gamma Title", description: "Gamma desc" }), + makeTool({ name: "delta", annotations: { title: "GAMMA-ish" } }), + ]; + const { lastFrame, stdin } = render( + <ToolsTab + tools={filterable} + isConnected={false} + width={120} + height={30} + focusedPane="list" + onFilterEditingChange={onFilterEditingChange} + />, + ); + await tick(); + expect(lastFrame()).toContain("/ to filter"); + for (const k of ["/", "g", "a", "m"]) { + stdin.write(k); + await tick(); + } + let frame = lastFrame() ?? ""; + expect(onFilterEditingChange).toHaveBeenLastCalledWith(true); + expect(frame).toContain("Tools (2/3)"); + // The untitled row keeps its unfiltered number. + expect(frame).toContain("Tool 2"); + expect(frame).toContain("Gamma desc"); + expect(frame).not.toContain("alpha"); + + // Enter keeps the filter rather than testing the selected tool; the + // arrows then move within the narrowed list. + stdin.write("\r"); + await tick(); + stdin.write(DOWN); + await tick(); + frame = lastFrame() ?? ""; + expect(frame).toContain("(/ to edit)"); + expect(frame).toContain("▶ delta"); + + // A query nothing matches. + for (const k of ["/", "z"]) { + stdin.write(k); + await tick(); + } + expect(lastFrame()).toContain("No tools match the filter"); + + stdin.write(ESC); + await tick(); + frame = lastFrame() ?? ""; + expect(frame).toContain("Tools (3)"); + expect(frame).toContain("alpha"); + expect(onFilterEditingChange).toHaveBeenLastCalledWith(false); + }); +}); + +describe("ToolsTab list filter — duplicate names (#2430)", () => { + it("renders both copies of a repeated tool name, filtered or not", async () => { + const dupes: Tool[] = [ + makeTool({ name: "dup", description: "first copy" }), + makeTool({ name: "other" }), + makeTool({ name: "dup", description: "second copy" }), + ]; + const { lastFrame, stdin } = render( + <ToolsTab + tools={dupes} + isConnected={false} + width={120} + height={30} + focusedPane="list" + />, + ); + await tick(); + for (const k of ["/", "d", "u", "p", "\r", DOWN]) { + stdin.write(k); + await tick(); + } + const frame = lastFrame() ?? ""; + expect(frame).toContain("Tools (2/3)"); + expect(frame.match(/dup/g)?.length).toBeGreaterThanOrEqual(2); + expect(frame).toContain("second copy"); + }); +}); diff --git a/clients/tui/__tests__/clipboard.test.ts b/clients/tui/__tests__/clipboard.test.ts new file mode 100644 index 0000000000..965070f72a --- /dev/null +++ b/clients/tui/__tests__/clipboard.test.ts @@ -0,0 +1,86 @@ +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { afterEach, describe, expect, it } from "vitest"; +import { + OSC52_LARGE_PAYLOAD_BYTES, + buildOsc52Sequence, + copyViaOsc52, + toCopyText, + writeCopyFile, +} from "../src/utils/clipboard.js"; + +describe("toCopyText", () => { + it("passes strings through verbatim", () => { + expect(toCopyText("raw value")).toBe("raw value"); + }); + + it("pretty-prints anything else as JSON", () => { + expect(toCopyText({ a: 1, b: [true] })).toBe( + JSON.stringify({ a: 1, b: [true] }, null, 2), + ); + }); + + it("falls back to String() when JSON has no representation", () => { + expect(toCopyText(undefined)).toBe("undefined"); + }); + + it("falls back to String() when JSON.stringify throws", () => { + const cyclic: Record<string, unknown> = {}; + cyclic.self = cyclic; + expect(toCopyText(cyclic)).toBe("[object Object]"); + expect(toCopyText(10n)).toBe("10"); + }); +}); + +describe("buildOsc52Sequence", () => { + it("frames the base64 UTF-8 payload as OSC 52 for the clipboard", () => { + const seq = buildOsc52Sequence("héllo"); + expect(seq).toBe( + `\u001b]52;c;${Buffer.from("héllo", "utf8").toString("base64")}\u0007`, + ); + }); +}); + +describe("copyViaOsc52", () => { + it("writes the sequence to the sink and reports its size", () => { + const written: string[] = []; + const result = copyViaOsc52("abc", { write: (s) => written.push(s) }); + expect(written).toEqual([buildOsc52Sequence("abc")]); + expect(result).toEqual({ chars: 3, payloadBytes: 4, large: false }); + }); + + it("flags a payload over the large-payload threshold", () => { + const text = "x".repeat(OSC52_LARGE_PAYLOAD_BYTES); + const result = copyViaOsc52(text, { write: () => undefined }); + expect(result.large).toBe(true); + expect(result.payloadBytes).toBeGreaterThan(OSC52_LARGE_PAYLOAD_BYTES); + }); +}); + +describe("writeCopyFile", () => { + const created: string[] = []; + afterEach(() => { + for (const dir of created.splice(0)) { + fs.rmSync(dir, { recursive: true, force: true }); + } + }); + + it("saves the text to an owner-only file in a fresh temp directory", () => { + const base = fs.mkdtempSync(path.join(os.tmpdir(), "clip-test-")); + created.push(base); + const file = writeCopyFile("secret-token", base); + expect(path.dirname(path.dirname(file))).toBe(base); + expect(fs.readFileSync(file, "utf8")).toBe("secret-token"); + if (process.platform !== "win32") { + expect(fs.statSync(file).mode & 0o777).toBe(0o600); + expect(fs.statSync(path.dirname(file)).mode & 0o777).toBe(0o700); + } + }); + + it("defaults to the OS temp directory", () => { + const file = writeCopyFile("v"); + created.push(path.dirname(file)); + expect(file.startsWith(os.tmpdir())).toBe(true); + }); +}); diff --git a/clients/tui/__tests__/errorText.test.ts b/clients/tui/__tests__/errorText.test.ts new file mode 100644 index 0000000000..275354bcc3 --- /dev/null +++ b/clients/tui/__tests__/errorText.test.ts @@ -0,0 +1,53 @@ +import { describe, it, expect } from "vitest"; +import { + errorMessage, + redactErrorText, + redactedJson, +} from "../src/utils/errorText.js"; + +// #2490: every error the TUI draws goes through these, so a server or SDK +// error quoting a URL with an OAuth code / token never reaches the screen. +const SECRET_URL = "https://srv.example/cb?code=abc123&state=ok"; +const REDACTED_URL = "https://srv.example/cb?code=%5BREDACTED%5D&state=ok"; + +describe("errorMessage", () => { + it("returns an Error's message, redacted", () => { + expect(errorMessage(new Error(`Callback ${SECRET_URL} failed.`))).toBe( + `Callback ${REDACTED_URL} failed.`, + ); + }); + + it("stringifies a non-Error, redacted", () => { + expect(errorMessage(`at ${SECRET_URL}`)).toBe(`at ${REDACTED_URL}`); + expect(errorMessage(42)).toBe("42"); + expect(errorMessage(undefined)).toBe("undefined"); + }); + + it("leaves text without a sensitive URL untouched", () => { + expect(errorMessage(new Error("plain failure"))).toBe("plain failure"); + }); +}); + +describe("redactErrorText", () => { + it("redacts query secrets in free text", () => { + expect(redactErrorText(`lost ${SECRET_URL}`)).toBe(`lost ${REDACTED_URL}`); + }); +}); + +describe("redactedJson", () => { + it("pretty-prints and redacts a URL inside a string value", () => { + const out = redactedJson({ message: `boom ${SECRET_URL}`, code: -32000 }); + expect(out).toBe( + JSON.stringify( + { message: `boom ${REDACTED_URL}`, code: -32000 }, + null, + 2, + ), + ); + expect(out).not.toContain("abc123"); + }); + + it("falls back to String() when JSON.stringify yields undefined", () => { + expect(redactedJson(undefined)).toBe("undefined"); + }); +}); diff --git a/clients/tui/__tests__/keybindings.test.ts b/clients/tui/__tests__/keybindings.test.ts new file mode 100644 index 0000000000..bedfcd9826 --- /dev/null +++ b/clients/tui/__tests__/keybindings.test.ts @@ -0,0 +1,78 @@ +import { describe, it, expect } from "vitest"; +import { tabs } from "../src/components/tabsConfig.js"; +import { + HELP_BINDINGS, + TAB_BINDINGS, + globalBindings, + keyColumnWidth, + keybindingSections, +} from "../src/utils/keybindings.js"; + +describe("keybindings", () => { + it("documents every tab with at least one binding", () => { + for (const tab of tabs) { + expect(TAB_BINDINGS[tab.id].length).toBeGreaterThan(0); + } + }); + + it("titles every tab's section with its tab-bar label", () => { + for (const id of Object.keys( + TAB_BINDINGS, + ) as (keyof typeof TAB_BINDINGS)[]) { + const label = tabs.find((tab) => tab.id === id)?.label; + expect(label).toBeDefined(); + expect(keybindingSections(id)[1]!.title).toBe(`${label} tab`); + } + }); + + it("lists every tab accelerator in the global section, from tabsConfig", () => { + const accelerators = globalBindings().find((b) => + b.action.includes("underlined letter"), + ); + expect(accelerators?.keys.split(" ")).toEqual( + tabs.map((tab) => tab.accelerator), + ); + }); + + it("lists only the accelerators of the tabs it is told are visible", () => { + const visible = tabs.filter( + (tab) => tab.id === "info" || tab.id === "tools", + ); + const sections = keybindingSections("tools", visible); + const row = sections[0]!.bindings.find((b) => + b.action.includes("underlined letter"), + ); + expect(row?.keys).toBe("i t"); + }); + + it("documents the help toggle itself", () => { + expect(globalBindings().some((b) => b.keys === "?")).toBe(true); + expect(HELP_BINDINGS.some((b) => b.keys.includes("Esc"))).toBe(true); + }); + + it("returns Global, the active tab's section by label, then the help section", () => { + const sections = keybindingSections("messages"); + expect(sections.map((s) => s.title)).toEqual([ + "Global", + "Protocol tab", + "This help", + ]); + expect(sections[1]!.bindings).toBe(TAB_BINDINGS.messages); + }); + + it("measures the widest keys column across all sections", () => { + expect( + keyColumnWidth([ + { title: "a", bindings: [{ keys: "ab", action: "x" }] }, + { + title: "b", + bindings: [ + { keys: "abcd", action: "y" }, + { keys: "a", action: "z" }, + ], + }, + ]), + ).toBe(4); + expect(keyColumnWidth([])).toBe(0); + }); +}); diff --git a/clients/tui/__tests__/openUrl.test.ts b/clients/tui/__tests__/openUrl.test.ts index 6c30c22e8e..4eea8d6109 100644 --- a/clients/tui/__tests__/openUrl.test.ts +++ b/clients/tui/__tests__/openUrl.test.ts @@ -1,16 +1,20 @@ import { describe, it, expect, vi, beforeEach } from "vitest"; -import open from "open"; -import { openUrl } from "../src/utils/openUrl.js"; +import { openUrl as openInBrowser } from "@inspector/core/node/openUrl.js"; +import { browserOpenFailedMessage, openUrl } from "../src/utils/openUrl.js"; -// `open` shells out to the OS browser opener — stub it so the test only checks -// that openUrl forwards the right string. -vi.mock("open", () => ({ default: vi.fn().mockResolvedValue(undefined) })); +// Core's opener spawns the OS browser — stub it so these tests only check what +// the TUI wrapper forwards and how it owns a failure. The spawn-failure +// mechanics themselves are tested against core (web suite, #2533). +vi.mock("@inspector/core/node/openUrl.js", () => ({ + openUrl: vi.fn().mockResolvedValue(undefined), +})); -const openMock = vi.mocked(open); +const openMock = vi.mocked(openInBrowser); describe("openUrl", () => { beforeEach(() => { - openMock.mockClear(); + openMock.mockReset(); + openMock.mockResolvedValue(undefined); }); it("passes a string URL straight through", async () => { @@ -24,4 +28,34 @@ describe("openUrl", () => { "https://example.com/callback?code=1", ); }); + + it("does not report a failure when the browser opens", async () => { + const onFailure = vi.fn(); + await openUrl("https://example.com/auth", onFailure); + expect(onFailure).not.toHaveBeenCalled(); + }); + + it("resolves and reports the manual-open note when the opener fails", async () => { + openMock.mockRejectedValue(new Error("spawn xdg-open ENOENT")); + const onFailure = vi.fn(); + await expect( + openUrl(new URL("https://example.com/auth?x=1"), onFailure), + ).resolves.toBeUndefined(); + expect(onFailure).toHaveBeenCalledWith( + browserOpenFailedMessage("https://example.com/auth?x=1"), + ); + }); + + it("swallows a failure even with no callback", async () => { + openMock.mockRejectedValue(new Error("spawn xdg-open ENOENT")); + await expect(openUrl("https://example.com/auth")).resolves.toBeUndefined(); + }); +}); + +describe("browserOpenFailedMessage", () => { + it("names the URL to open by hand", () => { + expect(browserOpenFailedMessage("https://example.com/auth")).toBe( + "Could not open a browser automatically. Open this URL manually to authorize: https://example.com/auth", + ); + }); }); diff --git a/clients/tui/__tests__/resourceUpdates.test.ts b/clients/tui/__tests__/resourceUpdates.test.ts new file mode 100644 index 0000000000..26bf68e66e --- /dev/null +++ b/clients/tui/__tests__/resourceUpdates.test.ts @@ -0,0 +1,92 @@ +import { describe, it, expect } from "vitest"; +import type { + InspectorResourceSubscription, + MessageEntry, +} from "@inspector/core/mcp/types.js"; +import { + RESOURCE_UPDATED_METHOD, + resourceUpdateFeed, + subscribableResources, +} from "../src/utils/resourceUpdates.js"; + +const at = (s: number) => new Date(Date.UTC(2026, 0, 1, 0, 0, s)); + +const notification = ( + id: string, + method: string, + params: Record<string, unknown> | undefined, + seconds: number, +): MessageEntry => ({ + id, + timestamp: at(seconds), + direction: "notification", + message: params + ? { jsonrpc: "2.0", method, params } + : { jsonrpc: "2.0", method }, +}); + +describe("resourceUpdateFeed", () => { + it("keeps only resources/updated notifications with a URI, newest first", () => { + const messages: MessageEntry[] = [ + notification("1", RESOURCE_UPDATED_METHOD, { uri: "file:///a" }, 1), + { + id: "2", + timestamp: at(2), + direction: "request", + message: { jsonrpc: "2.0", id: 1, method: RESOURCE_UPDATED_METHOD }, + }, + notification("3", "notifications/tools/list_changed", undefined, 3), + notification("4", RESOURCE_UPDATED_METHOD, undefined, 4), + notification("5", RESOURCE_UPDATED_METHOD, { uri: 42 }, 5), + { + id: "6", + timestamp: at(6), + direction: "response", + message: { jsonrpc: "2.0", id: 1, result: {} }, + }, + notification("7", RESOURCE_UPDATED_METHOD, { uri: "file:///b" }, 7), + ]; + expect(resourceUpdateFeed(messages)).toEqual([ + { id: "7", timestamp: at(7), uri: "file:///b" }, + { id: "1", timestamp: at(1), uri: "file:///a" }, + ]); + }); + + it("is empty for an empty log", () => { + expect(resourceUpdateFeed([])).toEqual([]); + }); +}); + +describe("subscribableResources", () => { + it("marks listed resources and appends unlisted subscriptions", () => { + const updated = at(9); + const subs: InspectorResourceSubscription[] = [ + { resource: { uri: "file:///a", name: "a" }, lastUpdated: updated }, + { resource: { uri: "file:///x", name: "file:///x" } }, + ]; + const rows = subscribableResources( + [ + { uri: "file:///a", name: "a" }, + { uri: "file:///b", name: "b" }, + ], + subs, + ); + expect(rows).toEqual([ + { + resource: { uri: "file:///a", name: "a" }, + subscribed: true, + lastUpdated: updated, + }, + { + resource: { uri: "file:///b", name: "b" }, + subscribed: false, + lastUpdated: undefined, + }, + { + resource: { uri: "file:///x", name: "file:///x" }, + subscribed: true, + lastUpdated: undefined, + }, + ]); + }); +}); diff --git a/clients/tui/__tests__/saveResult.queue.test.ts b/clients/tui/__tests__/saveResult.queue.test.ts new file mode 100644 index 0000000000..c004bf91d0 --- /dev/null +++ b/clients/tui/__tests__/saveResult.queue.test.ts @@ -0,0 +1,60 @@ +import { describe, it, expect, vi } from "vitest"; + +// A writeFile double whose calls settle only when the test says so, so the +// queue is observed directly: a second write must not BEGIN while the first +// is pending, whatever the filesystem's own timing would have been. +const writes: Array<{ + path: string; + resolve: () => void; + reject: (e: unknown) => void; +}> = []; +vi.mock("node:fs/promises", () => ({ + writeFile: (path: string) => + new Promise<void>((resolve, reject) => { + writes.push({ path, resolve, reject }); + }), +})); + +import { saveResultToFile } from "../src/utils/saveResult.js"; + +const TEXT = { content: [{ type: "text", text: "hi" }] }; + +const flush = async () => { + for (let i = 0; i < 5; i++) + await new Promise((resolve) => setImmediate(resolve)); +}; + +describe("saveResultToFile queueing (#2571)", () => { + it("starts each write only after the previous one has settled", async () => { + const order: string[] = []; + const first = saveResultToFile(TEXT, "a.json", "json", "/d").then(() => { + order.push("a"); + }); + const second = saveResultToFile(TEXT, "b.json", "json", "/d").then(() => { + order.push("b"); + }); + await flush(); + expect(writes.map((w) => w.path)).toEqual(["/d/a.json"]); + writes[0]!.resolve(); + await first; + await flush(); + expect(writes.map((w) => w.path)).toEqual(["/d/a.json", "/d/b.json"]); + writes[1]!.resolve(); + await second; + expect(order).toEqual(["a", "b"]); + }); + + it("a failed save does not hold up the ones queued after it", async () => { + writes.length = 0; + const failed = saveResultToFile(TEXT, "x.json", "json", "/d"); + const next = saveResultToFile(TEXT, "y.json", "json", "/d"); + await flush(); + expect(writes).toHaveLength(1); + writes[0]!.reject(new Error("EACCES")); + await expect(failed).rejects.toThrow("Could not write /d/x.json: EACCES"); + await flush(); + expect(writes).toHaveLength(2); + writes[1]!.resolve(); + await expect(next).resolves.toMatchObject({ path: "/d/y.json" }); + }); +}); diff --git a/clients/tui/__tests__/saveResult.test.ts b/clients/tui/__tests__/saveResult.test.ts new file mode 100644 index 0000000000..a8927c0536 --- /dev/null +++ b/clients/tui/__tests__/saveResult.test.ts @@ -0,0 +1,94 @@ +import { describe, it, expect, beforeEach, afterEach } from "vitest"; +import { mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { + defaultResultFileName, + saveResultToFile, +} from "../src/utils/saveResult.js"; + +const TEXT = { content: [{ type: "text", text: "hi" }] }; +const IMAGE = { + content: [{ type: "image", data: btoa("\x01\x02"), mimeType: "image/png" }], +}; +const LINK_ONLY = { content: [{ type: "resource_link", uri: "x://a" }] }; + +describe("defaultResultFileName (#2571)", () => { + it("is <tool>-result.json for json", () => { + expect(defaultResultFileName("alpha", "json", TEXT)).toBe( + "alpha-result.json", + ); + }); + + it("picks .txt or .bin for raw by what the result renders as", () => { + expect(defaultResultFileName("alpha", "raw", TEXT)).toBe( + "alpha-result.txt", + ); + expect(defaultResultFileName("alpha", "raw", IMAGE)).toBe( + "alpha-result.bin", + ); + expect(defaultResultFileName("alpha", "raw", LINK_ONLY)).toBe( + "alpha-result.txt", + ); + }); + + it("makes the tool name safe as a bare file name", () => { + expect(defaultResultFileName("fs/read file", "json", TEXT)).toBe( + "fs_read_file-result.json", + ); + expect(defaultResultFileName("", "json", TEXT)).toBe("tool-result.json"); + }); +}); + +describe("saveResultToFile (#2571)", () => { + let dir: string; + beforeEach(() => { + dir = mkdtempSync(join(tmpdir(), "tui-save-util-")); + }); + afterEach(() => { + rmSync(dir, { recursive: true, force: true }); + }); + + it("writes json relative to cwd and reports the absolute path and size", async () => { + const saved = await saveResultToFile(TEXT, " out.json ", "json", dir); + const expected = JSON.stringify(TEXT, null, 2) + "\n"; + expect(saved).toEqual({ + path: join(dir, "out.json"), + format: "json", + bytes: expected.length, + }); + expect(readFileSync(join(dir, "out.json"), "utf8")).toBe(expected); + }); + + it("writes decoded bytes for a single binary block as raw", async () => { + const saved = await saveResultToFile(IMAGE, "img.bin", "raw", dir); + expect(saved.bytes).toBe(2); + expect([...readFileSync(join(dir, "img.bin"))]).toEqual([1, 2]); + }); + + it("defaults cwd to the process working directory", async () => { + const target = join(dir, "abs.json"); + const saved = await saveResultToFile(TEXT, target, "json"); + expect(saved.path).toBe(target); + }); + + it("refuses an empty path", async () => { + await expect(saveResultToFile(TEXT, " ", "json", dir)).rejects.toThrow( + "Enter a file path to save to.", + ); + }); + + it("explains a result with no raw form and points at json", async () => { + await expect( + saveResultToFile(LINK_ONLY, "x.txt", "raw", dir), + ).rejects.toThrow( + /no text or binary content.*as json instead \(w, then Enter\)/, + ); + }); + + it("reports a failed write with the resolved path", async () => { + await expect( + saveResultToFile(TEXT, "missing/x.json", "json", dir), + ).rejects.toThrow(`Could not write ${join(dir, "missing/x.json")}:`); + }); +}); diff --git a/clients/tui/__tests__/tabsConfig.test.ts b/clients/tui/__tests__/tabsConfig.test.ts index 2bdb41b5fe..ab3975b095 100644 --- a/clients/tui/__tests__/tabsConfig.test.ts +++ b/clients/tui/__tests__/tabsConfig.test.ts @@ -29,6 +29,8 @@ describe("visibleTabs", () => { expect(ids).not.toContain("logging"); expect(ids).not.toContain("requests"); expect(ids).not.toContain("skills"); + expect(ids).not.toContain("subscriptions"); + expect(ids).not.toContain("tasks"); // The unconditional ones remain. expect(ids).toContain("info"); expect(ids).toContain("tools"); @@ -40,6 +42,8 @@ describe("visibleTabs", () => { showLogging: true, showRequests: true, showSkills: true, + showSubscriptions: true, + showTasks: true, }).map((t) => t.id); expect(ids).toEqual(tabs.map((t) => t.id)); }); diff --git a/clients/tui/__tests__/tui-entry-flags.test.ts b/clients/tui/__tests__/tui-entry-flags.test.ts new file mode 100644 index 0000000000..a68cc609b3 --- /dev/null +++ b/clients/tui/__tests__/tui-entry-flags.test.ts @@ -0,0 +1,63 @@ +import { describe, it, expect, vi } from "vitest"; +import type { ServerLoadOptions } from "@inspector/core/mcp/node/servers.js"; + +/** + * `runTui`'s Commander registration and the forwarding of parsed flags into + * the shared server loader (#2420). `tui.tsx` sits outside the TUI coverage + * `include` (`src/**`), and `tui-servers.test.ts` calls the loader directly, + * so without this a misspelt option or a dropped forwarding field would pass + * every other test. + * + * The loader is mocked to throw a sentinel carrying the options it received, + * which stops `runTui` before it touches stdout or renders Ink. + */ +class Captured extends Error { + constructor(readonly options: ServerLoadOptions) { + super("captured"); + } +} + +vi.mock("../src/tui-servers.js", () => ({ + loadTuiServers: (options: ServerLoadOptions) => { + throw new Captured(options); + }, +})); + +async function loaderOptionsFor(argv: string[]): Promise<ServerLoadOptions> { + const { runTui } = await import("../tui.js"); + try { + await runTui(["node", "mcp-inspector-tui", ...argv]); + } catch (err) { + if (err instanceof Captured) return err.options; + throw err; + } + throw new Error("runTui returned without calling the server loader"); +} + +describe("runTui --skill-catalog-max-* forwarding (#2420)", () => { + it("forwards both budget flags, parsed as numbers, to the server loader", async () => { + const options = await loaderOptionsFor([ + "--server-url", + "http://127.0.0.1:1/mcp", + "--transport", + "http", + "--skill-catalog-max-skills", + "3", + "--skill-catalog-max-bytes", + "4096", + ]); + expect(options.skillCatalogMaxSkills).toBe(3); + expect(options.skillCatalogMaxBytes).toBe(4096); + }); + + it("leaves both unset when the flags are absent", async () => { + const options = await loaderOptionsFor([ + "--server-url", + "http://127.0.0.1:1/mcp", + "--transport", + "http", + ]); + expect(options.skillCatalogMaxSkills).toBeUndefined(); + expect(options.skillCatalogMaxBytes).toBeUndefined(); + }); +}); diff --git a/clients/tui/__tests__/useCopyKeys.test.tsx b/clients/tui/__tests__/useCopyKeys.test.tsx new file mode 100644 index 0000000000..997c2e9f80 --- /dev/null +++ b/clients/tui/__tests__/useCopyKeys.test.tsx @@ -0,0 +1,151 @@ +import React from "react"; +import fs from "node:fs"; +import path from "node:path"; +import { afterEach, describe, expect, it, vi } from "vitest"; +import { Text, useInput } from "ink"; +import { render } from "./helpers/renderTui"; +import { + COPY_KEY, + SAVE_KEY, + copyStatusColor, + useCopyKeys, +} from "../src/hooks/useCopyKeys.js"; +import { + OSC52_LARGE_PAYLOAD_BYTES, + buildOsc52Sequence, +} from "../src/utils/clipboard.js"; + +const tick = async () => { + for (let i = 0; i < 8; i++) + await new Promise((resolve) => setTimeout(resolve, 4)); +}; + +const squashWhitespace = (s: string) => s.replace(/\s+/g, ""); + +/** Renders the hook's status and records whether each key was consumed. */ +function Harness({ + value, + consumed, +}: { + value: string | undefined; + consumed: boolean[]; +}) { + const { handleCopyKey, status } = useCopyKeys(value, "thing"); + useInput((input) => { + consumed.push(handleCopyKey(input)); + }); + return ( + <Text>{status ? `[${status.tone}] ${status.message}` : "no status"}</Text> + ); +} + +const savedPaths: string[] = []; +afterEach(() => { + vi.restoreAllMocks(); + for (const file of savedPaths.splice(0)) { + fs.rmSync(path.dirname(file), { recursive: true, force: true }); + } +}); + +describe("useCopyKeys", () => { + it("Y emits OSC 52 to stdout and reports success", async () => { + const consumed: boolean[] = []; + const { stdin, stdout, lastFrame } = render( + <Harness value="hello" consumed={consumed} />, + ); + await tick(); + stdin.write(COPY_KEY.toUpperCase()); + await tick(); + expect(stdout.frames).toContain(buildOsc52Sequence("hello")); + expect(lastFrame()).toContain("[success] Copied thing (5 chars) via OSC"); + expect(consumed).toEqual([true]); + }); + + it("warns when the payload is large", async () => { + const big = "x".repeat(OSC52_LARGE_PAYLOAD_BYTES); + const { stdin, lastFrame } = render(<Harness value={big} consumed={[]} />); + await tick(); + stdin.write(COPY_KEY); + await tick(); + expect(lastFrame()).toContain("[warning] Sent thing"); + }); + + it("W saves the value to a file and shows its path", async () => { + // Take the path from the directory mkdtemp actually created rather than + // parsing it out of the frame: Ink wraps the status line at the frame + // width, and a long os.tmpdir() (macOS's /var/folders/…/T/) moves the + // wrap to wherever it falls, so no regex over the frame is reliable (#2609). + const mkdtemp = vi.spyOn(fs, "mkdtempSync"); + const { stdin, lastFrame } = render( + <Harness value="saved-value" consumed={[]} />, + ); + await tick(); + stdin.write(SAVE_KEY); + await tick(); + const file = path.join(String(mkdtemp.mock.results[0]?.value), "value.txt"); + savedPaths.push(file); + expect(fs.readFileSync(file, "utf8")).toBe("saved-value"); + // Compare with all whitespace removed, so a wrap anywhere in the line — + // including mid-path — cannot fail the check. + expect(squashWhitespace(lastFrame() ?? "")).toContain( + squashWhitespace(`[success] Saved thing to ${file}`), + ); + }); + + it("reports a save failure (Error and non-Error)", async () => { + const spy = vi.spyOn(fs, "mkdtempSync").mockImplementationOnce(() => { + throw new Error("disk full"); + }); + const { stdin, lastFrame } = render(<Harness value="v" consumed={[]} />); + await tick(); + stdin.write(SAVE_KEY); + await tick(); + expect(lastFrame()).toContain("[error] Could not save thing: disk full"); + + spy.mockImplementationOnce(() => { + throw "plain string"; + }); + stdin.write(SAVE_KEY); + await tick(); + expect(lastFrame()).toContain("Could not save thing: plain string"); + }); + + it("consumes nothing when there is no value, and ignores other keys", async () => { + const consumed: boolean[] = []; + const { stdin, lastFrame } = render( + <Harness value={undefined} consumed={consumed} />, + ); + await tick(); + stdin.write(COPY_KEY); + await tick(); + const withValue: boolean[] = []; + const second = render(<Harness value="v" consumed={withValue} />); + await tick(); + second.stdin.write("q"); + await tick(); + expect(consumed).toEqual([false]); + expect(withValue).toEqual([false]); + expect(lastFrame()).toContain("no status"); + }); + + it("clears the status when the value changes", async () => { + const { stdin, lastFrame, rerender } = render( + <Harness value="first" consumed={[]} />, + ); + await tick(); + stdin.write(COPY_KEY); + await tick(); + expect(lastFrame()).toContain("Copied thing"); + rerender(<Harness value="second" consumed={[]} />); + await tick(); + expect(lastFrame()).toContain("no status"); + }); +}); + +describe("copyStatusColor", () => { + it("maps each tone to an Ink colour", () => { + expect(copyStatusColor("success")).toBe("green"); + expect(copyStatusColor("warning")).toBe("yellow"); + expect(copyStatusColor("error")).toBe("red"); + }); +}); diff --git a/clients/tui/__tests__/useListFilter.test.tsx b/clients/tui/__tests__/useListFilter.test.tsx new file mode 100644 index 0000000000..c3e89bc660 --- /dev/null +++ b/clients/tui/__tests__/useListFilter.test.tsx @@ -0,0 +1,179 @@ +import React, { useState } from "react"; +import { describe, it, expect, vi } from "vitest"; +import { Text, useInput } from "ink"; +import { render } from "./helpers/renderTui"; +import { useListFilter } from "../src/hooks/useListFilter.js"; + +const tick = async () => { + for (let i = 0; i < 8; i++) + await new Promise((resolve) => setTimeout(resolve, 4)); +}; + +const ESC = String.fromCharCode(27); +const UP = `${ESC}[A`; +const ENTER = "\r"; +const BACKSPACE = "\x7f"; +const CTRL_B = "\x02"; + +interface Item { + name: string; + title?: string; +} + +const items: Item[] = [ + { name: "alpha", title: "First Thing" }, + { name: "beta" }, + { name: "gamma", title: "ALPHA-ish" }, +]; + +const fields = (item: Item) => [item.name, item.title]; + +/** + * Mounts the hook behind a `useInput` the way a list tab does, and prints its + * state as one line so a test can read every field back from the frame. + */ +function Harness({ + enabled = true, + onEditingChange, + onUnconsumed, +}: { + enabled?: boolean; + onEditingChange?: (editing: boolean) => void; + onUnconsumed?: (input: string) => void; +}) { + const filter = useListFilter(items, fields, { enabled, onEditingChange }); + useInput((input, key) => { + if (!filter.handleInput(input, key)) onUnconsumed?.(input); + }); + return ( + <Text> + q=[{filter.query}] editing={String(filter.editing)} active= + {String(filter.active)} items={filter.items.map((i) => i.name).join(",")}{" "} + indices={filter.indices.join(",")} + </Text> + ); +} + +async function type(stdin: { write: (s: string) => void }, keys: string[]) { + for (const k of keys) { + stdin.write(k); + await tick(); + } +} + +describe("useListFilter", () => { + it("shows every item until a query narrows them", () => { + const { lastFrame } = render(<Harness />); + expect(lastFrame()).toContain("items=alpha,beta,gamma"); + expect(lastFrame()).toContain("indices=0,1,2"); + expect(lastFrame()).toContain("active=false"); + }); + + it("opens on '/', narrows case-insensitively over every field, and keeps on Enter", async () => { + const onEditingChange = vi.fn(); + const { lastFrame, stdin } = render( + <Harness onEditingChange={onEditingChange} />, + ); + await tick(); + await type(stdin, ["/"]); + expect(lastFrame()).toContain("editing=true"); + expect(onEditingChange).toHaveBeenLastCalledWith(true); + + // "Alpha" matches alpha by name and gamma by its title. + await type(stdin, ["A", "l", "p", "h", "a"]); + expect(lastFrame()).toContain("q=[Alpha]"); + expect(lastFrame()).toContain("items=alpha,gamma"); + expect(lastFrame()).toContain("indices=0,2"); + expect(lastFrame()).toContain("active=true"); + + await type(stdin, [ENTER]); + expect(lastFrame()).toContain("editing=false"); + expect(lastFrame()).toContain("items=alpha,gamma"); + expect(onEditingChange).toHaveBeenLastCalledWith(false); + }); + + it("trims with backspace and clears on Esc", async () => { + const { lastFrame, stdin } = render(<Harness />); + await tick(); + await type(stdin, ["/", "b", "x"]); + expect(lastFrame()).toContain("items= "); + await type(stdin, [BACKSPACE]); + expect(lastFrame()).toContain("q=[b]"); + expect(lastFrame()).toContain("items=beta"); + await type(stdin, [ESC]); + expect(lastFrame()).toContain("q=[]"); + expect(lastFrame()).toContain("editing=false"); + expect(lastFrame()).toContain("items=alpha,beta,gamma"); + }); + + it("leaves navigation and unrelated keys to the list", async () => { + const onUnconsumed = vi.fn(); + const { lastFrame, stdin } = render( + <Harness onUnconsumed={onUnconsumed} />, + ); + await tick(); + // Not editing: an ordinary letter is the list's to handle. + await type(stdin, ["x"]); + expect(onUnconsumed).toHaveBeenCalledWith("x"); + onUnconsumed.mockClear(); + + await type(stdin, ["/"]); + // Editing: arrows still reach the list, ctrl chords and control + // characters are swallowed without touching the query. + await type(stdin, [UP, CTRL_B, "\t"]); + expect(onUnconsumed).toHaveBeenCalledTimes(1); + expect(lastFrame()).toContain("q=[]"); + expect(lastFrame()).toContain("editing=true"); + }); + + it("ignores '/' and suspends editing while the list lacks the keyboard", async () => { + const onEditingChange = vi.fn(); + const { lastFrame, stdin, rerender } = render( + <Harness enabled={false} onEditingChange={onEditingChange} />, + ); + await tick(); + await type(stdin, ["/"]); + expect(lastFrame()).toContain("editing=false"); + + rerender(<Harness enabled onEditingChange={onEditingChange} />); + await tick(); + await type(stdin, ["/"]); + expect(onEditingChange).toHaveBeenLastCalledWith(true); + + // A modal opening (or focus moving) withdraws the report… + rerender(<Harness enabled={false} onEditingChange={onEditingChange} />); + await tick(); + expect(lastFrame()).toContain("editing=false"); + expect(onEditingChange).toHaveBeenLastCalledWith(false); + }); + + it("withdraws the editing report when the list unmounts mid-query", async () => { + const onEditingChange = vi.fn(); + const { stdin, unmount } = render( + <Harness onEditingChange={onEditingChange} />, + ); + await tick(); + await type(stdin, ["/"]); + expect(onEditingChange).toHaveBeenLastCalledWith(true); + unmount(); + expect(onEditingChange).toHaveBeenLastCalledWith(false); + }); + + it("works without an editing listener", async () => { + function Bare() { + const [, force] = useState(0); + const filter = useListFilter(items, fields, { enabled: true }); + useInput((input, key) => { + filter.handleInput(input, key); + force((n) => n + 1); + }); + return <Text>editing={String(filter.editing)}</Text>; + } + const { lastFrame, stdin } = render(<Bare />); + await tick(); + await type(stdin, ["/"]); + expect(lastFrame()).toContain("editing=true"); + await type(stdin, [ENTER]); + expect(lastFrame()).toContain("editing=false"); + }); +}); diff --git a/clients/tui/package-lock.json b/clients/tui/package-lock.json index 211ff8bc98..33617e0a94 100644 --- a/clients/tui/package-lock.json +++ b/clients/tui/package-lock.json @@ -1143,9 +1143,6 @@ "arm64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1163,9 +1160,6 @@ "arm64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -1183,9 +1177,6 @@ "ppc64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1203,9 +1194,6 @@ "s390x" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1223,9 +1211,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1243,9 +1228,6 @@ "x64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -3084,9 +3066,6 @@ "arm64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -3108,9 +3087,6 @@ "arm64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -3132,9 +3108,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -3156,9 +3129,6 @@ "x64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -3850,9 +3820,9 @@ } }, "node_modules/source-map-js": { - "version": "1.2.1", - "resolved": "https://registry.npmjs.org/source-map-js/-/source-map-js-1.2.1.tgz", - "integrity": "sha512-UXWMKhLOwVKb728IUtQPXxfYU+usdybtUrK/8uGE8CQMvrhOpwvzDBwj0QhSL7MQc7vIsISBG8VQ8+IDQxpfQA==", + "version": "1.2.2", + "resolved": "https://registry.npmjs.org/source-map-js/-/source-map-js-1.2.2.tgz", + "integrity": "sha512-KGj/8Y43x35aZVDtt+J4mK1hoLGHULMYfSkODJNQjNDC3oW1PqPoxMwo0pLUsWM/UEGzON/NxeHywEfNXNP3Vw==", "dev": true, "license": "BSD-3-Clause", "engines": { @@ -4430,9 +4400,9 @@ "license": "MIT" }, "node_modules/zod": { - "version": "4.4.3", - "resolved": "https://registry.npmjs.org/zod/-/zod-4.4.3.tgz", - "integrity": "sha512-ytENFjIJFl2UwYglde2jchW2Hwm4GJFLDiSXWdTrJQBIN9Fcyp7n4DhxJEiWNAJMV1/BqWfW/kkg71UDcHJyTQ==", + "version": "4.6.5", + "resolved": "https://registry.npmjs.org/zod/-/zod-4.6.5.tgz", + "integrity": "sha512-v5l/aFXZQeai4awLbOpSoHecE9UiMrnfx75tEXLjNonXVARxQ5mOeipTjROUchszUNCqnE+hqAMujRsRHsut2Q==", "dev": true, "license": "MIT", "funding": { diff --git a/clients/tui/src/App.tsx b/clients/tui/src/App.tsx index 9a463d2efa..754cd859e0 100644 --- a/clients/tui/src/App.tsx +++ b/clients/tui/src/App.tsx @@ -18,6 +18,9 @@ import type { Prompt, PromptArgument, GetPromptResult, + Root, + Task, + CallToolResult, } from "@modelcontextprotocol/client"; import { InspectorClient } from "@inspector/core/mcp/index.js"; import { cleanRoots } from "@inspector/core/mcp/serverList.js"; @@ -31,6 +34,8 @@ import { MessageLogState, FetchRequestLogState, StderrLogState, + ManagedRequestorTasksState, + ResourceSubscriptionsState, } from "@inspector/core/mcp/state/index.js"; import { createProxyFetch, @@ -45,6 +50,9 @@ import { useManagedSkills } from "@inspector/core/react/useManagedSkills.js"; import { useMessageLog } from "@inspector/core/react/useMessageLog.js"; import { useFetchRequestLog } from "@inspector/core/react/useFetchRequestLog.js"; import { useStderrLog } from "@inspector/core/react/useStderrLog.js"; +import { useManagedRequestorTasks } from "@inspector/core/react/useManagedRequestorTasks.js"; +import { useResourceSubscriptions } from "@inspector/core/react/useResourceSubscriptions.js"; +import { useStoreSnapshot } from "@inspector/core/react/useStoreSnapshot.js"; import { CallbackNavigation, MutableRedirectUrlProvider, @@ -91,8 +99,15 @@ import { ToolTestModal } from "./components/ToolTestModal.js"; import { ResourceTestModal } from "./components/ResourceTestModal.js"; import { PromptTestModal } from "./components/PromptTestModal.js"; import { DetailsModal } from "./components/DetailsModal.js"; +import { toCopyText } from "./utils/clipboard.js"; +import { HelpOverlay } from "./components/HelpOverlay.js"; +import { keybindingSections } from "./utils/keybindings.js"; +import { TasksTab } from "./components/TasksTab.js"; +import { SubscriptionsTab } from "./components/SubscriptionsTab.js"; +import { RootsModal } from "./components/RootsModal.js"; import { BodyLines } from "./components/BodyLines.js"; import type { TuiServer } from "./tui-servers.js"; +import { errorMessage, redactErrorText } from "./utils/errorText.js"; // Header branding. The version is the single source of truth — the root // package.json — read via the shared core reader; the name/description are the @@ -106,6 +121,13 @@ const APP_VERSION = readInspectorVersion(import.meta.url); /** Client identity name the TUI reports to servers. */ const TUI_CLIENT_NAME = "inspector-tui"; +/** + * Roots snapshot for `useStoreSnapshot`: module scope so both the reader and + * the fallback are referentially stable across renders (#2432). + */ +const NO_ROOTS: Root[] = []; +const readRoots = (client: InspectorClient): Root[] => client.getRoots(); + // Focus management types type FocusArea = | "serverList" @@ -119,7 +141,11 @@ type FocusArea = | "messagesDetail" // Used only when activeTab === 'requests' | "requestsList" - | "requestsDetail"; + | "requestsDetail" + // What `focus` reads while the `?` help overlay is open (#2436). No pane + // matches it, so every tab's own key handler goes inert underneath. Never + // stored in `paneFocus`. + | "help"; interface AppProps { mcpServers: Record<string, TuiServer>; @@ -153,13 +179,26 @@ function App({ const [selectedServer, setSelectedServer] = useState<string | null>(null); const [activeTab, setActiveTab] = useState<TabType>("info"); - const [focus, setFocus] = useState<FocusArea>("serverList"); + const [paneFocus, setFocus] = useState<FocusArea>("serverList"); + const [helpOpen, setHelpOpen] = useState(false); + // While the `?` help overlay is open, every pane sees focus as "help" — which + // none matches — so their key handlers go inert underneath it. Derived rather + // than set, so a focus move made while it is open (an async step-up landing + // on the Auth tab) updates `paneFocus` and takes effect only once the help + // closes, instead of re-activating a hidden pane (Copilot). + const focus: FocusArea = helpOpen ? "help" : paneFocus; + // True while a list tab's `/` filter is capturing keystrokes (#2430). The + // global accelerators below stand down then, or typing a query would switch + // tabs, connect, or quit on Esc. Reported by `useListFilter`. + const [listFilterEditing, setListFilterEditing] = useState(false); const [tabCounts, setTabCounts] = useState<{ info?: number; resources?: number; + subscriptions?: number; prompts?: number; skills?: number; tools?: number; + tasks?: number; messages?: number; requests?: number; logging?: number; @@ -232,8 +271,13 @@ function App({ const [detailsModal, setDetailsModal] = useState<{ title: string; content: React.ReactNode; + /** Raw value behind `content`, for the modal's copy keys (#2421). */ + copyText: string; } | null>(null); + // Roots editor (#2432), opened from the Info tab. + const [rootsModalOpen, setRootsModalOpen] = useState(false); + // InspectorClient instances for each server const [inspectorClients, setInspectorClients] = useState< Record<string, InspectorClient> @@ -262,6 +306,13 @@ function App({ const [stderrLogStates, setStderrLogStates] = useState< Record<string, StderrLogState> >({}); + // Created with the client, like the stores above, so a task status update or + // a resources/updated that lands while another tab is showing is not lost. + const [requestorTasksStates, setRequestorTasksStates] = useState< + Record<string, ManagedRequestorTasksState> + >({}); + const [resourceSubscriptionsStates, setResourceSubscriptionsStates] = + useState<Record<string, ResourceSubscriptionsState>>({}); const [dimensions, setDimensions] = useState({ width: process.stdout.columns || 80, height: process.stdout.rows || 24, @@ -294,6 +345,19 @@ function App({ // Create InspectorClient and state managers for each server on mount useEffect(() => { + // The browser could not be launched for `serverName`'s authorization page + // (#2533). Show the manual-open note on the Auth tab, raised as a warning + // so it is coloured as one (see oauthMessageToneFor) — unless the user has + // since selected another server, where that URL would be the wrong one. + const showBrowserOpenFailure = ( + serverName: string, + message: string, + ): void => { + if (selectedServerRef.current !== serverName) return; + setOauthWarningText(message); + setOauthMessage(message); + setActiveTab("auth"); + }; const newClients: Record<string, InspectorClient> = {}; const newManagers: Record<string, ManagedToolsState> = {}; const newManagedResourcesStates: Record<string, ManagedResourcesState> = {}; @@ -306,6 +370,12 @@ function App({ const newMessageLogStates: Record<string, MessageLogState> = {}; const newFetchRequestLogStates: Record<string, FetchRequestLogState> = {}; const newStderrLogStates: Record<string, StderrLogState> = {}; + const newRequestorTasksStates: Record<string, ManagedRequestorTasksState> = + {}; + const newResourceSubscriptionsStates: Record< + string, + ResourceSubscriptionsState + > = {}; for (const serverName of serverNames) { if (!(serverName in inspectorClients)) { const { config: serverConfig, settings: savedSettings } = @@ -362,8 +432,12 @@ function App({ formatRunnerOAuthRedirectUrl(callbackUrlConfig); environment.oauth = { storage: new NodeOAuthStorage(), - navigation: new CallbackNavigation( - async (url) => await openUrl(url), + // openUrl never rejects; a browser that could not be launched + // (#2533) surfaces as the manual-open note instead. + navigation: new CallbackNavigation((url) => + openUrl(url, (message) => + showBrowserOpenFailure(serverName, message), + ), ), redirectUrlProvider, }; @@ -371,9 +445,8 @@ function App({ const client = new InspectorClient(serverConfig, opts); newClients[serverName] = client; newManagers[serverName] = new ManagedToolsState(client); - newManagedResourcesStates[serverName] = new ManagedResourcesState( - client, - ); + const resourcesState = new ManagedResourcesState(client); + newManagedResourcesStates[serverName] = resourcesState; newManagedResourceTemplatesStates[serverName] = new ManagedResourceTemplatesState(client); newManagedPromptsStates[serverName] = new ManagedPromptsState(client); @@ -381,6 +454,13 @@ function App({ newMessageLogStates[serverName] = new MessageLogState(client); newFetchRequestLogStates[serverName] = new FetchRequestLogState(client); newStderrLogStates[serverName] = new StderrLogState(client); + newRequestorTasksStates[serverName] = new ManagedRequestorTasksState( + client, + ); + // Given the resources store so a subscription carries the listed + // Resource's name rather than a bare URI. + newResourceSubscriptionsStates[serverName] = + new ResourceSubscriptionsState(client, resourcesState); } } if (Object.keys(newClients).length > 0) { @@ -408,6 +488,14 @@ function App({ ...newFetchRequestLogStates, })); setStderrLogStates((prev) => ({ ...prev, ...newStderrLogStates })); + setRequestorTasksStates((prev) => ({ + ...prev, + ...newRequestorTasksStates, + })); + setResourceSubscriptionsStates((prev) => ({ + ...prev, + ...newResourceSubscriptionsStates, + })); } // Omitted on purpose: `inspectorClients` is this effect's own output, so // depending on it would re-run the effect after every client it creates @@ -453,6 +541,12 @@ function App({ Object.values(stderrLogStates).forEach((manager) => { manager.destroy(); }); + Object.values(requestorTasksStates).forEach((manager) => { + manager.destroy(); + }); + Object.values(resourceSubscriptionsStates).forEach((manager) => { + manager.destroy(); + }); Object.values(inspectorClients).forEach((client) => { client.disconnect().catch(() => { // Ignore errors during cleanup @@ -469,6 +563,8 @@ function App({ messageLogStates, fetchRequestLogStates, stderrLogStates, + requestorTasksStates, + resourceSubscriptionsStates, ]); // Preselect the first server on mount @@ -659,6 +755,62 @@ function App({ } }, [activeTab, inspectorStatus, showSkillsTab]); + // Tasks, resource subscriptions and roots (#2432) — all read from core + // stores, so nothing about their lifecycle is decided in the TUI. + const selectedRequestorTasksState = useMemo( + () => + selectedServer && requestorTasksStates[selectedServer] + ? requestorTasksStates[selectedServer] + : null, + [selectedServer, requestorTasksStates], + ); + const { + tasks: requestorTasks, + refresh: refreshRequestorTasks, + clearCompleted: clearCompletedRequestorTasks, + } = useManagedRequestorTasks( + selectedInspectorClient, + selectedRequestorTasksState, + ); + const selectedResourceSubscriptionsState = useMemo( + () => + selectedServer && resourceSubscriptionsStates[selectedServer] + ? resourceSubscriptionsStates[selectedServer] + : null, + [selectedServer, resourceSubscriptionsStates], + ); + const { + subscriptions: resourceSubscriptions, + streamState: subscriptionStreamState, + } = useResourceSubscriptions(selectedResourceSubscriptionsState); + const advertisedRoots = useStoreSnapshot( + selectedInspectorClient ?? null, + "rootsChange", + readRoots, + NO_ROOTS, + ); + // Server-declared, so only knowable once connected — the same gate the web + // client uses for its Resources subscribe control and its Tasks screen. + const showSubscriptionsTab = + inspectorStatus === "connected" && + inspectorCapabilities?.resources?.subscribe === true; + const showTasksTab = + inspectorStatus === "connected" && + (!!inspectorCapabilities?.tasks || + (selectedInspectorClient?.isTasksExtensionNegotiated() ?? false)); + + // Switch away from a tab the server stops serving, for the Skills reason + // above: the bar drops it, but `activeTab` would keep rendering its pane. + useEffect(() => { + if (inspectorStatus !== "connected") return; + if ( + (activeTab === "subscriptions" && !showSubscriptionsTab) || + (activeTab === "tasks" && !showTasksTab) + ) { + setActiveTab("info"); + } + }, [activeTab, inspectorStatus, showSubscriptionsTab, showTasksTab]); + // Connect — on 401 or mid-session auth recovery, run OAuth then retry. type TuiOAuthRunResult = | "success" @@ -838,8 +990,7 @@ function App({ setOauthMessage("OAuth already in progress."); } } catch (authErr) { - const authMsg = - authErr instanceof Error ? authErr.message : String(authErr); + const authMsg = errorMessage(authErr); setOauthStatus("error"); setOauthMessage(authMsg); } @@ -882,12 +1033,12 @@ function App({ try { await finishConnect(); } catch (err) { - const msg = err instanceof Error ? err.message : String(err); + const msg = errorMessage(err); setConnectError(msg); if (isEmaClientNotConfiguredError(err)) { setOauthStatus("error"); - setOauthMessage(err.message); + setOauthMessage(errorMessage(err)); return; } @@ -914,12 +1065,11 @@ function App({ handleAuthRecoveryRequired(selectedServer, authErr); return; } - const authMsg = - authErr instanceof Error ? authErr.message : String(authErr); + const authMsg = errorMessage(authErr); setConnectError(authMsg); if (isEmaClientNotConfiguredError(authErr)) { setOauthStatus("error"); - setOauthMessage(authErr.message); + setOauthMessage(authMsg); return; } setOauthStatus("error"); @@ -970,10 +1120,7 @@ function App({ }; const onOAuthError = (event: TypedEvent<"oauthError">): void => { if (selectedServerRef.current !== serverName) return; - const message = - event.detail.error instanceof Error - ? event.detail.error.message - : String(event.detail.error); + const message = errorMessage(event.detail.error); setOauthStatus("error"); setOauthMessage(message); }; @@ -1023,7 +1170,7 @@ function App({ ) { return; } - setDisconnectError(err instanceof Error ? err.message : String(err)); + setDisconnectError(errorMessage(err)); } }, [selectedServer, disconnectInspector]); @@ -1052,7 +1199,9 @@ function App({ } setOauthStatus("idle"); if (revocation.status === "failed") { - const warning = `Cleared locally, but revoking the grant at the authorization server failed: ${revocation.detail}. It may still be valid there.`; + // The detail is a caught revocation/network error's text, so it can quote + // a secret-bearing URL (#2490). + const warning = `Cleared locally, but revoking the grant at the authorization server failed: ${redactErrorText(revocation.detail)}. It may still be valid there.`; setOauthWarningText(warning); setOauthMessage(warning); } else { @@ -1067,7 +1216,7 @@ function App({ try { await disconnectInspector(); } catch (err) { - setDisconnectError(err instanceof Error ? err.message : String(err)); + setDisconnectError(errorMessage(err)); } // Revalidate: the disconnect is a second await, and a switch during it // would make the revision bump below land on the new selection. @@ -1092,7 +1241,12 @@ function App({ if (!selectedServer) return null; return { status: inspectorStatus, - error: connectError ?? inspectorLastError ?? null, + // `connectError` is already redacted where it is set; `lastError` comes + // from core's client event untouched, so it is redacted here, at the one + // place it reaches the screen (#2490). + error: + connectError ?? + (inspectorLastError ? redactErrorText(inspectorLastError) : null), capabilities: inspectorCapabilities, serverInfo: inspectorServerInfo, instructions: inspectorInstructions, @@ -1318,6 +1472,25 @@ function App({ </> ); + const renderTaskDetails = (task: Task, result: CallToolResult | null) => ( + <> + <Box flexShrink={0} flexDirection="column"> + <Text bold>Task:</Text> + <Box paddingLeft={2}> + <Text dimColor>{JSON.stringify(task, null, 2)}</Text> + </Box> + </Box> + {result && ( + <Box marginTop={1} flexShrink={0} flexDirection="column"> + <Text bold>Result:</Text> + <Box paddingLeft={2}> + <Text dimColor>{JSON.stringify(result, null, 2)}</Text> + </Box> + </Box> + )} + </> + ); + const renderMessageDetails = (message: MessageEntry) => ( <> <Box flexShrink={0}> @@ -1329,34 +1502,41 @@ function App({ {message.duration !== undefined && ` (${message.duration}ms)`} </Text> </Box> + {/* Bodies go through the same capped BodyLines as the Network zoom + above, so a huge payload cannot render unbounded here (#2539). */} {message.direction === "request" ? ( <> - <Box marginTop={1} flexShrink={0} flexDirection="column"> + <Box marginTop={1} flexShrink={0}> <Text bold>Request:</Text> - <Box paddingLeft={2}> - <Text dimColor>{JSON.stringify(message.message, null, 2)}</Text> - </Box> </Box> + <BodyLines + body={JSON.stringify(message.message)} + keyPrefix="zoom-req" + /> {message.response && ( - <Box marginTop={1} flexShrink={0} flexDirection="column"> - <Text bold>Response:</Text> - <Box paddingLeft={2}> - <Text dimColor> - {JSON.stringify(message.response, null, 2)} - </Text> + <> + <Box marginTop={1} flexShrink={0}> + <Text bold>Response:</Text> </Box> - </Box> + <BodyLines + body={JSON.stringify(message.response)} + keyPrefix="zoom-resp" + /> + </> )} </> ) : ( - <Box marginTop={1} flexShrink={0} flexDirection="column"> - <Text bold> - {message.direction === "response" ? "Response:" : "Notification:"} - </Text> - <Box paddingLeft={2}> - <Text dimColor>{JSON.stringify(message.message, null, 2)}</Text> + <> + <Box marginTop={1} flexShrink={0}> + <Text bold> + {message.direction === "response" ? "Response:" : "Notification:"} + </Text> </Box> - </Box> + <BodyLines + body={JSON.stringify(message.message)} + keyPrefix="zoom-msg" + /> + </> )} </> ); @@ -1373,6 +1553,8 @@ function App({ prompts: managedPrompts.length || 0, skills: managedSkills.length || 0, tools: managedTools.length || 0, + subscriptions: resourceSubscriptions.length, + tasks: requestorTasks.length, messages: inspectorMessages.length || 0, requests: inspectorFetchRequests.length || 0, logging: inspectorStderrLogs.length || 0, @@ -1383,6 +1565,8 @@ function App({ managedPrompts, managedSkills, managedTools, + resourceSubscriptions, + requestorTasks, inspectorMessages, inspectorFetchRequests, inspectorStderrLogs, @@ -1391,25 +1575,25 @@ function App({ // Keep focus state consistent when switching tabs (only adjust if focus is already in tab content) useEffect(() => { if (activeTab === "messages") { - if (focus === "tabContentList" || focus === "tabContentDetails") { + if (paneFocus === "tabContentList" || paneFocus === "tabContentDetails") { setFocus("messagesList"); } } else if (activeTab === "requests") { - if (focus === "tabContentList" || focus === "tabContentDetails") { + if (paneFocus === "tabContentList" || paneFocus === "tabContentDetails") { setFocus("requestsList"); } } else { if ( - focus === "messagesList" || - focus === "messagesDetail" || - focus === "requestsList" || - focus === "requestsDetail" + paneFocus === "messagesList" || + paneFocus === "messagesDetail" || + paneFocus === "requestsList" || + paneFocus === "requestsDetail" ) { setFocus("tabContentList"); } } - // Runs on a tab switch only. `focus` is read, not reacted to: this is a - // one-time adjustment when the tab changes, and depending on `focus` + // Runs on a tab switch only. `paneFocus` is read, not reacted to: this is + // a one-time adjustment when the tab changes, and depending on it // would re-run it on every focus move, turning it into a standing // constraint on focus that no caller asked for. // eslint-disable-next-line react-hooks/exhaustive-deps -- react to tab switches, not focus moves @@ -1427,7 +1611,14 @@ function App({ useInput((input: string, key: Key) => { // Don't process input when modal is open - if (toolTestModal || resourceTestModal || promptTestModal || detailsModal) { + if ( + toolTestModal || + resourceTestModal || + promptTestModal || + detailsModal || + helpOpen || + rootsModalOpen + ) { return; } @@ -1435,6 +1626,19 @@ function App({ exit(); } + // A list filter owns the keyboard while it is being edited (#2430) — + // including `?`, which is a legitimate query character there. + if (listFilterEditing) { + return; + } + + // Open the keybinding help. It closes itself (on `?` or Esc); see + // `focus` above for how the panes underneath are kept inert meanwhile. + if (input === "?") { + setHelpOpen(true); + return; + } + // Exit accelerators if (key.escape) { exit(); @@ -1460,6 +1664,8 @@ function App({ if (tab.id === "logging" && !showLoggingTab) return false; if (tab.id === "requests" && !showRequestsTab) return false; if (tab.id === "skills" && !showSkillsTab) return false; + if (tab.id === "subscriptions" && !showSubscriptionsTab) return false; + if (tab.id === "tasks" && !showTasksTab) return false; return true; }) .map((tab: { id: TabType; label: string; accelerator: string }) => [ @@ -1474,7 +1680,13 @@ function App({ nextTab === "auth" && activeTab === "auth" && pendingStepUp?.serverName === selectedServer; - if (!authStepUpAccelerator) { + // AuthTab binds `s` to "clear OAuth state" while its pane is focused, so + // the Tasks accelerator must not also fire there (#2432). + const authClearKey = + nextTab === "tasks" && + activeTab === "auth" && + (focus === "tabContentList" || focus === "tabContentDetails"); + if (!authStepUpAccelerator && !authClearKey) { setActiveTab(nextTab); setFocus(nextTab === "auth" ? "tabContentList" : "tabs"); } @@ -1546,9 +1758,11 @@ function App({ "info", "auth", "resources", + "subscriptions", "prompts", "skills", "tools", + "tasks", "messages", "requests", "logging", @@ -1558,6 +1772,8 @@ function App({ if (t === "logging" && !showLoggingTab) return false; if (t === "requests" && !showRequestsTab) return false; if (t === "skills" && !showSkillsTab) return false; + if (t === "subscriptions" && !showSubscriptionsTab) return false; + if (t === "tasks" && !showTasksTab) return false; return true; }); const currentIndex = tabs.indexOf(activeTab); @@ -1597,26 +1813,27 @@ function App({ // terminal width — which a stdio server with Skills does at any ordinary // width — and a hard-coded 1 sized every pane below it one row too tall, // clipping the bottom of the TUI (Copilot). - const tabsHeight = tabBarRows( - visibleTabs({ - showAuth: !!( - selectedServer && - selectedServerConfig && - isOAuthCapableServerConfig(selectedServerConfig) - ), - showLogging: - !!selectedServer && - inspectorClients[selectedServer]?.getServerType() === "stdio", - showRequests: - !!selectedServer && - (inspectorClients[selectedServer]?.getServerType() === "sse" || - inspectorClients[selectedServer]?.getServerType() === - "streamable-http"), - showSkills: showSkillsTab, - }), - tabCounts, - contentWidth, - ); + // Also what the help overlay lists accelerators for: a hidden tab's letter + // does nothing, so advertising it would be wrong (Copilot). + const shownTabs = visibleTabs({ + showAuth: !!( + selectedServer && + selectedServerConfig && + isOAuthCapableServerConfig(selectedServerConfig) + ), + showLogging: + !!selectedServer && + inspectorClients[selectedServer]?.getServerType() === "stdio", + showRequests: + !!selectedServer && + (inspectorClients[selectedServer]?.getServerType() === "sse" || + inspectorClients[selectedServer]?.getServerType() === + "streamable-http"), + showSkills: showSkillsTab, + showSubscriptions: showSubscriptionsTab, + showTasks: showTasksTab, + }); + const tabsHeight = tabBarRows(shownTabs, tabCounts, contentWidth); // Server details will be flexible - calculate remaining space for content const availableHeight = dimensions.height - headerHeight - tabsHeight; // Reserve space for server details (will grow as needed, but we'll use flexGrow) @@ -1724,7 +1941,7 @@ function App({ backgroundColor="gray" > <Text bold color="white"> - ESC to exit + ? help · ESC exit </Text> </Box> </Box> @@ -1819,6 +2036,8 @@ function App({ : false } showSkills={showSkillsTab} + showSubscriptions={showSubscriptionsTab} + showTasks={showTasksTab} showRequests={ selectedServer && inspectorClients[selectedServer] ? (() => { @@ -1851,7 +2070,15 @@ function App({ width={contentWidth} height={contentHeight} focused={ - focus === "tabContentList" || focus === "tabContentDetails" + (focus === "tabContentList" || + focus === "tabContentDetails") && + !rootsModalOpen + } + roots={advertisedRoots} + onEditRoots={ + selectedInspectorClient + ? () => setRootsModalOpen(true) + : undefined } /> )} @@ -1923,7 +2150,9 @@ function App({ if (outcome.kind === "failed") { setOauthStatus("error"); setOauthMessage( - emaStepUpFailureMessage(outcome.error.message), + emaStepUpFailureMessage( + errorMessage(outcome.error), + ), ); return; } @@ -1950,10 +2179,7 @@ function App({ setOauthMessage("OAuth already in progress."); } } catch (authErr) { - const authMsg = - authErr instanceof Error - ? authErr.message - : String(authErr); + const authMsg = errorMessage(authErr); setOauthStatus("error"); setOauthMessage(authMsg); } @@ -1982,6 +2208,7 @@ function App({ selectedInspectorClient ? ( <ResourcesTab key={`resources-${selectedServer}`} + onFilterEditingChange={setListFilterEditing} resources={currentServerState.resources} resourceTemplates={currentServerState.resourceTemplates} inspectorClient={selectedInspectorClient} @@ -2001,6 +2228,7 @@ function App({ setDetailsModal({ title: `Resource: ${"uri" in resource ? resource.name || resource.uri || "Unknown" : "Resource content"}`, content: renderResourceDetails(resource), + copyText: toCopyText(resource), }) } onFetchResource={() => { @@ -2023,11 +2251,64 @@ function App({ ) } /> + ) : activeTab === "subscriptions" && + showSubscriptionsTab && + selectedInspectorClient ? ( + <SubscriptionsTab + key={`subscriptions-${selectedServer}`} + resources={managedResources} + subscriptions={resourceSubscriptions} + streamState={subscriptionStreamState} + messages={inspectorMessages} + inspectorClient={selectedInspectorClient} + width={contentWidth} + height={contentHeight} + focusedPane={ + focus === "tabContentDetails" + ? "details" + : focus === "tabContentList" + ? "list" + : null + } + modalOpen={!!detailsModal} + onAuthRecoveryRequired={onAuthRecoveryRequired} + /> + ) : activeTab === "tasks" && + showTasksTab && + selectedInspectorClient ? ( + <TasksTab + key={`tasks-${selectedServer}`} + tasks={requestorTasks} + inspectorClient={selectedInspectorClient} + width={contentWidth} + height={contentHeight} + focusedPane={ + focus === "tabContentDetails" + ? "details" + : focus === "tabContentList" + ? "list" + : null + } + modalOpen={!!detailsModal} + onRefresh={refreshRequestorTasks} + onClearCompleted={clearCompletedRequestorTasks} + onViewDetails={(task, result) => + setDetailsModal({ + title: `Task: ${task.taskId}`, + content: renderTaskDetails(task, result), + copyText: toCopyText( + result === null ? task : { task, result }, + ), + }) + } + onAuthRecoveryRequired={onAuthRecoveryRequired} + /> ) : activeTab === "skills" && currentServerState?.status === "connected" && selectedInspectorClient ? ( <SkillsTab key={`skills-${selectedServer}`} + onFilterEditingChange={setListFilterEditing} skills={managedSkills} pageCount={managedSkillsPageCount} loadError={managedSkillsError} @@ -2056,6 +2337,7 @@ function App({ selectedInspectorClient ? ( <PromptsTab key={`prompts-${selectedServer}`} + onFilterEditingChange={setListFilterEditing} prompts={currentServerState.prompts} inspectorClient={selectedInspectorClient} width={contentWidth} @@ -2074,6 +2356,7 @@ function App({ setDetailsModal({ title: `Prompt: ${prompt.name || "Unknown"}`, content: renderPromptDetails(prompt), + copyText: toCopyText(prompt), }) } onFetchPrompt={(prompt) => { @@ -2097,6 +2380,7 @@ function App({ selectedInspectorClient ? ( <ToolsTab key={`tools-${selectedServer}`} + onFilterEditingChange={setListFilterEditing} tools={currentServerState.tools} isConnected={inspectorStatus === "connected"} width={contentWidth} @@ -2121,6 +2405,7 @@ function App({ setDetailsModal({ title: `Tool: ${tool.name || "Unknown"}`, content: renderToolDetails(tool), + copyText: toCopyText(tool), }) } modalOpen={!!(toolTestModal || detailsModal)} @@ -2156,6 +2441,7 @@ function App({ setDetailsModal({ title: `Message: ${label}`, content: renderMessageDetails(message), + copyText: toCopyText(message), }); }} /> @@ -2183,6 +2469,7 @@ function App({ setDetailsModal({ title: `Request: ${request.method} ${request.url}`, content: renderRequestDetails(request), + copyText: toCopyText(request), }); }} /> @@ -2247,14 +2534,40 @@ function App({ /> )} - {/* Details Modal - rendered at App level for full screen overlay */} - {detailsModal && ( + {/* Keybinding help (#2436) - rendered at App level for full screen overlay */} + {helpOpen && ( + <HelpOverlay + sections={keybindingSections(activeTab, shownTabs)} + width={dimensions.width} + height={dimensions.height} + onClose={() => setHelpOpen(false)} + /> + )} + + {/* Roots editor (#2432) - rendered at App level for full screen overlay */} + {rootsModalOpen && ( + <RootsModal + roots={advertisedRoots} + inspectorClient={selectedInspectorClient} + connected={inspectorStatus === "connected"} + width={dimensions.width} + height={dimensions.height} + onClose={() => setRootsModalOpen(false)} + /> + )} + + {/* Details Modal - rendered at App level for full screen overlay. Held + back while the help is open: one can arrive asynchronously (a + no-argument prompt fetch completing), and both overlays would then + take the same Esc. It appears once the help closes (Copilot). */} + {detailsModal && !helpOpen && ( <DetailsModal title={detailsModal.title} content={detailsModal.content} width={dimensions.width} height={dimensions.height} onClose={() => setDetailsModal(null)} + copyText={detailsModal.copyText} /> )} </Box> diff --git a/clients/tui/src/components/AuthTab.tsx b/clients/tui/src/components/AuthTab.tsx index d2160135c4..f18bf8e2d3 100644 --- a/clients/tui/src/components/AuthTab.tsx +++ b/clients/tui/src/components/AuthTab.tsx @@ -2,6 +2,7 @@ import React, { useState, useEffect, useRef, useCallback } from "react"; import { Box, Text, useInput, type Key } from "ink"; import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; import { SelectableItem } from "./SelectableItem.js"; +import { copyStatusColor, useCopyKeys } from "../hooks/useCopyKeys.js"; import type { MCPServerConfig, InspectorClient, @@ -21,6 +22,7 @@ import { stepUpFollowUpMessage, stepUpModalTitle, } from "../utils/tuiOAuth.js"; +import { errorMessage } from "../utils/errorText.js"; interface AuthTabProps { serverName: string | null; @@ -87,9 +89,21 @@ export function AuthTab({ const isLiveConnection = connectionStatus === "connected" || connectionStatus === "connecting"; const scrollViewRef = useRef<ScrollViewRef>(null); - const [oauthState, setOauthState] = useState< - OAuthConnectionState | undefined - >(undefined); + // The OAuth state is stored with the client it was read from, and shown only + // while that client is still the selected one. AuthTab is reused across + // server switches, so without this the previous server's state — bearer + // token included — stays on screen, and copyable with Y/W, until the new + // server's read lands (#2421). + const [oauthEntry, setOauthEntry] = useState<{ + client: InspectorClient | null; + state: OAuthConnectionState | undefined; + }>({ client: null, state: undefined }); + const oauthState = + oauthEntry.client === inspectorClient ? oauthEntry.state : undefined; + const inspectorClientRef = useRef(inspectorClient); + useEffect(() => { + inspectorClientRef.current = inspectorClient; + }, [inspectorClient]); const [clearState, setClearState] = useState< "idle" | "clearing" | "cleared" | "failed" >("idle"); @@ -136,11 +150,14 @@ export function AuthTab({ const refreshOAuthState = useCallback(async () => { if (!inspectorClient) { - setOauthState(undefined); + setOauthEntry({ client: null, state: undefined }); return; } const state = await inspectorClient.getOAuthState(); - setOauthState(state); + // A read that settles after the selection moved on belongs to a server + // no longer on screen; storing it would overwrite the current one's. + if (inspectorClientRef.current !== inspectorClient) return; + setOauthEntry({ client: inspectorClient, state }); }, [inspectorClient]); useEffect(() => { @@ -174,6 +191,14 @@ export function AuthTab({ }; }, [inspectorClient, refreshOAuthState]); + const accessToken = oauthState?.tokens?.access_token; + // Y / W copy or save the full access token — the row below shows only a + // prefix, and terminal mouse-selection cannot reach the rest (#2421). + const { handleCopyKey, status: copyStatus } = useCopyKeys( + accessToken, + "access token", + ); + useInput( (input: string, key: Key) => { if (!focused) return; @@ -206,6 +231,8 @@ export function AuthTab({ return; } + if (handleCopyKey(input)) return; + if (key.upArrow && scrollViewRef.current) { scrollViewRef.current.scrollBy(-1); } else if (key.downArrow && scrollViewRef.current) { @@ -256,7 +283,7 @@ export function AuthTab({ // Left set, it would swallow the next *unrelated* revision change // and strand this banner after the OAuth state moved on. ownClearRef.current = false; - setClearFailure(err instanceof Error ? err.message : String(err)); + setClearFailure(errorMessage(err)); setClearState("failed"); }, ); @@ -274,7 +301,6 @@ export function AuthTab({ } const scopes = oauthState ? formatScopes(oauthState) : undefined; - const accessToken = oauthState?.tokens?.access_token; return ( <Box width={width} height={height} flexDirection="column" paddingX={1}> @@ -289,6 +315,14 @@ export function AuthTab({ {oauthStatus === "authenticating" && ( <Text color="yellow">Authenticating…</Text> )} + {/* A note raised mid-flow — e.g. the browser could not be opened + and the URL must be visited by hand (#2533) — has to stay + visible while the flow waits for the callback. */} + {oauthStatus === "authenticating" && oauthMessage && ( + <Text color={oauthMessageTone === "warning" ? "yellow" : "cyan"}> + {oauthMessage} + </Text> + )} {oauthStatus === "error" && oauthMessage && ( <Text color="red">{oauthMessage}</Text> )} @@ -407,6 +441,11 @@ export function AuthTab({ value={`${accessToken.slice(0, 24)}…`} /> )} + {accessToken && copyStatus && ( + <Text color={copyStatusColor(copyStatus.tone)}> + {copyStatus.message} + </Text> + )} </Box> </Box> ) : ( @@ -455,7 +494,7 @@ export function AuthTab({ <Text bold color="white"> {pendingStepUp ? "↑/↓ select, Enter confirm, A authorize, C cancel" - : `S ${isLiveConnection ? "clear+disconnect" : "clear"}, ↑/↓ scroll`} + : `S ${isLiveConnection ? "clear+disconnect" : "clear"}, ${accessToken ? "Y copy token, W save token, " : ""}↑/↓ scroll`} </Text> </Box> )} diff --git a/clients/tui/src/components/BodyLines.tsx b/clients/tui/src/components/BodyLines.tsx index 2a9d2d51d6..3656172d09 100644 --- a/clients/tui/src/components/BodyLines.tsx +++ b/clients/tui/src/components/BodyLines.tsx @@ -5,8 +5,8 @@ import { layoutBody } from "../utils/bodyLines.js"; /** * Renders a request/response body as indented, dimmed lines with a bounded * component count (#2407) — see `utils/bodyLines.ts` for the caps and why both - * exist. Shared by the Requests tab and the App details view, which previously - * carried four copies of an uncapped line map. + * exist. Shared by the Requests tab, the Protocol tab (#2539) and both of + * their App zoom views, which previously each carried an uncapped copy. */ export function BodyLines({ body, diff --git a/clients/tui/src/components/DetailsModal.tsx b/clients/tui/src/components/DetailsModal.tsx index e01b555d3c..0b8da4a899 100644 --- a/clients/tui/src/components/DetailsModal.tsx +++ b/clients/tui/src/components/DetailsModal.tsx @@ -1,6 +1,7 @@ import React, { useRef } from "react"; import { Box, Text, useInput, type Key } from "ink"; import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; +import { copyStatusColor, useCopyKeys } from "../hooks/useCopyKeys.js"; interface DetailsModalProps { title: string; @@ -8,6 +9,11 @@ interface DetailsModalProps { width: number; height: number; onClose: () => void; + /** + * The raw text behind `content`, offered to Y (copy via OSC 52) and W (save + * to a file) — #2421. Omit it and the modal offers no copy affordance. + */ + copyText?: string; } export function DetailsModal({ @@ -16,8 +22,13 @@ export function DetailsModal({ width, height, onClose, + copyText, }: DetailsModalProps) { const scrollViewRef = useRef<ScrollViewRef>(null); + const { handleCopyKey, status: copyStatus } = useCopyKeys( + copyText, + "details", + ); // Use full terminal dimensions const [terminalDimensions, setTerminalDimensions] = React.useState({ @@ -44,6 +55,8 @@ export function DetailsModal({ (input: string, key: Key) => { if (key.escape) { onClose(); + } else if (handleCopyKey(input)) { + return; } else if (key.downArrow) { scrollViewRef.current?.scrollBy(1); } else if (key.upArrow) { @@ -89,9 +102,21 @@ export function DetailsModal({ {title} </Text> <Text> </Text> - <Text dimColor>(Press ESC to close)</Text> + <Text dimColor> + {copyText === undefined + ? "(Press ESC to close)" + : "(Press ESC to close, Y to copy, W to save to a file)"} + </Text> </Box> + {copyStatus && ( + <Box flexShrink={0} marginBottom={1}> + <Text color={copyStatusColor(copyStatus.tone)}> + {copyStatus.message} + </Text> + </Box> + )} + {/* Content Area */} <Box flexGrow={1} flexDirection="column" overflow="hidden"> <ScrollView ref={scrollViewRef}>{content}</ScrollView> diff --git a/clients/tui/src/components/HelpOverlay.tsx b/clients/tui/src/components/HelpOverlay.tsx new file mode 100644 index 0000000000..5d8d9ccd79 --- /dev/null +++ b/clients/tui/src/components/HelpOverlay.tsx @@ -0,0 +1,107 @@ +/** + * The `?` keybinding help overlay (#2436). + * + * A full-screen overlay, drawn the same way as `DetailsModal`, listing the + * bindings `utils/keybindings.ts` returns for the active tab. It renders data + * and owns no knowledge of which keys exist: a new binding is a row in that + * table, never an edit here. + * + * It closes on `?` (the key that opened it) or Esc, and scrolls when the list + * is taller than the terminal. Keeping the panes underneath inert while it is + * open is `App`'s job — it moves focus off them — because the overlay cannot + * stop other `useInput` handlers from seeing a key. + */ +import React, { useRef } from "react"; +import { Box, Text, useInput, type Key } from "ink"; +import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; +import { + keyColumnWidth, + type KeyBindingSection, +} from "../utils/keybindings.js"; + +interface HelpOverlayProps { + sections: readonly KeyBindingSection[]; + width: number; + height: number; + onClose: () => void; +} + +/** Gap between the keys column and the action column. */ +const COLUMN_GAP = 2; + +export function HelpOverlay({ + sections, + width, + height, + onClose, +}: HelpOverlayProps) { + const scrollViewRef = useRef<ScrollViewRef>(null); + const keysWidth = keyColumnWidth(sections) + COLUMN_GAP; + + useInput((input: string, key: Key) => { + // The ref is set once the ScrollView mounts; `?.` short-circuits the + // whole call, arguments included, on the render before that. + const view = scrollViewRef.current; + if (key.escape || input === "?") { + onClose(); + } else if (key.downArrow) { + view?.scrollBy(1); + } else if (key.upArrow) { + view?.scrollBy(-1); + } else if (key.pageDown) { + view?.scrollBy(view.getViewportHeight()); + } else if (key.pageUp) { + view?.scrollBy(-view.getViewportHeight()); + } + }); + + return ( + <Box + position="absolute" + width={width} + height={height} + flexDirection="column" + justifyContent="center" + alignItems="center" + > + <Box + width={width - 2} + height={height - 2} + borderStyle="single" + borderColor="cyan" + flexDirection="column" + paddingX={1} + paddingY={1} + backgroundColor="black" + > + <Box flexShrink={0} marginBottom={1}> + <Text bold color="cyan"> + Keyboard shortcuts + </Text> + <Text> </Text> + <Text dimColor>(Press ? or ESC to close)</Text> + </Box> + + <Box flexGrow={1} flexDirection="column" overflow="hidden"> + <ScrollView ref={scrollViewRef}> + {sections.map((section) => ( + <Box key={section.title} flexDirection="column" marginBottom={1}> + <Text bold underline> + {section.title} + </Text> + {section.bindings.map((binding) => ( + <Box key={`${binding.keys}:${binding.action}`}> + <Box width={keysWidth} flexShrink={0}> + <Text color="yellow">{binding.keys}</Text> + </Box> + <Text>{binding.action}</Text> + </Box> + ))} + </Box> + ))} + </ScrollView> + </Box> + </Box> + </Box> + ); +} diff --git a/clients/tui/src/components/HistoryTab.tsx b/clients/tui/src/components/HistoryTab.tsx index 6798203e1d..e594e3d30a 100644 --- a/clients/tui/src/components/HistoryTab.tsx +++ b/clients/tui/src/components/HistoryTab.tsx @@ -3,6 +3,7 @@ import { Box, Text, useInput, type Key } from "ink"; import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; import type { MessageEntry } from "@inspector/core/mcp/index.js"; import { useSelectableList } from "../hooks/useSelectableList.js"; +import { BodyLines } from "./BodyLines.js"; interface HistoryTabProps { serverName: string | null; @@ -237,18 +238,10 @@ export function HistoryTab({ </Box> {/* Request content */} - {JSON.stringify(selectedMessage.message, null, 2) - .split("\n") - .map((line: string, idx: number) => ( - <Box - key={`req-${idx}`} - marginTop={idx === 0 ? 1 : 0} - paddingLeft={2} - flexShrink={0} - > - <Text dimColor>{line}</Text> - </Box> - ))} + <BodyLines + body={JSON.stringify(selectedMessage.message)} + keyPrefix="req" + /> {/* Response section */} {selectedMessage.response ? ( @@ -256,18 +249,10 @@ export function HistoryTab({ <Box marginTop={1} flexShrink={0}> <Text bold>Response:</Text> </Box> - {JSON.stringify(selectedMessage.response, null, 2) - .split("\n") - .map((line: string, idx: number) => ( - <Box - key={`resp-${idx}`} - marginTop={idx === 0 ? 1 : 0} - paddingLeft={2} - flexShrink={0} - > - <Text dimColor>{line}</Text> - </Box> - ))} + <BodyLines + body={JSON.stringify(selectedMessage.response)} + keyPrefix="resp" + /> </> ) : ( <Box marginTop={1} flexShrink={0}> @@ -289,18 +274,10 @@ export function HistoryTab({ </Box> {/* Message content */} - {JSON.stringify(selectedMessage.message, null, 2) - .split("\n") - .map((line: string, idx: number) => ( - <Box - key={`msg-${idx}`} - marginTop={idx === 0 ? 1 : 0} - paddingLeft={2} - flexShrink={0} - > - <Text dimColor>{line}</Text> - </Box> - ))} + <BodyLines + body={JSON.stringify(selectedMessage.message)} + keyPrefix="msg" + /> </> )} </ScrollView> diff --git a/clients/tui/src/components/InfoTab.tsx b/clients/tui/src/components/InfoTab.tsx index 6d7dbadd04..6b3b652eea 100644 --- a/clients/tui/src/components/InfoTab.tsx +++ b/clients/tui/src/components/InfoTab.tsx @@ -6,6 +6,7 @@ import type { ServerState, } from "@inspector/core/mcp/index.js"; import type { InspectorServerSettings } from "@inspector/core/mcp/types.js"; +import type { Root } from "@modelcontextprotocol/client"; interface InfoTabProps { serverName: string | null; @@ -19,6 +20,10 @@ interface InfoTabProps { width: number; height: number; focused?: boolean; + /** The roots the client currently advertises to this server (#2432). */ + roots?: Root[]; + /** Open the roots editor — bound to `e` while this pane is focused. */ + onEditRoots?: () => void; } export function InfoTab({ @@ -29,6 +34,8 @@ export function InfoTab({ width, height, focused = false, + roots = [], + onEditRoots, }: InfoTabProps) { const headerPairs = serverSettings?.headers ?? []; // Shared header display for the sse / streamable-http branches (identical for @@ -48,7 +55,9 @@ export function InfoTab({ useInput( (input: string, key: Key) => { if (focused) { - if (key.upArrow) { + if (input === "e" && onEditRoots) { + onEditRoots(); + } else if (key.upArrow) { scrollViewRef.current?.scrollBy(-1); } else if (key.downArrow) { scrollViewRef.current?.scrollBy(1); @@ -202,6 +211,27 @@ export function InfoTab({ <Text dimColor>Server not connected</Text> </Box> )} + {/* Roots advertised to the server (#2432) */} + <Box flexShrink={0} marginTop={2}> + <Text bold>Roots ({roots.length})</Text> + </Box> + <Box + flexShrink={0} + marginTop={1} + paddingLeft={2} + flexDirection="column" + > + {roots.length === 0 ? ( + <Text dimColor>None</Text> + ) : ( + roots.map((root, idx) => ( + <Text key={`root-${idx}`} dimColor> + {root.uri} + {root.name ? ` (${root.name})` : ""} + </Text> + )) + )} + </Box> </ScrollView> </Box> @@ -215,6 +245,7 @@ export function InfoTab({ > <Text bold color="white"> ↑/↓ to scroll, + to zoom + {onEditRoots ? ", e to edit roots" : ""} </Text> </Box> )} diff --git a/clients/tui/src/components/ListFilterBar.tsx b/clients/tui/src/components/ListFilterBar.tsx new file mode 100644 index 0000000000..0bfe13829f --- /dev/null +++ b/clients/tui/src/components/ListFilterBar.tsx @@ -0,0 +1,68 @@ +/** + * The one-row filter line under a list pane's heading (#2430) — the visible half + * of `useListFilter`, shared by the Tools, Resources, Prompts and Skills panes + * so all four read the same way. + * + * The row is **always reserved**, even when it is blank, so opening or clearing + * a filter never changes how many list rows fit: a list that grew and shrank + * by one row as the user pressed `/` would shift the selection's on-screen + * position under the cursor. `LIST_FILTER_ROWS` is that reservation, which each + * pane subtracts from its visible-row budget. + * + * What it shows, in priority order: + * + * - **editing** — the query with a cursor block, and the two ways out; + * - **a kept filter** — the query, and how to change it; + * - **the list is focused** — a dim `/ to filter`, so the feature is + * discoverable from the pane itself rather than only from documentation; + * - otherwise nothing. + */ +import React from "react"; +import { Box, Text } from "ink"; + +/** Rows the filter line occupies; subtract from a pane's visible-row count. */ +export const LIST_FILTER_ROWS = 1; + +interface ListFilterBarProps { + query: string; + editing: boolean; + /** Whether the owning list has the keyboard, which is when the hint shows. */ + focused: boolean; +} + +export function ListFilterBar({ query, editing, focused }: ListFilterBarProps) { + return ( + <Box height={LIST_FILTER_ROWS} flexShrink={0} overflow="hidden"> + {editing ? ( + <Text wrap="truncate-end"> + <Text color="cyan">/</Text> + {query} + <Text inverse> </Text> + <Text dimColor> Enter keep · Esc clear</Text> + </Text> + ) : query.trim() !== "" ? ( + <Text wrap="truncate-end"> + <Text color="cyan">/</Text> + {query} + <Text dimColor> (/ to edit)</Text> + </Text> + ) : focused ? ( + <Text dimColor>/ to filter</Text> + ) : ( + <Text> </Text> + )} + </Box> + ); +} + +/** + * The heading count: the total alone, or `matches/total` while a filter + * narrows the list, so the user can always see how much is hidden. + */ +export function filterCount( + active: boolean, + shown: number, + total: number, +): string { + return active ? `${shown}/${total}` : String(total); +} diff --git a/clients/tui/src/components/PromptTestModal.tsx b/clients/tui/src/components/PromptTestModal.tsx index c48fdd8a03..a78eec4927 100644 --- a/clients/tui/src/components/PromptTestModal.tsx +++ b/clients/tui/src/components/PromptTestModal.tsx @@ -6,6 +6,7 @@ import { AuthRecoveryRequiredError } from "@inspector/core/auth/challenge.js"; import type { Prompt, GetPromptResult } from "@modelcontextprotocol/client"; import { promptArgsToForm } from "../utils/promptArgsToForm.js"; import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; +import { redactErrorText, redactedJson } from "../utils/errorText.js"; // Helper to extract error message from various error types function getErrorMessage(error: unknown): string { @@ -267,7 +268,9 @@ export function PromptTestModal({ </Text> </Box> <Box marginTop={1} paddingLeft={2} flexShrink={0}> - <Text color="red">{String(result.error)}</Text> + <Text color="red"> + {redactErrorText(String(result.error))} + </Text> </Box> {result.errorDetails != null ? ( <> @@ -278,7 +281,7 @@ export function PromptTestModal({ </Box> <Box marginTop={1} paddingLeft={2} flexShrink={0}> <Text dimColor> - {JSON.stringify(result.errorDetails, null, 2)} + {redactedJson(result.errorDetails)} </Text> </Box> </> diff --git a/clients/tui/src/components/PromptsTab.tsx b/clients/tui/src/components/PromptsTab.tsx index 46db93b567..c43008bdb4 100644 --- a/clients/tui/src/components/PromptsTab.tsx +++ b/clients/tui/src/components/PromptsTab.tsx @@ -9,6 +9,13 @@ import type { GetPromptResult, } from "@modelcontextprotocol/client"; import { useSelectableList } from "../hooks/useSelectableList.js"; +import { errorMessage } from "../utils/errorText.js"; +import { useListFilter } from "../hooks/useListFilter.js"; +import { + ListFilterBar, + LIST_FILTER_ROWS, + filterCount, +} from "./ListFilterBar.js"; interface PromptsTabProps { prompts: Prompt[]; @@ -21,6 +28,13 @@ interface PromptsTabProps { onFetchPrompt?: (prompt: Prompt) => void; onAuthRecoveryRequired?: (error: AuthRecoveryRequiredError) => void; modalOpen?: boolean; + /** Told when the list filter starts or stops capturing keys (#2430). */ + onFilterEditingChange?: (editing: boolean) => void; +} + +/** What the `/` filter matches a prompt against (#2430). */ +function promptFilterFields(prompt: Prompt): ReadonlyArray<string | undefined> { + return [prompt.name, prompt.title]; } export function PromptsTab({ @@ -33,12 +47,18 @@ export function PromptsTab({ onFetchPrompt, onAuthRecoveryRequired, modalOpen = false, + onFilterEditingChange, }: PromptsTabProps) { - const visibleCount = Math.max(1, height - 7); + const visibleCount = Math.max(1, height - 7 - LIST_FILTER_ROWS); + const filter = useListFilter(prompts, promptFilterFields, { + enabled: !modalOpen && focusedPane === "list", + onEditingChange: onFilterEditingChange, + }); + const shownPrompts = filter.items; const { selectedIndex, firstVisible, setSelection } = useSelectableList( - prompts.length, + shownPrompts.length, visibleCount, - { resetWhen: prompts }, + { resetWhen: shownPrompts }, ); const [error, setError] = useState<string | null>(null); const scrollViewRef = useRef<ScrollViewRef>(null); @@ -46,6 +66,10 @@ export function PromptsTab({ // Handle arrow key navigation when focused useInput( (input: string, key: Key) => { + // The filter goes first: while it is editing, Enter and letters belong + // to the query, not to the list. + if (filter.handleInput(input, key)) return; + // Handle Enter key to fetch prompt (works from both list and details) if (key.return && selectedPrompt && inspectorClient && onFetchPrompt) { // If prompt has arguments, open modal to collect them @@ -74,7 +98,9 @@ export function PromptsTab({ return; } setError( - error instanceof Error ? error.message : "Failed to get prompt", + error instanceof Error + ? errorMessage(error) + : "Failed to get prompt", ); } })(); @@ -85,7 +111,7 @@ export function PromptsTab({ if (focusedPane === "list") { if (key.upArrow && selectedIndex > 0) { setSelection(selectedIndex - 1); - } else if (key.downArrow && selectedIndex < prompts.length - 1) { + } else if (key.downArrow && selectedIndex < shownPrompts.length - 1) { setSelection(selectedIndex + 1); } return; @@ -125,7 +151,7 @@ export function PromptsTab({ scrollViewRef.current?.scrollTo(0); }, [selectedIndex]); - const selectedPrompt = prompts[selectedIndex] || null; + const selectedPrompt = shownPrompts[selectedIndex] || null; const listWidth = Math.floor(width * 0.4); const detailWidth = width - listWidth; @@ -149,9 +175,15 @@ export function PromptsTab({ bold backgroundColor={focusedPane === "list" ? "yellow" : undefined} > - Prompts ({prompts.length}) + Prompts ( + {filterCount(filter.active, shownPrompts.length, prompts.length)}) </Text> </Box> + <ListFilterBar + query={filter.query} + editing={filter.editing} + focused={focusedPane === "list" && !modalOpen} + /> {error ? ( <Box paddingY={1}> <Text color="red">{error}</Text> @@ -160,6 +192,10 @@ export function PromptsTab({ <Box paddingY={1}> <Text dimColor>No prompts available</Text> </Box> + ) : shownPrompts.length === 0 ? ( + <Box paddingY={1}> + <Text dimColor>No prompts match the filter</Text> + </Box> ) : ( <Box flexDirection="column" @@ -167,16 +203,22 @@ export function PromptsTab({ overflow="hidden" flexShrink={0} > - {prompts + {shownPrompts .slice(firstVisible, firstVisible + visibleCount) .map((prompt, i) => { const index = firstVisible + i; const isSelected = index === selectedIndex; + // Unfiltered position: unique even across repeated names. + const ordinal = filter.indices[index]; return ( - <Box key={prompt.name || index} paddingY={0} flexShrink={0}> + <Box + key={`${ordinal}:${prompt.name}`} + paddingY={0} + flexShrink={0} + > <Text> {isSelected ? "▶ " : " "} - {prompt.name || `Prompt ${index + 1}`} + {prompt.name || `Prompt ${ordinal + 1}`} </Text> </Box> ); diff --git a/clients/tui/src/components/ResourceTestModal.tsx b/clients/tui/src/components/ResourceTestModal.tsx index 87c3ddeb2d..1bbf305828 100644 --- a/clients/tui/src/components/ResourceTestModal.tsx +++ b/clients/tui/src/components/ResourceTestModal.tsx @@ -11,6 +11,7 @@ import { unmetRequiredGroups, } from "@inspector/core/mcp/uriTemplate.js"; import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; +import { redactErrorText, redactedJson } from "../utils/errorText.js"; // Helper to extract error message from various error types function getErrorMessage(error: unknown): string { @@ -326,7 +327,9 @@ export function ResourceTestModal({ </Text> </Box> <Box marginTop={1} paddingLeft={2} flexShrink={0}> - <Text color="red">{String(result.error)}</Text> + <Text color="red"> + {redactErrorText(String(result.error))} + </Text> </Box> {result.errorDetails != null ? ( <> @@ -337,7 +340,7 @@ export function ResourceTestModal({ </Box> <Box marginTop={1} paddingLeft={2} flexShrink={0}> <Text dimColor> - {JSON.stringify(result.errorDetails, null, 2)} + {redactedJson(result.errorDetails)} </Text> </Box> </> diff --git a/clients/tui/src/components/ResourcesTab.tsx b/clients/tui/src/components/ResourcesTab.tsx index 4d0fa05dbb..689c4babf5 100644 --- a/clients/tui/src/components/ResourcesTab.tsx +++ b/clients/tui/src/components/ResourcesTab.tsx @@ -8,9 +8,17 @@ import type { ReadResourceResult, } from "@modelcontextprotocol/client"; import { useSelectableList } from "../hooks/useSelectableList.js"; +import { errorMessage } from "../utils/errorText.js"; +import { useListFilter } from "../hooks/useListFilter.js"; +import { + ListFilterBar, + LIST_FILTER_ROWS, + filterCount, +} from "./ListFilterBar.js"; interface ResourceTemplate { name: string; + title?: string; uriTemplate: string; description?: string; } @@ -30,6 +38,24 @@ interface ResourcesTabProps { onFetchTemplate?: (template: ResourceTemplate) => void; onAuthRecoveryRequired?: (error: AuthRecoveryRequiredError) => void; modalOpen?: boolean; + /** Told when the list filter starts or stops capturing keys (#2430). */ + onFilterEditingChange?: (editing: boolean) => void; +} + +type ResourceListItem = + | { type: "resource"; data: Resource } + | { type: "template"; data: ResourceTemplate }; + +/** + * What the `/` filter matches a row against (#2430): its names and its URI, + * since a resource is as often recognized by its URI as by its name. + */ +function resourceFilterFields( + item: ResourceListItem, +): ReadonlyArray<string | undefined> { + return item.type === "resource" + ? [item.data.name, item.data.title, item.data.uri] + : [item.data.name, item.data.title, item.data.uriTemplate]; } export function ResourcesTab({ @@ -45,6 +71,7 @@ export function ResourcesTab({ onFetchTemplate, onAuthRecoveryRequired, modalOpen = false, + onFilterEditingChange, }: ResourcesTabProps) { const [error, setError] = useState<string | null>(null); const [resourceContent, setResourceContent] = @@ -57,7 +84,7 @@ export function ResourcesTab({ // Combined list: resources first, then templates - memoized to prevent unnecessary recalculations const allItems = useMemo( - () => [ + (): ResourceListItem[] => [ ...resources.map((r) => ({ type: "resource" as const, data: r })), ...resourceTemplates.map((t) => ({ type: "template" as const, data: t })), ], @@ -68,20 +95,29 @@ export function ResourcesTab({ [resources.length, resourceTemplates.length], ); - const visibleCount = Math.max(1, height - 7); + const visibleCount = Math.max(1, height - 7 - LIST_FILTER_ROWS); + const filter = useListFilter(allItems, resourceFilterFields, { + enabled: !modalOpen && focusedPane === "list", + onEditingChange: onFilterEditingChange, + }); + const shownItems = filter.items; const { selectedIndex, firstVisible, setSelection } = useSelectableList( - totalCount, + shownItems.length, visibleCount, - { resetWhen: resources }, + { resetWhen: shownItems }, ); const selectedItem = useMemo( - () => allItems[selectedIndex] || null, - [allItems, selectedIndex], + () => shownItems[selectedIndex] || null, + [shownItems, selectedIndex], ); // Handle arrow key navigation when focused useInput( (input: string, key: Key) => { + // The filter goes first: while it is editing, Enter and letters belong + // to the query, not to the list. + if (filter.handleInput(input, key)) return; + // Handle Enter key to fetch resource (works from both list and details) if ( key.return && @@ -105,7 +141,7 @@ export function ResourcesTab({ if (focusedPane === "list") { if (key.upArrow && selectedIndex > 0) { setSelection(selectedIndex - 1); - } else if (key.downArrow && selectedIndex < totalCount - 1) { + } else if (key.downArrow && selectedIndex < shownItems.length - 1) { setSelection(selectedIndex + 1); } return; @@ -177,7 +213,7 @@ export function ResourcesTab({ return; } setError( - err instanceof Error ? err.message : "Failed to read resource", + err instanceof Error ? errorMessage(err) : "Failed to read resource", ); setResourceContent(null); } finally { @@ -222,9 +258,15 @@ export function ResourcesTab({ bold backgroundColor={focusedPane === "list" ? "yellow" : undefined} > - Resources ({totalCount}) + Resources ( + {filterCount(filter.active, shownItems.length, totalCount)}) </Text> </Box> + <ListFilterBar + query={filter.query} + editing={filter.editing} + focused={focusedPane === "list" && !modalOpen} + /> {error ? ( <Box paddingY={1}> <Text color="red">{error}</Text> @@ -233,6 +275,10 @@ export function ResourcesTab({ <Box paddingY={1}> <Text dimColor>No resources available</Text> </Box> + ) : shownItems.length === 0 ? ( + <Box paddingY={1}> + <Text dimColor>No resources match the filter</Text> + </Box> ) : ( <Box flexDirection="column" @@ -240,20 +286,29 @@ export function ResourcesTab({ overflow="hidden" flexShrink={0} > - {allItems + {shownItems .slice(firstVisible, firstVisible + visibleCount) .map((item, i) => { const index = firstVisible + i; const isSelected = index === selectedIndex; + // Fallback ordinals count from the UNFILTERED list, so a row + // keeps its number when a filter hides its neighbours. + const ordinal = filter.indices[index]; const label = item.type === "resource" - ? item.data.name || item.data.uri || `Resource ${index + 1}` + ? item.data.name || + item.data.uri || + `Resource ${ordinal + 1}` : item.data.name || - `Template ${index - resources.length + 1}`; - const key = + `Template ${ordinal - resources.length + 1}`; + // Keyed by the unfiltered position too: a server can repeat a + // URI, and colliding keys would let React reuse the wrong row + // once a filter hides one of the duplicates (Copilot). + const key = `${ordinal}:${ item.type === "resource" - ? item.data.uri || index - : item.data.uriTemplate || index; + ? item.data.uri + : item.data.uriTemplate + }`; return ( <Box key={key} paddingY={0} flexShrink={0}> <Text> diff --git a/clients/tui/src/components/RootsModal.tsx b/clients/tui/src/components/RootsModal.tsx new file mode 100644 index 0000000000..1514329ceb --- /dev/null +++ b/clients/tui/src/components/RootsModal.tsx @@ -0,0 +1,237 @@ +/** + * Roots editor (#2432): list, add and remove the roots this client advertises + * to the selected server. + * + * A modal rather than a tab because adding a root needs free-text input, and + * only a modal suppresses App's global accelerators (a typed `r` would + * otherwise jump to Resources). It is opened from the Info tab with `e`. + * + * Every change goes through `InspectorClient.setRoots`, which normalizes the + * list with `cleanRoots` and sends `notifications/roots/list_changed`, so the + * server re-requests `roots/list` — the same path the web client's settings + * save takes. The change is for this session only: the TUI does not write + * `mcp.json`, so the configured roots return on the next launch. + */ +import React, { useRef, useState } from "react"; +import { Box, Text, useInput, type Key } from "ink"; +import { Form, type FormStructure } from "ink-form"; +import type { Root } from "@modelcontextprotocol/client"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import { useSelectableList } from "../hooks/useSelectableList.js"; +import { errorMessage } from "../utils/errorText.js"; + +export const ADD_ROOT_FORM: FormStructure = { + title: "Add Root", + sections: [ + { + title: "Root", + fields: [ + { + name: "uri", + label: "URI (e.g. file:///path)", + type: "string", + required: true, + }, + { name: "name", label: "Name (optional)", type: "string" }, + ], + }, + ], +}; + +/** Build the root a submitted form describes, or explain why it can't. */ +export function rootFromForm( + values: Record<string, unknown>, +): Root | { error: string } { + const uri = typeof values.uri === "string" ? values.uri.trim() : ""; + if (!uri) return { error: "A root needs a URI" }; + const name = typeof values.name === "string" ? values.name.trim() : ""; + return name ? { uri, name } : { uri }; +} + +interface RootsModalProps { + roots: Root[]; + inspectorClient: InspectorClient | null; + connected: boolean; + width: number; + height: number; + onClose: () => void; +} + +export function RootsModal({ + roots, + inspectorClient, + connected, + width, + height, + onClose, +}: RootsModalProps) { + const [mode, setMode] = useState<"list" | "add">("list"); + const [error, setError] = useState<string | null>(null); + const [saving, setSaving] = useState(false); + const modalWidth = width - 2; + const modalHeight = height - 2; + const visibleCount = Math.max(1, modalHeight - 8); + const { selectedIndex, firstVisible, setSelection } = useSelectableList( + roots.length, + visibleCount, + ); + + // `saving` drives the status line; this ref is the guard, because two keys + // can land before React re-renders and both would save from the same stale + // `roots`, announcing roots/list_changed twice for one edit. + const savingRef = useRef(false); + + const save = async (next: Root[]) => { + if (!inspectorClient || savingRef.current) return; + savingRef.current = true; + setSaving(true); + setError(null); + try { + await inspectorClient.setRoots(next); + setMode("list"); + } catch (err) { + setError(errorMessage(err)); + } finally { + savingRef.current = false; + setSaving(false); + } + }; + + const handleSubmit = (values: Record<string, unknown>) => { + const root = rootFromForm(values); + if ("error" in root) { + setError(root.error); + setMode("list"); + return; + } + // `save` owns every rejection (its catch surfaces the message). + void save([...roots, root]); + }; + + useInput((input: string, key: Key) => { + if (key.escape) { + // Esc backs out of the form first, and closes the editor from the list. + if (mode === "add") { + setMode("list"); + } else { + onClose(); + } + return; + } + // The form owns every other key while it is open. + if (mode === "add" || saving) return; + + if (key.upArrow && selectedIndex > 0) { + setSelection(selectedIndex - 1); + } else if (key.downArrow && selectedIndex < roots.length - 1) { + setSelection(selectedIndex + 1); + } else if (input === "n" || input === "+") { + if (!connected) { + setError("Connect to the server to change its roots"); + return; + } + setError(null); + setMode("add"); + } else if ((input === "x" || key.delete) && roots[selectedIndex]) { + if (!connected) { + setError("Connect to the server to change its roots"); + return; + } + void save(roots.filter((_, i) => i !== selectedIndex)); + } + }); + + return ( + <Box + position="absolute" + width={width} + height={height} + flexDirection="column" + justifyContent="center" + alignItems="center" + > + <Box + width={modalWidth} + height={modalHeight} + borderStyle="single" + borderColor="cyan" + flexDirection="column" + paddingX={1} + paddingY={1} + backgroundColor="black" + > + <Box flexShrink={0} marginBottom={1}> + <Text bold color="cyan"> + Roots ({roots.length}) + </Text> + <Text> </Text> + <Text dimColor> + {mode === "add" ? "(ESC to go back)" : "(ESC to close)"} + </Text> + </Box> + + {mode === "add" ? ( + <Box flexGrow={1} flexDirection="column"> + <Form + form={ADD_ROOT_FORM} + onSubmit={(values: object) => + handleSubmit(values as Record<string, unknown>) + } + /> + </Box> + ) : ( + <Box flexGrow={1} flexDirection="column" overflow="hidden"> + {roots.length === 0 ? ( + <Text dimColor>No roots are advertised to this server</Text> + ) : ( + roots + .slice(firstVisible, firstVisible + visibleCount) + .map((root, i) => { + const index = firstVisible + i; + return ( + <Box key={`${root.uri}-${index}`} flexShrink={0}> + <Text wrap="truncate"> + {index === selectedIndex ? "▶ " : " "} + {root.uri} + {root.name && <Text dimColor> ({root.name})</Text>} + </Text> + </Box> + ); + }) + )} + </Box> + )} + + {saving && ( + <Box flexShrink={0}> + <Text color="yellow">Updating roots…</Text> + </Box> + )} + {error && ( + <Box flexShrink={0}> + <Text color="red">{error}</Text> + </Box> + )} + {!connected && ( + <Box flexShrink={0}> + <Text dimColor> + Read-only while disconnected — connect to edit. + </Text> + </Box> + )} + {mode === "list" && ( + <Box + flexShrink={0} + height={1} + justifyContent="center" + backgroundColor="gray" + > + <Text bold color="white" wrap="truncate"> + n to add, x to remove, ↑/↓ to select. Changes last this session. + </Text> + </Box> + )} + </Box> + </Box> + ); +} diff --git a/clients/tui/src/components/SaveResultBar.tsx b/clients/tui/src/components/SaveResultBar.tsx new file mode 100644 index 0000000000..b678358d29 --- /dev/null +++ b/clients/tui/src/components/SaveResultBar.tsx @@ -0,0 +1,58 @@ +/** + * The bottom line of the tool result view while saving a result (#2571): the + * filename prompt `w` opens, or — once it closes — the outcome of the save. + * + * Its own component, rather than inline in `ToolTestModal`, because the modal + * renders `position="absolute"`, which ink-testing-library lays out as an empty + * frame. Split out, the prompt text and the confirmation / failure line can be + * asserted on directly instead of only through their side effects. + */ +import React from "react"; +import { Box, Text } from "ink"; +import type { ResultFileFormat } from "@inspector/core/mcp/resultFile.js"; + +/** The open `w` filename prompt. */ +export interface SavePrompt { + path: string; + format: ResultFileFormat; + /** Whether the user has typed into the path, so Tab stops rewriting it. */ + edited: boolean; +} + +/** The outcome line shown after a save attempt. */ +export interface SaveStatus { + ok: boolean; + message: string; +} + +interface SaveResultBarProps { + prompt: SavePrompt | null; + status: SaveStatus | null; +} + +export function SaveResultBar({ prompt, status }: SaveResultBarProps) { + if (prompt) { + return ( + <Box flexShrink={0} flexDirection="column"> + <Text> + <Text bold color="cyan"> + Save as {prompt.format}:{" "} + </Text> + <Text>{prompt.path}</Text> + <Text inverse> </Text> + </Text> + <Text dimColor> + Enter to save, Tab to switch json/raw, ESC to cancel + </Text> + </Box> + ); + } + if (status) { + return ( + <Box flexShrink={0}> + <Text color={status.ok ? "green" : "red"}>{status.message}</Text> + </Box> + ); + } + return null; +} diff --git a/clients/tui/src/components/SkillsTab.tsx b/clients/tui/src/components/SkillsTab.tsx index 1f3fee30da..3eb3863f97 100644 --- a/clients/tui/src/components/SkillsTab.tsx +++ b/clients/tui/src/components/SkillsTab.tsx @@ -18,7 +18,13 @@ * that gate, and the reset that leaves the tab when it goes false, live in * `App.tsx` because they are navigation concerns rather than this pane's. */ -import React, { useCallback, useEffect, useRef, useState } from "react"; +import React, { + useCallback, + useEffect, + useLayoutEffect, + useRef, + useState, +} from "react"; import { Box, Text, useInput, type Key } from "ink"; import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; import type { InspectorClient } from "@inspector/core/mcp/index.js"; @@ -41,6 +47,13 @@ import { type SkillVerifyReport, } from "@inspector/core/mcp/skillsVerification.js"; import { useSelectableList } from "../hooks/useSelectableList.js"; +import { errorMessage } from "../utils/errorText.js"; +import { useListFilter } from "../hooks/useListFilter.js"; +import { + ListFilterBar, + LIST_FILTER_ROWS, + filterCount, +} from "./ListFilterBar.js"; interface SkillsTabProps { skills: SkillEntry[]; @@ -54,6 +67,18 @@ interface SkillsTabProps { focusedPane?: "list" | "details" | null; onAuthRecoveryRequired?: (error: AuthRecoveryRequiredError) => void; modalOpen?: boolean; + /** Told when the list filter starts or stops capturing keys (#2430). */ + onFilterEditingChange?: (editing: boolean) => void; +} + +/** + * What the `/` filter matches a skill against (#2430): the name the row + * shows, and the declared frontmatter name when that differs from it. + */ +function skillFilterFields( + skill: SkillEntry, +): ReadonlyArray<string | undefined> { + return [skillDisplayName(skill), skill.frontmatter.name]; } /** @@ -119,6 +144,80 @@ function failureDetail(file: SkillFileReport): string | undefined { return `expected ${short(file.expectedDigest)}, got ${short(file.actualDigest)}`; } +/** + * Fold one run's reports into the held set (#2590). + * + * ⚠️ **The first occurrence of a key in a run wins.** A malformed listing can + * repeat an entry verbatim, and both copies share a `skillEntryKey`. Under a + * catalog budget the walk reads the first and stops before the second, so a + * last-write-wins merge replaced a real verdict — a digest failure, say — with + * the second copy's unread `incomplete` (Copilot). Identical entries are served + * identical bytes, so the first, read occurrence speaks for both. + * + * Verdicts for entries no longer in the listing (`live`) are dropped — held + * ones and this run's alike — so a pane left open across refreshes does not + * accumulate one per entry snapshot. + */ +function mergeReports( + previous: ReadonlyMap<string, SkillVerifyReport>, + entries: readonly SkillEntry[], + results: readonly SkillVerifyReport[], + live: ReadonlySet<string>, +): Map<string, SkillVerifyReport> { + const next = new Map([...previous].filter(([key]) => live.has(key))); + const seen = new Set<string>(); + results.forEach((result, index) => { + const key = skillEntryKey(entries[index]!); + // `live` filters this run's results too: an entry a refresh removed while + // the run was in flight would otherwise be cached, and its verdict would + // reappear unverified if the entry came back (Copilot). + if (seen.has(key) || !live.has(key)) return; + seen.add(key); + next.set(key, result); + }); + return next; +} + +/** + * The catalog line under the list (#2590): the hint for `v` until something has + * been verified, then a tally of the verdicts held for the CURRENT listing. + * + * Derived from the report map on every render rather than stored by the + * verify-all run, so a refresh that changes an entry drops its verdict from the + * tally the same way it drops it from the detail pane — a stored summary would + * keep counting verdicts for entries that are no longer what the server lists. + */ +function catalogSummary( + skills: readonly SkillEntry[], + reports: ReadonlyMap<string, SkillVerifyReport>, +): string { + const counts: Record<SkillVerifyReport["outcome"], number> = { + verified: 0, + failed: 0, + incomplete: 0, + unverifiable: 0, + }; + let unchecked = 0; + for (const skill of skills) { + const report = reports.get(skillEntryKey(skill)); + if (report) counts[report.outcome] += 1; + else unchecked += 1; + } + if (unchecked === skills.length) return "[v: verify all]"; + // Compact, because at an 80-column terminal the list pane has 19 columns of + // text (Copilot). The glyphs are the pane's own: `✓`/`✗`/`?` as in the + // manifest rows, `·` for not looked at, and `…` for cut short by the budget. + // Verified and failed always show; the rest only when non-zero. + const parts = [ + `✓${counts.verified}`, + `✗${counts.failed}`, + ...(counts.incomplete > 0 ? [`…${counts.incomplete}`] : []), + ...(counts.unverifiable > 0 ? [`?${counts.unverifiable}`] : []), + ...(unchecked > 0 ? [`·${unchecked}`] : []), + ]; + return parts.join(" "); +} + /** The file name a manifest URI ends in, for a list that must fit 40 columns. */ function fileNameOf(uri: string): string { const cut = uri.lastIndexOf("/"); @@ -134,6 +233,13 @@ function fileNameOf(uri: string): string { * so it is one `resources/read` per manifest entry and must be asked for. That * split is the same one the web screen makes and the same one SEP-2640 makes: * hosts MUST NOT retrieve a skill's files ahead of need. + * + * **`v` verifies the whole listing** (#2590), in one `verifySkills` run — the + * same call the CLI's `--verify` makes. That is what gives the per-server + * catalog budget (`skillCatalogMaxSkills` / `skillCatalogMaxBytes`) something + * to bound here: it is a RUN-level limit, and a run of one skill always reads + * that skill in full. Skills past the budget come back `incomplete`, unread, + * and Enter on one of them verifies it on its own. */ export function SkillsTab({ skills, @@ -145,17 +251,24 @@ export function SkillsTab({ focusedPane = null, onAuthRecoveryRequired, modalOpen = false, + onFilterEditingChange, }: SkillsTabProps) { - const visibleCount = Math.max(1, height - 7); + // One more row than the bare list: the catalog summary line (#2590). + const visibleCount = Math.max(1, height - 8 - LIST_FILTER_ROWS); + const filter = useListFilter(skills, skillFilterFields, { + enabled: !modalOpen && focusedPane === "list", + onEditingChange: onFilterEditingChange, + }); + const shownSkills = filter.items; const { selectedIndex, firstVisible, setSelection } = useSelectableList( - skills.length, + shownSkills.length, visibleCount, - { resetWhen: skills }, + { resetWhen: shownSkills }, ); const [error, setError] = useState<string | null>(null); const [verifying, setVerifying] = useState(false); /** - * The last verification, keyed by the **entry it was computed against**. + * Every verification held, keyed by the **entry it was computed against**. * * Keyed rather than cleared on selection change, so moving off a skill and * back does not silently discard a verdict the user just paid a round trip @@ -169,17 +282,31 @@ export function SkillsTab({ * The same key the web screen uses, for the same reason: re-verifying after a * metadata-only refresh is the cheap direction to be wrong in; showing a * verdict computed against a different entry is not. + * + * A map rather than a single slot since verify-all (#2590): one run yields a + * verdict for every entry, and each must survive moving the selection. */ - const [report, setReport] = useState<{ - key: string; - result: SkillVerifyReport; - } | null>(null); + const [reports, setReports] = useState< + ReadonlyMap<string, SkillVerifyReport> + >(() => new Map()); const scrollViewRef = useRef<ScrollViewRef>(null); + // The listing as of the latest commit, for a verification that resolves + // after it changed. A LAYOUT effect, so the ref is synchronized during the + // commit itself: a passive effect can run later, and a run resolving in that + // gap would still see a removed entry as live (Copilot). + const skillsRef = useRef(skills); + useLayoutEffect(() => { + skillsRef.current = skills; + }, [skills]); - const selectedSkill = skills[selectedIndex] ?? null; + const selectedSkill = shownSkills[selectedIndex] ?? null; + /** + * Verify `entries` in ONE `verifySkills` run, so the run-level catalog budget + * applies across them — Enter passes the selected skill, `v` the listing. + */ const runVerify = useCallback( - (skill: SkillEntry) => { + (entries: readonly SkillEntry[]) => { if (!inspectorClient || verifying) return; setVerifying(true); setError(null); @@ -187,8 +314,13 @@ export function SkillsTab({ // this key handler — which cannot await — to own. void (async () => { try { - const [result] = await verifySkills(inspectorClient, [skill]); - setReport({ key: skillEntryKey(skill), result }); + const results = await verifySkills(inspectorClient, entries); + // Read AFTER the run, so a listing that changed while it was in + // flight is the one pruned against. + const live = new Set(skillsRef.current.map(skillEntryKey)); + setReports((previous) => + mergeReports(previous, entries, results, live), + ); } catch (err) { if (err instanceof AuthRecoveryRequiredError) { onAuthRecoveryRequired?.(err); @@ -202,7 +334,7 @@ export function SkillsTab({ visible message. Exercising it would mean faking a throw the walk cannot make, which tests the fake rather than the code. */ setError( - err instanceof Error ? err.message : "Failed to verify skill", + err instanceof Error ? errorMessage(err) : "Failed to verify skill", ); /* v8 ignore stop */ } finally { @@ -215,14 +347,23 @@ export function SkillsTab({ useInput( (input: string, key: Key) => { + // The filter goes first: while it is editing, Enter and letters belong + // to the query, not to the list. + if (filter.handleInput(input, key)) return; if (key.return && selectedSkill && inspectorClient) { - runVerify(selectedSkill); + runVerify([selectedSkill]); + return; + } + // The WHOLE listing, not the filtered rows: the CLI's `--verify` covers + // the catalog, and the budget is sized against the catalog (#2590). + if (input === "v" && skills.length > 0 && inspectorClient) { + runVerify(skills); return; } if (focusedPane === "list") { if (key.upArrow && selectedIndex > 0) { setSelection(selectedIndex - 1); - } else if (key.downArrow && selectedIndex < skills.length - 1) { + } else if (key.downArrow && selectedIndex < shownSkills.length - 1) { setSelection(selectedIndex + 1); } return; @@ -267,10 +408,9 @@ export function SkillsTab({ return [...checkSkillConformance(skill), ...(collision ? [collision] : [])]; }; const issues = selectedSkill ? findingsFor(selectedSkill) : []; - const activeReport = - selectedSkill && report?.key === skillEntryKey(selectedSkill) - ? report.result - : null; + const activeReport = selectedSkill + ? (reports.get(skillEntryKey(selectedSkill)) ?? null) + : null; const manifest = selectedSkill && selectedSkill.resources !== DYNAMIC_RESOURCES ? selectedSkill.resources @@ -302,18 +442,37 @@ export function SkillsTab({ bold backgroundColor={focusedPane === "list" ? "yellow" : undefined} > - Skills ({skills.length} + Skills ( + {filterCount(filter.active, shownSkills.length, skills.length)} {pageCount > 1 ? `, ${pageCount} pages` : ""}) </Text> </Box> + <ListFilterBar + query={filter.query} + editing={filter.editing} + focused={focusedPane === "list" && !modalOpen} + /> + <Box flexShrink={0} height={1}> + <Text dimColor wrap="truncate"> + {skills.length === 0 + ? "" + : verifying + ? "[Verifying…]" + : catalogSummary(skills, reports)} + </Text> + </Box> {loadError ? ( <Box paddingY={1}> - <Text color="red">{loadError.message}</Text> + <Text color="red">{errorMessage(loadError)}</Text> </Box> ) : skills.length === 0 ? ( <Box paddingY={1}> <Text dimColor>No skills available</Text> </Box> + ) : shownSkills.length === 0 ? ( + <Box paddingY={1}> + <Text dimColor>No skills match the filter</Text> + </Box> ) : ( <Box flexDirection="column" @@ -321,10 +480,13 @@ export function SkillsTab({ overflow="hidden" flexShrink={0} > - {skills + {shownSkills .slice(firstVisible, firstVisible + visibleCount) .map((skill, i) => { const index = firstVisible + i; + // Keyed by the UNFILTERED position, which stays unique even + // when a malformed listing repeats a URI (see below). + const ordinal = filter.indices[index]; const isSelected = index === selectedIndex; // The per-row mark is the static conformance verdict, which // costs nothing — it is what makes a bad skill visible in the @@ -342,7 +504,7 @@ export function SkillsTab({ // collide them and let React drop or reuse the wrong row // (Copilot). <Box - key={`${index}:${skill.uri}`} + key={`${ordinal}:${skill.uri}`} paddingY={0} flexShrink={0} > diff --git a/clients/tui/src/components/SubscriptionsTab.tsx b/clients/tui/src/components/SubscriptionsTab.tsx new file mode 100644 index 0000000000..6742df5592 --- /dev/null +++ b/clients/tui/src/components/SubscriptionsTab.tsx @@ -0,0 +1,340 @@ +/** + * Subscriptions tab (#2432): subscribe to and unsubscribe from resources, and + * watch the `notifications/resources/updated` feed they produce. + * + * Everything stateful is core's: the subscribed set and each one's + * `lastUpdated` come from `ResourceSubscriptionsState` (through + * `useResourceSubscriptions` in `App`), the toggle is one + * `subscribeToResource` / `unsubscribeFromResource` call — which picks the + * legacy `resources/subscribe` or the modern `subscriptions/listen` filter by + * era — and the feed is read from the protocol message log. The tab owns only + * its selection and the last error. + * + * Its own keys avoid every tab accelerator and the global `c`/`d`, since App's + * handler sees each keypress too: Enter toggles, arrows navigate or scroll. + */ +import React, { useState, useEffect, useRef, useMemo } from "react"; +import { Box, Text, useInput, type Key } from "ink"; +import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; +import type { Resource } from "@modelcontextprotocol/client"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import type { + InspectorResourceSubscription, + MessageEntry, + ResourceSubscriptionStreamState, +} from "@inspector/core/mcp/types.js"; +import { + AuthRecoveryRequiredError, + findNestedAuthError, +} from "@inspector/core/auth/challenge.js"; +import { useSelectableList } from "../hooks/useSelectableList.js"; +import { errorMessage } from "../utils/errorText.js"; +import { + resourceUpdateFeed, + subscribableResources, +} from "../utils/resourceUpdates.js"; + +/** How each modern-era listen-stream status reads in the details pane. */ +const STREAM_STATUS_TEXT: Record< + ResourceSubscriptionStreamState["status"], + { text: string; color: string } +> = { + connecting: { text: "connecting…", color: "yellow" }, + acknowledged: { text: "acknowledged", color: "green" }, + reconnecting: { text: "reconnecting…", color: "yellow" }, + ended: { text: "ended", color: "gray" }, + "never-acknowledged": { + text: "closed without acknowledging (server answered listen with a result)", + color: "red", + }, +}; + +interface SubscriptionsTabProps { + resources: Resource[]; + subscriptions: InspectorResourceSubscription[]; + streamState: ResourceSubscriptionStreamState; + messages: MessageEntry[]; + inspectorClient: InspectorClient | null; + width: number; + height: number; + focusedPane?: "list" | "details" | null; + modalOpen?: boolean; + onAuthRecoveryRequired?: (error: AuthRecoveryRequiredError) => void; +} + +export function SubscriptionsTab({ + resources, + subscriptions, + streamState, + messages, + inspectorClient, + width, + height, + focusedPane = null, + modalOpen = false, + onAuthRecoveryRequired, +}: SubscriptionsTabProps) { + const rows = useMemo( + () => subscribableResources(resources, subscriptions), + [resources, subscriptions], + ); + const feed = useMemo(() => resourceUpdateFeed(messages), [messages]); + const visibleCount = Math.max(1, height - 7); + const { selectedIndex, firstVisible, setSelection } = useSelectableList( + rows.length, + visibleCount, + ); + const [error, setError] = useState<string | null>(null); + const [pendingUri, setPendingUri] = useState<string | null>(null); + const scrollViewRef = useRef<ScrollViewRef>(null); + const listWidth = Math.floor(width * 0.4); + const detailWidth = width - listWidth; + const selected = rows[selectedIndex] ?? null; + + // `pendingUri` drives the status line; this ref is the guard, because two + // Enter presses can land before React re-renders and both would start a + // request (two `resources/subscribe` on the legacy era). + const inFlightRef = useRef(false); + + const toggle = async (row: (typeof rows)[number]) => { + if (!inspectorClient || inFlightRef.current) return; + inFlightRef.current = true; + const { uri } = row.resource; + setPendingUri(uri); + setError(null); + try { + if (row.subscribed) { + await inspectorClient.unsubscribeFromResource(uri); + } else { + await inspectorClient.subscribeToResource(uri); + } + } catch (err) { + // `subscribeToResource` / `unsubscribeFromResource` wrap every failure in + // a new Error with the original as `cause`, so the recovery signal has to + // be dug out of the chain rather than matched on the thrown value. + const authErr = findNestedAuthError(err); + if (authErr instanceof AuthRecoveryRequiredError) { + onAuthRecoveryRequired?.(authErr); + return; + } + setError(errorMessage(err)); + } finally { + inFlightRef.current = false; + setPendingUri(null); + } + }; + + useInput( + (_input: string, key: Key) => { + if (key.return && selected) { + // `toggle` owns every rejection (its catch surfaces the message), and + // a key handler cannot await. + void toggle(selected); + return; + } + + if (focusedPane === "list") { + if (key.upArrow && selectedIndex > 0) { + setSelection(selectedIndex - 1); + } else if (key.downArrow && selectedIndex < rows.length - 1) { + setSelection(selectedIndex + 1); + } + return; + } + + // Only "details" remains: the hook is inactive for any other pane. + if (key.upArrow) { + scrollViewRef.current?.scrollBy(-1); + } else if (key.downArrow) { + scrollViewRef.current?.scrollBy(1); + } else if (key.pageUp) { + const viewportHeight = scrollViewRef.current?.getViewportHeight() || 1; + scrollViewRef.current?.scrollBy(-viewportHeight); + } else if (key.pageDown) { + const viewportHeight = scrollViewRef.current?.getViewportHeight() || 1; + scrollViewRef.current?.scrollBy(viewportHeight); + } + }, + { + isActive: + !modalOpen && (focusedPane === "list" || focusedPane === "details"), + }, + ); + + useEffect(() => { + scrollViewRef.current?.scrollTo(0); + }, [selectedIndex]); + + const streamText = STREAM_STATUS_TEXT[streamState.status]; + // On the modern era a server may honor only part of the filter; say so for + // the selected URI once the stream has been acknowledged. + const notHonored = + !!selected && + selected.subscribed && + streamState.active && + streamState.status === "acknowledged" && + !streamState.honoredUris.includes(selected.resource.uri); + + return ( + <Box flexDirection="row" width={width} height={height}> + {/* Resource list with subscription markers */} + <Box + width={listWidth} + height={height} + borderStyle="single" + borderTop={false} + borderBottom={false} + borderLeft={false} + borderRight={true} + flexDirection="column" + paddingX={1} + > + <Box paddingY={1}> + <Text + bold + backgroundColor={focusedPane === "list" ? "yellow" : undefined} + > + Subscriptions ({subscriptions.length}/{rows.length}) + </Text> + </Box> + {rows.length === 0 ? ( + <Box paddingY={1}> + <Text dimColor>No resources to subscribe to</Text> + </Box> + ) : ( + <Box + flexDirection="column" + height={visibleCount} + overflow="hidden" + flexShrink={0} + > + {rows + .slice(firstVisible, firstVisible + visibleCount) + .map((row, i) => { + const index = firstVisible + i; + const isSelected = index === selectedIndex; + return ( + <Box key={row.resource.uri} paddingY={0} flexShrink={0}> + <Text wrap="truncate"> + {isSelected ? "▶ " : " "} + <Text color={row.subscribed ? "green" : "gray"}> + {row.subscribed ? "●" : "○"} + </Text>{" "} + {row.resource.name || row.resource.uri} + </Text> + </Box> + ); + })} + </Box> + )} + </Box> + + {/* Selected resource + live update feed */} + <Box + width={detailWidth} + height={height} + paddingX={1} + flexDirection="column" + overflow="hidden" + > + <Box flexShrink={0} paddingTop={1}> + <Text + bold + wrap="truncate" + backgroundColor={focusedPane === "details" ? "yellow" : undefined} + {...(focusedPane === "details" ? {} : { color: "cyan" })} + > + {selected + ? selected.resource.name || selected.resource.uri + : "Resource updates"} + </Text> + </Box> + + <ScrollView ref={scrollViewRef} height={height - 5}> + {selected && ( + <> + <Box marginTop={1} flexShrink={0}> + <Text dimColor wrap="truncate"> + URI: {selected.resource.uri} + </Text> + </Box> + <Box flexShrink={0}> + <Text> + {pendingUri === selected.resource.uri ? ( + <Text color="yellow">Updating subscription…</Text> + ) : selected.subscribed ? ( + <Text color="green">Subscribed</Text> + ) : ( + <Text dimColor>Not subscribed</Text> + )} + {selected.lastUpdated && ( + <Text dimColor> + {" "} + — last updated {selected.lastUpdated.toLocaleTimeString()} + </Text> + )} + </Text> + </Box> + {notHonored && ( + <Box flexShrink={0}> + <Text color="yellow"> + Server did not honor this URI in its listen filter + </Text> + </Box> + )} + </> + )} + {streamState.active && ( + <Box flexShrink={0}> + <Text> + Listen stream:{" "} + <Text color={streamText.color}>{streamText.text}</Text> + </Text> + </Box> + )} + {error && ( + <Box marginTop={1} flexShrink={0}> + <Text color="red">{error}</Text> + </Box> + )} + <Box marginTop={1} flexShrink={0}> + <Text bold>Updates ({feed.length}):</Text> + </Box> + {feed.length === 0 ? ( + <Box paddingLeft={2} flexShrink={0}> + <Text dimColor>No resources/updated notifications yet</Text> + </Box> + ) : ( + feed.map((event) => ( + <Box key={event.id} paddingLeft={2} flexShrink={0}> + <Text + wrap="truncate" + {...(selected && event.uri === selected.resource.uri + ? { color: "cyan" } + : { dimColor: true })} + > + {event.timestamp.toLocaleTimeString()} {event.uri} + </Text> + </Box> + )) + )} + </ScrollView> + + {selected && ( + <Box + flexShrink={0} + height={1} + justifyContent="center" + backgroundColor="gray" + > + <Text bold color="white" wrap="truncate"> + {selected.subscribed + ? "Enter to unsubscribe, ↑/↓ to navigate" + : "Enter to subscribe, ↑/↓ to navigate"} + </Text> + </Box> + )} + </Box> + </Box> + ); +} diff --git a/clients/tui/src/components/Tabs.tsx b/clients/tui/src/components/Tabs.tsx index 2a23d21689..d8aa74f73f 100644 --- a/clients/tui/src/components/Tabs.tsx +++ b/clients/tui/src/components/Tabs.tsx @@ -29,9 +29,11 @@ interface TabsProps { info?: number; auth?: number; resources?: number; + subscriptions?: number; prompts?: number; skills?: number; tools?: number; + tasks?: number; messages?: number; requests?: number; logging?: number; @@ -47,6 +49,10 @@ interface TabsProps { * known after connecting. */ showSkills?: boolean; + /** Server advertised `resources.subscribe` (#2432). */ + showSubscriptions?: boolean; + /** Server supports Tasks (legacy capability or SEP-2663 extension) (#2432). */ + showTasks?: boolean; } export function Tabs({ @@ -58,6 +64,8 @@ export function Tabs({ showLogging = true, showRequests = false, showSkills = false, + showSubscriptions = false, + showTasks = false, }: TabsProps) { // Shared with `App`, which sizes the pane below this bar from the same list — // see `tabBarRows`. @@ -66,6 +74,8 @@ export function Tabs({ showLogging, showRequests, showSkills, + showSubscriptions, + showTasks, }); return ( diff --git a/clients/tui/src/components/TasksTab.tsx b/clients/tui/src/components/TasksTab.tsx new file mode 100644 index 0000000000..6dbb94634b --- /dev/null +++ b/clients/tui/src/components/TasksTab.tsx @@ -0,0 +1,357 @@ +/** + * Tasks tab (#2432): the TUI's view of requestor tasks — the tasks this client + * created on the server, whether through a task-augmented call or a server that + * answered an ordinary call with a task handle. + * + * Deliberately thin: the list and its live updates come from core's + * `ManagedRequestorTasksState` (through `useManagedRequestorTasks` in `App`), + * and every action here is one `InspectorClient` call. Nothing about task + * lifecycle is decided in this file — a cancelled task stays cancelled because + * the store pins it, not because this view remembers. + * + * Keys are chosen to avoid every tab accelerator and the global `c`/`d`, since + * App's handler sees each keypress too: `x` cancel, `f` refresh, `l` clear + * completed, Enter fetch the result, `+` zoom. + */ +import React, { useState, useEffect, useRef } from "react"; +import { Box, Text, useInput, type Key } from "ink"; +import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; +import type { CallToolResult, Task } from "@modelcontextprotocol/client"; +import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import { AuthRecoveryRequiredError } from "@inspector/core/auth/challenge.js"; +import { useSelectableList } from "../hooks/useSelectableList.js"; +import { errorMessage } from "../utils/errorText.js"; + +/** Glyph and color per task status; unknown statuses fall back to gray. */ +const STATUS_STYLE: Record<string, { glyph: string; color: string }> = { + working: { glyph: "◐", color: "yellow" }, + input_required: { glyph: "?", color: "magenta" }, + completed: { glyph: "●", color: "green" }, + failed: { glyph: "✗", color: "red" }, + cancelled: { glyph: "○", color: "gray" }, +}; + +export function taskStatusStyle(status: string): { + glyph: string; + color: string; +} { + return STATUS_STYLE[status] ?? { glyph: "·", color: "gray" }; +} + +/** A task the server may still be working on — the only kind worth cancelling. */ +export function isTaskActive(status: string): boolean { + return status === "working" || status === "input_required"; +} + +/** A task with a terminal outcome the server can hand back via `tasks/result`. */ +export function hasTaskResult(status: string): boolean { + return status === "completed" || status === "failed"; +} + +interface TasksTabProps { + tasks: Task[]; + inspectorClient: InspectorClient | null; + width: number; + height: number; + focusedPane?: "list" | "details" | null; + modalOpen?: boolean; + /** Re-list (or re-poll) tasks — `useManagedRequestorTasks().refresh`. */ + onRefresh: () => Promise<unknown>; + /** Drop terminal tasks — `useManagedRequestorTasks().clearCompleted`. */ + onClearCompleted: () => void; + onViewDetails?: (task: Task, result: CallToolResult | null) => void; + onAuthRecoveryRequired?: (error: AuthRecoveryRequiredError) => void; +} + +export function TasksTab({ + tasks, + inspectorClient, + width, + height, + focusedPane = null, + modalOpen = false, + onRefresh, + onClearCompleted, + onViewDetails, + onAuthRecoveryRequired, +}: TasksTabProps) { + const visibleCount = Math.max(1, height - 7); + const { selectedIndex, firstVisible, setSelection } = useSelectableList( + tasks.length, + visibleCount, + ); + const [error, setError] = useState<string | null>(null); + const [busy, setBusy] = useState<string | null>(null); + // Result fetched for one task id; a different selection shows none. + const [result, setResult] = useState<{ + taskId: string; + value: CallToolResult; + } | null>(null); + const scrollViewRef = useRef<ScrollViewRef>(null); + const listWidth = Math.floor(width * 0.4); + const detailWidth = width - listWidth; + + const selectedTask = tasks[selectedIndex] ?? null; + const selectedResult = + result && selectedTask && result.taskId === selectedTask.taskId + ? result.value + : null; + + // One operation at a time. A ref rather than `busy`, because two keypresses + // can land before the state update re-renders, and overlapping refreshes + // walk the store's paginated list concurrently and duplicate rows. + const inFlightRef = useRef(false); + + /** Run one task operation, routing auth recovery and errors uniformly. */ + const run = async (label: string, op: () => Promise<void>) => { + if (inFlightRef.current) return; + inFlightRef.current = true; + setBusy(label); + setError(null); + try { + await op(); + } catch (err) { + if (err instanceof AuthRecoveryRequiredError) { + onAuthRecoveryRequired?.(err); + return; + } + setError(errorMessage(err)); + } finally { + inFlightRef.current = false; + setBusy(null); + } + }; + + useInput( + (input: string, key: Key) => { + if (input === "f") { + // `run` owns every rejection (its catch surfaces the message), and a + // key handler cannot await. + void run("Refreshing…", async () => { + await onRefresh(); + }); + return; + } + if (input === "l") { + // Not during a refresh: the store has already emptied its list for the + // page walk, so a clear now is lost and the tasks reappear after it. + if (inFlightRef.current) return; + onClearCompleted(); + return; + } + if (input === "x" && selectedTask && inspectorClient) { + if (!isTaskActive(selectedTask.status)) return; + const { taskId } = selectedTask; + void run("Cancelling…", () => + inspectorClient.cancelRequestorTask(taskId), + ); + return; + } + if (key.return && selectedTask && inspectorClient) { + if (!hasTaskResult(selectedTask.status)) return; + const { taskId } = selectedTask; + void run("Fetching result…", async () => { + const value = await inspectorClient.getRequestorTaskResult(taskId); + setResult({ taskId, value }); + }); + return; + } + + if (focusedPane === "list") { + if (key.upArrow && selectedIndex > 0) { + setSelection(selectedIndex - 1); + } else if (key.downArrow && selectedIndex < tasks.length - 1) { + setSelection(selectedIndex + 1); + } + return; + } + + // Only "details" remains: the hook is inactive for any other pane. + if (input === "+" && selectedTask && onViewDetails) { + onViewDetails(selectedTask, selectedResult); + return; + } + if (key.upArrow) { + scrollViewRef.current?.scrollBy(-1); + } else if (key.downArrow) { + scrollViewRef.current?.scrollBy(1); + } else if (key.pageUp) { + const viewportHeight = scrollViewRef.current?.getViewportHeight() || 1; + scrollViewRef.current?.scrollBy(-viewportHeight); + } else if (key.pageDown) { + const viewportHeight = scrollViewRef.current?.getViewportHeight() || 1; + scrollViewRef.current?.scrollBy(viewportHeight); + } + }, + { + isActive: + !modalOpen && (focusedPane === "list" || focusedPane === "details"), + }, + ); + + // Reset scroll when selection changes + useEffect(() => { + scrollViewRef.current?.scrollTo(0); + }, [selectedIndex]); + + const selectedStyle = selectedTask + ? taskStatusStyle(selectedTask.status) + : null; + + return ( + <Box flexDirection="row" width={width} height={height}> + {/* Task list */} + <Box + width={listWidth} + height={height} + borderStyle="single" + borderTop={false} + borderBottom={false} + borderLeft={false} + borderRight={true} + flexDirection="column" + paddingX={1} + > + <Box paddingY={1}> + <Text + bold + backgroundColor={focusedPane === "list" ? "yellow" : undefined} + > + Tasks ({tasks.length}) + </Text> + </Box> + {tasks.length === 0 ? ( + <Box paddingY={1}> + <Text dimColor>No tasks yet (f to refresh)</Text> + </Box> + ) : ( + <Box + flexDirection="column" + height={visibleCount} + overflow="hidden" + flexShrink={0} + > + {tasks + .slice(firstVisible, firstVisible + visibleCount) + .map((task, i) => { + const index = firstVisible + i; + const isSelected = index === selectedIndex; + const style = taskStatusStyle(task.status); + return ( + <Box key={task.taskId} paddingY={0} flexShrink={0}> + <Text wrap="truncate"> + {isSelected ? "▶ " : " "} + <Text color={style.color}>{style.glyph}</Text>{" "} + {task.taskId} + </Text> + </Box> + ); + })} + </Box> + )} + </Box> + + {/* Task details */} + <Box + width={detailWidth} + height={height} + paddingX={1} + flexDirection="column" + overflow="hidden" + > + {selectedTask && selectedStyle ? ( + <> + <Box flexShrink={0} paddingTop={1}> + <Text + bold + wrap="truncate" + backgroundColor={ + focusedPane === "details" ? "yellow" : undefined + } + {...(focusedPane === "details" ? {} : { color: "cyan" })} + > + {selectedTask.taskId} + </Text> + </Box> + + <ScrollView ref={scrollViewRef} height={height - 5}> + <Box marginTop={1} flexShrink={0}> + <Text> + Status:{" "} + <Text color={selectedStyle.color} bold> + {selectedTask.status} + </Text> + </Text> + </Box> + {selectedTask.statusMessage && ( + <Box flexShrink={0}> + <Text dimColor>{selectedTask.statusMessage}</Text> + </Box> + )} + <Box flexShrink={0}> + <Text dimColor>Created: {selectedTask.createdAt}</Text> + </Box> + {selectedTask.lastUpdatedAt && ( + <Box flexShrink={0}> + <Text dimColor>Updated: {selectedTask.lastUpdatedAt}</Text> + </Box> + )} + <Box flexShrink={0}> + <Text dimColor> + TTL: {selectedTask.ttl === null ? "none" : selectedTask.ttl} + {selectedTask.pollInterval !== undefined && + ` Poll: ${selectedTask.pollInterval}ms`} + </Text> + </Box> + {busy && ( + <Box marginTop={1} flexShrink={0}> + <Text color="yellow">{busy}</Text> + </Box> + )} + {error && ( + <Box marginTop={1} flexShrink={0}> + <Text color="red">{error}</Text> + </Box> + )} + {selectedResult && ( + <Box marginTop={1} flexShrink={0} flexDirection="column"> + <Text bold>Result:</Text> + <Box paddingLeft={2}> + <Text dimColor> + {JSON.stringify(selectedResult, null, 2)} + </Text> + </Box> + </Box> + )} + </ScrollView> + + <Box + flexShrink={0} + height={1} + justifyContent="center" + backgroundColor="gray" + > + <Text bold color="white" wrap="truncate"> + {[ + hasTaskResult(selectedTask.status) && "Enter result", + isTaskActive(selectedTask.status) && "x cancel", + "f refresh", + "l clear done", + focusedPane === "details" && "+ zoom", + ] + .filter(Boolean) + .join(", ")} + </Text> + </Box> + </> + ) : ( + <Box paddingY={1} flexShrink={0} flexDirection="column"> + <Text dimColor>Select a task to view details</Text> + {busy && <Text color="yellow">{busy}</Text>} + {error && <Text color="red">{error}</Text>} + </Box> + )} + </Box> + </Box> + ); +} diff --git a/clients/tui/src/components/ToolTestModal.tsx b/clients/tui/src/components/ToolTestModal.tsx index 4e7f693dff..30f7833e8e 100644 --- a/clients/tui/src/components/ToolTestModal.tsx +++ b/clients/tui/src/components/ToolTestModal.tsx @@ -12,6 +12,17 @@ import { } from "../utils/schemaToForm.js"; import { ScrollView, type ScrollViewRef } from "ink-scroll-view"; import { inlineLocalRefs } from "@inspector/core/json/localRefs.js"; +import { redactErrorText, redactedJson } from "../utils/errorText.js"; +import type { ResultFileFormat } from "@inspector/core/mcp/resultFile.js"; +import { + defaultResultFileName, + saveResultToFile, +} from "../utils/saveResult.js"; +import { + SaveResultBar, + type SavePrompt, + type SaveStatus, +} from "./SaveResultBar.js"; interface ToolTestModalProps { tool: Tool; @@ -30,6 +41,12 @@ interface ToolResult { error?: string; errorDetails?: unknown; duration: number; + /** + * The server's own `CallToolResult`, whether it succeeded or carried + * `isError` — what `w` saves. Absent when no result came back at all (a + * thrown call, a failed invocation, a missing-argument refusal). + */ + callResult?: CallToolResult; } export function ToolTestModal({ @@ -42,7 +59,24 @@ export function ToolTestModal({ }: ToolTestModalProps) { const [state, setState] = useState<ModalState>("form"); const [result, setResult] = useState<ToolResult | null>(null); + const [savePrompt, setSavePromptState] = useState<SavePrompt | null>(null); + // The prompt as of the last keystroke, not the last render. Ink keeps the + // previous render's input handler attached until effects flush, so a key + // typed right after `w` (or a second key in the same burst) would otherwise + // see a stale `savePrompt` — Enter dropped, or two Backspaces deleting one + // character. Every update goes through `setSavePrompt` so the two agree. + const savePromptRef = React.useRef<SavePrompt | null>(null); + const setSavePrompt = (next: SavePrompt | null) => { + savePromptRef.current = next; + setSavePromptState(next); + }; + const [saveStatus, setSaveStatus] = useState<SaveStatus | null>(null); const scrollViewRef = React.useRef<ScrollViewRef>(null); + // Numbers this view's saves so only the latest may report. The queue in + // saveResultToFile keeps the writes themselves ordered; this keeps an older + // save settling from replacing the newer one's "Saving…" with its "Saved". + const saveAttemptRef = React.useRef(0); + const toolName = tool?.name || "tool"; // Use full terminal dimensions instead of passed dimensions const [terminalDimensions, setTerminalDimensions] = React.useState({ @@ -88,41 +122,140 @@ export function ToolTestModal({ // Handle all input when modal is open - prevents input from reaching underlying components // When in form mode, only handle escape (form handles its own input) // When in results mode, handle scrolling keys - useInput( - (input: string, key: Key) => { - // Always handle escape to close modal - if (key.escape) { - setState("form"); - setResult(null); - onClose(); - return; - } + // Ink attaches useInput's handler in an effect, so until effects flush a key + // still reaches the previous render's handler. Routing every key through a + // ref read at call time means it always sees the latest render's state. + const handleInputRef = React.useRef<(input: string, key: Key) => void>( + () => {}, + ); + handleInputRef.current = (input: string, key: Key) => { + // While the save prompt is open it owns every key, Escape included — + // Escape cancels the prompt rather than closing the whole modal. + const prompt = savePromptRef.current; + if (prompt) { + handleSavePromptInput(prompt, input, key); + return; + } - if (state === "form") { - // In form mode, let the form handle all other input - // Don't process anything else - this prevents input from reaching underlying components + // Always handle escape to close modal + if (key.escape) { + setState("form"); + setResult(null); + onClose(); + return; + } + + if (state === "form") { + // In form mode, let the form handle all other input + // Don't process anything else - this prevents input from reaching underlying components + return; + } + + if (state === "results") { + // Plain w only: Ink reports a chord's letter in `input` too. + if (input === "w" && !key.ctrl && !key.meta) { + openSavePrompt(); return; } - - if (state === "results") { - // Allow scrolling in results view - if (key.downArrow) { - scrollViewRef.current?.scrollBy(1); - } else if (key.upArrow) { - scrollViewRef.current?.scrollBy(-1); - } else if (key.pageDown) { - const viewportHeight = - scrollViewRef.current?.getViewportHeight() || 1; - scrollViewRef.current?.scrollBy(viewportHeight); - } else if (key.pageUp) { - const viewportHeight = - scrollViewRef.current?.getViewportHeight() || 1; - scrollViewRef.current?.scrollBy(-viewportHeight); - } + // Allow scrolling in results view + if (key.downArrow) { + scrollViewRef.current?.scrollBy(1); + } else if (key.upArrow) { + scrollViewRef.current?.scrollBy(-1); + } else if (key.pageDown) { + const viewportHeight = scrollViewRef.current?.getViewportHeight() || 1; + scrollViewRef.current?.scrollBy(viewportHeight); + } else if (key.pageUp) { + const viewportHeight = scrollViewRef.current?.getViewportHeight() || 1; + scrollViewRef.current?.scrollBy(-viewportHeight); } - }, - { isActive: true }, - ); + } + }; + useInput((input: string, key: Key) => handleInputRef.current(input, key), { + isActive: true, + }); + + const openSavePrompt = () => { + if (!result?.callResult) { + setSaveStatus({ + ok: false, + message: + "No tool result to save — no result came back from the server.", + }); + return; + } + setSaveStatus(null); + setSavePrompt({ + path: defaultResultFileName(toolName, "json", result.callResult), + format: "json", + edited: false, + }); + }; + + const handleSavePromptInput = ( + prompt: SavePrompt, + input: string, + key: Key, + ) => { + if (key.escape) { + setSavePrompt(null); + return; + } + if (key.return) { + setSavePrompt(null); + // A key handler cannot await; saveCurrentResult owns its failures and + // surfaces them on the status line. + void saveCurrentResult(prompt); + return; + } + if (key.tab) { + const format: ResultFileFormat = + prompt.format === "json" ? "raw" : "json"; + setSavePrompt({ + ...prompt, + format, + path: prompt.edited + ? prompt.path + : defaultResultFileName(toolName, format, result?.callResult), + }); + return; + } + if (key.backspace || key.delete) { + setSavePrompt({ + ...prompt, + path: prompt.path.slice(0, -1), + edited: true, + }); + return; + } + if (input && !key.ctrl && !key.meta) { + setSavePrompt({ ...prompt, path: prompt.path + input, edited: true }); + } + }; + + const saveCurrentResult = async (prompt: SavePrompt) => { + const attempt = ++saveAttemptRef.current; + const report = (status: SaveStatus) => { + if (attempt === saveAttemptRef.current) setSaveStatus(status); + }; + setSaveStatus({ ok: true, message: `Saving to ${prompt.path}…` }); + try { + const saved = await saveResultToFile( + result?.callResult, + prompt.path, + prompt.format, + ); + report({ + ok: true, + message: `Saved ${saved.format} result to ${saved.path} (${saved.bytes} ${saved.bytes === 1 ? "byte" : "bytes"})`, + }); + } catch (err) { + report({ + ok: false, + message: err instanceof Error ? err.message : String(err), + }); + } + }; const handleFormSubmit = async (rawValues: Record<string, JsonValue>) => { if (!inspectorClient || !tool) return; @@ -182,6 +315,7 @@ export function ToolTestModal({ error: isError ? "Tool returned an error" : undefined, errorDetails: isError ? result : undefined, duration, + callResult: result, }); } setState("results"); @@ -238,7 +372,11 @@ export function ToolTestModal({ {formStructure.title} </Text> <Text> </Text> - <Text dimColor>(Press ESC to close)</Text> + <Text dimColor> + {state === "results" && result?.callResult + ? "(w to save result, ESC to close)" + : "(Press ESC to close)"} + </Text> </Box> {/* Content Area */} @@ -289,7 +427,9 @@ export function ToolTestModal({ Error: </Text> <Box paddingLeft={2}> - <Text color="red">{String(result.error)}</Text> + <Text color="red"> + {redactErrorText(String(result.error))} + </Text> </Box> {result.errorDetails != null ? ( <> @@ -300,7 +440,7 @@ export function ToolTestModal({ </Box> <Box paddingLeft={2}> <Text dimColor> - {JSON.stringify(result.errorDetails, null, 2)} + {redactedJson(result.errorDetails)} </Text> </Box> </> @@ -321,6 +461,8 @@ export function ToolTestModal({ </ScrollView> </Box> )} + + <SaveResultBar prompt={savePrompt} status={saveStatus} /> </Box> </Box> </Box> diff --git a/clients/tui/src/components/ToolsTab.tsx b/clients/tui/src/components/ToolsTab.tsx index 0ac47a5bcb..5bab4c5f23 100644 --- a/clients/tui/src/components/ToolsTab.tsx +++ b/clients/tui/src/components/ToolsTab.tsx @@ -8,6 +8,12 @@ import { type SchemaFinding, } from "@inspector/core/json/schemaLint.js"; import { useSelectableList } from "../hooks/useSelectableList.js"; +import { useListFilter } from "../hooks/useListFilter.js"; +import { + ListFilterBar, + LIST_FILTER_ROWS, + filterCount, +} from "./ListFilterBar.js"; /** * How each severity renders in the terminal (#1005). One table, read by both @@ -48,6 +54,13 @@ interface ToolsTabProps { onTestTool?: (tool: Tool) => void; onViewDetails?: (tool: Tool) => void; modalOpen?: boolean; + /** Told when the list filter starts or stops capturing keys (#2430). */ + onFilterEditingChange?: (editing: boolean) => void; +} + +/** What the `/` filter matches a tool against (#2430). */ +function toolFilterFields(tool: Tool): ReadonlyArray<string | undefined> { + return [tool.name, tool.title, tool.annotations?.title]; } export function ToolsTab({ @@ -59,19 +72,26 @@ export function ToolsTab({ onTestTool, onViewDetails, modalOpen = false, + onFilterEditingChange, }: ToolsTabProps) { - const visibleCount = Math.max(1, height - 7); + const visibleCount = Math.max(1, height - 7 - LIST_FILTER_ROWS); + const filter = useListFilter(tools, toolFilterFields, { + enabled: !modalOpen && focusedPane === "list", + onEditingChange: onFilterEditingChange, + }); + const shownTools = filter.items; const { selectedIndex, firstVisible, setSelection } = useSelectableList( - tools.length, + shownTools.length, visibleCount, - { resetWhen: tools }, + { resetWhen: shownTools }, ); const [error] = useState<string | null>(null); // Tool-schema portability findings, one entry per tool, same order (#1005). // Pure walk over data already in memory, so it is recomputed only when the // list itself changes rather than on every keypress. + // Keyed by the tool object so a filtered view still finds each row's entry. const findingsByTool = useMemo( - () => tools.map((tool) => lintToolSchemas(tool)), + () => new Map(tools.map((tool) => [tool, lintToolSchemas(tool)])), [tools], ); const scrollViewRef = useRef<ScrollViewRef>(null); @@ -81,6 +101,10 @@ export function ToolsTab({ // Handle arrow key navigation when focused useInput( (input: string, key: Key) => { + // The filter goes first: while it is editing, Enter and letters belong + // to the query, not to the list. + if (filter.handleInput(input, key)) return; + // Handle Enter key to test tool (works from both list and details) if (key.return && selectedTool && isConnected && onTestTool) { onTestTool(selectedTool); @@ -90,7 +114,7 @@ export function ToolsTab({ if (focusedPane === "list") { if (key.upArrow && selectedIndex > 0) { setSelection(selectedIndex - 1); - } else if (key.downArrow && selectedIndex < tools.length - 1) { + } else if (key.downArrow && selectedIndex < shownTools.length - 1) { setSelection(selectedIndex + 1); } return; @@ -130,8 +154,9 @@ export function ToolsTab({ scrollViewRef.current?.scrollTo(0); }, [selectedIndex]); - const selectedTool = tools[selectedIndex] || null; - const selectedFindings = findingsByTool[selectedIndex] ?? []; + const selectedTool = shownTools[selectedIndex] || null; + const selectedFindings = + (selectedTool && findingsByTool.get(selectedTool)) || []; return ( <Box flexDirection="row" width={width} height={height}> @@ -152,9 +177,15 @@ export function ToolsTab({ bold backgroundColor={focusedPane === "list" ? "yellow" : undefined} > - Tools ({tools.length}) + Tools ({filterCount(filter.active, shownTools.length, tools.length)} + ) </Text> </Box> + <ListFilterBar + query={filter.query} + editing={filter.editing} + focused={focusedPane === "list" && !modalOpen} + /> {error ? ( <Box paddingY={1}> <Text color="red">{error}</Text> @@ -163,6 +194,10 @@ export function ToolsTab({ <Box paddingY={1}> <Text dimColor>No tools available</Text> </Box> + ) : shownTools.length === 0 ? ( + <Box paddingY={1}> + <Text dimColor>No tools match the filter</Text> + </Box> ) : ( <Box flexDirection="column" @@ -170,17 +205,24 @@ export function ToolsTab({ overflow="hidden" flexShrink={0} > - {tools + {shownTools .slice(firstVisible, firstVisible + visibleCount) .map((tool, i) => { const index = firstVisible + i; const isSelected = index === selectedIndex; - const marker = schemaMarker(findingsByTool[index]); + const marker = schemaMarker(findingsByTool.get(tool)); + // The unfiltered position keeps the key unique when a server + // repeats a tool name and the filter hides one copy (Copilot). + const ordinal = filter.indices[index]; return ( - <Box key={tool.name || index} paddingY={0} flexShrink={0}> + <Box + key={`${ordinal}:${tool.name}`} + paddingY={0} + flexShrink={0} + > <Text> {isSelected ? "▶ " : " "} - {tool.name || `Tool ${index + 1}`} + {tool.name || `Tool ${ordinal + 1}`} {marker && ( <Text color={marker.color}> {marker.glyph}</Text> )} diff --git a/clients/tui/src/components/tabsConfig.ts b/clients/tui/src/components/tabsConfig.ts index 7d52bd377a..bfb954aa55 100644 --- a/clients/tui/src/components/tabsConfig.ts +++ b/clients/tui/src/components/tabsConfig.ts @@ -2,9 +2,11 @@ export type TabType = | "info" | "auth" | "resources" + | "subscriptions" | "prompts" | "skills" | "tools" + | "tasks" | "messages" | "requests" | "logging"; @@ -23,13 +25,19 @@ export const tabs: { id: TabType; label: string; accelerator: string }[] = [ { id: "info", label: "Info", accelerator: "i" }, { id: "auth", label: "Auth", accelerator: "a" }, { id: "resources", label: "Resources", accelerator: "r" }, + // `u`, not `s`: `s` is the only free letter in `Ta**s**ks`, so Subscriptions + // takes the earliest remaining letter in its own word instead (#2432). + { id: "subscriptions", label: "Subscriptions", accelerator: "u" }, { id: "prompts", label: "Prompts", accelerator: "m" }, - // `k`, not `s`: `s` is not in conflict today, but the accelerator has to - // appear in the label and be unique, and `S`kills against a future `S`ampling - // or `S`ettings is the collision this rule anticipates. `k` is the earliest - // remaining letter in the word after `s` and `i` (Info). + // `k`, not `s`: `s` is reserved for Tasks (Ta**s**ks has no other free + // letter), and the accelerator has to appear in the label and be unique. + // `k` is the earliest remaining letter in the word after `s` and `i` (Info). { id: "skills", label: "Skills", accelerator: "k" }, { id: "tools", label: "Tools", accelerator: "t" }, + // `s` is the only letter of `Tasks` not already taken (t/a/k). AuthTab binds + // `s` locally to "clear OAuth state", so App suppresses this accelerator + // while the Auth pane holds focus — the same carve-out the step-up `a` gets. + { id: "tasks", label: "Tasks", accelerator: "s" }, { id: "messages", label: "Protocol", accelerator: "p" }, { id: "requests", label: "Network", accelerator: "n" }, { id: "logging", label: "Console", accelerator: "o" }, @@ -41,6 +49,16 @@ export interface TabVisibility { showLogging: boolean; showRequests: boolean; showSkills: boolean; + /** + * Shown only when the connected server advertised `resources.subscribe` — + * there is nothing to subscribe to otherwise (#2432). + */ + showSubscriptions?: boolean; + /** + * Shown only when the connected server supports Tasks: the legacy + * `capabilities.tasks` or the negotiated SEP-2663 extension (#2432). + */ + showTasks?: boolean; } /** @@ -56,6 +74,8 @@ export function visibleTabs(v: TabVisibility): typeof tabs { if (tab.id === "logging") return v.showLogging; if (tab.id === "requests") return v.showRequests; if (tab.id === "skills") return v.showSkills; + if (tab.id === "subscriptions") return v.showSubscriptions ?? false; + if (tab.id === "tasks") return v.showTasks ?? false; return true; }); } diff --git a/clients/tui/src/hooks/useCopyKeys.ts b/clients/tui/src/hooks/useCopyKeys.ts new file mode 100644 index 0000000000..6e5631939d --- /dev/null +++ b/clients/tui/src/hooks/useCopyKeys.ts @@ -0,0 +1,80 @@ +import { useCallback, useState } from "react"; +import { useStdout } from "ink"; +import { copyViaOsc52, writeCopyFile } from "../utils/clipboard.js"; + +/** Copy to the terminal clipboard (OSC 52). Vim's "yank". */ +export const COPY_KEY = "y"; +/** Save to a temp file — the fallback for a terminal without OSC 52. */ +export const SAVE_KEY = "w"; + +export interface CopyStatus { + tone: "success" | "warning" | "error"; + message: string; +} + +/** + * The copy affordance shared by every TUI surface that offers one (#2421): + * `Y` sends the value to the terminal clipboard over OSC 52, and `W` saves it + * to a private temp file for terminals that ignore OSC 52. + * + * The caller owns its `useInput` handler (focus and modal gating differ per + * surface) and forwards keys to `handleCopyKey`, which reports whether it + * consumed the key. `value` is read at keypress time, so it may be undefined + * when there is nothing to copy — the keys are then not consumed. + */ +export function useCopyKeys(value: string | undefined, label: string) { + const { stdout } = useStdout(); + const [status, setStatus] = useState<CopyStatus | null>(null); + // A status describes the value it was produced for; once the value changes + // (another server's token, another entry) it would describe the wrong one. + // Cleared during render — React's "adjusting state on a prop change" + // pattern — so the stale line is never painted. + const [statusFor, setStatusFor] = useState(value); + if (statusFor !== value) { + setStatusFor(value); + setStatus(null); + } + + const handleCopyKey = useCallback( + (input: string): boolean => { + if (value === undefined) return false; + const key = input.toLowerCase(); + if (key === COPY_KEY) { + const result = copyViaOsc52(value, stdout); + setStatus( + result.large + ? { + tone: "warning", + message: `Sent ${label} (${result.chars} chars) to the clipboard via OSC 52. Large payloads may be truncated or dropped by the terminal — press W to save it to a file instead.`, + } + : { + tone: "success", + message: `Copied ${label} (${result.chars} chars) via OSC 52. Nothing pasted? Your terminal may not support OSC 52 — press W to save it to a file.`, + }, + ); + return true; + } + if (key === SAVE_KEY) { + try { + const file = writeCopyFile(value); + setStatus({ tone: "success", message: `Saved ${label} to ${file}` }); + } catch (err) { + setStatus({ + tone: "error", + message: `Could not save ${label}: ${err instanceof Error ? err.message : String(err)}`, + }); + } + return true; + } + return false; + }, + [value, label, stdout], + ); + + return { handleCopyKey, status }; +} + +/** Ink colour for a {@link CopyStatus} tone. */ +export function copyStatusColor(tone: CopyStatus["tone"]): string { + return tone === "success" ? "green" : tone === "warning" ? "yellow" : "red"; +} diff --git a/clients/tui/src/hooks/useListFilter.ts b/clients/tui/src/hooks/useListFilter.ts new file mode 100644 index 0000000000..9cbf7ebf82 --- /dev/null +++ b/clients/tui/src/hooks/useListFilter.ts @@ -0,0 +1,166 @@ +/** + * A type-to-filter query over one of the TUI's list panes (#2430). + * + * The web client's Tools / Resources / Prompts screens carry a search box; the + * TUI's equivalents had nothing, so a server exposing a few hundred tools could + * only be read by scrolling. This hook is the one place that behaviour lives, + * so the four list tabs (`ToolsTab`, `ResourcesTab`, `PromptsTab`, + * `SkillsTab`) cannot disagree about what `/` does or what a match is. + * + * **Keys** — the `/` convention from `less`, `vim` and most TUIs: + * + * - `/` on a focused list opens the filter for editing; + * - printable characters extend the query and backspace trims it, narrowing + * the list as you type; + * - **Enter keeps** the query and leaves editing, so the arrows, Enter and the + * app-wide accelerators work on the narrowed list; + * - **Esc clears** the query and leaves editing. A kept filter is cleared the + * same way — `/` then Esc — because a bare Esc outside editing is the app's + * exit key and stays that. + * + * Up/Down are deliberately **not** consumed while editing, so the selection can + * be moved without leaving the query. + * + * ⚠️ **While editing, the app-wide accelerators must stand down**, or typing + * `connect` into the filter would connect (`c`), switch to the Prompts tab + * (`p`)… and Esc would quit. Ink delivers every key to every active `useInput`, + * so the owning tab cannot swallow a key on App's behalf; instead the hook + * reports its editing state through `onEditingChange` and App skips its own + * handler while it is `true`. That report is an effect with a cleanup, so it is + * withdrawn on every way editing can end — the keys above, the pane losing + * focus or a modal opening (`enabled` going false), or the tab unmounting — + * and App can never be left deaf to the keyboard. + * + * A match is a case-insensitive substring of any of the strings `fields` + * returns for an item. `fields` is called during render, so it must be pure; + * declare it at module scope so the memoized result is not rebuilt every + * render. + */ +import { useCallback, useEffect, useMemo, useState } from "react"; +import type { Key } from "ink"; + +/** The strings an item is matched against; `undefined` entries are skipped. */ +export type ListFilterFields<T> = ( + item: T, +) => ReadonlyArray<string | undefined>; + +export interface UseListFilterOptions { + /** + * Whether the owning list currently has the keyboard — the pane is focused + * and no modal is open. `/` is ignored while it is false, and editing that + * was in progress is suspended (and reported as not editing) until it is + * true again. + */ + enabled: boolean; + /** + * Told whenever the filter starts or stops capturing keys, so App can mute + * its global accelerators. Must be referentially stable (a `useState` + * setter is), since it is an effect dependency. + */ + onEditingChange?: (editing: boolean) => void; +} + +export interface ListFilter<T> { + /** The current query, as typed. */ + query: string; + /** Whether keystrokes are currently going into the query. */ + editing: boolean; + /** Whether the list is narrowed — a non-empty query. */ + active: boolean; + /** The matching items, in their original order. */ + items: readonly T[]; + /** `indices[i]` is the position of `items[i]` in the unfiltered list. */ + indices: readonly number[]; + /** + * Feed a keystroke from the owning tab's `useInput`. Returns `true` when the + * filter consumed it, in which case the tab must not act on it. + */ + handleInput: (input: string, key: Key) => boolean; +} + +/** + * Strip control characters (and the ESC sequences that carry them) from typed + * or pasted text, so a stray escape code can never land in the query. + */ +function printable(input: string): string { + // eslint-disable-next-line no-control-regex -- matching control characters is the point + return input.replace(/[\u0000-\u001f\u007f]/g, ""); +} + +export function useListFilter<T>( + items: readonly T[], + fields: ListFilterFields<T>, + { enabled, onEditingChange }: UseListFilterOptions, +): ListFilter<T> { + const [query, setQuery] = useState(""); + const [editingRequested, setEditingRequested] = useState(false); + // Editing is suspended, not cancelled, while the list lacks the keyboard: + // derived rather than reset in an effect, so nothing renders a stale frame. + const editing = editingRequested && enabled; + + useEffect(() => { + if (!editing) return; + onEditingChange?.(true); + return () => onEditingChange?.(false); + }, [editing, onEditingChange]); + + const { filtered, indices } = useMemo(() => { + const needle = query.trim().toLowerCase(); + const filtered: T[] = []; + const indices: number[] = []; + items.forEach((item, index) => { + if ( + needle === "" || + fields(item).some((field) => field?.toLowerCase().includes(needle)) + ) { + filtered.push(item); + indices.push(index); + } + }); + return { filtered, indices }; + }, [items, fields, query]); + + const handleInput = useCallback( + (input: string, key: Key): boolean => { + if (!enabled) return false; + if (!editingRequested) { + if (input === "/" && !key.ctrl && !key.meta) { + setEditingRequested(true); + return true; + } + return false; + } + if (key.escape) { + setQuery(""); + setEditingRequested(false); + return true; + } + if (key.return) { + setEditingRequested(false); + return true; + } + if (key.backspace || key.delete) { + setQuery((prev) => prev.slice(0, -1)); + return true; + } + // Navigation stays with the list, so the selection can move mid-query. + if (key.upArrow || key.downArrow || key.pageUp || key.pageDown) { + return false; + } + if (key.ctrl || key.meta) return true; + const text = printable(input); + if (text) setQuery((prev) => prev + text); + return true; + }, + [enabled, editingRequested], + ); + + return { + query, + editing, + active: query.trim() !== "", + items: filtered, + indices, + handleInput, + }; +} diff --git a/clients/tui/src/utils/clipboard.ts b/clients/tui/src/utils/clipboard.ts new file mode 100644 index 0000000000..3c93335d06 --- /dev/null +++ b/clients/tui/src/utils/clipboard.ts @@ -0,0 +1,99 @@ +// Getting a value out of the TUI and onto the user's clipboard (#2421). +// +// The web client copies through the browser Clipboard API (`CopyButton`). A +// terminal app has no such API, and the obvious substitutes — shelling out to +// `pbcopy` / `xclip` / `clip.exe` — write to the clipboard of the machine the +// *process* runs on, which over SSH is the remote host rather than the machine +// the user is sitting at. OSC 52 is the escape sequence that asks the +// *terminal emulator* to set its clipboard, so it reaches the local clipboard +// through any number of SSH hops; that is why it is the mechanism here and why +// no clipboard package is needed. +// +// It is fire-and-forget: the terminal sends no acknowledgement, so a terminal +// that ignores OSC 52 (or caps its payload size) fails silently. The fallback +// for that case is `writeCopyFile`, which saves the raw value to a private temp +// file whose path the TUI then shows — readable with `cat` or `scp` from +// anywhere, and not bounded by the terminal's line width or scrollback. +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; + +/** + * Encoded payload size above which a copy is flagged as "large". Several + * terminals cap or drop OSC 52 payloads around this size (hterm's default is + * 100,000 bytes; others are lower), so the status line suggests the file + * fallback rather than letting a truncated paste go unexplained. + */ +export const OSC52_LARGE_PAYLOAD_BYTES = 100_000; + +/** The text to copy for a value: strings verbatim, anything else as JSON. */ +export function toCopyText(value: unknown): string { + if (typeof value === "string") return value; + try { + // JSON.stringify returns undefined for undefined / functions / symbols. + return JSON.stringify(value, null, 2) ?? String(value); + } catch { + // A cycle or a BigInt: still give the user something rather than nothing. + return String(value); + } +} + +/** + * The OSC 52 "set clipboard" sequence for `text`, terminated with BEL (the + * terminator the widest set of terminals accepts). `c` selects the clipboard + * rather than the primary selection. + */ +export function buildOsc52Sequence(text: string): string { + const encoded = Buffer.from(text, "utf8").toString("base64"); + return `\u001b]52;c;${encoded}\u0007`; +} + +export interface Osc52CopyResult { + /** Characters in the copied text. */ + chars: number; + /** Bytes of base64 payload sent to the terminal. */ + payloadBytes: number; + /** True when the payload exceeds {@link OSC52_LARGE_PAYLOAD_BYTES}. */ + large: boolean; +} + +/** The slice of a writable stream `copyViaOsc52` needs. */ +export interface EscapeSink { + write(chunk: string): unknown; +} + +/** + * Emit `text` to the terminal as an OSC 52 clipboard write. + * + * Written straight to the stream rather than through Ink's `write`, which + * would erase and redraw the frame around it: the sequence has no visible + * output and moves no cursor, so it can go out between frames untouched. + */ +export function copyViaOsc52(text: string, sink: EscapeSink): Osc52CopyResult { + const sequence = buildOsc52Sequence(text); + sink.write(sequence); + // Everything but the payload is fixed framing (ESC ] 5 2 ; c ; … BEL), and + // base64 is ASCII, so string length is byte length. + const payloadBytes = sequence.length - buildOsc52Sequence("").length; + return { + chars: text.length, + payloadBytes, + large: payloadBytes > OSC52_LARGE_PAYLOAD_BYTES, + }; +} + +/** + * Save `text` to a fresh, owner-only temp file and return its path — the + * fallback for a terminal that ignores OSC 52. The directory is created with + * `mkdtemp` (mode 0700) and the file with mode 0600, because the value may be + * a bearer token. + */ +export function writeCopyFile( + text: string, + baseDir: string = os.tmpdir(), +): string { + const dir = fs.mkdtempSync(path.join(baseDir, "mcp-inspector-copy-")); + const file = path.join(dir, "value.txt"); + fs.writeFileSync(file, text, { encoding: "utf8", mode: 0o600 }); + return file; +} diff --git a/clients/tui/src/utils/errorText.ts b/clients/tui/src/utils/errorText.ts new file mode 100644 index 0000000000..8e06f38683 --- /dev/null +++ b/clients/tui/src/utils/errorText.ts @@ -0,0 +1,36 @@ +import { redactUrlsInText } from "@inspector/core/mcp/fetchTracking.js"; + +/** + * The TUI's display boundary for error text (#2490). + * + * A server or SDK error whose text quotes `https://…?code=…` would otherwise be + * drawn on screen verbatim — a screenshot or screen-share away from leaking an + * OAuth code, access token or client secret. Everything the TUI shows from a + * caught error goes through one of these, which apply core's + * {@link redactUrlsInText} (the same redaction the CLI's stderr envelope and + * the web client's toasts use): sensitive query values are replaced, while the + * path and non-sensitive parameters stay readable. + * + * Redaction is applied only to what is displayed, never to what is classified + * — callers keep reading the original error for `instanceof` checks. + */ + +/** The display text of a caught value: an Error's message, else `String(err)`. */ +export function errorMessage(err: unknown): string { + return redactUrlsInText(err instanceof Error ? err.message : String(err)); +} + +/** Redact URL query secrets in free text that is about to be displayed. */ +export function redactErrorText(text: string): string { + return redactUrlsInText(text); +} + +/** + * Pretty-printed JSON of an error-details value, with URL query secrets + * redacted. The URL pattern stops at a double quote, so a URL inside a + * serialized string value is redacted without disturbing the JSON around it. + */ +export function redactedJson(value: unknown): string { + // `JSON.stringify` returns undefined for `undefined` / a function. + return redactUrlsInText(JSON.stringify(value, null, 2) ?? String(value)); +} diff --git a/clients/tui/src/utils/keybindings.ts b/clients/tui/src/utils/keybindings.ts new file mode 100644 index 0000000000..ca42c9cb4e --- /dev/null +++ b/clients/tui/src/utils/keybindings.ts @@ -0,0 +1,195 @@ +/** + * The TUI's keybinding reference, as data (#2436). + * + * The `?` help overlay (`components/HelpOverlay.tsx`) renders whatever this + * module returns, so **adding a keybinding anywhere in the TUI means adding a + * row here** — that is the whole of what keeps the overlay honest. The table is + * one flat entry per binding rather than prose so a new key is a one-line, + * append-only diff that does not collide with an unrelated one. + * + * `TAB_BINDINGS` is a full `Record<TabType, …>` on purpose: a new tab that adds + * no row fails to typecheck instead of silently showing an empty help section. + * + * Pure by design (utils = compute): no Ink, no React, no I/O. + */ +import { tabs, type TabType } from "../components/tabsConfig.js"; + +export interface KeyBinding { + /** The key or keys, as the user would type them (`"↑/↓"`, `"Shift+Tab"`). */ + keys: string; + /** What pressing it does, phrased as an action. */ + action: string; +} + +export interface KeyBindingSection { + title: string; + bindings: readonly KeyBinding[]; +} + +/** + * The tab accelerators, derived from the tabs the caller says are visible + * rather than restated, so a new tab's letter appears the moment it is added to + * the tab bar — and a hidden tab's letter, which does nothing, does not. + */ +function tabAcceleratorBinding( + visible: readonly { accelerator: string }[], +): KeyBinding { + return { + keys: visible.map((tab) => tab.accelerator).join(" "), + action: "Jump to a tab by its underlined letter", + }; +} + +/** Bindings that work wherever no dialog is open, whatever the active tab. */ +export function globalBindings( + visible: readonly { accelerator: string }[] = tabs, +): readonly KeyBinding[] { + return [ + { keys: "?", action: "Show or hide this help" }, + { keys: "Esc / Ctrl+C", action: "Exit (Esc closes a dialog first)" }, + { + keys: "Tab / Shift+Tab", + action: "Move focus: servers → tabs → list → details", + }, + { keys: "↑/↓", action: "Select a server (server list focused)" }, + { keys: "←/→", action: "Switch tab (tab bar focused)" }, + tabAcceleratorBinding(visible), + { keys: "c", action: "Connect the selected server" }, + { keys: "d", action: "Disconnect the selected server" }, + ]; +} + +const DETAILS_SCROLL: readonly KeyBinding[] = [ + { keys: "↑/↓", action: "Scroll the details pane (details focused)" }, + { keys: "PgUp/PgDn", action: "Scroll the details pane a page" }, + { keys: "+", action: "Open the details full screen (details focused)" }, + { keys: "y / w", action: "In a details dialog: copy / save the value" }, +]; + +const LIST_FILTER: KeyBinding = { + keys: "/", + action: "Filter the list (Enter keeps it, Esc clears it)", +}; + +const PANE_SCROLL: readonly KeyBinding[] = [ + { keys: "↑/↓", action: "Scroll (content focused)" }, + { keys: "PgUp/PgDn", action: "Scroll a page" }, +]; + +/** Bindings specific to one tab, shown only while that tab is active. */ +export const TAB_BINDINGS: Readonly<Record<TabType, readonly KeyBinding[]>> = { + info: [ + ...PANE_SCROLL, + { keys: "e", action: "Edit the advertised roots (content focused)" }, + ], + auth: [ + ...PANE_SCROLL, + { keys: "s", action: "Clear OAuth state (disconnects if connected)" }, + { keys: "↑/↓ + Enter", action: "Choose Authorize or Cancel (step-up)" }, + { keys: "a", action: "Authorize a pending step-up" }, + { keys: "c", action: "Cancel a pending step-up" }, + { keys: "y / w", action: "Copy / save the access token" }, + ], + resources: [ + { keys: "↑/↓", action: "Select a resource (list focused)" }, + { keys: "Enter", action: "Fetch the resource, or fill in a template" }, + LIST_FILTER, + ...DETAILS_SCROLL, + ], + prompts: [ + { keys: "↑/↓", action: "Select a prompt (list focused)" }, + { keys: "Enter", action: "Get the prompt (asks for arguments if any)" }, + LIST_FILTER, + ...DETAILS_SCROLL, + ], + skills: [ + { keys: "↑/↓", action: "Select a skill (list focused)" }, + { keys: "Enter", action: "Verify the skill's digests and frontmatter" }, + { + keys: "v", + action: "Verify every listed skill, under the catalog budget", + }, + LIST_FILTER, + { keys: "↑/↓", action: "Scroll the details pane (details focused)" }, + { keys: "PgUp/PgDn", action: "Scroll the details pane a page" }, + ], + tools: [ + { keys: "↑/↓", action: "Select a tool (list focused)" }, + { keys: "Enter", action: "Test the tool" }, + LIST_FILTER, + { + keys: "w", + action: "In the tool's result view: save the result to a file", + }, + ...DETAILS_SCROLL, + ], + messages: [ + { keys: "↑/↓", action: "Select a message (list focused)" }, + { keys: "PgUp/PgDn", action: "Move the selection a page (list focused)" }, + ...DETAILS_SCROLL, + ], + requests: [ + { keys: "↑/↓", action: "Select a request (list focused)" }, + { keys: "PgUp/PgDn", action: "Move the selection a page (list focused)" }, + ...DETAILS_SCROLL, + ], + logging: PANE_SCROLL, + subscriptions: [ + { keys: "↑/↓", action: "Select a resource (list focused)" }, + { keys: "Enter", action: "Subscribe to or unsubscribe from the resource" }, + ...PANE_SCROLL, + ], + tasks: [ + { keys: "↑/↓", action: "Select a task (list focused)" }, + { keys: "Enter", action: "Fetch the task's result" }, + { keys: "x", action: "Cancel the selected task" }, + { keys: "f", action: "Refresh the task list" }, + { keys: "l", action: "Clear finished tasks" }, + ...DETAILS_SCROLL, + ], +}; + +/** Bindings inside the help overlay itself. */ +export const HELP_BINDINGS: readonly KeyBinding[] = [ + { keys: "? / Esc", action: "Close this help" }, + { keys: "↑/↓, PgUp/PgDn", action: "Scroll this help" }, +]; + +/** + * Each tab's bar label, so a section is titled the way the tab bar shows it + * (`messages` reads "Protocol"). `tabs` lists every `TabType` — which is what + * makes the cast sound, and `keybindings.test.ts` pins it per tab. + */ +const TAB_LABELS = Object.fromEntries( + tabs.map((tab) => [tab.id, tab.label]), +) as Record<TabType, string>; + +/** + * The sections the help overlay shows for the active tab: what works + * everywhere, then what this tab adds, then how to leave the overlay. + * `visible` is the tab bar's current tabs (default: all of them). + */ +export function keybindingSections( + activeTab: TabType, + visible: readonly { accelerator: string }[] = tabs, +): readonly KeyBindingSection[] { + return [ + { title: "Global", bindings: globalBindings(visible) }, + { + title: `${TAB_LABELS[activeTab]} tab`, + bindings: TAB_BINDINGS[activeTab], + }, + { title: "This help", bindings: HELP_BINDINGS }, + ]; +} + +/** Width of the widest `keys` across `sections`, for aligning the columns. */ +export function keyColumnWidth(sections: readonly KeyBindingSection[]): number { + let width = 0; + for (const section of sections) { + for (const binding of section.bindings) { + width = Math.max(width, binding.keys.length); + } + } + return width; +} diff --git a/clients/tui/src/utils/openUrl.ts b/clients/tui/src/utils/openUrl.ts index a2a7d96398..0a4adca869 100644 --- a/clients/tui/src/utils/openUrl.ts +++ b/clients/tui/src/utils/openUrl.ts @@ -1,12 +1,33 @@ -import open from "open"; +import { openUrl as openInBrowser } from "@inspector/core/node/openUrl.js"; + +/** The note the TUI shows when it could not launch a browser for `href`. */ +export function browserOpenFailedMessage(href: string): string { + return `Could not open a browser automatically. Open this URL manually to authorize: ${href}`; +} /** - * Opens a URL in the user's default browser. - * Used when handling oauthAuthorizationRequired to launch the OAuth authorization page. + * Open the OAuth authorization page in the user's default browser, never + * rejecting. + * + * The launch goes through core's shared `openUrl`, which turns a spawn + * failure (the opener binary missing from `PATH`) into a rejection instead of + * an unlistened `'error'` event that crashed the TUI mid-flow (#2533). This + * wrapper then owns that rejection: `CallbackNavigation` discards the + * callback's promise, so a rejection escaping here would be unhandled and + * crash the process just the same. On failure it hands `onFailure` the + * "open the URL manually" note, which App renders on the Auth tab. * * @param url - URL to open (string or URL) - * @returns Promise that resolves when the opener completes (or rejects on error) + * @param onFailure - receives the fallback note when the browser did not open */ -export async function openUrl(url: string | URL): Promise<void> { - await open(typeof url === "string" ? url : url.href); +export async function openUrl( + url: string | URL, + onFailure?: (message: string) => void, +): Promise<void> { + const href = typeof url === "string" ? url : url.href; + try { + await openInBrowser(href); + } catch { + onFailure?.(browserOpenFailedMessage(href)); + } } diff --git a/clients/tui/src/utils/resourceUpdates.ts b/clients/tui/src/utils/resourceUpdates.ts new file mode 100644 index 0000000000..fac2c66ad1 --- /dev/null +++ b/clients/tui/src/utils/resourceUpdates.ts @@ -0,0 +1,80 @@ +/** + * Pure helpers behind the TUI Subscriptions tab (#2432). + * + * The live feed is read out of the protocol message log the TUI already keeps + * (`MessageLogState`), rather than from a new store: every + * `notifications/resources/updated` the server sends is already recorded there + * with its timestamp, in both protocol eras, so a feed derived from it cannot + * disagree with the Protocol tab about what arrived. + */ +import type { Resource } from "@modelcontextprotocol/client"; +import type { + InspectorResourceSubscription, + MessageEntry, +} from "@inspector/core/mcp/types.js"; + +export const RESOURCE_UPDATED_METHOD = "notifications/resources/updated"; + +export interface ResourceUpdateEvent { + id: string; + timestamp: Date; + uri: string; +} + +/** + * Every `notifications/resources/updated` in the log, newest first. Entries + * without a string `params.uri` are skipped — they name nothing to show. + */ +export function resourceUpdateFeed( + messages: readonly MessageEntry[], +): ResourceUpdateEvent[] { + const feed: ResourceUpdateEvent[] = []; + for (const entry of messages) { + if (entry.direction !== "notification") continue; + const message = entry.message; + if (!("method" in message) || message.method !== RESOURCE_UPDATED_METHOD) { + continue; + } + const uri = message.params?.uri; + if (typeof uri !== "string") continue; + feed.push({ id: entry.id, timestamp: entry.timestamp, uri }); + } + return feed.reverse(); +} + +export interface SubscribableResource { + resource: Resource; + subscribed: boolean; + lastUpdated?: Date; +} + +/** + * The rows the Subscriptions tab lists: every listed resource, then any + * subscribed URI the list does not contain (a template-expanded URI, or one the + * server has since dropped), so an active subscription can always be found and + * cancelled from here. + */ +export function subscribableResources( + resources: readonly Resource[], + subscriptions: readonly InspectorResourceSubscription[], +): SubscribableResource[] { + const byUri = new Map(subscriptions.map((s) => [s.resource.uri, s])); + const rows: SubscribableResource[] = resources.map((resource) => { + const sub = byUri.get(resource.uri); + return { + resource, + subscribed: sub !== undefined, + lastUpdated: sub?.lastUpdated, + }; + }); + const listed = new Set(resources.map((r) => r.uri)); + for (const sub of subscriptions) { + if (listed.has(sub.resource.uri)) continue; + rows.push({ + resource: sub.resource, + subscribed: true, + lastUpdated: sub.lastUpdated, + }); + } + return rows; +} diff --git a/clients/tui/src/utils/saveResult.ts b/clients/tui/src/utils/saveResult.ts new file mode 100644 index 0000000000..145bfb97d8 --- /dev/null +++ b/clients/tui/src/utils/saveResult.ts @@ -0,0 +1,118 @@ +/** + * Writing a tool call result to a file from the TUI's result view — the `w` + * keybinding (#2571). + * + * The bytes come from core's `renderResultForFile`: the `raw` / `json` + * encodings planned for the CLI's `--output` (#2431), kept in `core/` so that + * once it lands both clients render a result through the same code. What lives here is the TUI-only half: the default + * filename the prompt opens with, and the write itself, which resolves a + * relative path against the working directory the TUI was launched from and + * reports a failure as a message rather than letting it reach Ink. + */ +import { writeFile } from "node:fs/promises"; +import { resolve } from "node:path"; +import { + renderResultForFile, + renderedByteLength, + type ResultFileFormat, +} from "@inspector/core/mcp/resultFile.js"; + +/** What a successful save wrote, for the confirmation line. */ +export interface SavedResult { + path: string; + format: ResultFileFormat; + bytes: number; +} + +/** + * A tool name made safe as a bare filename: anything outside a conservative + * portable set becomes `_`, so a name like `fs/read` cannot point the default + * at a subdirectory. + */ +function fileSafe(name: string): string { + const safe = name.replace(/[^A-Za-z0-9._-]/g, "_"); + return safe === "" ? "tool" : safe; +} + +/** + * The filename the save prompt opens with: `<tool-name>-result.json` for + * `json`, and for `raw` `.txt` when the result renders as text or `.bin` when + * it renders as a single decoded binary block (an image or audio result). + */ +export function defaultResultFileName( + toolName: string, + format: ResultFileFormat, + result: unknown, +): string { + const base = `${fileSafe(toolName)}-result`; + if (format === "json") return `${base}.json`; + try { + const data = renderResultForFile(result, "tools/call", "raw"); + return `${base}.${typeof data === "string" ? "txt" : "bin"}`; + } catch { + // No raw form; the save itself reports why. The name only needs to be + // plausible. + return `${base}.txt`; + } +} + +/** + * The tail of every save started in this process. Saves run strictly one after + * another, so two writes can never overlap on one file and they settle in the + * order they were started — whichever reports last is the latest. It lives at + * module scope rather than in the result view because a save outlives the view + * that started it: closing the modal mid-write and saving again from a fresh + * one must still queue behind the first write. + */ +let saveQueue: Promise<unknown> = Promise.resolve(); + +/** + * Render `result` in `format` and write it to `path` (relative paths resolve + * against `cwd`), queued behind any save already in flight. An existing file is + * replaced and the parent directory must already exist. Any failure — a result + * with no raw form, or the write itself — rejects with an `Error` whose message + * is fit to show in the TUI, and does not hold up the saves queued after it. + */ +export function saveResultToFile( + result: unknown, + path: string, + format: ResultFileFormat, + cwd: string = process.cwd(), +): Promise<SavedResult> { + const run = saveQueue.then(() => writeResult(result, path, format, cwd)); + // The queue only orders saves; each caller handles its own rejection. + saveQueue = run.catch(() => undefined); + return run; +} + +async function writeResult( + result: unknown, + path: string, + format: ResultFileFormat, + cwd: string, +): Promise<SavedResult> { + const trimmed = path.trim(); + if (trimmed === "") throw new Error("Enter a file path to save to."); + const target = resolve(cwd, trimmed); + let data: string | Uint8Array; + try { + data = renderResultForFile(result, "tools/call", format); + } catch (err) { + throw new Error( + `${errorMessage(err)} Save it as json instead (w, then Enter).`, + { cause: err }, + ); + } + try { + await writeFile(target, data); + } catch (err) { + throw new Error(`Could not write ${target}: ${errorMessage(err)}`, { + cause: err, + }); + } + return { path: target, format, bytes: renderedByteLength(data) }; +} + +function errorMessage(err: unknown): string { + return err instanceof Error ? err.message : String(err); +} diff --git a/clients/tui/tsup.config.ts b/clients/tui/tsup.config.ts index e28cf3bc19..4d2814ab6f 100644 --- a/clients/tui/tsup.config.ts +++ b/clients/tui/tsup.config.ts @@ -133,6 +133,7 @@ export default defineConfig({ // client's own code, so all three lists carry it (AGENTS.md). The CLI was // inlining it; the #2067 guard surfaced that. "@modelcontextprotocol/ext-apps", + "@modelcontextprotocol/ext-tasks", "@napi-rs/keyring", // Root-declared (see the repo's dependency-placement rule) and CJS, which // is the combination that bites: tsup externalizes what the *client's* diff --git a/clients/tui/tui.tsx b/clients/tui/tui.tsx index 1768614868..4d3d604fc0 100644 --- a/clients/tui/tui.tsx +++ b/clients/tui/tui.tsx @@ -6,6 +6,7 @@ import { parseKeyValuePair, parseHeaderPair, parseProtocolEra, + skillCatalogLimitParser, } from "@inspector/core/mcp/node/index.js"; import type { ServerProtocolEra } from "@inspector/core/mcp/types.js"; import { loadRunnerClientConfig } from "@inspector/core/client/runner.js"; @@ -48,6 +49,16 @@ export async function runTui(args?: string[]): Promise<void> { "Protocol era to negotiate: legacy, auto, or modern (overrides the file's protocolEra; default legacy)", parseProtocolEra, ) + .option( + "--skill-catalog-max-skills <n>", + "Skills catalog budget: the most skills one multi-skill verification run reads (positive integer; overrides the file's skillCatalogMaxSkills; default 256). Bounds the Skills pane's v (verify all); Enter verifies one skill in full", + skillCatalogLimitParser("--skill-catalog-max-skills"), + ) + .option( + "--skill-catalog-max-bytes <n>", + "Skills catalog budget: the most bytes one multi-skill verification run reads (positive integer; overrides the file's skillCatalogMaxBytes; default 64 MiB). Bounds the Skills pane's v (verify all); Enter verifies one skill in full", + skillCatalogLimitParser("--skill-catalog-max-bytes"), + ) .option( "--client-id <id>", "OAuth client ID (static client) for HTTP servers", @@ -86,6 +97,8 @@ export async function runTui(args?: string[]): Promise<void> { cwd?: string; header?: Record<string, string>; protocolEra?: ServerProtocolEra; + skillCatalogMaxSkills?: number; + skillCatalogMaxBytes?: number; clientId?: string; clientSecret?: string; clientMetadataUrl?: string; @@ -104,6 +117,8 @@ export async function runTui(args?: string[]): Promise<void> { env: options.e, headers: options.header, protocolEra: options.protocolEra, + skillCatalogMaxSkills: options.skillCatalogMaxSkills, + skillCatalogMaxBytes: options.skillCatalogMaxBytes, transport: options.transport, serverUrl: options.serverUrl?.trim() || undefined, }; diff --git a/clients/web/README.md b/clients/web/README.md index bdce4099fb..fd64ea5dfd 100644 --- a/clients/web/README.md +++ b/clients/web/README.md @@ -19,6 +19,7 @@ The `server/` directory holds the Node-only backend: - **`web-server-config.ts`** — env parsing, the `GET /api/config` payload, the startup banner, the default origin allow-list. - **`resolve-bind-host.ts`** — the shared bind-host guard (refuses an all-interfaces `HOST` unless `DANGEROUSLY_BIND_ALL_INTERFACES`), used by both bind points (`web-server-config.ts` + `vite.config.ts`); see [Host binding & the origin allow-list](#host-binding--the-origin-allow-list). - **`inject-auth-token.ts`** — embeds the API token into the served `index.html` (see [Auth token](#auth-token)). +- **`health.ts`** — the unauthenticated `GET /healthz` liveness/readiness probe both backends answer (see [Health check](#health-check)). - **`sandbox-controller.ts`** — the MCP Apps sandbox HTTP server; **`app-origin-controller.ts`** — the dedicated app-origin server for `_meta.ui.domain` (see [MCP App dedicated origins](#mcp-app-dedicated-origins-metauidomain)); **`public-address.ts`** — validates `MCP_SANDBOX_FULL_ADDRESS` / `MCP_APP_ORIGIN_FULL_ADDRESS`, the public addresses those two advertise behind a reverse proxy; **`ensure-web-build.ts`** — builds `dist/` on demand for prod `--web`; **`vite-base-config.ts`** — shared `optimizeDeps` exclusions. - **`browser-externalized-builtin-gate.ts`** — Vite-agnostic build-gate logic that fails `vite build` when a Node built-in reaches the browser bundle (#1769); the thin Vite plugin wiring lives in `vite.config.ts`. It sits under `server/` (rather than `src/`) as the home for Node-only, build-time tooling — it's imported by the Vite config, never by the browser — alongside the other `vite-*` config helpers here. @@ -386,6 +387,18 @@ Storybook is first-class here because the components are presentational — each The dev/prod backend guards every `/api/*` route with `x-mcp-remote-auth: Bearer <MCP_INSPECTOR_API_TOKEN>`. The browser recovers the token, in priority order (see `App.tsx` `getAuthToken()`): the `window.__INSPECTOR_API_TOKEN__` global injected into `index.html` on every page load (`server/inject-auth-token.ts`), then a `?MCP_INSPECTOR_API_TOKEN=…` query param, then `sessionStorage`. Injection is a no-op when auth is disabled (`DANGEROUSLY_OMIT_AUTH`). See the root [AGENTS.md](../../AGENTS.md) for the full rationale, and [Environment variables](../../docs/environment-variables.md) for every variable the backend reads. +## Health check + +Both the prod server and the dev Vite backend answer **`GET /healthz`** (and `HEAD`) with `200` and a fixed `{"status":"ok"}` body, `Cache-Control: no-store` (#2438). It is meant for an orchestrator — a Docker or Kubernetes probe, a process manager — that needs to know the backend is up without exercising a real proxy/connect flow. + +- **It is unauthenticated, by design.** It sits outside `/api/*`, so neither the bearer-token check nor the origin allow-list applies; a probe has no way to learn a token that is generated per start. +- **It discloses nothing beyond "up"** — no version, uptime, servers, config or storage state. Anything that can reach the port can read it, so it says nothing `GET /` does not already say. +- **Liveness and readiness are one answer.** The backend only starts listening once the sandbox, the app-origin listener and the API are constructed, so any response means it is ready. + +```sh +curl -fsS http://127.0.0.1:6274/healthz # {"status":"ok"} +``` + ## Host binding & the origin allow-list Both the prod backend (`server/web-server-config.ts`) and the dev Vite server (`vite.config.ts`) resolve their bind host through one shared guard, `server/resolve-bind-host.ts`. It binds **`127.0.0.1`** by default and **refuses an all-interfaces host** (`0.0.0.0`, `::`, empty, or any equivalent spelling — `0`, `0x0`, `0.0`, `::0`, `::ffff:0.0.0.0`, … are all folded to the wildcard and refused) — which would expose the process-spawning backend to the whole network, the exposure DNS-rebinding attacks target — unless `DANGEROUSLY_BIND_ALL_INTERFACES=true` is set. The Docker image sets that flag (a container must bind `0.0.0.0` to be reachable through `-p`); a bare `HOST=0.0.0.0` anywhere else exits with an actionable error. diff --git a/clients/web/package-lock.json b/clients/web/package-lock.json index 324c16c8a7..ff27af3554 100644 --- a/clients/web/package-lock.json +++ b/clients/web/package-lock.json @@ -1393,9 +1393,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1561,9 +1558,6 @@ "arm64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1581,9 +1575,6 @@ "arm64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -1601,9 +1592,6 @@ "ppc64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1621,9 +1609,6 @@ "riscv64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1641,9 +1626,6 @@ "riscv64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -1661,9 +1643,6 @@ "s390x" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1681,9 +1660,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1701,9 +1677,6 @@ "x64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -1916,9 +1889,6 @@ "arm64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1933,9 +1903,6 @@ "arm64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -1950,9 +1917,6 @@ "ppc64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1967,9 +1931,6 @@ "riscv64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1984,9 +1945,6 @@ "riscv64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2001,9 +1959,6 @@ "s390x" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2018,9 +1973,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2035,9 +1987,6 @@ "x64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2239,9 +2188,6 @@ "arm64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2259,9 +2205,6 @@ "arm64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2279,9 +2222,6 @@ "ppc64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2299,9 +2239,6 @@ "s390x" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2319,9 +2256,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2339,9 +2273,6 @@ "x64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2574,9 +2505,6 @@ "arm" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2591,9 +2519,6 @@ "arm" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2608,9 +2533,6 @@ "arm64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2625,9 +2547,6 @@ "arm64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2642,9 +2561,6 @@ "loong64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2659,9 +2575,6 @@ "loong64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2676,9 +2589,6 @@ "ppc64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2693,9 +2603,6 @@ "ppc64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2710,9 +2617,6 @@ "riscv64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2727,9 +2631,6 @@ "riscv64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -2744,9 +2645,6 @@ "s390x" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2761,9 +2659,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -2778,9 +2673,6 @@ "x64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -5892,9 +5784,6 @@ "arm64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -5916,9 +5805,6 @@ "arm64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -5940,9 +5826,6 @@ "x64" ], "dev": true, - "libc": [ - "glibc" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -5964,9 +5847,6 @@ "x64" ], "dev": true, - "libc": [ - "musl" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -8272,9 +8152,9 @@ } }, "node_modules/source-map-js": { - "version": "1.2.1", - "resolved": "https://registry.npmjs.org/source-map-js/-/source-map-js-1.2.1.tgz", - "integrity": "sha512-UXWMKhLOwVKb728IUtQPXxfYU+usdybtUrK/8uGE8CQMvrhOpwvzDBwj0QhSL7MQc7vIsISBG8VQ8+IDQxpfQA==", + "version": "1.2.2", + "resolved": "https://registry.npmjs.org/source-map-js/-/source-map-js-1.2.2.tgz", + "integrity": "sha512-KGj/8Y43x35aZVDtt+J4mK1hoLGHULMYfSkODJNQjNDC3oW1PqPoxMwo0pLUsWM/UEGzON/NxeHywEfNXNP3Vw==", "dev": true, "license": "BSD-3-Clause", "engines": { @@ -9456,9 +9336,9 @@ } }, "node_modules/zod": { - "version": "4.4.3", - "resolved": "https://registry.npmjs.org/zod/-/zod-4.4.3.tgz", - "integrity": "sha512-ytENFjIJFl2UwYglde2jchW2Hwm4GJFLDiSXWdTrJQBIN9Fcyp7n4DhxJEiWNAJMV1/BqWfW/kkg71UDcHJyTQ==", + "version": "4.6.5", + "resolved": "https://registry.npmjs.org/zod/-/zod-4.6.5.tgz", + "integrity": "sha512-v5l/aFXZQeai4awLbOpSoHecE9UiMrnfx75tEXLjNonXVARxQ5mOeipTjROUchszUNCqnE+hqAMujRsRHsut2Q==", "dev": true, "license": "MIT", "funding": { diff --git a/clients/web/server/health.ts b/clients/web/server/health.ts new file mode 100644 index 0000000000..7e01db3d6f --- /dev/null +++ b/clients/web/server/health.ts @@ -0,0 +1,99 @@ +/** + * The web backend's liveness/readiness endpoint, `GET /healthz` (#2438). + * + * An orchestrator (Docker, Kubernetes, a process manager) needs a cheap probe + * that says "the backend is up" without exercising a real proxy/connect flow. + * `GET /` works but reads `index.html` from disk and embeds the API token into + * the response on every probe; this answers from memory and carries nothing. + * + * Three decisions, each deliberate: + * + * - **It lives outside `/api/*`, so it is unauthenticated.** Every `/api/*` + * route sits behind the `x-mcp-remote-auth` bearer check (and the origin + * allow-list), and a probe has no way to learn a token that is generated + * fresh per start. Putting the route under `/api` would mean carving an + * exception into that middleware; a top-level path needs none. + * - **It discloses nothing beyond "up".** The body is a fixed + * `{"status":"ok"}` — no version, uptime, connected servers, config or + * storage state. An unauthenticated route is readable by anything that can + * reach the port (including a DNS-rebinding page, which the origin check + * does not cover outside `/api`), and a version string is a fingerprinting + * aid. "Is it running" is already answerable from `GET /`, so this adds no + * new information. + * - **Liveness and readiness are the same answer here.** Both servers only + * start listening after the sandbox and app-origin listeners and the API app + * are fully constructed, so any response at all means the backend is ready. + * + * `Cache-Control: no-store` keeps an intermediary from answering a probe for a + * backend that has since died. + */ + +import type { IncomingMessage, ServerResponse } from "node:http"; + +/** The path the health route is served at, in both the prod and dev backends. */ +export const HEALTH_PATH = "/healthz"; + +/** The fixed health response body. Frozen: it is shared across requests. */ +export const HEALTH_BODY: Readonly<{ status: "ok" }> = Object.freeze({ + status: "ok", +}); + +/** Response headers sent with every health response. */ +export const HEALTH_HEADERS: Readonly<Record<string, string>> = Object.freeze({ + "Content-Type": "application/json; charset=utf-8", + "Cache-Control": "no-store", +}); + +/** Methods the route answers. `HEAD` is the cheapest probe some tools use. */ +const HEALTH_METHODS = new Set(["GET", "HEAD"]); + +/** + * True when a raw request target (`req.url` — a path plus optional query, as + * Node's `IncomingMessage` carries it) addresses the health route with a + * method it answers. The query string is ignored, so a cache-busting probe + * (`/healthz?t=…`) still matches; a trailing-slash or sub-path does not. + */ +export function isHealthRequest( + method: string | undefined, + url: string | undefined, +): boolean { + if (!method || !HEALTH_METHODS.has(method.toUpperCase())) return false; + const path = (url ?? "").split(/[?#]/, 1)[0]; + return path === HEALTH_PATH; +} + +/** The serialized body, or none for `HEAD`, per HTTP semantics. */ +function healthBodyFor(method: string): string | null { + return method.toUpperCase() === "HEAD" ? null : JSON.stringify(HEALTH_BODY); +} + +/** + * Build the health response (the prod Hono route). `HEAD` gets the same + * status and headers with no body. + */ +export function healthResponse(method = "GET"): Response { + return new Response(healthBodyFor(method), { + status: 200, + headers: { ...HEALTH_HEADERS }, + }); +} + +/** + * Answer the health route on a raw Node request/response pair (the dev Vite + * middleware, which sees Node's `IncomingMessage`, not a fetch `Request`). + * Returns `true` when it handled the request, `false` to let the caller pass + * it on. Kept here rather than inline in `vite-hono-plugin.ts` so the dev + * path is tested directly — that plugin is excluded from coverage as glue. + */ +export function handleNodeHealthRequest( + req: Pick<IncomingMessage, "method" | "url">, + res: Pick<ServerResponse, "writeHead" | "end">, +): boolean { + const { method } = req; + if (method === undefined || !isHealthRequest(method, req.url)) return false; + res.writeHead(200, { ...HEALTH_HEADERS }); + const body = healthBodyFor(method); + if (body === null) res.end(); + else res.end(body); + return true; +} diff --git a/clients/web/server/server.ts b/clients/web/server/server.ts index 21b344c0a2..efe82b7bdd 100644 --- a/clients/web/server/server.ts +++ b/clients/web/server/server.ts @@ -21,6 +21,7 @@ import { appDocumentEmbedders, } from "./app-origin-controller.js"; import { injectAuthToken } from "./inject-auth-token.js"; +import { HEALTH_PATH, healthResponse } from "./health.js"; import type { WebServerConfig } from "./web-server-config.js"; import { getSecretStorageInfo } from "../../../core/auth/node/secret-store-selection.ts"; import { @@ -94,6 +95,12 @@ export async function startHonoServer( }); const app = new Hono(); + // Unauthenticated liveness/readiness probe (#2438). Outside `/api/*`, so the + // auth middleware never sees it, and registered ahead of the static and SPA + // fallbacks below so it is not answered with index.html. See health.ts for + // what it does and does not disclose. Hono serves HEAD from this GET route. + app.get(HEALTH_PATH, (c) => healthResponse(c.req.method)); + app.use("/api/*", async (c) => { return apiApp.fetch(c.req.raw); }); diff --git a/clients/web/server/vite-hono-plugin.ts b/clients/web/server/vite-hono-plugin.ts index 19f92b7781..0852fcb061 100644 --- a/clients/web/server/vite-hono-plugin.ts +++ b/clients/web/server/vite-hono-plugin.ts @@ -21,6 +21,7 @@ import { appDocumentEmbedders, } from "./app-origin-controller.js"; import { injectAuthToken } from "./inject-auth-token.js"; +import { handleNodeHealthRequest } from "./health.js"; import type { WebServerConfig } from "./web-server-config.js"; import { getSecretStorageInfo } from "../../../core/auth/node/secret-store-selection.ts"; import { @@ -170,6 +171,11 @@ export function honoMiddlewarePlugin(config: WebServerConfig): Plugin { ) => { try { const pathname = req.url || ""; + // The same unauthenticated `/healthz` probe the prod server + // answers (#2438), so dev and prod agree; see health.ts. + if (handleNodeHealthRequest(req, res)) { + return; + } if (!pathname.startsWith("/api")) { return next(); } diff --git a/clients/web/server/web-server-config.ts b/clients/web/server/web-server-config.ts index 55283953e3..d8f8a24f5a 100644 --- a/clients/web/server/web-server-config.ts +++ b/clients/web/server/web-server-config.ts @@ -2,7 +2,6 @@ * Config object for the web server (dev and prod). Passed in-process; no env handoff. */ -import open from "open"; import pino from "pino"; import type { Logger } from "pino"; import type { MCPConfig, MCPServerConfig } from "../../../core/mcp/types.ts"; @@ -18,6 +17,7 @@ import type { InitialConfigPayload } from "../../../core/mcp/remote/node/server. import type { SecretStorageInfo } from "../../../core/auth/secret-storage-info.ts"; import { secretStorageSummary } from "../../../core/auth/secret-storage-info.ts"; import { readInspectorVersionSafe } from "../../../core/node/version.ts"; +import { openUrl } from "../../../core/node/openUrl.ts"; import { resolveSandboxPort } from "./sandbox-controller.js"; import { resolveAppOriginPort } from "./app-origin-controller.js"; import { isEnvFlagEnabled, resolveBindHostname } from "./resolve-bind-host.js"; @@ -262,7 +262,11 @@ export function printServerBanner( * `.catch` is two chances for one of them to drift back to a bare `open(url)`. */ export function openBrowser(url: string): void { - open(url).catch((err: unknown) => { + // Through core's `openUrl`, not a bare `open(url)`: `open` resolves before a + // spawn failure (the opener missing from `PATH`) arrives as an `'error'` + // event on the child, which no `.catch` here could see and which crashed + // the server at startup (#2533). `openUrl` turns it into a rejection. + openUrl(url).catch((err: unknown) => { console.warn("Could not open a browser automatically:", err); }); } diff --git a/clients/web/src/App.tsx b/clients/web/src/App.tsx index 3315d3381c..90f233ebb7 100644 --- a/clients/web/src/App.tsx +++ b/clients/web/src/App.tsx @@ -128,6 +128,7 @@ import { FetchBodyDroppedToastMessage } from "./components/elements/Toasts/Fetch import { HeadersReconnectToastMessage } from "./components/elements/Toasts/HeadersReconnectToastMessage"; import { OutputValidationToastMessage } from "./components/elements/Toasts/OutputValidationToastMessage"; import { ReAuthBannerBar } from "./components/groups/ReAuthBanner/ReAuthBannerBar"; +import { errorMessage } from "./utils/errorFormat"; /** * Terminates a dispatched handler's promise by reporting the failure, for the @@ -138,7 +139,7 @@ function reportDispatchFailure(title: string) { return (err: unknown) => { notifications.show({ title, - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", }); }; @@ -821,7 +822,7 @@ function App() { const name = sessionRef.current.activeServerName; notifications.show({ title: name ? `Connection to "${name}" lost` : "Connection lost", - message: lastError, + message: errorMessage(lastError), color: "red", }); }, [sessionRef, lastError]); @@ -968,7 +969,7 @@ function App() { const onAppError = useCallback((err: Error) => { notifications.show({ title: "MCP App error", - message: err.message, + message: errorMessage(err), color: "red", }); }, []); @@ -1385,7 +1386,7 @@ function App() { err instanceof ServerListReloadError ? `Saved settings for "${id}", but the server list did not reload` : `Failed to save settings for "${id}"`, - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", }); }, @@ -1498,7 +1499,7 @@ function App() { .catch((err: unknown) => { notifications.show({ title: "Could not clear the stored OAuth state", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", // The tokens may still be on disk and the session may still be up, // so this is not a notice to let time out. @@ -1723,7 +1724,7 @@ function App() { reorderServers(orderedIds).catch((err: unknown) => { notifications.show({ title: "Failed to reorder servers", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", }); }); diff --git a/clients/web/src/components/elements/ContentViewer/ContentViewer.test.tsx b/clients/web/src/components/elements/ContentViewer/ContentViewer.test.tsx index 66d5ace345..0268edf42d 100644 --- a/clients/web/src/components/elements/ContentViewer/ContentViewer.test.tsx +++ b/clients/web/src/components/elements/ContentViewer/ContentViewer.test.tsx @@ -5,7 +5,7 @@ import type { TextResourceContents, } from "@modelcontextprotocol/client"; import { renderWithMantine, screen } from "../../../test/renderWithMantine"; -import { getAceText } from "../../../test/aceEditor"; +import { getAceEditor, getAceText } from "../../../test/aceEditor"; import { ContentViewer } from "./ContentViewer"; // Stub the lazy highlighter so JSON/XML/CSS branches are assertable @@ -249,6 +249,44 @@ describe("ContentViewer", () => { ).toBeInTheDocument(); }); + // #2525: the row cap is the default, so a host that scrolls a single payload + // itself can lift it rather than nest Ace's scrollbar inside its own. + it("caps the JSON editor's height by default", () => { + const block: ContentBlock = { type: "text", text: '{"a":1}' }; + renderWithMantine(<ContentViewer block={block} />); + expect(getAceEditor().getOption("maxLines")).toBe(200); + }); + + it("passes jsonMaxLines through to a declared-JSON editor", () => { + const block: ContentBlock = { type: "text", text: '{"a":1}' }; + renderWithMantine( + <ContentViewer + block={block} + mimeType="application/json" + jsonMaxLines={Infinity} + />, + ); + expect(getAceEditor().getOption("maxLines")).toBe(Infinity); + }); + + it("passes jsonMaxLines through to heuristically-detected JSON", () => { + const block: ContentBlock = { type: "text", text: '{"a":1}' }; + renderWithMantine(<ContentViewer block={block} jsonMaxLines={500} />); + expect(getAceEditor().getOption("maxLines")).toBe(500); + }); + + it("passes jsonMaxLines through to JSON resource contents", () => { + const contents: TextResourceContents = { + uri: "demo://r.json", + mimeType: "application/json", + text: '{"a":1}', + }; + renderWithMantine( + <ContentViewer contents={contents} jsonMaxLines={Infinity} />, + ); + expect(getAceEditor().getOption("maxLines")).toBe(Infinity); + }); + it("falls back to a generic name when no label is given", () => { const block: ContentBlock = { type: "text", text: '{"a":1}' }; renderWithMantine(<ContentViewer block={block} />); diff --git a/clients/web/src/components/elements/ContentViewer/ContentViewer.tsx b/clients/web/src/components/elements/ContentViewer/ContentViewer.tsx index 65d149eb84..4f0f6ed3c5 100644 --- a/clients/web/src/components/elements/ContentViewer/ContentViewer.tsx +++ b/clients/web/src/components/elements/ContentViewer/ContentViewer.tsx @@ -73,6 +73,21 @@ export interface ContentViewerProps { * and takes no name. */ jsonLabel?: string; + /** + * Rows a JSON payload grows to before the Ace editor starts scrolling + * internally. Defaults to {@link JSON_DISPLAY_MAX_LINES}. + * + * Pass `Infinity` when the host already scrolls the payload itself and shows + * only one at a time — the Tools result panel (#2525). With the default cap a + * payload longer than it gets an Ace scrollbar *inside* the host's, two + * vertical scrollbars over the same content. The default stays capped for + * the hosts the cap exists for: the Protocol and Network lists, which keep + * every entry's payload mounted, and where an uncapped editor would put + * every row of every response in the DOM. + * + * Reaches the JSON branch only, like `jsonLabel`. + */ + jsonMaxLines?: number; } /** @@ -203,6 +218,7 @@ function JsonContent({ formatted, copyable, label, + maxLines, }: { /** The original payload — what the copy button copies. */ text: string; @@ -210,6 +226,8 @@ function JsonContent({ formatted: string; copyable: boolean; label: string; + /** Row cap before Ace scrolls internally — see `ContentViewerProps.jsonMaxLines`. */ + maxLines: number; }) { return ( <CopyableWrapper copyable={copyable} copyValue={text}> @@ -234,9 +252,11 @@ function JsonContent({ // editor would put a whole 10k-line response in the DOM per entry. // // 200 is chosen so essentially no real payload nests a scrollbar while - // the pathological one stays bounded. + // the pathological one stays bounded. It is the *default*: a host that + // mounts a single payload at a time and owns the scroll container can + // lift it via `jsonMaxLines` (the Tools result panel does, #2525). minLines={1} - maxLines={JSON_DISPLAY_MAX_LINES} + maxLines={maxLines} /> </CopyableWrapper> ); @@ -247,11 +267,13 @@ function PlainTextContent({ copyable, wrap, jsonLabel, + jsonMaxLines, }: { text: string; copyable: boolean; wrap: boolean; jsonLabel: string; + jsonMaxLines: number; }) { // Untyped text that really parses as JSON gets the JSON renderer too — the // heuristic is how a server that sent no MIME type still reads well. Not when @@ -268,6 +290,7 @@ function PlainTextContent({ formatted={asJson} copyable={copyable} label={jsonLabel} + maxLines={jsonMaxLines} /> ); } @@ -301,12 +324,14 @@ function TextualContent({ copyable, wrap, jsonLabel, + jsonMaxLines, }: { text: string; mimeType: string | undefined; copyable: boolean; wrap: boolean; jsonLabel: string; + jsonMaxLines: number; }) { const kind = mimeType ? getMimeKind(mimeType) : "text"; switch (kind) { @@ -327,6 +352,7 @@ function TextualContent({ formatted={formatJson(text)} copyable={copyable} label={jsonLabel} + maxLines={jsonMaxLines} /> ) : ( <PlainTextContent @@ -334,6 +360,7 @@ function TextualContent({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); case "xml": @@ -373,6 +400,7 @@ function TextualContent({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); } @@ -403,12 +431,14 @@ function ResourceContent({ copyable, wrap, jsonLabel, + jsonMaxLines, }: { contents: TextResourceContents | BlobResourceContents; mimeType: string; copyable: boolean; wrap: boolean; jsonLabel: string; + jsonMaxLines: number; }) { if ("text" in contents) { return ( @@ -418,6 +448,7 @@ function ResourceContent({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); } @@ -447,6 +478,7 @@ function ResourceContent({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); } @@ -460,12 +492,14 @@ function BlockContent({ copyable, wrap, jsonLabel, + jsonMaxLines, }: { block: ContentBlock; mimeType: string | undefined; copyable: boolean; wrap: boolean; jsonLabel: string; + jsonMaxLines: number; }) { switch (block.type) { case "text": @@ -476,6 +510,7 @@ function BlockContent({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); case "image": @@ -503,6 +538,7 @@ function BlockContent({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); } @@ -539,6 +575,7 @@ export function ContentViewer({ mimeType, wrap = true, jsonLabel = "JSON content", + jsonMaxLines = JSON_DISPLAY_MAX_LINES, }: ContentViewerProps) { if (contents) { const effective = @@ -550,6 +587,7 @@ export function ContentViewer({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); } @@ -561,6 +599,7 @@ export function ContentViewer({ copyable={copyable} wrap={wrap} jsonLabel={jsonLabel} + jsonMaxLines={jsonMaxLines} /> ); } diff --git a/clients/web/src/components/elements/ListLoadError/ListLoadError.tsx b/clients/web/src/components/elements/ListLoadError/ListLoadError.tsx index fdb4ef8392..b753941b48 100644 --- a/clients/web/src/components/elements/ListLoadError/ListLoadError.tsx +++ b/clients/web/src/components/elements/ListLoadError/ListLoadError.tsx @@ -1,4 +1,5 @@ import { Alert, Button, Code, ScrollArea, Stack } from "@mantine/core"; +import { errorMessage } from "../../../utils/errorFormat"; export interface ListLoadErrorProps { /** @@ -56,7 +57,7 @@ export function ListLoadError({ error, what, onRetry }: ListLoadErrorProps) { <ErrorAlert title={`Couldn't load ${what}`}> <Stack gap="xs"> <MessageScroll> - <ErrorMessage>{error.message}</ErrorMessage> + <ErrorMessage>{errorMessage(error)}</ErrorMessage> </MessageScroll> {onRetry && <RetryButton onClick={onRetry}>Retry</RetryButton>} </Stack> diff --git a/clients/web/src/components/groups/ConnectionInfoContent/ConnectionInfoContent.tsx b/clients/web/src/components/groups/ConnectionInfoContent/ConnectionInfoContent.tsx index 6c60488029..6228c15c40 100644 --- a/clients/web/src/components/groups/ConnectionInfoContent/ConnectionInfoContent.tsx +++ b/clients/web/src/components/groups/ConnectionInfoContent/ConnectionInfoContent.tsx @@ -26,7 +26,7 @@ import { NO_OUTSTANDING_REQUESTS_LABEL, } from "../../../utils/connectionActivity"; import { useTickingClock } from "../../../hooks/useTickingClock"; -import { TASKS_EXTENSION_KEY } from "@inspector/core/mcp/modernTaskSchemas.js"; +import { TASKS_EXTENSION_KEY } from "@inspector/core/extension/tasks/constants.js"; import { getSkillsExtension } from "@inspector/core/mcp/skills.js"; import type { OAuthClientRegistrationKind } from "@inspector/core/auth/types.js"; import { diff --git a/clients/web/src/components/groups/ProtocolListPanel/ProtocolListPanel.tsx b/clients/web/src/components/groups/ProtocolListPanel/ProtocolListPanel.tsx index 0ebe496912..fbc8d12981 100644 --- a/clients/web/src/components/groups/ProtocolListPanel/ProtocolListPanel.tsx +++ b/clients/web/src/components/groups/ProtocolListPanel/ProtocolListPanel.tsx @@ -316,7 +316,7 @@ export function ProtocolListPanel({ // a single unit; everything else stays a plain ProtocolEntry. `sectionPinned` // is the section's pin state (used for a lone entry's pin label). const renderRows = (sectionEntries: MessageEntry[], sectionPinned: boolean) => - groupProtocolEntries(sectionEntries).map((row) => + groupProtocolEntries(sectionEntries, sortDirection).map((row) => row.kind === "mrtr" ? ( <MrtrConversation key={`mrtr-${row.requestState}-${row.rounds[0].id}`} diff --git a/clients/web/src/components/groups/ResourceLink/ResourceLink.tsx b/clients/web/src/components/groups/ResourceLink/ResourceLink.tsx index f00de8b4b7..9321b10eca 100644 --- a/clients/web/src/components/groups/ResourceLink/ResourceLink.tsx +++ b/clients/web/src/components/groups/ResourceLink/ResourceLink.tsx @@ -4,6 +4,7 @@ import type { ReadResourceResult } from "@modelcontextprotocol/client"; import { ContentViewer } from "../../elements/ContentViewer/ContentViewer"; import { ExpandToggle } from "../../elements/ExpandToggle/ExpandToggle"; import { ResourceLinkInfo } from "../../elements/ResourceLinkInfo/ResourceLinkInfo"; +import { errorMessage } from "../../../utils/errorFormat"; export interface ResourceLinkProps { /** The linked resource's URI (always shown). */ @@ -96,7 +97,7 @@ export function ResourceLink({ try { setResult(await onReadResource(uri)); } catch (err) { - setError(err instanceof Error ? err.message : String(err)); + setError(errorMessage(err)); } finally { setLoading(false); } diff --git a/clients/web/src/components/groups/ServerConfigModal/ServerConfigModal.tsx b/clients/web/src/components/groups/ServerConfigModal/ServerConfigModal.tsx index c89c2e06a7..8d8ddec523 100644 --- a/clients/web/src/components/groups/ServerConfigModal/ServerConfigModal.tsx +++ b/clients/web/src/components/groups/ServerConfigModal/ServerConfigModal.tsx @@ -17,6 +17,7 @@ import type { MCPServerConfig, StdioServerConfig, } from "@inspector/core/mcp/types.js"; +import { errorMessage } from "../../../utils/errorFormat"; /** Allowed id pattern — mirrors validateStoreId in core/storage/store-io.ts */ const ID_PATTERN = /^[a-zA-Z0-9_-]+$/; @@ -304,7 +305,7 @@ export function ServerConfigModal({ await onSubmit(trimmedId, built.config); onClose(); } catch (err) { - setSubmitError(err instanceof Error ? err.message : String(err)); + setSubmitError(errorMessage(err)); } finally { setSubmitting(false); } diff --git a/clients/web/src/components/groups/ServerRemoveConfirmModal/ServerRemoveConfirmModal.tsx b/clients/web/src/components/groups/ServerRemoveConfirmModal/ServerRemoveConfirmModal.tsx index 7ce9cf7386..acf1308d02 100644 --- a/clients/web/src/components/groups/ServerRemoveConfirmModal/ServerRemoveConfirmModal.tsx +++ b/clients/web/src/components/groups/ServerRemoveConfirmModal/ServerRemoveConfirmModal.tsx @@ -1,6 +1,7 @@ import { useState } from "react"; import { Button, Group, Modal, Paper, Stack, Text } from "@mantine/core"; import type { ServerEntry } from "@inspector/core/mcp/types.js"; +import { errorMessage } from "../../../utils/errorFormat"; export interface ServerRemoveConfirmModalProps { opened: boolean; @@ -55,7 +56,7 @@ export function ServerRemoveConfirmModal({ try { await onConfirm(); } catch (err) { - setError(err instanceof Error ? err.message : String(err)); + setError(errorMessage(err)); } finally { setSubmitting(false); } diff --git a/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.test.tsx b/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.test.tsx index fe5c9330cf..d77095c26c 100644 --- a/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.test.tsx +++ b/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.test.tsx @@ -4,6 +4,7 @@ import { within } from "@testing-library/react"; import type { InspectorServerSettings } from "@inspector/core/mcp/types.js"; import { renderWithMantine, screen } from "../../../test/renderWithMantine"; import { getAceText, setAceText } from "../../../test/aceEditor"; +import { redirectUrlProvider } from "../../../lib/authToken"; import { ServerSettingsForm, type ServerSettingsSection, @@ -1567,6 +1568,54 @@ describe("ServerSettingsForm", () => { ).toBeInTheDocument(); }); + // #2524 — a pre-registered client must register the redirect URI on its + // authorization server, so the form shows the exact value the flow sends. + it("shows the read-only OAuth redirect URI with a copy control (#2524)", async () => { + const user = userEvent.setup(); + const writeText = vi.fn().mockResolvedValue(undefined); + Object.defineProperty(navigator, "clipboard", { + value: { writeText }, + configurable: true, + }); + renderWithMantine( + <ServerSettingsForm + {...baseHandlers} + settings={emptySettings} + expandedSections={["oauth"]} + />, + ); + const expected = redirectUrlProvider.getRedirectUrl(); + expect(expected).toBe(`${window.location.origin}/oauth/callback`); + const input = screen.getByLabelText("Redirect URI"); + expect(input).toHaveValue(expected); + expect(input).toHaveAttribute("readonly"); + expect( + screen.getByText(/depends on the origin the Inspector is opened from/i), + ).toBeInTheDocument(); + expect( + screen.getByText(/when using a pre-registered client ID/i), + ).toBeInTheDocument(); + await user.click(screen.getByRole("button", { name: "Copy Redirect URI" })); + expect(writeText).toHaveBeenCalledWith(expected); + }); + + it("still shows the redirect URI when EMA is on (#2524)", () => { + renderWithMantine( + <ServerSettingsForm + {...baseHandlers} + settings={{ ...emptySettings, enterpriseManaged: true }} + expandedSections={["oauth"]} + />, + ); + expect(screen.getByLabelText("Redirect URI")).toHaveValue( + redirectUrlProvider.getRedirectUrl(), + ); + // Under EMA the URI is registered on the IdP client, not the Resource AS. + expect( + screen.getByText(/enterprise IdP client configured in Client Settings/i), + ).toBeInTheDocument(); + }); + it("hides the OAuth Settings section for stdio servers", () => { renderWithMantine( <ServerSettingsForm diff --git a/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.tsx b/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.tsx index c9569691a4..18dd96ab2c 100644 --- a/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.tsx +++ b/clients/web/src/components/groups/ServerSettingsForm/ServerSettingsForm.tsx @@ -15,6 +15,8 @@ import { } from "@mantine/core"; import { ClearButton } from "../../elements/ClearButton/ClearButton"; import { JsonObjectInput } from "../../elements/JsonObjectInput/JsonObjectInput"; +import { CopyButton } from "../../elements/CopyButton/CopyButton"; +import { redirectUrlProvider } from "../../../lib/authToken"; import type { ChangeEvent } from "react"; import type { ProtocolEra } from "@modelcontextprotocol/client"; import type { @@ -192,6 +194,13 @@ const ClearableTextInput = TextInput.withProps({ rightSectionPointerEvents: "auto", }); +// Read-only display of the OAuth redirect URI (#2524). `rightSectionPointerEvents` +// keeps the CopyButton in the right section clickable, as for ClearableTextInput. +const RedirectUriInput = TextInput.withProps({ + readOnly: true, + rightSectionPointerEvents: "auto", +}); + const LogSizeInput = NumberInput.withProps({ min: 0, step: 100, @@ -523,6 +532,19 @@ export function ServerSettingsForm({ const clientSecretLabel = enterpriseManaged ? "Resource AS Client Secret" : "Client Secret"; + // The exact redirect URI the connect path sends (#2524). Read from the same + // `redirectUrlProvider` the OAuth flow uses, so this field cannot drift from + // the value the authorization server actually receives. It follows the + // address bar's origin, which is why it is computed rather than hard-coded. + const oauthRedirectUri = redirectUrlProvider.getRedirectUrl(); + // Under EMA the authorization request carrying this URI goes to the + // enterprise IdP, so it is registered on the IdP client from Client Settings + // — not on the Resource AS client named by the fields beside it. + const redirectUriOriginNote = + "It depends on the origin the Inspector is opened from — localhost vs 127.0.0.1, a different port or host each change it."; + const redirectUriDescription = enterpriseManaged + ? `Register this exact URI on the enterprise IdP client configured in Client Settings — the IdP authorization request carries it, not the Resource AS client above. ${redirectUriOriginNote}` + : `Register this exact URI with your authorization server when using a pre-registered client ID. ${redirectUriOriginNote}`; const resourceAsDescription = enterpriseManaged ? "The resource authorization server's registered client credential (EMA leg 3) — not the app client id/secret, which belong in Client Settings." : undefined; @@ -1005,6 +1027,14 @@ export function ServerSettingsForm({ ) : null } /> + <RedirectUriInput + label="Redirect URI" + description={redirectUriDescription} + value={oauthRedirectUri} + rightSection={ + <CopyButton value={oauthRedirectUri} label="Redirect URI" /> + } + /> <ClearableTextInput label="Scopes" description="Space-separated OAuth scopes (RFC 6749). Do not use commas — a comma-separated entry is sent as one invalid token and rejected by the authorization server." diff --git a/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.stories.tsx b/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.stories.tsx index 683bb42381..b1f2ad5bf8 100644 --- a/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.stories.tsx +++ b/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.stories.tsx @@ -1,7 +1,7 @@ import type { Decorator, Meta, StoryObj } from "@storybook/react-vite"; import type { CallToolResult } from "@modelcontextprotocol/client"; import { Card, Flex } from "@mantine/core"; -import { fn } from "storybook/test"; +import { expect, fn, waitFor } from "storybook/test"; import { ToolResultPanel } from "./ToolResultPanel"; const meta: Meta<typeof ToolResultPanel> = { @@ -134,6 +134,52 @@ export const JsonResult: Story = { }, }; +// A JSON result longer than ContentViewer's default 200-row editor cap. The +// panel's own scroll area is the only vertical scrollbar: the editor grows to +// the payload's full height rather than scrolling inside it (#2525). +export const LongJsonResult: Story = { + args: { + result: { + content: [ + { + type: "text", + text: JSON.stringify( + Array.from({ length: 120 }, (_, id) => ({ + id, + name: `Item ${id}`, + })), + ), + }, + ], + }, + }, + decorators: fillHeightDecorators, + // Padded rather than the default centered layout, so the card takes the + // canvas width the way it does on the Tools screen instead of shrinking to + // the editor's gutter. + parameters: { layout: "padded" }, + // Asserted on real geometry: the editor's own vertical scrollbar stays + // hidden (Ace sets `display: none` on it when it has nothing to scroll), + // while the panel's scroll viewport is what overflows. + play: async ({ canvasElement }) => { + const aceScrollbar = await waitFor(() => { + const node = canvasElement.querySelector(".ace_scrollbar-v"); + if (!(node instanceof HTMLElement)) { + throw new Error("JSON editor not rendered"); + } + return node; + }); + await expect(aceScrollbar.style.display).toBe("none"); + const viewport = canvasElement.querySelector( + ".mantine-ScrollArea-viewport", + ); + if (!(viewport instanceof HTMLElement)) { + throw new Error("panel scroll viewport not found"); + } + await expect(viewport.scrollHeight).toBeGreaterThan(viewport.clientHeight); + }, +}; + export const ImageResult: Story = { args: { result: imageResult, diff --git a/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.test.tsx b/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.test.tsx index e1d783a4c0..f599498b1e 100644 --- a/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.test.tsx +++ b/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.test.tsx @@ -6,7 +6,7 @@ import { screen, waitFor, } from "../../../test/renderWithMantine"; -import { getAceText } from "../../../test/aceEditor"; +import { getAceEditor, getAceText } from "../../../test/aceEditor"; import { ToolResultPanel } from "./ToolResultPanel"; import { resultHasResourceLinks } from "./toolResultUtils"; @@ -119,6 +119,42 @@ describe("ToolResultPanel", () => { expect(screen.getByText("divider")).toBeInTheDocument(); }); + // #2525: the panel's ScrollArea already scrolls a long JSON result, so an + // Ace editor that also stopped at its row cap and scrolled internally put two + // vertical scrollbars side by side. The editor grows to full height instead. + describe("long JSON results (#2525)", () => { + const longJson = JSON.stringify( + Array.from({ length: 400 }, (_, id) => ({ id })), + null, + 2, + ); + + it("lets the JSON editor grow so the panel is the only scroller", () => { + renderWithMantine( + <ToolResultPanel + result={{ content: [{ type: "text", text: longJson }] }} + onClear={() => {}} + />, + ); + expect(getAceEditor().getOption("maxLines")).toBe(Infinity); + }); + + it("lets it grow inside the capped block beside Resource Links too", () => { + renderWithMantine( + <ToolResultPanel + result={{ + content: [ + { type: "text", text: longJson }, + { type: "resource_link", uri: "demo://r/1", name: "One" }, + ], + }} + onClear={() => {}} + />, + ); + expect(getAceEditor().getOption("maxLines")).toBe(Infinity); + }); + }); + describe("structuredContent (#1908)", () => { const structured = { items: [{ id: 1, name: "Item A" }], total: 1 }; diff --git a/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.tsx b/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.tsx index 772965dc34..200fededb7 100644 --- a/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.tsx +++ b/clients/web/src/components/groups/ToolResultPanel/ToolResultPanel.tsx @@ -94,6 +94,15 @@ const ResultScroll = ScrollArea.withProps({ offsetScrollbars: true, }); +// Every scroll region in this panel (`ResultScroll`, `NonLinkCap`) already +// scrolls a JSON block whole, so the block's Ace editor must grow to its full +// height rather than stop at `ContentViewer`'s default row cap and scroll +// inside — which is what put two vertical scrollbars side by side on a long +// result (#2525). Lifting the cap is affordable here because the panel shows +// one result at a time; the cap exists for the Protocol and Network lists, +// which keep every payload mounted. +const UNCAPPED_JSON_LINES = Infinity; + const ResultStack = Stack.withProps({ gap: "md", }); @@ -253,6 +262,7 @@ export function ToolResultPanel({ key={segment.index} block={segment.block} copyable={segment.block.type === "text"} + jsonMaxLines={UNCAPPED_JSON_LINES} /> ); // Alongside a Resource Links box, cap the block at half the height (and let diff --git a/clients/web/src/components/groups/protocolUtils.test.ts b/clients/web/src/components/groups/protocolUtils.test.ts index fff5c5e650..bae6e54b2f 100644 --- a/clients/web/src/components/groups/protocolUtils.test.ts +++ b/clients/web/src/components/groups/protocolUtils.test.ts @@ -390,4 +390,120 @@ describe("groupProtocolEntries", () => { { kind: "mrtr", requestState: "A", rounds: [a2] }, ]); }); + + // #2608: `mrtr_two_step` mints a fresh token each round, so no single key + // spans the conversation — each retry echoes the token the round before it + // was issued. + function rotatingConversation(): MessageEntry[] { + return [ + callEntry( + "orig", + 1, + { name: "mrtr_two_step" }, + { resultType: "input_required", requestState: "step2" }, + 1, + ), + callEntry( + "retry1", + 2, + { name: "mrtr_two_step", requestState: "step2", inputResponses: {} }, + { resultType: "input_required", requestState: "done" }, + 2, + ), + callEntry( + "retry2", + 3, + { name: "mrtr_two_step", requestState: "done", inputResponses: {} }, + { resultType: "complete", content: [] }, + 3, + ), + ]; + } + + it("chains rounds when the server rotates requestState each round", () => { + const rounds = rotatingConversation(); + expect(groupProtocolEntries(rounds)).toEqual([ + { kind: "mrtr", requestState: "step2", rounds }, + ]); + }); + + it("chains a rotating conversation sorted newest-first, named by its first token", () => { + const rounds = rotatingConversation().reverse(); + expect(groupProtocolEntries(rounds, "newest-first")).toEqual([ + { kind: "mrtr", requestState: "step2", rounds }, + ]); + }); + + it("names a lone retry by the token it echoed", () => { + const [, retry1] = rotatingConversation(); + expect(groupProtocolEntries([retry1])).toEqual([ + { kind: "mrtr", requestState: "step2", rounds: [retry1] }, + ]); + }); + + it("clusters two retries that echo the same token after an error round", () => { + const original = callEntry( + "orig", + 1, + {}, + { resultType: "input_required", requestState: "tok" }, + 1, + ); + const failed: MessageEntry = { + ...callEntry("failed", 2, { requestState: "tok" }, undefined, 2), + response: { + jsonrpc: "2.0", + id: 2, + error: { code: -32603, message: "boom" }, + }, + }; + const retried = callEntry( + "retried", + 3, + { requestState: "tok" }, + { resultType: "complete", content: [] }, + 3, + ); + expect(groupProtocolEntries([original, failed, retried])).toEqual([ + { + kind: "mrtr", + requestState: "tok", + rounds: [original, failed, retried], + }, + ]); + }); + + // A server may reuse an opaque token, so a hand-off can run back to a value + // already seen. Naming must not depend on the sort order (Copilot, #2610). + it("names a conversation with a reused token the same in both sort orders", () => { + const ab = callEntry( + "ab", + 1, + { requestState: "A" }, + { resultType: "input_required", requestState: "B" }, + 1, + ); + const ba = callEntry( + "ba", + 2, + { requestState: "B" }, + { resultType: "input_required", requestState: "A" }, + 2, + ); + expect(groupProtocolEntries([ab, ba], "oldest-first")).toEqual([ + { kind: "mrtr", requestState: "A", rounds: [ab, ba] }, + ]); + expect(groupProtocolEntries([ba, ab], "newest-first")).toEqual([ + { kind: "mrtr", requestState: "A", rounds: [ba, ab] }, + ]); + }); + + it("does not chain backwards in the opposite sort order", () => { + const [orig, retry1] = rotatingConversation(); + // Oldest-first, a later entry that the earlier one would continue is not + // its successor, so the two stay separate rows. + expect(groupProtocolEntries([retry1, orig], "oldest-first")).toHaveLength( + 2, + ); + }); }); diff --git a/clients/web/src/components/groups/protocolUtils.ts b/clients/web/src/components/groups/protocolUtils.ts index cc778ff225..65aaff4aa3 100644 --- a/clients/web/src/components/groups/protocolUtils.ts +++ b/clients/web/src/components/groups/protocolUtils.ts @@ -1,4 +1,5 @@ import type { MessageEntry, MessageMethod } from "@inspector/core/mcp/types.js"; +import type { SortDirection } from "../elements/SortToggle/SortToggle.js"; import { isInputRequiredResult, SUBSCRIPTION_ID_META_KEY, @@ -64,6 +65,21 @@ export function extractResultType( return result.resultType === "complete" ? "complete" : undefined; } +// A non-empty `requestState` string, or undefined. +function asRequestState(value: unknown): string | undefined { + return typeof value === "string" && value.length > 0 ? value : undefined; +} + +// The `requestState` a round's `input_required` result *issued* to the client. +function issuedRequestState(entry: MessageEntry): string | undefined { + return asRequestState(getResponseResult(entry)?.requestState); +} + +// The `requestState` a retried round *echoed* back in its params. +function echoedRequestState(entry: MessageEntry): string | undefined { + return asRequestState(getMessageParams(entry)?.requestState); +} + /** * The opaque MRTR `requestState` token that links the rounds of one logical * operation across multiple JSON-RPC ids. It appears on the `input_required` @@ -71,15 +87,7 @@ export function extractResultType( * request (spec §7.3). Returns undefined for non-MRTR traffic. */ export function extractRequestState(entry: MessageEntry): string | undefined { - const result = getResponseResult(entry); - const fromResult = result?.requestState; - if (typeof fromResult === "string" && fromResult.length > 0) - return fromResult; - const params = getMessageParams(entry); - const fromParams = params?.requestState; - if (typeof fromParams === "string" && fromParams.length > 0) - return fromParams; - return undefined; + return issuedRequestState(entry) ?? echoedRequestState(entry); } /** @@ -96,7 +104,7 @@ export function extractSubscriptionId(entry: MessageEntry): string | undefined { /** * A rendered row in the Protocol list: either a single message entry, or an - * MRTR conversation — the contiguous run of entries sharing one `requestState` + * MRTR conversation — the contiguous run of entries linked by `requestState` * (original call → `input_required` → retried call → final result), grouped so * one logical operation renders as one expandable unit. */ @@ -104,22 +112,64 @@ export type ProtocolRow = | { kind: "single"; entry: MessageEntry } | { kind: "mrtr"; requestState: string; rounds: MessageEntry[] }; +// Whether `later` is the next round of `earlier`'s conversation: it echoes the +// token `earlier`'s result issued. The token is opaque, so a server is free to +// mint a fresh one each round (#2608) — rounds are linked by this hand-off, not +// by one shared key. Two retries echoing the same token (a round whose result +// was an error, retried) are siblings of one conversation too. +function continuesConversation( + earlier: MessageEntry, + later: MessageEntry, +): boolean { + const echoed = echoedRequestState(later); + if (echoed === undefined) return false; + return ( + echoed === issuedRequestState(earlier) || + echoed === echoedRequestState(earlier) + ); +} + +// The token a conversation is named by: the one its chronologically first +// round received (a lone retry) or else issued (the original call). +function conversationRequestState(first: MessageEntry): string | undefined { + return echoedRequestState(first) ?? issuedRequestState(first); +} + /** * Fold a (already filtered/sorted) entry list into rows, clustering contiguous - * entries that share a non-empty `requestState` into one MRTR row. Contiguity is - * safe because the SDK auto-fulfils MRTR input in-process (no intervening wire - * frames) so an operation's rounds are adjacent in the log. Order is preserved; - * everything without a `requestState` stays a `single` row. + * MRTR rounds into one MRTR row. Each round joins the row before it when the + * two hand a `requestState` from one to the other (see + * `continuesConversation`). `sortDirection` says which of the two is the + * earlier round: the hand-off is checked in that direction only, because an + * opaque token may be reused (`A → B`, then `B → A`) and checking both ways + * would then name or group a conversation differently per sort order. + * Contiguity is safe because the SDK auto-fulfils MRTR input in-process (no + * intervening wire frames) so an operation's rounds are adjacent in the log. + * Order is preserved; everything without a `requestState` stays a `single` + * row. */ -export function groupProtocolEntries(entries: MessageEntry[]): ProtocolRow[] { +export function groupProtocolEntries( + entries: MessageEntry[], + sortDirection: SortDirection = "oldest-first", +): ProtocolRow[] { + const newestFirst = sortDirection === "newest-first"; const rows: ProtocolRow[] = []; for (const entry of entries) { - const requestState = extractRequestState(entry); + const requestState = conversationRequestState(entry); if (requestState) { const last = rows[rows.length - 1]; - if (last?.kind === "mrtr" && last.requestState === requestState) { - last.rounds.push(entry); - continue; + if (last?.kind === "mrtr") { + const neighbour = last.rounds[last.rounds.length - 1]; + const joins = newestFirst + ? continuesConversation(entry, neighbour) + : continuesConversation(neighbour, entry); + if (joins) { + last.rounds.push(entry); + // Newest-first: this entry is the earlier round, so the + // conversation is now named by its token. + if (newestFirst) last.requestState = requestState; + continue; + } } rows.push({ kind: "mrtr", requestState, rounds: [entry] }); continue; diff --git a/clients/web/src/components/screens/AppsScreen/AppsScreen.tsx b/clients/web/src/components/screens/AppsScreen/AppsScreen.tsx index ab56111f04..7b89b56e4e 100644 --- a/clients/web/src/components/screens/AppsScreen/AppsScreen.tsx +++ b/clients/web/src/components/screens/AppsScreen/AppsScreen.tsx @@ -43,6 +43,7 @@ import { ContentViewer } from "../../elements/ContentViewer/ContentViewer"; import { LogLevelBadge } from "../../elements/LogLevelBadge/LogLevelBadge"; import { hasInputFields, resolveDisplayLabel } from "../../../utils/toolUtils"; import { collectSchemaDefaults, toFormSchema } from "../../../utils/jsonUtils"; +import { errorMessage } from "../../../utils/errorFormat"; export interface AppsScreenProps { tools: Tool[]; @@ -605,7 +606,9 @@ export function AppsScreen({ <ContentCard data-testid="apps-form" data-app-status={running ? appStatus : "idle"} - data-app-error={running ? appError?.message : undefined} + data-app-error={ + running && appError ? errorMessage(appError) : undefined + } > {selectedTool ? ( <ContentStack> @@ -683,7 +686,7 @@ export function AppsScreen({ {appError && ( <AppErrorPanel data-testid="apps-error"> <AppErrorTitle>App failed to load</AppErrorTitle> - <AppErrorMessage>{appError.message}</AppErrorMessage> + <AppErrorMessage>{errorMessage(appError)}</AppErrorMessage> </AppErrorPanel> )} </RendererContainer> diff --git a/clients/web/src/components/screens/SkillsScreen/SkillsScreen.tsx b/clients/web/src/components/screens/SkillsScreen/SkillsScreen.tsx index ac0e2173ff..5159b16c11 100644 --- a/clients/web/src/components/screens/SkillsScreen/SkillsScreen.tsx +++ b/clients/web/src/components/screens/SkillsScreen/SkillsScreen.tsx @@ -56,6 +56,7 @@ import { isMarkdownMime, } from "../../../utils/inferMimeFromUri"; import { tryDecodeBase64ToUtf8 } from "../../elements/ContentViewer/contentViewerUtils"; +import { errorMessage } from "../../../utils/errorFormat"; /** * How many skill files are read at once by "Verify all". A conforming manifest @@ -1012,7 +1013,7 @@ export function SkillsScreen({ write({ attempt, status: "error", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), }); } }, @@ -1121,7 +1122,7 @@ export function SkillsScreen({ .catch((err: unknown) => { writePreview({ uri, - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), }); }); }, @@ -1269,7 +1270,7 @@ export function SkillsScreen({ ...(cursor === undefined ? {} : { children: prev.children, nextCursor: cursor }), - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), })); }); }, @@ -1321,7 +1322,7 @@ export function SkillsScreen({ }) .catch((err: unknown) => { writeFetched({ - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), }); }); }, [manifestKey, onGetSkill, openConformance, selected]); @@ -1627,7 +1628,7 @@ export function SkillsScreen({ </ControlsRow> {loadError && ( <Alert color="red" title="Could not load skills"> - {loadError.message} + {errorMessage(loadError)} </Alert> )} {filtered.length === 0 ? ( diff --git a/clients/web/src/components/views/InspectorView/InspectorView.tsx b/clients/web/src/components/views/InspectorView/InspectorView.tsx index 9ee980f8cc..0ab8f3b1e6 100644 --- a/clients/web/src/components/views/InspectorView/InspectorView.tsx +++ b/clients/web/src/components/views/InspectorView/InspectorView.tsx @@ -12,7 +12,7 @@ import type { Implementation, Tool } from "@modelcontextprotocol/client"; import type { ServerEntry } from "@inspector/core/mcp/types.js"; import { isTerminalStatus } from "@inspector/core/mcp/types.js"; import { isAppTool } from "@inspector/core/mcp/apps.js"; -import { TASKS_EXTENSION_KEY } from "@inspector/core/mcp/modernTaskSchemas.js"; +import { TASKS_EXTENSION_KEY } from "@inspector/core/extension/tasks/constants.js"; import { isSkillsExtensionSupported } from "@inspector/core/mcp/skills.js"; import { ViewHeader } from "../../groups/ViewHeader/ViewHeader"; import { VersionBadge } from "../../elements/VersionBadge/VersionBadge"; diff --git a/clients/web/src/hooks/useConnectionLifecycle.test.tsx b/clients/web/src/hooks/useConnectionLifecycle.test.tsx index 8cd7d8c0b9..e66c4c7bbc 100644 --- a/clients/web/src/hooks/useConnectionLifecycle.test.tsx +++ b/clients/web/src/hooks/useConnectionLifecycle.test.tsx @@ -1206,6 +1206,38 @@ describe("useConnectionLifecycle", () => { ); }); + // #2490: the recorded message is redacted for display, but the 409 + // classification reads the raw text — so a phrase that only survives in + // the unredacted message still swallows, and the record is redacted. + it("classifies on the raw message and records the redacted one", async () => { + const swallow = vi + .fn() + .mockRejectedValue( + new Error("rejected https://s.example/?token=already exists"), + ); + const h1 = harness({ + servers: [], + deepLink: deepLink(), + addServerImpl: swallow, + }); + await waitFor(() => expect(swallow).toHaveBeenCalled()); + expect(h1.api().connectErrorMessage).toBeUndefined(); + + const record = vi + .fn() + .mockRejectedValue(new Error("rejected https://s.example/?code=abc")); + const h2 = harness({ + servers: [], + deepLink: deepLink(), + addServerImpl: record, + }); + await waitFor(() => + expect(h2.api().connectErrorMessage).toBe( + "rejected https://s.example/?code=%5BREDACTED%5D", + ), + ); + }); + it("records an update failure", async () => { const link = deepLink(); const updateServerImpl = vi.fn().mockRejectedValue("backend 500"); diff --git a/clients/web/src/hooks/useConnectionLifecycle.ts b/clients/web/src/hooks/useConnectionLifecycle.ts index dc554e7ff0..bdf21cf7c2 100644 --- a/clients/web/src/hooks/useConnectionLifecycle.ts +++ b/clients/web/src/hooks/useConnectionLifecycle.ts @@ -54,6 +54,7 @@ import { OAUTH_CALLBACK_PATH, isUnauthorizedError } from "../utils/oauthFlow"; import { authRecoveryRestoredMessage } from "../utils/oauthUx"; import { deepLinkConfigEquals } from "../utils/deepLink"; import type { DeepLink } from "../utils/deepLink"; +import { errorMessage } from "../utils/errorFormat"; /** * Client identity name the web client reports to servers. It matches core's @@ -736,10 +737,7 @@ export function useConnectionLifecycle({ return; } setFailedServerId(id); - const message = - recoveryErr instanceof Error - ? recoveryErr.message - : String(recoveryErr); + const message = errorMessage(recoveryErr); setConnectErrorMessage(message); notifications.show({ title: `Failed to connect to "${target.name}"`, @@ -805,8 +803,7 @@ export function useConnectionLifecycle({ // `disconnect()` above settles it at `"disconnected"`), so this // flag is the only signal the view has that a connect attempt died. setFailedServerId(id); - const message = - authErr instanceof Error ? authErr.message : String(authErr); + const message = errorMessage(authErr); setConnectErrorMessage(message); notifications.show({ title: `OAuth authorization failed for "${target.name}"`, @@ -821,7 +818,7 @@ export function useConnectionLifecycle({ // instead of the ConnectionToggle silently reverting to // "disconnected", and flag the card with a red border (#1621). setFailedServerId(id); - const message = err instanceof Error ? err.message : String(err); + const message = errorMessage(err); setConnectErrorMessage(message); notifications.show({ title: `Failed to connect to "${target.name}"`, @@ -938,13 +935,16 @@ export function useConnectionLifecycle({ if (deepLinkEnsureRef.current) return; deepLinkEnsureRef.current = true; void addServer(deepLink.serverId, deepLink.serverConfig).catch((err) => { - const message = err instanceof Error ? err.message : String(err); + // Classify on the raw text; only the recorded (displayed) copy is + // redacted, so a phrase inside a redacted query value cannot flip it. + const raw = err instanceof Error ? err.message : String(err); // A 409 ("already exists") means the row is on disk and hydration will // surface it on a later render, so the connect phase still proceeds — // swallow it. Any other failure (read-only catalog, backend 5xx) would // otherwise leave the deep link permanently stuck at this guard with no // signal, so record it on the machine-readable error surface. - if (!message.includes("already exists")) recordConnectError(message); + if (!raw.includes("already exists")) + recordConnectError(errorMessage(err)); }); return; } @@ -957,7 +957,7 @@ export function useConnectionLifecycle({ deepLink.serverId, deepLink.serverConfig, ).catch((err) => { - const message = err instanceof Error ? err.message : String(err); + const message = errorMessage(err); recordConnectError(message); }); return; @@ -985,7 +985,7 @@ export function useConnectionLifecycle({ void onToggleConnection(deepLink.serverId).catch((err) => { // The toast fires from inside `onToggleConnection` for the common // cases; this catch covers the rest (surfaced on `data-error-message`). - const message = err instanceof Error ? err.message : String(err); + const message = errorMessage(err); recordConnectError(message); }); } @@ -1018,7 +1018,7 @@ export function useConnectionLifecycle({ try { await onToggleConnection(serverId); } catch (err) { - recordConnectError(err instanceof Error ? err.message : String(err)); + recordConnectError(errorMessage(err)); } }, [onToggleConnection, recordConnectError], @@ -1059,7 +1059,7 @@ export function useConnectionLifecycle({ } catch (err) { notifications.show({ title: "Could not clear the stored authorization state", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", // The banner is already dismissed and the flow is dead, so this // is the only remaining explanation — don't time it out. @@ -1115,7 +1115,7 @@ export function useConnectionLifecycle({ ) { return; } - const message = err instanceof Error ? err.message : String(err); + const message = errorMessage(err); notifications.show({ title: server ? `OAuth authorization failed for "${server.name}"` diff --git a/clients/web/src/hooks/useExportActions.ts b/clients/web/src/hooks/useExportActions.ts index f367df0a73..aadbf448d2 100644 --- a/clients/web/src/hooks/useExportActions.ts +++ b/clients/web/src/hooks/useExportActions.ts @@ -16,6 +16,7 @@ import { replayProtocolRequest, type ReplayParamsOverride, } from "../lib/protocolReplay"; +import { errorMessage } from "../utils/errorFormat"; /** One Protocol panel section, split by pin membership. */ export type ProtocolSection = "pinned" | "history"; @@ -213,7 +214,7 @@ export function useExportActions({ .catch((err: unknown) => { notifications.show({ title: "Replay failed", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", }); }); diff --git a/clients/web/src/hooks/useImportClientConfig.ts b/clients/web/src/hooks/useImportClientConfig.ts index c5e439f8e1..b5470f057f 100644 --- a/clients/web/src/hooks/useImportClientConfig.ts +++ b/clients/web/src/hooks/useImportClientConfig.ts @@ -10,6 +10,7 @@ import { import { validateStoreId } from "@inspector/core/storage/store-id.js"; import { ServerListReloadError } from "@inspector/core/react/useServers.js"; import type { MCPConfig, MCPServerConfig } from "@inspector/core/mcp/types.js"; +import { errorMessage } from "../utils/errorFormat"; export type ImportPhase = "select" | "loading" | "review" | "summary"; @@ -161,7 +162,7 @@ export function useImportClientConfig({ } beginReview(result.config); } catch (err) { - setError(err instanceof Error ? err.message : String(err)); + setError(errorMessage(err)); setPhase("select"); } } @@ -174,7 +175,7 @@ export function useImportClientConfig({ const raw = await file.text(); beginReview(parseClientConfig(raw)); } catch (err) { - setError(err instanceof Error ? err.message : String(err)); + setError(errorMessage(err)); } } @@ -199,7 +200,7 @@ export function useImportClientConfig({ await write(); return { id, status: outcome }; } catch (err) { - const detail = err instanceof Error ? err.message : String(err); + const detail = errorMessage(err); // A `ServerListReloadError` means the write landed and only reading the // list back failed, so this entry really was imported. Reporting it as // `failed` would contradict the error's own message and invite a retry diff --git a/clients/web/src/hooks/useMcpApps.ts b/clients/web/src/hooks/useMcpApps.ts index 8a90be2caf..5dc8a8e8cc 100644 --- a/clients/web/src/hooks/useMcpApps.ts +++ b/clients/web/src/hooks/useMcpApps.ts @@ -22,6 +22,7 @@ import { type AppElicitationSession, } from "../lib/appElicitationController"; import { getAuthToken } from "../lib/authToken"; +import { errorMessage } from "../utils/errorFormat"; export interface UseMcpAppsOptions { /** The live client. Null while disconnected; the bridges read it lazily. */ @@ -156,7 +157,7 @@ export function useMcpApps({ onResourceError: (err) => { notifications.show({ title: "App resource failed to load", - message: err.message, + message: errorMessage(err), color: "red", }); }, @@ -207,7 +208,7 @@ export function useMcpApps({ onResourceError: (err) => { notifications.show({ title: "Elicitation app failed to load", - message: err.message, + message: errorMessage(err), color: "red", }); }, diff --git a/clients/web/src/hooks/useOAuthRecovery.test.tsx b/clients/web/src/hooks/useOAuthRecovery.test.tsx index 757db8a60f..6ba9f18078 100644 --- a/clients/web/src/hooks/useOAuthRecovery.test.tsx +++ b/clients/web/src/hooks/useOAuthRecovery.test.tsx @@ -2008,6 +2008,15 @@ describe("useOAuthRecovery", () => { expect(text).toContain("may still be valid"); }); + it("redacts URL query secrets in the failure detail (#2490)", () => { + const text = revocationSuffix({ + status: "failed", + detail: "POST https://as.example/revoke?token=s3cr3t failed", + }); + expect(text).toContain("token=%5BREDACTED%5D"); + expect(text).not.toContain("s3cr3t"); + }); + it("says nothing for a skip or an absent outcome", () => { expect( revocationSuffix({ status: "skipped", reason: "no_endpoint" }), diff --git a/clients/web/src/hooks/useOAuthRecovery.ts b/clients/web/src/hooks/useOAuthRecovery.ts index 292152f6fc..313143a6a2 100644 --- a/clients/web/src/hooks/useOAuthRecovery.ts +++ b/clients/web/src/hooks/useOAuthRecovery.ts @@ -79,6 +79,8 @@ import { } from "../utils/stepUp"; import type { SessionRef } from "./useSessionRef"; import type { TabUiState, TabUiStateSetters } from "./useTabUiState"; +import { errorMessage } from "../utils/errorFormat"; +import { redactUrlsInText } from "@inspector/core/mcp/fetchTracking.js"; /** The banner raised when a session needs the user to authorize again. */ export interface ReAuthBannerState { @@ -158,7 +160,8 @@ export function revocationSuffix( return " The grant was also revoked at the authorization server."; } if (outcome?.status === "failed") { - return ` Revoking the grant at the authorization server failed (${outcome.detail}), so it may still be valid there.`; + // The detail is a caught revocation/network error's text (#2490). + return ` Revoking the grant at the authorization server failed (${redactUrlsInText(outcome.detail)}), so it may still be valid there.`; } return ""; } @@ -436,7 +439,9 @@ export function useOAuthRecovery({ const message = reAuthBannerMessage({ serverName: server?.name, detail: - detail !== undefined ? formatOAuthFailureDetail(detail) : undefined, + detail !== undefined + ? redactUrlsInText(formatOAuthFailureDetail(detail)) + : undefined, }); const reason = options?.reason; if (reason !== undefined && !isReAuthBannerReason(reason)) { @@ -916,7 +921,7 @@ export function useOAuthRecovery({ if (!errorTitle) return; notifications.show({ title: errorTitle, - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", }); }, @@ -1037,7 +1042,7 @@ export function useOAuthRecovery({ if (stillThisSession) { setPendingReauth((prev) => prev ?? pending); } - const detail = err instanceof Error ? err.message : String(err); + const detail = errorMessage(err); notifications.show({ title: "Could not continue authorization", // The copy has to match what actually happens. When the session is @@ -1594,7 +1599,7 @@ export function useOAuthRecovery({ // cleared-successfully toast below still goes out. notifications.show({ title: "Cleared, but the session did not disconnect cleanly", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "yellow", }); } finally { @@ -1692,7 +1697,7 @@ export function useOAuthRecovery({ // command and is reported as that command's — routed to the // panel that issued it rather than dressed up as a step-up // failure (#2165). - const message = err instanceof Error ? err.message : String(err); + const message = errorMessage(err); notifications.show({ title: "Retry failed", message, @@ -1721,7 +1726,9 @@ export function useOAuthRecovery({ // Terminal, and this attempt's retry is already out of the shared // ref — so it dies with the attempt, and whatever a later prompt // has installed is left alone. - const failureMessage = emaStepUpFailureMessage(outcome.error.message); + const failureMessage = emaStepUpFailureMessage( + errorMessage(outcome.error), + ); notifications.show({ title: "Organization permissions", message: failureMessage, @@ -1736,9 +1743,7 @@ export function useOAuthRecovery({ // latch and the prompt is already dismissed, leaving the panel that // asked for the permissions never told (#2165). Reported exactly as // the `failed` outcome above is, so the two cannot disagree. - const failureMessage = emaStepUpFailureMessage( - err instanceof Error ? err.message : String(err), - ); + const failureMessage = emaStepUpFailureMessage(errorMessage(err)); notifications.show({ title: "Organization permissions", message: failureMessage, diff --git a/clients/web/src/hooks/useServerCommands.tsx b/clients/web/src/hooks/useServerCommands.tsx index 2f02bec753..e84d54f190 100644 --- a/clients/web/src/hooks/useServerCommands.tsx +++ b/clients/web/src/hooks/useServerCommands.tsx @@ -508,7 +508,7 @@ export function useServerCommands({ setGetPromptState({ status: "error", promptName: name, - error: err instanceof Error ? err.message : String(err), + error: errorMessage(err), }); } }, @@ -546,7 +546,7 @@ export function useServerCommands({ setReadResourceState({ status: "error", uri, - error: err instanceof Error ? err.message : String(err), + error: errorMessage(err), }); } }, @@ -666,7 +666,7 @@ export function useServerCommands({ } notifications.show({ title: "Failed to cancel task", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", }); } @@ -869,7 +869,7 @@ export function useServerCommands({ notifications.show({ title: "Pagination setting saved, but the server list did not reload", - message: err.message, + message: errorMessage(err), color: "red", }); return; @@ -913,7 +913,7 @@ export function useServerCommands({ } notifications.show({ title: "Failed to save pagination setting", - message: err instanceof Error ? err.message : String(err), + message: errorMessage(err), color: "red", }); }); diff --git a/clients/web/src/hooks/useServerJsonImport.ts b/clients/web/src/hooks/useServerJsonImport.ts index 5692140e38..5929f8a62b 100644 --- a/clients/web/src/hooks/useServerJsonImport.ts +++ b/clients/web/src/hooks/useServerJsonImport.ts @@ -14,6 +14,7 @@ import type { PackageInfo, ValidationResult, } from "../components/groups/ImportServerJsonPanel/ImportServerJsonPanel"; +import { errorMessage } from "../utils/errorFormat"; /** Debounce (ms) before a textarea edit re-triggers parse/validation. */ export const VALIDATE_DEBOUNCE_MS = 300; @@ -48,7 +49,7 @@ function parseDraft(rawText: string): ParseState { } catch (err) { return { ok: false, - error: err instanceof Error ? err.message : String(err), + error: errorMessage(err), }; } } @@ -252,7 +253,7 @@ export function useServerJsonImport({ const text = await file.text(); setDraft((d) => ({ ...d, rawText: text })); } catch (err) { - setSubmitError(err instanceof Error ? err.message : String(err)); + setSubmitError(errorMessage(err)); } } @@ -283,7 +284,7 @@ export function useServerJsonImport({ await onAddServer(sel.serverId, config); return true; } catch (err) { - setSubmitError(err instanceof Error ? err.message : String(err)); + setSubmitError(errorMessage(err)); return false; } } diff --git a/clients/web/src/test/core/auth/ema/idpOidc.test.ts b/clients/web/src/test/core/auth/ema/idpOidc.test.ts index 1f9fec6263..71376f4537 100644 --- a/clients/web/src/test/core/auth/ema/idpOidc.test.ts +++ b/clients/web/src/test/core/auth/ema/idpOidc.test.ts @@ -227,9 +227,47 @@ describe("discoverIdpMetadata", () => { vi.restoreAllMocks(); }); - it("returns parsed OAuth metadata from OIDC discovery", async () => { + it("prefers the OIDC discovery document and preserves end_session_endpoint", async () => { const fetchFn = vi.fn(async (input: RequestInfo | URL) => { const url = String(input); + if (url === `${IDP_ISSUER}/.well-known/openid-configuration`) { + return new Response( + JSON.stringify({ + ...minimalOAuthAsMetadata(IDP_ISSUER), + end_session_endpoint: `${IDP_ISSUER}/session/end`, + }), + ); + } + throw new Error(`unexpected fetch: ${url}`); + }); + const metadata = await discoverIdpMetadata(IDP_ISSUER, fetchFn); + expect(metadata.issuer).toBe(IDP_ISSUER); + expect( + (metadata as { end_session_endpoint?: unknown }).end_session_endpoint, + ).toBe(`${IDP_ISSUER}/session/end`); + // The RFC 8414 path is never consulted when the OIDC document answers. + expect(fetchFn).toHaveBeenCalledTimes(1); + }); + + it("appends the well-known path after a path-suffixed issuer (OIDC Discovery §4.1)", async () => { + const issuer = "https://idp.refresh.test/tenant"; + const fetchFn = vi.fn(async (input: RequestInfo | URL) => { + const url = String(input); + if (url === `${issuer}/.well-known/openid-configuration`) { + return new Response(JSON.stringify(minimalOAuthAsMetadata(issuer))); + } + throw new Error(`unexpected fetch: ${url}`); + }); + const metadata = await discoverIdpMetadata(issuer, fetchFn); + expect(metadata.issuer).toBe(issuer); + }); + + it("falls back to SDK discovery when the OIDC document is unavailable", async () => { + const fetchFn = vi.fn(async (input: RequestInfo | URL) => { + const url = String(input); + if (url.includes("/.well-known/openid-configuration")) { + return new Response("not found", { status: 404 }); + } if (url.includes("/.well-known/oauth-authorization-server")) { return new Response(JSON.stringify(minimalOAuthAsMetadata(IDP_ISSUER))); } @@ -240,6 +278,59 @@ describe("discoverIdpMetadata", () => { expect(metadata.issuer).toBe(IDP_ISSUER); }); + it("rejects an OIDC document whose issuer does not match (OIDC Discovery §4.3), falling back to SDK discovery", async () => { + const fetchFn = vi.fn(async (input: RequestInfo | URL) => { + const url = String(input); + if (url.includes("/.well-known/openid-configuration")) { + // Valid metadata, wrong issuer — must be treated as invalid. + return new Response( + JSON.stringify(minimalOAuthAsMetadata("https://evil.example")), + ); + } + if (url.includes("/.well-known/oauth-authorization-server")) { + return new Response(JSON.stringify(minimalOAuthAsMetadata(IDP_ISSUER))); + } + throw new Error(`unexpected fetch: ${url}`); + }); + const metadata = await discoverIdpMetadata(IDP_ISSUER, fetchFn); + expect(metadata.issuer).toBe(IDP_ISSUER); + }); + + it("rejects an OIDC document whose issuer differs even by a trailing slash (exact match), falling back to SDK discovery", async () => { + const fetchFn = vi.fn(async (input: RequestInfo | URL) => { + const url = String(input); + if (url === `${IDP_ISSUER}/.well-known/openid-configuration`) { + return new Response( + JSON.stringify(minimalOAuthAsMetadata(`${IDP_ISSUER}/`)), + ); + } + if (url.includes("/.well-known/oauth-authorization-server")) { + return new Response(JSON.stringify(minimalOAuthAsMetadata(IDP_ISSUER))); + } + throw new Error(`unexpected fetch: ${url}`); + }); + const metadata = await discoverIdpMetadata(IDP_ISSUER, fetchFn); + expect(metadata.issuer).toBe(IDP_ISSUER); + // The OIDC document was rejected; the SDK path answered. + expect(fetchFn.mock.calls.length).toBeGreaterThan(1); + }); + + it("falls back to SDK discovery when the OIDC document is not valid metadata", async () => { + const fetchFn = vi.fn(async (input: RequestInfo | URL) => { + const url = String(input); + if (url.includes("/.well-known/openid-configuration")) { + // Parses as JSON but fails OAuthMetadataSchema (missing endpoints). + return new Response(JSON.stringify({ hello: "world" })); + } + if (url.includes("/.well-known/oauth-authorization-server")) { + return new Response(JSON.stringify(minimalOAuthAsMetadata(IDP_ISSUER))); + } + throw new Error(`unexpected fetch: ${url}`); + }); + const metadata = await discoverIdpMetadata(IDP_ISSUER, fetchFn); + expect(metadata.issuer).toBe(IDP_ISSUER); + }); + it("throws when discovery yields no metadata (all probes 404)", async () => { const errorSpy = vi.spyOn(console, "error").mockImplementation(() => {}); const fetchFn = vi.fn( diff --git a/clients/web/src/test/core/auth/ema/idpSession.test.ts b/clients/web/src/test/core/auth/ema/idpSession.test.ts index 465be78532..6647db4a28 100644 --- a/clients/web/src/test/core/auth/ema/idpSession.test.ts +++ b/clients/web/src/test/core/auth/ema/idpSession.test.ts @@ -5,6 +5,23 @@ import { getEmaIdpLoginState, normalizeIdpIssuer, } from "@inspector/core/auth/ema/idpSession.js"; +import type { OAuthMetadata } from "@modelcontextprotocol/client"; + +/** + * Minimal valid RFC 8414 metadata plus extras. `end_session_endpoint` is an + * OIDC-layer field outside the typed shape, so it rides in via spread (the + * SDK parses discovery responses with a loose object, so the metadata cached + * at login carries it the same way). + */ +function idpMetadata(extra: Record<string, unknown> = {}): OAuthMetadata { + return { + issuer: "https://idp.test", + authorization_endpoint: "https://idp.test/authorize", + token_endpoint: "https://idp.test/token", + response_types_supported: ["code"], + ...extra, + }; +} function jwtWithExp(expSec: number): string { const payload = btoa(JSON.stringify({ exp: expSec })) @@ -20,11 +37,12 @@ describe("idpSession", () => { beforeEach(() => { storage = { load: vi.fn().mockResolvedValue(undefined), - getIdpSession: vi.fn(), + getIdpSession: vi.fn().mockResolvedValue(undefined), saveIdpSession: vi.fn(), clearIdpSession: vi.fn(), clear: vi.fn(), clearEnterpriseManagedResourceServers: vi.fn(), + getServerMetadata: vi.fn().mockResolvedValue(null), takeRevocationSnapshot: vi.fn().mockResolvedValue({ byIssuer: {} }), } as unknown as OAuthStorage; }); @@ -76,10 +94,12 @@ describe("idpSession", () => { }); it("clearEmaIdpSession clears idp session, leg-1 key, and tagged resource servers", async () => { - await clearEmaIdpSession(storage, "https://idp.test/"); + const result = await clearEmaIdpSession(storage, "https://idp.test/"); expect(storage.clearIdpSession).toHaveBeenCalledWith("https://idp.test"); expect(storage.clear).toHaveBeenCalledWith("ema-idp:https://idp.test"); expect(storage.clearEnterpriseManagedResourceServers).toHaveBeenCalled(); + // No session and no cached metadata: nothing to build a URL from. + expect(result).toEqual({}); }); it("clearEmaIdpSession no-ops when issuer normalizes to empty", async () => { @@ -90,4 +110,84 @@ describe("idpSession", () => { storage.clearEnterpriseManagedResourceServers, ).not.toHaveBeenCalled(); }); + + describe("clearEmaIdpSession end-session URL", () => { + it("returns the end-session URL with id_token_hint from cached metadata", async () => { + vi.mocked(storage.getIdpSession).mockResolvedValue({ + idToken: "a.b.c", + }); + vi.mocked(storage.getServerMetadata).mockResolvedValue( + idpMetadata({ end_session_endpoint: "https://idp.test/session/end" }), + ); + const result = await clearEmaIdpSession(storage, "https://idp.test"); + expect(result.endSessionUrl).toBe( + "https://idp.test/session/end?id_token_hint=a.b.c", + ); + // The reads hit the leg-1 cache key, and the clear still happened in full. + expect(storage.getServerMetadata).toHaveBeenCalledWith( + "ema-idp:https://idp.test", + ); + expect(storage.clearIdpSession).toHaveBeenCalledWith("https://idp.test"); + expect(storage.clearEnterpriseManagedResourceServers).toHaveBeenCalled(); + }); + + it("returns no URL when there is no IdP session", async () => { + vi.mocked(storage.getServerMetadata).mockResolvedValue( + idpMetadata({ end_session_endpoint: "https://idp.test/session/end" }), + ); + const result = await clearEmaIdpSession(storage, "https://idp.test"); + expect(result).toEqual({}); + expect(storage.clearIdpSession).toHaveBeenCalled(); + }); + + it("returns no URL when no metadata is cached (clear still happens)", async () => { + vi.mocked(storage.getIdpSession).mockResolvedValue({ idToken: "a.b.c" }); + const result = await clearEmaIdpSession(storage, "https://idp.test"); + expect(result).toEqual({}); + expect(storage.clearIdpSession).toHaveBeenCalled(); + }); + + it.each([ + ["absent", {}], + ["not a string", { end_session_endpoint: 42 }], + ["not a URL", { end_session_endpoint: "not a url" }], + ["non-http scheme", { end_session_endpoint: "javascript:alert(1)" }], + // The URL carries the ID token, so plain http is rejected for any + // non-loopback host — never offer a cleartext logout URL. + [ + "plain http on a non-loopback host", + { end_session_endpoint: "http://idp.test/session/end" }, + ], + ])( + "returns no URL when end_session_endpoint is %s", + async (_label, metadata) => { + vi.mocked(storage.getIdpSession).mockResolvedValue({ + idToken: "a.b.c", + }); + vi.mocked(storage.getServerMetadata).mockResolvedValue( + idpMetadata(metadata), + ); + const result = await clearEmaIdpSession(storage, "https://idp.test"); + expect(result).toEqual({}); + }, + ); + + it.each(["localhost", "127.0.0.1", "[::1]"])( + "allows plain http when the host is loopback (%s)", + async (host) => { + vi.mocked(storage.getIdpSession).mockResolvedValue({ + idToken: "a.b.c", + }); + vi.mocked(storage.getServerMetadata).mockResolvedValue( + idpMetadata({ + end_session_endpoint: `http://${host}:8800/session/end`, + }), + ); + const result = await clearEmaIdpSession(storage, "https://idp.test"); + expect(result.endSessionUrl).toBe( + `http://${host}:8800/session/end?id_token_hint=a.b.c`, + ); + }, + ); + }); }); diff --git a/clients/web/src/test/core/auth/ema/storage.test.ts b/clients/web/src/test/core/auth/ema/storage.test.ts new file mode 100644 index 0000000000..cceb0f047c --- /dev/null +++ b/clients/web/src/test/core/auth/ema/storage.test.ts @@ -0,0 +1,39 @@ +import { describe, it, expect } from "vitest"; +import { + IDP_OAUTH_KEY_PREFIX, + idpOAuthStorageKey, + normalizeIdpIssuer, + parseIdpOAuthStorageKey, +} from "@inspector/core/auth/ema/storage.js"; + +describe("normalizeIdpIssuer", () => { + it("strips a single trailing slash", () => { + expect(normalizeIdpIssuer("https://idp.example.com/")).toBe( + "https://idp.example.com", + ); + }); + + it("leaves an issuer without a trailing slash unchanged", () => { + expect(normalizeIdpIssuer("https://idp.example.com")).toBe( + "https://idp.example.com", + ); + }); +}); + +describe("idpOAuthStorageKey / parseIdpOAuthStorageKey", () => { + it("builds a prefixed key and round-trips back to the issuer", () => { + const key = idpOAuthStorageKey("https://idp.example.com"); + expect(key).toBe(`${IDP_OAUTH_KEY_PREFIX}https://idp.example.com`); + expect(parseIdpOAuthStorageKey(key)).toBe("https://idp.example.com"); + }); + + it("normalises a trailing slash before prefixing", () => { + expect(idpOAuthStorageKey("https://idp.example.com/")).toBe( + "ema-idp:https://idp.example.com", + ); + }); + + it("returns null for a plain server URL (not an IdP key)", () => { + expect(parseIdpOAuthStorageKey("https://mcp.example.com/mcp")).toBeNull(); + }); +}); diff --git a/clients/web/src/test/core/auth/oauth-namespace-ledger.test.ts b/clients/web/src/test/core/auth/oauth-namespace-ledger.test.ts new file mode 100644 index 0000000000..ada92ca161 --- /dev/null +++ b/clients/web/src/test/core/auth/oauth-namespace-ledger.test.ts @@ -0,0 +1,474 @@ +/** + * Unit tests for the secrets-namespace ledger (#2560): its tolerant parse, + * record-only-when-new writes, per-namespace purge bookkeeping, and the + * best-effort failure handling that keeps a ledger problem from ever + * failing the OAuth save or removal it rides on. The end-to-end scenarios + * (a stripped stamp re-adopted, a stripped file removed) live in the + * integration suite's `oauth-secrets-namespace.test.ts`. + */ + +import { describe, it, expect, beforeEach, afterEach, vi } from "vitest"; +import { + existsSync, + mkdirSync, + mkdtempSync, + readFileSync, + rmSync, + writeFileSync, +} from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +// Passthrough, so one test can fail the ledger's write-back on demand. +vi.mock("@inspector/core/storage/store-io.js", async (importOriginal) => { + const actual = + await importOriginal< + typeof import("@inspector/core/storage/store-io.js") + >(); + return { ...actual, writeStoreFile: vi.fn(actual.writeStoreFile) }; +}); + +import { writeStoreFile } from "@inspector/core/storage/store-io.js"; +import { + namespaceLedgerPath, + purgeSupersededNamespaces, + recordNamespaceKeys, + resetNamespaceLedgerWarnings, + secretStoreLocation, +} from "@inspector/core/auth/node/oauth-namespace-ledger.js"; +import { + InMemorySecretStore, + KeyringSecretStore, + SessionSecretStore, +} from "@inspector/core/auth/node/secret-store.js"; +import { FileSecretStore } from "@inspector/core/auth/node/file-secret-store.js"; +import { + defaultSecretStore, + SECRET_FILE_ENV, + SECRET_STORE_ENV, +} from "@inspector/core/auth/node/secret-store-selection.js"; + +/** + * The keychain, as far as `instanceof` is concerned, backed by a map — the + * native module is absent in CI. + */ +class FakeKeyring extends KeyringSecretStore { + readonly deleted: string[] = []; + async deleteAllForServer(serverId: string): Promise<void> { + this.deleted.push(serverId); + } +} +import { + LEGACY_TOKENS_FIELD, + IDP_SESSION_FIELD, + oauthIdpSecretServerId, + oauthSecretServerId, +} from "@inspector/core/auth/node/oauth-secrets.js"; + +const NS1 = "11111111-1111-4111-8111-111111111111"; +const NS2 = "22222222-2222-4222-8222-222222222222"; +const SERVER = "https://api.example/mcp"; +const ISSUER = "https://as.example"; + +let tempDir: string; +let stateFile: string; +let ledgerFile: string; +let store: InMemorySecretStore; + +beforeEach(() => { + tempDir = mkdtempSync(join(tmpdir(), "inspector-ns-ledger-")); + stateFile = join(tempDir, "oauth.json"); + ledgerFile = namespaceLedgerPath(stateFile); + store = new InMemorySecretStore(); +}); + +afterEach(() => { + resetNamespaceLedgerWarnings(); + vi.restoreAllMocks(); + rmSync(tempDir, { recursive: true, force: true }); +}); + +type LedgerJson = Record< + string, + Record<string, { servers: string[]; idpSessions: string[] }> +>; + +function readLedger(): LedgerJson { + return ( + JSON.parse(readFileSync(ledgerFile, "utf8")) as { namespaces: LedgerJson } + ).namespaces; +} + +describe("namespaceLedgerPath", () => { + it("sits beside the state file", () => { + expect(namespaceLedgerPath("/x/oauth.json")).toBe( + "/x/oauth.json.namespaces.json", + ); + }); +}); + +describe("recordNamespaceKeys", () => { + it("creates the ledger and accumulates keys per namespace", async () => { + await recordNamespaceKeys(stateFile, store, NS1, [SERVER], []); + await recordNamespaceKeys( + stateFile, + store, + NS1, + ["https://b.example"], + [ISSUER], + ); + await recordNamespaceKeys(stateFile, store, NS2, [SERVER], []); + expect(readLedger()).toEqual({ + [NS1]: { + memory: { + servers: [SERVER, "https://b.example"], + idpSessions: [ISSUER], + }, + }, + [NS2]: { memory: { servers: [SERVER], idpSessions: [] } }, + }); + }); + + it("does not rewrite the ledger when every key is already recorded", async () => { + await recordNamespaceKeys(stateFile, store, NS1, [SERVER], [ISSUER]); + const before = readFileSync(ledgerFile, "utf8"); + // A sentinel the rewrite would replace: one space after the opening brace. + writeFileSync(ledgerFile, before.replace(/^\{/, "{ ")); + const sentinel = readFileSync(ledgerFile, "utf8"); + expect(sentinel).not.toBe(before); + await recordNamespaceKeys(stateFile, store, NS1, [SERVER], [ISSUER]); + expect(readFileSync(ledgerFile, "utf8")).toBe(sentinel); + }); + + it("warns once per reason instead of throwing when the ledger is unusable", async () => { + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + // A directory where the ledger should be: every read fails (EISDIR). + mkdirSync(ledgerFile); + await expect( + recordNamespaceKeys(stateFile, store, NS1, [SERVER], []), + ).resolves.toBeUndefined(); + await recordNamespaceKeys(stateFile, store, NS1, [SERVER], []); + expect(warn).toHaveBeenCalledTimes(1); + expect(warn.mock.calls[0][0]).toContain( + "Could not update the OAuth secrets-namespace ledger", + ); + }); + + it("warns with a non-Error rejection's string form", async () => { + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + // Purge's per-id failure path receives whatever the store threw. + await recordNamespaceKeys(stateFile, store, NS1, [SERVER], []); + vi.spyOn(store, "deleteAllForServer").mockRejectedValue("plain string"); + await purgeSupersededNamespaces(stateFile, store, undefined); + expect(warn.mock.calls[0][0]).toContain("(plain string)"); + }); +}); + +describe("tolerant parse", () => { + async function purgeAllFrom(raw: string): Promise<void> { + writeFileSync(ledgerFile, raw); + await purgeSupersededNamespaces(stateFile, store, undefined); + } + + it.each([ + ["unparseable JSON", "{not json"], + ["a non-object", "42"], + ["no namespaces map", JSON.stringify({ other: 1 })], + ["a non-object namespaces value", JSON.stringify({ namespaces: "x" })], + ])("reads %s as an empty ledger", async (_label, raw) => { + const del = vi.spyOn(store, "deleteAllForServer"); + await purgeAllFrom(raw); + expect(del).not.toHaveBeenCalled(); + // Nothing changed, so the junk is left as found rather than rewritten. + expect(readFileSync(ledgerFile, "utf8")).toBe(raw); + }); + + it("skips an invalid namespace and non-string keys, never building their ids", async () => { + const del = vi.spyOn(store, "deleteAllForServer"); + await purgeAllFrom( + JSON.stringify({ + namespaces: { + "bad+ns": { memory: { servers: [SERVER] } }, + [NS1]: { memory: { servers: [SERVER, 7], idpSessions: "nope" } }, + [NS2]: "not-an-object", + "33333333-3333-4333-8333-333333333333": { memory: ["array"] }, + }, + }), + ); + expect(del.mock.calls.map(([id]) => id)).toEqual([ + oauthSecretServerId(SERVER, NS1), + ]); + // Both valid namespaces purged cleanly, so the ledger is gone. + expect(existsSync(ledgerFile)).toBe(false); + }); + + it("keeps a `__proto__` location an own key through a rewrite", async () => { + writeFileSync( + ledgerFile, + `{"namespaces":{"${NS1}":{"__proto__":{"servers":["${SERVER}"]}},"${NS2}":{"memory":{"servers":["${SERVER}"]}}}}`, + ); + await purgeSupersededNamespaces(stateFile, store, NS1); + // NS2 purged and rewritten out; NS1's odd location survives as data. + expect( + Object.getOwnPropertyNames(readLedger()[NS1] ?? {}).includes("__proto__"), + ).toBe(true); + }); +}); + +describe("secretStoreLocation", () => { + it("names the keychain, a specific secrets file, or memory", async () => { + expect(await secretStoreLocation(new KeyringSecretStore())).toBe("keyring"); + expect( + await secretStoreLocation( + new FileSecretStore({ filePath: "rel/secrets.json", passphrase: "" }), + ), + ).toBe(`file:${join(process.cwd(), "rel/secrets.json")}`); + expect(await secretStoreLocation(new InMemorySecretStore())).toBe("memory"); + expect(await secretStoreLocation(new SessionSecretStore())).toBe("memory"); + }); + + it("names the store a production default resolves to, not the wrapper", async () => { + // `defaultSecretStore()` is a deferred wrapper; read as-is it would + // name nothing and every production store would read as `memory`. + const secretsFile = join(tempDir, "secrets.json"); + const saved = { + kind: process.env[SECRET_STORE_ENV], + file: process.env[SECRET_FILE_ENV], + }; + process.env[SECRET_STORE_ENV] = "file"; + process.env[SECRET_FILE_ENV] = secretsFile; + vi.spyOn(console, "warn").mockImplementation(() => {}); + try { + expect(await secretStoreLocation(defaultSecretStore())).toBe( + `file:${secretsFile}`, + ); + } finally { + for (const [name, value] of [ + [SECRET_STORE_ENV, saved.kind], + [SECRET_FILE_ENV, saved.file], + ] as const) { + if (value === undefined) delete process.env[name]; + else process.env[name] = value; + } + } + }); +}); + +describe("purgeSupersededNamespaces", () => { + async function seed(namespace: string): Promise<void> { + await recordNamespaceKeys(stateFile, store, namespace, [SERVER], [ISSUER]); + await store.set( + oauthSecretServerId(SERVER, namespace), + LEGACY_TOKENS_FIELD, + "t", + ); + await store.set( + oauthIdpSecretServerId(ISSUER, namespace), + IDP_SESSION_FIELD, + "s", + ); + } + + it("purges every namespace but the current one, server and IdP ids alike", async () => { + await seed(NS1); + await seed(NS2); + await purgeSupersededNamespaces(stateFile, store, NS2); + expect( + await store.get(oauthSecretServerId(SERVER, NS1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect( + await store.get(oauthIdpSecretServerId(ISSUER, NS1), IDP_SESSION_FIELD), + ).toBeNull(); + expect( + await store.get(oauthSecretServerId(SERVER, NS2), LEGACY_TOKENS_FIELD), + ).toBe("t"); + expect(Object.keys(readLedger())).toEqual([NS2]); + }); + + it("deletes the ledger once nothing is left recorded", async () => { + await seed(NS1); + await purgeSupersededNamespaces(stateFile, store, undefined); + expect(existsSync(ledgerFile)).toBe(false); + }); + + it("is a no-op without a ledger", async () => { + const del = vi.spyOn(store, "deleteAllForServer"); + await purgeSupersededNamespaces(stateFile, store, undefined); + expect(del).not.toHaveBeenCalled(); + expect(existsSync(ledgerFile)).toBe(false); + }); + + it("keeps a namespace recorded when one of its ids fails, still attempting the rest", async () => { + await seed(NS1); + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + const failing = oauthSecretServerId(SERVER, NS1); + const realDelete = store.deleteAllForServer.bind(store); + vi.spyOn(store, "deleteAllForServer").mockImplementation(async (id) => { + if (id === failing) throw new Error("keychain locked"); + return realDelete(id); + }); + await purgeSupersededNamespaces(stateFile, store, undefined); + expect( + await store.get(oauthIdpSecretServerId(ISSUER, NS1), IDP_SESSION_FIELD), + ).toBeNull(); + expect(Object.keys(readLedger())).toEqual([NS1]); + expect(warn.mock.calls[0][0]).toContain("keychain locked"); + }); + + it("leaves another location's record for a run that uses that store", async () => { + // Entries written to the keychain; this run fell back to another + // store. Deleting there "succeeds" against nothing, so the keychain + // record must survive for a later keychain run. + const keyring = new KeyringSecretStore(); + await recordNamespaceKeys(stateFile, keyring, NS1, [SERVER], []); + const del = vi.spyOn(store, "deleteAllForServer"); + await purgeSupersededNamespaces(stateFile, store, undefined); + expect(del).not.toHaveBeenCalled(); + expect(readLedger()[NS1]).toEqual({ + keyring: { servers: [SERVER], idpSessions: [] }, + }); + }); + + it("from the keychain, also purges a secrets file's records but keeps them", async () => { + // The hand-off may have copied them in — or may still be copying, or + // was interrupted — so the keychain run purges and keeps the record + // for the next save rather than guessing that it is finished. + const keyring = new FakeKeyring(); + const fileStore = new FileSecretStore({ + filePath: join(tempDir, "secrets.json"), + passphrase: "", + }); + await recordNamespaceKeys(stateFile, fileStore, NS1, [SERVER], []); + await recordNamespaceKeys(stateFile, keyring, NS2, [SERVER], []); + + await purgeSupersededNamespaces(stateFile, keyring, undefined); + expect(keyring.deleted.sort()).toEqual( + [ + oauthSecretServerId(SERVER, NS1), + oauthSecretServerId(SERVER, NS2), + ].sort(), + ); + // The keychain's own record is done; the file's record stays. + expect(readLedger()).toEqual({ + [NS1]: { + [`file:${fileStore.filePath}`]: { servers: [SERVER], idpSessions: [] }, + }, + }); + }); + + it("a secrets file's own purge hands its record to the keychain instead of forgetting it", async () => { + // Its entries may already have been copied into the keychain by a + // hand-off, so the keys stay tracked there for a keychain run. + const fileStore = new FileSecretStore({ + filePath: join(tempDir, "secrets.json"), + passphrase: "", + }); + const fileDel = vi + .spyOn(fileStore, "deleteAllForServer") + .mockResolvedValue(undefined); + await recordNamespaceKeys(stateFile, fileStore, NS1, [SERVER], [ISSUER]); + await recordNamespaceKeys( + stateFile, + new FakeKeyring(), + NS1, + ["https://b.example"], + [], + ); + + await purgeSupersededNamespaces(stateFile, fileStore, undefined); + expect(fileDel).toHaveBeenCalledTimes(2); + expect(readLedger()).toEqual({ + [NS1]: { + keyring: { + servers: ["https://b.example", SERVER], + idpSessions: [ISSUER], + }, + }, + }); + + const keyring = new FakeKeyring(); + await purgeSupersededNamespaces(stateFile, keyring, undefined); + expect(existsSync(ledgerFile)).toBe(false); + }); + + it("does not extend a secrets file's or memory run to another file's records", async () => { + const other = new FileSecretStore({ + filePath: join(tempDir, "other.json"), + passphrase: "", + }); + await recordNamespaceKeys(stateFile, other, NS1, [SERVER], []); + const del = vi.spyOn(store, "deleteAllForServer"); + const mine = new FileSecretStore({ + filePath: join(tempDir, "mine.json"), + passphrase: "", + }); + const mineDel = vi.spyOn(mine, "deleteAllForServer"); + await purgeSupersededNamespaces(stateFile, store, undefined); + await purgeSupersededNamespaces(stateFile, mine, undefined); + expect(del).not.toHaveBeenCalled(); + expect(mineDel).not.toHaveBeenCalled(); + expect(Object.keys(readLedger())).toEqual([NS1]); + }); + + it("drops only the purged location's record, keeping the namespace for the rest", async () => { + await seed(NS1); + await recordNamespaceKeys( + stateFile, + new KeyringSecretStore(), + NS1, + [SERVER], + [], + ); + await purgeSupersededNamespaces(stateFile, store, undefined); + expect(Object.keys(readLedger()[NS1])).toEqual(["keyring"]); + }); + + it("survives a key whose id cannot be built, purging the rest and keeping the record", async () => { + // JSON-escaped unpaired surrogate: parses to a string that makes + // `encodeURIComponent` throw URIError. + writeFileSync( + ledgerFile, + `{"namespaces":{"${NS1}":{"memory":{"servers":["\\ud800","${SERVER}"]}}}}`, + ); + await store.set(oauthSecretServerId(SERVER, NS1), LEGACY_TOKENS_FIELD, "t"); + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + await expect( + purgeSupersededNamespaces(stateFile, store, undefined), + ).resolves.toBeUndefined(); + expect( + await store.get(oauthSecretServerId(SERVER, NS1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect(Object.keys(readLedger())).toEqual([NS1]); + expect(warn).toHaveBeenCalledTimes(1); + }); + + it("warns and gives up when the ledger cannot be read", async () => { + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + mkdirSync(ledgerFile); + const del = vi.spyOn(store, "deleteAllForServer"); + await expect( + purgeSupersededNamespaces(stateFile, store, undefined), + ).resolves.toBeUndefined(); + expect(del).not.toHaveBeenCalled(); + expect(warn.mock.calls[0][0]).toContain( + "Could not read the OAuth secrets-namespace ledger", + ); + }); + + it("warns when the purged ledger cannot be written back", async () => { + await seed(NS1); + await seed(NS2); + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + vi.mocked(writeStoreFile).mockRejectedValueOnce(new Error("disk full")); + await expect( + purgeSupersededNamespaces(stateFile, store, NS2), + ).resolves.toBeUndefined(); + // The purge itself happened; only the bookkeeping is stale. + expect( + await store.get(oauthSecretServerId(SERVER, NS1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect(warn.mock.calls.at(-1)?.[0]).toContain( + "Could not update the OAuth secrets-namespace ledger", + ); + }); +}); diff --git a/clients/web/src/test/core/auth/oauth-persist-file.test.ts b/clients/web/src/test/core/auth/oauth-persist-file.test.ts index eab3ea482d..2f160335e8 100644 --- a/clients/web/src/test/core/auth/oauth-persist-file.test.ts +++ b/clients/web/src/test/core/auth/oauth-persist-file.test.ts @@ -85,7 +85,7 @@ describe("writeOAuthSections lock failures", () => { // the wrong file. Only acquisition failures (the mocks above, which // reject before the callback runs) get the OAuth wording. vi.mocked(withSecretFileLock).mockImplementation( - async (_path, fn) => fn() as Promise<never>, + async (_path, fn) => fn(true) as Promise<never>, ); const original = new SecretFileLockHeldError( "Could not lock the secrets file at /home/u/.mcp-inspector/secrets.json", @@ -142,7 +142,7 @@ describe("readOAuthStore locking", () => { // residue with its already-committed new secrets. The whole read must // execute inside the same lock the writers hold. vi.mocked(withSecretFileLock).mockImplementation( - async (_path, fn) => fn() as Promise<never>, + async (_path, fn) => fn(true) as Promise<never>, ); const result = await readOAuthStore( @@ -190,7 +190,7 @@ describe("persistEntrySecrets partial-commit compensation", () => { beforeEach(() => { vi.mocked(withSecretFileLock).mockReset(); vi.mocked(withSecretFileLock).mockImplementation( - async (_path, fn) => fn() as Promise<never>, + async (_path, fn) => fn(true) as Promise<never>, ); }); @@ -218,13 +218,18 @@ describe("persistEntrySecrets partial-commit compensation", () => { failWhen: (field: string, value: string) => boolean, ) => { const store = new InMemorySecretStore(); - const serverId = oauthSecretServerId(url); await writeOAuthSections( file, { servers: { [url]: SEED_STATE }, idpSessions: {} }, { servers: [url] }, store, ); + // The seed write adopted a secrets namespace (#2549); the entry's store + // id is scoped by it, so read it back from the written file. + const { secretsNamespace } = JSON.parse(await readFile(file, "utf8")) as { + secretsNamespace: string; + }; + const serverId = oauthSecretServerId(url, secretsNamespace); const realSet = store.set.bind(store); store.set = async (sid: string, field: string, value: string) => { if (failWhen(field, value)) diff --git a/clients/web/src/test/core/auth/oauth-secrets.test.ts b/clients/web/src/test/core/auth/oauth-secrets.test.ts index b47ae2d89f..cdbb1539e1 100644 --- a/clients/web/src/test/core/auth/oauth-secrets.test.ts +++ b/clients/web/src/test/core/auth/oauth-secrets.test.ts @@ -10,6 +10,8 @@ import { resetPersistTokensPolicyWarnings, oauthSecretServerId, oauthIdpSecretServerId, + isValidSecretsNamespace, + newSecretsNamespace, issuerTokensField, issuerClientSecretField, issuerRegistrationTokenField, @@ -103,6 +105,43 @@ describe("id and field schemes", () => { const withPort = `${oauthSecretServerId("https://a.example:8080")}:tokens`; expect(withPort.startsWith(plain)).toBe(false); }); + + it("scopes ids by secrets namespace with an unforgeable delimiter", () => { + const ns = "9a3c2e1f-0b4d-4c5e-8f6a-7b8c9d0e1f2a"; + expect(oauthSecretServerId("https://s.example/mcp", ns)).toBe( + `oauth+${ns}+https%3A%2F%2Fs.example%2Fmcp`, + ); + expect(oauthIdpSecretServerId("https://idp.example", ns)).toBe( + `oauth-idp+${ns}+https%3A%2F%2Fidp.example`, + ); + // Namespaced ids stay colon-free (same keyring purge constraint). + expect(oauthSecretServerId("https://s.example:8443/mcp", ns)).not.toContain( + ":", + ); + // encodeURIComponent escapes `+`, so a URL cannot forge the delimiter: + // a legacy id over a `+`-bearing URL never collides with a namespaced id. + expect(oauthSecretServerId(`${ns}+https://s.example/mcp`)).not.toBe( + oauthSecretServerId("https://s.example/mcp", ns), + ); + }); + + it("validates and mints namespaces", () => { + expect(isValidSecretsNamespace(newSecretsNamespace())).toBe(true); + expect(isValidSecretsNamespace("abc-123.DEF_456")).toBe(true); + for (const bad of [ + undefined, + null, + 42, + "", + "-leading-separator", + "has:colon", + "has+plus", + "has space", + "a".repeat(129), + ]) { + expect(isValidSecretsNamespace(bad)).toBe(false); + } + }); }); describe("splitServerOAuthState", () => { diff --git a/clients/web/src/test/core/auth/runner-interactive-oauth.test.ts b/clients/web/src/test/core/auth/runner-interactive-oauth.test.ts index b141c1cba8..e885e06d26 100644 --- a/clients/web/src/test/core/auth/runner-interactive-oauth.test.ts +++ b/clients/web/src/test/core/auth/runner-interactive-oauth.test.ts @@ -522,4 +522,222 @@ describe("runRunnerInteractiveOAuth", () => { ).rejects.toThrow("bind failed"); expect(mockServer.stop).toHaveBeenCalled(); }); + + it("rejects cleanly on SIGINT while waiting on the callback, instead of hanging or killing the process", async () => { + const redirectUrlProvider = { redirectUrl: "" }; + const mockServer = createMockCallbackServer(handlers); + const client = mockClient({ + authenticate: vi.fn(async () => new URL("https://as.example/authorize")), + }); + + const promise = runRunnerInteractiveOAuth({ + client, + redirectUrlProvider, + callbackListen: { + hostname: "127.0.0.1", + port: 6276, + pathname: "/oauth/callback", + }, + createCallbackServer: () => mockServer, + handleSignals: true, + }); + + // Give beginInteractiveAuthorization/authenticate a tick to register the + // listener before the signal fires. + await Promise.resolve(); + await Promise.resolve(); + process.emit("SIGINT", "SIGINT"); + + await expect(promise).rejects.toThrow( + "OAuth authorization cancelled (SIGINT).", + ); + expect(mockServer.stop).toHaveBeenCalled(); + // The handler must be removed once the wait settles, so a later SIGINT + // elsewhere in the process isn't accidentally swallowed by a stale + // listener from this call. + expect(process.listenerCount("SIGINT")).toBe(0); + }); + + it("rejects cleanly on SIGTERM the same way", async () => { + const redirectUrlProvider = { redirectUrl: "" }; + const mockServer = createMockCallbackServer(handlers); + const client = mockClient({ + authenticate: vi.fn(async () => new URL("https://as.example/authorize")), + }); + + const promise = runRunnerInteractiveOAuth({ + client, + redirectUrlProvider, + callbackListen: { + hostname: "127.0.0.1", + port: 6276, + pathname: "/oauth/callback", + }, + createCallbackServer: () => mockServer, + handleSignals: true, + }); + + await Promise.resolve(); + await Promise.resolve(); + process.emit("SIGTERM", "SIGTERM"); + + await expect(promise).rejects.toThrow( + "OAuth authorization cancelled (SIGTERM).", + ); + expect(process.listenerCount("SIGTERM")).toBe(0); + }); + + it("handles a signal during server startup without an unhandled rejection", async () => { + const redirectUrlProvider = { redirectUrl: "" }; + let releaseStart!: () => void; + const startGate = new Promise<void>((resolve) => (releaseStart = resolve)); + const mockServer = { + start: vi.fn(async (opts: OAuthCallbackServerStartOptions) => { + handlers.current = { + onCallback: opts.onCallback, + onError: opts.onError, + }; + await startGate; + return { + port: 6276, + redirectUrl: "http://127.0.0.1:6276/oauth/callback", + }; + }), + stop: vi.fn(async () => {}), + } as unknown as OAuthCallbackServer; + const client = mockClient({ + authenticate: vi.fn(async () => new URL("https://as.example/authorize")), + }); + + const unhandled: unknown[] = []; + const onUnhandled = (reason: unknown) => unhandled.push(reason); + process.on("unhandledRejection", onUnhandled); + try { + const promise = runRunnerInteractiveOAuth({ + client, + redirectUrlProvider, + callbackListen: { + hostname: "127.0.0.1", + port: 6276, + pathname: "/oauth/callback", + }, + createCallbackServer: () => mockServer, + handleSignals: true, + }); + + // The signal listeners are installed before `server.start()` is + // awaited, so a signal in that window rejects flowDone before the + // Promise.race ever subscribes to it. Cancellation now also aborts the + // stalled start() itself, so the runner promise rejects immediately — + // subscribe before emitting the signal. + const expectation = expect(promise).rejects.toThrow( + "OAuth authorization cancelled (SIGINT).", + ); + await Promise.resolve(); + process.emit("SIGINT", "SIGINT"); + // A full macrotask turn: Node reports any unhandled rejection here. + await new Promise((resolve) => setImmediate(resolve)); + expect(unhandled).toEqual([]); + + releaseStart(); + await expectation; + expect(mockServer.stop).toHaveBeenCalled(); + expect(process.listenerCount("SIGINT")).toBe(0); + } finally { + process.off("unhandledRejection", onUnhandled); + } + }); + + it("cancels a stalled server.start() on SIGINT instead of blocking", async () => { + const redirectUrlProvider = { redirectUrl: "" }; + const mockServer = { + // Never resolves: a callback server stalled on listen(). + start: vi.fn(() => new Promise<never>(() => {})), + stop: vi.fn(async () => {}), + } as unknown as OAuthCallbackServer; + const client = mockClient(); + + const promise = runRunnerInteractiveOAuth({ + client, + redirectUrlProvider, + callbackListen: { + hostname: "127.0.0.1", + port: 6276, + pathname: "/oauth/callback", + }, + createCallbackServer: () => mockServer, + handleSignals: true, + }); + await Promise.resolve(); + process.emit("SIGINT", "SIGINT"); + await expect(promise).rejects.toThrow( + "OAuth authorization cancelled (SIGINT).", + ); + expect(mockServer.stop).toHaveBeenCalled(); + expect(process.listenerCount("SIGINT")).toBe(0); + }); + + it("cancels a stalled authenticate() on SIGTERM, never reporting already_authorized", async () => { + const redirectUrlProvider = { redirectUrl: "" }; + const mockServer = createMockCallbackServer(handlers); + let releaseAuthenticate!: () => void; + const authGate = new Promise<void>( + (resolve) => (releaseAuthenticate = resolve), + ); + const client = mockClient({ + // Resolves undefined — but only after the signal has already fired; + // the cancellation must win, not be discarded as already_authorized. + authenticate: vi.fn(async () => { + await authGate; + return undefined; + }), + }); + + const promise = runRunnerInteractiveOAuth({ + client, + redirectUrlProvider, + callbackListen: { + hostname: "127.0.0.1", + port: 6276, + pathname: "/oauth/callback", + }, + createCallbackServer: () => mockServer, + handleSignals: true, + }); + await vi.waitFor(() => expect(client.authenticate).toHaveBeenCalled()); + process.emit("SIGTERM", "SIGTERM"); + releaseAuthenticate(); + await expect(promise).rejects.toThrow( + "OAuth authorization cancelled (SIGTERM).", + ); + expect(mockServer.stop).toHaveBeenCalled(); + expect(process.listenerCount("SIGTERM")).toBe(0); + }); + + it("installs no signal listeners unless handleSignals is set (TUI owns Ctrl-C via Ink)", async () => { + const redirectUrlProvider = { redirectUrl: "" }; + const mockServer = createMockCallbackServer(handlers); + const client = mockClient({ + authenticate: vi.fn(async () => new URL("https://as.example/authorize")), + }); + const before = process.listenerCount("SIGINT"); + + const promise = runRunnerInteractiveOAuth({ + client, + redirectUrlProvider, + callbackListen: { + hostname: "127.0.0.1", + port: 6276, + pathname: "/oauth/callback", + }, + createCallbackServer: () => mockServer, + }); + + await Promise.resolve(); + await Promise.resolve(); + expect(process.listenerCount("SIGINT")).toBe(before); + + await simulateCallback(handlers.current); + await expect(promise).resolves.toEqual({ kind: "success" }); + }); }); diff --git a/clients/web/src/test/core/auth/storage-browser.test.ts b/clients/web/src/test/core/auth/storage-browser.test.ts index 82074ab1a2..0c083e596c 100644 --- a/clients/web/src/test/core/auth/storage-browser.test.ts +++ b/clients/web/src/test/core/auth/storage-browser.test.ts @@ -177,7 +177,7 @@ describe("BrowserOAuthStorage", () => { await storage.saveTokens(testServerUrl, { refresh_token: "rt-only", token_type: "Bearer", - } as unknown as OAuthTokens); + } as OAuthTokens); await expect(storage.getTokens(testServerUrl)).resolves.toBeUndefined(); }); }); diff --git a/clients/web/src/test/core/extension/tasks/errors.test.ts b/clients/web/src/test/core/extension/tasks/errors.test.ts new file mode 100644 index 0000000000..fa2e484182 --- /dev/null +++ b/clients/web/src/test/core/extension/tasks/errors.test.ts @@ -0,0 +1,73 @@ +import { describe, expect, it } from "vitest"; +import { ProtocolError, ProtocolErrorCode } from "@modelcontextprotocol/client"; +import { + DispatchError, + JsonRpcResponseError, + TaskFailedError, +} from "@modelcontextprotocol/ext-tasks/client"; +import { + abortError, + toProtocolError, + unwrapTaskDispatchError, +} from "@inspector/core/extension/tasks/errors.js"; + +describe("abortError", () => { + it("returns the signal's own Error reason", () => { + const reason = new Error("stop"); + const controller = new AbortController(); + controller.abort(reason); + expect(abortError(controller.signal)).toBe(reason); + }); + + it("builds an AbortError when the reason is not an Error", () => { + const controller = new AbortController(); + controller.abort("user cancelled"); + const error = abortError(controller.signal); + expect(error.name).toBe("AbortError"); + }); +}); + +describe("unwrapTaskDispatchError", () => { + it("restores a DispatchError's Error cause", () => { + const cause = new Error("socket closed"); + expect( + unwrapTaskDispatchError(new DispatchError("wrapped", false, { cause })), + ).toBe(cause); + }); + + it("leaves a DispatchError without an Error cause as is", () => { + const error = new DispatchError("bare"); + expect(unwrapTaskDispatchError(error)).toBe(error); + }); + + it("maps a JSON-RPC response error to a ProtocolError", () => { + const unwrapped = unwrapTaskDispatchError( + new JsonRpcResponseError({ code: -32602, message: "bad", data: 1 }), + ); + expect(unwrapped).toBeInstanceOf(ProtocolError); + expect(unwrapped).toMatchObject({ code: -32602, data: 1 }); + }); + + it("maps a coded task failure to a ProtocolError, uncoded passes through", () => { + expect( + unwrapTaskDispatchError(new TaskFailedError("boom", { code: -32000 })), + ).toMatchObject({ code: -32000 }); + const uncoded = new TaskFailedError("boom"); + expect(unwrapTaskDispatchError(uncoded)).toBe(uncoded); + }); +}); + +describe("toProtocolError", () => { + it("keeps a ProtocolError and wraps anything else as InternalError", () => { + const protocol = new ProtocolError(-32601, "missing"); + expect(toProtocolError(protocol)).toBe(protocol); + expect(toProtocolError(new Error("plain"))).toMatchObject({ + code: ProtocolErrorCode.InternalError, + message: expect.stringContaining("plain"), + }); + expect(toProtocolError("text")).toMatchObject({ + code: ProtocolErrorCode.InternalError, + message: expect.stringContaining("text"), + }); + }); +}); diff --git a/clients/web/src/test/core/taskNotificationSchemas.test.ts b/clients/web/src/test/core/extension/tasks/notificationSchemas.test.ts similarity index 96% rename from clients/web/src/test/core/taskNotificationSchemas.test.ts rename to clients/web/src/test/core/extension/tasks/notificationSchemas.test.ts index e8e7c1f756..d8cff75fb6 100644 --- a/clients/web/src/test/core/taskNotificationSchemas.test.ts +++ b/clients/web/src/test/core/extension/tasks/notificationSchemas.test.ts @@ -1,5 +1,5 @@ import { describe, it, expect } from "vitest"; -import { TasksListChangedNotificationSchema } from "@inspector/core/mcp/taskNotificationSchemas.js"; +import { TasksListChangedNotificationSchema } from "@inspector/core/extension/tasks/notificationSchemas.js"; describe("TasksListChangedNotificationSchema", () => { it("parses a notification with no params", () => { diff --git a/clients/web/src/test/core/extension/tasks/progress.test.ts b/clients/web/src/test/core/extension/tasks/progress.test.ts new file mode 100644 index 0000000000..e615f53f45 --- /dev/null +++ b/clients/web/src/test/core/extension/tasks/progress.test.ts @@ -0,0 +1,45 @@ +import { describe, expect, it } from "vitest"; +import { TaskProgressRouter } from "@inspector/core/extension/tasks/progress.js"; + +describe("TaskProgressRouter", () => { + it("routes nothing for a token no task call owns", () => { + expect(new TaskProgressRouter().route("t")).toBeUndefined(); + }); + + it("owns a token from acquire, before any task is correlated", () => { + const router = new TaskProgressRouter(); + router.acquire("t"); + expect(router.route("t")).toEqual([]); + }); + + it("delivers a shared token's progress to every correlated task", () => { + const router = new TaskProgressRouter(); + router.acquire("t"); + router.acquire("t"); + router.correlate("t", "a"); + router.correlate("t", "b"); + expect(router.route("t")).toEqual(["a", "b"]); + + // One owner settles: only its own task is released. + router.release("t", "a"); + expect(router.route("t")).toEqual(["b"]); + + router.release("t", "b"); + expect(router.route("t")).toBeUndefined(); + }); + + it("releases an owner that never saw a task", () => { + const router = new TaskProgressRouter(); + router.acquire("t"); + router.release("t"); + expect(router.route("t")).toBeUndefined(); + }); + + it("forgets everything on clear", () => { + const router = new TaskProgressRouter(); + router.acquire(1); + router.correlate(1, "a"); + router.clear(); + expect(router.route(1)).toBeUndefined(); + }); +}); diff --git a/clients/web/src/test/core/extension/tasks/rawWireChannel.test.ts b/clients/web/src/test/core/extension/tasks/rawWireChannel.test.ts new file mode 100644 index 0000000000..eb2300cd41 --- /dev/null +++ b/clients/web/src/test/core/extension/tasks/rawWireChannel.test.ts @@ -0,0 +1,181 @@ +import { afterEach, describe, expect, it, vi } from "vitest"; +import type { JSONRPCMessage, Transport } from "@modelcontextprotocol/client"; +import { DispatchError } from "@modelcontextprotocol/ext-tasks/client"; +import { + RAW_WIRE_ID_PREFIX, + RawWireChannel, + type RawWireChannelHost, +} from "@inspector/core/extension/tasks/rawWireChannel.js"; + +type Sent = { id?: string | number; method: string; params?: unknown }; + +function setup(overrides: Partial<RawWireChannelHost> = {}) { + const sent: Sent[] = []; + const transport = { + start: async () => {}, + close: async () => {}, + send: vi.fn(async (message: JSONRPCMessage) => { + // The fields these tests read; every frame the channel sends has them. + sent.push(message as Sent); + }), + } satisfies Transport; + const channel = new RawWireChannel({ + transport: () => transport, + defaultTimeoutMs: () => 1_000, + resetTimeoutOnProgress: () => true, + annotateTimeout: (error) => error, + ...overrides, + }); + return { channel, sent, transport }; +} + +afterEach(() => { + vi.useRealTimers(); +}); + +describe("RawWireChannel", () => { + it("rejects without a transport, retryably", async () => { + const { channel } = setup({ transport: () => null }); + await expect(channel.dispatch({ method: "tasks/get" })).rejects.toSatisfy( + (error) => error instanceof DispatchError && error.retryable, + ); + }); + + it("rejects malformed requests before sending", async () => { + const { channel, sent } = setup(); + await expect(channel.dispatch([])).rejects.toThrow(/JSON object/); + await expect(channel.dispatch({ method: 1 })).rejects.toThrow(/method/); + await expect( + channel.dispatch({ method: "x", params: [1] }), + ).rejects.toThrow(/params/); + expect(sent).toHaveLength(0); + }); + + it("resolves the matching response and leaves foreign ids to the SDK", async () => { + const { channel, sent } = setup(); + const promise = channel.dispatch({ method: "tasks/get", params: {} }); + await Promise.resolve(); + const id = String(sent[0]!.id); + expect(id.startsWith(RAW_WIRE_ID_PREFIX)).toBe(true); + expect(channel.consume({ jsonrpc: "2.0", id: 7, result: {} })).toBe(false); + // A late answer to an id this channel issued is swallowed, not forwarded. + expect( + channel.consume({ jsonrpc: "2.0", id: `${id}-late`, result: {} }), + ).toBe(true); + expect(channel.consume({ jsonrpc: "2.0", id, result: { ok: true } })).toBe( + true, + ); + await expect(promise).resolves.toEqual({ + kind: "result", + result: { ok: true }, + }); + }); + + it("returns error responses, dropping non-JSON data", async () => { + const { channel, sent } = setup(); + const promise = channel.dispatch({ method: "tasks/get" }); + await Promise.resolve(); + channel.consume({ + jsonrpc: "2.0", + id: String(sent[0]!.id), + error: { code: -1, message: "no", data: Number.NaN }, + }); + await expect(promise).resolves.toEqual({ + kind: "error", + error: { code: -1, message: "no" }, + }); + }); + + it("rejects a non-JSON result", async () => { + const { channel, sent } = setup(); + const promise = channel.dispatch({ method: "tasks/get" }); + await Promise.resolve(); + channel.consume({ + jsonrpc: "2.0", + id: String(sent[0]!.id), + result: { n: Number.POSITIVE_INFINITY }, + }); + await expect(promise).rejects.toThrow(/non-JSON result/); + }); + + it("times out with the annotated timeout as its cause, re-armed by progress", async () => { + vi.useFakeTimers(); + const annotateTimeout = vi.fn((error: unknown) => error); + const { channel, sent } = setup({ annotateTimeout }); + const promise = channel.dispatch({ + method: "tools/call", + params: { _meta: { progressToken: "p" } }, + }); + let settled = false; + promise.catch(() => { + settled = true; + }); + await vi.advanceTimersByTimeAsync(800); + channel.noteProgress("p"); + channel.noteProgress("other"); + await vi.advanceTimersByTimeAsync(800); + expect(settled).toBe(false); + await vi.advanceTimersByTimeAsync(300); + await expect(promise).rejects.toSatisfy( + (error) => + error instanceof DispatchError && + (error.cause as { code?: unknown }).code !== undefined, + ); + expect(annotateTimeout).toHaveBeenCalledWith( + expect.anything(), + "tools/call", + ); + // stdio-style transport: the timeout reaches the server as a cancel. + expect(sent[1]).toMatchObject({ method: "notifications/cancelled" }); + }); + + it("honors an explicit timeout override", async () => { + vi.useFakeTimers(); + const { channel } = setup(); + const promise = channel.dispatch({ method: "tasks/get" }, {}, 5); + const assertion = expect(promise).rejects.toThrow(/after 5 ms/); + await vi.advanceTimersByTimeAsync(5); + await assertion; + }); + + it("rejects an already-aborted signal without sending", async () => { + const { channel, sent } = setup(); + const controller = new AbortController(); + controller.abort(new Error("gone")); + await expect( + channel.dispatch({ method: "tasks/get" }, { signal: controller.signal }), + ).rejects.toThrow("gone"); + expect(sent).toHaveLength(0); + }); + + it("cancels on the wire when the caller aborts", async () => { + const { channel, sent } = setup(); + const controller = new AbortController(); + const promise = channel.dispatch( + { method: "tools/call" }, + { signal: controller.signal }, + ); + await Promise.resolve(); + controller.abort("user"); + await expect(promise).rejects.toMatchObject({ name: "AbortError" }); + expect(sent[1]).toMatchObject({ + method: "notifications/cancelled", + params: { reason: "user" }, + }); + }); + + it("annotates a send failure and rejects everything on rejectAll", async () => { + const boom = new Error("closed"); + const annotateTimeout = vi.fn(() => boom); + const failing = setup({ annotateTimeout }); + failing.transport.send.mockRejectedValueOnce("raw failure"); + await expect( + failing.channel.dispatch({ method: "tasks/get" }), + ).rejects.toBe(boom); + + const { channel } = setup(); + const pending = channel.dispatch({ method: "tasks/get" }); + channel.rejectAll("Disconnected"); + await expect(pending).rejects.toThrow("Disconnected"); + }); +}); diff --git a/clients/web/src/test/core/extension/tasks/session.test.ts b/clients/web/src/test/core/extension/tasks/session.test.ts new file mode 100644 index 0000000000..19a44293ef --- /dev/null +++ b/clients/web/src/test/core/extension/tasks/session.test.ts @@ -0,0 +1,58 @@ +import { describe, expect, it } from "vitest"; +import { + isTasksExtensionNegotiated, + taskSessionEndpointId, +} from "@inspector/core/extension/tasks/session.js"; +import { TASKS_EXTENSION_KEY } from "@inspector/core/extension/tasks/constants.js"; + +const clientInfo = { name: "inspector", version: "1.0.0" }; + +describe("taskSessionEndpointId", () => { + it("is stable for one target and differs across targets", async () => { + const http = { type: "streamable-http" as const, url: "http://h/mcp" }; + const a = await taskSessionEndpointId(http, clientInfo); + expect(await taskSessionEndpointId(http, clientInfo)).toBe(a); + expect( + await taskSessionEndpointId( + { type: "sse", url: "http://h/mcp" }, + clientInfo, + ), + ).not.toBe(a); + }); + + it("covers stdio targets, with and without a cwd", async () => { + const stdio = { type: "stdio" as const, command: "node", args: ["s.js"] }; + const bare = await taskSessionEndpointId(stdio, clientInfo); + const withCwd = await taskSessionEndpointId( + { ...stdio, cwd: "/tmp" }, + clientInfo, + ); + expect(bare).not.toBe(withCwd); + }); +}); + +describe("isTasksExtensionNegotiated", () => { + const caps = { extensions: { [TASKS_EXTENSION_KEY]: {} } }; + + it("needs both the modern era and the advertised extension", () => { + expect(isTasksExtensionNegotiated("modern", caps)).toBe(true); + expect(isTasksExtensionNegotiated("legacy", caps)).toBe(false); + expect(isTasksExtensionNegotiated("modern", {})).toBe(false); + expect(isTasksExtensionNegotiated(undefined, undefined)).toBe(false); + }); + + it("accepts only the empty-object shape ext-tasks starts a session on", () => { + const withValue = (value: unknown) => + // Single cast: the SDK types the value as an object; these are the + // malformed wire shapes it cannot express. + ({ extensions: { [TASKS_EXTENSION_KEY]: value } }) as Parameters< + typeof isTasksExtensionNegotiated + >[1]; + expect(isTasksExtensionNegotiated("modern", withValue({ x: 1 }))).toBe( + false, + ); + expect(isTasksExtensionNegotiated("modern", withValue(null))).toBe(false); + expect(isTasksExtensionNegotiated("modern", withValue([]))).toBe(false); + expect(isTasksExtensionNegotiated("modern", withValue("yes"))).toBe(false); + }); +}); diff --git a/clients/web/src/test/core/extension/tasks/wire.test.ts b/clients/web/src/test/core/extension/tasks/wire.test.ts new file mode 100644 index 0000000000..faff54f33c --- /dev/null +++ b/clients/web/src/test/core/extension/tasks/wire.test.ts @@ -0,0 +1,91 @@ +import { describe, expect, it } from "vitest"; +import { RELATED_TASK_META_KEY } from "@modelcontextprotocol/client"; +import type { TaskView } from "@modelcontextprotocol/ext-tasks/client"; +import { taskId } from "@modelcontextprotocol/ext-tasks/core"; +import { + jsonObject, + paramsWithRelatedTask, + taskInputOrigin, + taskToolResultCodec, + toInspectorTask, +} from "@inspector/core/extension/tasks/wire.js"; + +describe("taskToolResultCodec", () => { + it("accepts a CallToolResult and reports issues for anything else", () => { + const ok = taskToolResultCodec.parse({ + content: [{ type: "text", text: "hi" }], + }); + expect(ok).toMatchObject({ success: true }); + const bad = taskToolResultCodec.parse({ content: "nope" }); + expect(bad).toMatchObject({ success: false }); + }); +}); + +describe("jsonObject", () => { + it("projects an object and rejects non-objects", () => { + expect(jsonObject({ a: 1, b: undefined })).toEqual({ a: 1 }); + expect(() => jsonObject([1])).toThrow(TypeError); + expect(() => jsonObject("x")).toThrow(TypeError); + }); +}); + +describe("paramsWithRelatedTask", () => { + it("stamps the related task and keeps existing _meta", () => { + const params = paramsWithRelatedTask( + { message: "m", _meta: { keep: true } }, + "task-1", + ); + expect(params._meta).toMatchObject({ + keep: true, + [RELATED_TASK_META_KEY]: { taskId: "task-1" }, + }); + }); + + it("ignores a non-object _meta", () => { + const params = paramsWithRelatedTask({ _meta: [1] }, "task-2"); + expect(params._meta).toEqual({ + [RELATED_TASK_META_KEY]: { taskId: "task-2" }, + }); + }); +}); + +describe("toInspectorTask", () => { + function view( + times: Pick<TaskView, "createdAt" | "lastUpdatedAt">, + ): TaskView { + return { + taskId: taskId("a"), + status: "working", + retentionMs: null, + ttl: null, + raw: {}, + extensions: {}, + ...times, + }; + } + + it("fills whichever timestamp is missing from the other", () => { + expect(toInspectorTask(view({ lastUpdatedAt: "t2" }))).toMatchObject({ + createdAt: "t2", + lastUpdatedAt: "t2", + }); + expect(toInspectorTask(view({ createdAt: "t1" }))).toMatchObject({ + createdAt: "t1", + lastUpdatedAt: "t1", + }); + expect(toInspectorTask(view({}))).toMatchObject({ + createdAt: "", + lastUpdatedAt: "", + }); + }); +}); + +describe("taskInputOrigin", () => { + it("maps each delivery to its pending-request origin", () => { + expect( + (["peer-request", "request-retry", "task-update"] as const).map( + taskInputOrigin, + ), + ).toEqual(["server-request", "input-required", "task-input-required"]); + }); +}); diff --git a/clients/web/src/test/core/mcp/extensions.test.ts b/clients/web/src/test/core/mcp/extensions.test.ts index 1695aaca3a..8ec3ff6540 100644 --- a/clients/web/src/test/core/mcp/extensions.test.ts +++ b/clients/web/src/test/core/mcp/extensions.test.ts @@ -7,7 +7,7 @@ import { buildClientExtensions, isAdvertisedByDefault, } from "@inspector/core/mcp/extensions.js"; -import { TASKS_EXTENSION_KEY } from "@inspector/core/mcp/modernTaskSchemas.js"; +import { TASKS_EXTENSION_KEY } from "@inspector/core/extension/tasks/constants.js"; import { SKILLS_EXTENSION_KEY } from "@inspector/core/mcp/skillsSchemas.js"; // The `ui` extension carries a non-empty advertisement value; the others are diff --git a/clients/web/src/test/core/mcp/fetchTracking.test.ts b/clients/web/src/test/core/mcp/fetchTracking.test.ts index 4a5f2d094e..d9019501b3 100644 --- a/clients/web/src/test/core/mcp/fetchTracking.test.ts +++ b/clients/web/src/test/core/mcp/fetchTracking.test.ts @@ -9,6 +9,7 @@ import { redactSensitiveHeaders, redactBody, redactUrlQuery, + redactUrlsInText, REDACTED_HEADER_VALUE, REDACTED_VALUE, } from "@inspector/core/mcp/fetchTracking.js"; @@ -864,6 +865,70 @@ describe("redactUrlQuery", () => { }); }); +describe("redactUrlsInText", () => { + const R = encodeURIComponent(REDACTED_VALUE); + + it("redacts a URL embedded in prose, keeping trailing punctuation", () => { + expect( + redactUrlsInText( + "Request to https://srv.example/mcp?code=abc&tenant=acme. Retry later", + ), + ).toBe( + `Request to https://srv.example/mcp?code=${R}&tenant=acme. Retry later`, + ); + }); + + it("keeps a punctuation run inside the URL and splits only the trailing one", () => { + // A long run that does not end the match was quadratic under an unanchored + // /[…]+$/ (#2540); the backward scan must still stop at `x`. + const run = "!".repeat(50_000); + const out = redactUrlsInText( + `see https://srv.example/cb?note=${run}x&code=abc123!?`, + ); + expect(out).toMatch(new RegExp(`x&code=${R}!\\?$`)); + expect(out).not.toContain("abc123"); + }); + + it("redacts a URL whose scheme is upper- or mixed-case", () => { + expect( + redactUrlsInText( + "a HTTP://s.example/cb?access_token=t b HttpS://s.example/cb?code=c", + ), + ).toBe( + `a HTTP://s.example/cb?access_token=${R} b HttpS://s.example/cb?code=${R}`, + ); + }); + + it("redacts each of two comma-joined URLs separately", () => { + expect( + redactUrlsInText( + "https://one.example/cb?state=ok,https://two.example/cb?code=secret", + ), + ).toBe(`https://one.example/cb?state=ok,https://two.example/cb?code=${R}`); + }); + + it("redacts through an apostrophe and keeps a closing quote", () => { + expect(redactUrlsInText("at https://s.example/cb?code=abc'def now")).toBe( + `at https://s.example/cb?code=${R} now`, + ); + expect(redactUrlsInText("at 'https://s.example/cb?code=abc123'.")).toBe( + `at 'https://s.example/cb?code=${R}'.`, + ); + }); + + it("stops at a double quote, so a URL inside serialized JSON is redacted", () => { + expect( + redactUrlsInText(JSON.stringify({ url: "https://s.example/?token=t" })), + ).toBe(`{"url":"https://s.example/?token=${R}"}`); + }); + + it("leaves text without a sensitive URL untouched", () => { + const text = "Failed at https://srv.example/mcp?tenant=acme and nowhere"; + expect(redactUrlsInText(text)).toBe(text); + expect(redactUrlsInText("no url here")).toBe("no url here"); + }); +}); + describe("redactBody", () => { it("returns undefined / empty bodies unchanged", () => { expect(redactBody(undefined, "application/json")).toBeUndefined(); diff --git a/clients/web/src/test/core/mcp/inspectorClient-list-cursor.test.ts b/clients/web/src/test/core/mcp/inspectorClient-list-cursor.test.ts index 345f60b7c3..17f4540564 100644 --- a/clients/web/src/test/core/mcp/inspectorClient-list-cursor.test.ts +++ b/clients/web/src/test/core/mcp/inspectorClient-list-cursor.test.ts @@ -18,6 +18,12 @@ import { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; * pins is that `""` survives and that a genuinely absent cursor still sends no * `cursor` key at all. The SDK client is stubbed rather than connected — the * decision under test is made entirely in `InspectorClient`. + * + * `listRequestorTasks` is asserted against the ext-tasks session instead of + * the SDK client because this branch routes it through + * `runTaskSessionOperation`; the session owns the wire params, and its own + * suite pins the verbatim-cursor behavior there. The Inspector-side property + * is that the cursor reaches `session.listTasks` unchanged. */ describe("InspectorClient list cursor handling (#2220)", () => { /** @@ -106,12 +112,6 @@ describe("InspectorClient list cursor handling (#2220)", () => { result: { resourceTemplates: [] }, call: (client, cursor) => client.listResourceTemplates(cursor), }, - { - name: "listRequestorTasks", - method: "tasks/list", - result: { tasks: [] }, - call: (client, cursor) => client.listRequestorTasks(cursor), - }, ]; it.each(ADAPTERS)( @@ -154,4 +154,27 @@ describe("InspectorClient list cursor handling (#2220)", () => { expect(sent.params.cursor).toBe("page-2"); }, ); + + it.each([[""], [undefined], ["page-2"]])( + "listRequestorTasks forwards cursor %j to the ext-tasks session unchanged", + async (cursor) => { + const client = makeClient(); + const listTasks = vi.fn(async (received?: string) => { + // The session owns the wire params; the Inspector-side property is + // that the cursor arrives here verbatim. + void received; + return { tasks: [] }; + }); + ( + client as unknown as { + taskSession: { listTasks: typeof listTasks } | null; + } + ).taskSession = { listTasks }; + + await client.listRequestorTasks(cursor); + + expect(listTasks).toHaveBeenCalledTimes(1); + expect(listTasks.mock.calls[0][0]).toBe(cursor); + }, + ); }); diff --git a/clients/web/src/test/core/mcp/inspectorClient-peer-handler-timing.test.ts b/clients/web/src/test/core/mcp/inspectorClient-peer-handler-timing.test.ts index cfc22003e4..c1ddd36c8d 100644 --- a/clients/web/src/test/core/mcp/inspectorClient-peer-handler-timing.test.ts +++ b/clients/web/src/test/core/mcp/inspectorClient-peer-handler-timing.test.ts @@ -5,7 +5,7 @@ import type { Transport, } from "@modelcontextprotocol/client"; import { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; -import { ModernGetTaskResultSchema } from "@inspector/core/mcp/modernTaskSchemas.js"; +import { GetTaskResultV2Schema as ModernGetTaskResultSchema } from "@modelcontextprotocol/ext-tasks/core/v2"; /** * Regression coverage for #1797: a server may talk to us the instant it is @@ -790,13 +790,8 @@ describe("InspectorClient peer-handler timing (#1797)", () => { ); await client.connect(); - // Seeded directly: subscribing for real needs a server that answers - // `resources/subscribe` and cancelling needs a live task, neither of which - // adds to what is under test — that a new session starts empty. The cast is - // the only route to `cancelledTaskIds`, which has no public reader. const internals = client as unknown as { subscribedResources: Set<string>; - cancelledTaskIds: Set<string>; modernStreamState: { active: boolean; status: string; @@ -804,7 +799,6 @@ describe("InspectorClient peer-handler timing (#1797)", () => { }; }; internals.subscribedResources.add("file:///watched"); - internals.cancelledTaskIds.add("task-1"); // The stream state a live modern subscription would have left behind. internals.modernStreamState = { active: true, @@ -829,7 +823,6 @@ describe("InspectorClient peer-handler timing (#1797)", () => { await client.connect(); expect(client.getSubscribedResources()).toEqual([]); - expect(internals.cancelledTaskIds.size).toBe(0); // Cleared with the set it is derived from, not left reading `active` for // an empty one. expect(client.getResourceSubscriptionStreamState()).toMatchObject({ @@ -880,35 +873,6 @@ describe("InspectorClient peer-handler timing (#1797)", () => { await client.disconnect(); }); - it("aborts a paused task-input wait when the session ends", async () => { - // The bounded-window member: both registration sites release in a - // `finally`, so nothing leaks permanently — this closes the gap between a - // crash and the loop unwinding on its own. - const transport = new SampleAfterConnectTransport(); - const client = new InspectorClient( - { type: "stdio", command: "noop", args: [] }, - { environment: { transport: () => ({ transport }) } }, - ); - await client.connect(); - - // Seeded directly: reaching this map for real needs a modern task paused at - // `input_required`, which adds nothing to what is under test. No public - // reader, hence the cast. - const controller = new AbortController(); - ( - client as unknown as { - taskInputAbortControllers: Map<string, AbortController>; - } - ).taskInputAbortControllers.set("task-1", controller); - - transport.onclose?.(); - await client.connect(); - - expect(controller.signal.aborted).toBe(true); - - await client.disconnect(); - }); - it("closes a live listen stream the next connect drops", async () => { // An `onerror` without an `onclose` leaves the transport up, and `connect()` // reuses it — so the reference the reset drops can be the last one to a @@ -1001,25 +965,15 @@ describe("InspectorClient peer-handler timing (#1797)", () => { ); await client.connect(); - // Seeded directly, for the reasons `closes a live listen stream the - // next connect drops` and `aborts a paused task-input wait when the - // session ends` give. No public writer for either, hence the casts. + // Seed the live subscription directly; it has no public writer. ( client as unknown as { modernSubscription: { close: () => Promise<void> } | null; } ).modernSubscription = { close }; - // A downstream teardown step, to witness that teardown continued. - const controller = new AbortController(); - ( - client as unknown as { - taskInputAbortControllers: Map<string, AbortController>; - } - ).taskInputAbortControllers.set("task-1", controller); - await expect(client.disconnect()).resolves.toBeUndefined(); - expect(controller.signal.aborted).toBe(true); + expect(client.getStatus()).toBe("disconnected"); // Node reports an unhandled rejection after the microtask checkpoint, // so yield to the macrotask queue before reading the listener — nothing diff --git a/clients/web/src/test/core/mcp/inspectorClient-raw-wire.test.ts b/clients/web/src/test/core/mcp/inspectorClient-raw-wire.test.ts index 638891b573..2a6f9f06a8 100644 --- a/clients/web/src/test/core/mcp/inspectorClient-raw-wire.test.ts +++ b/clients/web/src/test/core/mcp/inspectorClient-raw-wire.test.ts @@ -1,7 +1,18 @@ import { describe, it, expect, vi } from "vitest"; +import { + ProtocolError, + type CallToolResult, + type Tool, +} from "@modelcontextprotocol/client"; +import { + DispatchError, + JsonRpcResponseError, +} from "@modelcontextprotocol/ext-tasks/client"; import { SdkError, SdkErrorCode } from "@modelcontextprotocol/client"; import { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; -import { ModernGetTaskResultSchema } from "@inspector/core/mcp/modernTaskSchemas.js"; +import type { TaskProgressRouter } from "@inspector/core/extension/tasks/progress.js"; +import type { TaskWithOptionalCreatedAt } from "@inspector/core/mcp/inspectorClientEventTarget.js"; +import { GetTaskResultV2Schema as ModernGetTaskResultSchema } from "@modelcontextprotocol/ext-tasks/core/v2"; /** * Unit coverage for the raw-wire request channel (#1631) that drives the modern @@ -21,8 +32,28 @@ describe("InspectorClient raw-wire channel (#1631)", () => { } interface RawWireInternals { - transport: { send: (m: unknown) => Promise<void> } | null; + transport: { + send: ( + message: unknown, + options?: { + headers?: Readonly<Record<string, string>>; + requestSignal?: AbortSignal; + }, + ) => Promise<void>; + hasPerRequestStream?: boolean; + } | null; requestTimeout?: number; + resetTimeoutOnProgress?: boolean; + dispatchTaskRequest: ( + request: unknown, + options?: { + signal?: AbortSignal; + context?: { + headers?: Readonly<Record<string, string>>; + requestTimeoutMs?: number; + }; + }, + ) => Promise<unknown>; rawWireRequest: ( method: string, params: Record<string, unknown>, @@ -32,10 +63,83 @@ describe("InspectorClient raw-wire channel (#1631)", () => { rejectPendingRawWireRequests: (reason: string) => void; } + interface TaskSessionCallOptions { + requestTimeoutMs?: number; + metadata?: Readonly<Record<string, unknown>>; + task: { preference: "allow" | "prefer"; retentionMs?: number }; + } + + interface TaskExecutionSettleOptions { + signal?: AbortSignal; + onEvent: (event: unknown) => void; + } + + interface TaskBoundaryInternals { + client: object | null; + status: string; + disconnecting: boolean; + protocolVersion?: string; + transportConfig: { type: string; command?: string; args?: string[] }; + attachTaskSession: () => Promise<void>; + closeTaskSession: () => Promise<void>; + protocolEra?: "legacy" | "modern"; + taskSession: { + callTool: ( + name: string, + args: Readonly<Record<string, unknown>>, + options: TaskSessionCallOptions, + ) => Promise<{ + kind?: string; + serializeReference?: () => Record<string, unknown>; + settle: ( + options: TaskExecutionSettleOptions, + ) => Promise<{ outcome: unknown; lastTask?: unknown }>; + }>; + resumeTask?: ( + reference: Record<string, unknown>, + options: Record<string, unknown>, + ) => Promise<{ + kind?: string; + settle: ( + options: TaskExecutionSettleOptions, + ) => Promise<{ outcome: unknown; lastTask?: unknown }>; + }>; + } | null; + emitTaskExecutionEvent: (event: unknown) => unknown; + emitTaskError: (lastTask: unknown, reason: unknown) => void; + dispatchTaskProgress: (notification: unknown) => void; + } + function internals(client: InspectorClient): RawWireInternals { return client as unknown as RawWireInternals; } + function taskInternals(client: InspectorClient): TaskBoundaryInternals { + // Private session fields deliberately have no public mutation API; this narrow + // structurally matching cast injects only the ext-tasks boundary under test. + return client as unknown as TaskBoundaryInternals; + } + + const taskTool: Tool = { + name: "boundary_task", + description: "Exercises the Inspector/ext-tasks boundary", + inputSchema: { type: "object" }, + }; + + const successfulResult: CallToolResult = { + content: [{ type: "text", text: "done" }], + }; + + function attachTaskBoundary( + client: InspectorClient, + callTool: NonNullable<TaskBoundaryInternals["taskSession"]>["callTool"], + ): void { + const boundary = taskInternals(client); + boundary.client = {}; + boundary.protocolEra = "modern"; + boundary.taskSession = { callTool }; + } + it("throws when there is no transport", async () => { const client = makeClient(); internals(client).transport = null; @@ -67,12 +171,13 @@ describe("InspectorClient raw-wire channel (#1631)", () => { const consumed = internals(client).consumeRawWireResponse({ id: sent!.id, result: { + resultType: "complete", taskId: "x", status: "completed", createdAt: "a", lastUpdatedAt: "b", ttlMs: null, - result: { content: [] }, + result: { resultType: "complete", content: [] }, }, }); expect(consumed).toBe(true); @@ -151,6 +256,237 @@ describe("InspectorClient raw-wire channel (#1631)", () => { } }); + it("sends notifications/cancelled on timeout for transports without a per-request stream", async () => { + vi.useFakeTimers(); + try { + // A timed-out raw request must reach the server-side cancellation + // path too, or the orphaned call keeps running (round-11 finding). + const client = makeClient(); + internals(client).requestTimeout = 10; + const sent: unknown[] = []; + internals(client).transport = { + send: vi.fn(async (message) => { + sent.push(message); + }), + }; + const promise = internals(client).rawWireRequest( + "tasks/get", + {}, + ModernGetTaskResultSchema, + ); + const assertion = expect(promise).rejects.toThrow(/timed out/); + await vi.advanceTimersByTimeAsync(20); + await assertion; + const requestId = (sent[0] as { id: string }).id; + expect(sent[1]).toEqual({ + jsonrpc: "2.0", + method: "notifications/cancelled", + params: { + requestId, + reason: "Request timed out after 10 ms", + }, + }); + + // With a per-request stream the SDK-mirrored fork stays silent: the + // transport's own stream teardown is the cancellation signal: the + // timeout aborts the forwarded per-request wire signal instead. + const streaming = makeClient(); + internals(streaming).requestTimeout = 10; + const streamingSent: unknown[] = []; + let wireSignal: AbortSignal | undefined; + internals(streaming).transport = { + send: vi.fn(async (message, options) => { + streamingSent.push(message); + wireSignal ??= options?.requestSignal; + }), + hasPerRequestStream: true, + }; + const streamingPromise = internals(streaming).rawWireRequest( + "tasks/get", + {}, + ModernGetTaskResultSchema, + ); + const streamingAssertion = + expect(streamingPromise).rejects.toThrow(/timed out/); + await vi.advanceTimersByTimeAsync(20); + await streamingAssertion; + expect(streamingSent).toHaveLength(1); + expect(wireSignal?.aborted).toBe(true); + } finally { + vi.useRealTimers(); + } + }); + + it("honors the ext-tasks operation timeout context", async () => { + vi.useFakeTimers(); + try { + const client = makeClient(); + internals(client).requestTimeout = 10_000; + internals(client).transport = { + send: vi.fn().mockResolvedValue(undefined), + }; + const promise = internals(client).dispatchTaskRequest( + { method: "tasks/get" }, + { context: { requestTimeoutMs: 25 } }, + ); + const assertion = expect(promise).rejects.toThrow(/25 ms/); + await vi.advanceTimersByTimeAsync(25); + await assertion; + } finally { + vi.useRealTimers(); + } + }); + + it("re-arms the raw request timeout on progress when resetTimeoutOnProgress is enabled", async () => { + vi.useFakeTimers(); + try { + const client = makeClient(); + internals(client).requestTimeout = 10; + internals(client).transport = { + send: vi.fn().mockResolvedValue(undefined), + }; + const promise = internals(client).dispatchTaskRequest({ + method: "tools/call", + params: { name: "x", _meta: { progressToken: "raw-progress" } }, + }); + let settled = false; + void promise.catch(() => { + settled = true; + }); + await vi.advanceTimersByTimeAsync(8); + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken: "raw-progress", progress: 1 }, + }); + // 16ms elapsed exceeds the 10ms budget, but progress at 8ms re-armed it. + await vi.advanceTimersByTimeAsync(8); + expect(settled).toBe(false); + const assertion = expect(promise).rejects.toThrow( + /timed out after 10 ms/, + ); + await vi.advanceTimersByTimeAsync(10); + await assertion; + } finally { + vi.useRealTimers(); + } + }); + + it("does not re-arm the raw request timeout when resetTimeoutOnProgress is disabled", async () => { + vi.useFakeTimers(); + try { + const client = makeClient(); + internals(client).requestTimeout = 10; + internals(client).resetTimeoutOnProgress = false; + internals(client).transport = { + send: vi.fn().mockResolvedValue(undefined), + }; + const promise = internals(client).dispatchTaskRequest({ + method: "tools/call", + params: { name: "x", _meta: { progressToken: "raw-progress" } }, + }); + const assertion = expect(promise).rejects.toThrow( + /timed out after 10 ms/, + ); + await vi.advanceTimersByTimeAsync(8); + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken: "raw-progress", progress: 1 }, + }); + await vi.advanceTimersByTimeAsync(4); + await assertion; + } finally { + vi.useRealTimers(); + } + }); + + it("does not install a task session when a disconnect overtakes attachTaskSession", async () => { + const client = makeClient(); + const boundary = taskInternals(client); + // Minimal SDK-client surface createTaskSessionFromClient reads on install. + const sdkClient = { + getServerCapabilities: () => ({}), + getProtocolEra: () => "modern" as const, + }; + boundary.client = sdkClient; + boundary.status = "connected"; + boundary.protocolVersion = "2026-07-28"; + + // No SDK client at all: attach is a no-op. + boundary.client = null; + await boundary.attachTaskSession(); + expect(boundary.taskSession).toBeNull(); + boundary.client = sdkClient; + + // A disconnect that claimed teardown ownership while attach was suspended. + boundary.disconnecting = true; + await boundary.attachTaskSession(); + expect(boundary.taskSession).toBeNull(); + + // A disconnect that already settled the status. + boundary.disconnecting = false; + boundary.status = "disconnected"; + await boundary.attachTaskSession(); + expect(boundary.taskSession).toBeNull(); + + // A reconnect that replaced the SDK client while attach was suspended on + // its first await (the endpoint-id derivation). + boundary.status = "connected"; + const staleAttach = boundary.attachTaskSession(); + boundary.client = {}; + await staleAttach; + expect(boundary.taskSession).toBeNull(); + + // The same attach with no overtaking teardown installs the session. + boundary.client = sdkClient; + await boundary.attachTaskSession(); + expect(boundary.taskSession).not.toBeNull(); + await boundary.closeTaskSession(); + }); + + it("projects configured roots to the typed ext-tasks handler shape", () => { + const client = makeClient(); + const boundary = client as unknown as { + roots?: readonly Record<string, unknown>[]; + applicationRoots: () => readonly Record<string, unknown>[]; + }; + // No roots configured: an empty list, not undefined. + expect(boundary.applicationRoots()).toEqual([]); + boundary.roots = [ + { uri: "file:///bare" }, + { uri: "file:///named", name: "Named" }, + { uri: "file:///meta", _meta: { vendor: true } }, + ]; + expect(boundary.applicationRoots()).toEqual([ + { uri: "file:///bare" }, + { uri: "file:///named", name: "Named" }, + { uri: "file:///meta", _meta: { vendor: true } }, + ]); + }); + + it("uses the SDK 60-second default when no timeout is configured", async () => { + vi.useFakeTimers(); + try { + const client = makeClient(); + internals(client).transport = { + send: vi.fn().mockResolvedValue(undefined), + }; + const promise = internals(client).dispatchTaskRequest({ + method: "tasks/get", + }); + let settled = false; + void promise.catch(() => { + settled = true; + }); + await vi.advanceTimersByTimeAsync(30_000); + expect(settled).toBe(false); + const assertion = expect(promise).rejects.toThrow(/60000 ms/); + await vi.advanceTimersByTimeAsync(30_000); + await assertion; + } finally { + vi.useRealTimers(); + } + }); + it("annotates a request timeout the transport's send rejects with (#2318)", async () => { // The browser's remote transport awaits the response inside `send`, so // its relay wait can expire before the local timer; it arrives here as @@ -190,7 +526,6 @@ describe("InspectorClient raw-wire channel (#1631)", () => { ), ).rejects.toBe(boom); }); - it("rejects all pending requests on teardown", async () => { const client = makeClient(); internals(client).transport = { @@ -205,4 +540,907 @@ describe("InspectorClient raw-wire channel (#1631)", () => { internals(client).rejectPendingRawWireRequests("Disconnected"); await expect(promise).rejects.toThrow(/Disconnected/); }); + + it.each([ + [null, /JSON object/], + ["not-an-object", /JSON object/], + [7, /JSON object/], + [[], /JSON object/], + [{ method: 42 }, /method must be a string/], + [{ method: "tasks/get", params: null }, /params must be a JSON object/], + [{ method: "tasks/get", params: [] }, /params must be a JSON object/], + [{ method: "tasks/get", params: "bad" }, /params must be a JSON object/], + [{ method: "tasks/get", params: 7 }, /params must be a JSON object/], + ])("validates ext-tasks dispatch input %#", async (request, message) => { + const client = makeClient(); + internals(client).transport = { + send: vi.fn().mockResolvedValue(undefined), + }; + await expect( + internals(client).dispatchTaskRequest(request), + ).rejects.toThrow(message); + }); + + it("rejects a dispatch that is already aborted without sending", async () => { + const client = makeClient(); + const send = vi.fn().mockResolvedValue(undefined); + internals(client).transport = { send }; + const controller = new AbortController(); + controller.abort(new Error("stop before dispatch")); + + await expect( + internals(client).dispatchTaskRequest( + { method: "tasks/get", params: { taskId: "x" } }, + { signal: controller.signal }, + ), + ).rejects.toThrow(/stop before dispatch/); + expect(send).not.toHaveBeenCalled(); + }); + + it("uses the standard abort error for a non-Error reason", async () => { + const client = makeClient(); + const send = vi.fn().mockResolvedValue(undefined); + internals(client).transport = { send }; + const controller = new AbortController(); + controller.abort("stop"); + + await expect( + internals(client).dispatchTaskRequest( + { method: "tasks/get" }, + { signal: controller.signal }, + ), + ).rejects.toMatchObject({ name: "AbortError" }); + expect(send).not.toHaveBeenCalled(); + }); + + it("frames an omitted-params dispatch and forwards headers and its signal", async () => { + const client = makeClient(); + let sent: { id: string; method: string; params?: unknown } | undefined; + let sendOptions: + | { + headers?: Readonly<Record<string, string>>; + requestSignal?: AbortSignal; + } + | undefined; + internals(client).transport = { + send: vi.fn(async (message, options) => { + sent = message as { id: string; method: string; params?: unknown }; + sendOptions = options; + }), + }; + const controller = new AbortController(); + const promise = internals(client).dispatchTaskRequest( + { method: "tasks/list" }, + { + signal: controller.signal, + context: { headers: { "x-route": "blue" } }, + }, + ); + await Promise.resolve(); + + expect(sent).toEqual({ + jsonrpc: "2.0", + id: expect.stringMatching(/^inspector-ext-/), + method: "tasks/list", + }); + expect(sendOptions).toEqual({ + headers: { "x-route": "blue" }, + // The wire signal is a per-request controller linked to the caller's, + // not the caller's own: a timeout must also be able to abort the + // per-request stream (round-12). Forwarding and the timeout abort are + // asserted in the transport-fork tests. + requestSignal: expect.any(AbortSignal), + }); + expect(sendOptions?.requestSignal).not.toBe(controller.signal); + internals(client).consumeRawWireResponse({ id: sent!.id, result: {} }); + await expect(promise).resolves.toEqual({ kind: "result", result: {} }); + }); + + it("aborts an in-flight dispatch and ignores its late response", async () => { + const client = makeClient(); + let sentId = ""; + internals(client).transport = { + send: vi.fn(async (message) => { + // Only the request carries an id; the abort's notifications/cancelled + // does not, and must not overwrite it. + sentId ||= (message as { id?: string }).id ?? ""; + }), + }; + const controller = new AbortController(); + const promise = internals(client).dispatchTaskRequest( + { method: "tasks/get", params: { taskId: "x" } }, + { signal: controller.signal }, + ); + await Promise.resolve(); + controller.abort(new Error("stop in flight")); + + await expect(promise).rejects.toThrow(/stop in flight/); + expect(sentId).toMatch(/^inspector-ext-/); + // Swallowed rather than forwarded: the SDK would report an unknown id. + expect( + internals(client).consumeRawWireResponse({ id: sentId, result: {} }), + ).toBe(true); + }); + + it("forks cancellation by transport: notifications/cancelled without a per-request stream, requestSignal with one (#2140)", async () => { + // stdio/SSE (no per-request stream): the abort must POST the correlated + // cancellation frame, because these transports ignore requestSignal and + // this path bypasses Client.request, which would otherwise send it. + const client = makeClient(); + const sent: unknown[] = []; + internals(client).transport = { + send: vi.fn(async (message) => { + sent.push(message); + }), + }; + const controller = new AbortController(); + const promise = internals(client).dispatchTaskRequest( + { method: "tasks/get", params: { taskId: "x" } }, + { signal: controller.signal }, + ); + await Promise.resolve(); + controller.abort("user cancelled"); + await expect(promise).rejects.toThrow(); + const requestId = (sent[0] as { id: string }).id; + expect(sent[1]).toEqual({ + jsonrpc: "2.0", + method: "notifications/cancelled", + params: { requestId, reason: "user cancelled" }, + }); + + // A per-request-stream transport (2026-era Streamable HTTP) already + // treats the forwarded requestSignal abort as the wire cancellation, so + // no notification is sent. + const streaming = makeClient(); + const streamingSent: unknown[] = []; + internals(streaming).transport = { + send: vi.fn(async (message) => { + streamingSent.push(message); + }), + hasPerRequestStream: true, + }; + const streamingController = new AbortController(); + const streamingPromise = internals(streaming).dispatchTaskRequest( + { method: "tasks/get", params: { taskId: "x" } }, + { signal: streamingController.signal }, + ); + await Promise.resolve(); + streamingController.abort(new Error("stop")); + await expect(streamingPromise).rejects.toThrow("stop"); + expect(streamingSent).toHaveLength(1); + }); + + it("ignores a transport rejection after abort already settled the dispatch", async () => { + const client = makeClient(); + let rejectSend: ((reason: unknown) => void) | undefined; + internals(client).transport = { + send: vi.fn( + () => + new Promise<void>((_resolve, reject) => { + rejectSend = reject; + }), + ), + }; + const controller = new AbortController(); + const promise = internals(client).dispatchTaskRequest( + { method: "tasks/get" }, + { signal: controller.signal }, + ); + await Promise.resolve(); + controller.abort(new Error("caller stopped")); + await expect(promise).rejects.toThrow("caller stopped"); + + rejectSend?.(new Error("late socket failure")); + await Promise.resolve(); + }); + + it("normalizes a non-Error transport rejection", async () => { + const client = makeClient(); + internals(client).transport = { + send: vi.fn().mockRejectedValue("socket vanished"), + }; + await expect( + internals(client).dispatchTaskRequest({ method: "tasks/get" }), + ).rejects.toThrow("socket vanished"); + }); + + it.each([ + [{ retryAfter: 5 }, { retryAfter: 5 }], + [10n, undefined], + ])("projects serializable error data %#", async (data, expectedData) => { + const client = makeClient(); + let sentId = ""; + internals(client).transport = { + send: vi.fn(async (message) => { + sentId = (message as { id: string }).id; + }), + }; + const promise = internals(client).dispatchTaskRequest({ + method: "tasks/get", + }); + await Promise.resolve(); + internals(client).consumeRawWireResponse({ + id: sentId, + error: { code: -32001, message: "task failed", data }, + }); + + await expect(promise).resolves.toEqual({ + kind: "error", + error: { + code: -32001, + message: "task failed", + ...(expectedData === undefined ? {} : { data: expectedData }), + }, + }); + }); + + it("rejects a non-JSON raw result", async () => { + const client = makeClient(); + let sentId = ""; + internals(client).transport = { + send: vi.fn(async (message) => { + sentId = (message as { id: string }).id; + }), + }; + const promise = internals(client).dispatchTaskRequest({ + method: "tasks/get", + }); + await Promise.resolve(); + internals(client).consumeRawWireResponse({ id: sentId, result: 10n }); + + await expect(promise).rejects.toThrow(/non-JSON result/); + }); + + it.each([ + [undefined, "allow", undefined], + [{ ttl: 2500 }, "prefer", 2500], + ] as const)( + "maps %s task options to the ext-tasks preference contract", + async (taskOptions, expectedPreference, expectedRetention) => { + const client = makeClient(); + internals(client).requestTimeout = 12_345; + const settle = vi.fn(async (options: TaskExecutionSettleOptions) => { + options.onEvent({ + type: "task", + task: { + taskId: "task-preference", + status: "working", + lastUpdatedAt: "2026-01-02T03:04:05.000Z", + }, + }); + options.onEvent({ + type: "outcome", + outcome: { + status: "completed", + result: successfulResult, + task: { + taskId: "task-preference", + status: "completed", + createdAt: "2026-01-02T03:04:05.000Z", + }, + }, + }); + return { + outcome: { status: "completed", result: successfulResult }, + }; + }); + const callTool = vi.fn(async () => ({ settle })); + attachTaskBoundary(client, callTool); + const updates: Array<{ + task: TaskWithOptionalCreatedAt; + result?: CallToolResult; + }> = []; + client.addEventListener("requestorTaskUpdated", (event) => { + updates.push(event.detail); + }); + + const invocation = await client.callTool( + taskTool, + {}, + undefined, + undefined, + taskOptions, + ); + + expect(invocation.result).toEqual(successfulResult); + expect(callTool).toHaveBeenCalledWith( + taskTool.name, + {}, + expect.objectContaining({ + requestTimeoutMs: 12_345, + task: { + preference: expectedPreference, + retentionMs: expectedRetention, + }, + }), + ); + expect(settle).toHaveBeenCalledWith( + expect.objectContaining({ + onEvent: expect.any(Function), + }), + ); + expect(updates).toEqual([ + { + taskId: "task-preference", + task: expect.objectContaining({ + createdAt: "2026-01-02T03:04:05.000Z", + lastUpdatedAt: "2026-01-02T03:04:05.000Z", + }), + }, + { + taskId: "task-preference", + task: expect.objectContaining({ + createdAt: "2026-01-02T03:04:05.000Z", + lastUpdatedAt: "2026-01-02T03:04:05.000Z", + }), + result: successfulResult, + }, + ]); + }, + ); + + it("delivers raw-call progress before a task snapshot and releases the token", async () => { + const client = makeClient(); + const progress: unknown[] = []; + client.addEventListener("progressNotification", (event) => { + progress.push(event.detail); + }); + let progressToken: string | number | undefined; + attachTaskBoundary( + client, + vi.fn(async (_name, _args, options) => { + progressToken = options.metadata?.progressToken as + | string + | number + | undefined; + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken, progress: 1, total: 2 }, + }); + return { + settle: vi.fn(async () => ({ + outcome: { status: "completed", result: successfulResult }, + })), + }; + }), + ); + + const invocation = await client.callTool(taskTool, {}); + + expect(invocation.metadata?.progressToken).toBe(progressToken); + expect(progress).toEqual([{ progressToken, progress: 1, total: 2 }]); + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken, progress: 2, total: 2 }, + }); + expect(progress).toHaveLength(1); + }); + + const makeBoundaryClient = ( + options?: ConstructorParameters<typeof InspectorClient>[1], + ): InspectorClient => + new InspectorClient( + { type: "stdio", command: "noop", args: [] }, + options ?? { environment: { transport: () => ({}) as never } }, + ); + + it("omits the progress token when progress is disabled", async () => { + const client = makeBoundaryClient({ + environment: { transport: () => ({}) as never }, + progress: false, + }); + let sentMetadata: Readonly<Record<string, unknown>> | undefined; + attachTaskBoundary( + client, + vi.fn(async (_name, _args, options) => { + sentMetadata = options.metadata; + return { + settle: vi.fn(async () => ({ + outcome: { status: "completed", result: successfulResult }, + })), + }; + }), + ); + + const invocation = await client.callTool(taskTool, {}); + + expect(sentMetadata?.progressToken).toBeUndefined(); + expect(invocation.metadata?.progressToken).toBeUndefined(); + }); + + it("reuses a caller progress token across concurrent raw calls", async () => { + const client = makeBoundaryClient(); + let releaseSettle: (() => void) | undefined; + const gate = new Promise<void>((resolve) => { + releaseSettle = resolve; + }); + const callTool = vi.fn(async () => ({ + settle: vi.fn(async (options: TaskExecutionSettleOptions) => { + options.onEvent({ + type: "task", + task: { + taskId: "shared-token", + status: "working", + lastUpdatedAt: "2026-01-02T03:04:05.000Z", + }, + }); + await gate; + return { + outcome: { status: "completed", result: successfulResult }, + }; + }), + })); + attachTaskBoundary(client, callTool); + const progresses: unknown[] = []; + client.addEventListener("progressNotification", (event) => { + progresses.push(event.detail); + }); + + const sharedToken = "caller-token"; + const first = client.callTool( + taskTool, + {}, + { + progressToken: sharedToken, + }, + ); + const second = client.callTool( + taskTool, + {}, + { + progressToken: sharedToken, + }, + ); + await vi.waitFor(() => expect(callTool).toHaveBeenCalledTimes(2)); + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken: sharedToken, progress: 1 }, + }); + expect(progresses).toHaveLength(1); + + releaseSettle?.(); + const [firstDone, secondDone] = await Promise.all([first, second]); + expect(firstDone.metadata?.progressToken).toBe(sharedToken); + expect(secondDone.metadata?.progressToken).toBe(sharedToken); + + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken: sharedToken, progress: 2 }, + }); + expect(progresses).toHaveLength(1); + }); + + it("routes shared-token progress to every correlated task, not only the latest", async () => { + const client = makeBoundaryClient(); + let releaseSettle: (() => void) | undefined; + const gate = new Promise<void>((resolve) => { + releaseSettle = resolve; + }); + // Each concurrent call correlates the shared token to its OWN task id. + let taskCounter = 0; + const callTool = vi.fn(async () => { + taskCounter += 1; + const taskId = `shared-token-task-${taskCounter}`; + return { + settle: vi.fn(async (options: TaskExecutionSettleOptions) => { + options.onEvent({ + type: "task", + task: { + taskId, + status: "working", + lastUpdatedAt: "2026-01-02T03:04:05.000Z", + }, + }); + await gate; + return { + outcome: { status: "completed", result: successfulResult }, + }; + }), + }; + }); + attachTaskBoundary(client, callTool); + const taskProgress: { taskId: string }[] = []; + client.addEventListener("requestorTaskProgress", (event) => { + taskProgress.push({ taskId: event.detail.taskId }); + }); + + const sharedToken = "caller-token"; + const first = client.callTool(taskTool, {}, { progressToken: sharedToken }); + const second = client.callTool( + taskTool, + {}, + { progressToken: sharedToken }, + ); + await vi.waitFor(() => expect(callTool).toHaveBeenCalledTimes(2)); + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken: sharedToken, progress: 1 }, + }); + // The wire cannot say which owner the progress belongs to, so both + // correlated tasks receive it — the earlier one is not overwritten. + expect(taskProgress.map((p) => p.taskId).sort()).toEqual([ + "shared-token-task-1", + "shared-token-task-2", + ]); + + releaseSettle?.(); + await Promise.all([first, second]); + // Both calls settled, so the correlation set is fully released. + taskProgress.length = 0; + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken: sharedToken, progress: 2 }, + }); + expect(taskProgress).toHaveLength(0); + }); + + it("resumes the created task instead of re-calling the tool on a recovery rerun", async () => { + const client = makeBoundaryClient(); + const reference = { + endpointId: "e", + generation: "v2", + taskId: "recovered-task", + originalOperation: "tools/call", + }; + const callTool = vi.fn(async () => ({ + kind: "task", + serializeReference: () => reference, + settle: vi.fn(async () => ({ + outcome: { status: "completed", result: successfulResult }, + })), + })); + const resumeTask = vi.fn(async (reference: Record<string, unknown>) => ({ + kind: "task", + // A resumed execution re-serializes to the same task identity. + serializeReference: () => reference, + settle: vi.fn(async () => ({ + outcome: { status: "completed", result: successfulResult }, + })), + })); + const boundary = taskInternals(client); + boundary.client = {}; + boundary.protocolEra = "modern"; + boundary.taskSession = { callTool, resumeTask }; + const invoke = ( + client as unknown as { + callTaskToolAndSettle: ( + tool: Tool, + args: Record<string, unknown>, + metadata: undefined, + preference: "prefer", + retentionMs: undefined, + signal: undefined, + progressToken: undefined, + recovery?: { reference?: Record<string, unknown> }, + ) => Promise<unknown>; + } + ).callTaskToolAndSettle.bind(client); + + // First attempt: no reference yet, so the tool is called and the box is + // populated the moment the task exists. + const recovery: { reference?: Record<string, unknown> } = {}; + await invoke( + taskTool, + {}, + undefined, + "prefer", + undefined, + undefined, + undefined, + recovery, + ); + expect(callTool).toHaveBeenCalledTimes(1); + expect(resumeTask).not.toHaveBeenCalled(); + expect(recovery.reference).toBe(reference); + + // Recovery rerun: the populated box routes through resumeTask, so no + // second remote task is created. + await invoke( + taskTool, + {}, + undefined, + "prefer", + undefined, + undefined, + undefined, + recovery, + ); + expect(callTool).toHaveBeenCalledTimes(1); + expect(resumeTask).toHaveBeenCalledTimes(1); + expect(resumeTask.mock.calls[0]?.[0]).toBe(reference); + }); + + it("projects a task-scoped failure onto the public task event", async () => { + const client = makeClient(); + attachTaskBoundary( + client, + vi.fn(async () => ({ + settle: vi.fn(async (options: TaskExecutionSettleOptions) => { + options.onEvent({ + type: "task", + task: { + taskId: "task-failed", + status: "working", + createdAt: "2026-01-02T03:04:05.000Z", + lastUpdatedAt: "2026-01-02T03:04:06.000Z", + }, + }); + throw new Error("worker exploded"); + }), + })), + ); + const updates: Array<{ error?: Error }> = []; + client.addEventListener("requestorTaskUpdated", (event) => { + updates.push(event.detail); + }); + + await expect(client.callTool(taskTool, {})).rejects.toThrow( + "worker exploded", + ); + expect(updates.at(-1)?.error?.message).toBe("worker exploded"); + }); + + it("enforces and can explicitly bypass output validation on task results", async () => { + const client = makeClient(); + const invalidResult: CallToolResult = { + content: [], + structuredContent: { count: "not-a-number" }, + }; + attachTaskBoundary( + client, + vi.fn(async () => ({ + settle: vi.fn(async () => ({ + outcome: { status: "completed", result: invalidResult }, + })), + })), + ); + const toolWithOutput: Tool = { + ...taskTool, + outputSchema: { + type: "object", + properties: { count: { type: "number" } }, + required: ["count"], + }, + }; + + await expect(client.callTool(toolWithOutput, {})).rejects.toThrow( + /output schema|must be number/i, + ); + const advisory = await client.callTool( + toolWithOutput, + {}, + undefined, + undefined, + undefined, + { skipOutputValidation: true }, + ); + expect(advisory.success).toBe(true); + expect(advisory.outputValidationError).toMatch( + /output schema|must be number/i, + ); + }); + + it("projects every ext-tasks event boundary", () => { + const client = makeClient(); + const boundary = taskInternals(client); + + const updates: Array<{ + task: TaskWithOptionalCreatedAt; + result?: CallToolResult; + error?: Error; + }> = []; + client.addEventListener("requestorTaskUpdated", (event) => { + updates.push(event.detail); + }); + + expect( + boundary.emitTaskExecutionEvent({ + type: "outcome", + outcome: { status: "cancelled" }, + }), + ).toBeUndefined(); + boundary.emitTaskError(undefined, "ignored without a task"); + + const emptyTimestampTask = { + taskId: "task-projection", + status: "working", + }; + expect( + boundary.emitTaskExecutionEvent({ + type: "task", + task: emptyTimestampTask, + }), + ).toEqual({ + ...emptyTimestampTask, + createdAt: "", + lastUpdatedAt: "", + }); + boundary.emitTaskExecutionEvent({ + type: "outcome", + outcome: { + status: "failed", + error: "string failure", + task: { ...emptyTimestampTask, status: "failed" }, + }, + }); + boundary.emitTaskExecutionEvent({ + type: "outcome", + outcome: { + status: "cancelled", + task: { ...emptyTimestampTask, status: "cancelled" }, + }, + }); + + expect(updates).toHaveLength(3); + expect(updates[1]?.error?.message).toBe("string failure"); + expect(updates[2]?.result).toBeUndefined(); + expect(updates[2]?.error).toBeUndefined(); + }); + + it("restores DispatchError cause identity from ext-tasks", async () => { + const client = makeClient(); + const cause = new Error("host transport failed"); + attachTaskBoundary( + client, + vi.fn(async () => ({ + settle: vi.fn(async () => { + throw new DispatchError("dispatch policy wrapper", false, { cause }); + }), + })), + ); + + await expect(client.callTool(taskTool, {})).rejects.toBe(cause); + }); + + it("preserves ext-tasks protocol error code and data", async () => { + const client = makeClient(); + const errorData = { reason: "task rejected" }; + attachTaskBoundary( + client, + vi.fn(async () => ({ + settle: vi.fn(async (options: TaskExecutionSettleOptions) => { + options.onEvent({ + type: "task", + task: { + taskId: "task-protocol-error", + status: "working", + createdAt: "2026-01-02T03:04:05.000Z", + lastUpdatedAt: "2026-01-02T03:04:05.000Z", + }, + }); + throw new JsonRpcResponseError({ + code: -32099, + message: "Task protocol failure", + data: errorData, + }); + }), + })), + ); + const taskErrors: ProtocolError[] = []; + client.addEventListener("requestorTaskUpdated", (event) => { + if (event.detail.error) taskErrors.push(event.detail.error); + }); + + const rejection = client.callTool(taskTool, {}); + + await expect(rejection).rejects.toMatchObject({ + code: -32099, + data: errorData, + }); + expect(taskErrors.at(-1)).toMatchObject({ + code: -32099, + data: errorData, + }); + }); + it("validates output from the public streaming task path", async () => { + const client = makeClient(); + const invalidResult: CallToolResult = { + content: [], + structuredContent: { count: "not-a-number" }, + }; + attachTaskBoundary( + client, + vi.fn(async () => ({ + settle: vi.fn(async () => ({ + outcome: { status: "completed", result: invalidResult }, + })), + })), + ); + const toolWithOutput: Tool = { + ...taskTool, + outputSchema: { + type: "object", + properties: { count: { type: "number" } }, + required: ["count"], + }, + }; + + await expect(client.callToolStream(toolWithOutput, {})).rejects.toThrow( + /output schema|must be number/i, + ); + const advisory = await client.callToolStream( + toolWithOutput, + {}, + undefined, + undefined, + undefined, + { skipOutputValidation: true }, + ); + expect(advisory.outputValidationError).toMatch( + /output schema|must be number/i, + ); + }); + + it("stringifies a non-Error streaming task failure", async () => { + const client = makeClient(); + attachTaskBoundary( + client, + vi.fn(async () => { + throw "worker vanished"; + }), + ); + const invocations: Array<{ error?: string }> = []; + client.addEventListener("toolCallResultChange", (event) => { + invocations.push(event.detail); + }); + + await expect(client.callToolStream(taskTool, {})).rejects.toBe( + "worker vanished", + ); + expect(invocations.at(-1)?.error).toBe("worker vanished"); + }); + it("routes owned task progress without re-emitting it when progress is off", () => { + const client = new InspectorClient( + { type: "stdio", command: "noop", args: [] }, + { environment: { transport: () => ({}) as never }, progress: false }, + ); + // Double cast: the router is a private field with no public seam, and this + // test seeds it directly so the dispatch branches run without a live call. + const router = (client as unknown as { taskProgress: TaskProgressRouter }) + .taskProgress; + router.acquire("owned"); + router.correlate("owned", "task-1"); + const taskProgress: unknown[] = []; + const generic: unknown[] = []; + client.addEventListener("requestorTaskProgress", (event) => { + taskProgress.push(event.detail); + }); + client.addEventListener("progressNotification", (event) => { + generic.push(event.detail); + }); + + // A progress frame without a token belongs to no request at all. + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progress: 1 }, + }); + taskInternals(client).dispatchTaskProgress({ + method: "notifications/progress", + params: { progressToken: "owned", progress: 2 }, + }); + + expect(generic).toEqual([]); + expect(taskProgress).toEqual([ + { taskId: "task-1", progress: { progressToken: "owned", progress: 2 } }, + ]); + }); + it("cancels only the pending input belonging to a cancelled task", () => { + const client = makeClient(); + const own = { taskId: "task-1", cancel: vi.fn() }; + const other = { taskId: "task-2", cancel: vi.fn() }; + // Double cast: the pending queues and the cancel helper are private, and + // the queue entries need only the two members the helper reads. + const pending = client as unknown as { + pendingElicitations: unknown[]; + pendingSamples: unknown[]; + cancelPendingTaskInput: (taskId: string) => void; + }; + pending.pendingElicitations = [own, other]; + pending.pendingSamples = [other]; + pending.cancelPendingTaskInput("task-1"); + expect(own.cancel).toHaveBeenCalledTimes(1); + expect(other.cancel).not.toHaveBeenCalled(); + expect(pending.pendingElicitations).toEqual([other]); + expect(pending.pendingSamples).toEqual([other]); + }); }); diff --git a/clients/web/src/test/core/mcp/inspectorClientUrlElicitation.test.ts b/clients/web/src/test/core/mcp/inspectorClientUrlElicitation.test.ts index 764c6baf2f..fe2e0bb466 100644 --- a/clients/web/src/test/core/mcp/inspectorClientUrlElicitation.test.ts +++ b/clients/web/src/test/core/mcp/inspectorClientUrlElicitation.test.ts @@ -4,6 +4,7 @@ import { ProtocolError, UrlElicitationRequiredError, } from "@modelcontextprotocol/client"; +import { JsonRpcResponseError } from "@modelcontextprotocol/ext-tasks/client"; import type { ElicitRequestURLParams, Tool, @@ -80,7 +81,20 @@ describe("InspectorClient URL-elicitation error path", () => { request: vi.fn(async () => { attempt += 1; if (attempt === 1) { - throw new UrlElicitationRequiredError([elicitation]); + throw new JsonRpcResponseError({ + code: ProtocolErrorCode.UrlElicitationRequired, + message: "Authorization required", + data: { + elicitations: [ + { + mode: elicitation.mode, + elicitationId: elicitation.elicitationId, + url: elicitation.url, + message: elicitation.message, + }, + ], + }, + }); } return okResult; }), diff --git a/clients/web/src/test/core/mcp/modernTaskSchemas.test.ts b/clients/web/src/test/core/mcp/modernTaskSchemas.test.ts deleted file mode 100644 index 2bc9020a6b..0000000000 --- a/clients/web/src/test/core/mcp/modernTaskSchemas.test.ts +++ /dev/null @@ -1,136 +0,0 @@ -import { describe, it, expect } from "vitest"; -import { - TASKS_EXTENSION_KEY, - MODERN_PROTOCOL_VERSION, - MODERN_TASK_HANDLE_META, - TASKS_EXTENSION_CLIENT_CAPABILITY, - ModernDetailedTaskSchema, - ModernGetTaskResultSchema, - ModernUpdateTaskResultSchema, - ModernCancelTaskResultSchema, - normalizeModernTask, - readInputRequests, - isModernCreateTaskResult, -} from "@inspector/core/mcp/modernTaskSchemas.js"; - -describe("modernTaskSchemas (#1631)", () => { - const baseTask = { - taskId: "abc", - status: "working" as const, - createdAt: "2026-07-20T00:00:00Z", - lastUpdatedAt: "2026-07-20T00:00:01Z", - }; - - describe("constants", () => { - it("exposes the SEP-2663 identifiers", () => { - expect(TASKS_EXTENSION_KEY).toBe("io.modelcontextprotocol/tasks"); - expect(MODERN_PROTOCOL_VERSION).toBe("2026-07-28"); - expect(MODERN_TASK_HANDLE_META).toContain("modernTaskHandle"); - expect( - TASKS_EXTENSION_CLIENT_CAPABILITY.extensions[TASKS_EXTENSION_KEY], - ).toEqual({}); - }); - }); - - describe("ModernDetailedTaskSchema", () => { - it("parses a working task and passes unknown fields through (loose)", () => { - const parsed = ModernDetailedTaskSchema.parse({ - ...baseTask, - ttlMs: 60000, - pollIntervalMs: 500, - somethingNew: "kept", - }); - expect(parsed.taskId).toBe("abc"); - expect((parsed as Record<string, unknown>).somethingNew).toBe("kept"); - }); - - it("accepts a null ttlMs and inline result/error/inputRequests", () => { - const completed = ModernGetTaskResultSchema.parse({ - ...baseTask, - status: "completed", - ttlMs: null, - result: { content: [{ type: "text", text: "done" }] }, - }); - expect(completed.result).toBeDefined(); - const failed = ModernDetailedTaskSchema.parse({ - ...baseTask, - status: "failed", - error: { code: -1, message: "boom" }, - }); - expect(failed.error).toBeDefined(); - }); - - it("rejects a task missing required identity fields", () => { - expect(() => - ModernDetailedTaskSchema.parse({ status: "working" }), - ).toThrow(); - }); - - it("accepts empty update/cancel acks", () => { - expect(ModernUpdateTaskResultSchema.parse({})).toEqual({}); - expect( - ModernCancelTaskResultSchema.parse({ resultType: "complete" }), - ).toBeDefined(); - }); - }); - - describe("normalizeModernTask", () => { - it("maps ttlMs → ttl and pollIntervalMs → pollInterval", () => { - const task = normalizeModernTask({ - ...baseTask, - ttlMs: 60000, - pollIntervalMs: 250, - }); - expect(task.ttl).toBe(60000); - expect((task as { pollInterval?: number }).pollInterval).toBe(250); - }); - - it("defaults ttl to null and omits pollInterval when absent", () => { - const task = normalizeModernTask({ ...baseTask }); - expect(task.ttl).toBeNull(); - expect((task as { pollInterval?: number }).pollInterval).toBeUndefined(); - }); - - it("carries the status-specific members structurally", () => { - const task = normalizeModernTask({ - ...baseTask, - status: "completed", - ttlMs: null, - result: { content: [] }, - }); - expect((task as { result?: unknown }).result).toEqual({ content: [] }); - }); - }); - - describe("readInputRequests", () => { - it("returns the inputRequests map when present", () => { - const requests = { confirm: { method: "elicitation/create" } }; - const out = readInputRequests({ - ...baseTask, - status: "input_required", - inputRequests: requests, - }); - expect(out).toBe(requests); - }); - - it("returns undefined when absent", () => { - expect(readInputRequests({ ...baseTask })).toBeUndefined(); - }); - }); - - describe("isModernCreateTaskResult", () => { - it("is true only for a resultType:task frame with a taskId", () => { - expect( - isModernCreateTaskResult({ resultType: "task", taskId: "x" }), - ).toBe(true); - }); - - it("is false for complete results, non-objects, and missing taskId", () => { - expect(isModernCreateTaskResult({ resultType: "complete" })).toBe(false); - expect(isModernCreateTaskResult({ resultType: "task" })).toBe(false); - expect(isModernCreateTaskResult({ content: [] })).toBe(false); - expect(isModernCreateTaskResult(null)).toBe(false); - expect(isModernCreateTaskResult("task")).toBe(false); - }); - }); -}); diff --git a/clients/web/src/test/core/mcp/node/config.test.ts b/clients/web/src/test/core/mcp/node/config.test.ts index 7e1e1689ec..d107e985b8 100644 --- a/clients/web/src/test/core/mcp/node/config.test.ts +++ b/clients/web/src/test/core/mcp/node/config.test.ts @@ -13,6 +13,7 @@ import { parseKeyValuePair, parseHeaderPair, parseProtocolEra, + skillCatalogLimitParser, withDefaultCatalogPath, resolveServerConfigs, resolveServerSource, @@ -57,6 +58,28 @@ describe("parseProtocolEra", () => { }); }); +describe("skillCatalogLimitParser", () => { + const parse = skillCatalogLimitParser("--skill-catalog-max-skills"); + + it.each([ + ["1", 1], + ["256", 256], + [" 42 ", 42], + ["67108864", 67108864], + ])("accepts %j as %d", (value, expected) => { + expect(parse(value)).toBe(expected); + }); + + it.each(["0", "-1", "1.5", "1e3", "0x10", "", "abc", "9007199254740993"])( + "rejects %j, naming the flag", + (value) => { + expect(() => parse(value)).toThrow( + `Invalid --skill-catalog-max-skills: ${value}. Expected a positive integer written as plain decimal digits, at most 9007199254740991.`, + ); + }, + ); +}); + describe("parseHeaderPair", () => { it("splits on the first colon and trims both sides", () => { expect(parseHeaderPair("X-Test: value")).toEqual({ @@ -588,6 +611,45 @@ describe("resolveServerConfigs — single mode", () => { ).toThrow(/Server 'bar' not found/); }); + // #2537: the lookup must be an own-property read, so an absent server whose + // name is an inherited Object.prototype member reports "not found" instead + // of resolving to the inherited function and failing downstream. + it.each(["constructor", "toString", "hasOwnProperty", "__proto__"])( + "reports an absent server named %s as not found", + (serverName) => { + writeFileSync( + configPath, + JSON.stringify({ mcpServers: { real: { command: "node" } } }), + ); + expect(() => + resolveServerConfigs({ configPath, serverName }, "single"), + ).toThrow( + `Server '${serverName}' not found in config file. Available servers: real`, + ); + }, + ); + + it("resolves a server genuinely named constructor", () => { + writeFileSync( + configPath, + JSON.stringify({ + mcpServers: { + constructor: { command: "node", args: ["ctor.js"] }, + real: { command: "node" }, + }, + }), + ); + const [config] = resolveServerConfigs( + { configPath, serverName: "constructor" }, + "single", + ); + expect(config).toEqual({ + type: "stdio", + command: "node", + args: ["ctor.js"], + }); + }); + it("applies env/cwd overrides when loading from config", () => { writeFileSync( configPath, diff --git a/clients/web/src/test/core/mcp/node/servers.test.ts b/clients/web/src/test/core/mcp/node/servers.test.ts index 3cd372d8d7..36b293db5c 100644 --- a/clients/web/src/test/core/mcp/node/servers.test.ts +++ b/clients/web/src/test/core/mcp/node/servers.test.ts @@ -297,6 +297,74 @@ describe("loadServerEntries", () => { }); }); + it("overrides the disk skills catalog budget with --skill-catalog-max-*, preserving the rest", async () => { + const configPath = join(tempDir, "mcp.json"); + writeFileSync( + configPath, + JSON.stringify({ + mcpServers: { + web: { + type: "streamable-http", + url: "http://x/mcp", + skillCatalogMaxSkills: 10, + skillCatalogMaxBytes: 2048, + requestTimeout: 9000, + }, + }, + }), + ); + + const fromFile = await loadServerEntries({ configPath }); + expect(fromFile.web?.settings).toMatchObject({ + skillCatalogMaxSkills: 10, + skillCatalogMaxBytes: 2048, + }); + + const servers = await loadServerEntries({ + configPath, + skillCatalogMaxSkills: 3, + skillCatalogMaxBytes: 4096, + }); + expect(servers.web?.settings).toMatchObject({ + skillCatalogMaxSkills: 3, + skillCatalogMaxBytes: 4096, + requestTimeout: 9000, + }); + + const onlyBytes = await loadServerEntries({ + configPath, + skillCatalogMaxBytes: 1, + }); + expect(onlyBytes.web?.settings).toMatchObject({ + skillCatalogMaxSkills: 10, + skillCatalogMaxBytes: 1, + }); + }); + + it("applies --skill-catalog-max-* to an ad-hoc URL with no other settings", async () => { + const servers = await loadServerEntries({ + serverUrl: "http://x/mcp", + transport: "http", + skillCatalogMaxSkills: 5, + }); + expect(servers.default?.settings).toMatchObject({ + skillCatalogMaxSkills: 5, + headers: [], + connectionTimeout: DEFAULT_CONNECTION_TIMEOUT_MS, + }); + expect(servers.default?.settings?.skillCatalogMaxBytes).toBeUndefined(); + + const bytesOnly = await loadServerEntries({ + serverUrl: "http://x/mcp", + transport: "http", + skillCatalogMaxBytes: 7, + }); + expect(bytesOnly.default?.settings).toMatchObject({ + skillCatalogMaxBytes: 7, + headers: [], + }); + }); + it("gives an ad-hoc target with neither flag no settings (legacy default)", async () => { const servers = await loadServerEntries({ target: ["my-server"] }); expect(servers.default?.settings).toBeUndefined(); @@ -337,6 +405,22 @@ describe("selectServerEntry", () => { ); }); + // #2537: `--server constructor` against a source without that server used to + // return the inherited Object.prototype member and fail with an unrelated + // error; it must report "not found" like any other absent name. + it.each(["constructor", "toString", "hasOwnProperty", "__proto__"])( + "reports an absent entry named %s as not found", + (name) => { + expect(() => selectServerEntry({ a, b }, name)).toThrow( + `Server '${name}' not found. Available servers: a, b`, + ); + }, + ); + + it("returns an entry genuinely named constructor", () => { + expect(selectServerEntry({ constructor: a, b }, "constructor")).toBe(a); + }); + it("returns the only entry when no name is given", () => { expect(selectServerEntry({ a })).toBe(a); }); diff --git a/clients/web/src/test/core/mcp/resultFile.test.ts b/clients/web/src/test/core/mcp/resultFile.test.ts new file mode 100644 index 0000000000..61592fadba --- /dev/null +++ b/clients/web/src/test/core/mcp/resultFile.test.ts @@ -0,0 +1,120 @@ +import { describe, it, expect } from "vitest"; +import { + RawResultUnavailableError, + RESULT_FILE_FORMATS, + renderRawResult, + renderResultForFile, + renderedByteLength, +} from "@inspector/core/mcp/resultFile"; + +const PNG_BYTES = [0x89, 0x50, 0x4e, 0x47]; +const PNG_B64 = btoa(String.fromCharCode(...PNG_BYTES)); + +describe("renderResultForFile", () => { + it("lists both formats", () => { + expect(RESULT_FILE_FORMATS).toEqual(["raw", "json"]); + }); + + it("json is the whole result, two-space indented, newline-terminated", () => { + const result = { content: [{ type: "text", text: "hi" }] }; + expect(renderResultForFile(result, "tools/call", "json")).toBe( + JSON.stringify(result, null, 2) + "\n", + ); + }); + + it("raw delegates to the raw renderer", () => { + expect( + renderResultForFile( + { content: [{ type: "text", text: "hi" }] }, + "tools/call", + "raw", + ), + ).toBe("hi"); + }); +}); + +describe("renderRawResult", () => { + it("joins the text of every text-bearing tool block, embedded resources included", () => { + const result = { + content: [ + { type: "text", text: "one" }, + { type: "image", data: PNG_B64, mimeType: "image/png" }, + { type: "resource", resource: { uri: "x://a", text: "two" } }, + { type: "resource_link", uri: "x://b", name: "b" }, + "not a block", + null, + ], + }; + expect(renderRawResult(result, "tools/call")).toBe("one\ntwo"); + }); + + it("decodes a single binary tool block when there is no text", () => { + const out = renderRawResult( + { content: [{ type: "image", data: PNG_B64, mimeType: "image/png" }] }, + "tools/call", + ); + expect(Array.from(out as Uint8Array)).toEqual(PNG_BYTES); + }); + + it("decodes a single blob from a resources/read result", () => { + const out = renderRawResult( + { contents: [{ uri: "x://a", blob: PNG_B64 }] }, + "resources/read", + ); + expect(Array.from(out as Uint8Array)).toEqual(PNG_BYTES); + }); + + it("reads resources/read text from `contents`, not `content`", () => { + expect( + renderRawResult( + { + contents: [{ uri: "x://a", text: "r" }], + content: [{ text: "ignored" }], + }, + "resources/read", + ), + ).toBe("r"); + }); + + it("refuses a result with no payload", () => { + expect(() => renderRawResult({ content: [] }, "tools/call")).toThrow( + new RawResultUnavailableError("tools/call", 0), + ); + expect(() => renderRawResult(null, "tools/call")).toThrow( + /no text or binary content/, + ); + expect(() => + renderRawResult({ content: "not a list" }, "tools/call"), + ).toThrow(RawResultUnavailableError); + }); + + it("refuses several binaries with no text", () => { + let caught: unknown; + try { + renderRawResult( + { + content: [ + { type: "image", data: PNG_B64 }, + { type: "audio", data: PNG_B64 }, + ], + }, + "tools/call", + ); + } catch (err) { + caught = err; + } + expect(caught).toBeInstanceOf(RawResultUnavailableError); + const err = caught as RawResultUnavailableError; + expect(err.name).toBe("RawResultUnavailableError"); + expect(err.binaryCount).toBe(2); + expect(err.method).toBe("tools/call"); + expect(err.message).toMatch(/2 binary blocks and no text/); + }); +}); + +describe("renderedByteLength", () => { + it("counts UTF-8 bytes for strings and length for bytes", () => { + expect(renderedByteLength("é")).toBe(2); + expect(renderedByteLength(new Uint8Array(3))).toBe(3); + }); +}); diff --git a/clients/web/src/test/core/mcp/serverList.test.ts b/clients/web/src/test/core/mcp/serverList.test.ts index d660ecd00e..6dd9dc9782 100644 --- a/clients/web/src/test/core/mcp/serverList.test.ts +++ b/clients/web/src/test/core/mcp/serverList.test.ts @@ -610,6 +610,60 @@ describe("serverEntriesToMcpConfig", () => { expect("protocolEra" in (round.mcpServers["era-legacy"] ?? {})).toBe(false); }); + it("round-trips elicitCapability: lifts a non-default value to settings and back to disk (#1783)", () => { + const original: MCPConfig = { + mcpServers: { + "elicit-url": { + type: "streamable-http", + url: "https://x.test/mcp", + elicitCapability: "url", + }, + }, + }; + const [entry] = mcpConfigToServerEntries(original); + expect(entry?.settings?.elicitCapability).toBe("url"); + const round = serverEntriesToMcpConfig(mcpConfigToServerEntries(original)); + expect(round).toEqual(original); + }); + + it("drops an unknown elicitCapability literal on read (hand-edited file)", () => { + // Like protocolEra: garbage from a hand-edited mcp.json is dropped here + // rather than reaching the connect-time capability wiring. + const badElicit: object = { elicitCapability: "everything" }; + const original: MCPConfig = { + mcpServers: { + "elicit-bad": { + type: "streamable-http", + url: "https://x.test/mcp", + ...badElicit, + }, + }, + }; + const [entry] = mcpConfigToServerEntries(original); + expect(entry?.settings?.elicitCapability).toBeUndefined(); + }); + + it("omits elicitCapability from disk when it equals the default (both)", () => { + // "both" is the default — writing it back must NOT inject the field. + // A benign inspector field keeps `settings` materialized. + const original: MCPConfig = { + mcpServers: { + "elicit-both": { + type: "streamable-http", + url: "https://x.test/mcp", + elicitCapability: "both", + connectionTimeout: 5000, + }, + }, + }; + const [entry] = mcpConfigToServerEntries(original); + expect(entry?.settings?.elicitCapability).toBe("both"); + const round = serverEntriesToMcpConfig(mcpConfigToServerEntries(original)); + expect("elicitCapability" in (round.mcpServers["elicit-both"] ?? {})).toBe( + false, + ); + }); + it("round-trips modernLogLevel: lifts a non-default value to settings and back to disk (#1629)", () => { const original: MCPConfig = { mcpServers: { diff --git a/clients/web/src/test/core/mcp/skillsVerification.test.ts b/clients/web/src/test/core/mcp/skillsVerification.test.ts index ca4eda44df..5c848244a7 100644 --- a/clients/web/src/test/core/mcp/skillsVerification.test.ts +++ b/clients/web/src/test/core/mcp/skillsVerification.test.ts @@ -793,6 +793,13 @@ describe("verifySkills (#2248)", () => { for (const report of past) { expect(report.outcome).toBe("incomplete"); expect(report.incomplete).toMatch(/catalog budget/); + // The escape hatch carries the `--verify` that produces a verdict — + // `--uri` alone fetches the skill and checks nothing (#2428) — and a + // placeholder rather than the server-controlled URI itself. + expect(report.incomplete).toContain( + "`--method skills/get --uri <uri> --verify`", + ); + expect(report.incomplete).not.toContain(report.uri); expect(report.files).toHaveLength(0); } expect(reports[0].outcome).toBe("verified"); diff --git a/clients/web/src/test/core/mcp/state/managedRequestorTasksState.test.ts b/clients/web/src/test/core/mcp/state/managedRequestorTasksState.test.ts index 349e5296d1..9d18f11457 100644 --- a/clients/web/src/test/core/mcp/state/managedRequestorTasksState.test.ts +++ b/clients/web/src/test/core/mcp/state/managedRequestorTasksState.test.ts @@ -23,10 +23,19 @@ describe("ManagedRequestorTasksState", () => { let state: ManagedRequestorTasksState; beforeEach(() => { - // Default to a server that advertises `tasks` so the existing flow tests - // exercise the live `listRequestorTasks` path; capability-absent tests - // below override this. + // Default to a session whose inventory is server-list so the existing + // flow tests exercise the live `listRequestorTasks` path; the refresh + // gate reads the authoritative task-session capabilities (round-12 + // finding), so the fake presents those rather than the legacy + // `capabilities.tasks` shape. Capability-absent tests below override it. client = new FakeInspectorClient({ capabilities: { tasks: {} } }); + client.taskSessionCapabilities = { + inventory: "server-list", + execution: true, + cancellation: true, + inputResponses: false, + requestedRetention: true, + }; state = new ManagedRequestorTasksState(client); }); @@ -77,6 +86,39 @@ describe("ManagedRequestorTasksState", () => { expect(tasklessState.getTasks()).toEqual([]); }); + it("routes a known-handles legacy session to re-polling, never tasks/list", async () => { + // A legacy Tasks session without `list` capability has inventory + // "known-handles": ext-tasks rejects listTasks() there, so refresh must + // re-poll known handles like a modern session (round-12 finding). + const knownHandles = new FakeInspectorClient({ + capabilities: { tasks: { cancel: {} } }, + }); + knownHandles.taskSessionCapabilities = { + inventory: "known-handles", + execution: true, + cancellation: true, + inputResponses: false, + requestedRetention: true, + }; + knownHandles.setStatus("connected"); + const knownHandlesState = new ManagedRequestorTasksState(knownHandles); + + // Seed a known task first: with an empty store this test would pass even + // if the branch returned immediately, and re-polling known handles is the + // behavior this branch exists to protect (round-15 finding). + const changePromise = waitForChange(knownHandlesState); + knownHandles.dispatchTypedEvent("taskStatusChange", { + taskId: "legacy-1", + task: task("legacy-1", "working"), + }); + await changePromise; + + const result = await knownHandlesState.refresh(); + expect(result.map((t) => t.taskId)).toEqual(["legacy-1"]); + expect(knownHandles.getRequestorTask).toHaveBeenCalledWith("legacy-1"); + expect(knownHandles.listRequestorTasks).not.toHaveBeenCalled(); + }); + it("refresh fetches a single page and dispatches tasksChange", async () => { client.setStatus("connected"); client.queueTaskPages({ tasks: [task("t1"), task("t2")] }); @@ -140,6 +182,21 @@ describe("ManagedRequestorTasksState", () => { expect(next.map((t) => t.taskId)).toEqual(["t1", "t2"]); }); + it("owns a rejected automatic refresh instead of leaking it (#2432)", async () => { + // A leaked rejection fails this run as an unhandled error; the explicit + // refresh afterwards proves the store still works and the list held. + client.setStatus("connected"); + client.listRequestorTasks.mockRejectedValueOnce(new Error("tasks/list")); + client.listRequestorTasks.mockRejectedValueOnce(new Error("tasks/list")); + client.dispatchTypedEvent("connect"); + client.dispatchTypedEvent("tasksListChanged"); + await new Promise((resolve) => setTimeout(resolve, 0)); + expect(client.listRequestorTasks).toHaveBeenCalledTimes(2); + expect(state.getTasks()).toEqual([]); + client.queueTaskPages({ tasks: [task("t1")] }); + await expect(state.refresh()).resolves.toHaveLength(1); + }); + it("statusChange to disconnected clears tasks and dispatches tasksChange", async () => { client.setStatus("connected"); client.queueTaskPages({ tasks: [task("t1")] }); diff --git a/clients/cli/__tests__/open-url.test.ts b/clients/web/src/test/core/node/openUrl.test.ts similarity index 85% rename from clients/cli/__tests__/open-url.test.ts rename to clients/web/src/test/core/node/openUrl.test.ts index c7cbe85247..1fbe59da9a 100644 --- a/clients/cli/__tests__/open-url.test.ts +++ b/clients/web/src/test/core/node/openUrl.test.ts @@ -22,9 +22,6 @@ vi.mock("open", () => ({ default: (...args: unknown[]) => openMock(...args), })); -// Suite-wide setup mocks open-url; this file exercises the real wrapper. -vi.unmock("../src/open-url.js"); - describe("openUrl", () => { beforeEach(() => { openMock.mockClear(); @@ -36,13 +33,13 @@ describe("openUrl", () => { }); it("forwards a string URL to open", async () => { - const { openUrl } = await import("../src/open-url.js"); + const { openUrl } = await import("@inspector/core/node/openUrl.js"); await openUrl("https://example.com/auth"); expect(openMock).toHaveBeenCalledWith("https://example.com/auth"); }); it("forwards URL.href for URL instances", async () => { - const { openUrl } = await import("../src/open-url.js"); + const { openUrl } = await import("@inspector/core/node/openUrl.js"); await openUrl(new URL("https://example.com/callback?code=1")); expect(openMock).toHaveBeenCalledWith( "https://example.com/callback?code=1", @@ -51,7 +48,7 @@ describe("openUrl", () => { it("rejects when the opener rejects", async () => { openMock.mockRejectedValue(new Error("xdg-open not found")); - const { openUrl } = await import("../src/open-url.js"); + const { openUrl } = await import("@inspector/core/node/openUrl.js"); await expect(openUrl("https://example.com/auth")).rejects.toThrow( "xdg-open not found", ); @@ -62,7 +59,7 @@ describe("openUrl", () => { code: "ENOENT", }); openMock.mockImplementation(async () => fakeChild(enoent)); - const { openUrl } = await import("../src/open-url.js"); + const { openUrl } = await import("@inspector/core/node/openUrl.js"); await expect(openUrl("https://example.com/auth")).rejects.toThrow( "spawn open ENOENT", ); @@ -76,7 +73,7 @@ describe("openUrl", () => { code: "ENOENT", }); openMock.mockImplementation(async () => fakeChild(enoent)); - const { openUrl } = await import("../src/open-url.js"); + const { openUrl } = await import("@inspector/core/node/openUrl.js"); const outcome = await new Promise<unknown>((resolve) => { setImmediate(() => { openUrl("https://example.com/auth").then( @@ -91,7 +88,7 @@ describe("openUrl", () => { it("absorbs an opener error that arrives after launch", async () => { let child: EventEmitter | undefined; openMock.mockImplementation(async () => (child = fakeChild())); - const { openUrl } = await import("../src/open-url.js"); + const { openUrl } = await import("@inspector/core/node/openUrl.js"); await openUrl("https://example.com/auth"); expect(() => child!.emit("error", new Error("late"))).not.toThrow(); }); @@ -99,7 +96,7 @@ describe("openUrl", () => { it("rejects when the opener does not settle within the timeout", async () => { vi.useFakeTimers(); openMock.mockReturnValue(new Promise(() => {})); - const { openUrl } = await import("../src/open-url.js"); + const { openUrl } = await import("@inspector/core/node/openUrl.js"); const pending = openUrl("https://example.com/auth", 2_000); const assertion = expect(pending).rejects.toThrow( "browser did not open within 2s", @@ -110,7 +107,8 @@ describe("openUrl", () => { it("clears its timer once the opener resolves", async () => { vi.useFakeTimers(); - const { openUrl, OPEN_URL_TIMEOUT_MS } = await import("../src/open-url.js"); + const { openUrl, OPEN_URL_TIMEOUT_MS } = + await import("@inspector/core/node/openUrl.js"); await openUrl("https://example.com/auth"); expect(vi.getTimerCount()).toBe(0); expect(OPEN_URL_TIMEOUT_MS).toBe(5_000); diff --git a/clients/web/src/test/core/react/useEmaIdpLoginState.test.tsx b/clients/web/src/test/core/react/useEmaIdpLoginState.test.tsx index d346677d17..e1e502fc74 100644 --- a/clients/web/src/test/core/react/useEmaIdpLoginState.test.tsx +++ b/clients/web/src/test/core/react/useEmaIdpLoginState.test.tsx @@ -10,6 +10,9 @@ describe("useEmaIdpLoginState", () => { storage = { load: vi.fn().mockResolvedValue(undefined), getIdpSession: vi.fn().mockResolvedValue(undefined), + // clearEmaIdpSession reads the cached IdP metadata (for the end-session + // URL) before clearing; absent here, since these tests assert the clears. + getServerMetadata: vi.fn().mockResolvedValue(null), clearIdpSession: vi.fn().mockResolvedValue(undefined), clear: vi.fn().mockResolvedValue(undefined), clearEnterpriseManagedResourceServers: vi diff --git a/clients/web/src/test/core/react/useManagedRequestorTasks.test.tsx b/clients/web/src/test/core/react/useManagedRequestorTasks.test.tsx index 3260b255f1..ea6177a2ab 100644 --- a/clients/web/src/test/core/react/useManagedRequestorTasks.test.tsx +++ b/clients/web/src/test/core/react/useManagedRequestorTasks.test.tsx @@ -20,12 +20,20 @@ describe("useManagedRequestorTasks", () => { let state: ManagedRequestorTasksState; beforeEach(() => { - // Capabilities include `tasks` so refresh() reaches the live - // listRequestorTasks path; the state manager gates on capability. + // The fake presents a server-list task session so refresh() reaches the + // live listRequestorTasks path; the state manager routes on the + // authoritative task-session inventory (round-12). client = new FakeInspectorClient({ status: "connected", capabilities: { tasks: {} }, }); + client.taskSessionCapabilities = { + inventory: "server-list", + execution: true, + cancellation: true, + inputResponses: false, + requestedRetention: true, + }; state = new ManagedRequestorTasksState(client); }); diff --git a/clients/web/src/test/integration/auth/node/file-lock.test.ts b/clients/web/src/test/integration/auth/node/file-lock.test.ts index 5dccf460eb..c74b646c7d 100644 --- a/clients/web/src/test/integration/auth/node/file-lock.test.ts +++ b/clients/web/src/test/integration/auth/node/file-lock.test.ts @@ -184,11 +184,15 @@ describe("withSecretFileLock across processes", () => { await expect(fs.stat(target)).rejects.toThrow(); let ran = false; - await withSecretFileLock(target, async () => { + let sawLocked: boolean | undefined; + await withSecretFileLock(target, async (locked) => { ran = true; + sawLocked = locked; }); expect(ran).toBe(true); + // The body is told it holds a real lock (adoption gates on this). + expect(sawLocked).toBe(true); expect(warnings()).toBe(""); }); @@ -271,14 +275,20 @@ describe("withSecretFileLock degrades rather than failing", () => { const target = path.join(tmpDir, "not-a-dir", "secrets.json"); let ran = 0; - await withSecretFileLock(target, async () => { + const sawLocked: boolean[] = []; + await withSecretFileLock(target, async (locked) => { ran += 1; + sawLocked.push(locked); }); - await withSecretFileLock(target, async () => { + await withSecretFileLock(target, async (locked) => { ran += 1; + sawLocked.push(locked); }); expect(ran).toBe(2); + // The body is told the run is unlocked, so work that is only safe + // under real exclusion (legacy adoption) can refuse instead of racing. + expect(sawLocked).toEqual([false, false]); expect(warnings()).toContain("Could not take a lock on the secrets file"); // Once per reason per process — a warning on every save would be noise // on precisely the deployment that cannot act on it. diff --git a/clients/web/src/test/integration/auth/node/oauth-adoption-locking.test.ts b/clients/web/src/test/integration/auth/node/oauth-adoption-locking.test.ts new file mode 100644 index 0000000000..d298f8e518 --- /dev/null +++ b/clients/web/src/test/integration/auth/node/oauth-adoption-locking.test.ts @@ -0,0 +1,208 @@ +/** + * Degraded-lock gate on legacy namespace adoption (#2549): migrating a + * legacy file's secret-store entries deletes its sources, so it is only + * safe under the real cross-process file lock — two unlocked adopters can + * each observe the other's half-finished move and strand credentials. + * These tests mock `withSecretFileLock` to simulate the degraded + * (unlocked) run and assert the save refuses the migration before + * touching anything, while mint-only adoption (fresh file) and + * already-stamped files keep saving unlocked. + */ + +import { describe, it, expect, beforeEach, afterEach, vi } from "vitest"; +import { mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +// Every lock in this suite degrades: the body runs, told it is unlocked. +vi.mock("@inspector/core/auth/node/file-lock.js", async (importOriginal) => { + const actual = + await importOriginal< + typeof import("@inspector/core/auth/node/file-lock.js") + >(); + return { + ...actual, + withSecretFileLock: async <T>( + _filePath: string, + fn: (locked: boolean) => Promise<T>, + ): Promise<T> => fn(false), + }; +}); + +import { + writeOAuthSections, + removeOAuthStore, + SECRETS_NAMESPACE_KEY, +} from "@inspector/core/auth/node/oauth-persist-file.js"; +import { recordNamespaceKeys } from "@inspector/core/auth/node/oauth-namespace-ledger.js"; +import { + InMemorySecretStore, + SecretStoreUnavailableError, +} from "@inspector/core/auth/node/secret-store.js"; +import { + PERSIST_TOKENS_ENV, + oauthSecretServerId, + isValidSecretsNamespace, + LEGACY_TOKENS_FIELD, + resetPersistTokensPolicyWarnings, +} from "@inspector/core/auth/node/oauth-secrets.js"; +import { + writeStoreFile, + flushStoreFileWrites, +} from "@inspector/core/storage/store-io.js"; +import type { OAuthPersistSnapshot } from "@inspector/core/auth/oauth-persist.js"; + +const SERVER = "https://api.example/mcp"; +const TOKENS = { + access_token: "at-legacy", + token_type: "Bearer", + refresh_token: "rt-legacy", +}; + +function snapshotFor(tag: string): OAuthPersistSnapshot { + return { + servers: { + [SERVER]: { + scope: "read", + tokens: { access_token: `at-${tag}`, token_type: "Bearer" }, + }, + }, + idpSessions: {}, + }; +} + +function fileNamespace(filePath: string): string { + const parsed = JSON.parse(readFileSync(filePath, "utf8")) as Record< + string, + unknown + >; + return parsed[SECRETS_NAMESPACE_KEY] as string; +} + +let tempDir: string; +let filePath: string; +let store: InMemorySecretStore; +let savedPolicy: string | undefined; + +beforeEach(() => { + tempDir = mkdtempSync(join(tmpdir(), "inspector-oauth-adopt-lock-")); + filePath = join(tempDir, "oauth.json"); + store = new InMemorySecretStore(); + savedPolicy = process.env[PERSIST_TOKENS_ENV]; + delete process.env[PERSIST_TOKENS_ENV]; +}); + +afterEach(() => { + if (savedPolicy === undefined) delete process.env[PERSIST_TOKENS_ENV]; + else process.env[PERSIST_TOKENS_ENV] = savedPolicy; + resetPersistTokensPolicyWarnings(); + rmSync(tempDir, { recursive: true, force: true }); +}); + +describe("legacy adoption under a degraded (unlocked) file lock", () => { + it("refuses the migration and leaves the file and legacy entries untouched", async () => { + const legacyBlob = JSON.stringify({ + servers: { [SERVER]: { scope: "read" } }, + idpSessions: {}, + }); + await writeStoreFile(filePath, legacyBlob); + await flushStoreFileWrites(filePath); + await store.set( + oauthSecretServerId(SERVER), + LEGACY_TOKENS_FIELD, + JSON.stringify(TOKENS), + ); + + await expect( + writeOAuthSections(filePath, snapshotFor("new"), undefined, store), + ).rejects.toThrow(SecretStoreUnavailableError); + + // Nothing moved, nothing stamped: the legacy entry still resolves and + // the file carries no namespace, so a locked retry migrates cleanly. + expect(readFileSync(filePath, "utf8")).toBe(legacyBlob); + expect( + await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(TOKENS)); + }); + + it("still mints for a fresh file — nothing to migrate, nothing to lose", async () => { + await writeOAuthSections(filePath, snapshotFor("fresh"), undefined, store); + await flushStoreFileWrites(filePath); + + const ns = fileNamespace(filePath); + expect(isValidSecretsNamespace(ns)).toBe(true); + expect( + await store.get(oauthSecretServerId(SERVER, ns), LEGACY_TOKENS_FIELD), + ).not.toBeNull(); + }); + + it("still mints for a recognized but entry-less legacy file", async () => { + // `{ servers: {}, idpSessions: {} }` indexes no store ids, so there is + // no destructive race to guard — refusing it would leave users on + // lock-hostile filesystems unable to save forever. + await writeStoreFile( + filePath, + JSON.stringify({ servers: {}, idpSessions: {} }), + ); + await flushStoreFileWrites(filePath); + + await writeOAuthSections(filePath, snapshotFor("empty"), undefined, store); + await flushStoreFileWrites(filePath); + + const ns = fileNamespace(filePath); + expect(isValidSecretsNamespace(ns)).toBe(true); + expect( + await store.get(oauthSecretServerId(SERVER, ns), LEGACY_TOKENS_FIELD), + ).not.toBeNull(); + }); + + it("still saves against an already-stamped file under its namespace", async () => { + await writeOAuthSections(filePath, snapshotFor("first"), undefined, store); + await flushStoreFileWrites(filePath); + const ns = fileNamespace(filePath); + + await writeOAuthSections(filePath, snapshotFor("second"), undefined, store); + await flushStoreFileWrites(filePath); + + expect(fileNamespace(filePath)).toBe(ns); + expect( + JSON.parse( + (await store.get( + oauthSecretServerId(SERVER, ns), + LEGACY_TOKENS_FIELD, + ))!, + ), + ).toMatchObject({ access_token: "at-second" }); + }); +}); + +describe("namespace-ledger purge under a degraded (unlocked) file lock (#2560)", () => { + const OTHER_NS = "11111111-2222-4333-8444-555555555555"; + + /** A recorded namespace holding a live entry, as a racing adopter leaves it. */ + async function seedRecordedNamespace(): Promise<string> { + await recordNamespaceKeys(filePath, store, OTHER_NS, [SERVER], []); + const id = oauthSecretServerId(SERVER, OTHER_NS); + await store.set(id, LEGACY_TOKENS_FIELD, JSON.stringify(TOKENS)); + return id; + } + + it("a mint-only adoption does not purge recorded namespaces", async () => { + // Unlocked, a recorded-but-unstamped namespace may be a concurrent + // adopter's, so its entries must survive. + const id = await seedRecordedNamespace(); + await writeOAuthSections(filePath, snapshotFor("fresh"), undefined, store); + await flushStoreFileWrites(filePath); + expect(await store.get(id, LEGACY_TOKENS_FIELD)).toBe( + JSON.stringify(TOKENS), + ); + }); + + it("an unlocked removal does not purge recorded namespaces", async () => { + const id = await seedRecordedNamespace(); + await removeOAuthStore(filePath, store); + expect(await store.get(id, LEGACY_TOKENS_FIELD)).toBe( + JSON.stringify(TOKENS), + ); + }); +}); diff --git a/clients/web/src/test/integration/auth/node/oauth-secrets-namespace.test.ts b/clients/web/src/test/integration/auth/node/oauth-secrets-namespace.test.ts new file mode 100644 index 0000000000..3dda230dbe --- /dev/null +++ b/clients/web/src/test/integration/auth/node/oauth-secrets-namespace.test.ts @@ -0,0 +1,847 @@ +/** + * Integration tests for the per-state-file secrets namespace (#2549): + * isolation between two state files sharing a server, adoption of a legacy + * (un-namespaced) file's store entries, fresh-file minting, invalid + * namespaces, adoption rollback, and keyring-backend isolation (the issue's + * acceptance criterion — exercised via the mocked `@napi-rs/keyring` + * bindings, the same pattern as secret-store.test.ts, since the real + * keychain is unreachable in CI). + */ + +import { describe, it, expect, beforeEach, afterEach, vi } from "vitest"; +import { existsSync, mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +// Hoisted above the secret-store import so KeyringSecretStore binds to the +// in-memory fake instead of the native module (absent in CI). +const keyringMocks = vi.hoisted(() => { + const password = new Map<string, string | null>(); + class AsyncEntry { + private readonly key: string; + constructor(_service: string, username: string) { + this.key = username; + } + async getPassword(): Promise<string | undefined> { + const v = password.get(this.key); + return v === undefined || v === null ? undefined : v; + } + async setPassword(value: string): Promise<void> { + password.set(this.key, value); + } + async deleteCredential(): Promise<boolean> { + return password.delete(this.key); + } + } + const findCredentialsAsync = async (): Promise< + Array<{ account: string; password: string }> + > => { + const out: Array<{ account: string; password: string }> = []; + for (const [k, v] of password.entries()) { + if (v !== null) out.push({ account: k, password: v }); + } + return out; + }; + return { AsyncEntry, findCredentialsAsync, password }; +}); + +vi.mock("@napi-rs/keyring", () => ({ + AsyncEntry: keyringMocks.AsyncEntry, + findCredentialsAsync: keyringMocks.findCredentialsAsync, +})); + +// Passthrough mock so one test can make adoption's stamp write fail at the +// commit point; every other call runs the real implementation. +vi.mock("@inspector/core/storage/store-io.js", async (importOriginal) => { + const actual = + await importOriginal< + typeof import("@inspector/core/storage/store-io.js") + >(); + return { + ...actual, + writeStoreFile: vi.fn(actual.writeStoreFile), + }; +}); + +import { + writeOAuthSections, + readOAuthStore, + removeOAuthStore, + resetOAuthSecretStoreWarnings, + SECRETS_NAMESPACE_KEY, +} from "@inspector/core/auth/node/oauth-persist-file.js"; +import { + InMemorySecretStore, + KeyringSecretStore, + type SecretStore, +} from "@inspector/core/auth/node/secret-store.js"; +import { FileSecretStore } from "@inspector/core/auth/node/file-secret-store.js"; +import { + absorbFileSecretsIntoKeyring, + SECRET_FILE_ENV, +} from "@inspector/core/auth/node/secret-store-selection.js"; +import { + namespaceLedgerPath, + resetNamespaceLedgerWarnings, +} from "@inspector/core/auth/node/oauth-namespace-ledger.js"; +import { + PERSIST_TOKENS_ENV, + oauthSecretServerId, + oauthIdpSecretServerId, + isValidSecretsNamespace, + LEGACY_TOKENS_FIELD, + IDP_SESSION_FIELD, + resetPersistTokensPolicyWarnings, +} from "@inspector/core/auth/node/oauth-secrets.js"; +import { + writeStoreFile, + flushStoreFileWrites, +} from "@inspector/core/storage/store-io.js"; +import type { OAuthPersistSnapshot } from "@inspector/core/auth/oauth-persist.js"; + +const SERVER = "https://api.example/mcp"; +const ISSUER = "https://as.example"; + +function tokensFor(tag: string) { + return { + access_token: `at-${tag}`, + token_type: "Bearer", + refresh_token: `rt-${tag}`, + }; +} + +function snapshotFor(tag: string): OAuthPersistSnapshot { + return { + servers: { [SERVER]: { scope: "read", tokens: tokensFor(tag) } }, + idpSessions: {}, + }; +} + +function namespaceOf(filePath: string): string { + const parsed = JSON.parse(readFileSync(filePath, "utf8")) as Record< + string, + unknown + >; + return parsed[SECRETS_NAMESPACE_KEY] as string; +} + +let tempDir: string; +let fileA: string; +let fileB: string; +let savedPolicy: string | undefined; + +beforeEach(() => { + tempDir = mkdtempSync(join(tmpdir(), "inspector-oauth-ns-")); + fileA = join(tempDir, "profile-a.json"); + fileB = join(tempDir, "profile-b.json"); + savedPolicy = process.env[PERSIST_TOKENS_ENV]; + delete process.env[PERSIST_TOKENS_ENV]; + keyringMocks.password.clear(); +}); + +afterEach(() => { + if (savedPolicy === undefined) delete process.env[PERSIST_TOKENS_ENV]; + else process.env[PERSIST_TOKENS_ENV] = savedPolicy; + resetPersistTokensPolicyWarnings(); + resetOAuthSecretStoreWarnings(); + resetNamespaceLedgerWarnings(); + vi.restoreAllMocks(); + rmSync(tempDir, { recursive: true, force: true }); +}); + +async function flushBoth(): Promise<void> { + await flushStoreFileWrites(fileA); + await flushStoreFileWrites(fileB); +} + +/** The issue's core scenario, parameterized over the shared store backend. */ +async function assertTwoProfileIsolation(store: SecretStore): Promise<void> { + await writeOAuthSections(fileA, snapshotFor("a"), undefined, store); + await writeOAuthSections(fileB, snapshotFor("b"), undefined, store); + await flushBoth(); + + const nsA = namespaceOf(fileA); + const nsB = namespaceOf(fileB); + expect(nsA).not.toBe(nsB); + expect(oauthSecretServerId(SERVER, nsA)).not.toBe( + oauthSecretServerId(SERVER, nsB), + ); + + // B's save must not have clobbered A's entry for the same server. + const readA = await readOAuthStore(fileA, store); + const readB = await readOAuthStore(fileB, store); + expect(readA?.servers[SERVER]?.tokens).toEqual(tokensFor("a")); + expect(readB?.servers[SERVER]?.tokens).toEqual(tokensFor("b")); +} + +describe("secrets namespace isolation (#2549)", () => { + it("keeps two state files' tokens for the same server apart in one shared store", async () => { + await assertTwoProfileIsolation(new InMemorySecretStore()); + }); + + it("keeps them apart on the keyring backend too (acceptance criterion)", async () => { + await assertTwoProfileIsolation(new KeyringSecretStore()); + // Both namespaced accounts coexist in the shared keychain. + const accounts = [...keyringMocks.password.keys()]; + expect( + accounts.filter((a) => a.includes(encodeURIComponent(SERVER))), + ).toHaveLength(2); + }); + + it("keeps them apart on the file backend too (acceptance criterion)", async () => { + // The real FileSecretStore: nested secrets-file locking and serialized + // whole-file mutations are backend-specific and not represented by the + // in-memory double. + await assertTwoProfileIsolation( + new FileSecretStore({ filePath: join(tempDir, "secrets.json") }), + ); + }); + + it("mints a valid namespace on a fresh file's first write and keeps it on later saves", async () => { + const store = new InMemorySecretStore(); + await writeOAuthSections(fileA, snapshotFor("a"), undefined, store); + await flushStoreFileWrites(fileA); + const ns = namespaceOf(fileA); + expect(isValidSecretsNamespace(ns)).toBe(true); + + await writeOAuthSections( + fileA, + { servers: { [SERVER]: { scope: "write" } }, idpSessions: {} }, + { servers: [SERVER] }, + store, + ); + await flushStoreFileWrites(fileA); + expect(namespaceOf(fileA)).toBe(ns); + }); + + it("removing one profile's state purges only its own entries, not the other's", async () => { + const store = new InMemorySecretStore(); + await writeOAuthSections(fileA, snapshotFor("a"), undefined, store); + await writeOAuthSections(fileB, snapshotFor("b"), undefined, store); + await flushBoth(); + const idB = oauthSecretServerId(SERVER, namespaceOf(fileB)); + + await removeOAuthStore(fileA, store); + + // B's scoped entry survives A's removal, and B still reads back whole. + expect(await store.get(idB, LEGACY_TOKENS_FIELD)).not.toBeNull(); + const readB = await readOAuthStore(fileB, store); + expect(readB?.servers[SERVER]?.tokens).toEqual(tokensFor("b")); + }); +}); + +describe("legacy adoption (#2549)", () => { + /** A pre-namespace file plus its legacy unscoped store entries. */ + async function seedLegacy(store: SecretStore): Promise<void> { + await writeStoreFile( + fileA, + JSON.stringify({ + servers: { [SERVER]: { scope: "read" } }, + idpSessions: { [ISSUER]: { clientInformation: { client_id: "idp" } } }, + }), + ); + await flushStoreFileWrites(fileA); + await store.set( + oauthSecretServerId(SERVER), + LEGACY_TOKENS_FIELD, + JSON.stringify(tokensFor("legacy")), + ); + await store.set( + oauthIdpSecretServerId(ISSUER), + IDP_SESSION_FIELD, + JSON.stringify({ tokens: tokensFor("idp") }), + ); + } + + it("first write stamps a namespace, moves legacy entries under it, and deletes the originals", async () => { + const store = new InMemorySecretStore(); + await seedLegacy(store); + + // A sectioned save touching an unrelated server triggers adoption. + await writeOAuthSections( + fileA, + { + servers: { "https://other.example": { scope: "x" } }, + idpSessions: {}, + }, + { servers: ["https://other.example"] }, + store, + ); + await flushStoreFileWrites(fileA); + + const ns = namespaceOf(fileA); + expect(isValidSecretsNamespace(ns)).toBe(true); + expect( + await store.get(oauthSecretServerId(SERVER, ns), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(tokensFor("legacy"))); + expect( + await store.get(oauthIdpSecretServerId(ISSUER, ns), IDP_SESSION_FIELD), + ).toBe(JSON.stringify({ tokens: tokensFor("idp") })); + // The shared legacy slots are retired. + expect( + await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect( + await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), + ).toBeNull(); + + // Joined read-back still sees the moved tokens. + const read = await readOAuthStore(fileA, store); + expect(read?.servers[SERVER]?.tokens).toEqual(tokensFor("legacy")); + }); + + it("re-adopts under a fresh namespace when the stamp is stripped", async () => { + const store = new InMemorySecretStore(); + await seedLegacy(store); + // Simulate a file whose namespace stamp was lost (hand-edit, partial + // restore): earlier scoped entries exist, the file reads as legacy, and + // the next save re-adopts under a brand-new UUID. + await writeOAuthSections(fileA, snapshotFor("scoped"), undefined, store); + await flushStoreFileWrites(fileA); + const ns = namespaceOf(fileA); + const parsed = JSON.parse(readFileSync(fileA, "utf8")) as Record< + string, + unknown + >; + delete parsed[SECRETS_NAMESPACE_KEY]; + await writeStoreFile(fileA, JSON.stringify(parsed)); + await flushStoreFileWrites(fileA); + await store.set( + oauthSecretServerId(SERVER), + LEGACY_TOKENS_FIELD, + JSON.stringify(tokensFor("stale-legacy")), + ); + + await writeOAuthSections( + fileA, + { + servers: { "https://other.example": { scope: "x" } }, + idpSessions: {}, + }, + { servers: ["https://other.example"] }, + store, + ); + await flushStoreFileWrites(fileA); + + // The re-adoption minted a new UUID, so read it back from the file. + const ns2 = namespaceOf(fileA); + expect(ns2).not.toBe(ns); + expect( + await store.get(oauthSecretServerId(SERVER, ns2), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(tokensFor("stale-legacy"))); + }); + + it("a failed legacy purge still attempts every remaining legacy id", async () => { + const store = new InMemorySecretStore(); + await seedLegacy(store); + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + // The first purge (the server's legacy id) fails; the idp session's + // must still be attempted rather than abandoned (best-effort per id). + const realDelete = store.deleteAllForServer.bind(store); + vi.spyOn(store, "deleteAllForServer").mockImplementation(async (id) => { + if (id === oauthSecretServerId(SERVER)) { + throw new Error("purge refused"); + } + return realDelete(id); + }); + + await writeOAuthSections( + fileA, + { + servers: { "https://other.example": { scope: "x" } }, + idpSessions: {}, + }, + { servers: ["https://other.example"] }, + store, + ); + await flushStoreFileWrites(fileA); + + const ns = namespaceOf(fileA); + // Both moves landed, and the idp legacy original was purged despite the + // earlier server purge failing. + expect( + await store.get(oauthSecretServerId(SERVER, ns), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(tokensFor("legacy"))); + expect( + await store.get(oauthIdpSecretServerId(ISSUER, ns), IDP_SESSION_FIELD), + ).toBe(JSON.stringify({ tokens: tokensFor("idp") })); + expect( + await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), + ).toBeNull(); + // The failed purge left its legacy original behind, warned not thrown. + expect( + await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(tokensFor("legacy"))); + expect(warn).toHaveBeenCalledWith( + expect.stringContaining("legacy un-namespaced secret-store entries"), + ); + }); + + it("an invalid secretsNamespace is ignored on read (legacy ids) and replaced on save", async () => { + const store = new InMemorySecretStore(); + await writeStoreFile( + fileA, + JSON.stringify({ + [SECRETS_NAMESPACE_KEY]: "bad+delimiter", + servers: { [SERVER]: { scope: "read" } }, + idpSessions: {}, + }), + ); + await flushStoreFileWrites(fileA); + await store.set( + oauthSecretServerId(SERVER), + LEGACY_TOKENS_FIELD, + JSON.stringify(tokensFor("legacy")), + ); + + // Read resolves through legacy ids, never the invalid value. + const read = await readOAuthStore(fileA, store); + expect(read?.servers[SERVER]?.tokens).toEqual(tokensFor("legacy")); + + // A save (touching an unrelated server) treats the file as legacy: + // mints a fresh valid namespace and moves the legacy entry under it. + await writeOAuthSections( + fileA, + { + servers: { "https://other.example": { scope: "x" } }, + idpSessions: {}, + }, + { servers: ["https://other.example"] }, + store, + ); + await flushStoreFileWrites(fileA); + const ns = namespaceOf(fileA); + expect(isValidSecretsNamespace(ns)).toBe(true); + expect( + await store.get(oauthSecretServerId(SERVER, ns), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(tokensFor("legacy"))); + }); + + it("rolls the copied entries back when adoption cannot complete", async () => { + const store = new InMemorySecretStore(); + await seedLegacy(store); + // First scoped copy lands, second throws → the first must be restored + // and the legacy entries left untouched for the retry. + const realSet = store.set.bind(store); + let sets = 0; + vi.spyOn(store, "set").mockImplementation(async (id, field, value) => { + sets += 1; + if (sets === 2) throw new Error("keychain write refused"); + await realSet(id, field, value); + }); + + await expect( + writeOAuthSections(fileA, snapshotFor("new"), undefined, store), + ).rejects.toThrow("keychain write refused"); + + // Legacy entries are intact; no stamped namespace reached the file. + expect( + await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(tokensFor("legacy"))); + expect( + await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), + ).toBe(JSON.stringify({ tokens: tokensFor("idp") })); + const parsed = JSON.parse(readFileSync(fileA, "utf8")) as Record< + string, + unknown + >; + expect(parsed[SECRETS_NAMESPACE_KEY]).toBeUndefined(); + }); + + it("rolls the scoped copies back when the commit-point stamp write fails", async () => { + const store = new InMemorySecretStore(); + await seedLegacy(store); + // Copies land, then the state-file stamp — the migration's commit + // point — rejects. The copies must be removed and the legacy file and + // ids left authoritative for the retry. Only the state file's write + // fails: the namespace ledger (#2560) is written first, to its own path. + const real = vi.mocked(writeStoreFile).getMockImplementation()!; + vi.mocked(writeStoreFile).mockImplementation(async (path, data) => { + if (path === fileA) throw new Error("disk full during stamp"); + return real(path, data); + }); + + try { + await expect( + writeOAuthSections(fileA, snapshotFor("new"), undefined, store), + ).rejects.toThrow("disk full during stamp"); + } finally { + vi.mocked(writeStoreFile).mockImplementation(real); + } + + // The namespace the failed stamp would have committed (from the blob + // handed to the rejected write) holds no copies. + const attempted = vi + .mocked(writeStoreFile) + .mock.calls.filter(([path]) => path === fileA) + .at(-1)?.[1]; + const ns = (JSON.parse(attempted as string) as Record<string, unknown>)[ + SECRETS_NAMESPACE_KEY + ] as string; + expect(isValidSecretsNamespace(ns)).toBe(true); + expect( + await store.get(oauthSecretServerId(SERVER, ns), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect( + await store.get(oauthIdpSecretServerId(ISSUER, ns), IDP_SESSION_FIELD), + ).toBeNull(); + // Legacy entries are intact and the file is still un-stamped. + expect( + await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD), + ).toBe(JSON.stringify(tokensFor("legacy"))); + expect( + await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), + ).toBe(JSON.stringify({ tokens: tokensFor("idp") })); + const parsed = JSON.parse(readFileSync(fileA, "utf8")) as Record< + string, + unknown + >; + expect(parsed[SECRETS_NAMESPACE_KEY]).toBeUndefined(); + }); + + it("removing a still-legacy profile purges the shared legacy ids — deliberately", async () => { + // A pre-namespace file's live index IS the shared legacy ids, and the + // file being deleted is the store's only index of them: skipping the + // purge would strand credentials in the shared store with nothing left + // able to find or clear them. So removal keeps the pre-namespace + // world's semantics — another still-legacy profile sharing the server + // re-authorizes once, the same cost it pays when a sibling adopts. + // Isolation on removal is a property of *stamped* files (covered + // above), not a retroactive one. + const store = new InMemorySecretStore(); + await seedLegacy(store); + + await removeOAuthStore(fileA, store); + + expect( + await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect( + await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), + ).toBeNull(); + expect(() => readFileSync(fileA, "utf8")).toThrow(); + }); +}); + +describe("namespace ledger: superseded namespaces are purged (#2560)", () => { + const OTHER = "https://other.example/mcp"; + + /** + * What a ≤ 2.9.x save leaves behind: the same entries, no stamp. `drop` + * also removes servers, as the old version may have in the meantime. + */ + async function stripStamp(filePath: string, drop: string[] = []) { + const parsed = JSON.parse(readFileSync(filePath, "utf8")) as { + servers: Record<string, unknown>; + } & Record<string, unknown>; + delete parsed[SECRETS_NAMESPACE_KEY]; + for (const url of drop) delete parsed.servers[url]; + await writeStoreFile(filePath, JSON.stringify(parsed)); + await flushStoreFileWrites(filePath); + } + + function ledgerNamespaces(filePath: string): string[] { + const parsed = JSON.parse( + readFileSync(namespaceLedgerPath(filePath), "utf8"), + ) as { namespaces: Record<string, unknown> }; + return Object.keys(parsed.namespaces); + } + + /** Save once, strip the stamp, save again: the alternating-versions cycle. */ + async function alternate(store: SecretStore, drop: string[] = []) { + await writeOAuthSections( + fileA, + { + servers: { + [SERVER]: { scope: "read", tokens: tokensFor("one") }, + [OTHER]: { scope: "read", tokens: tokensFor("other") }, + }, + idpSessions: {}, + }, + undefined, + store, + ); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + expect(ledgerNamespaces(fileA)).toEqual([ns1]); + + await stripStamp(fileA, drop); + await writeOAuthSections( + fileA, + { servers: { [SERVER]: { scope: "write" } }, idpSessions: {} }, + { servers: [SERVER] }, + store, + ); + await flushStoreFileWrites(fileA); + const ns2 = namespaceOf(fileA); + expect(ns2).not.toBe(ns1); + return { ns1, ns2 }; + } + + it("re-adoption purges the stripped namespace's entries and drops it from the ledger", async () => { + const store = new InMemorySecretStore(); + const { ns1, ns2 } = await alternate(store); + expect( + await store.get(oauthSecretServerId(SERVER, ns1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect( + await store.get(oauthSecretServerId(OTHER, ns1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect(ledgerNamespaces(fileA)).toEqual([ns2]); + }); + + it("also purges servers the old version removed from the file meanwhile", async () => { + // The stripped file no longer names OTHER, so only the ledger can. + const store = new InMemorySecretStore(); + const { ns1 } = await alternate(store, [OTHER]); + expect( + await store.get(oauthSecretServerId(OTHER, ns1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + }); + + it("leaves nothing of the stripped namespace in the keychain (acceptance criterion)", async () => { + const store = new KeyringSecretStore(); + const { ns1 } = await alternate(store); + const accounts = [...keyringMocks.password.keys()]; + expect(accounts.some((a) => a.includes(ns1))).toBe(false); + }); + + it("purges IdP session entries recorded under a stripped namespace", async () => { + const store = new InMemorySecretStore(); + await writeOAuthSections( + fileA, + { + servers: {}, + idpSessions: { [ISSUER]: { idToken: "id-1", refreshToken: "rt-1" } }, + }, + undefined, + store, + ); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + expect( + await store.get(oauthIdpSecretServerId(ISSUER, ns1), IDP_SESSION_FIELD), + ).not.toBeNull(); + + await stripStamp(fileA); + await writeOAuthSections( + fileA, + snapshotFor("x"), + { servers: [SERVER] }, + store, + ); + await flushStoreFileWrites(fileA); + expect( + await store.get(oauthIdpSecretServerId(ISSUER, ns1), IDP_SESSION_FIELD), + ).toBeNull(); + }); + + it("keeps a namespace whose purge failed recorded, and retries it at the next adoption", async () => { + const store = new InMemorySecretStore(); + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + await writeOAuthSections(fileA, snapshotFor("one"), undefined, store); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + const id1 = oauthSecretServerId(SERVER, ns1); + + const realDelete = store.deleteAllForServer.bind(store); + const del = vi + .spyOn(store, "deleteAllForServer") + .mockImplementation(async (id) => { + if (id === id1) throw new Error("keychain locked"); + return realDelete(id); + }); + await stripStamp(fileA); + await writeOAuthSections(fileA, snapshotFor("two"), undefined, store); + await flushStoreFileWrites(fileA); + const ns2 = namespaceOf(fileA); + expect(await store.get(id1, LEGACY_TOKENS_FIELD)).not.toBeNull(); + expect(ledgerNamespaces(fileA).sort()).toEqual([ns1, ns2].sort()); + expect(warn).toHaveBeenCalledWith( + expect.stringContaining("keychain locked"), + ); + + // Purge works again; the next adoption picks up both stale namespaces. + del.mockRestore(); + await stripStamp(fileA); + await writeOAuthSections(fileA, snapshotFor("three"), undefined, store); + await flushStoreFileWrites(fileA); + expect(await store.get(id1, LEGACY_TOKENS_FIELD)).toBeNull(); + expect(ledgerNamespaces(fileA)).toEqual([namespaceOf(fileA)]); + }); + + it("retries a failed purge on the next ordinary save, without another strip", async () => { + const store = new InMemorySecretStore(); + vi.spyOn(console, "warn").mockImplementation(() => {}); + await writeOAuthSections(fileA, snapshotFor("one"), undefined, store); + await flushStoreFileWrites(fileA); + const id1 = oauthSecretServerId(SERVER, namespaceOf(fileA)); + + const realDelete = store.deleteAllForServer.bind(store); + const del = vi + .spyOn(store, "deleteAllForServer") + .mockImplementation(async (id) => { + if (id === id1) throw new Error("keychain locked"); + return realDelete(id); + }); + await stripStamp(fileA); + await writeOAuthSections(fileA, snapshotFor("two"), undefined, store); + await flushStoreFileWrites(fileA); + expect(await store.get(id1, LEGACY_TOKENS_FIELD)).not.toBeNull(); + + // The store recovers; the file keeps its new stamp. + del.mockRestore(); + const ns2 = namespaceOf(fileA); + await writeOAuthSections(fileA, snapshotFor("three"), undefined, store); + await flushStoreFileWrites(fileA); + expect(namespaceOf(fileA)).toBe(ns2); + expect(await store.get(id1, LEGACY_TOKENS_FIELD)).toBeNull(); + expect(ledgerNamespaces(fileA)).toEqual([ns2]); + }); + + it("a run on another backend leaves the keychain's record for a keychain run", async () => { + // Saved to the keychain, stamp stripped, then a run that fell back to + // a different store re-adopts: its no-op purge must not consume the + // keychain record. The next keychain run cleans up. + const keyring = new KeyringSecretStore(); + await writeOAuthSections(fileA, snapshotFor("one"), undefined, keyring); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + const hasNs1 = () => + [...keyringMocks.password.keys()].some((a) => a.includes(ns1)); + expect(hasNs1()).toBe(true); + + await stripStamp(fileA); + await writeOAuthSections( + fileA, + snapshotFor("fallback"), + undefined, + new InMemorySecretStore(), + ); + await flushStoreFileWrites(fileA); + expect(hasNs1()).toBe(true); + expect(ledgerNamespaces(fileA)).toContain(ns1); + + await writeOAuthSections(fileA, snapshotFor("back"), undefined, keyring); + await flushStoreFileWrites(fileA); + expect(hasNs1()).toBe(false); + expect(ledgerNamespaces(fileA)).not.toContain(ns1); + }); + + it("cleans up a stripped namespace whose file-store entries the keychain absorbed", async () => { + // File-backed save → stamp stripped → keychain becomes available and + // absorbs secrets.json (superseded namespace included) → re-adoption + // on the keychain must purge the absorbed copies. + const secretsFile = join(tempDir, "secrets.json"); + const savedFileEnv = process.env[SECRET_FILE_ENV]; + process.env[SECRET_FILE_ENV] = secretsFile; + vi.spyOn(console, "warn").mockImplementation(() => {}); + try { + const fileStore = new FileSecretStore({ + filePath: secretsFile, + passphrase: "", + }); + await writeOAuthSections(fileA, snapshotFor("one"), undefined, fileStore); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + await stripStamp(fileA); + + const keyring = new KeyringSecretStore(); + await absorbFileSecretsIntoKeyring(keyring); + const hasNs1 = () => + [...keyringMocks.password.keys()].some((a) => a.includes(ns1)); + expect(existsSync(secretsFile)).toBe(false); + expect(hasNs1()).toBe(true); + + await writeOAuthSections(fileA, snapshotFor("two"), undefined, keyring); + await flushStoreFileWrites(fileA); + expect(hasNs1()).toBe(false); + } finally { + if (savedFileEnv === undefined) delete process.env[SECRET_FILE_ENV]; + else process.env[SECRET_FILE_ENV] = savedFileEnv; + } + }); + + it("keeps tracking hand-off copies through a file-backed re-adoption, until the keychain returns", async () => { + // File-backed save → strip → hand-off copies into the keychain → the + // next run is file-backed again and re-adopts → back on the keychain. + // The file run's purge must not be the end of the keychain copies. + const secretsFile = join(tempDir, "secrets.json"); + const savedFileEnv = process.env[SECRET_FILE_ENV]; + process.env[SECRET_FILE_ENV] = secretsFile; + vi.spyOn(console, "warn").mockImplementation(() => {}); + try { + const fileStore = () => + new FileSecretStore({ filePath: secretsFile, passphrase: "" }); + await writeOAuthSections( + fileA, + snapshotFor("one"), + undefined, + fileStore(), + ); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + await stripStamp(fileA); + + const keyring = new KeyringSecretStore(); + await absorbFileSecretsIntoKeyring(keyring); + const hasNs1 = () => + [...keyringMocks.password.keys()].some((a) => a.includes(ns1)); + expect(hasNs1()).toBe(true); + + await writeOAuthSections( + fileA, + snapshotFor("two"), + undefined, + fileStore(), + ); + await flushStoreFileWrites(fileA); + expect(hasNs1()).toBe(true); + + await writeOAuthSections(fileA, snapshotFor("three"), undefined, keyring); + await flushStoreFileWrites(fileA); + expect(hasNs1()).toBe(false); + } finally { + if (savedFileEnv === undefined) delete process.env[SECRET_FILE_ENV]; + else process.env[SECRET_FILE_ENV] = savedFileEnv; + } + }); + + it("removing a stripped file purges the namespace it lost, and the ledger with it", async () => { + const store = new InMemorySecretStore(); + await writeOAuthSections(fileA, snapshotFor("one"), undefined, store); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + await stripStamp(fileA); + + await removeOAuthStore(fileA, store); + expect( + await store.get(oauthSecretServerId(SERVER, ns1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + expect(existsSync(namespaceLedgerPath(fileA))).toBe(false); + }); + + it("a state file deleted by hand has its entries purged by the next save", async () => { + const store = new InMemorySecretStore(); + await writeOAuthSections(fileA, snapshotFor("one"), undefined, store); + await flushStoreFileWrites(fileA); + const ns1 = namespaceOf(fileA); + rmSync(fileA); + + await writeOAuthSections(fileA, snapshotFor("two"), undefined, store); + await flushStoreFileWrites(fileA); + expect( + await store.get(oauthSecretServerId(SERVER, ns1), LEGACY_TOKENS_FIELD), + ).toBeNull(); + }); + + it("does not touch another state file's namespace", async () => { + const store = new InMemorySecretStore(); + await writeOAuthSections(fileB, snapshotFor("b"), undefined, store); + await flushStoreFileWrites(fileB); + const idB = oauthSecretServerId(SERVER, namespaceOf(fileB)); + + await alternate(store); + expect(await store.get(idB, LEGACY_TOKENS_FIELD)).not.toBeNull(); + }); +}); diff --git a/clients/web/src/test/integration/auth/node/secret-store-selection.test.ts b/clients/web/src/test/integration/auth/node/secret-store-selection.test.ts index 7aca1a2e16..1c340d90c5 100644 --- a/clients/web/src/test/integration/auth/node/secret-store-selection.test.ts +++ b/clients/web/src/test/integration/auth/node/secret-store-selection.test.ts @@ -27,6 +27,7 @@ import { isOnMountPoint, parseSecretStoreEnv, SECRET_STORAGE_DOCS_URL, + setSecretStorageWarningsQuiet, warnAboutSecretStorage, } from "@inspector/core/auth/node/secret-store-selection.js"; import { @@ -140,6 +141,26 @@ describe("chooseFallbackKind", () => { "file", ); }); + + it("never uses memory when allowMemory is false (multi-process consumers)", () => { + // mcpdo's front-end, OAuth helper, and daemon are separate processes; a + // per-process memory store can never serve them, so the container + // special case collapses to file. + expect( + chooseFallbackKind({ + container: true, + mounted: false, + allowMemory: false, + }), + ).toBe("file"); + expect( + chooseFallbackKind({ + container: false, + mounted: false, + allowMemory: false, + }), + ).toBe("file"); + }); }); describe("isContainer", () => { @@ -527,6 +548,46 @@ describe("warnAboutSecretStorage", () => { }); }); +describe("setSecretStorageWarningsQuiet", () => { + const fallback: SecretStorageInfo = { + kind: "file", + reason: "fallback", + durable: true, + path: "/home/u/.mcp-inspector/secrets.json", + plaintext: true, + detail: "no Secret Service", + }; + + afterEach(() => { + // The flag is process-wide; never leak quiet into a later test. + setSecretStorageWarningsQuiet(false); + }); + + it("suppresses the automatic warning when quiet", () => { + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + setSecretStorageWarningsQuiet(true); + warnAboutSecretStorage(fallback); + expect(warn).not.toHaveBeenCalled(); + }); + + it("still prints when quiet is cleared again", () => { + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + setSecretStorageWarningsQuiet(true); + setSecretStorageWarningsQuiet(false); + warnAboutSecretStorage(fallback); + expect(warn).toHaveBeenCalled(); + }); + + it("force bypasses quiet so connect can re-surface it once", () => { + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + setSecretStorageWarningsQuiet(true); + warnAboutSecretStorage(fallback, { force: true }); + const output = warn.mock.calls.flat().join("\n"); + expect(output).toContain("Secrets are stored unencrypted"); + expect(output).toContain(SECRET_STORAGE_DOCS_URL); + }); +}); + describe("resolveSecretStore", () => { it("uses the keychain when the probe succeeds", async () => { const mod = await loadWithProbe(true); @@ -566,6 +627,29 @@ describe("resolveSecretStore", () => { }); }); + it("explicit memory still wins after disallowMemorySecretStoreFallback", async () => { + // The disallow shapes only the automatic fallback; a user override is a + // statement of intent and keeps winning outright. + process.env.MCP_INSPECTOR_SECRET_STORE = "memory"; + vi.spyOn(console, "warn").mockImplementation(() => {}); + const mod = await loadWithProbe(true); + mod.disallowMemorySecretStoreFallback(); + const { info } = await mod.resolveSecretStore(); + expect(info).toMatchObject({ kind: "memory", reason: "configured" }); + }); + + it("falls back to file, never memory, once memory fallback is disallowed", async () => { + vi.spyOn(console, "warn").mockImplementation(() => {}); + process.env.MCP_INSPECTOR_SECRET_FILE = path.join(tmpDir, "secrets.json"); + // Look like the one environment whose automatic answer is memory. + process.env.KUBERNETES_SERVICE_HOST = "10.0.0.1"; + const mod = await loadWithProbe(false); + mod.disallowMemorySecretStoreFallback(); + const { info } = await mod.resolveSecretStore(); + expect(info.reason).toBe("fallback"); + expect(info.kind).toBe("file"); + }); + it("honors an explicit file store, reporting its path and encryption state", async () => { const target = path.join(tmpDir, "custom", "secrets.json"); process.env.MCP_INSPECTOR_SECRET_STORE = "file"; diff --git a/clients/web/src/test/integration/mcp/inspectorClient-advertised-extensions.test.ts b/clients/web/src/test/integration/mcp/inspectorClient-advertised-extensions.test.ts index fc7f2c4420..58b22f0126 100644 --- a/clients/web/src/test/integration/mcp/inspectorClient-advertised-extensions.test.ts +++ b/clients/web/src/test/integration/mcp/inspectorClient-advertised-extensions.test.ts @@ -2,7 +2,7 @@ import { describe, it, expect, afterEach } from "vitest"; import * as z from "zod/v4"; import { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; import { createTransportNode } from "@inspector/core/mcp/node/transport.js"; -import { TASKS_EXTENSION_KEY } from "@inspector/core/mcp/modernTaskSchemas.js"; +import { TASKS_EXTENSION_KEY } from "@inspector/core/extension/tasks/constants.js"; import { createTestServerHttp, type TestServerHttp, diff --git a/clients/web/src/test/integration/mcp/inspectorClient-ambient-signal.test.ts b/clients/web/src/test/integration/mcp/inspectorClient-ambient-signal.test.ts new file mode 100644 index 0000000000..c67166a05e --- /dev/null +++ b/clients/web/src/test/integration/mcp/inspectorClient-ambient-signal.test.ts @@ -0,0 +1,154 @@ +import { describe, it, expect, afterEach } from "vitest"; +import { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; +import { createTransportNode } from "@inspector/core/mcp/node/transport.js"; +import { + createTestServerHttp, + createTestServerInfo, + type TestServerHttp, + type ResourceTemplateDefinition, +} from "@modelcontextprotocol/inspector-test-server"; + +/** + * #1783 — the ambient request signal must cancel an in-flight *non-tool* + * request, the same way `cancelToolCall()` cancels a tool call. + * + * The daemon serializes every method on a connection, so a `resources/read` + * (or any list, `prompts/get`, …) the server never answers held the + * per-connection rpc queue slot forever once the caller hung up — wedging the + * whole connection until a reconnect. Core only ever threaded a cancellation + * signal into *tool* calls; every other method called `getRequestOptions` + * with no signal, so there was nothing for a caller disconnect to abort. + * + * The fix is an ambient request signal the daemon sets for the span of one + * command: `getRequestOptions` folds it into *every* request's options. This + * drives the whole chain a real caller drives — InspectorClient -> + * transport -> a real test server — and asserts at the far end, on the + * server's own request abort signal. Nothing shallower reaches it: a unit + * test of `getRequestOptions` cannot prove the signal survives the SDK and the + * wire, which is the entire failure mode. + */ +describe("ambient request signal cancels non-tool requests (#1783)", () => { + let client: InspectorClient | null = null; + let server: TestServerHttp | null = null; + + afterEach(async () => { + if (client) { + try { + await client.disconnect(); + } catch { + // ignore + } + client = null; + } + if (server) { + try { + await server.stop(); + } catch { + // ignore + } + server = null; + } + }); + + function deferred<T>(): { + promise: Promise<T>; + resolve: (value: T) => void; + } { + let resolve!: (value: T) => void; + const promise = new Promise<T>((r) => { + resolve = r; + }); + return { promise, resolve }; + } + + /** + * A resource template whose read never returns on its own. It reports when + * the read started and whether the server-side request signal — which the + * SDK aborts when the client cancels the request — ever fired. + */ + function slowTemplate(): { + template: ResourceTemplateDefinition; + started: Promise<void>; + aborted: Promise<boolean>; + } { + const started = deferred<void>(); + const aborted = deferred<boolean>(); + const template: ResourceTemplateDefinition = { + name: "slow", + uriTemplate: "slow://item/{id}", + handler: (_uri, _params, _context, extra) => { + extra?.signal?.addEventListener("abort", () => aborted.resolve(true), { + once: true, + }); + started.resolve(); + // Never resolves on its own; only the abort settles the read. + return new Promise(() => {}); + }, + }; + return { template, started: started.promise, aborted: aborted.promise }; + } + + async function connect( + template: ResourceTemplateDefinition, + ): Promise<InspectorClient> { + const started = createTestServerHttp({ + serverInfo: createTestServerInfo("ambient-signal-test", "1.0.0"), + tools: [], + resourceTemplates: [template], + }); + await started.start(); + server = started; + + const connected = new InspectorClient( + { type: "streamable-http", url: started.url }, + { environment: { transport: createTransportNode } }, + ); + await connected.connect(); + client = connected; + return connected; + } + + it("aborts an in-flight resources/read when the ambient signal fires", async () => { + const slow = slowTemplate(); + const connected = await connect(slow.template); + + const controller = new AbortController(); + const dispose = connected.setAmbientRequestSignal(controller.signal); + + // A non-tool read that the server never answers. + const read = connected.readResource("slow://item/1"); + const rejected = expect(read).rejects.toThrow(); + + await slow.started; + controller.abort(); + + // The read rejects (the caller is unwedged) and the server observed the + // cancellation at the far end of the chain. + await rejected; + expect(await slow.aborted).toBe(true); + + dispose(); + }); + + it("leaves a normal request untouched when the ambient signal never fires", async () => { + // A control: setting an ambient signal that is never aborted must not break + // an ordinary request. Uses a template that answers immediately. + const template: ResourceTemplateDefinition = { + name: "fast", + uriTemplate: "fast://item/{id}", + handler: async (uri) => ({ + contents: [{ uri: uri.href, text: "ok" }], + }), + }; + const connected = await connect(template); + + const controller = new AbortController(); + const dispose = connected.setAmbientRequestSignal(controller.signal); + + const { result } = await connected.readResource("fast://item/1"); + const first = result.contents[0]; + expect(first && "text" in first ? first.text : undefined).toBe("ok"); + + dispose(); + }); +}); diff --git a/clients/web/src/test/integration/mcp/inspectorClient-coverage-backfill.test.ts b/clients/web/src/test/integration/mcp/inspectorClient-coverage-backfill.test.ts index afbdff41bc..25d5b1d3a9 100644 --- a/clients/web/src/test/integration/mcp/inspectorClient-coverage-backfill.test.ts +++ b/clients/web/src/test/integration/mcp/inspectorClient-coverage-backfill.test.ts @@ -569,34 +569,6 @@ describe("InspectorClient coverage backfill", () => { ).rejects.toThrow(/forced-call-failure/); }); - it("callToolStream error path covers metadata combinations", async () => { - client = stdioClient(); - await client.connect(); - // Force the streaming API to throw synchronously so the error-path - // metadata-merge branches run for each combination. `callToolStream` - // drives tasks via the private `pollTaskToolCall` generator (SDK v2 - // removed `client.experimental.tasks`), so patch that instead. - const c = client as unknown as { - pollTaskToolCall: () => never; - }; - c.pollTaskToolCall = () => { - throw new Error("forced-stream-failure"); - }; - - await expect( - client.callToolStream(badTool, {}, { g: "1" }, undefined), - ).rejects.toThrow(/forced-stream-failure/); - await expect( - client.callToolStream(badTool, {}, undefined, { t: "1" }), - ).rejects.toThrow(/forced-stream-failure/); - await expect( - client.callToolStream(badTool, {}, { g: "1" }, { t: "2" }), - ).rejects.toThrow(/forced-stream-failure/); - await expect( - client.callToolStream(badTool, {}, undefined, undefined), - ).rejects.toThrow(/forced-stream-failure/); - }); - it("callToolStream success path covers metadata combinations", async () => { client = stdioClient(); await client.connect(); @@ -726,21 +698,6 @@ describe("InspectorClient coverage backfill", () => { ); }); - it("callToolStream error dispatch handles a non-Error rejection", async () => { - client = stdioClient(); - await client.connect(); - const echo = await getTool(client, "echo"); - const c = client as unknown as { - pollTaskToolCall: () => never; - }; - c.pollTaskToolCall = () => { - throw "string-stream-failure"; - }; - await expect( - client.callToolStream(echo, { message: "x" }, { g: "1" }), - ).rejects.toBe("string-stream-failure"); - }); - it("list methods fall back to empty arrays when the response omits them", async () => { client = stdioClient(); await client.connect(); @@ -772,114 +729,7 @@ describe("InspectorClient coverage backfill", () => { }); }); - describe("callToolStream task-result fallback and terminal branches", () => { - function patchStream( - c: InspectorClient, - gen: () => AsyncGenerator<unknown>, - getTaskResult?: () => Promise<unknown>, - ): void { - // SDK v2 removed `client.experimental.tasks`; `callToolStream` now drives - // tasks via the private `pollTaskToolCall` generator, and the no-result - // fallback fetches the payload with a raw `client.request({method:"tasks/result"})`. - const internal = c as unknown as { - pollTaskToolCall: () => AsyncGenerator<unknown>; - client: { - request: ( - req: { method: string }, - schema: unknown, - opts: unknown, - ) => Promise<unknown>; - }; - }; - internal.pollTaskToolCall = gen; - if (getTaskResult) { - const origRequest = internal.client.request.bind(internal.client); - internal.client.request = (req, schema, opts) => - req.method === "tasks/result" - ? getTaskResult() - : origRequest(req, schema, opts); - } - } - - const fakeTask = (taskId: string) => ({ - taskId, - status: "working" as const, - ttl: null, - createdAt: new Date().toISOString(), - lastUpdatedAt: new Date().toISOString(), - }); - - it("falls back to getTaskResult when the stream yields no result message", async () => { - client = stdioClient(); - await client.connect(); - const echo = await getTool(client, "echo"); - patchStream( - client, - async function* () { - // taskCreated then taskStatus, but NO result → triggers the fallback. - yield { type: "taskCreated", task: fakeTask("T1") }; - yield { type: "taskStatus", task: fakeTask("T1") }; - }, - async () => ({ content: [{ type: "text", text: "from-fallback" }] }), - ); - const result = await client.callToolStream(echo, { message: "x" }); - expect(result.success).toBe(true); - expect(result.result?.content?.[0]).toMatchObject({ - text: "from-fallback", - }); - }); - - it("throws when the getTaskResult fallback itself fails", async () => { - client = stdioClient(); - await client.connect(); - const echo = await getTool(client, "echo"); - patchStream( - client, - async function* () { - yield { type: "taskCreated", task: fakeTask("T2") }; - }, - async () => { - throw new Error("fallback-failed"); - }, - ); - await expect( - client.callToolStream(echo, { message: "x" }), - ).rejects.toThrow(/Tool call did not return a result: fallback-failed/); - }); - - it("surfaces a stream error message and marks the task failed", async () => { - client = stdioClient(); - await client.connect(); - const echo = await getTool(client, "echo"); - patchStream(client, async function* () { - yield { type: "taskCreated", task: fakeTask("T3") }; - // Empty error message → `message.error.message || "Task execution failed"` - // falsy branch. - yield { type: "error", error: { message: "" } }; - }); - await expect( - client.callToolStream(echo, { message: "x" }), - ).rejects.toThrow(/Task execution failed/); - }); - - it("treats a stream taskStatus before taskCreated as the task id", async () => { - client = stdioClient(); - await client.connect(); - const echo = await getTool(client, "echo"); - patchStream(client, async function* () { - // taskStatus arrives first (no prior taskCreated) → `if (!taskId)` true. - yield { type: "taskStatus", task: fakeTask("T4") }; - yield { - type: "result", - result: { content: [{ type: "text", text: "ok" }] }, - }; - }); - const result = await client.callToolStream(echo, { message: "x" }); - expect(result.success).toBe(true); - }); - }); - - describe("constructor capability branches and createReceiverTask options", () => { + describe("constructor capability branches", () => { it("still advertises the registry extensions when no other capabilities are set", async () => { // sample:false + elicit:false + no roots + no receiverTasks → the only // advertised capabilities are the default-on registry extensions (Tasks @@ -906,117 +756,5 @@ describe("InspectorClient coverage backfill", () => { await client.connect(); expect(client.getStatus()).toBe("connected"); }); - - it("createReceiverTask honors pollInterval and statusMessage options", async () => { - client = stdioClient(); - await client.connect(); - const internal = client as unknown as { - createReceiverTask: (opts: { - initialStatus: string; - ttl?: number; - pollInterval?: number; - statusMessage?: string; - }) => { - task: { - taskId: string; - pollInterval?: number; - statusMessage?: string; - }; - }; - }; - const record = internal.createReceiverTask({ - initialStatus: "working", - ttl: 5000, - pollInterval: 250, - statusMessage: "in progress", - }); - expect(record.task.pollInterval).toBe(250); - expect(record.task.statusMessage).toBe("in progress"); - - // Omitting ttl falls back to the configured numeric receiverTaskTtlMs - // (default 60_000) → the non-function branch of the ttl resolution. - const recordNoTtl = internal.createReceiverTask({ - initialStatus: "working", - }) as unknown as { task: { ttl: number } }; - expect(recordNoTtl.task.ttl).toBe(60_000); - }); - - it("createReceiverTask falls back to the configured TTL when ttl is omitted (function form)", async () => { - // receiverTaskTtlMs as a function → exercises the typeof === 'function' - // branch of the ttl resolution. - client = new InspectorClient( - { - type: "stdio", - command: serverCommand.command, - args: serverCommand.args, - }, - { - environment: { transport: createTransportNode }, - receiverTasks: true, - receiverTaskTtlMs: () => 1234, - }, - ); - await client.connect(); - const internal = client as unknown as { - createReceiverTask: (opts: { initialStatus: string }) => { - task: { ttl: number }; - }; - }; - const record = internal.createReceiverTask({ initialStatus: "working" }); - expect(record.task.ttl).toBe(1234); - }); - }); - - describe("emitReceiverTaskStatus guards", () => { - it("is a no-op when there is no connected client", () => { - const c = stdioClient(); - (c as unknown as { client: unknown }).client = null; - // Should not throw even though client is null. - expect(() => - ( - c as unknown as { emitReceiverTaskStatus: (t: unknown) => void } - ).emitReceiverTaskStatus({ taskId: "t" }), - ).not.toThrow(); - }); - - it("swallows notification-build errors via the catch path", async () => { - client = stdioClient(); - await client.connect(); - const internal = client as unknown as { - emitReceiverTaskStatus: (t: unknown) => void; - }; - // Passing a malformed task makes TaskStatusNotificationSchema.parse throw, - // which is caught and logged (no throw). - expect(() => - internal.emitReceiverTaskStatus({ not: "a task" }), - ).not.toThrow(); - }); - - it("upsertReceiverTask emits status for an existing record", async () => { - client = stdioClient(); - await client.connect(); - const internal = client as unknown as { - createReceiverTask: (opts: { initialStatus: string; ttl?: number }) => { - task: { taskId: string; status: string }; - }; - upsertReceiverTask: (t: { taskId: string; status: string }) => void; - getReceiverTask: ( - id: string, - ) => { task: { status: string } } | undefined; - }; - const record = internal.createReceiverTask({ - initialStatus: "working", - ttl: 5000, - }); - const updated = { ...record.task, status: "completed" }; - internal.upsertReceiverTask(updated); - expect(internal.getReceiverTask(record.task.taskId)?.task.status).toBe( - "completed", - ); - // upsert on an unknown id is a no-op (record undefined branch). - expect(() => - internal.upsertReceiverTask({ taskId: "missing", status: "completed" }), - ).not.toThrow(); - }); }); }); diff --git a/clients/web/src/test/integration/mcp/inspectorClient-crash-reconciliation.test.ts b/clients/web/src/test/integration/mcp/inspectorClient-crash-reconciliation.test.ts new file mode 100644 index 0000000000..077fdc700e --- /dev/null +++ b/clients/web/src/test/integration/mcp/inspectorClient-crash-reconciliation.test.ts @@ -0,0 +1,270 @@ +/** + * Mid-session crash reconciliation (#2437), driven by a real server process + * that dies at a point the test picks. + * + * `InspectorClient` has a whole teardown route for a connection that ends + * without `disconnect()` — the transport `onclose` in + * `attachTransportListeners`: settle the status, drop the server's queued + * sampling/elicitation requests before announcing the `disconnect`, reject + * what we had asked the server, fire `disconnect` exactly once. Until this + * file, that route was only reached by unit tests calling `onclose` on a fake + * transport, so nothing checked it against what a real transport does when + * the process at the other end goes away. + * + * The server is the stdio test server started with `--crashable`, which adds + * a `crash_server` tool (see `createCrashServerTool`): `respond: false` exits + * with the call still in flight, `respond: true` answers first and exits on an + * idle session. It runs as its own process, so the exit is a real one. + */ + +import { describe, it, expect, afterEach } from "vitest"; +import type { Tool } from "@modelcontextprotocol/client"; +import { InspectorClient } from "@inspector/core/mcp/inspectorClient.js"; +import { createTransportNode } from "@inspector/core/mcp/node/transport.js"; +import type { + ConnectionStatus, + InspectorClientOptions, +} from "@inspector/core/mcp/types.js"; +import { + CRASH_SERVER_TOOL_NAME, + getCrashableTestMcpServerCommand, + waitForEvent, +} from "@modelcontextprotocol/inspector-test-server"; + +/** One entry per teardown-relevant event, in dispatch order. */ +type RecordedEvent = + | { type: "statusChange"; status: ConnectionStatus } + | { type: "disconnect"; pendingSamples: number; pendingElicitations: number } + | { type: "pendingSamplesChange"; count: number } + | { type: "pendingElicitationsChange"; count: number }; + +let client: InspectorClient | null = null; + +function createCrashableClient( + options: Partial<InspectorClientOptions> = {}, +): InspectorClient { + const { command, args } = getCrashableTestMcpServerCommand(); + return new InspectorClient( + { type: "stdio", command, args }, + { environment: { transport: createTransportNode }, ...options }, + ); +} + +/** + * Record the events the crash route dispatches, in order. The `disconnect` + * entry snapshots both peer-request queues *as its listener sees them*, which + * is the ordering guarantee under test: a consumer handling `disconnect` + * already sees the queues empty. + */ +function recordEvents(target: InspectorClient): RecordedEvent[] { + const events: RecordedEvent[] = []; + target.addEventListener("statusChange", (e) => + events.push({ type: "statusChange", status: e.detail }), + ); + target.addEventListener("disconnect", () => + events.push({ + type: "disconnect", + pendingSamples: target.getPendingSamples().length, + pendingElicitations: target.getPendingElicitations().length, + }), + ); + target.addEventListener("pendingSamplesChange", (e) => + events.push({ type: "pendingSamplesChange", count: e.detail.length }), + ); + target.addEventListener("pendingElicitationsChange", (e) => + events.push({ type: "pendingElicitationsChange", count: e.detail.length }), + ); + return events; +} + +async function getTool(target: InspectorClient, name: string): Promise<Tool> { + const { tools } = await target.listTools(); + const tool = tools.find((t) => t.name === name); + if (!tool) throw new Error(`Tool ${name} not found`); + return tool; +} + +function countOf(events: RecordedEvent[], type: RecordedEvent["type"]) { + return events.filter((e) => e.type === type).length; +} + +/** + * Exit with the `crash_server` call itself still in flight, handing back that + * call's (rejecting) promise with its rejection already handled — so it + * cannot surface as an unhandled rejection if an assertion fails first. + * + * Wrapped in an object because an `async` function returning a promise + * adopts it: returned bare, awaiting this helper would rethrow the crash. + */ +async function crashWithCallInFlight( + target: InspectorClient, +): Promise<{ call: Promise<unknown> }> { + const crashTool = await getTool(target, CRASH_SERVER_TOOL_NAME); + const call = target.callTool(crashTool, { respond: false }); + call.catch(() => {}); + return { call }; +} + +afterEach(async () => { + // A crashed client is already torn down, but a test that failed before its + // crash leaves a live process behind; `disconnect()` is safe either way. + await client?.disconnect().catch(() => {}); + client = null; +}); + +describe("InspectorClient mid-session crash reconciliation (#2437)", () => { + it("settles to disconnected and fires disconnect once when an idle session's server exits", async () => { + client = createCrashableClient(); + await client.connect(); + const events = recordEvents(client); + + const disconnected = waitForEvent(client, "disconnect"); + const crashTool = await getTool(client, CRASH_SERVER_TOOL_NAME); + const result = await client.callTool(crashTool, { + respond: true, + delayMs: 20, + }); + expect(result.success).toBe(true); + await disconnected; + + expect(client.getStatus()).toBe("disconnected"); + expect(events).toEqual([ + { type: "statusChange", status: "disconnected" }, + { type: "disconnect", pendingSamples: 0, pendingElicitations: 0 }, + ]); + + // An explicit disconnect() after the crash must not announce the + // teardown a second time: the crash route already did. + await client.disconnect(); + expect(countOf(events, "disconnect")).toBe(1); + expect(client.getStatus()).toBe("disconnected"); + }); + + it("rejects the call in flight when the server exits without answering it", async () => { + client = createCrashableClient(); + await client.connect(); + const events = recordEvents(client); + + const disconnected = waitForEvent(client, "disconnect"); + const { call } = await crashWithCallInFlight(client); + + await expect(call).rejects.toThrow(/closed/i); + await disconnected; + expect(client.getStatus()).toBe("disconnected"); + expect(countOf(events, "disconnect")).toBe(1); + }); + + it("drops a pending elicitation and announces it before the disconnect", async () => { + client = createCrashableClient({ elicit: true }); + await client.connect(); + + const pendingArrived = waitForEvent(client, "newPendingElicitation"); + const elicitTool = await getTool(client, "collect_elicitation"); + const elicitCall = client.callTool(elicitTool, { + message: "Never answered", + schema: { type: "object", properties: { name: { type: "string" } } }, + }); + elicitCall.catch(() => {}); + await pendingArrived; + expect(client.getPendingElicitations()).toHaveLength(1); + + const events = recordEvents(client); + const disconnected = waitForEvent(client, "disconnect"); + await crashWithCallInFlight(client); + await disconnected; + + expect(client.getPendingElicitations()).toHaveLength(0); + // Cleared *and announced* before `disconnect`: the web pending-request + // modal tracks its own state off the change event, so a clear without it + // would leave the modal up for a connection that is gone. + const dropped = events.findIndex( + (e) => e.type === "pendingElicitationsChange" && e.count === 0, + ); + const disconnect = events.findIndex((e) => e.type === "disconnect"); + expect(dropped).toBeGreaterThanOrEqual(0); + expect(dropped).toBeLessThan(disconnect); + expect(events[disconnect]).toEqual({ + type: "disconnect", + pendingSamples: 0, + pendingElicitations: 0, + }); + // The tool call that raised the elicitation dies with the connection. + await expect(elicitCall).rejects.toThrow(); + }); + + it("drops a pending sampling request and announces it before the disconnect", async () => { + client = createCrashableClient({ sample: true }); + await client.connect(); + + const pendingArrived = waitForEvent(client, "newPendingSample"); + const sampleTool = await getTool(client, "collect_sample"); + const sampleCall = client.callTool(sampleTool, { text: "Never answered" }); + sampleCall.catch(() => {}); + await pendingArrived; + expect(client.getPendingSamples()).toHaveLength(1); + + const events = recordEvents(client); + const disconnected = waitForEvent(client, "disconnect"); + await crashWithCallInFlight(client); + await disconnected; + + expect(client.getPendingSamples()).toHaveLength(0); + const dropped = events.findIndex( + (e) => e.type === "pendingSamplesChange" && e.count === 0, + ); + const disconnect = events.findIndex((e) => e.type === "disconnect"); + expect(dropped).toBeGreaterThanOrEqual(0); + expect(dropped).toBeLessThan(disconnect); + expect(events[disconnect]).toEqual({ + type: "disconnect", + pendingSamples: 0, + pendingElicitations: 0, + }); + await expect(sampleCall).rejects.toThrow(); + }); + + it("captures what the server wrote to stderr on its way down", async () => { + client = createCrashableClient({ pipeStderr: true }); + await client.connect(); + + const lines: string[] = []; + client.addEventListener("stderrLog", (e) => lines.push(e.detail.message)); + const disconnected = waitForEvent(client, "disconnect"); + const dyingWords = `fatal: crash-${Date.now()}`; + const crashTool = await getTool(client, CRASH_SERVER_TOOL_NAME); + const call = client.callTool(crashTool, { stderr: dyingWords }); + call.catch(() => {}); + await disconnected; + + // The child's stderr is read out-of-band from its exit, so allow the + // chunk a moment to land rather than asserting on a single sample. + await expect + .poll(() => lines.some((line) => line.includes(dyingWords))) + .toBe(true); + await expect(call).rejects.toThrow(); + }); + + it("reconnects the same client after a crash, against a fresh server process", async () => { + client = createCrashableClient(); + await client.connect(); + const disconnected = waitForEvent(client, "disconnect"); + await crashWithCallInFlight(client); + await disconnected; + + // The crash route leaves the transport object cached; the next connect() + // has to bring up a new process through it rather than fail on the dead + // one. + await client.connect(); + expect(client.getStatus()).toBe("connected"); + expect(client.getCapabilities()?.tools).toBeDefined(); + + // And the new session is live end to end: it can itself be crashed, and + // that second crash is reconciled the same way as the first. + const events = recordEvents(client); + const disconnectedAgain = waitForEvent(client, "disconnect"); + await crashWithCallInFlight(client); + await disconnectedAgain; + expect(client.getStatus()).toBe("disconnected"); + expect(countOf(events, "disconnect")).toBe(1); + }); +}); diff --git a/clients/web/src/test/integration/mcp/inspectorClient-modern-era.test.ts b/clients/web/src/test/integration/mcp/inspectorClient-modern-era.test.ts index 5c11dcb787..0727cf83f6 100644 --- a/clients/web/src/test/integration/mcp/inspectorClient-modern-era.test.ts +++ b/clients/web/src/test/integration/mcp/inspectorClient-modern-era.test.ts @@ -97,6 +97,28 @@ describe("modern-era negotiation (2026-07-28)", () => { return connected; } + it("fails connect when the required Tasks session cannot attach", async () => { + const started = await startServer(); + const connected = new InspectorClient( + { type: "streamable-http", url: started.url }, + { + environment: { transport: createTransportNode }, + versionNegotiation: eraToVersionNegotiation("modern"), + }, + ); + const boundary = connected as unknown as { + attachTaskSession: () => Promise<void>; + }; + boundary.attachTaskSession = () => + Promise.reject(new Error("task session attach failed")); + client = connected; + + await expect(connected.connect()).rejects.toThrow( + "task session attach failed", + ); + expect(connected.getStatus()).toBe("error"); + }); + it("negotiates the modern era under 'auto' with a populated discover result", async () => { const started = await startServer(); const connected = await connectWithEra(started.url, "auto"); @@ -357,7 +379,7 @@ describe("modern-era negotiation (2026-07-28)", () => { const { tools } = await connected.listTools(); const tool = tools.find((t) => t.name === "mrtr_loop"); await expect(connected.callTool(tool!, {})).rejects.toThrow( - /exceeded .* input_required rounds/, + /exceeded .* input[-_]required rounds/, ); }); diff --git a/clients/web/src/test/integration/mcp/inspectorClient-tasks-era.test.ts b/clients/web/src/test/integration/mcp/inspectorClient-tasks-era.test.ts index 6cb8413254..e8b53169c1 100644 --- a/clients/web/src/test/integration/mcp/inspectorClient-tasks-era.test.ts +++ b/clients/web/src/test/integration/mcp/inspectorClient-tasks-era.test.ts @@ -260,27 +260,6 @@ describe("tasks era fork (#1631)", () => { expect(methodsSent(messages)).toContain("tasks/update"); }); - it("bounds a never-completing input_required task with the round cap (#1631 review)", async () => { - const started = await startModernTasksServer(); - const { connected } = await connect(started.url, "modern"); - const { tools } = await connected.listTools(); - const tool = tools.find((t) => t.name === "modern_loop_task")!; - - // The server never advances past input_required, so the client re-prompts - // each poll; auto-answer, and the round cap must eventually abort instead - // of looping forever. - connected.addEventListener("newPendingElicitation", (event) => { - void event.detail.respond({ - action: "accept", - content: { approved: true }, - }); - }); - - await expect(connected.callToolStream(tool, {})).rejects.toThrow( - /exceeded \d+ input_required rounds/, - ); - }); - it("cancels a task paused at input_required — aborts the pending elicitation and unblocks the poll (#1631)", async () => { const started = await startModernTasksServer(); const { connected, messages } = await connect(started.url, "modern"); @@ -319,6 +298,40 @@ describe("tasks era fork (#1631)", () => { expect(connected.getPendingElicitations()).toHaveLength(0); }); + it("re-prompts every round of a looping input task, then stops at the round cap (R1)", async () => { + const started = await startModernTasksServer(); + const { connected, messages } = await connect(started.url, "modern"); + const { tools } = await connected.listTools(); + const tool = tools.find((t) => t.name === "modern_loop_task")!; + + // Auto-approve every round. A correct server emits a *distinct* input key + // per round; the ext-tasks SDK fingerprints keys and silently skips a + // recurring one, so a fixture that reused a single key (the R1 bug) would + // deliver round 1, swallow the answer, and poll forever — this call would + // hang rather than reject. + let rounds = 0; + connected.addEventListener("newPendingElicitation", (event) => { + rounds += 1; + void event.detail.respond({ + action: "accept", + content: { approved: true }, + }); + }); + + messages.length = 0; + // The client caps a modern task at MRTR_MAX_ROUNDS (10) input rounds and + // rejects once the server keeps asking past it. + await expect(connected.callTool(tool, {})).rejects.toThrow( + /input-required rounds/, + ); + // Every round must actually have reached us and been answered — not a + // single round followed by a dedup-induced infinite poll. + expect(rounds).toBeGreaterThan(1); + expect( + methodsSent(messages).filter((m) => m === "tasks/update").length, + ).toBeGreaterThan(1); + }); + it("passes a synchronous (non-task) tool result straight through", async () => { const started = await startModernTasksServer(); const { connected } = await connect(started.url, "modern"); diff --git a/clients/web/src/test/integration/mcp/inspectorClient-timeout-diagnostics.test.ts b/clients/web/src/test/integration/mcp/inspectorClient-timeout-diagnostics.test.ts index fd4649b642..c6db61a46d 100644 --- a/clients/web/src/test/integration/mcp/inspectorClient-timeout-diagnostics.test.ts +++ b/clients/web/src/test/integration/mcp/inspectorClient-timeout-diagnostics.test.ts @@ -169,7 +169,16 @@ describe("InspectorClient request-timeout diagnostics (#2318)", () => { let client: InspectorClient | null = null; let server: HangingServer | null = null; - async function connectTo(url: string): Promise<InspectorClient> { + /** + * `beforeConnect` runs on the constructed client before `connect()`, so a + * listener it attaches sees every fetch — including the notification-stream + * GET, which the SDK can open and the transport can log before `connect()` + * resolves (#2580). + */ + async function connectTo( + url: string, + beforeConnect?: (c: InspectorClient) => void, + ): Promise<InspectorClient> { client = new InspectorClient( { type: "streamable-http", url }, { @@ -180,6 +189,7 @@ describe("InspectorClient request-timeout diagnostics (#2318)", () => { timeout: REQUEST_TIMEOUT_MS, }, ); + beforeConnect?.(client); await client.connect(); return client; } @@ -260,8 +270,12 @@ describe("InspectorClient request-timeout diagnostics (#2318)", () => { it("emits connectionDiagnosticsChange as the state moves, and counts stream events", async () => { server = await startHangingServer(); - const c = await connectTo(server.url); - const log = new FetchRequestLogState(c); + // Subscribe the log before connecting: the GET can be logged before + // `connect()` resolves, and a log attached afterwards misses it (#2580). + let log!: FetchRequestLogState; + const c = await connectTo(server.url, (created) => { + log = new FetchRequestLogState(created); + }); const snapshots: ConnectionDiagnostics[] = []; c.addEventListener("connectionDiagnosticsChange", (event) => snapshots.push(event.detail), @@ -274,7 +288,10 @@ describe("InspectorClient request-timeout diagnostics (#2318)", () => { // Open from when the headers arrived, not from when the GET went out. const getEntryAtOpen = log .getFetchRequests() - .find((entry) => entry.method === "GET")!; + .find((entry) => entry.method === "GET"); + if (!getEntryAtOpen) { + throw new Error("the notification-stream GET was never logged"); + } expect(stream.openedAt).toBe( getEntryAtOpen.timestamp.getTime() + (getEntryAtOpen.duration ?? 0), ); diff --git a/clients/web/src/test/integration/mcp/inspectorClient.test.ts b/clients/web/src/test/integration/mcp/inspectorClient.test.ts index 0d89558d9c..728cd7a144 100644 --- a/clients/web/src/test/integration/mcp/inspectorClient.test.ts +++ b/clients/web/src/test/integration/mcp/inspectorClient.test.ts @@ -395,6 +395,23 @@ describe("InspectorClient", () => { }; } + it("sets error status when transport creation fails", async () => { + client = new InspectorClient( + { type: "stdio", command: "missing", args: [] }, + { + environment: { + transport: () => { + throw new Error("transport factory failed"); + }, + }, + }, + ); + + await expect(client.connect()).rejects.toThrow( + "transport factory failed", + ); + expect(client.getStatus()).toBe("error"); + }); it("rejects connect() with a timeout error when serverSettings.connectionTimeout fires", async () => { // Stub transport whose start() never resolves — simulates a slow / // unreachable upstream. InspectorClient.connect() should race against @@ -4549,68 +4566,6 @@ describe("InspectorClient", () => { expect(Array.isArray(result.tasks)).toBe(true); }); - it("should run tool as task (callTool with taskOptions returns task reference, poll getRequestorTask/getRequestorTaskResult yields result)", async () => { - // Same path as web App "Run as task": callTool with taskOptions -> task reference -> poll until completed - const optionalTaskTool = await getTool(client!, "optional_task"); - const invocation = await client!.callTool( - optionalTaskTool, - { message: "e2e-run-as-task" }, - undefined, - undefined, - { ttl: 5000 }, - ); - - expect(invocation.success).toBe(true); - expect(invocation.result).toBeDefined(); - expect(typeof invocation.result).toBe("object"); - const rawResult = invocation.result as Record<string, unknown>; - expect(rawResult.task).toBeDefined(); - const taskRef = rawResult.task as { - taskId: string; - status: string; - pollInterval?: number; - }; - expect(taskRef.taskId).toBeDefined(); - expect(typeof taskRef.taskId).toBe("string"); - expect(taskRef.taskId.length).toBeGreaterThan(0); - expect(taskRef.status).toBeDefined(); - expect(typeof taskRef.status).toBe("string"); - - const taskId = taskRef.taskId; - const pollIntervalMs = taskRef.pollInterval ?? 1000; - const timeoutMs = 12000; - const start = Date.now(); - let task = await client!.getRequestorTask(taskId); - while ( - task.status !== "completed" && - task.status !== "failed" && - task.status !== "cancelled" - ) { - expect(Date.now() - start).toBeLessThan(timeoutMs); - await new Promise((r) => setTimeout(r, pollIntervalMs)); - task = await client!.getRequestorTask(taskId); - } - - expect(task.status).toBe("completed"); - - const result = await client!.getRequestorTaskResult(taskId); - expect(result).toBeDefined(); - expect(result).toHaveProperty("content"); - expect(Array.isArray(result.content)).toBe(true); - expect(result.content.length).toBe(1); - const firstContent = result.content[0]; - expect(firstContent).toBeDefined(); - expect(firstContent!.type).toBe("text"); - expect(firstContent!).toHaveProperty("text"); - const resultText = JSON.parse((firstContent as { text: string }).text); - expect(resultText.message).toBe("Task completed: e2e-run-as-task"); - expect(resultText.taskId).toBe(taskId); - - const listResult = await client!.listRequestorTasks(); - const found = listResult.tasks.some((t) => t.taskId === taskId); - expect(found).toBe(true); - }); - it("should call tool with task support using callToolStream", async () => { const toolCallTaskUpdatedEvents: Array<{ taskId: string; @@ -5971,88 +5926,6 @@ describe("InspectorClient", () => { ); expect(c.getStatus()).toBe("disconnected"); }); - - it("receiver-task internals: TTL cleanup and cancel terminate via private surface", async () => { - // Drive the private createReceiverTask + cancelReceiverTask + TTL-cleanup - // paths by reaching into the instance. These are server-driven in - // practice (tasks/cancel from server), but the existing receiver-task - // e2e tests don't exercise the cancel path; this is the focused unit - // pass the issue suggested. - const c = new InspectorClient( - { - type: "stdio", - command: serverCommand.command, - args: serverCommand.args, - }, - { environment: { transport: createTransportNode } }, - ); - const internal = c as unknown as { - createReceiverTask: (opts: { - ttl?: number; - initialStatus: "input_required" | "working"; - statusMessage?: string; - }) => { - task: { taskId: string; status: string }; - payloadPromise: Promise<unknown>; - }; - cancelReceiverTask: (taskId: string) => { - taskId: string; - status: string; - }; - listReceiverTasks: () => Array<{ taskId: string; status: string }>; - getReceiverTask: (taskId: string) => unknown; - getReceiverTaskPayload: (taskId: string) => Promise<unknown>; - receiverTaskRecords: Map<string, unknown>; - }; - - // Short TTL so the cleanup setTimeout fires in-test - const record = internal.createReceiverTask({ - ttl: 50, - initialStatus: "working", - statusMessage: "running", - }); - // Capture the rejection's message for the assertion below. (Not for - // unhandled-rejection suppression — `createReceiverTask` marks the - // promise handled at the source.) - const payloadResult = record.payloadPromise.catch( - (e) => (e as Error).message, - ); - expect(record.task.taskId).toBeDefined(); - // listReceiverTasks contains the new task - const list = internal.listReceiverTasks(); - expect(list.some((t) => t.taskId === record.task.taskId)).toBe(true); - expect(internal.getReceiverTask(record.task.taskId)).toBeDefined(); - - // getReceiverTaskPayload on an unknown id throws InvalidParams - await expect( - internal.getReceiverTaskPayload("does-not-exist"), - ).rejects.toThrow(/Unknown taskId/); - - // Cancel before TTL fires - const cancelled = internal.cancelReceiverTask(record.task.taskId); - expect(cancelled.status).toBe("cancelled"); - await expect(payloadResult).resolves.toBe("Task cancelled"); - // Cancel again — record is in terminal state, returns existing task - const reCancel = internal.cancelReceiverTask(record.task.taskId); - expect(reCancel.status).toBe("cancelled"); - - // cancelReceiverTask on an unknown id throws InvalidParams - expect(() => internal.cancelReceiverTask("nope")).toThrow( - /Unknown taskId/, - ); - - // Drive the TTL-cleanup path: create a record with very short ttl and - // let setTimeout fire — receiverTaskRecords drops the entry. - const ttlRecord = internal.createReceiverTask({ - ttl: 20, - initialStatus: "working", - statusMessage: "running", - }); - await new Promise<void>((r) => setTimeout(r, 80)); - expect(internal.receiverTaskRecords.has(ttlRecord.task.taskId)).toBe( - false, - ); - }); }); describe("defensive guards when client is uninitialized", () => { diff --git a/clients/web/src/test/integration/server/health.test.ts b/clients/web/src/test/integration/server/health.test.ts new file mode 100644 index 0000000000..cc1973a672 --- /dev/null +++ b/clients/web/src/test/integration/server/health.test.ts @@ -0,0 +1,239 @@ +/** + * Tests for the web backend's `GET /healthz` probe (#2438). + * + * Two layers: the pure helpers in `server/health.ts` (the request matcher and + * the response builder, which the dev Vite middleware and the prod Hono route + * both build on), and the route as the production server actually serves it, + * started for real via `startHonoServer`. The second layer is what pins the + * contract the route exists for: unauthenticated, answered ahead of the SPA + * fallback, disclosing nothing but `{"status":"ok"}` (never the API token that + * `GET /` embeds), and leaving every `/api/*` route behind its auth check. + * + * The dev Vite middleware's path is covered through `handleNodeHealthRequest`, + * the Node-level adapter it delegates to, driven by a real `node:http` server + * — `vite-hono-plugin.ts` itself is excluded from coverage as runtime glue + * that needs a live Vite server. + * + * It lives in the `integration` project because it binds real listeners (the + * HTTP server plus the sandbox and app-origin servers it starts). + */ + +import { describe, it, expect, beforeAll, afterAll } from "vitest"; +import { createServer } from "node:net"; +import { createServer as createHttpServer, type Server } from "node:http"; +import { mkdtemp, writeFile, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { + HEALTH_BODY, + HEALTH_PATH, + handleNodeHealthRequest, + healthResponse, + isHealthRequest, +} from "../../../../server/health.js"; +import { startHonoServer } from "../../../../server/server.js"; +import type { WebServerConfig } from "../../../../server/web-server-config.js"; +import type { WebServerHandle } from "../../../../server/types.js"; +import { INSPECTOR_API_TOKEN_GLOBAL } from "../../../../../../core/mcp/remote/constants.js"; + +// Ask the OS for an ephemeral port, then release it for the server to claim +// (same pattern as server-token-injection.test.ts). +async function findFreePort(): Promise<number> { + return new Promise((resolve, reject) => { + const srv = createServer(); + srv.unref(); + srv.on("error", reject); + srv.listen(0, "127.0.0.1", () => { + const addr = srv.address(); + if (addr && typeof addr === "object") { + const { port } = addr; + srv.close(() => resolve(port)); + } else { + srv.close(() => reject(new Error("Could not resolve a free port"))); + } + }); + }); +} + +describe("isHealthRequest", () => { + it("matches GET and HEAD on the health path, case-insensitively by method", () => { + expect(isHealthRequest("GET", HEALTH_PATH)).toBe(true); + expect(isHealthRequest("HEAD", HEALTH_PATH)).toBe(true); + expect(isHealthRequest("get", HEALTH_PATH)).toBe(true); + }); + + it("ignores a query string or fragment", () => { + expect(isHealthRequest("GET", `${HEALTH_PATH}?t=123`)).toBe(true); + expect(isHealthRequest("GET", `${HEALTH_PATH}#x`)).toBe(true); + }); + + it("rejects other methods", () => { + expect(isHealthRequest("POST", HEALTH_PATH)).toBe(false); + expect(isHealthRequest(undefined, HEALTH_PATH)).toBe(false); + }); + + it("rejects other paths, including a trailing slash, a sub-path and a missing url", () => { + expect(isHealthRequest("GET", `${HEALTH_PATH}/`)).toBe(false); + expect(isHealthRequest("GET", `${HEALTH_PATH}/x`)).toBe(false); + expect(isHealthRequest("GET", "/api/healthz")).toBe(false); + expect(isHealthRequest("GET", "/")).toBe(false); + expect(isHealthRequest("GET", undefined)).toBe(false); + }); +}); + +describe("healthResponse", () => { + it("answers GET with 200, the fixed body, and no-store", async () => { + const res = healthResponse("GET"); + expect(res.status).toBe(200); + expect(res.headers.get("cache-control")).toBe("no-store"); + expect(res.headers.get("content-type")).toContain("application/json"); + expect(await res.json()).toEqual({ status: "ok" }); + }); + + it("defaults to GET", async () => { + expect(await healthResponse().json()).toEqual(HEALTH_BODY); + }); + + it("answers HEAD with the same status and headers but no body", async () => { + const res = healthResponse("head"); + expect(res.status).toBe(200); + expect(res.headers.get("cache-control")).toBe("no-store"); + expect(await res.text()).toBe(""); + }); + + it("keeps the shared body immutable", () => { + expect(Object.isFrozen(HEALTH_BODY)).toBe(true); + }); +}); + +// The dev Vite middleware's path: a real node:http server whose handler calls +// `handleNodeHealthRequest` first and falls through (`next()`) otherwise. +describe("handleNodeHealthRequest (dev middleware path)", () => { + let server: Server; + let baseUrl: string; + + beforeAll(async () => { + server = createHttpServer((req, res) => { + if (handleNodeHealthRequest(req, res)) return; + res.writeHead(418); + res.end("fell through"); + }); + await new Promise<void>((resolve) => + server.listen(0, "127.0.0.1", () => resolve()), + ); + const addr = server.address(); + if (!addr || typeof addr !== "object") throw new Error("no address"); + baseUrl = `http://127.0.0.1:${addr.port}`; + }); + + afterAll(async () => { + await new Promise<void>((resolve) => server.close(() => resolve())); + }); + + it("answers GET with 200, the fixed body and the health headers", async () => { + const res = await fetch(`${baseUrl}${HEALTH_PATH}?t=1`); + expect(res.status).toBe(200); + expect(res.headers.get("cache-control")).toBe("no-store"); + expect(res.headers.get("content-type")).toContain("application/json"); + expect(await res.json()).toEqual({ status: "ok" }); + }); + + it("answers HEAD with 200 and no body", async () => { + const res = await fetch(`${baseUrl}${HEALTH_PATH}`, { method: "HEAD" }); + expect(res.status).toBe(200); + expect(res.headers.get("cache-control")).toBe("no-store"); + expect(await res.text()).toBe(""); + }); + + it("returns false and leaves other requests to the caller", async () => { + const other = await fetch(`${baseUrl}/api/config`); + expect(other.status).toBe(418); + expect(await other.text()).toBe("fell through"); + const post = await fetch(`${baseUrl}${HEALTH_PATH}`, { method: "POST" }); + expect(post.status).toBe(418); + }); + + it("writes nothing for a request with no method", () => { + const untouched = (): never => { + throw new Error("the response must not be written"); + }; + expect( + handleNodeHealthRequest( + { method: undefined, url: HEALTH_PATH }, + { writeHead: untouched, end: untouched }, + ), + ).toBe(false); + }); +}); + +const TOKEN = "test-health-token-1234567890"; + +describe("startHonoServer GET /healthz", () => { + let handle: WebServerHandle; + let baseUrl: string; + let staticRoot: string; + + beforeAll(async () => { + staticRoot = await mkdtemp(join(tmpdir(), "inspector-health-")); + await writeFile( + join(staticRoot, "index.html"), + "<!doctype html><html><head></head><body></body></html>", + "utf-8", + ); + const port = await findFreePort(); + baseUrl = `http://127.0.0.1:${port}`; + const config: WebServerConfig = { + port, + hostname: "127.0.0.1", + authToken: TOKEN, + dangerouslyOmitAuth: false, + initialMcpConfig: null, + mcpConfigPath: undefined, + writable: true, + initialServers: null, + storageDir: undefined, + allowedOrigins: [baseUrl], + sandboxPort: 0, + appOriginPort: 0, + sandboxHost: "127.0.0.1", + logger: undefined, + autoOpen: false, + staticRoot, + }; + handle = await startHonoServer(config); + }); + + afterAll(async () => { + await handle?.close(); + if (staticRoot) await rm(staticRoot, { recursive: true, force: true }); + }); + + it("answers 200 with only the fixed status body and no auth token", async () => { + const res = await fetch(`${baseUrl}${HEALTH_PATH}`); + expect(res.status).toBe(200); + expect(res.headers.get("cache-control")).toBe("no-store"); + const text = await res.text(); + expect(JSON.parse(text)).toEqual({ status: "ok" }); + // Unauthenticated, so it must not leak the token the way `/` embeds it. + expect(text).not.toContain(TOKEN); + expect(text).not.toContain(INSPECTOR_API_TOKEN_GLOBAL); + }); + + it("answers HEAD with 200 and no body", async () => { + const res = await fetch(`${baseUrl}${HEALTH_PATH}`, { method: "HEAD" }); + expect(res.status).toBe(200); + expect(await res.text()).toBe(""); + }); + + it("is answered from any Origin, since it sits outside the /api allow-list", async () => { + const res = await fetch(`${baseUrl}${HEALTH_PATH}`, { + headers: { origin: "http://not-allowed.example" }, + }); + expect(res.status).toBe(200); + }); + + it("leaves /api/* authenticated", async () => { + const res = await fetch(`${baseUrl}/api/config`); + expect(res.status).toBe(401); + }); +}); diff --git a/clients/web/src/test/integration/server/server-auto-open.test.ts b/clients/web/src/test/integration/server/server-auto-open.test.ts index 003ece9f6d..c152ba3de6 100644 --- a/clients/web/src/test/integration/server/server-auto-open.test.ts +++ b/clients/web/src/test/integration/server/server-auto-open.test.ts @@ -15,6 +15,7 @@ * `vite-hono-plugin.ts`) — but it calls the same tested helper. */ import { describe, it, expect, afterEach, vi } from "vitest"; +import { EventEmitter } from "node:events"; import { createServer } from "node:net"; import { mkdtemp, writeFile, rm } from "node:fs/promises"; import { tmpdir } from "node:os"; @@ -31,6 +32,19 @@ const { startHonoServer } = await import("../../../../server/server.js"); const WARNING = "Could not open a browser automatically:"; +/** + * A stand-in for the `ChildProcess` `open` resolves with: like a real spawn it + * reports on a later `process.nextTick` — `'spawn'` when the opener launched, + * `'error'` when it could not be spawned (#2533). + */ +function fakeChild(outcome: "spawn" | Error = "spawn"): EventEmitter { + const child = new EventEmitter(); + process.nextTick(() => + outcome === "spawn" ? child.emit("spawn") : child.emit("error", outcome), + ); + return child; +} + async function findFreePort(): Promise<number> { return new Promise((resolve, reject) => { const srv = createServer(); @@ -87,9 +101,22 @@ describe("openBrowser", () => { expect(warn).toHaveBeenCalledWith(WARNING, expect.any(Error)); }); + it("warns instead of crashing when the opener cannot be spawned", async () => { + // `open` resolves first and the ENOENT arrives afterwards as an 'error' + // event on the child; unlistened, that event throws and kills the server. + const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); + const enoent = Object.assign(new Error("spawn xdg-open ENOENT"), { + code: "ENOENT", + }); + openMock.mockImplementation(async () => fakeChild(enoent)); + expect(openBrowser("http://127.0.0.1:6274")).toBeUndefined(); + await settle(); + expect(warn).toHaveBeenCalledWith(WARNING, enoent); + }); + it("stays quiet when the browser opens", async () => { const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); - openMock.mockResolvedValue(undefined); + openMock.mockImplementation(async () => fakeChild()); openBrowser("http://127.0.0.1:6274"); await settle(); expect(openMock).toHaveBeenCalledWith("http://127.0.0.1:6274"); diff --git a/clients/web/src/test/integration/storage/adapters.test.ts b/clients/web/src/test/integration/storage/adapters.test.ts index 2cb50c25db..9005f4697e 100644 --- a/clients/web/src/test/integration/storage/adapters.test.ts +++ b/clients/web/src/test/integration/storage/adapters.test.ts @@ -97,7 +97,7 @@ describe("OAuth persistence", () => { expect( JSON.parse( (await secretStore.get( - oauthSecretServerId("https://example.com"), + oauthSecretServerId("https://example.com", parsed.secretsNamespace), LEGACY_TOKENS_FIELD, ))!, ), @@ -319,7 +319,9 @@ describe("OAuth persistence", () => { ); await flushStoreFileWrites(filePath); const parsed = JSON.parse(readFileSync(filePath, "utf-8")); + // The first write also stamps the file's secrets namespace (#2549). expect(parsed).toEqual({ + secretsNamespace: expect.any(String), servers: { "https://mine.example": { scope: "mine" } }, idpSessions: {}, }); @@ -569,7 +571,7 @@ describe("OAuth persistence", () => { expect(raw.servers["https://example.com"].tokens).toBeUndefined(); expect( await secretStore.get( - oauthSecretServerId("https://example.com"), + oauthSecretServerId("https://example.com", raw.secretsNamespace), LEGACY_TOKENS_FIELD, ), ).not.toBeNull(); @@ -634,12 +636,13 @@ describe("OAuth persistence", () => { idpSessions: {}, }), }); - expect( - await secretStore.get( - oauthSecretServerId("https://example.com"), - LEGACY_TOKENS_FIELD, - ), - ).not.toBeNull(); + // The id is scoped by the file's adopted namespace; capture it while + // the file still exists (#2549). + const { secretsNamespace } = JSON.parse( + readFileSync(join(tempDir, "oauth.json"), "utf-8"), + ) as { secretsNamespace: string }; + const id = oauthSecretServerId("https://example.com", secretsNamespace); + expect(await secretStore.get(id, LEGACY_TOKENS_FIELD)).not.toBeNull(); const del = await fetch(`${baseUrl}/api/storage/oauth`, { method: "DELETE", @@ -647,12 +650,7 @@ describe("OAuth persistence", () => { }); expect(del.status).toBe(200); expect(existsSync(join(tempDir, "oauth.json"))).toBe(false); - expect( - await secretStore.get( - oauthSecretServerId("https://example.com"), - LEGACY_TOKENS_FIELD, - ), - ).toBeNull(); + expect(await secretStore.get(id, LEGACY_TOKENS_FIELD)).toBeNull(); }); }); }); diff --git a/clients/web/src/test/integration/storage/oauth-secret-split.test.ts b/clients/web/src/test/integration/storage/oauth-secret-split.test.ts index 24afda57d1..0857a572a1 100644 --- a/clients/web/src/test/integration/storage/oauth-secret-split.test.ts +++ b/clients/web/src/test/integration/storage/oauth-secret-split.test.ts @@ -95,6 +95,24 @@ function readRawFile(): OAuthPersistSnapshot { return JSON.parse(readFileSync(filePath, "utf8")) as OAuthPersistSnapshot; } +/** The secrets namespace the file's first write adopted (#2549). */ +function fileNamespace(): string { + const { secretsNamespace } = JSON.parse(readFileSync(filePath, "utf8")) as { + secretsNamespace: string; + }; + return secretsNamespace; +} + +/** Store id for a server under the file's adopted namespace. */ +function idOf(url: string): string { + return oauthSecretServerId(url, fileNamespace()); +} + +/** Store id for an IdP issuer under the file's adopted namespace. */ +function idpIdOf(issuer: string): string { + return oauthIdpSecretServerId(issuer, fileNamespace()); +} + describe("writeOAuthSections secret split", () => { it("writes only residue to the file and secrets to the store", async () => { const store = new InMemorySecretStore(); @@ -107,7 +125,7 @@ describe("writeOAuthSections secret split", () => { expect(raw.servers[SERVER]!.clientInformation).toEqual({ client_id: "cid", }); - const id = oauthSecretServerId(SERVER); + const id = idOf(SERVER); expect(JSON.parse((await store.get(id, LEGACY_TOKENS_FIELD))!)).toEqual( TOKENS, ); @@ -134,9 +152,7 @@ describe("writeOAuthSections secret split", () => { const raw = readRawFile(); expect(raw.idpSessions[ISSUER]).toEqual({ idTokenExpiresAt: 9 }); - expect( - await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), - ).not.toBeNull(); + expect(await store.get(idpIdOf(ISSUER), IDP_SESSION_FIELD)).not.toBeNull(); const joined = await readOAuthStore(filePath, store); expect(joined?.idpSessions[ISSUER]).toEqual({ @@ -157,7 +173,7 @@ describe("writeOAuthSections secret split", () => { ); await flushStoreFileWrites(filePath); - const id = oauthSecretServerId(SERVER); + const id = idOf(SERVER); expect(await store.get(id, LEGACY_TOKENS_FIELD)).toBeNull(); expect(await store.get(id, LEGACY_CLIENT_SECRET_FIELD)).toBeNull(); expect(readRawFile().servers[SERVER]).toBeUndefined(); @@ -183,7 +199,7 @@ describe("writeOAuthSections secret split", () => { { servers: [SERVER] }, store, ); - const id = oauthSecretServerId(SERVER); + const id = idOf(SERVER); expect(await store.get(id, issuerTokensField(ISSUER))).not.toBeNull(); await writeOAuthSections( @@ -198,10 +214,10 @@ describe("writeOAuthSections secret split", () => { it("enforces the persist-tokens policy and self-cleans on downgrade", async () => { const store = new InMemorySecretStore(); - const id = oauthSecretServerId(SERVER); process.env[PERSIST_TOKENS_ENV] = "access"; await writeOAuthSections(filePath, snapshotWith(), undefined, store); + const id = idOf(SERVER); expect(JSON.parse((await store.get(id, LEGACY_TOKENS_FIELD))!)).toEqual({ access_token: "at", token_type: "Bearer", @@ -222,7 +238,7 @@ describe("writeOAuthSections secret split", () => { undefined, store, ); - const id = oauthIdpSecretServerId(ISSUER); + const id = idpIdOf(ISSUER); expect(await store.get(id, IDP_SESSION_FIELD)).not.toBeNull(); // Sections naming only idpSessions also exercises the servers-omitted @@ -395,7 +411,7 @@ describe("writeOAuthSections secret split", () => { // The store holds the *old* secrets again, matching the old residue // still on disk — no cid/cs2 mismatch on the next joined read. - const id = oauthSecretServerId(SERVER); + const id = idOf(SERVER); expect(JSON.parse((await store.get(id, LEGACY_TOKENS_FIELD))!)).toEqual( TOKENS, ); @@ -434,9 +450,9 @@ describe("writeOAuthSections secret split", () => { chmodSync(tempDir, 0o755); } - expect( - await store.get(oauthSecretServerId(SERVER), LEGACY_CLIENT_SECRET_FIELD), - ).toBe("cs"); + expect(await store.get(idOf(SERVER), LEGACY_CLIENT_SECRET_FIELD)).toBe( + "cs", + ); }); it("aborts the write when a store delete fails, keeping the old residue", async () => { @@ -1016,15 +1032,13 @@ describe("removeOAuthStore", () => { store, ); await flushStoreFileWrites(filePath); + const id = idOf(SERVER); + const idpId = idpIdOf(ISSUER); await removeOAuthStore(filePath, store); expect(existsSync(filePath)).toBe(false); - expect( - await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD), - ).toBeNull(); - expect( - await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), - ).toBeNull(); + expect(await store.get(id, LEGACY_TOKENS_FIELD)).toBeNull(); + expect(await store.get(idpId, IDP_SESSION_FIELD)).toBeNull(); }); it("propagates a failed purge and leaves the file as the index", async () => { @@ -1078,13 +1092,9 @@ describe("removeOAuthStore", () => { // The first target's purged secrets were restored — a retry of the // removal (or a plain read) still finds everything the file indexes. expect( - JSON.parse( - (await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD))!, - ), + JSON.parse((await store.get(idOf(SERVER), LEGACY_TOKENS_FIELD))!), ).toEqual(TOKENS); - expect( - await store.get(oauthIdpSecretServerId(ISSUER), IDP_SESSION_FIELD), - ).not.toBeNull(); + expect(await store.get(idpIdOf(ISSUER), IDP_SESSION_FIELD)).not.toBeNull(); }); it("restores purged secrets when the file delete fails", async () => { @@ -1104,13 +1114,11 @@ describe("removeOAuthStore", () => { // The file survives as the index and the store matches it again. expect(existsSync(filePath)).toBe(true); expect( - JSON.parse( - (await store.get(oauthSecretServerId(SERVER), LEGACY_TOKENS_FIELD))!, - ), + JSON.parse((await store.get(idOf(SERVER), LEGACY_TOKENS_FIELD))!), ).toEqual(TOKENS); - expect( - await store.get(oauthSecretServerId(SERVER), LEGACY_CLIENT_SECRET_FIELD), - ).toBe("cs"); + expect(await store.get(idOf(SERVER), LEGACY_CLIENT_SECRET_FIELD)).toBe( + "cs", + ); }); it("is a no-op purge for a missing file", async () => { @@ -1268,12 +1276,12 @@ describe("partial token payloads round-trip through the store", () => { it("a save moves a partial token payload to the store and serves it back", async () => { const store = new InMemorySecretStore(); - const id = oauthSecretServerId(SERVER); const snapshot = snapshotWith(); snapshot.servers[SERVER]!.tokens = { ...PARTIAL } as never; await writeOAuthSections(filePath, snapshot, { servers: [SERVER] }, store); await flushStoreFileWrites(filePath); + const id = idOf(SERVER); // The bearer-grade refresh token is in the store, not the file. expect(readRawFile().servers[SERVER]!.tokens).toBeUndefined(); @@ -1386,7 +1394,7 @@ describe("saves only touch changed store fields", () => { expect(raw.servers[SERVER]!.tokens).toBeUndefined(); // The store still holds the unchanged credentials, untouched. - const id = oauthSecretServerId(SERVER); + const id = idOf(SERVER); expect(JSON.parse((await store.get(id, LEGACY_TOKENS_FIELD))!)).toEqual( TOKENS, ); diff --git a/clients/web/src/test/integration/storage/oauth-write-convergence.test.ts b/clients/web/src/test/integration/storage/oauth-write-convergence.test.ts index 22b7985c09..f733d8dc7c 100644 --- a/clients/web/src/test/integration/storage/oauth-write-convergence.test.ts +++ b/clients/web/src/test/integration/storage/oauth-write-convergence.test.ts @@ -44,13 +44,16 @@ vi.mock("@inspector/core/storage/store-io.js", async (importOriginal) => { >(); return { ...actual, + // The hooks simulate a writer racing on the *state file*; the + // namespace ledger beside it (#2560) is not part of that race. writeStoreFile: vi.fn(async (filePath: string, data: string) => { - hook.beforeWrite?.(filePath, data); + const raced = !filePath.endsWith(".namespaces.json"); + if (raced) hook.beforeWrite?.(filePath, data); await actual.writeStoreFile(filePath, data); - await hook.afterWrite?.(filePath, data); + if (raced) await hook.afterWrite?.(filePath, data); }), readStoreFile: vi.fn(async (filePath: string) => { - hook.beforeRead?.(filePath); + if (!filePath.endsWith(".namespaces.json")) hook.beforeRead?.(filePath); return actual.readStoreFile(filePath); }), }; @@ -82,6 +85,15 @@ let filePath: string; let store: InMemorySecretStore; /** File bytes holding only server A, as the racing writer would leave them. */ let onlyA: string; +/** Store id for a url under the seed write's adopted secrets namespace. */ +let idOf: (url: string) => string; + +/** Writes to the state file itself, excluding its namespace ledger. */ +function stateFileWrites(): number { + return vi + .mocked(writeStoreFile) + .mock.calls.filter(([path]) => path === filePath).length; +} beforeEach(async () => { tempDir = mkdtempSync(join(tmpdir(), "inspector-oauth-converge-")); @@ -98,6 +110,10 @@ beforeEach(async () => { store, ); onlyA = readFileSync(filePath, "utf-8"); + const { secretsNamespace } = JSON.parse(onlyA) as { + secretsNamespace: string; + }; + idOf = (url) => oauthSecretServerId(url, secretsNamespace); }); afterEach(() => { @@ -124,12 +140,106 @@ describe("writeOAuthSections convergence verification", () => { ); // Seed + first (clobbered) attempt + converging retry. - expect(vi.mocked(writeStoreFile)).toHaveBeenCalledTimes(3); + expect(stateFileWrites()).toBe(3); const read = await readOAuthStore(filePath, store); expect(read?.servers[SERVER_A]?.tokens?.access_token).toBe("at-a"); expect(read?.servers[SERVER_B]?.tokens?.access_token).toBe("at-b"); }); + it("converges onto a concurrent adopter's namespace instead of re-stamping its own", async () => { + // The degraded-lock first-write race: another writer adopted a different + // namespace and won the file between our write and read-back. The retry + // must re-key to the namespace observed on disk — re-stamping our own + // mint would ping-pong and strand the other writer's secrets. + const racingNs = "99999999-9999-4999-8999-999999999999"; + const racing = JSON.parse(onlyA) as Record<string, unknown>; + racing.secretsNamespace = racingNs; + // The racing writer's own store entries live under its namespace. + await store.set( + oauthSecretServerId(SERVER_A, racingNs), + LEGACY_TOKENS_FIELD, + JSON.stringify({ access_token: "at-racing", token_type: "Bearer" }), + ); + let clobbers = 0; + hook.afterWrite = (path) => { + if (clobbers++ === 0) writeFileSync(path, JSON.stringify(racing)); + }; + + await writeOAuthSections( + filePath, + snapshotOf({ [SERVER_B]: serverState("b") }), + { servers: [SERVER_B] }, + store, + ); + + // Seed + clobbered attempt + converging retry. + expect(stateFileWrites()).toBe(3); + const final = JSON.parse(readFileSync(filePath, "utf-8")) as { + secretsNamespace: string; + }; + expect(final.secretsNamespace).toBe(racingNs); + // B's secrets were written under the adopted namespace… + expect( + await store.get( + oauthSecretServerId(SERVER_B, racingNs), + LEGACY_TOKENS_FIELD, + ), + ).not.toBeNull(); + // …and the abandoned attempt's writes under our own mint were unwound. + expect(await store.get(idOf(SERVER_B), LEGACY_TOKENS_FIELD)).toBeNull(); + expect( + await store.get(idOf(SERVER_B), LEGACY_CLIENT_SECRET_FIELD), + ).toBeNull(); + // Joined read-back sees both writers' entries under the one namespace. + const read = await readOAuthStore(filePath, store); + expect(read?.servers[SERVER_B]?.tokens?.access_token).toBe("at-b"); + expect(read?.servers[SERVER_A]?.tokens?.access_token).toBe("at-racing"); + }); + + it("aborts the namespace re-key when the baseline restore fails, instead of reporting success", async () => { + // Same race as above, but unwinding the abandoned attempt's store + // writes fails. Carrying on would clear the rollback baseline and let + // the save report success with those writes stranded under a namespace + // no file references — the re-key must abort instead, leaving a loud + // failure a retried save can converge from. + const racingNs = "99999999-9999-4999-8999-999999999999"; + const racing = JSON.parse(onlyA) as Record<string, unknown>; + racing.secretsNamespace = racingNs; + let clobbers = 0; + hook.afterWrite = (path) => { + if (clobbers++ === 0) writeFileSync(path, JSON.stringify(racing)); + }; + // The re-key restore deletes the abandoned attempt's new-entry writes + // (their baseline is "absent"). Refuse the first such delete once; the + // failure-path rollback that follows retries it and succeeds. + const realDelete = store.delete.bind(store); + let refused = false; + vi.spyOn(store, "delete").mockImplementation(async (serverId, field) => { + if (!refused && serverId === idOf(SERVER_B)) { + refused = true; + throw new Error("keychain delete refused"); + } + await realDelete(serverId, field); + }); + + await expect( + writeOAuthSections( + filePath, + snapshotOf({ [SERVER_B]: serverState("b") }), + { servers: [SERVER_B] }, + store, + ), + ).rejects.toThrow("keychain delete refused"); + + // The failure-path rollback unwound the abandoned writes after all — + // nothing is stranded under our mint, and the racing file stands. + expect(await store.get(idOf(SERVER_B), LEGACY_TOKENS_FIELD)).toBeNull(); + const final = JSON.parse(readFileSync(filePath, "utf-8")) as { + secretsNamespace: string; + }; + expect(final.secretsNamespace).toBe(racingNs); + }); + it("gives up with a typed, retryable error when the file keeps changing", async () => { hook.afterWrite = (path) => writeFileSync(path, onlyA); @@ -143,7 +253,7 @@ describe("writeOAuthSections convergence verification", () => { ).rejects.toThrow(SecretStoreUnavailableError); // Seed + five attempts, then the bounded loop reports instead of spinning. - expect(vi.mocked(writeStoreFile)).toHaveBeenCalledTimes(6); + expect(stateFileWrites()).toBe(6); }); it("unwinds a new entry's store secrets when it gives up, so nothing is stranded without a file index", async () => { @@ -160,11 +270,11 @@ describe("writeOAuthSections convergence verification", () => { // Server B never made it into the file, so its secrets must not linger // in the store (they would have no index for removeOAuthStore to find). - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); expect(await store.get(idB, LEGACY_TOKENS_FIELD)).toBeNull(); expect(await store.get(idB, LEGACY_CLIENT_SECRET_FIELD)).toBeNull(); // Server A's stored secrets are untouched. - const idA = oauthSecretServerId(SERVER_A); + const idA = idOf(SERVER_A); expect(await store.get(idA, LEGACY_TOKENS_FIELD)).not.toBeNull(); expect(await store.get(idA, LEGACY_CLIENT_SECRET_FIELD)).toBe("cs-a"); }); @@ -192,10 +302,10 @@ describe("writeOAuthSections convergence verification", () => { ), ).rejects.toThrow(/disk full/); - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); expect(await store.get(idB, LEGACY_TOKENS_FIELD)).toBeNull(); expect(await store.get(idB, LEGACY_CLIENT_SECRET_FIELD)).toBeNull(); - const idA = oauthSecretServerId(SERVER_A); + const idA = idOf(SERVER_A); expect(await store.get(idA, LEGACY_CLIENT_SECRET_FIELD)).toBe("cs-a"); expect(readFileSync(filePath, "utf-8")).toBe(onlyA); }); @@ -216,10 +326,10 @@ describe("writeOAuthSections convergence verification", () => { ), ).rejects.toThrow(/refusing/i); - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); expect(await store.get(idB, LEGACY_TOKENS_FIELD)).toBeNull(); expect(await store.get(idB, LEGACY_CLIENT_SECRET_FIELD)).toBeNull(); - const idA = oauthSecretServerId(SERVER_A); + const idA = idOf(SERVER_A); expect(await store.get(idA, LEGACY_CLIENT_SECRET_FIELD)).toBe("cs-a"); }); @@ -230,7 +340,7 @@ describe("writeOAuthSections convergence verification", () => { // writer. The rollback baseline folds each attempt's priors, telling our // own earlier attempt's writes (equal to what this call writes — they // are constant across attempts) apart from foreign values. - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); await writeOAuthSections( filePath, snapshotOf({ [SERVER_B]: serverState("b") }), @@ -278,7 +388,7 @@ describe("writeOAuthSections convergence verification", () => { // pairs with — so the failure escalates into the reconciling exit, which // finds the file changed and restores the pre-operation secrets. const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); await writeOAuthSections( filePath, snapshotOf({ [SERVER_B]: serverState("b") }), @@ -337,7 +447,7 @@ describe("writeOAuthSections convergence verification", () => { // so attempt 2's store writes must be rolled back to the baseline, not // left in place under the old residue. const warn = vi.spyOn(console, "warn").mockImplementation(() => {}); - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); await writeOAuthSections( filePath, snapshotOf({ [SERVER_B]: serverState("b") }), @@ -358,11 +468,12 @@ describe("writeOAuthSections convergence verification", () => { let reads = 0; hook.beforeRead = () => { reads += 1; - // Read 1: attempt 1's disk read. Read 2: its failing read-back. - // Read 3: attempt 2's disk read — the store has recovered by now. - // Read 4: the reconciling exit's confirmation read. - if (reads === 2) throw new Error("EIO: read failed"); - if (reads === 3) failNewSets = false; + // Read 1: the save's namespace-adoption read. Read 2: attempt 1's + // disk read. Read 3: its failing read-back. Read 4: attempt 2's disk + // read — the store has recovered by now. Read 5: the reconciling + // exit's confirmation read. + if (reads === 3) throw new Error("EIO: read failed"); + if (reads === 4) failNewSets = false; }; let writes = 0; hook.beforeWrite = () => { @@ -393,7 +504,7 @@ describe("writeOAuthSections convergence verification", () => { // reconciling exit. The file is confirmed to still hold attempt 1's // write, so the save is committed: the store is re-pointed at attempt // 1's values and the call reports success. - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); await writeOAuthSections( filePath, snapshotOf({ [SERVER_B]: serverState("b") }), @@ -410,12 +521,13 @@ describe("writeOAuthSections convergence verification", () => { let reads = 0; hook.beforeRead = () => { reads += 1; - // Read 1: attempt 1's disk read. Read 2: its failing read-back. - // Read 3: attempt 2's disk read — the store starts flaking here. - // Read 4: the confirmation read — the flake has passed. - if (reads === 2) throw new Error("EIO: read failed"); - if (reads === 3) failSets = true; - if (reads === 4) failSets = false; + // Read 1: the save's namespace-adoption read. Read 2: attempt 1's + // disk read. Read 3: its failing read-back. Read 4: attempt 2's disk + // read — the store starts flaking here. Read 5: the confirmation + // read — the flake has passed. + if (reads === 3) throw new Error("EIO: read failed"); + if (reads === 4) failSets = true; + if (reads === 5) failSets = false; }; await writeOAuthSections( @@ -441,10 +553,11 @@ describe("writeOAuthSections convergence verification", () => { let reads = 0; hook.beforeRead = () => { reads += 1; - // Read 1: attempt 1's disk read. Reads 2-3: attempt 1's verifying - // read-back and attempt 2's disk read, both failing. Read 4: the - // reconciling exit's confirmation read, which succeeds. - if (reads === 2 || reads === 3) throw new Error("EIO: read failed"); + // Read 1: the save's namespace-adoption read. Read 2: attempt 1's + // disk read. Reads 3-4: attempt 1's verifying read-back and attempt + // 2's disk read, both failing. Read 5: the reconciling exit's + // confirmation read, which succeeds. + if (reads === 3 || reads === 4) throw new Error("EIO: read failed"); }; await writeOAuthSections( @@ -454,7 +567,7 @@ describe("writeOAuthSections convergence verification", () => { store, ); - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); expect(await store.get(idB, LEGACY_TOKENS_FIELD)).toContain("at-b"); expect(await store.get(idB, LEGACY_CLIENT_SECRET_FIELD)).toBe("cs-b"); const read = await readOAuthStore(filePath, store); @@ -470,7 +583,9 @@ describe("writeOAuthSections convergence verification", () => { let reads = 0; hook.beforeRead = () => { reads += 1; - if (reads >= 2) throw new Error("EIO: read failed"); + // Read 1 is the save's namespace-adoption read, read 2 attempt 1's + // disk read; everything after fails. + if (reads >= 3) throw new Error("EIO: read failed"); }; await expect( @@ -482,7 +597,7 @@ describe("writeOAuthSections convergence verification", () => { ), ).rejects.toThrow(/EIO/); - const idB = oauthSecretServerId(SERVER_B); + const idB = idOf(SERVER_B); expect(await store.get(idB, LEGACY_TOKENS_FIELD)).toBeNull(); expect(await store.get(idB, LEGACY_CLIENT_SECRET_FIELD)).toBeNull(); expect( diff --git a/clients/web/src/utils/deepLink.test.ts b/clients/web/src/utils/deepLink.test.ts index 14d32dd151..e03cb00ea3 100644 --- a/clients/web/src/utils/deepLink.test.ts +++ b/clients/web/src/utils/deepLink.test.ts @@ -3,6 +3,7 @@ import { parseDeepLink, deepLinkConfigEquals, deepLinkParseStatus, + constantTimeEqual, DEEP_LINK_SERVER_ID, } from "./deepLink"; @@ -78,6 +79,22 @@ describe("parseDeepLink", () => { TOKEN, ); expect(wrong?.autoOpen).toBe(false); + const prefix = parseDeepLink( + `?serverUrl=https%3A%2F%2Fexample.com%2Fmcp&autoConnect=${TOKEN}&autoOpen=${TOKEN.slice(0, -1)}`, + TOKEN, + ); + expect(prefix?.autoOpen).toBe(false); + }); + + it("rejects an autoConnect that is a strict prefix or extension of the token", () => { + for (const guess of [TOKEN.slice(0, -1), TOKEN + "x", "Xok-abc"]) { + expect( + parseDeepLink( + `?serverUrl=https%3A%2F%2Fexample.com%2Fmcp&autoConnect=${guess}`, + TOKEN, + ), + ).toBeUndefined(); + } }); it("honors transport=sse and ignores unknown transport values", () => { @@ -197,3 +214,34 @@ describe("deepLinkConfigEquals", () => { ).toBe(false); }); }); + +describe("constantTimeEqual", () => { + it("is true only for identical strings", () => { + expect(constantTimeEqual("tok-abc", "tok-abc")).toBe(true); + expect(constantTimeEqual("", "")).toBe(true); + expect(constantTimeEqual("tok-abd", "tok-abc")).toBe(false); + expect(constantTimeEqual("Xok-abc", "tok-abc")).toBe(false); + }); + + it("rejects a length mismatch in either direction, including prefixes", () => { + expect(constantTimeEqual("tok-ab", "tok-abc")).toBe(false); + expect(constantTimeEqual("tok-abcd", "tok-abc")).toBe(false); + expect(constantTimeEqual("", "tok-abc")).toBe(false); + expect(constantTimeEqual("tok-abc", "")).toBe(false); + }); + + it("does not mistake a missing code unit for a NUL one", () => { + // charCodeAt past the candidate's end is NaN -> 0, the same value as + // "\0"; the length term is what must reject this. + expect(constantTimeEqual("ab", "ab\0")).toBe(false); + }); + + it("compares UTF-16 code units, so unnormalized forms differ", () => { + // Built from code points so the two spellings stay visibly distinct in + // source: U+00E9 (precomposed) vs "e" + U+0301 (combining acute). + const precomposed = "caf" + String.fromCharCode(0xe9); + const decomposed = "cafe" + String.fromCharCode(0x301); + expect(constantTimeEqual(precomposed, precomposed)).toBe(true); + expect(constantTimeEqual(decomposed, precomposed)).toBe(false); + }); +}); diff --git a/clients/web/src/utils/deepLink.ts b/clients/web/src/utils/deepLink.ts index c663fafccd..22e5117fc5 100644 --- a/clients/web/src/utils/deepLink.ts +++ b/clients/web/src/utils/deepLink.ts @@ -82,6 +82,37 @@ function validateServerUrl(raw: string): string | undefined { return url.href; } +/** + * Compare a caller-supplied `candidate` against the `secret` in time that does + * not depend on where the two first differ (#2429). A plain `===` on strings + * returns at the first mismatched character, so in principle the time it takes + * leaks how long a correct prefix a guessed token has. + * + * The browser has no `crypto.timingSafeEqual` (that is Node-only — the backend + * uses it for the `x-mcp-remote-auth` header), and `crypto.subtle` is async and + * absent on non-secure origins, so this is a synchronous XOR accumulator: every + * UTF-16 code unit of `secret` is visited whatever `candidate` holds, and a + * length mismatch is folded into the result rather than returned early. Only + * the secret's length drives the loop, so the time reveals nothing about the + * candidate beyond what the caller already chose. + * + * The secret's **length** is not hidden: the loop runs `secret.length` times. + * For the default per-launch token that length is fixed by its generator, but + * a user-supplied token (`MCP_INSPECTOR_API_TOKEN` / `--auth-token`) can have + * any length, which therefore stays observable in principle. That is the + * accepted limit of this helper. What it removes is the prefix leak, which is + * what would let a guess be refined one character at a time. + */ +export function constantTimeEqual(candidate: string, secret: string): boolean { + let diff = candidate.length ^ secret.length; + for (let i = 0; i < secret.length; i++) { + // Past the candidate's end charCodeAt is NaN, which `| 0` maps to 0; the + // length term above has already recorded that mismatch. + diff |= (candidate.charCodeAt(i) | 0) ^ secret.charCodeAt(i); + } + return diff === 0; +} + /** * Shallow equality for the subset of {@link MCPServerConfig} a deep link can * produce. Used by the auto-connect effect to decide whether the persisted @@ -149,7 +180,9 @@ export function parseDeepLink( const autoConnect = params.get("autoConnect"); if (!rawServerUrl || !autoConnect) return undefined; - if (!authToken || autoConnect !== authToken) return undefined; + if (!authToken || !constantTimeEqual(autoConnect, authToken)) { + return undefined; + } const serverUrl = validateServerUrl(rawServerUrl); if (!serverUrl) return undefined; @@ -165,10 +198,11 @@ export function parseDeepLink( const openApp = params.get("openApp") ?? undefined; const appArgs = decodeAppArgs(params.get("appArgs")); - // autoOpen is gated on the same per-launch token as autoConnect (already - // validated above), so the mere presence of the param is sufficient here — - // a link that reached this line has already proven knowledge of the token. - const autoOpen = params.get("autoOpen") === authToken; + // autoOpen is gated on the same per-launch token as autoConnect, and + // compared the same constant-time way (#2429). + const autoOpenParam = params.get("autoOpen"); + const autoOpen = + autoOpenParam !== null && constantTimeEqual(autoOpenParam, authToken); return { serverId: DEEP_LINK_SERVER_ID, diff --git a/clients/web/src/utils/errorFormat.test.ts b/clients/web/src/utils/errorFormat.test.ts index f56121a97b..980de963b5 100644 --- a/clients/web/src/utils/errorFormat.test.ts +++ b/clients/web/src/utils/errorFormat.test.ts @@ -11,6 +11,19 @@ describe("errorMessage", () => { expect(errorMessage(undefined)).toBe("undefined"); expect(errorMessage(42)).toBe("42"); }); + + // #2490: the result is shown on screen, so URL query secrets are redacted. + it("redacts query secrets in a URL quoted by the message", () => { + const out = errorMessage( + new Error("Callback https://srv.example/cb?code=abc123&state=ok failed."), + ); + expect(out).toBe( + "Callback https://srv.example/cb?code=%5BREDACTED%5D&state=ok failed.", + ); + expect(errorMessage("see https://s.example/?access_token=zzz")).toBe( + "see https://s.example/?access_token=%5BREDACTED%5D", + ); + }); }); describe("errorCodeOf", () => { @@ -35,6 +48,17 @@ describe("errorCodeOf", () => { }); describe("formatErrorDetails", () => { + it("redacts query secrets in the message and data it prints (#2490)", () => { + const err = Object.assign(new Error("at https://s.example/?code=abc123"), { + code: -32000, + data: { url: "https://s.example/?client_secret=shh" }, + }); + const out = formatErrorDetails(err); + expect(out).not.toContain("abc123"); + expect(out).not.toContain("shh"); + expect(out).toContain("%5BREDACTED%5D"); + }); + it("pretty-prints code/message/data when a code is present", () => { const err = Object.assign(new Error("bad params"), { code: -32602, diff --git a/clients/web/src/utils/errorFormat.ts b/clients/web/src/utils/errorFormat.ts index cd6f2fb6a3..e2639282e1 100644 --- a/clients/web/src/utils/errorFormat.ts +++ b/clients/web/src/utils/errorFormat.ts @@ -2,9 +2,17 @@ // Deliberately duck-typed rather than coupled to the SDK's error classes: the // values reaching here come off a `catch`, so their only guaranteed type is // `unknown`. +// +// Everything these return is for DISPLAY, so it is URL-query-redacted (#2490): +// a server or SDK error quoting `https://…?code=…` would otherwise land in a +// toast verbatim — a screenshot or screen-share away from leaking. Routing every +// on-screen error through `errorMessage` makes that one boundary rather than a +// rule each call site has to remember. + +import { redactUrlsInText } from "@inspector/core/mcp/fetchTracking.js"; export function errorMessage(err: unknown): string { - return err instanceof Error ? err.message : String(err); + return redactUrlsInText(err instanceof Error ? err.message : String(err)); } // The numeric JSON-RPC code of a thrown protocol error (e.g. a `ProtocolError` @@ -26,10 +34,12 @@ export function formatErrorDetails(err: unknown): string { if (err && typeof err === "object") { const e = err as { code?: unknown; message?: unknown; data?: unknown }; if (e.code !== undefined || e.data !== undefined) { - return JSON.stringify( - { code: e.code, message: e.message, data: e.data }, - null, - 2, + return redactUrlsInText( + JSON.stringify( + { code: e.code, message: e.message, data: e.data }, + null, + 2, + ), ); } } diff --git a/clients/web/tsup.runner.config.ts b/clients/web/tsup.runner.config.ts index 49a7a7c44c..a6e6afd773 100644 --- a/clients/web/tsup.runner.config.ts +++ b/clients/web/tsup.runner.config.ts @@ -50,6 +50,7 @@ export default defineConfig({ // client's own code, so all three lists carry it (AGENTS.md). The CLI was // inlining it; the #2067 guard surfaced that. "@modelcontextprotocol/ext-apps", + "@modelcontextprotocol/ext-tasks", // Consolidated to the ROOT manifest by #2195, along with every other // runtime dependency `core/` imports. tsup externalizes only what the // *nearest* package.json declares, so once a client stops declaring one it diff --git a/clients/web/vite.config.ts b/clients/web/vite.config.ts index 0799c9bd77..0678e242f3 100644 --- a/clients/web/vite.config.ts +++ b/clients/web/vite.config.ts @@ -222,6 +222,7 @@ export default defineConfig(({ command }) => { path.join(repoRoot, "core/storage/**/*.{ts,tsx}"), path.join(repoRoot, "core/logging/**/*.{ts,tsx}"), path.join(repoRoot, "core/node/**/*.{ts,tsx}"), + path.join(repoRoot, "core/extension/**/*.{ts,tsx}"), ], exclude: [ "**/*.stories.{ts,tsx}", diff --git a/core/auth/ema/idpOidc.ts b/core/auth/ema/idpOidc.ts index 60f3bdbb69..fb15cfbe2c 100644 --- a/core/auth/ema/idpOidc.ts +++ b/core/auth/ema/idpOidc.ts @@ -42,11 +42,70 @@ async function resolveIdpMetadata( return discoverIdpMetadata(issuer, fetchFn); } +/** The OpenID Connect Discovery 1.0 well-known path. */ +const OIDC_WELL_KNOWN = "/.well-known/openid-configuration"; + +/** + * Fetch the issuer's OpenID Connect Discovery document directly. + * + * The EMA IdP leg is an OIDC flow (it requests the `openid` scope and consumes + * an ID token), so OIDC Discovery 1.0 is the spec-correct metadata document for + * it — and the only one that can carry `end_session_endpoint`, which + * `clearEmaIdpSession` reads from the login-time cache to offer RP-initiated + * logout. The SDK's `discoverAuthorizationServerMetadata` tries the RFC 8414 + * path first, and an IdP that serves both documents answers there with plain + * OAuth metadata that legitimately omits the OIDC-only fields. + * + * Returns `undefined` on any failure — non-2xx, network error, or a body that + * is not valid authorization-server metadata — so the caller can fall back to + * the SDK's discovery, which raises its own errors for an issuer that is + * genuinely unreachable. + */ +async function fetchOpenIdConfiguration( + issuer: string, + issuerUrl: URL, + fetchFn?: typeof fetch, +): Promise<OAuthMetadata | undefined> { + // OIDC Discovery 1.0 §4.1: the well-known path is appended after the + // issuer's path component. + const url = new URL( + `${issuerUrl.pathname.replace(/\/$/, "")}${OIDC_WELL_KNOWN}`, + issuerUrl.origin, + ); + try { + const response = await (fetchFn ?? fetch)(url); + if (!response.ok) return undefined; + const parsed = OAuthMetadataSchema.safeParse(await response.json()); + if (!parsed.success) return undefined; + // OIDC Discovery 1.0 §4.3: the document's `issuer` MUST be identical to + // the issuer the document was fetched for — an exact string match against + // the configured issuer, not a normalized one (`issuerUrl.href` won't do: + // URL canonicalization appends a trailing slash to an origin-only URL). + // The SDK's RFC 8414 fallback enforces its own issuer-echo check, so this + // direct path must too — a mismatched document is treated as invalid, + // falling back to the SDK. + if (parsed.data.issuer !== issuer) { + return undefined; + } + return parsed.data; + } catch { + return undefined; + } +} + export async function discoverIdpMetadata( issuer: string, fetchFn?: typeof fetch, ): Promise<OAuthMetadata> { const issuerUrl = parseHttpUrl(issuer, "EMA IdP issuer (Client Settings)"); + const oidcMetadata = await fetchOpenIdConfiguration( + issuer.trim(), + issuerUrl, + fetchFn, + ); + if (oidcMetadata) { + return oidcMetadata; + } const metadata = await discoverAuthorizationServerMetadata(issuerUrl, { fetchFn, }); diff --git a/core/auth/ema/idpSession.ts b/core/auth/ema/idpSession.ts index d736549ae3..6b490b125a 100644 --- a/core/auth/ema/idpSession.ts +++ b/core/auth/ema/idpSession.ts @@ -21,13 +21,81 @@ export async function getEmaIdpLoginState( return "expired"; } +export interface ClearEmaIdpSessionResult { + /** + * OIDC RP-initiated logout URL (`end_session_endpoint` + + * `id_token_hint`), for the caller to relay to a human whose browser + * holds the IdP SSO cookie the local clear cannot touch. Present only + * when a session with an idToken existed and the IdP's login-time + * discovery metadata (cached in storage) advertises the endpoint. + */ + endSessionUrl?: string; +} + +/** + * Clear the inspector-local EMA IdP state: the cached IdP OIDC session, the + * leg-1 pending key, and every EMA-tagged resource-server entry. Local-only + * and network-free — the IdP's own browser SSO session is untouched; the + * returned `endSessionUrl` (when the IdP advertises one) lets the caller + * offer that step. + */ export async function clearEmaIdpSession( storage: OAuthStorage, issuer: string, -): Promise<void> { +): Promise<ClearEmaIdpSessionResult> { const normalized = normalizeIdpIssuer(issuer); - if (!normalized) return; + if (!normalized) return {}; + // Read before clearing: the end-session URL needs the session's idToken + // (as id_token_hint) and the login-time discovery metadata cached under + // the leg-1 key — the clears below destroy both. + const idToken = (await storage.getIdpSession(normalized))?.idToken; + const metadata = await storage.getServerMetadata( + idpOAuthStorageKey(normalized), + ); await storage.clearIdpSession(normalized); await storage.clear(idpOAuthStorageKey(normalized)); await storage.clearEnterpriseManagedResourceServers(); + if (!idToken || !metadata) return {}; + const endSessionUrl = buildEndSessionUrl(metadata, idToken); + return endSessionUrl === undefined ? {} : { endSessionUrl }; +} + +/** + * RP-initiated logout URL from the IdP's cached discovery metadata. + * `end_session_endpoint` is an OIDC RP-Initiated Logout field; the SDK's + * RFC 8414 schema does not declare it but parses with a loose object, so it + * survives into the cached metadata as an untyped extra key — narrow it + * ourselves. The URL carries the ID token (`id_token_hint`), so a plain-http + * endpoint is rejected unless its host is loopback (the same exemption the + * SDK applies to token endpoints) — never offer a logout URL that would send + * the token in cleartext. + */ +function buildEndSessionUrl( + metadata: object, + idToken: string, +): string | undefined { + const endpoint = (metadata as { end_session_endpoint?: unknown }) + .end_session_endpoint; + if (typeof endpoint !== "string") return undefined; + let url: URL; + try { + url = new URL(endpoint); + } catch { + return undefined; + } + if ( + url.protocol !== "https:" && + !(url.protocol === "http:" && isLoopbackHost(url.hostname)) + ) { + return undefined; + } + url.searchParams.set("id_token_hint", idToken); + return url.toString(); +} + +/** The SDK's loopback exemption list: localhost, 127.0.0.1, ::1. */ +function isLoopbackHost(hostname: string): boolean { + return ( + hostname === "localhost" || hostname === "127.0.0.1" || hostname === "[::1]" + ); } diff --git a/core/auth/ema/index.ts b/core/auth/ema/index.ts index e5e5a847e2..7940a5b3b1 100644 --- a/core/auth/ema/index.ts +++ b/core/auth/ema/index.ts @@ -20,6 +20,7 @@ export { clearEmaIdpSession, getEmaIdpLoginState, normalizeIdpIssuer, + type ClearEmaIdpSessionResult, type EmaIdpLoginState, } from "./idpSession.js"; export { diff --git a/core/auth/ema/storage.ts b/core/auth/ema/storage.ts index 0f4148ae8b..5cf7346df9 100644 --- a/core/auth/ema/storage.ts +++ b/core/auth/ema/storage.ts @@ -3,7 +3,17 @@ export function normalizeIdpIssuer(issuer: string): string { return issuer.replace(/\/$/, ""); } +/** Prefix that distinguishes an EMA IdP OIDC record from a server URL in the shared OAuth store. */ +export const IDP_OAUTH_KEY_PREFIX = "ema-idp:"; + /** OAuth storage key for in-flight IdP OIDC (PKCE, metadata). Not an OAuth `state` param prefix. */ export function idpOAuthStorageKey(issuer: string): string { - return `ema-idp:${normalizeIdpIssuer(issuer)}`; + return `${IDP_OAUTH_KEY_PREFIX}${normalizeIdpIssuer(issuer)}`; +} + +/** Inverse of {@link idpOAuthStorageKey}: the issuer for an IdP key, else `null`. */ +export function parseIdpOAuthStorageKey(key: string): string | null { + return key.startsWith(IDP_OAUTH_KEY_PREFIX) + ? key.slice(IDP_OAUTH_KEY_PREFIX.length) + : null; } diff --git a/core/auth/node/file-lock.ts b/core/auth/node/file-lock.ts index b1eede68b8..295bb401b9 100644 --- a/core/auth/node/file-lock.ts +++ b/core/auth/node/file-lock.ts @@ -391,17 +391,6 @@ export async function isFileLockHeld(filePath: string): Promise<boolean> { } } -/** - * Run `fn` holding an exclusive cross-process lock on `filePath`. - * - * The lock is `<filePath>.lock`, a directory beside the secrets file rather - * than inside it — `proper-lockfile` never opens or truncates the file it - * guards, so a lock that outlives its holder can only ever block a write, - * never damage one. - * - * Returns whatever `fn` returns. `fn` runs exactly once either way — the - * lock's absence changes the guarantee, never whether the work happens. - */ /** * Take the lock and hand back its release, or `null` when locking is * unavailable here and the caller should proceed unprotected. @@ -513,14 +502,29 @@ export async function openSecretFileLock( }; } +/** + * Run `fn` holding an exclusive cross-process lock on `filePath`. + * + * The lock is `<filePath>.lock`, a directory beside the secrets file rather + * than inside it — `proper-lockfile` never opens or truncates the file it + * guards, so a lock that outlives its holder can only ever block a write, + * never damage one. + * + * Returns whatever `fn` returns. `fn` runs exactly once either way — the + * lock's absence changes the guarantee, never whether the work happens. + * `fn` receives whether the lock is actually held (`false` = degraded, + * unlocked run), so a caller whose work is only safe under real exclusion + * — legacy secret-entry migration, which deletes its sources — can refuse + * instead of racing (#2556 review). + */ export async function withSecretFileLock<T>( filePath: string, - fn: () => Promise<T>, + fn: (locked: boolean) => Promise<T>, ): Promise<T> { const release = await openSecretFileLock(filePath); - if (release === null) return fn(); + if (release === null) return fn(false); try { - return await fn(); + return await fn(true); } finally { await release(); } diff --git a/core/auth/node/oauth-namespace-ledger.ts b/core/auth/node/oauth-namespace-ledger.ts new file mode 100644 index 0000000000..3f964db13e --- /dev/null +++ b/core/auth/node/oauth-namespace-ledger.ts @@ -0,0 +1,339 @@ +/** + * The secrets-namespace ledger (#2560): a sidecar to the OAuth state file + * recording every secrets namespace that file has used, and which server + * URLs and IdP issuers have had store entries written under each one. + * + * Why a sidecar and not a key in the state file: the `secretsNamespace` + * stamp (#2549) does not survive a save by an Inspector older than it + * (≤ 2.9.x), whose parse → serialize round trip drops unknown keys. The + * next save by a current Inspector then sees an unstamped file and + * re-adopts under a fresh UUID, and the previous namespace's entries — + * possibly still-valid refresh tokens in the OS keychain — are left with + * no state file referencing them. Anything recorded *inside* `oauth.json` + * is lost the same way, so the record lives beside it, where an old + * version never looks. + * + * Why it records keys and not just the namespace: the store cannot be + * enumerated on every backend (the keyring cannot list), so a purge has to + * name each id it deletes. And the stripped file is not a sufficient index + * of what lived under the old namespace — the old version may have removed + * servers in the meantime. So keys are recorded *before* their entries are + * written; recording extra is harmless (purging an id that holds nothing + * is a no-op), recording too little is the leak this exists to close. + * + * Why keys are recorded per store *location*: the backend is selected per + * run, and a run can land on a different one than the run that wrote the + * entries (a locked keychain falling back to `secrets.json`, an explicit + * `MCP_INSPECTOR_SECRET_STORE` switch). Deleting from the wrong backend + * succeeds against nothing, and dropping the record on that "success" + * would strand the real entries for good. So a record is dropped only by a + * purge in the location it was written to, and survives every other run + * until one of those comes along — except that a `secrets.json` record is + * handed to the keychain record rather than dropped, because a + * file-to-keychain hand-off may already have copied its entries there. + * + * Everything here is best-effort and never throws: the ledger is a cleanup + * aid, not part of the credential path, so a failure warns once and the + * save or removal it rides on carries on. Purging deletes credentials, so + * the callers in `oauth-persist-file.ts` run it only under the real file + * lock — an unlocked purge could delete a concurrent adopter's live + * entries. + */ + +import { resolve } from "node:path"; +import { + deleteStoreFile, + readStoreFile, + writeStoreFile, +} from "../../storage/store-io.js"; +import { serializeStore } from "../../storage/store-serialize.js"; +import { KeyringSecretStore, type SecretStore } from "./secret-store.js"; +import { FileSecretStore } from "./file-secret-store.js"; +import { resolveConcreteSecretStore } from "./secret-store-selection.js"; +import { + isValidSecretsNamespace, + oauthIdpSecretServerId, + oauthSecretServerId, +} from "./oauth-secrets.js"; + +/** The ledger's path, derived from the state file it belongs to. */ +export const namespaceLedgerPath = (stateFilePath: string): string => + `${stateFilePath}.namespaces.json`; + +const KEYRING_LOCATION = "keyring"; +const FILE_LOCATION_PREFIX = "file:"; + +/** + * Where a store's entries live, as a stable string another process can + * compare: the OS keychain, one particular `secrets.json`, or RAM (the + * test doubles and the session-scoped container fallback — entries there + * die with the process, so purging them from any memory store is moot). + * Production stores arrive wrapped (`defaultSecretStore()`), so the + * concrete selection is resolved first; the wrapper itself names nothing. + */ +export async function secretStoreLocation(store: SecretStore): Promise<string> { + const concrete = await resolveConcreteSecretStore(store); + if (concrete instanceof KeyringSecretStore) return KEYRING_LOCATION; + if (concrete instanceof FileSecretStore) { + return `${FILE_LOCATION_PREFIX}${resolve(concrete.filePath)}`; + } + return "memory"; +} + +/** + * The recorded locations a purge from `location` covers: always its own, + * and — for a keychain run — every `secrets.json` record too, because the + * file-to-keychain hand-off (`absorbFileSecretsIntoKeyring`) copies a + * file's entries, superseded namespaces included, into the keychain + * without telling any ledger. Deleting ids that were never copied is a + * no-op. + * + * Only the run's own record is ever dropped. A keychain run cannot know + * whether a hand-off is still copying, was interrupted, or left the file + * in place, so a `file:` record it purges stays recorded and is purged + * again on the next locked save; it is consumed only by a run against that + * file — and then folded into the keychain record rather than forgotten + * (see {@link purgeSupersededNamespaces}), since copies may already sit in + * the keychain. + */ +function coveredLocations( + location: string, + byLocation: Map<string, LedgerKeys>, +): Array<{ location: string; keys: LedgerKeys }> { + return [...byLocation] + .filter( + ([recorded]) => + recorded === location || + (location === KEYRING_LOCATION && + recorded.startsWith(FILE_LOCATION_PREFIX)), + ) + .map(([recorded, keys]) => ({ location: recorded, keys })); +} + +/** The keys one namespace has had store entries written under. */ +interface LedgerKeys { + servers: Set<string>; + idpSessions: Set<string>; +} + +/** + * Namespace → store location → its keys. Namespaces are validated, and + * both maps serialize through `Object.fromEntries`, so a hostile key such + * as `__proto__` stays an own property rather than touching a prototype. + */ +type Ledger = Map<string, Map<string, LedgerKeys>>; + +const warnedLedgerFailures = new Set<string>(); + +function warnLedgerFailure(what: string, error: unknown): void { + const reason = error instanceof Error ? error.message : String(error); + const key = `${what}:${reason}`; + if (warnedLedgerFailures.has(key)) return; + warnedLedgerFailures.add(key); + console.warn( + `[mcp-inspector] Could not ${what} the OAuth secrets-namespace ledger (${reason}). OAuth state is unaffected, but secret-store entries left behind if this state file's namespace is ever lost (for example, after a save by an older Inspector) may not be cleaned up automatically.`, + ); +} + +/** Test seam: forget which ledger-failure warnings have been emitted. */ +export function resetNamespaceLedgerWarnings(): void { + warnedLedgerFailures.clear(); +} + +function isRecord(value: unknown): value is Record<string, unknown> { + return typeof value === "object" && value !== null && !Array.isArray(value); +} + +function stringsOf(value: unknown): string[] { + return Array.isArray(value) + ? value.filter((item): item is string => typeof item === "string") + : []; +} + +/** + * Parse a raw ledger, tolerantly: anything unrecognized reads as empty, and + * an invalid namespace is skipped — it would otherwise reach store ids, + * where it could forge the id delimiter (see `isValidSecretsNamespace`). + * Keys are URLs/issuers rather than raw store ids for the same reason: ids + * are always rebuilt here, so a tampered ledger can only ever name OAuth + * ids, never a catalog server's or `client.json`'s. + */ +function parseLedger(raw: string | null): Ledger { + const ledger: Ledger = new Map(); + if (raw === null) return ledger; + let parsed: unknown; + try { + parsed = JSON.parse(raw); + } catch { + return ledger; + } + const namespaces = isRecord(parsed) ? parsed.namespaces : undefined; + if (!isRecord(namespaces)) return ledger; + for (const [namespace, locations] of Object.entries(namespaces)) { + if (!isValidSecretsNamespace(namespace) || !isRecord(locations)) continue; + const byLocation = new Map<string, LedgerKeys>(); + for (const [location, keys] of Object.entries(locations)) { + if (!isRecord(keys)) continue; + byLocation.set(location, { + servers: new Set(stringsOf(keys.servers)), + idpSessions: new Set(stringsOf(keys.idpSessions)), + }); + } + if (byLocation.size > 0) ledger.set(namespace, byLocation); + } + return ledger; +} + +function serializeLedger(ledger: Ledger): string { + const namespaces = Object.fromEntries( + [...ledger].map(([namespace, byLocation]) => [ + namespace, + Object.fromEntries( + [...byLocation].map(([location, keys]) => [ + location, + { servers: [...keys.servers], idpSessions: [...keys.idpSessions] }, + ]), + ), + ]), + ); + return serializeStore({ namespaces }); +} + +async function writeLedger( + stateFilePath: string, + ledger: Ledger, +): Promise<void> { + const path = namespaceLedgerPath(stateFilePath); + if (ledger.size === 0) await deleteStoreFile(path); + else await writeStoreFile(path, serializeLedger(ledger)); +} + +/** + * Record that store entries for these keys are about to be written under + * `namespace` in `secretStore`. Call it *before* the store writes, so a + * crash between the two leaves an over-recorded ledger rather than an + * unrecorded entry. Rewrites the ledger only when a key is new, so a + * steady-state save costs one small read. + */ +export async function recordNamespaceKeys( + stateFilePath: string, + secretStore: SecretStore, + namespace: string, + servers: Iterable<string>, + idpSessions: Iterable<string>, +): Promise<void> { + try { + const ledger = parseLedger( + await readStoreFile(namespaceLedgerPath(stateFilePath)), + ); + const location = await secretStoreLocation(secretStore); + let byLocation = ledger.get(namespace); + if (!byLocation) { + byLocation = new Map(); + ledger.set(namespace, byLocation); + } + let keys = byLocation.get(location); + let changed = false; + if (!keys) { + keys = { servers: new Set(), idpSessions: new Set() }; + byLocation.set(location, keys); + changed = true; + } + for (const url of servers) { + if (keys.servers.has(url)) continue; + keys.servers.add(url); + changed = true; + } + for (const issuer of idpSessions) { + if (keys.idpSessions.has(issuer)) continue; + keys.idpSessions.add(issuer); + changed = true; + } + if (changed) await writeLedger(stateFilePath, ledger); + } catch (error) { + warnLedgerFailure("update", error); + } +} + +/** + * Purge, from `secretStore`, every recorded namespace other than `current` + * — the namespace the state file actually carries, or `undefined` when it + * carries none (a stripped stamp, a file deleted by hand, a whole-file + * removal), in which case every recorded namespace is superseded. Only the + * keys recorded for the locations this store covers are purged (see + * {@link coveredLocations}), and a location's record is dropped only once + * all its ids are gone; one whose purge failed stays recorded, so the next + * locked save or removal retries it. Records for other locations are left + * for a run that uses them. + * + * Only call this under the real file lock: it deletes credentials, and an + * unlocked caller could be racing an adopter whose namespace is recorded + * but not yet stamped. + */ +export async function purgeSupersededNamespaces( + stateFilePath: string, + secretStore: SecretStore, + current: string | undefined, +): Promise<void> { + let ledger: Ledger; + let location: string; + try { + ledger = parseLedger( + await readStoreFile(namespaceLedgerPath(stateFilePath)), + ); + location = await secretStoreLocation(secretStore); + } catch (error) { + warnLedgerFailure("read", error); + return; + } + let changed = false; + for (const [namespace, byLocation] of ledger) { + if (namespace === current) continue; + for (const { location: recorded, keys } of coveredLocations( + location, + byLocation, + )) { + const purges = [ + ...[...keys.servers].map( + (url) => () => oauthSecretServerId(url, namespace), + ), + ...[...keys.idpSessions].map( + (issuer) => () => oauthIdpSecretServerId(issuer, namespace), + ), + ]; + let purged = true; + // Per-id catch, like adoption's legacy purge: one failure must not + // abandon the remaining ids. The id is built inside it too — a key + // with an unpaired surrogate makes `encodeURIComponent` throw. + for (const idOf of purges) { + try { + await secretStore.deleteAllForServer(idOf()); + } catch (error) { + purged = false; + warnLedgerFailure("purge an orphaned namespace recorded in", error); + } + } + if (!purged || recorded !== location) continue; + byLocation.delete(recorded); + // A file's entries may also have been copied into the keychain by a + // hand-off this ledger never saw; keep tracking them there. + if (recorded.startsWith(FILE_LOCATION_PREFIX)) { + const keyring = byLocation.get(KEYRING_LOCATION) ?? { + servers: new Set<string>(), + idpSessions: new Set<string>(), + }; + for (const url of keys.servers) keyring.servers.add(url); + for (const issuer of keys.idpSessions) keyring.idpSessions.add(issuer); + byLocation.set(KEYRING_LOCATION, keyring); + } + changed = true; + } + if (byLocation.size === 0) ledger.delete(namespace); + } + if (!changed) return; + try { + await writeLedger(stateFilePath, ledger); + } catch (error) { + warnLedgerFailure("update", error); + } +} diff --git a/core/auth/node/oauth-persist-file.ts b/core/auth/node/oauth-persist-file.ts index 8c707bf257..35bf7da60b 100644 --- a/core/auth/node/oauth-persist-file.ts +++ b/core/auth/node/oauth-persist-file.ts @@ -21,6 +21,17 @@ * store is durable, the same guard the mcp.json/client.json migrations * use. A store write failure degrades those tokens to memory-only with a * loud warning; it never falls back to writing them into the file. + * 3. **Secret-store namespace** (#2549): each state file carries a + * `secretsNamespace` UUID, baked into every secret-store id its entries + * use, so two state files (profiles) naming the same server hold + * separate store entries instead of overwriting one shared slot. The + * namespace is minted — and a legacy file's unscoped entries moved under + * it — on the file's first write (`adoptSecretsNamespace`); reads and + * removes honor whatever the file says and never adopt. A sidecar + * ledger (`oauth-namespace-ledger.ts`, #2560) records every namespace + * the file has used, so entries under one the file no longer carries — + * its stamp stripped by an older Inspector's save — are purged on the + * next locked adoption or removal instead of orphaned in the store. */ import { @@ -28,6 +39,7 @@ import { writeStoreFile, deleteStoreFile, } from "../../storage/store-io.js"; +import { serializeStore } from "../../storage/store-serialize.js"; import { setOwnEntry, getOwnEntry } from "../../storage/own-entry.js"; import { mergeOAuthSections, @@ -53,13 +65,19 @@ import { type SecretStore, } from "./secret-store.js"; import { defaultSecretStore } from "./secret-store-selection.js"; +import { + purgeSupersededNamespaces, + recordNamespaceKeys, +} from "./oauth-namespace-ledger.js"; import { MAX_WRITE_ATTEMPTS } from "./file-secret-store.js"; import { IDP_SESSION_FIELD, getPersistTokensPolicy, isUsableStoredSecret, + isValidSecretsNamespace, joinIdpSession, joinServerOAuthState, + newSecretsNamespace, oauthIdpSecretServerId, oauthSecretServerId, serverSecretFields, @@ -187,13 +205,13 @@ function rethrowLockError( async function withOAuthStateLock<T>( filePath: string, action: "save" | "read" | "remove", - body: () => Promise<T>, + body: (locked: boolean) => Promise<T>, ): Promise<T> { let entered = false; try { - return await withSecretFileLock(filePath, async () => { + return await withSecretFileLock(filePath, async (locked) => { entered = true; - return body(); + return body(locked); }); } catch (error) { if (!entered) rethrowLockError(filePath, error, action); @@ -225,20 +243,210 @@ export class OAuthStateFileUnrecognizedError extends Error { } /** - * Locked-read helper for the mutation paths: parse the OAuth state file, - * distinguishing "absent" (null) from "present but unrecognized" (refuse — - * see {@link OAuthStateFileUnrecognizedError}). + * Top-level state-file key holding the file's secrets namespace (#2549): a + * UUID minted per state file and baked into every secret-store id the file's + * entries use (see `oauthSecretServerId`). It is what keeps two state files + * (profiles) that connect to the same server from sharing — and silently + * overwriting — one store entry. Node-file-backend-only: the browser and + * sessionStorage backends never see it ({@link parseOAuthPersistBlob} + * ignores unknown keys), and every writer re-reads it from disk under the + * file lock, so a snapshot round-tripped through the API cannot strip it. + */ +export const SECRETS_NAMESPACE_KEY = "secretsNamespace"; + +/** + * Extract the secrets namespace from a raw state-file blob. Tolerant like + * the read path: an unparseable file or an invalid value reads as "no + * namespace" (legacy unscoped ids) rather than failing — an invalid value + * must not reach store ids, where it could forge the `+` delimiter or break + * the keyring's colon parse (see `isValidSecretsNamespace`). + */ +function parseSecretsNamespace(raw: string | null): string | undefined { + if (raw === null) return undefined; + try { + const parsed: unknown = JSON.parse(raw); + if (typeof parsed !== "object" || parsed === null) return undefined; + const value = (parsed as Record<string, unknown>)[SECRETS_NAMESPACE_KEY]; + return isValidSecretsNamespace(value) ? value : undefined; + } catch { + return undefined; + } +} + +/** + * Serialize a snapshot for the state file, stamping the secrets namespace + * first so it survives every rewrite (sectioned saves, the plaintext-secret + * migration, adoption itself). Key order is fixed — namespace, then the + * snapshot — because the write path compares raw file strings to detect + * concurrent writers. + */ +function serializeOAuthFileBlob( + snapshot: OAuthPersistSnapshot, + namespace: string | undefined, +): string { + if (namespace === undefined) return serializeOAuthPersistBlob(snapshot); + return serializeStore({ [SECRETS_NAMESPACE_KEY]: namespace, ...snapshot }); +} + +/** + * Cleanup-failure variant of {@link warnStoreWriteFailure}: namespace + * adoption copied the legacy unscoped entries to their namespaced ids and + * stamped the file, but deleting the legacy originals failed. Nothing is + * lost — the namespaced ids are authoritative from here on — but the + * leftovers sit in the store un-indexed (no state file will purge them) and + * could serve stale credentials to a profile that has not adopted yet. */ -async function readDiskForMutation( +function warnNamespaceCleanupFailure(error: unknown): void { + const reason = error instanceof Error ? error.message : String(error); + const key = `namespace-cleanup:${reason}`; + if (warnedStoreFailures.has(key)) return; + warnedStoreFailures.add(key); + console.warn( + `[mcp-inspector] Could not remove legacy un-namespaced secret-store entries after scoping them to this state file (${reason}). This file uses the scoped copies, but another pre-namespace state file could still read the stale leftovers until it adopts or re-authorizes — and they will not be cleaned up automatically.`, + ); +} + +/** + * Ensure the state file has a secrets namespace, minting and stamping one — + * and moving its legacy unscoped store entries under it — when it does not + * (#2549). Runs under the caller's file lock, on the mutation path only: + * reads keep working against whatever the file currently says, so a + * pre-adoption profile loses nothing until it first writes. + * + * For a legacy file (recognized, entries, no namespace) the move is + * copy → stamp → delete, in that order, because each step's failure mode + * differs: + * + * - **Copy** (legacy id → namespaced id) lands on vacant scoped ids — the + * namespace is freshly minted, so nothing can already live under it. A + * failure deletes the copies it made and rethrows — the file is + * unstamped, so everything still resolves through the legacy ids and + * the next write retries under a new UUID. + * - **Stamp** (rewrite the file with the namespace) is the commit point: a + * failure triggers the same restore, because a successful copy with no + * stamp would be re-run under a *different* UUID next time, stranding + * this one's copies forever. + * - **Delete** (the legacy originals) is best-effort after the commit: the + * namespaced ids are already authoritative for this file, so a failure + * here never loses a token — but the leftovers are not harmless to + * everyone: they are unindexed here, and a pre-namespace profile could + * still read them as stale credentials until it adopts or re-authorizes + * ({@link warnNamespaceCleanupFailure}). Deleting is deliberate, not + * cautious copying: the legacy entry is exactly the shared slot this + * change exists to retire, and `removeOAuthStore` purges by the file's + * own ids, so a leftover would otherwise be orphaned forever. Another + * profile still reading the legacy ids re-authorizes once — its copy of + * those tokens was already being overwritten by every other profile, + * which is the bug. + * + * A fresh file (absent, or no entries on disk) just mints: the namespace + * reaches disk with the write that follows, and there is nothing to move. + * + * Migration requires the real file lock (`locked`). The move deletes its + * sources, so two unlocked adopters racing on one legacy file can each + * observe the other's half-finished move — one copies and deletes, the + * other strict-reads nothing, stamps an empty namespace of its own, and + * can win the file, stranding the first's copies under an abandoned + * namespace. Under the lock adopters serialize (the second sees the + * first's stamp and returns it); degraded, this refuses the one-time + * migration loudly rather than risking that loss. Mint-only adoption — + * a file that is absent, or recognized but indexing no entries — stays + * allowed unlocked: there is nothing to move, and a concurrent + * mint converges via the namespace re-key in `writeOAuthSections`. + * + * Under the lock, adoption first purges every namespace the ledger records + * (#2560): reaching this point means the file carries none, so each + * recorded one is superseded — most often a stamp an older Inspector's save + * stripped, whose entries nothing else would ever find. A stamped file's + * locked saves purge every recorded namespace but its own, which is how a + * purge that failed here is retried once the store recovers. Unlocked, the purge + * is skipped: a concurrent adopter's namespace can be recorded before it is + * stamped, and purging it would delete live credentials. + */ +async function adoptSecretsNamespace( filePath: string, - action: "save" | "remove", -): Promise<OAuthPersistSnapshot | null> { + secretStore: SecretStore, + locked: boolean, +): Promise<string> { const raw = await readStoreFile(filePath); - const parsed = parseOAuthPersistBlob(raw); - if (raw !== null && parsed === null) { - throw new OAuthStateFileUnrecognizedError(filePath, action); + const existing = parseSecretsNamespace(raw); + if (existing !== undefined) { + // An ordinary save also retries any superseded namespace whose purge + // failed at adoption — the new stamp means adoption will not run again. + if (locked) { + await purgeSupersededNamespaces(filePath, secretStore, existing); + } + return existing; + } + const snapshot = parseOAuthPersistBlob(raw); + if (raw !== null && snapshot === null) { + throw new OAuthStateFileUnrecognizedError(filePath, "save"); + } + if (locked) { + await purgeSupersededNamespaces(filePath, secretStore, undefined); + } + const namespace = newSecretsNamespace(); + if (snapshot === null) return namespace; + + const moves = [ + ...Object.entries(snapshot.servers).map(([url, state]) => ({ + legacyId: oauthSecretServerId(url), + scopedId: oauthSecretServerId(url, namespace), + fields: serverSecretFields(state), + })), + ...Object.keys(snapshot.idpSessions).map((issuer) => ({ + legacyId: oauthIdpSecretServerId(issuer), + scopedId: oauthIdpSecretServerId(issuer, namespace), + fields: [IDP_SESSION_FIELD], + })), + ]; + // A recognized but entry-less legacy file indexes no store ids, so no + // destructive race exists — it mints like a fresh file, locked or not + // (the save that follows writes the namespace to disk). + if (moves.length === 0) return namespace; + if (!locked) { + throw new SecretStoreUnavailableError( + `Could not save OAuth state: ${filePath} predates per-state-file secret namespaces, and migrating its secret-store entries needs the file lock, which is unavailable here (see the lock warning above). Migrating without it could lose credentials if two processes migrate at once. Nothing was changed; make the lock directory writable and retry — or, if you cannot, clear this file's stored OAuth state and re-authorize (clearing does not migrate, and the fresh state file mints its namespace without the lock).`, + ); } - return parsed; + + // Recorded before the copies, so an interrupted move's copies are still + // findable if this namespace is later superseded. + await recordNamespaceKeys( + filePath, + secretStore, + namespace, + Object.keys(snapshot.servers), + Object.keys(snapshot.idpSessions), + ); + // Rollback baseline: the scoped ids are vacant before this call — the + // namespace is a UUID minted moments ago, so nothing can already live + // under it — which makes "restore" simply "delete what we copied". + const touched: SecretFieldSnapshot[] = []; + try { + for (const { legacyId, scopedId, fields } of moves) { + for (const field of fields) { + const value = await secretStoreGetStrict(secretStore, legacyId, field); + if (value === null) continue; + touched.push({ serverId: scopedId, field, value: null }); + await secretStore.set(scopedId, field, value); + } + } + await writeStoreFile(filePath, serializeOAuthFileBlob(snapshot, namespace)); + } catch (error) { + await restoreSecretFields(secretStore, touched, warnRestoreFailure); + throw error; + } + // Per-id catch: one failed purge must not abandon the remaining legacy + // ids — each gets its own best-effort attempt. + for (const { legacyId } of moves) { + try { + await secretStore.deleteAllForServer(legacyId); + } catch (error) { + warnNamespaceCleanupFailure(error); + } + } + return namespace; } /** @@ -389,7 +597,16 @@ export async function writeOAuthSections( ): Promise<void> { const policy = getPersistTokensPolicy(); const durable = await secretStoreIsDurable(secretStore); - await withOAuthStateLock(filePath, "save", async () => { + await withOAuthStateLock(filePath, "save", async (locked) => { + // The namespace scoping every store id below; minted (and legacy + // entries moved — under the real lock only, see adoptSecretsNamespace) + // on this file's first namespaced write. Resolved once per save: an + // attempt retry never re-ADOPTS — but it may re-KEY. Under degraded + // (unlocked) locking a concurrent first writer can mint a different + // namespace and win the file between this read and an attempt's write, + // so each attempt below re-checks the namespace observed on disk and + // converges onto it (`let`, not `const`). + let namespace = await adoptSecretsNamespace(filePath, secretStore, locked); // Restore baseline for every failure exit below. For each touched // (server, field) it holds the latest store value NOT written by this // call: the pre-operation value, superseded by a concurrent writer's @@ -536,10 +753,49 @@ export async function writeOAuthSections( // The disk read sits inside the try too: on a retry the store already // holds an earlier attempt's writes, and a concurrent writer replacing - // the file with something unrecognized would otherwise make - // `readDiskForMutation` throw past the loop without any rollback. + // the file with something unrecognized would otherwise throw past the + // loop without any rollback. Read raw, not just parsed: the namespace + // check below needs the unparsed blob. try { - const disk = await readDiskForMutation(filePath, "save"); + const rawDisk = await readStoreFile(filePath); + const disk = parseOAuthPersistBlob(rawDisk); + if (rawDisk !== null && disk === null) { + throw new OAuthStateFileUnrecognizedError(filePath, "save"); + } + // Namespace convergence: under degraded (unlocked) locking a + // concurrent first writer can adopt a different namespace and win + // the file after our adoption read. The merge below already + // converges the *data* onto what they left; the namespace must + // converge the same way, or every retry re-stamps our own mint and + // the two writers ping-pong, stranding the loser's secrets under a + // namespace the final file no longer references. Re-key to the + // namespace observed on disk, first rolling earlier attempts' store + // writes (all keyed under the abandoned namespace) back to baseline. + const diskNamespace = parseSecretsNamespace(rawDisk); + if (diskNamespace !== undefined && diskNamespace !== namespace) { + // Strict, unlike the failure exits' best-effort restores: this is + // normal control flow with no original error to preserve, and + // carrying on past a failed restore would clear the baseline and + // let the save report success with earlier attempts' writes + // stranded under the abandoned namespace, unindexed by any file. + // Aborting keeps the baseline for the rethrow's reconciliation, + // and a retried save converges cleanly. + let restoreFailure: unknown; + await restoreSecretFields( + secretStore, + [...restoreBaseline.values()], + (error) => { + restoreFailure ??= error; + }, + ); + if (restoreFailure !== undefined) throw restoreFailure; + restoreBaseline.clear(); + ourWrites.clear(); + // Our unconfirmed write carried the abandoned namespace, and the + // read above proves the file no longer holds it. + unconfirmed = null; + namespace = diskNamespace; + } // Deduplicated: caller-passed sections may repeat a URL/issuer, and a // second pass over the same entry would snapshot the value the first // pass just wrote — a rollback would then "restore" that intermediate @@ -567,9 +823,18 @@ export async function writeOAuthSections( ], }; const merged = mergeOAuthSections(disk, snapshot, effective); + // Ledger before store writes (#2560): an entry written under this + // namespace must be findable once the namespace is superseded. + await recordNamespaceKeys( + filePath, + secretStore, + namespace, + Object.keys(merged.servers), + Object.keys(merged.idpSessions), + ); for (const url of effective.servers ?? []) { - const serverId = oauthSecretServerId(url); + const serverId = oauthSecretServerId(url, namespace); // Own-property reads: with a `__proto__` key a plain lookup on a // map that lacks it returns the inherited prototype, so a clear // would read as an update and skip the purge below. @@ -661,7 +926,7 @@ export async function writeOAuthSections( } for (const issuer of effective.idpSessions ?? []) { - const serverId = oauthIdpSecretServerId(issuer); + const serverId = oauthIdpSecretServerId(issuer, namespace); const next = getOwnEntry(snapshot.idpSessions, issuer); const entryPrior = await snapshotSecretFields(secretStore, serverId, [ IDP_SESSION_FIELD, @@ -724,7 +989,7 @@ export async function writeOAuthSections( }); } - written = serializeOAuthPersistBlob(merged); + written = serializeOAuthFileBlob(merged, namespace); await writeStoreFile(filePath, written); unconfirmed = { written, writes: attemptWrites }; } catch (error) { @@ -765,17 +1030,18 @@ export async function writeOAuthSections( /** Build the bulk-read request list for everything a snapshot could hold. */ function secretRequestsFor( snapshot: OAuthPersistSnapshot, + namespace: string | undefined, ): SecretBulkRequest[] { const requests: SecretBulkRequest[] = []; for (const [url, state] of Object.entries(snapshot.servers)) { requests.push({ - serverId: oauthSecretServerId(url), + serverId: oauthSecretServerId(url, namespace), fields: serverSecretFields(state), }); } for (const issuer of Object.keys(snapshot.idpSessions)) { requests.push({ - serverId: oauthIdpSecretServerId(issuer), + serverId: oauthIdpSecretServerId(issuer, namespace), fields: [IDP_SESSION_FIELD], }); } @@ -786,8 +1052,9 @@ function secretRequestsFor( async function joinSnapshot( snapshot: OAuthPersistSnapshot, secretStore: SecretStore, + namespace: string | undefined, ): Promise<OAuthPersistSnapshot> { - const requests = secretRequestsFor(snapshot); + const requests = secretRequestsFor(snapshot, namespace); // Strict: this read hydrates the memory state that later sectioned writes // diff against, so a tolerant read during a store outage would present // every credential as absent — and the next save would *delete* them from @@ -801,13 +1068,19 @@ async function joinSnapshot( servers: Object.fromEntries( Object.entries(snapshot.servers).map(([url, state]) => [ url, - joinServerOAuthState(state, values[oauthSecretServerId(url)] ?? {}), + joinServerOAuthState( + state, + values[oauthSecretServerId(url, namespace)] ?? {}, + ), ]), ), idpSessions: Object.fromEntries( Object.entries(snapshot.idpSessions).map(([issuer, session]) => [ issuer, - joinIdpSession(session, values[oauthIdpSecretServerId(issuer)] ?? {}), + joinIdpSession( + session, + values[oauthIdpSecretServerId(issuer, namespace)] ?? {}, + ), ]), ), }; @@ -846,6 +1119,20 @@ async function migratePlaintextSecrets( const raw = await readStoreFile(filePath); const fresh = parseOAuthPersistBlob(raw); if (!fresh || !snapshotHasPlaintextSecrets(fresh)) return; + // Migration honors — and preserves — the file's own namespace; it never + // mints one. Scoping is a write-path decision (`adoptSecretsNamespace`), + // and a read-triggered strip that also re-keyed the entries would be the + // adoption without its legacy-entry move. + const namespace = parseSecretsNamespace(raw); + if (namespace !== undefined) { + await recordNamespaceKeys( + filePath, + secretStore, + namespace, + Object.keys(fresh.servers), + Object.keys(fresh.idpSessions), + ); + } const migrateEntrySecrets = async ( serverId: string, secrets: OAuthSecretValues, @@ -873,12 +1160,18 @@ async function migratePlaintextSecrets( for (const [url, state] of Object.entries(fresh.servers)) { const split = splitServerOAuthState(state, "all"); setOwnEntry(residue.servers, url, split.residue); - await migrateEntrySecrets(oauthSecretServerId(url), split.secrets); + await migrateEntrySecrets( + oauthSecretServerId(url, namespace), + split.secrets, + ); } for (const [issuer, session] of Object.entries(fresh.idpSessions)) { const split = splitIdpSession(session, "all"); setOwnEntry(residue.idpSessions, issuer, split.residue); - await migrateEntrySecrets(oauthIdpSecretServerId(issuer), split.secrets); + await migrateEntrySecrets( + oauthIdpSecretServerId(issuer, namespace), + split.secrets, + ); } // The split leaves only type-corrupt token payloads in the residue (see // `splitTokens`), so a hand-edited junk entry stays plaintext and @@ -886,7 +1179,7 @@ async function migratePlaintextSecrets( // reject. Such a file re-enters migration on every read; skip the rewrite // when nothing would change so a steady-state file is not re-written (and // a concurrent writer not clobbered) per read. - const stripped = serializeOAuthPersistBlob(residue); + const stripped = serializeOAuthFileBlob(residue, namespace); if (stripped !== raw) { await writeStoreFile(filePath, stripped); } @@ -922,7 +1215,8 @@ export async function readOAuthStore( secretStore: SecretStore = defaultSecretStore(), ): Promise<OAuthPersistSnapshot | null> { return withOAuthStateLock(filePath, "read", async () => { - let snapshot = parseOAuthPersistBlob(await readStoreFile(filePath)); + let raw = await readStoreFile(filePath); + let snapshot = parseOAuthPersistBlob(raw); if (snapshot === null) return null; if ( @@ -931,14 +1225,16 @@ export async function readOAuthStore( ) { try { await migratePlaintextSecrets(filePath, secretStore); - snapshot = - parseOAuthPersistBlob(await readStoreFile(filePath)) ?? snapshot; + raw = await readStoreFile(filePath); + snapshot = parseOAuthPersistBlob(raw) ?? snapshot; } catch (error) { warnMigrationFailure(error); } } - return joinSnapshot(snapshot, secretStore); + // Reads never adopt: a legacy file's entries stay at their legacy ids + // until a write mints the namespace and moves them (#2549). + return joinSnapshot(snapshot, secretStore, parseSecretsNamespace(raw)); }); } @@ -962,16 +1258,25 @@ export async function removeOAuthStore( filePath: string, secretStore: SecretStore = defaultSecretStore(), ): Promise<void> { - await withOAuthStateLock(filePath, "remove", async () => { - const snapshot = await readDiskForMutation(filePath, "remove"); + await withOAuthStateLock(filePath, "remove", async (locked) => { + const rawBlob = await readStoreFile(filePath); + const snapshot = parseOAuthPersistBlob(rawBlob); + if (rawBlob !== null && snapshot === null) { + throw new OAuthStateFileUnrecognizedError(filePath, "remove"); + } if (snapshot) { + // Purge by the ids this file's entries actually use. A legacy + // (un-namespaced) file purges the legacy ids; a namespaced one must + // purge ONLY its own namespaced ids — the legacy ids may still be the + // live index of another, not-yet-adopted state file (#2549). + const namespace = parseSecretsNamespace(rawBlob); const targets = [ ...Object.entries(snapshot.servers).map(([url, state]) => ({ - id: oauthSecretServerId(url), + id: oauthSecretServerId(url, namespace), fields: serverSecretFields(state), })), ...Object.keys(snapshot.idpSessions).map((issuer) => ({ - id: oauthIdpSecretServerId(issuer), + id: oauthIdpSecretServerId(issuer, namespace), fields: [IDP_SESSION_FIELD], })), ]; @@ -995,6 +1300,13 @@ export async function removeOAuthStore( } else { await deleteStoreFile(filePath); } + // The file is gone, so every namespace the ledger records is now + // superseded — including ones an older Inspector's save stripped from + // it, which the purge above could not see (#2560). Locked only, for the + // reason adoption gives. + if (locked) { + await purgeSupersededNamespaces(filePath, secretStore, undefined); + } }); } diff --git a/core/auth/node/oauth-secrets.ts b/core/auth/node/oauth-secrets.ts index bf2c001f48..e62e991204 100644 --- a/core/auth/node/oauth-secrets.ts +++ b/core/auth/node/oauth-secrets.ts @@ -20,6 +20,7 @@ * route); the browser round-trips full snapshots over the authed local API. */ +import { randomUUID } from "node:crypto"; import { OAuthTokensSchema } from "@modelcontextprotocol/core"; import type { OAuthTokens } from "@modelcontextprotocol/client"; import { setOwnEntry } from "../../storage/own-entry.js"; @@ -74,11 +75,55 @@ export function resetPersistTokensPolicyWarnings(): void { * prefix of another's (`https://a` vs `https://a:8080`), letting * prefix-matching stores delete the wrong server's secrets. Encoding turns * `:` and `/` into `%3A`/`%2F`, which no other id can collide with. + * + * Ids are additionally scoped by the state file's secrets namespace + * (#2549): without it, two state files (profiles) that connect to the same + * server share one store entry, so profile B's login overwrites profile + * A's tokens and A silently acts as B. The namespace is a UUID stored in + * the state file itself (see `adoptSecretsNamespace` in + * `oauth-persist-file.ts`), inserted between the prefix and the encoded + * URL: `oauth+<namespace>+<encoded url>`. The delimiter stays unambiguous + * because `encodeURIComponent` escapes `+` (to `%2B`) and + * {@link isValidSecretsNamespace} rejects `+` (and `:`, keeping the id + * colon-free for the keyring account parse) — so a namespaced id can never + * equal a legacy unscoped one, and no (namespace, url) pair can produce + * another pair's id. `namespace === undefined` yields the legacy unscoped + * shape, still used to read (and migrate away from) pre-#2549 entries. */ -export const oauthSecretServerId = (serverUrl: string): string => - `oauth+${encodeURIComponent(serverUrl)}`; -export const oauthIdpSecretServerId = (issuer: string): string => - `oauth-idp+${encodeURIComponent(issuer)}`; +export const oauthSecretServerId = ( + serverUrl: string, + namespace?: string, +): string => + namespace === undefined + ? `oauth+${encodeURIComponent(serverUrl)}` + : `oauth+${namespace}+${encodeURIComponent(serverUrl)}`; +export const oauthIdpSecretServerId = ( + issuer: string, + namespace?: string, +): string => + namespace === undefined + ? `oauth-idp+${encodeURIComponent(issuer)}` + : `oauth-idp+${namespace}+${encodeURIComponent(issuer)}`; + +/** + * A usable secrets namespace: what `randomUUID()` produces, plus room for a + * hand-chosen value. The charset is what carries the id guarantees above — + * no `:` (keyring accounts parse at the first colon), no `+` (the id + * delimiter), and nothing `encodeURIComponent` leaves unescaped in a way + * that could forge a delimiter. Anything else in the file is ignored as if + * absent rather than propagated into store ids. + */ +export function isValidSecretsNamespace(value: unknown): value is string { + return ( + typeof value === "string" && + /^[A-Za-z0-9][A-Za-z0-9._-]{0,127}$/.test(value) + ); +} + +/** Mint a fresh secrets namespace for a state file that has none. */ +export function newSecretsNamespace(): string { + return randomUUID(); +} /** Field for one issuer's acquired tokens (JSON-serialized `OAuthTokens`). */ export const issuerTokensField = (issuer: string): string => `tokens:${issuer}`; diff --git a/core/auth/node/runner-interactive-oauth.ts b/core/auth/node/runner-interactive-oauth.ts index ce0dd199ee..6e084d85f8 100644 --- a/core/auth/node/runner-interactive-oauth.ts +++ b/core/auth/node/runner-interactive-oauth.ts @@ -40,6 +40,14 @@ export interface RunRunnerInteractiveOAuthOptions { onCallbackServer?: (server: OAuthCallbackServer) => void; /** Max wait for browser callback; defaults to {@link DEFAULT_RUNNER_INTERACTIVE_OAUTH_TIMEOUT_MS}. */ callbackTimeoutMs?: number; + /** + * Install process-wide SIGINT/SIGTERM handlers for the length of the wait + * so Ctrl-C rejects the flow cleanly (server stopped, classifiable error) + * instead of hanging or hitting Node's default abrupt exit. Opt-in + * because it is process-global state: the TUI owns Ctrl-C through Ink and + * must not have it intercepted here. CLI/mcpdo callers pass `true`. + */ + handleSignals?: boolean; } /** @@ -77,31 +85,79 @@ export async function runRunnerInteractiveOAuth( flowResolve = resolve; flowReject = reject; }); + // flowDone can reject before the Promise.race below ever subscribes — a + // signal (or an early callback error) while `server.start()` is still + // awaited would otherwise surface as an unhandled rejection. This no-op + // observer marks it handled for that window; the race still receives the + // rejection through its own subscription. + // void: intentional fire-and-forget rejection observer (see comment above) + void flowDone.catch(() => {}); + + // Ctrl-C / a caller killing the process while waiting on the loopback + // callback would otherwise either hang until the timeout below or (for + // SIGINT specifically, absent any handler) hit Node's default abrupt exit + // with no cleanup. Reject cleanly instead so the server is stopped and the + // caller gets a normal, classifiable error ("OAuth" in the message maps to + // AUTH_REQUIRED — see core/cli/error-handler.ts) rather than a raw + // process death. Opt-in (see handleSignals) — never installed under the + // TUI, which owns Ctrl-C through Ink. + // + // Every awaited phase — server.start(), authenticate() / + // beginInteractiveAuthorization(), the callback wait, and the challenge + // check — is raced against `signalAbort`: rejecting only flowDone would + // leave a signal during a stalled startup or authorization ignored until + // the final callback wait (and discarded entirely when authenticate() + // resolves undefined). + let signalAbortReject!: (err: Error) => void; + const signalAbort = new Promise<never>((_, reject) => { + signalAbortReject = reject; + }); + // Same pre-subscription window as flowDone above. + // void: intentional fire-and-forget rejection observer + void signalAbort.catch(() => {}); + const onSignal = (signal: NodeJS.Signals) => { + const err = new Error(`OAuth authorization cancelled (${signal}).`); + signalAbortReject(err); + flowReject(err); + }; + const racingSignals = <T>(work: Promise<T>): Promise<T> => { + if (!options.handleSignals) return work; + // If the signal wins the race, `work` is abandoned while still pending; + // observe its eventual rejection so it can't surface as unhandled. + void work.catch(() => {}); + return Promise.race([work, signalAbort]); + }; + if (options.handleSignals) { + process.on("SIGINT", onSignal); + process.on("SIGTERM", onSignal); + } let timeoutId: ReturnType<typeof setTimeout> | undefined; try { - const { redirectUrl } = await server.start({ - hostname: options.callbackListen.hostname, - port: options.callbackListen.port, - path: options.callbackListen.pathname, - onCallback: async (params) => { - try { - await options.client.completeOAuthFlow(params.code, params.iss); - flowResolve(); - } catch (err) { - flowReject(toRunnerOAuthError(err)); - } - }, - onError: (params) => { - flowReject( - new Error( - /* v8 ignore next -- params.error is a required non-null string, so the "OAuth error" fallback is unreachable */ - params.error_description ?? params.error ?? "OAuth error", - ), - ); - }, - }); + const { redirectUrl } = await racingSignals( + server.start({ + hostname: options.callbackListen.hostname, + port: options.callbackListen.port, + path: options.callbackListen.pathname, + onCallback: async (params) => { + try { + await options.client.completeOAuthFlow(params.code, params.iss); + flowResolve(); + } catch (err) { + flowReject(toRunnerOAuthError(err)); + } + }, + onError: (params) => { + flowReject( + new Error( + /* v8 ignore next -- params.error is a required non-null string, so the "OAuth error" fallback is unreachable */ + params.error_description ?? params.error ?? "OAuth error", + ), + ); + }, + }), + ); options.onCallbackServer?.(server); options.redirectUrlProvider.redirectUrl = redirectUrl; @@ -125,14 +181,20 @@ export async function runRunnerInteractiveOAuth( }, timeoutMs); }), ]); + // A signal can now abort the flow between this construction and the + // point waitForCallback is awaited (flowDone rejects but the racing + // authenticate()/beginInteractiveAuthorization() throws first); observe + // the rejection so that path can't surface it as unhandled. + // void: intentional fire-and-forget rejection observer + void waitForCallback.catch(() => {}); if (options.authorizationUrl) { - await options.client.beginInteractiveAuthorization( - options.authorizationUrl, + await racingSignals( + options.client.beginInteractiveAuthorization(options.authorizationUrl), ); await waitForCallback; } else { - const authUrl = await options.client.authenticate(); + const authUrl = await racingSignals(options.client.authenticate()); if (authUrl !== undefined) { await waitForCallback; } else { @@ -141,8 +203,8 @@ export async function runRunnerInteractiveOAuth( } if (options.authChallenge) { - const satisfied = await options.client.checkAuthChallengeSatisfied( - options.authChallenge, + const satisfied = await racingSignals( + options.client.checkAuthChallengeSatisfied(options.authChallenge), ); if (!satisfied) { return { @@ -154,6 +216,10 @@ export async function runRunnerInteractiveOAuth( return { kind: "success" }; } finally { + if (options.handleSignals) { + process.off("SIGINT", onSignal); + process.off("SIGTERM", onSignal); + } if (timeoutId !== undefined) { clearTimeout(timeoutId); } diff --git a/core/auth/node/secret-store-selection.ts b/core/auth/node/secret-store-selection.ts index 97b69b859f..2590b07585 100644 --- a/core/auth/node/secret-store-selection.ts +++ b/core/auth/node/secret-store-selection.ts @@ -443,14 +443,59 @@ async function describeFileStore( * Pick the fallback store for a host with no keychain. Exported for the * tests, which drive the container/mount predicates directly rather than * trying to make a real container appear. + * + * `allowMemory: false` (see {@link disallowMemorySecretStoreFallback}) + * collapses the container special case to `file`: a multi-process consumer + * can never be served by a per-process Map, so a plaintext file in an + * ephemeral layer — imperfect, but shared — is strictly better than a + * store two of its three processes cannot see. */ export function chooseFallbackKind(opts: { container: boolean; mounted: boolean; + allowMemory?: boolean; }): SecretStoreKind { + if (opts.allowMemory === false) return "file"; return opts.container && !opts.mounted ? "memory" : "file"; } +let memoryFallbackAllowed = true; + +/** + * Rule out `memory` as an automatic fallback for this process. + * + * For consumers that split one logical session across processes (the mcpdo + * front-end, its OAuth helper, and the connection daemon): an in-memory + * store is a per-process Map, so a token saved by one process is invisible + * to the others and OAuth can never complete. Call before the first store + * use — the selection is cached process-wide. An explicit + * `MCP_INSPECTOR_SECRET_STORE=memory` still wins: a user override is a + * statement of intent, this only shapes the automatic choice. + */ +export function disallowMemorySecretStoreFallback(): void { + memoryFallbackAllowed = false; +} + +let secretStorageWarningsQuiet = false; + +/** + * Silence the automatic {@link warnAboutSecretStorage} output for this + * process. + * + * For the same multi-process consumers as + * {@link disallowMemorySecretStoreFallback}: mcpdo runs a fresh front-end + * process for every command, and each one that touches a stored secret + * resolves the store and prints this banner — so an agent driving mcpdo + * sees the full fallback/caveat warning on *every* command (connect, + * `auth/list`, …). The front-end quiets it globally and re-surfaces it once, + * deliberately, on `connect` (passing `force`); the persistent daemon stays + * unquiet so the warning still lands once in its stderr log. A no-op for + * web/cli/tui, which never call this. + */ +export function setSecretStorageWarningsQuiet(quiet: boolean): void { + secretStorageWarningsQuiet = quiet; +} + /** * Print the fallback / plaintext warnings. * @@ -458,8 +503,16 @@ export function chooseFallbackKind(opts: { * who needs this is watching a terminal or `docker logs`, and store * selection happens before (and independently of) any logger being * configured. + * + * Suppressed when {@link setSecretStorageWarningsQuiet} is on, unless + * `force` is passed — the mcpdo front-end quiets the automatic (lazy) calls + * and re-emits once with `force` on `connect`. */ -export function warnAboutSecretStorage(info: SecretStorageInfo): void { +export function warnAboutSecretStorage( + info: SecretStorageInfo, + opts?: { force?: boolean }, +): void { + if (secretStorageWarningsQuiet && opts?.force !== true) return; if (info.reason === "fallback") { console.warn( `\n[mcp-inspector] The OS keychain is not available, so secrets will be kept in: ${secretStorageSummary(info)}.` + @@ -517,6 +570,7 @@ export function resolveSecretStore(): Promise<ResolvedSecretStore> { const kind = chooseFallbackKind({ container: isContainer(), mounted: isOnMountPoint(path.dirname(defaultSecretFilePath())), + allowMemory: memoryFallbackAllowed, }); const result = await buildStore(kind, "fallback", probe.detail); warnAboutSecretStorage(result.info); @@ -984,3 +1038,18 @@ class DeferredSecretStore implements SecretStore { export function defaultSecretStore(): SecretStore { return new DeferredSecretStore(); } + +/** + * The concrete store behind `store`: the selected one for a + * {@link defaultSecretStore}, otherwise `store` itself. For callers that + * must know *which* backend holds an entry — the OAuth namespace ledger + * (#2560) records keys per backend — since the deferred wrapper is + * deliberately indistinguishable from the store it forwards to. + */ +export async function resolveConcreteSecretStore( + store: SecretStore, +): Promise<SecretStore> { + return store instanceof DeferredSecretStore + ? (await resolveSecretStore()).store + : store; +} diff --git a/clients/cli/src/cli-oauth-navigation.ts b/core/cli/cli-oauth-navigation.ts similarity index 85% rename from clients/cli/src/cli-oauth-navigation.ts rename to core/cli/cli-oauth-navigation.ts index f805c08278..3e83e7edfb 100644 --- a/clients/cli/src/cli-oauth-navigation.ts +++ b/core/cli/cli-oauth-navigation.ts @@ -1,5 +1,5 @@ -import { CallbackNavigation } from "@inspector/core/auth/index.js"; -import { openUrl } from "./open-url.js"; +import { CallbackNavigation } from "../auth/index.js"; +import { openUrl } from "../node/openUrl.js"; import { createStyle, resolveAnsiEnabled } from "./style.js"; /** @@ -50,6 +50,15 @@ export type CliOAuthNavigationOptions = { * (`MCP_AUTO_OPEN_ENABLED=true`). */ forceAutoOpen?: boolean; + /** + * Build the printed prompt line for a given authorize URL. Receives the + * (possibly OSC-8-linked) display string and whether stderr is a TTY. + * Defaults to the CLI's own "Please navigate to: <url>" framing. Override + * when a different caller needs different wording — e.g. mcpdo, addressed to + * whatever is running it (which may be an agent that must relay the link to + * a human) rather than to a human reading the terminal directly. + */ + promptMessage?: (hrefDisplay: string, tty: boolean) => string; }; /** @@ -108,7 +117,10 @@ export function createCliOAuthNavigation( ); const write = options.write ?? ((line: string) => process.stderr.write(line)); - write(`Please navigate to: ${style.link(href)}\n`); + const promptMessage = + options.promptMessage ?? + ((hrefDisplay: string) => `Please navigate to: ${hrefDisplay}`); + write(`${promptMessage(style.link(href), tty)}\n`); const envAllows = options.autoOpenEnabled !== undefined diff --git a/clients/cli/src/cliOAuth.ts b/core/cli/cliOAuth.ts similarity index 91% rename from clients/cli/src/cliOAuth.ts rename to core/cli/cliOAuth.ts index 665c2d0a79..88265e44ed 100644 --- a/clients/cli/src/cliOAuth.ts +++ b/core/cli/cliOAuth.ts @@ -1,4 +1,4 @@ -import type { AuthChallenge } from "@inspector/core/auth/challenge.js"; +import type { AuthChallenge } from "../auth/challenge.js"; import { AuthRecoveryRequiredError, isStandardOAuthStepUp as isCoreStandardOAuthStepUp, @@ -6,16 +6,16 @@ import { stepUpConfirmMessage, stepUpInsufficientScopeMessage, MutableRedirectUrlProvider, -} from "@inspector/core/auth/index.js"; +} from "../auth/index.js"; import { createOAuthCallbackServer, runRunnerInteractiveOAuth, -} from "@inspector/core/auth/node/index.js"; -import type { RunnerInteractiveOAuthClient } from "@inspector/core/auth/node/runner-interactive-oauth.js"; -import type { RunnerOAuthCallbackConfig } from "@inspector/core/auth/node/runner-oauth-callback.js"; -import type { InspectorServerSettings } from "@inspector/core/mcp/types.js"; -import { isOAuthCapableServerConfig } from "@inspector/core/client/runner.js"; -import type { MCPServerConfig } from "@inspector/core/mcp/types.js"; +} from "../auth/node/index.js"; +import type { RunnerInteractiveOAuthClient } from "../auth/node/runner-interactive-oauth.js"; +import type { RunnerOAuthCallbackConfig } from "../auth/node/runner-oauth-callback.js"; +import type { InspectorServerSettings } from "../mcp/types.js"; +import { isOAuthCapableServerConfig } from "../client/runner.js"; +import type { MCPServerConfig } from "../mcp/types.js"; import { createInterface } from "node:readline/promises"; import { CliExitCodeError, EXIT_CODES } from "./error-handler.js"; import { @@ -63,6 +63,12 @@ export type CliOAuthConnectOptions = { * Default: {@link STEP_UP_PIPE_TIMEOUT_MS}. */ stepUpPromptTimeoutMs?: number; + /** + * `--quiet`: drop the "Authorization complete" status lines (#2435). The + * authorization URL and the step-up [y/N] still print — a human has to act + * on those for the flow to finish at all. + */ + quiet?: boolean; }; function authRequiredFailure(message: string): never { @@ -205,6 +211,7 @@ export async function runCliInteractiveOAuth( authorizationUrl?: URL; authChallenge?: AuthChallenge; autoOpenControl?: CliOAuthAutoOpenControl; + quiet?: boolean; }, ): Promise<void> { const result = await withArmedAutoOpen(options?.autoOpenControl, () => @@ -215,13 +222,15 @@ export async function runCliInteractiveOAuth( createCallbackServer: createOAuthCallbackServer, authorizationUrl: options?.authorizationUrl, authChallenge: options?.authChallenge, + // The CLI has no other Ctrl-C owner; cancel the wait cleanly. + handleSignals: true, }), ); if (result.kind === "insufficient_scope") { throw new Error(stepUpInsufficientScopeMessage(result.challenge)); } - if (result.kind === "success") { + if (result.kind === "success" && !options?.quiet) { process.stderr.write("Authorization complete.\n"); } } @@ -258,6 +267,7 @@ export async function handleCliAuthRecoveryRequired( await runCliInteractiveOAuth(client, redirectUrlProvider, callbackUrlConfig, { authorizationUrl: error.authorizationUrl, autoOpenControl: options?.autoOpenControl, + quiet: options?.quiet, ...(error.authChallenge.reason === "insufficient_scope" && { authChallenge: error.authChallenge, }), @@ -323,7 +333,7 @@ export async function connectInspectorWithOAuth( inspectorClient, redirectUrlProvider, callbackUrlConfig, - { autoOpenControl: options?.autoOpenControl }, + { autoOpenControl: options?.autoOpenControl, quiet: options?.quiet }, ); await inspectorClient.connect(); return; @@ -381,7 +391,9 @@ export async function withCliAuthRecoveryRetry<T>( // Belt-and-braces: this branch never disconnects today, so connect() is // usually a no-op (already connected). See connectInspectorWithOAuth. await inspectorClient.connect(); - process.stderr.write("Authorization complete. Retrying…\n"); + if (!options?.quiet) { + process.stderr.write("Authorization complete. Retrying…\n"); + } return await fn(); } @@ -399,12 +411,14 @@ export async function withCliAuthRecoveryRetry<T>( inspectorClient, redirectUrlProvider, callbackUrlConfig, - { autoOpenControl: options?.autoOpenControl }, + { autoOpenControl: options?.autoOpenControl, quiet: options?.quiet }, ); // Load-bearing: disconnect() above closed the session. // connect() is a no-op when already connected. await inspectorClient.connect(); - process.stderr.write("Authorization complete. Retrying…\n"); + if (!options?.quiet) { + process.stderr.write("Authorization complete. Retrying…\n"); + } return await fn(); } diff --git a/clients/cli/src/error-handler.ts b/core/cli/error-handler.ts similarity index 84% rename from clients/cli/src/error-handler.ts rename to core/cli/error-handler.ts index 9f49dbecc3..3bd6a318c5 100644 --- a/clients/cli/src/error-handler.ts +++ b/core/cli/error-handler.ts @@ -1,8 +1,8 @@ -import { redactUrlQuery } from "@inspector/core/mcp/fetchTracking.js"; +import { redactUrlQuery, redactUrlsInText } from "../mcp/fetchTracking.js"; import { awaitableError } from "./utils/awaitable-log.js"; -import { isUnauthorizedError } from "@inspector/core/auth/index.js"; -import { SecretStoreUnavailableError } from "@inspector/core/auth/node/secret-store.js"; -import { OAuthStateFileUnrecognizedError } from "@inspector/core/auth/node/oauth-persist-file.js"; +import { isUnauthorizedError } from "../auth/index.js"; +import { SecretStoreUnavailableError } from "../auth/node/secret-store.js"; +import { OAuthStateFileUnrecognizedError } from "../auth/node/oauth-persist-file.js"; /** * Exit-code map. Non-zero codes let an automated caller (CI, an agent) branch @@ -91,13 +91,18 @@ export interface ErrorEnvelope { * the real code and stderr. */ export class CliExitCodeError extends Error { + readonly exitCode: number; + readonly envelope?: Partial<ErrorEnvelope>; + constructor( - public readonly exitCode: number, + exitCode: number, message: string, - public readonly envelope?: Partial<ErrorEnvelope>, + envelope?: Partial<ErrorEnvelope>, ) { super(message); this.name = "CliExitCodeError"; + this.exitCode = exitCode; + this.envelope = envelope; } } @@ -142,59 +147,6 @@ function statusOf(error: unknown): number | undefined { const UNREACHABLE_PATTERN = /ENOTFOUND|ECONNREFUSED|ECONNRESET|EAI_AGAIN|ETIMEDOUT|fetch failed|getaddrinfo|connect(?:ion)? timed out|aborted/i; -/** - * An `http(s)://` URL embedded in free text. Stops at whitespace, at the - * double-quote/angle-bracket characters that commonly delimit a URL inside a - * message, and where a second `http(s)://` begins — so two URLs joined by a - * comma are redacted separately rather than the second one's query being read - * as part of the first one's last value (Copilot). An apostrophe is kept in the - * match because it is legal inside a query value; a *trailing* one is peeled - * off as punctuation below, which still handles a `'…'`-quoted URL. - * Case-insensitive because URI schemes are: `HTTPS://…?code=…` is the same - * URL and must not slip past the redaction (Copilot). - */ -const EMBEDDED_URL_PATTERN = /\bhttps?:\/\/(?:(?!https?:\/\/)[^\s"<>])+/gi; - -/** Sentence punctuation (or a closing quote) a message may put right after a URL. */ -const TRAILING_PUNCTUATION = new Set([ - ".", - ",", - ";", - ":", - "!", - "?", - ")", - "]", - "'", -]); - -/** - * Length of `match` once its trailing {@link TRAILING_PUNCTUATION} run is - * removed. A backward scan rather than an unanchored `/[…]+$/`: that regex - * rescans a punctuation run from every start position when the run does not - * end the string, which is quadratic, and the text here is server-controlled - * (an HTTP error body lands in the message), so a long `!!!…x` stalled the - * CLI's error path (#2540). - */ -function trailingPunctuationStart(match: string): number { - let end = match.length; - while (end > 0 && TRAILING_PUNCTUATION.has(match.charAt(end - 1))) end--; - return end; -} - -/** - * Apply {@link redactUrlQuery} to every URL embedded in `text`. Trailing - * sentence punctuation is split off first and re-appended, so a URL ending a - * sentence (`…?code=abc.`) keeps its full stop instead of having it folded into - * the redacted parameter value. - */ -function redactUrlsInText(text: string): string { - return text.replace(EMBEDDED_URL_PATTERN, (match) => { - const end = trailingPunctuationStart(match); - return redactUrlQuery(match.slice(0, end)) + match.slice(end); - }); -} - /** * Scrub query-string secrets out of every envelope field that can carry a URL * (#2423). The envelope is written verbatim to stderr — a terminal, a CI log, diff --git a/clients/cli/src/handlers/collect-app-info.ts b/core/cli/handlers/collect-app-info.ts similarity index 78% rename from clients/cli/src/handlers/collect-app-info.ts rename to core/cli/handlers/collect-app-info.ts index d46e37544c..af488dcf2c 100644 --- a/clients/cli/src/handlers/collect-app-info.ts +++ b/core/cli/handlers/collect-app-info.ts @@ -1,8 +1,8 @@ -import { InspectorClient } from "@inspector/core/mcp/index.js"; -import { extractAppInfo } from "@inspector/core/mcp/apps.js"; -import type { AppInfo } from "@inspector/core/mcp/apps.js"; +import { InspectorClient } from "../../mcp/index.js"; +import { extractAppInfo } from "../../mcp/apps.js"; +import type { AppInfo } from "../../mcp/apps.js"; import type { CliAppInfo } from "./method-types.js"; -import type { RequestMetadata } from "@inspector/core/mcp/types.js"; +import type { RequestMetadata } from "../../mcp/types.js"; /** * Build the CLI's app-info for a tool. Never throws — failures fold into diff --git a/clients/cli/src/handlers/connect-timeout.ts b/core/cli/handlers/connect-timeout.ts similarity index 96% rename from clients/cli/src/handlers/connect-timeout.ts rename to core/cli/handlers/connect-timeout.ts index 8bd8cf655c..90d81e0115 100644 --- a/clients/cli/src/handlers/connect-timeout.ts +++ b/core/cli/handlers/connect-timeout.ts @@ -2,7 +2,7 @@ import { DEFAULT_MAX_FETCH_REQUESTS, DEFAULT_TASK_TTL_MS, type InspectorServerSettings, -} from "@inspector/core/mcp/types.js"; +} from "../../mcp/types.js"; /** * Default connect timeout (ms) for ad-hoc server invocations. Without this an diff --git a/clients/cli/src/handlers/format-output.ts b/core/cli/handlers/format-output.ts similarity index 100% rename from clients/cli/src/handlers/format-output.ts rename to core/cli/handlers/format-output.ts diff --git a/clients/cli/src/handlers/method-types.ts b/core/cli/handlers/method-types.ts similarity index 71% rename from clients/cli/src/handlers/method-types.ts rename to core/cli/handlers/method-types.ts index 5bef12a05e..20e11773cc 100644 --- a/clients/cli/src/handlers/method-types.ts +++ b/core/cli/handlers/method-types.ts @@ -1,8 +1,9 @@ -import type { JsonValue } from "@inspector/core/mcp/index.js"; -import type { RequestMetadata } from "@inspector/core/mcp/types.js"; -import type { AppInfo } from "@inspector/core/mcp/apps.js"; +import type { JsonValue } from "../../mcp/index.js"; +import type { RequestMetadata } from "../../mcp/types.js"; +import type { AppInfo } from "../../mcp/apps.js"; import type { LoggingLevel } from "@modelcontextprotocol/client"; import type { OutputFormat } from "./format-output.js"; +import type { OutputFileFormat } from "./output-file.js"; export type { OutputFormat }; @@ -37,7 +38,22 @@ export type MethodArgs = { */ strict?: boolean; format?: OutputFormat; - /** Task id for tasks/get, tasks/cancel, tasks/result. */ + /** + * `--quiet` / `-q`: write only the result payload (or the error envelope). + * Advisory stderr lines — the schema-portability hint, the `--verify` + * summary, OAuth status lines — are dropped (#2435). Anything a human has to + * act on (an OAuth authorization URL, a step-up [y/N]) and anything the + * caller explicitly asked for (`--strict`'s report) still prints. + */ + quiet?: boolean; + /** + * `--output <path>`: write the result to this file instead of stdout + * (#2431). Only consulted on the single-result path (`emitResult`). + */ + output?: string; + /** `--output-format`: how the `--output` file is encoded (default `json`). */ + outputFormat?: OutputFileFormat; + /** Task id for tasks/get, tasks/cancel, tasks/result, tasks/update. */ taskId?: string; /** When true, tools/call uses callToolStream (task-augmented). */ task?: boolean; @@ -61,6 +77,12 @@ export type MethodArgs = { cursor?: string; /** roots/set payload (JSON array of {uri, name?}). */ rootsJson?: string; + /** + * tasks/update payload (JSON object keyed by the server's `inputRequests` + * ids). Resumes a modern (SEP-2663) task paused on `input_required` — + * modern-only, symmetric with `roots/set`'s JSON-blob convention. + */ + inputResponsesJson?: string; /** prompts/complete: argument name / value / ref. */ completeRefType?: "ref/prompt" | "ref/resource"; completeRef?: string; @@ -101,12 +123,18 @@ export type MethodOutcome = /** * Full method set supported by {@link runMethod}. * - * TODO(#1432): several of these (subscribe, tasks, roots, logging/tail, …) are - * not exposed by `mcp-inspector --cli` today; they exist for the experimental - * session CLI (`mcpi`) and other Node runners that share this dispatcher. + * Several of these (subscribe, tasks, roots, logging/tail, …) are not exposed + * by `mcp-inspector --cli`; they exist for the connection CLI (`mcpdo`) and + * other Node runners that share this dispatcher. + * + * Deliberately excludes `"initialize"` — that's still a valid {@link + * ONE_SHOT_METHODS} entry (scripting parity with the literal wire method + * name), but for `mcpdo` it read as "send another initialize", which it never + * did (it only replays cached connect-time state). `mcpdo connections/show` + * covers the same data (server info, capabilities, negotiated era) alongside + * daemon connection bookkeeping instead. */ -export const SESSION_RPC_METHODS = [ - "initialize", +export const CONNECTION_RPC_METHODS = [ "tools/list", "tools/call", "resources/list", @@ -124,13 +152,14 @@ export const SESSION_RPC_METHODS = [ "tasks/get", "tasks/cancel", "tasks/result", + "tasks/update", "roots/list", "roots/set", "skills/list", "skills/get", ] as const; -export type SessionRpcMethod = (typeof SESSION_RPC_METHODS)[number]; +export type ConnectionRpcMethod = (typeof CONNECTION_RPC_METHODS)[number]; /** * Methods accepted by `mcp-inspector --cli` (plus catalog-only diff --git a/core/cli/handlers/output-file.ts b/core/cli/handlers/output-file.ts new file mode 100644 index 0000000000..ee46c1719b --- /dev/null +++ b/core/cli/handlers/output-file.ts @@ -0,0 +1,206 @@ +/** + * `--output <path>` / `--output-format raw|json`: write a method result to a + * file instead of stdout (#2431). + * + * Kept in its own module, rather than inline in `cli.ts` and `emit-result.ts`, + * so the parse-time rules, the rendering and the write are one unit with one + * test file, and the two call sites stay a line each. + * + * The two formats mirror what the web client offers for the same result: + * + * - `json` (the default) is the whole result, pretty-printed with two-space + * indentation — the shape every web export (`useExportActions`) downloads. + * - `raw` is the result's own payload as a consumer would want it on disk: the + * text of its text-bearing blocks (what the web `ContentViewer`'s copy button + * copies), or — when the result carries no text and exactly one binary block — + * that block's decoded bytes, so an image or audio tool result saves as a + * playable file rather than as base64 inside JSON. + * + * The output flag name is deliberately not `--format`: that flag already exists + * and shapes **stdout** (`text` vs the `{ result }` envelope). Overloading it + * with a file encoding would make `--format json --output x` ambiguous. + */ +import { writeFile } from "node:fs/promises"; +import { base64ToBytes } from "../../mcp/skills.js"; +import { CliExitCodeError, EXIT_CODES } from "../error-handler.js"; +import type { McpResponse } from "./method-types.js"; + +export type OutputFileFormat = "raw" | "json"; + +export const OUTPUT_FILE_FORMATS: readonly OutputFileFormat[] = ["raw", "json"]; + +/** Methods whose result has a payload `raw` can extract. */ +const RAW_METHODS = ["tools/call", "resources/read"] as const; + +/** Commander parser for `--output-format`. */ +export function parseOutputFileFormat(value: string): OutputFileFormat { + if (value !== "raw" && value !== "json") { + throw new Error(`--output-format must be 'raw' or 'json'.`); + } + return value; +} + +/** The subset of the parsed CLI options the `--output` rules read. */ +export interface OutputOptionsInput { + output?: string; + outputFormat?: OutputFileFormat; + method?: string; + appInfo?: boolean; + verify?: boolean; + listStoredAuth?: boolean; + printHandoff?: boolean; +} + +/** + * Reject every `--output` combination that would be accepted and then silently + * do nothing. Called ahead of the CLI's short-circuit returns for the same + * reason `--strict` is: those paths never reach `emitResult`, so a later + * check would let them ignore the flag. + */ +export function validateOutputOptions(options: OutputOptionsInput): void { + if (options.output === undefined) { + if (options.outputFormat !== undefined) { + throw new Error("--output-format requires --output <path>."); + } + return; + } + if (options.output.trim() === "") { + throw new Error("--output requires a non-empty file path."); + } + if ( + options.listStoredAuth || + options.printHandoff || + options.method === "servers/list" || + options.method === "servers/show" + ) { + throw new Error( + "--output requires a method that calls a server; it has no effect with --list-stored-auth, --print-handoff, or --method servers/list / servers/show.", + ); + } + if (options.appInfo || options.verify) { + throw new Error( + "--output cannot be combined with --app-info or --verify, which emit a report rather than a result.", + ); + } + if ( + options.outputFormat === "raw" && + !(RAW_METHODS as readonly (string | undefined)[]).includes(options.method) + ) { + throw new Error( + `--output-format raw requires --method ${RAW_METHODS.join(" or ")}; use --output-format json for other methods.`, + ); + } +} + +type Payload = { text?: string; binary?: string }; + +function asRecord(value: unknown): Record<string, unknown> | undefined { + return value !== null && typeof value === "object" && !Array.isArray(value) + ? (value as Record<string, unknown>) + : undefined; +} + +/** + * One payload per block: its text, or its base64 binary. A `tools/call` block + * carries text on `text`, binary on `data` (image/audio), and an embedded + * resource nests either under `resource`; a `resources/read` entry carries + * `text` or `blob` directly. A `resource_link` carries neither and is skipped. + */ +function blockPayload(block: unknown): Payload | undefined { + const b = asRecord(block); + if (!b) return undefined; + const nested = asRecord(b.resource); + if (nested) return blockPayload(nested); + if (typeof b.text === "string") return { text: b.text }; + if (typeof b.data === "string") return { binary: b.data }; + if (typeof b.blob === "string") return { binary: b.blob }; + return undefined; +} + +/** + * Extract the `raw` bytes of a result: the text of its text-bearing blocks + * joined by newlines, or the decoded bytes of its single binary block when it + * has no text at all. Anything else (no payload, or several binaries with no + * text) has no unambiguous raw form and is refused with a pointer to `json`. + */ +export function renderRaw( + result: McpResponse, + method: string, +): string | Buffer { + const list = method === "resources/read" ? result.contents : result.content; + const payloads = (Array.isArray(list) ? list : []) + .map(blockPayload) + .filter((p): p is Payload => p !== undefined); + const texts = payloads.flatMap((p) => (p.text !== undefined ? [p.text] : [])); + if (texts.length > 0) return texts.join("\n"); + const binaries = payloads.flatMap((p) => + p.binary !== undefined ? [p.binary] : [], + ); + if (binaries.length === 1) { + // Strict decode: `Buffer.from(…, "base64")` silently drops invalid + // characters, so a malformed payload would be "written successfully" as + // unrelated bytes. `base64ToBytes` goes through `atob`, which throws. + try { + return Buffer.from(base64ToBytes(binaries[0]!)); + } catch { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `The ${method} result's binary content is not valid base64, so it has no raw form; use --output-format json.`, + { code: "output_not_raw" }, + ); + } + } + throw new CliExitCodeError( + EXIT_CODES.USAGE, + binaries.length === 0 + ? `The ${method} result has no text or binary content to write as raw; use --output-format json.` + : `The ${method} result has ${binaries.length} binary blocks and no text, so it has no single raw form; use --output-format json.`, + { code: "output_not_raw" }, + ); +} + +/** Render a result in the requested file format. */ +export function renderResultForFile( + result: McpResponse, + method: string, + format: OutputFileFormat, +): string | Buffer { + if (format === "raw") return renderRaw(result, method); + return JSON.stringify(result, null, 2) + "\n"; +} + +/** What was written, reported back on stdout (`--format json`) or stderr. */ +export interface WrittenOutput { + path: string; + format: OutputFileFormat; + bytes: number; +} + +/** + * Render and write a result to `path`. The parent directory must already exist + * (no implicit `mkdir`, as with `curl -o`), and an existing file is replaced. + * A failed write is a usage error with its own envelope code, so a script can + * tell "the tool failed" apart from "your path was wrong". + */ +export async function writeResultFile( + result: McpResponse, + method: string, + path: string, + format: OutputFileFormat = "json", +): Promise<WrittenOutput> { + const data = renderResultForFile(result, method, format); + try { + await writeFile(path, data); + } catch (err) { + throw new CliExitCodeError( + EXIT_CODES.USAGE, + `Could not write --output file ${path}: ${err instanceof Error ? err.message : String(err)}`, + { code: "output_write_failed" }, + ); + } + return { + path, + format, + bytes: typeof data === "string" ? Buffer.byteLength(data) : data.length, + }; +} diff --git a/clients/cli/src/handlers/run-method.ts b/core/cli/handlers/run-method.ts similarity index 68% rename from clients/cli/src/handlers/run-method.ts rename to core/cli/handlers/run-method.ts index 880c2b2572..e8ae570db9 100644 --- a/clients/cli/src/handlers/run-method.ts +++ b/core/cli/handlers/run-method.ts @@ -1,4 +1,4 @@ -import type { InspectorClient } from "@inspector/core/mcp/index.js"; +import type { InspectorClient } from "../../mcp/index.js"; import type { Root } from "@modelcontextprotocol/client"; import { ManagedToolsState, @@ -8,15 +8,15 @@ import { ManagedRequestorTasksState, ManagedSkillsState, MessageLogState, -} from "@inspector/core/mcp/state/index.js"; -import { SKILLS_EXTENSION_KEY } from "@inspector/core/mcp/skillsSchemas.js"; +} from "../../mcp/state/index.js"; +import { SKILLS_EXTENSION_KEY } from "../../mcp/skillsSchemas.js"; import { CliExitCodeError, EXIT_CODES } from "../error-handler.js"; import { collectAppInfo } from "./collect-app-info.js"; import { skillVerificationExitCode, summarizeSkillVerification, } from "./skills-verify.js"; -import { verifySkills } from "@inspector/core/mcp/skillsVerification.js"; +import { verifySkills } from "../../mcp/skillsVerification.js"; import type { CliAppInfo, McpResponse, @@ -24,6 +24,25 @@ import type { MethodOutcome, } from "./method-types.js"; +/** + * `JSON.parse` accepts numeric literals JSON cannot represent (`1e999` → + * `Infinity`); serializing the request for IPC/MCP would then silently send + * `null` instead of the value the user supplied. Reject anything that cannot + * round-trip. + */ +function assertJsonRoundTrips(value: unknown): void { + if (typeof value === "number" && !Number.isFinite(value)) { + throw new Error( + `${value} has no JSON representation (it would silently be sent as null)`, + ); + } + if (Array.isArray(value)) { + for (const item of value) assertJsonRoundTrips(item); + } else if (value !== null && typeof value === "object") { + for (const item of Object.values(value)) assertJsonRoundTrips(item); + } +} + /** * Refuse a `skills/*` call against a server that never declared the extension. * @@ -46,6 +65,32 @@ function assertSkillsSupported( } } +/** + * Live `resources/subscribe` stream consumers per client and URI. Streams + * for the same URI on one connection share a single core subscription + * (`subscribeToResource` is a no-op filter update when already subscribed), + * so the unsubscribe must be reference-counted: tearing it down when the + * first stream closes would leave the survivors open but silent. + */ +const resourceStreamRefs = new WeakMap< + InspectorClient, + Map<string, SharedResourceSubscription> +>(); + +/** + * One server-side subscription shared by every open subscribe stream for a + * given client + URI. The entry is the synchronization point for concurrent + * setups: it is reserved synchronously (before any await), so racing streams + * all join the same in-flight `ready` promise instead of each subscribing + * and corrupting the count. Consumers are counted from reservation; on + * subscribe failure each waiter rolls back its own reservation and the last + * one out removes the entry so a later subscribe can retry cleanly. + */ +type SharedResourceSubscription = { + count: number; + ready: Promise<void>; +}; + /** * Run one MCP method against a connected {@link InspectorClient}. * Core method dispatch used by the CLI (and other Inspector Node runners). @@ -190,23 +235,76 @@ export async function runMethod( "URI is required for resources/subscribe. Use --uri to specify the resource URI.", ); } - await inspectorClient.subscribeToResource(args.uri); + let refs = resourceStreamRefs.get(inspectorClient); + if (!refs) { + refs = new Map(); + resourceStreamRefs.set(inspectorClient, refs); + } + const uri = args.uri; + // Reserve before awaiting (see SharedResourceSubscription): the first + // arrival creates the entry with the in-flight subscribe, and every + // concurrent arrival joins it. The stream is only exposed once the + // shared subscribe has succeeded. + let shared = refs.get(uri); + if (!shared) { + shared = { count: 0, ready: inspectorClient.subscribeToResource(uri) }; + refs.set(uri, shared); + } + const entry = shared; + entry.count++; + // Attached BEFORE the subscribe handshake completes: a server may + // notify immediately after (or with) its subscribe response, and the + // stream's consumer only calls start() after this outcome crosses + // back through dispatch. Updates landing in that window are buffered + // and flushed to the first writeLine; ipc-glue guarantees every + // stream outcome is started (inert-started on a vanished caller), so + // stop() below always detaches this listener. + const buffered: Array<{ type: string; uri: string }> = []; + let sink: ((obj: unknown) => void) | undefined; + const onUpdate = (ev: Event) => { + const detail = (ev as CustomEvent<{ uri: string }>).detail; + // Multiple subscribe streams can share one connection; only + // forward updates for this stream's URI. Events without a uri + // (spec-noncompliant server) still pass through as before. + if (detail?.uri !== undefined && detail.uri !== uri) return; + const line = { type: "resources/updated", uri: detail?.uri ?? uri }; + if (sink) sink(line); + else buffered.push(line); + }; + inspectorClient.addEventListener("resourceUpdated", onUpdate); + try { + await entry.ready; + } catch (error) { + inspectorClient.removeEventListener("resourceUpdated", onUpdate); + entry.count--; + if (entry.count === 0 && refs.get(uri) === entry) refs.delete(uri); + throw error; + } return { kind: "stream", label: "resources/subscribe", start: (writeLine) => { writeLine({ type: "subscribed", uri: args.uri }); - const onUpdate = (ev: Event) => { - const detail = (ev as CustomEvent<{ uri: string }>).detail; - writeLine({ - type: "resources/updated", - uri: detail?.uri ?? args.uri, - }); - }; - inspectorClient.addEventListener("resourceUpdated", onUpdate); + for (const line of buffered) writeLine(line); + buffered.length = 0; + sink = writeLine; + let closed = false; return () => { + // A second stop from any caller must not double-decrement the + // shared count. + if (closed) return; + closed = true; inspectorClient.removeEventListener("resourceUpdated", onUpdate); - void inspectorClient.unsubscribeFromResource(args.uri!); + entry.count--; + if (entry.count > 0) return; + // Guard against deleting a successor generation: only remove + // the mapping if it is still this stream's entry. + if (refs.get(uri) === entry) refs.delete(uri); + // Catch the rejection here: this stop can run during daemon + // shutdown after disconnectAll has closed the client, where the + // unsubscribe rejects; a bare `void` would surface that as an + // unhandled rejection outside any caller's try/catch. + void inspectorClient.unsubscribeFromResource(uri).catch(() => {}); }; }, }; @@ -216,6 +314,17 @@ export async function runMethod( "URI is required for resources/unsubscribe. Use --uri to specify the resource URI.", ); } + // Subscribe streams share one server-side subscription per URI (see + // resourceStreamRefs above). An explicit unsubscribe here would tear + // that shared subscription down while the counted streams stay open + // and silent — and the last stream's cleanup would unsubscribe again. + const activeStreams = + resourceStreamRefs.get(inspectorClient)?.get(args.uri)?.count ?? 0; + if (activeStreams > 0) { + throw new Error( + `Cannot unsubscribe: ${activeStreams} active resources/subscribe stream(s) share this URI's subscription. Close those streams (Ctrl-C) instead; the subscription ends when the last one closes.`, + ); + } await inspectorClient.unsubscribeFromResource(args.uri); result = { unsubscribed: true, uri: args.uri }; } else if (args.method === "prompts/list") { @@ -313,6 +422,39 @@ export async function runMethod( result = (await inspectorClient.getRequestorTaskResult( args.taskId, )) as McpResponse; + } else if (args.method === "tasks/update") { + if (!args.taskId) { + throw new Error("Task id is required for tasks/update. Use --task-id."); + } + if (!args.inputResponsesJson) { + throw new Error( + "tasks/update requires --input-responses '<json object keyed by inputRequests id>'.", + ); + } + let inputResponses: Record<string, unknown>; + try { + const parsed: unknown = JSON.parse(args.inputResponsesJson); + if ( + typeof parsed !== "object" || + parsed === null || + Array.isArray(parsed) + ) { + throw new Error("must be a JSON object"); + } + assertJsonRoundTrips(parsed); + inputResponses = parsed as Record<string, unknown>; + } catch (e) { + throw new Error( + `--input-responses is invalid: ${e instanceof Error ? e.message : String(e)}`, + { cause: e }, + ); + } + await inspectorClient.updateRequestorTask(args.taskId, inputResponses); + // The server acks with an empty result and the task's status advances + // only on a subsequent tasks/get poll (updateRequestorTask says so) — + // so echo back what was actually sent rather than imply a fresher + // status is available here. + result = { updated: true, taskId: args.taskId }; } else if (args.method === "skills/list") { // The store's cursor walk is reused rather than re-implemented — it // carries the repeated-cursor and page-cap guards, and a second copy of diff --git a/clients/cli/src/handlers/servers-list.ts b/core/cli/handlers/servers-list.ts similarity index 79% rename from clients/cli/src/handlers/servers-list.ts rename to core/cli/handlers/servers-list.ts index d580f47c07..b09c88a264 100644 --- a/clients/cli/src/handlers/servers-list.ts +++ b/core/cli/handlers/servers-list.ts @@ -1,13 +1,15 @@ import type { InspectorServerSettings, MCPServerConfig, -} from "@inspector/core/mcp/types.js"; -import { InMemorySecretStore } from "@inspector/core/auth/node/secret-store.js"; +} from "../../mcp/types.js"; +import { InMemorySecretStore } from "../../auth/node/secret-store.js"; import { loadServerEntries, + resolveServerSource, selectServerEntry, + withDefaultCatalogPath, type ServerLoadOptions, -} from "@inspector/core/mcp/node/index.js"; +} from "../../mcp/node/index.js"; /** One catalog/config entry as returned by `servers/list`. */ export type ServerListEntry = { @@ -16,40 +18,61 @@ export type ServerListEntry = { /** Command line, URL, or other short identity for display. */ detail: string; /** - * Optional live-session name when a caller annotates catalog entries - * with connected sessions (omitted for plain catalog listing). + * Optional live-connection name when a caller annotates catalog entries + * with live connections (omitted for plain catalog listing). */ - session?: string; - /** True when that session is the most-recently-used connected session. */ + connection?: string; + /** True when that connection is the most-recently-used connection. */ isMru?: boolean; }; -/** Minimal session shape needed to annotate catalog entries. */ -export type SessionListRef = { +/** Minimal connection shape needed to annotate catalog entries. */ +export type ConnectionListRef = { name: string; isMru?: boolean; }; /** - * Mark catalog entries that have a live session with the same name. + * Where a server list came from: the writable catalog (default + * `~/.mcp-inspector/mcp.json`, or `--catalog` / `MCP_CATALOG_PATH`) or a + * read-only `--config` file. Surfaced by `servers/list` so users working + * across shells with different catalog env vars can see which file produced + * the entries. `null` for ad-hoc targets (no list source). + */ +export type ServerListSource = { kind: "catalog" | "config"; path: string }; + +/** + * Resolve the source `listServerEntries` would read for these options, + * applying the same default-catalog fallback. + */ +export function resolveServerListSource( + serverOptions: ServerLoadOptions = {}, +): ServerListSource | null { + const source = resolveServerSource(withDefaultCatalogPath(serverOptions)); + if (!source) return null; + return { kind: source.writable ? "catalog" : "config", path: source.path }; +} + +/** + * Mark catalog entries that have a live connection with the same name. * Does not mutate `entries`. * - * TODO(#1432): consumed by the experimental session CLI (`mcpi`); kept here so + * TODO(#1432): consumed by the experimental connection CLI (`mcpdo`); kept here so * that client can reuse catalog listing without duplicating this helper. */ -export function annotateServerEntriesWithSessions( +export function annotateServerEntriesWithConnections( entries: ServerListEntry[], - sessions: SessionListRef[], + connections: ConnectionListRef[], ): ServerListEntry[] { - if (sessions.length === 0) return entries; - const byName = new Map(sessions.map((s) => [s.name, s] as const)); + if (connections.length === 0) return entries; + const byName = new Map(connections.map((s) => [s.name, s] as const)); return entries.map((entry) => { - const session = byName.get(entry.name); - if (!session) return entry; + const connection = byName.get(entry.name); + if (!connection) return entry; return { ...entry, - session: session.name, - ...(session.isMru === true ? { isMru: true } : {}), + connection: connection.name, + ...(connection.isMru === true ? { isMru: true } : {}), }; }); } diff --git a/clients/cli/src/handlers/skills-verify.ts b/core/cli/handlers/skills-verify.ts similarity index 98% rename from clients/cli/src/handlers/skills-verify.ts rename to core/cli/handlers/skills-verify.ts index fe33940e1e..dc8148af8a 100644 --- a/clients/cli/src/handlers/skills-verify.ts +++ b/core/cli/handlers/skills-verify.ts @@ -17,7 +17,7 @@ import { anySkillFailed, anySkillUnverifiable, type SkillVerifyReport, -} from "@inspector/core/mcp/skillsVerification.js"; +} from "../../mcp/skillsVerification.js"; import { EXIT_CODES } from "../error-handler.js"; /** diff --git a/clients/cli/src/style.ts b/core/cli/style.ts similarity index 98% rename from clients/cli/src/style.ts rename to core/cli/style.ts index dc2ed15c53..748dc5a91b 100644 --- a/clients/cli/src/style.ts +++ b/core/cli/style.ts @@ -5,7 +5,7 @@ import type { OutputFormat } from "./handlers/format-output.js"; * * TODO(#1432): the CLI OAuth path only needs {@link Style.link} today; bold / * color helpers and {@link styleFromOpts} are used by the experimental session - * CLI (`mcpi`) human formatter. + * CLI (`mcpdo`) human formatter. */ export type Style = { /** Whether ANSI styling is enabled. */ diff --git a/clients/cli/src/utils/awaitable-log.ts b/core/cli/utils/awaitable-log.ts similarity index 100% rename from clients/cli/src/utils/awaitable-log.ts rename to core/cli/utils/awaitable-log.ts diff --git a/core/extension/tasks/constants.ts b/core/extension/tasks/constants.ts new file mode 100644 index 0000000000..aa7af9474b --- /dev/null +++ b/core/extension/tasks/constants.ts @@ -0,0 +1,7 @@ +/** + * Tasks extension identifiers the Inspector names outside the ext-tasks + * session — the advertised-extensions registry, the capability badges, and the + * negotiation check. Re-exported from `@modelcontextprotocol/ext-tasks` rather + * than restated, so the string has one owner and follows the package. + */ +export { TASKS_EXTENSION_ID_V2 as TASKS_EXTENSION_KEY } from "@modelcontextprotocol/ext-tasks/core/v2"; diff --git a/core/extension/tasks/errors.ts b/core/extension/tasks/errors.ts new file mode 100644 index 0000000000..847fe4273d --- /dev/null +++ b/core/extension/tasks/errors.ts @@ -0,0 +1,49 @@ +/** + * Error identity across the ext-tasks boundary. + * + * ext-tasks applies its own failure policy to whatever the host's raw + * dispatch throws: a transport or timeout failure comes back as a + * `DispatchError` whose `cause` is the original, and a JSON-RPC or task + * failure as `JsonRpcResponseError` / `TaskFailedError`. The Inspector's + * callers — and the UI that renders their messages — expect the SDK's own + * shapes (`ProtocolError`, an annotated `SdkError` timeout), so every task + * operation restores them on the way out through these helpers. + */ +import { ProtocolError, ProtocolErrorCode } from "@modelcontextprotocol/client"; +import { + DispatchError, + JsonRpcResponseError, + TaskFailedError, +} from "@modelcontextprotocol/ext-tasks/client"; + +/** The rejection for a signal that aborted: its own reason when that is an + * Error, else a standard `AbortError`. */ +export function abortError(signal: AbortSignal): Error { + return signal.reason instanceof Error + ? signal.reason + : new DOMException("The operation was aborted", "AbortError"); +} + +/** Restore host/transport and protocol error identity after ext-tasks policy. */ +export function unwrapTaskDispatchError(error: unknown): unknown { + const unwrapped = + error instanceof DispatchError && error.cause instanceof Error + ? error.cause + : error; + if (unwrapped instanceof JsonRpcResponseError) + return new ProtocolError(unwrapped.code, unwrapped.message, unwrapped.data); + if (unwrapped instanceof TaskFailedError && unwrapped.code !== undefined) + return new ProtocolError(unwrapped.code, unwrapped.message, unwrapped.data); + return unwrapped; +} + +/** Normalize any task failure to a `ProtocolError` for event payloads. */ +export function toProtocolError(reason: unknown): ProtocolError { + const normalized = unwrapTaskDispatchError(reason); + return normalized instanceof ProtocolError + ? normalized + : new ProtocolError( + ProtocolErrorCode.InternalError, + normalized instanceof Error ? normalized.message : String(normalized), + ); +} diff --git a/core/extension/tasks/index.ts b/core/extension/tasks/index.ts new file mode 100644 index 0000000000..4352d9092e --- /dev/null +++ b/core/extension/tasks/index.ts @@ -0,0 +1,29 @@ +/** + * The Inspector's adapter for the MCP Tasks extension. + * + * The protocol itself — both generations' schemas, the poll loop, task RPC, + * input rounds, the receiver side — is owned by `@modelcontextprotocol/ext-tasks` + * (#2316). What lives here is only what that package deliberately leaves to + * its host, plus the conversions between its shapes and the Inspector's: + * + * - `rawWireChannel` — the `rawDispatch` the package requires for 2026-07-28 + * traffic the SDK codec rejects. + * - `progress` — progress routing for task-backed calls, which the package + * does not handle. + * - `errors`, `wire`, `session` — error identity, JSON/task-view conversions, + * and the session's endpoint id and negotiation check. + * - `notificationSchemas` — `notifications/tasks/list_changed`, which neither + * the SDK nor the package defines. + * + * Nothing in this folder imports `InspectorClient`; host state reaches it + * through narrow interfaces. Keep it that way, so whatever the package later + * absorbs can be deleted here without touching the client, and so other + * extensions (Skills) can follow the same `core/extension/<name>/` shape. + */ +export * from "./constants.js"; +export * from "./errors.js"; +export * from "./notificationSchemas.js"; +export * from "./progress.js"; +export * from "./rawWireChannel.js"; +export * from "./session.js"; +export * from "./wire.js"; diff --git a/core/mcp/taskNotificationSchemas.ts b/core/extension/tasks/notificationSchemas.ts similarity index 100% rename from core/mcp/taskNotificationSchemas.ts rename to core/extension/tasks/notificationSchemas.ts diff --git a/core/extension/tasks/progress.ts b/core/extension/tasks/progress.ts new file mode 100644 index 0000000000..f8cf80af92 --- /dev/null +++ b/core/extension/tasks/progress.ts @@ -0,0 +1,67 @@ +/** + * Routes `notifications/progress` to the task-backed tool calls that own them. + * + * ext-tasks has no progress handling: once `tools/call` returns a task, the + * SDK's per-request progress handler is gone, so the Inspector watches the + * transport itself and needs to know which progress tokens belong to an + * in-flight task call and which task ids each token has been seen on. That + * bookkeeping lives here; `InspectorClient` only feeds it and dispatches the + * events it names. + */ +import type { ProgressToken } from "@modelcontextprotocol/client"; + +export class TaskProgressRouter { + /** + * Correlates task-call progress tokens to task ids after the first snapshot. + * A Set per token because concurrent calls may reuse a caller-supplied + * token; collapsing them to one task id would cross-wire progress between + * the calls. + */ + private readonly taskIds = new Map<ProgressToken, Set<string>>(); + /** Active task-call owners per progress token, before task correlation. */ + private readonly owners = new Map<ProgressToken, number>(); + + /** A task call carrying `token` has started. */ + acquire(token: ProgressToken): void { + this.owners.set(token, (this.owners.get(token) ?? 0) + 1); + } + + /** The call carrying `token` has observed task `taskId`. */ + correlate(token: ProgressToken, taskId: string): void { + const ids = this.taskIds.get(token) ?? new Set(); + ids.add(taskId); + this.taskIds.set(token, ids); + } + + /** + * A task call carrying `token` has settled. Releases only that call's own + * correlation; a concurrent call sharing the token keeps its entry. + */ + release(token: ProgressToken, taskId?: string): void { + const owners = this.owners.get(token) ?? 0; + if (owners <= 1) this.owners.delete(token); + else this.owners.set(token, owners - 1); + if (taskId === undefined) return; + const ids = this.taskIds.get(token); + ids?.delete(taskId); + if (ids?.size === 0) this.taskIds.delete(token); + } + + /** + * Who a progress notification for `token` belongs to: `undefined` when no + * task call owns it (the SDK's own handlers deliver it), otherwise every + * correlated task id — the wire cannot say which owner a shared token's + * progress is for, so each receives it. + */ + route(token: ProgressToken): readonly string[] | undefined { + const ids = this.taskIds.get(token); + if (!ids?.size && !this.owners.has(token)) return undefined; + return [...(ids ?? [])]; + } + + /** Forget every correlation (session reset). */ + clear(): void { + this.taskIds.clear(); + this.owners.clear(); + } +} diff --git a/core/extension/tasks/rawWireChannel.ts b/core/extension/tasks/rawWireChannel.ts new file mode 100644 index 0000000000..455062c5fc --- /dev/null +++ b/core/extension/tasks/rawWireChannel.ts @@ -0,0 +1,300 @@ +/** + * The raw request channel ext-tasks requires as `rawDispatch`. + * + * SDK v2's codec rejects the 2026-07-28 Tasks wire shapes — `tasks/*` throws + * `MethodNotSupportedByProtocolVersion` on the modern era, and a + * `resultType: "task"` result fails to decode — so ext-tasks deliberately + * leaves "send a JSON-RPC request below the SDK and hand back the response" to + * the host (`RawClientDispatch`). This is the Inspector's implementation: it + * writes the frame straight to the transport (which still logs it for the + * Protocol and Network tabs) under a string id the SDK's numeric ids cannot + * collide with, and `MessageTrackingTransport` hands the matching response to + * {@link RawWireChannel.consume} before the SDK would reject it as unknown. + * + * It mirrors what `Protocol.request` gives an ordinary request: a timeout + * (re-armed by progress when the session asks for that), caller aborts, and + * the SDK's cancellation fork — stream teardown on a per-request-stream + * transport, `notifications/cancelled` otherwise (#2140). + * + * Host state reaches it only through {@link RawWireChannelHost}, so the + * channel holds no reference to `InspectorClient`. + */ +import { SdkError, SdkErrorCode } from "@modelcontextprotocol/client"; +import type { + JSONRPCErrorResponse, + JSONRPCRequest, + JSONRPCResultResponse, + ProgressToken, + Transport, +} from "@modelcontextprotocol/client"; +import { DispatchError } from "@modelcontextprotocol/ext-tasks/client"; +import type { + DispatchOptions, + JsonRpcResponse, +} from "@modelcontextprotocol/ext-tasks/client"; +import type { JsonValue as TasksJsonValue } from "@modelcontextprotocol/ext-tasks/core"; +import { isSerializableJson, type JsonValue } from "../../json/jsonUtils.js"; +import { abortError } from "./errors.js"; + +/** Every raw request id starts with this, so responses can be told apart. */ +export const RAW_WIRE_ID_PREFIX = "inspector-ext-"; + +/** What the channel reads from its host, each time it needs it. */ +export interface RawWireChannelHost { + /** The connected transport, or null when there is none. */ + transport(): Transport | null; + /** The request budget when the call names none. */ + defaultTimeoutMs(): number; + /** Whether progress on a request's token re-arms its timeout. */ + resetTimeoutOnProgress(): boolean; + /** Decorate a timeout the way an SDK request's timeout is (#2318); any + * other error is returned untouched. */ + annotateTimeout(error: unknown, method: string): unknown; +} + +interface PendingRawRequest { + resolve: (response: JsonRpcResponse) => void; + reject: (error: unknown) => void; + cleanup: () => void; + // Present when the request carries a progress token and + // resetTimeoutOnProgress is enabled: re-arms the request's timeout. + progressToken?: ProgressToken; + resetTimeout?: () => void; +} + +export class RawWireChannel { + private readonly pending = new Map<string, PendingRawRequest>(); + private counter = 0; + private readonly host: RawWireChannelHost; + + constructor(host: RawWireChannelHost) { + this.host = host; + } + + /** Send one request and resolve with its JSON-RPC response. */ + async dispatch( + request: TasksJsonValue, + options: DispatchOptions = {}, + timeoutOverride?: number, + ): Promise<JsonRpcResponse> { + const transport = this.host.transport(); + if (!transport) + throw new DispatchError("MCP client is not connected", true); + if ( + request === null || + Array.isArray(request) || + typeof request !== "object" + ) { + throw new DispatchError("Raw MCP request must be a JSON object"); + } + const record = request as Readonly<Record<string, JsonValue>>; + if (typeof record.method !== "string") { + throw new DispatchError("Raw MCP request method must be a string"); + } + const params = record.params; + if ( + params !== undefined && + (params === null || Array.isArray(params) || typeof params !== "object") + ) { + throw new DispatchError("Raw MCP request params must be a JSON object"); + } + const signal = options.signal; + if (signal?.aborted) throw abortError(signal); + + const id = `${RAW_WIRE_ID_PREFIX}${(this.counter += 1)}`; + const message: JSONRPCRequest = { + jsonrpc: "2.0", + id, + method: record.method, + ...(params === undefined ? {} : { params }), + }; + const timeoutMs = + timeoutOverride ?? + options.context?.requestTimeoutMs ?? + this.host.defaultTimeoutMs(); + + // Extract the request's progress token (if any) because + // notifications/progress uses it to re-arm this request's timeout, + // matching the SDK path's resetTimeoutOnProgress. + const meta = (params as Readonly<Record<string, JsonValue>> | undefined)?.[ + "_meta" + ]; + const progressToken = + meta !== null && typeof meta === "object" && !Array.isArray(meta) + ? (meta as { progressToken?: ProgressToken }).progressToken + : undefined; + + return await new Promise<JsonRpcResponse>((resolve, reject) => { + let onAbort: (() => void) | undefined; + // The transport sees this controller's signal, not the caller's, + // because a timeout must also reach the wire: aborting it tears down a + // per-request stream (the 2026-era cancellation signal), which the + // caller's untouched signal cannot do. Caller aborts forward into it. + const wireController = new AbortController(); + const forwardAbort = () => { + wireController.abort(signal?.reason); + }; + signal?.addEventListener("abort", forwardAbort, { once: true }); + const cleanup = () => { + clearTimeout(timer); + if (signal && onAbort) signal.removeEventListener("abort", onAbort); + signal?.removeEventListener("abort", forwardAbort); + this.pending.delete(id); + }; + // Mirror the SDK's cancellation fork (#2140) for both local endings of + // a raw request: a per-request-stream transport (2026-era Streamable + // HTTP) treats the forwarded requestSignal abort as the wire + // cancellation, but stdio/SSE ignore requestSignal — and this path + // bypasses Client.request, so nothing else sends the + // notifications/cancelled frame they need. Without it a timed-out or + // aborted tools/call keeps running server-side (orphaning any task) and + // its late response is no longer consumed by this raw channel. + const sendWireCancellation = (reason?: string) => { + if (transport.hasPerRequestStream === true) return; + void transport + .send({ + jsonrpc: "2.0", + method: "notifications/cancelled", + params: { + requestId: id, + ...(reason === undefined ? {} : { reason }), + }, + }) + .catch(() => { + // Best effort: the local rejection is authoritative. + }); + }; + const onTimeout = () => { + cleanup(); + const timeoutReason = `Request timed out after ${String(timeoutMs)} ms`; + // Both wire paths, matching the abort fork: stream teardown for + // per-request-stream transports, notifications/cancelled otherwise. + wireController.abort(new DispatchError(timeoutReason)); + sendWireCancellation(timeoutReason); + // The same error, with the same annotation, as an SDK request that + // times out: this path bypasses `Protocol.request`, so it builds the + // SDK's own timeout shape and runs it through the decorator's + // annotation by hand — a raw-wire caller sees one kind of timeout, + // not two (#2318). It rides as the DispatchError's cause, which + // `unwrapTaskDispatchError` restores once ext-tasks hands it back. + reject( + new DispatchError( + `Raw MCP request "${message.method}" timed out after ${timeoutMs} ms`, + false, + { + cause: this.host.annotateTimeout( + new SdkError(SdkErrorCode.RequestTimeout, "Request timed out", { + timeout: timeoutMs, + }), + message.method, + ), + }, + ), + ); + }; + let timer = setTimeout(onTimeout, timeoutMs); + const resetTimeout = + progressToken !== undefined && this.host.resetTimeoutOnProgress() + ? () => { + clearTimeout(timer); + timer = setTimeout(onTimeout, timeoutMs); + } + : undefined; + + this.pending.set(id, { + resolve, + reject, + cleanup, + progressToken, + resetTimeout, + }); + if (signal) { + onAbort = () => { + const pending = this.pending.get(id); + if (!pending) return; + pending.cleanup(); + const reason = signal.reason; + sendWireCancellation(typeof reason === "string" ? reason : undefined); + reject(abortError(signal)); + }; + signal.addEventListener("abort", onAbort, { once: true }); + } + transport + .send(message, { + ...(options.context?.headers === undefined + ? {} + : { headers: options.context.headers }), + requestSignal: wireController.signal, + }) + .catch((error: unknown) => { + const pending = this.pending.get(id); + if (!pending) return; + pending.cleanup(); + // The browser's remote transport awaits the response inside `send`, + // so its relay wait can expire here first, as the SDK's timeout + // shape; annotate it exactly as the local timer above does (a + // non-timeout error passes through untouched). + reject( + this.host.annotateTimeout( + error instanceof Error ? error : new Error(String(error)), + message.method, + ), + ); + }); + }); + } + + /** + * Settle the raw request a response answers. Returns false — leaving the + * frame to the SDK — only for an id this channel did not issue. A response + * to one it issued but no longer holds (the request timed out or was + * aborted, and the server answered anyway) is swallowed: the SDK only knows + * numeric ids, so forwarding it would surface a spurious "unknown message + * ID" error. + */ + consume(message: JSONRPCResultResponse | JSONRPCErrorResponse): boolean { + const { id } = message; + if (typeof id !== "string" || !id.startsWith(RAW_WIRE_ID_PREFIX)) { + return false; + } + const pending = this.pending.get(id); + if (!pending) return true; + + pending.cleanup(); + if ("error" in message) { + const { error } = message; + pending.resolve({ + kind: "error", + error: { + code: error.code, + message: error.message, + ...(isSerializableJson(error.data) ? { data: error.data } : {}), + }, + }); + } else if (!isSerializableJson(message.result)) { + pending.reject( + new DispatchError(`Raw MCP request ${id} returned a non-JSON result`), + ); + } else { + pending.resolve({ kind: "result", result: message.result }); + } + return true; + } + + /** Re-arm the timeout of every pending request carrying `token`. */ + noteProgress(token: ProgressToken): void { + for (const pending of this.pending.values()) { + if (pending.progressToken === token) pending.resetTimeout?.(); + } + } + + /** Reject everything in flight (connection ended or disconnected). */ + rejectAll(reason: string): void { + const pendingRequests = [...this.pending.values()]; + this.pending.clear(); + for (const pending of pendingRequests) { + pending.cleanup(); + pending.reject(new DispatchError(reason)); + } + } +} diff --git a/core/extension/tasks/session.ts b/core/extension/tasks/session.ts new file mode 100644 index 0000000000..190918b8ab --- /dev/null +++ b/core/extension/tasks/session.ts @@ -0,0 +1,68 @@ +/** + * Inputs to an ext-tasks session that the Inspector derives from its own + * connection: the stable endpoint identity task references are scoped to, and + * whether the modern Tasks extension was negotiated at all. + */ +import type { + Implementation, + ProtocolEra, + ServerCapabilities, +} from "@modelcontextprotocol/client"; +import { createTaskSessionEndpointId } from "@modelcontextprotocol/ext-tasks/client"; +import type { TaskSessionEndpointId } from "@modelcontextprotocol/ext-tasks/client"; +import type { MCPServerConfig } from "../../mcp/types.js"; +import { TASKS_EXTENSION_KEY } from "./constants.js"; + +/** + * The endpoint id a task session scopes its serialized task references to: + * the client identity plus the transport target, so a reference taken against + * one server is never resumed against another. + */ +export function taskSessionEndpointId( + config: MCPServerConfig, + clientInfo: Implementation, +): Promise<TaskSessionEndpointId> { + return createTaskSessionEndpointId( + "inspector", + config.type === "sse" || config.type === "streamable-http" + ? { + host: clientInfo, + transport: { type: config.type, url: new URL(config.url).toString() }, + } + : { + host: clientInfo, + transport: { + type: "stdio", + command: config.command, + args: config.args, + cwd: config.cwd ?? null, + }, + }, + ); +} + +/** + * True when the connection is modern (2026-07-28) AND the server advertised + * the `io.modelcontextprotocol/tasks` extension (SEP-2663) in its + * `server/discover` capabilities. Legacy servers use `capabilities.tasks`. + * + * Accepts the value only in the shape ext-tasks itself accepts — a plain, + * empty object — because the package starts a 2026-07-28 task session on + * nothing else; a looser check here would show Tasks for a server the + * attached session treats as unsupported. + */ +export function isTasksExtensionNegotiated( + era: ProtocolEra | undefined, + capabilities: ServerCapabilities | undefined, +): boolean { + if (era !== "modern") return false; + // Typed as unknown: the SDK types the value as an object, but a malformed + // server can still send null, an array, or a scalar on the wire. + const extension: unknown = capabilities?.extensions?.[TASKS_EXTENSION_KEY]; + return ( + extension !== null && + typeof extension === "object" && + !Array.isArray(extension) && + Object.keys(extension).length === 0 + ); +} diff --git a/core/extension/tasks/wire.ts b/core/extension/tasks/wire.ts new file mode 100644 index 0000000000..e193f59bf1 --- /dev/null +++ b/core/extension/tasks/wire.ts @@ -0,0 +1,92 @@ +/** + * Pure conversions between the Inspector's SDK-typed values and the JSON + * shapes ext-tasks takes and returns: the tool-result codec, JSON-object + * projection, related-task `_meta` stamping, the task view the UI renders, and + * the pending-request origin for an input exchange. No state, no I/O — each is + * input → output, which is what lets the session wiring in `InspectorClient` + * stay a thin composition of them. + */ +import { CallToolResultSchema } from "@modelcontextprotocol/core"; +import type { CallToolResult } from "@modelcontextprotocol/client"; +import { withRelatedTaskMetadata } from "@modelcontextprotocol/ext-tasks/client"; +import type { + ResolvedInputExchangeContext, + TaskView, +} from "@modelcontextprotocol/ext-tasks/client"; +import { + runtimeCodecFromStandardSchema, + taskId as extTaskId, + toJsonValue, +} from "@modelcontextprotocol/ext-tasks/core"; +import type { JsonValue as TasksJsonValue } from "@modelcontextprotocol/ext-tasks/core"; +import type { InspectorTask, PendingRequestOrigin } from "../../mcp/types.js"; + +export type TasksJsonObject = Readonly<Record<string, TasksJsonValue>>; + +/** Validates a task's terminal `tools/call` payload with the SDK's own schema. */ +export const taskToolResultCodec = + runtimeCodecFromStandardSchema<CallToolResult>({ + "~standard": { + version: 1, + vendor: "mcp-inspector", + validate(value) { + const result = CallToolResultSchema.safeParse(value); + return result.success + ? { value: result.data as CallToolResult } + : { + issues: result.error.issues.map(({ message }) => ({ message })), + }; + }, + }, + }); + +/** Project a value to a JSON object, throwing when it is not one. */ +export function jsonObject(value: unknown): TasksJsonObject { + const json = toJsonValue(value); + if (json === null || Array.isArray(json) || typeof json !== "object") { + throw new TypeError("Expected a JSON object"); + } + const object: Record<string, TasksJsonValue> = {}; + for (const [key, member] of Object.entries(json)) object[key] = member; + return object; +} + +/** Stamp `io.modelcontextprotocol/related-task` onto request params, keeping + * any existing `_meta`, so the pending-request UI can tag it with its task. */ +export function paramsWithRelatedTask( + params: TasksJsonObject, + taskId: string, +): TasksJsonObject { + const rawMetadata = params._meta; + const metadata = + rawMetadata !== null && + !Array.isArray(rawMetadata) && + typeof rawMetadata === "object" + ? (rawMetadata as TasksJsonObject) + : undefined; + return { + ...params, + _meta: withRelatedTaskMetadata(metadata, { taskId: extTaskId(taskId) }), + }; +} + +/** The generation-neutral task view, with the timestamps the UI sorts on + * always present (a modern task may omit either). */ +export function toInspectorTask(view: TaskView): InspectorTask { + const timestamp = view.createdAt ?? view.lastUpdatedAt ?? ""; + return { + ...view, + createdAt: timestamp, + lastUpdatedAt: view.lastUpdatedAt ?? timestamp, + }; +} + +/** Which pending-request origin an ext-tasks input exchange surfaces as, so + * the UI can show era-accurate semantics. */ +export function taskInputOrigin( + delivery: ResolvedInputExchangeContext["delivery"], +): PendingRequestOrigin { + if (delivery === "task-update") return "task-input-required"; + if (delivery === "request-retry") return "input-required"; + return "server-request"; +} diff --git a/core/mcp/__tests__/fakeInspectorClient.ts b/core/mcp/__tests__/fakeInspectorClient.ts index 41bcfcc135..2c230df3be 100644 --- a/core/mcp/__tests__/fakeInspectorClient.ts +++ b/core/mcp/__tests__/fakeInspectorClient.ts @@ -38,6 +38,7 @@ import type { ExcludedTool, RequestMetadata, } from "../types.js"; +import type { TaskCapabilities } from "@modelcontextprotocol/ext-tasks/client"; import { INACTIVE_SUBSCRIPTION_STREAM_STATE } from "../types.js"; import type { MalformedListItem } from "../listSalvage.js"; import type { ConnectionDiagnostics } from "../connectionDiagnostics.js"; @@ -149,6 +150,14 @@ export class FakeInspectorClient return this.tasksExtensionNegotiated; } + // Authoritative task-session capabilities this fake presents. Tests assign + // a TaskCapabilities object to exercise inventory-gated refresh routing; + // undefined means "no task session" (inventory treated as unsupported). + taskSessionCapabilities: TaskCapabilities | undefined = undefined; + getTaskSessionCapabilities(): TaskCapabilities | undefined { + return this.taskSessionCapabilities; + } + // The Skills extension (SEP-2640) this fake presents. `undefined` means the // server declared none, which is what `getSkillsExtension` returns then — // tests assign a support object to exercise the skills paths. diff --git a/core/mcp/extensions.ts b/core/mcp/extensions.ts index 24ba650fae..538c4655d6 100644 --- a/core/mcp/extensions.ts +++ b/core/mcp/extensions.ts @@ -1,6 +1,6 @@ import type { ClientCapabilities } from "@modelcontextprotocol/client"; import { RESOURCE_MIME_TYPE } from "@modelcontextprotocol/ext-apps/app-bridge"; -import { TASKS_EXTENSION_KEY } from "./modernTaskSchemas.js"; +import { TASKS_EXTENSION_KEY } from "../extension/tasks/constants.js"; import { SKILLS_EXTENSION_KEY } from "./skillsSchemas.js"; /** diff --git a/core/mcp/fetchTracking.ts b/core/mcp/fetchTracking.ts index 4b5a1366b5..decc5faae7 100644 --- a/core/mcp/fetchTracking.ts +++ b/core/mcp/fetchTracking.ts @@ -107,6 +107,65 @@ export function redactUrlQuery(url: string): string { } } +/** + * An `http(s)://` URL embedded in free text. Stops at whitespace, at the + * double-quote/angle-bracket characters that commonly delimit a URL inside a + * message, and where a second `http(s)://` begins — so two URLs joined by a + * comma are redacted separately rather than the second one's query being read + * as part of the first one's last value (Copilot). An apostrophe is kept in the + * match because it is legal inside a query value; a *trailing* one is peeled + * off as punctuation below, which still handles a `'…'`-quoted URL. + * Case-insensitive because URI schemes are: `HTTPS://…?code=…` is the same + * URL and must not slip past the redaction (Copilot). + */ +const EMBEDDED_URL_PATTERN = /\bhttps?:\/\/(?:(?!https?:\/\/)[^\s"<>])+/gi; + +/** Sentence punctuation (or a closing quote) a message may put right after a URL. */ +const TRAILING_PUNCTUATION = new Set([ + ".", + ",", + ";", + ":", + "!", + "?", + ")", + "]", + "'", +]); + +/** + * Length of `match` once its trailing {@link TRAILING_PUNCTUATION} run is + * removed. A backward scan rather than an unanchored `/[…]+$/`: that regex + * rescans a punctuation run from every start position when the run does not + * end the string, which is quadratic, and the text here is server-controlled + * (an HTTP error body lands in the message), so a long `!!!…x` stalled the + * CLI's error path (#2540). + */ +function trailingPunctuationStart(match: string): number { + let end = match.length; + while (end > 0 && TRAILING_PUNCTUATION.has(match.charAt(end - 1))) end--; + return end; +} + +/** + * Apply {@link redactUrlQuery} to every URL embedded in `text`. Trailing + * sentence punctuation is split off first and re-appended, so a URL ending a + * sentence (`…?code=abc.`) keeps its full stop instead of having it folded into + * the redacted parameter value. + * + * This is the free-text counterpart of {@link redactUrlQuery}: every client + * routes error text it shows or writes through it — the CLI's stderr envelope + * (#2423) and the web and TUI clients' on-screen error messages (#2490) — so a + * server or SDK error quoting `https://…?code=…` never reaches a terminal, + * toast or screenshot verbatim. + */ +export function redactUrlsInText(text: string): string { + return text.replace(EMBEDDED_URL_PATTERN, (match) => { + const end = trailingPunctuationStart(match); + return redactUrlQuery(match.slice(0, end)) + match.slice(end); + }); +} + /** Recursively redact sensitive keys in a parsed JSON value (in place). */ function redactJsonValue(value: unknown): unknown { if (Array.isArray(value)) { diff --git a/core/mcp/inspectorClient.ts b/core/mcp/inspectorClient.ts index 3af2e7a379..bee4faeb27 100644 --- a/core/mcp/inspectorClient.ts +++ b/core/mcp/inspectorClient.ts @@ -1,4 +1,34 @@ -import { Client } from "@modelcontextprotocol/client"; +import { + Client, + DEFAULT_REQUEST_TIMEOUT_MSEC, + isInputRequiredResult, + withInputRequired, +} from "@modelcontextprotocol/client"; +import { + createApplicationInputHandler, + createTaskSessionFromClient, + resultFromTaskOutcome, + taskViewFromExecutionEvent, + toolDeclarationFromMcpTool, +} from "@modelcontextprotocol/ext-tasks/client"; +import type { + ApplicationElicitContentValue, + ApplicationSamplingContentBlock, + ApplicationRoot, + JsonRpcResponse, + RawClientDispatch, + SerializedTaskReference, + TaskCapabilities, + TaskEnabledSession, + TaskExecutionEvent, +} from "@modelcontextprotocol/ext-tasks/client"; +import { bindTaskReceiver } from "@modelcontextprotocol/ext-tasks/receiver"; +import type { TaskReceiverBinding } from "@modelcontextprotocol/ext-tasks/receiver"; +import { + taskId as extTaskId, + toJsonValue, +} from "@modelcontextprotocol/ext-tasks/core"; +import type { JsonValue as TasksJsonValue } from "@modelcontextprotocol/ext-tasks/core"; // The protocol's own schemas for the reserved `_meta` members, so the client // validates against the SDK rather than a restatement of it that can drift. import { @@ -25,6 +55,7 @@ import type { ResourceSubscriptionStreamState, ExcludedTool, RequestMetadata, + InspectorTask, } from "./types.js"; import { scanXMcpHeaderDeclarations, @@ -74,12 +105,11 @@ import { type MessageTrackingCallbacks, } from "./messageTrackingTransport.js"; import type { - CallToolRequest, JSONRPCRequest, JSONRPCNotification, JSONRPCResultResponse, JSONRPCErrorResponse, - JSONRPCMessage, + StandardSchemaV1, ServerCapabilities, ClientCapabilities, Implementation, @@ -91,7 +121,6 @@ import type { Root, CreateMessageRequest, CreateMessageResult, - CreateTaskResult, ElicitRequest, ElicitResult, ElicitRequestURLParams, @@ -119,32 +148,32 @@ import type { DiscoverResult, InputRequests, InputRequiredOptions, - StandardSchemaV1, McpSubscription, SubscriptionFilter, } from "@modelcontextprotocol/client"; -import { ProtocolError, ProtocolErrorCode } from "@modelcontextprotocol/client"; import { - isInputRequiredResult, - withInputRequired, + ProtocolError, + ProtocolErrorCode, LOG_LEVEL_META_KEY, CLIENT_CAPABILITIES_META_KEY, CLIENT_INFO_META_KEY, PROTOCOL_VERSION_META_KEY, - RELATED_TASK_META_KEY, } from "@modelcontextprotocol/client"; +import { MODERN_PROTOCOL_VERSION } from "./types.js"; import { - TASKS_EXTENSION_KEY, - MODERN_TASK_HANDLE_META, - MODERN_PROTOCOL_VERSION, - ModernGetTaskResultSchema, - ModernUpdateTaskResultSchema, - ModernCancelTaskResultSchema, - normalizeModernTask, - readInputRequests, - isModernCreateTaskResult, - type ModernDetailedTask, -} from "./modernTaskSchemas.js"; + TasksListChangedNotificationSchema, + TaskProgressRouter, + RawWireChannel, + isTasksExtensionNegotiated, + jsonObject, + paramsWithRelatedTask, + taskInputOrigin, + taskSessionEndpointId, + taskToolResultCodec, + toInspectorTask, + toProtocolError, + unwrapTaskDispatchError, +} from "../extension/tasks/index.js"; import { buildClientExtensions } from "./extensions.js"; import { DirectoryReadResultSchema, @@ -173,19 +202,7 @@ import { CallToolResultSchema, GetPromptResultSchema, ReadResourceResultSchema, - // Task request schemas — used for `.shape.params` in the 3-arg custom - // `setRequestHandler` form (tasks/* are excluded from v2's spec-method set). - ListTasksRequestSchema, - GetTaskRequestSchema, - GetTaskPayloadRequestSchema, - CancelTaskRequestSchema, TaskStatusNotificationSchema, - // Task result schemas — explicit result schemas for the raw requestor-task - // requests that replace the removed `client.experimental.tasks.*` helpers. - CreateTaskResultSchema, - GetTaskResultSchema, - CancelTaskResultSchema, - ListTasksResultSchema, // List result schemas — used by the single-page list methods below. SDK v2's // high-level `client.listTools()` etc. auto-aggregate ALL pages (returning // `nextCursor: undefined`), which defeats the Inspector's pagination-debugging @@ -203,7 +220,6 @@ import { ResourceTemplateSchema, PromptSchema, } from "@modelcontextprotocol/core"; -import type { ClientResult } from "@modelcontextprotocol/client"; import { AjvJsonSchemaValidator } from "@modelcontextprotocol/client/validators/ajv"; import { z } from "zod/v4"; import { validateToolOutput } from "./toolOutputValidation.js"; @@ -221,17 +237,13 @@ import { summarizeMalformed, type MalformedListItem, } from "./listSalvage.js"; -import { TasksListChangedNotificationSchema } from "./taskNotificationSchemas.js"; import { type JsonValue, convertToolParameters, convertPromptArguments, } from "../json/jsonUtils.js"; import { expandUriTemplateStrict } from "./uriTemplate.js"; -import { - InspectorClientEventTarget, - type TaskWithOptionalCreatedAt, -} from "./inspectorClientEventTarget.js"; +import { InspectorClientEventTarget } from "./inspectorClientEventTarget.js"; import { SamplingCreateMessage } from "./samplingCreateMessage.js"; import { ElicitationCreateMessage } from "./elicitationCreateMessage.js"; import { @@ -270,7 +282,6 @@ import { findHeader, isLongLivedStreamResponse, } from "./fetchTracking.js"; -import { SdkError, SdkErrorCode } from "@modelcontextprotocol/client"; import { annotateRequestTimeout, isRequestTimeoutError, @@ -290,23 +301,6 @@ interface TrackedNotificationStream extends NotificationStreamState { id: string; } -/** Internal record for a receiver task (server polls us for status/result). */ -interface ReceiverTaskRecord { - task: Task; - payloadPromise: Promise<ClientResult>; - resolvePayload: (payload: ClientResult) => void; - rejectPayload: (reason?: unknown) => void; - cleanupTimeoutId?: ReturnType<typeof setTimeout>; - /** - * Aborted when the task reaches a terminal state some way other than the - * user answering — a `tasks/cancel`, or session teardown. Whatever is - * collecting the answer (the native pending-request entry, or an app-rendered - * elicitation and its bridge) is torn down from this, so a cancelled task - * cannot leave a modal on screen waiting for an answer nothing will read. - */ - abort: AbortController; -} - /** * Cap on how many times a single `callTool` will surface URL elicitations and * retry after a `-32042` (UrlElicitationRequired) response. A spec-compliant @@ -350,11 +344,21 @@ function createPendingAbortError(): Error { const TOOL_CALL_CANCELLED_REASON = "Tool call cancelled by user"; /** - * Fallback poll cadence (ms) for {@link InspectorClient.pollTaskToolCall} when a - * task does not advertise its own `pollInterval`. Replaces the cadence the - * removed SDK `experimental.tasks.callToolStream` helper managed internally. + * Combine two optional abort signals into one for {@link RequestOptions}: the + * result aborts when either source does (#1783). Returns the lone signal when + * only one is present and `undefined` when neither is, so a request that needs + * no cancellation keeps carrying no signal. `AbortSignal.any` (Node 20.3+) + * creates a fresh composite per call, GC'd when the request settles. */ -const DEFAULT_TASK_POLL_INTERVAL_MS = 500; +function combineAbortSignals( + a?: AbortSignal, + b?: AbortSignal, +): AbortSignal | undefined { + if (a && b) { + return AbortSignal.any([a, b]); + } + return a ?? b; +} /** * Close a modern listen stream best-effort, absorbing both failure modes a @@ -489,28 +493,35 @@ const MODERN_RECONNECT_BASE_MS = 500; const MODERN_RECONNECT_MAX_MS = 15_000; const MODERN_RECONNECT_MAX_ATTEMPTS = 8; -/** - * InspectorClient wraps an MCP Client and provides: - * - Message tracking and storage - * - Stderr log tracking and storage (for stdio transports) - * - EventTarget interface for React hooks (cross-platform: works in browser and Node.js) - * - Access to client functionality (prompts, resources, tools) - */ export class InspectorClient extends InspectorClientEventTarget { /** - * Upper bound on MRTR (`input_required`) rounds for a single logical request - * before {@link requestWithInputRequired} gives up. We drive the loop - * ourselves (`inputRequired: { autoFulfill: false }`), so this is the manual - * counterpart to the SDK auto-driver's default `maxRounds` (10) and guards - * against a server that keeps returning `input_required` forever. + * We construct the v2 client with auto-fulfilment disabled and drive MRTR + * ourselves, so this mirrors the SDK auto-driver's default round bound. */ private static readonly MRTR_MAX_ROUNDS = 10; private client: Client | null = null; + /** Requester-side Tasks orchestration, attached once per connected SDK client. */ + private taskSession: TaskEnabledSession | null = null; + /** Receiver-side Tasks binding, replaced with each connected session. */ + private taskReceiverBinding: TaskReceiverBinding | null = null; // Lazily-built validator used only on the skipOutputValidation path to detect // (non-fatally) when a delivered result violates the tool's outputSchema. private outputValidator: AjvJsonSchemaValidator | null = null; private transport: Transport | MessageTrackingTransport | null = null; private baseTransport: Transport | null = null; + // The `rawDispatch` ext-tasks requires: below-SDK requests whose responses + // MessageTrackingTransport hands back before the SDK sees them. + private readonly rawWire = new RawWireChannel({ + transport: () => this.transport, + defaultTimeoutMs: () => this.requestTimeout ?? DEFAULT_REQUEST_TIMEOUT_MSEC, + resetTimeoutOnProgress: () => this.resetTimeoutOnProgress, + annotateTimeout: (error, method) => + annotateRequestTimeout(error, method, this.getConnectionDiagnostics()), + }); + private readonly dispatchTaskRequest: RawClientDispatch = ( + request, + options, + ) => this.rawWire.dispatch(request, options); // Every outbound request still awaiting a response, keyed by JSON-RPC id. // Serves two readers: the `markResponseRejected` correlation (#1953), which // needs the method, and the connection diagnostics (#2318), which need the @@ -620,7 +631,7 @@ export class InspectorClient extends InspectorClientEventTarget { private readonly elicitationCapabilityAdvertised: boolean; /** As above, for `capabilities.sampling`. */ private readonly samplingCapabilityAdvertised: boolean; - /** As above, for `capabilities.tasks` (the receiver-side `tasks/*` polls). */ + /** Whether receiver-side Tasks were advertised for at least one input method. */ private readonly tasksCapabilityAdvertised: boolean; /** As above, for `capabilities.elicitation.url` (the URL-mode completion). */ private readonly urlElicitationCapabilityAdvertised: boolean; @@ -676,33 +687,8 @@ export class InspectorClient extends InspectorClientEventTarget { // site to the two failure handlers. Cleared by any user-initiated refresh (a // subscribe/unsubscribe is a fresh attempt, and the server may have changed). private modernNeverAcknowledged = false; - // Task ids the user explicitly cancelled. A cancel makes the in-flight - // `callToolStream` reject with a generic -32603 error, which the stream's - // error path would otherwise report as a *failed* task — flashing "failed" - // in the UI until a refresh fetches the server's true "cancelled" state. - // Recording the id lets that path label the terminal task "cancelled" - // instead, so it lands in the right state immediately (#1455). Cleared on - // disconnect. - private cancelledTaskIds: Set<string> = new Set(); - // Per-task abort controllers for a modern task paused at `input_required`. - // While the poll loop blocks on the pending elicitation (the modal), the tool - // call's own abort path isn't in play — so `cancelRequestorTask` aborts this - // controller to reject the pending request, close the modal, and let the poll - // observe the cancellation. Keyed by taskId; created/removed by the poll loops. - private taskInputAbortControllers = new Map<string, AbortController>(); - // Pending raw-wire requests (modern tasks/* — see rawWireRequest). Keyed by a - // string JSON-RPC id we mint; the SDK Client only mints numeric ids, so ours - // never collide with (or reach) it. Resolved by the transport's - // consume-response hook and rejected on disconnect. - private pendingRawWireRequests = new Map< - string, - { - resolve: (result: unknown) => void; - reject: (err: Error) => void; - timer: ReturnType<typeof setTimeout>; - } - >(); - private rawWireRequestCounter = 0; + /** Which in-flight task calls own each progress token (#2316). */ + private readonly taskProgress = new TaskProgressRouter(); // Abort controller for the in-flight ordinary (non-task) tool call. Aborting // it hands the SDK the MCP cancellation flow for that request and rejects the // pending call, which `callTool` surfaces as a `ToolCallCancelledError`. Which @@ -713,7 +699,17 @@ export class InspectorClient extends InspectorClientEventTarget { // Task-augmented calls have a server-side task and are cancelled via // `cancelRequestorTask` instead, so they don't use this (#1458). private activeToolCallAbortController?: AbortController; - // Receiver tasks (server-initiated: server sends createMessage/elicit with params.task, server polls us) + // An "ambient" abort signal folded into *every* request's options by + // {@link getRequestOptions} while it is set (#1783). Unlike + // `activeToolCallAbortController`, which only cancels tool calls, this covers + // any method — a list, a read, a prompt get. The daemon sets it for the span + // of one serialized command so a caller disconnect cancels whatever request + // is in flight, freeing the per-connection rpc queue instead of wedging it + // behind a request the server never answers. It is opt-in per consumer: web + // and tui never set it, so their behavior is unchanged, and it is per-client + // rather than global, so it never cross-cancels concurrent requests. + private ambientRequestSignal?: AbortSignal; + /** Enable ext-tasks receiver ownership for advertised sampling/elicitation methods. */ private readonly receiverTasks: boolean; // Per-extension advertise overrides (#1738); undefined key falls back to the // registry default in ADVERTISABLE_EXTENSIONS. @@ -750,8 +746,7 @@ export class InspectorClient extends InspectorClientEventTarget { * cannot outlive the connection that asked for it. */ private activeAppElicitations = new Set<AbortController>(); - private receiverTaskTtlMs: number | (() => number); - private receiverTaskRecords: Map<string, ReceiverTaskRecord> = new Map(); + private readonly receiverTaskTtlMs: number | (() => number); // OAuth support (config owned by oauthManager; client delegates and uses !!oauthManager for "is OAuth configured") private oauthManager: OAuthManager | null = null; private logger: InspectorLogger; @@ -985,26 +980,16 @@ export class InspectorClient extends InspectorClientEventTarget { if (this.roots !== undefined) { capabilities.roots = { listChanged: true }; } - // Receiver tasks: advertise so server can send task-augmented createMessage/elicit and poll us if (this.receiverTasks) { - // `requests` declares which server→client requests we accept as tasks, so - // it must name only capabilities we actually advertised — both are decided - // above. Advertising a channel we then answer `-32601` on is the shape - // #1797 is about, and `{ receiverTasks: true, elicit: false }` would do - // exactly that. - const taskRequests: NonNullable< + const requests: NonNullable< NonNullable<ClientCapabilities["tasks"]>["requests"] > = {}; - if (capabilities.sampling) { - taskRequests.sampling = { createMessage: {} }; - } - if (capabilities.elicitation) { - taskRequests.elicitation = { create: {} }; - } + if (capabilities.sampling) requests.sampling = { createMessage: {} }; + if (capabilities.elicitation) requests.elicitation = { create: {} }; capabilities.tasks = { list: {}, cancel: {}, - ...(Object.keys(taskRequests).length > 0 && { requests: taskRequests }), + ...(Object.keys(requests).length > 0 && { requests }), }; } // Assemble the advertised-extensions map from one builder (the single @@ -1064,6 +1049,9 @@ export class InspectorClient extends InspectorClientEventTarget { Object.keys(clientOptions).length > 0 ? clientOptions : undefined, ); this.annotateSdkRequestTimeouts(this.client); + if (this.tasksCapabilityAdvertised) { + this.bindReceiverTasks(); + } } /** @@ -1203,6 +1191,7 @@ export class InspectorClient extends InspectorClientEventTarget { message: JSONRPCNotification, origin: MessageOrigin, ) => { + if (origin === "server") this.dispatchTaskProgress(message); const entry: MessageEntry = { id: crypto.randomUUID(), timestamp: new Date(), @@ -1215,6 +1204,27 @@ export class InspectorClient extends InspectorClientEventTarget { }; } + private dispatchTaskProgress(message: JSONRPCNotification): void { + if (message.method !== "notifications/progress") return; + const params = message.params as Progress & { + progressToken?: ProgressToken; + }; + const progressToken = params.progressToken; + if (progressToken === undefined) return; + // Re-arm any pending raw request's timeout on progress because a + // long-running immediate modern call that keeps reporting progress must + // not time out (same contract as the SDK path's resetTimeoutOnProgress). + this.rawWire.noteProgress(progressToken); + const taskIds = this.taskProgress.route(progressToken); + if (taskIds === undefined) return; + if (this.progress) this.dispatchTypedEvent("progressNotification", params); + for (const taskId of taskIds) + this.dispatchTypedEvent("requestorTaskProgress", { + taskId, + progress: params, + }); + } + private attachTransportListeners(baseTransport: Transport): void { baseTransport.onclose = () => { // An explicit disconnect() owns the teardown and will set the canonical @@ -1224,6 +1234,9 @@ export class InspectorClient extends InspectorClientEventTarget { // "error" would fire `disconnect`, then disconnect()'s own guard would // fire it again (#1490 re-review). if (this.disconnecting) return; + // A handshake-time close belongs to the awaited connect() failure path; + // changing status here races ahead of its catch and briefly reports disconnected. + if (this.status === "connecting") return; // Already fully torn down — nothing to do (avoids a duplicate // `disconnect` event after an explicit disconnect()). if (this.status === "disconnected") return; @@ -1255,6 +1268,8 @@ export class InspectorClient extends InspectorClientEventTarget { // would otherwise wait out its own 30s timeout and blame the timeout for // a crash. Rejecting a settled promise is a no-op and the helper clears // the map, so this can't double-settle with `disconnect()`. + // onclose is synchronous; this best-effort async cleanup owns and logs failures. + void this.closeTaskSessionBestEffort(); this.rejectPendingRawWireRequests("Connection closed"); this.dispatchTypedEvent("disconnect"); }; @@ -1356,9 +1371,16 @@ export class InspectorClient extends InspectorClientEventTarget { // When provided, aborting this signal cancels the request and rejects it // (#1458). The SDK picks the wire signal from the transport: an aborted // per-request SSE stream on a 2026-era Streamable HTTP connection, and - // `notifications/cancelled` everywhere else (#2140). - if (signal) { - opts.signal = signal; + // `notifications/cancelled` everywhere else (#2140). The per-call `signal` + // is combined with any ambient request signal (#1783) so a caller + // disconnect aborts the request even for methods that pass no signal of + // their own. + const effectiveSignal = combineAbortSignals( + signal, + this.ambientRequestSignal, + ); + if (effectiveSignal) { + opts.signal = effectiveSignal; } if (this.progress) { const token = progressToken; @@ -1414,198 +1436,51 @@ export class InspectorClient extends InspectorClientEventTarget { ); } - /** - * True when task status is completed, failed, or cancelled. - * We use this private helper instead of the SDK's experimental isTerminal() - * to avoid depending on experimental API and to get a type predicate so - * TypeScript narrows status to "completed" | "failed" | "cancelled" after the check. - */ - private static isTerminalTaskStatus( - status: Task["status"], - ): status is "completed" | "failed" | "cancelled" { - return ( - status === "completed" || status === "failed" || status === "cancelled" - ); - } - - /** - * Route a receiver (server-initiated) task-augmented `sampling/createMessage` - * or `elicitation/create` response around the v2 Client's result validation. - * - * SDK v2's `Client` wraps every spec request handler (`_wrapHandler`) to - * validate the result it returns — for sampling/elicitation it checks the - * value against `CreateMessageResult` / `ElicitResult` and rejects anything - * else with a `-32602`. The 2025-11-25 task flow answers a task-augmented - * request with a `CreateTaskResult` (`{ task }`), which that validation - * rejects — breaking server-initiated tasks that worked on the legacy client. - * - * There is no public seam to opt a handler out of result validation, so we - * swap the wrapped entry in the Protocol's private `_requestHandlers` map for - * one that dispatches the task-augmented branch straight through the raw - * handler (whose `{ task }` return then rides the legacy codec's pass-through - * `encodeResult` to the wire), while ordinary (non-task) requests keep the - * validating path. Mirrors the bypass a legacy server needs to emit `{ task }`. - * Delete once the SDK models task-augmented results natively (see #1624 stack). - */ - private installReceiverTaskResponseBypass( - method: "sampling/createMessage" | "elicitation/create", - rawHandler: ( - request: CreateMessageRequest & ElicitRequest, - ) => Promise<CreateMessageResult> | Promise<ElicitResult>, - ): void { - if (!this.client) return; - // SDK gap: `Client` exposes no public way to (a) read a registered request - // handler or (b) opt one out of the result validation its `_wrapHandler` - // installs, so we reach the private `_requestHandlers` map through a - // narrowed cast. A public "register a raw/unvalidated handler" API — or a - // handler-result type that includes `CreateTaskResult` — would remove both - // this cast and the ones on the sampling/elicit returns above. - const internal = this.client as unknown as { - _requestHandlers: Map< - string, - (request: unknown, ctx: unknown) => unknown - >; - }; - const validating = internal._requestHandlers.get(method); - if (!validating) return; - internal._requestHandlers.set(method, (request, ctx) => { - const task = (request as { params?: { task?: unknown } })?.params?.task; - // The advertisement check is redundant here — this wrapper only exists - // when tasks are advertised — but it mirrors the handler branch below - // deliberately: the two must agree, so they read one predicate. - if (this.tasksCapabilityAdvertised && task != null) { - return rawHandler(request as CreateMessageRequest & ElicitRequest); - } - return validating(request, ctx); - }); - } - - private createReceiverTask(opts: { - ttl?: number; - initialStatus: Task["status"]; - statusMessage?: string; - pollInterval?: number; - }): ReceiverTaskRecord { - const taskId = crypto.randomUUID(); - const ttlMs = - opts.ttl ?? - (typeof this.receiverTaskTtlMs === "function" - ? this.receiverTaskTtlMs() - : this.receiverTaskTtlMs); - const now = new Date().toISOString(); - const task: Task = { - taskId, - status: opts.initialStatus, - ttl: ttlMs, - createdAt: now, - lastUpdatedAt: now, - ...(opts.pollInterval != null && { pollInterval: opts.pollInterval }), - ...(opts.statusMessage != null && { statusMessage: opts.statusMessage }), - }; - let resolvePayload!: (payload: ClientResult) => void; - let rejectPayload!: (reason?: unknown) => void; - const payloadPromise = new Promise<ClientResult>((resolve, reject) => { - resolvePayload = resolve; - rejectPayload = reject; - }); - // Mark it handled. The real consumer is the server polling `tasks/result` - // (`getReceiverTaskPayload` returns this same promise, so a real awaiter - // still sees the rejection), but nothing has attached a handler while the - // task sits in `input_required` — and it can be rejected from there, by an - // explicit `tasks/cancel` or by teardown settling a queued sample. Without - // this, that reject surfaces as an unhandled rejection. - void payloadPromise.catch(() => {}); - const record: ReceiverTaskRecord = { - task, - payloadPromise, - resolvePayload, - rejectPayload, - abort: new AbortController(), + /** Install package-owned receiver handlers for the current SDK client. */ + private bindReceiverTasks(): void { + this.taskReceiverBinding?.close(); + this.taskReceiverBinding = null; + const client = this.client; + if (!client || !this.tasksCapabilityAdvertised) return; + const methods = { + "sampling/createMessage": this.samplingCapabilityAdvertised, + "elicitation/create": this.elicitationCapabilityAdvertised, }; - record.cleanupTimeoutId = setTimeout(() => { - record.cleanupTimeoutId = undefined; - this.receiverTaskRecords.delete(taskId); - }, ttlMs); - this.receiverTaskRecords.set(taskId, record); - return record; - } - - private emitReceiverTaskStatus(task: Task): void { - if (!this.client) return; - try { - const notification = TaskStatusNotificationSchema.parse({ - method: "notifications/tasks/status" as const, - params: task, - }); - this.client.notification(notification).catch((err) => { - this.logger.warn( - { err, taskId: task.taskId }, - "receiver task status notification failed", + const ttlMs = this.receiverTaskTtlMs; + const binding = bindTaskReceiver(client, { + methods, + ttlMs, + sampling: async (request, context) => { + const result = await this.enqueuePendingSample( + { + method: "sampling/createMessage", + params: paramsWithRelatedTask(request.params, context.taskId), + } as CreateMessageRequest, + "server-request", + context.signal, ); - }); - } catch (err) { - this.logger.warn( - { err, taskId: task.taskId }, - "receiver task status notification failed", - ); - } - } - - private upsertReceiverTask(updatedTask: Task): void { - const record = this.receiverTaskRecords.get(updatedTask.taskId); - if (record) { - record.task = updatedTask; - this.emitReceiverTaskStatus(updatedTask); - } - } - - private getReceiverTask(taskId: string): ReceiverTaskRecord | undefined { - return this.receiverTaskRecords.get(taskId); - } - - private listReceiverTasks(): Task[] { - return Array.from(this.receiverTaskRecords.values()).map((r) => r.task); - } - - private async getReceiverTaskPayload(taskId: string): Promise<ClientResult> { - const record = this.receiverTaskRecords.get(taskId); - if (!record) { - throw new ProtocolError( - ProtocolErrorCode.InvalidParams, - `Unknown taskId: ${taskId}`, - ); - } - return record.payloadPromise; + return toJsonValue(result) as Readonly<Record<string, TasksJsonValue>>; + }, + elicitation: async (request, context) => { + const result = await this.enqueuePendingElicitation( + { + method: "elicitation/create", + params: paramsWithRelatedTask(request.params, context.taskId), + } as ElicitRequest, + "server-request", + context.signal, + ); + return toJsonValue(result) as Readonly<Record<string, TasksJsonValue>>; + }, + onError: (error, context) => + this.logger.error({ error, ...context }, "ext-tasks receiver error"), + }); + this.taskReceiverBinding = binding; } - private cancelReceiverTask(taskId: string): Task { - const record = this.receiverTaskRecords.get(taskId); - if (!record) { - throw new ProtocolError( - ProtocolErrorCode.InvalidParams, - `Unknown taskId: ${taskId}`, - ); - } - if (InspectorClient.isTerminalTaskStatus(record.task.status)) { - return record.task; - } - const now = new Date().toISOString(); - const updatedTask: Task = { - ...record.task, - status: "cancelled", - lastUpdatedAt: now, - }; - record.task = updatedTask; - record.rejectPayload(new Error("Task cancelled")); - // Stop collecting an answer nobody will read: drops the native pending - // entry and tears down an app-rendered elicitation's renderer. - record.abort.abort(); - if (record.cleanupTimeoutId != null) { - clearTimeout(record.cleanupTimeoutId); - record.cleanupTimeoutId = undefined; - } - this.emitReceiverTaskStatus(updatedTask); - return updatedTask; + private closeTaskReceiver(): void { + this.taskReceiverBinding?.close(); + this.taskReceiverBinding = null; } /** @@ -1629,281 +1504,28 @@ export class InspectorClient extends InspectorClientEventTarget { this.transportHasAuthProvider = false; } - /** - * Register the handlers for requests the *server* makes of *us* — - * `roots/list`, `sampling/createMessage`, `elicitation/create`, and the - * receiver-side `tasks/*` polls. - * - * MUST be called before `client.connect()`. The matching capabilities are - * advertised on the `Client` at construction time, so from the moment - * `connect()` sends `notifications/initialized` the server is entitled to - * issue any of these requests. Registering afterwards leaves a window in - * which the SDK `Client` has no handler and answers `-32601 Method not - * found` — which is exactly what a server that asks for roots the instant it - * is initialized (e.g. `server-filesystem`, which learns its allowed - * directories that way) hits, while a server that asks later does not (#1797). - * - * Nothing here depends on the server's capabilities — only on constructor-set - * state — so there is nothing to wait for. Its sibling - * {@link registerPeerNotificationHandlers} does the same for the one - * notification handler in that position; the notification handlers that *do* - * gate on `this.capabilities` stay in `connect()`, after the handshake. - */ + /** Register ordinary server→client request handlers before the handshake. */ private registerPeerRequestHandlers(): void { - // Gated on what was advertised, like the others — see - // `rootsCapabilityAdvertised`. - if (this.samplingCapabilityAdvertised && this.client) { - const samplingHandler = ( - request: CreateMessageRequest, - ): Promise<CreateMessageResult> => { - const paramsTask = (request.params as { task?: { ttl?: number } }) - ?.task; - if (this.tasksCapabilityAdvertised && paramsTask != null) { - const record = this.createReceiverTask({ - ttl: paramsTask.ttl, - initialStatus: "input_required", - statusMessage: "Awaiting user input", - }); - void (async () => { - const samplingRequest = new SamplingCreateMessage( - request, - (result) => { - record.resolvePayload(result); - const now = new Date().toISOString(); - const updated: Task = { - ...record.task, - status: "completed", - lastUpdatedAt: now, - }; - record.task = updated; - this.upsertReceiverTask(updated); - }, - (error) => { - record.rejectPayload(error); - const now = new Date().toISOString(); - const updated: Task = { - ...record.task, - status: "failed", - lastUpdatedAt: now, - statusMessage: - error instanceof Error ? error.message : String(error), - }; - record.task = updated; - this.upsertReceiverTask(updated); - }, - (id) => this.removePendingSample(id), - ); - this.addPendingSample(samplingRequest); - })(); - // Task-augmented (2025-11-25) response: the server sent a - // task-augmented `sampling/createMessage`, so we reply with a - // `CreateTaskResult` (`{ task }`) rather than a `CreateMessageResult`. - // The v2 Client validates a spec handler's result and would reject - // `{ task }` with -32602; `installReceiverTaskResponseBypass` below - // routes this task-augmented branch around that validation so the - // legacy `{ task }` response reaches the wire. `taskResult` is typed - // as `CreateTaskResult` so its shape IS checked; the unavoidable - // `as unknown as CreateMessageResult` bridges the SDK gap — the 2-arg - // `setRequestHandler` overload types a sampling handler's return as - // `CreateMessageResult` only and doesn't model the (deprecated but - // wire-valid) task-augmented `CreateTaskResult`. A handler-result - // union `CreateMessageResult | CreateTaskResult` on the SDK side - // would remove this cast. - const taskResult: CreateTaskResult = { task: record.task }; - return Promise.resolve(taskResult as unknown as CreateMessageResult); - } - return this.enqueuePendingSample(request, "server-request"); - }; - this.client.setRequestHandler("sampling/createMessage", samplingHandler); - // Registration, like the `setRequestHandler` above it — and the whole - // bypass mechanism (install, wrapper branch, handler branch) reads this - // one predicate, so the install can't drift from the branch it controls. - if (this.tasksCapabilityAdvertised) { - this.installReceiverTaskResponseBypass( - "sampling/createMessage", - samplingHandler, - ); - } + if (!this.client) return; + if (this.samplingCapabilityAdvertised) { + this.client.setRequestHandler("sampling/createMessage", (request) => + this.enqueuePendingSample(request, "server-request"), + ); } - - // Gated on what was advertised, not on `this.elicit` — see the field's doc: - // an elicit option that enables no mode advertises nothing, and registering - // regardless throws before the handshake. - if (this.elicitationCapabilityAdvertised && this.client) { - const elicitHandler = ( - request: ElicitRequest, - // Structural, and only the one field this needs: the SDK's - // `ClientContext` carries much more, and naming it here would tie the - // handler to a type the bypass helper below does not thread through. - ctx?: { mcpReq?: { signal?: AbortSignal } }, - ): Promise<ElicitResult> => { - const paramsTask = (request.params as { task?: { ttl?: number } }) - ?.task; - if (this.tasksCapabilityAdvertised && paramsTask != null) { - const record = this.createReceiverTask({ - ttl: paramsTask.ttl, - initialStatus: "input_required", - statusMessage: "Awaiting user input", - }); - // Settling the receiver task, shared by both answer routes below so - // an app-rendered answer completes the task exactly as a native one - // does. - const completeTask = (result: ElicitResult) => { - // A cancelled (or otherwise terminal) task must not be re-settled: - // an answer that arrives after `tasks/cancel` would otherwise - // overwrite `cancelled` with `completed`. - if (InspectorClient.isTerminalTaskStatus(record.task.status)) - return; - record.resolvePayload(result); - const updated: Task = { - ...record.task, - status: "completed", - lastUpdatedAt: new Date().toISOString(), - }; - record.task = updated; - this.upsertReceiverTask(updated); - }; - const failTask = (error: Error) => { - if (InspectorClient.isTerminalTaskStatus(record.task.status)) - return; - record.rejectPayload(error); - const updated: Task = { - ...record.task, - status: "failed", - lastUpdatedAt: new Date().toISOString(), - statusMessage: error.message, - }; - record.task = updated; - this.upsertReceiverTask(updated); - }; - void (async () => { - // A task-augmented request is still an `elicitation/create`, so the - // app-rendering contract applies to it too (#1854). It cannot go - // through `enqueuePendingElicitation` — the response frame has - // already been sent as a `CreateTaskResult` and the answer settles - // the TASK rather than the request — so the same attempt is made - // here, falling back to the native queue exactly as that funnel - // does. An abort (disconnect) fails the task rather than reopening - // it natively. - let appResult: ElicitResult | null; - try { - appResult = await this.tryAppElicitation( - request, - record.abort.signal, - ); - } catch (error) { - failTask( - error instanceof Error ? error : new Error(String(error)), - ); - return; - } - if (appResult) { - completeTask(appResult); - return; - } - const elicitationRequest = new ElicitationCreateMessage( - request, - completeTask, - (id) => this.removePendingElicitation(id), - failTask, - ); - this.addPendingElicitation(elicitationRequest); - // A `tasks/cancel` (or teardown) drops the queued entry, so the - // modal does not outlive the task it belongs to. - this.wirePendingAbort(record.abort.signal, () => - this.removePendingElicitation(elicitationRequest.id), - ); - })(); - // Task-augmented (2025-11-25) response — see the sampling handler - // above. Reply with a `CreateTaskResult` (`{ task }`), routed around - // the v2 Client's result validation by - // `installReceiverTaskResponseBypass` below. `taskResult` is typed so - // its shape is checked; the `as unknown as ElicitResult` bridges the - // same SDK gap as the sampling handler — the 2-arg `setRequestHandler` - // overload types an elicitation handler's return as `ElicitResult` - // only and doesn't model the task-augmented `CreateTaskResult`. - const taskResult: CreateTaskResult = { task: record.task }; - return Promise.resolve(taskResult as unknown as ElicitResult); - } - // `ctx.mcpReq.signal` aborts when the server cancels this request - // (`notifications/cancelled`). Threading it through means both answer - // surfaces — the native queue entry and an app-rendered elicitation's - // renderer — are torn down with the request, instead of a modal - // outliving work the server abandoned. The task-augmented branch above - // deliberately does NOT use it: that request is answered immediately - // with a `CreateTaskResult`, so its lifetime is the task's, which - // carries its own abort (see `ReceiverTaskRecord.abort`). - return this.enqueuePendingElicitation( + if (this.elicitationCapabilityAdvertised) { + this.client.setRequestHandler("elicitation/create", (request, context) => + this.enqueuePendingElicitation( request, "server-request", - ctx?.mcpReq?.signal, - ); - }; - this.client.setRequestHandler("elicitation/create", elicitHandler); - // Registration, like the `setRequestHandler` above it — and the whole - // bypass mechanism (install, wrapper branch, handler branch) reads this - // one predicate, so the install can't drift from the branch it controls. - if (this.tasksCapabilityAdvertised) { - this.installReceiverTaskResponseBypass( - "elicitation/create", - elicitHandler, - ); - } - } - - // Gated on what was advertised at construction, and it has to be: the SDK - // asserts the matching client capability inside `setRequestHandler`, so - // registering this on a client built without `roots` throws "Client does - // not support roots capability". Since `capabilities.roots` is negotiated at - // `initialize` (set in the constructor) and `registerCapabilities` refuses - // to run after connect, a client that omits the option can never serve - // `roots/list` — which is why every client that may call `setRoots()` later - // must pass `roots` up front (web does; the CLI and TUI now do too — #1797). - if (this.rootsCapabilityAdvertised && this.client) { - this.client.setRequestHandler("roots/list", async () => { - return { roots: this.roots ?? [] }; - }); - } - - // Set up receiver-task request handlers (server polls us for tasks/list, - // tasks/get, tasks/result, tasks/cancel). SDK v2 removed tasks from the - // spec-method set, so these register through the 3-arg custom form with an - // explicit params schema (from the deprecated-but-importable task request - // schemas). The `result` schema is intentionally omitted so the SDK does - // not validate our responder return — matching v1, where only the - // requester validated (our receiver `Task` may omit fields a strict result - // schema would require). - if (this.tasksCapabilityAdvertised && this.client) { - this.client.setRequestHandler( - "tasks/list", - { params: ListTasksRequestSchema.shape.params }, - async () => ({ tasks: this.listReceiverTasks() }), - ); - this.client.setRequestHandler( - "tasks/get", - { params: GetTaskRequestSchema.shape.params }, - async (params) => { - const record = this.getReceiverTask(params.taskId); - if (!record) { - throw new ProtocolError( - ProtocolErrorCode.InvalidParams, - `Unknown taskId: ${params.taskId}`, - ); - } - return record.task; - }, - ); - this.client.setRequestHandler( - "tasks/result", - { params: GetTaskPayloadRequestSchema.shape.params }, - async (params) => this.getReceiverTaskPayload(params.taskId), - ); - this.client.setRequestHandler( - "tasks/cancel", - { params: CancelTaskRequestSchema.shape.params }, - async (params) => this.cancelReceiverTask(params.taskId), + context.mcpReq?.signal, + ), ); } + if (this.rootsCapabilityAdvertised) { + this.client.setRequestHandler("roots/list", async () => ({ + roots: this.roots ?? [], + })); + } } /** @@ -1940,29 +1562,6 @@ export class InspectorClient extends InspectorClientEventTarget { ); } - /** - * Stop the receiver tasks' TTL timers and drop the records. - * - * These are tasks a *server* created with us, so they belong to the session - * that created them: `listReceiverTasks()` is what the `tasks/list` handler - * answers with, and a record surviving into the next session would report a - * task the new server never created. `disconnect()` clears them, and so does - * `connect()` — the auth-recovery retry reconnects the *same* client - * instance, so ending the session isn't the only way a new one begins - * (#1797). - */ - private clearReceiverTasks(): void { - for (const record of this.receiverTaskRecords.values()) { - if (record.cleanupTimeoutId != null) { - clearTimeout(record.cleanupTimeoutId); - } - // Same reason as `cancelReceiverTask`: the session that owns whatever is - // collecting the answer is ending. - record.abort.abort(); - } - this.receiverTaskRecords.clear(); - } - /** * Reset the modern listen-stream cluster: the subscribed set, the stream * state derived from it, and the reconnect machinery that reports on it. @@ -2023,22 +1622,12 @@ export class InspectorClient extends InspectorClientEventTarget { * this same instance (the auth-recovery path), both leave it behind. Called * start-clean from `connect()` so every route in is covered. * - * Each member has a symptom, not just untidiness: a stale `subscribedResources` - * entry makes the modern `subscribeToResource` early-return, so the user's - * Subscribe click silently sends nothing to the new server; a stale - * `cancelledTaskIds` entry mislabels a *new* task sharing the id as - * `cancelled` rather than `failed`; a stale subscription stream state reads - * `active` for a set that is now empty, which every reader of it treats as - * impossible; a receiver-task record is reported to the new server by - * `tasks/list`; and an un-aborted `taskInputAbortControllers` - * entry delays a paused poll loop unwinding — both registration sites release - * in a `finally`, so nothing leaks permanently; the abort just closes the - * window between the crash and the unwind (#1797). + * The task receiver binding is session-owned too: closing it aborts pending + * callbacks, drops package-owned records, and restores prior request handlers. */ private resetSessionState(): void { - this.clearReceiverTasks(); + this.closeTaskReceiver(); this.resetSubscriptionStream(); - this.cancelledTaskIds.clear(); // Correlation data is per-session: JSON-RPC ids don't survive it, and // MessageLogState drops its entries on disconnect, so anything left here // could only point at an entry that no longer exists. Clearing also @@ -2050,6 +1639,7 @@ export class InspectorClient extends InspectorClientEventTarget { // (#1953). this.outstandingRequests.clear(); this.lastAnsweredRequestByMethod.clear(); + this.taskProgress.clear(); this.lastResponse = undefined; this.notificationStream = undefined; this.dispatchConnectionDiagnosticsChange(); @@ -2070,10 +1660,6 @@ export class InspectorClient extends InspectorClientEventTarget { this.excludedTools = []; this.dispatchTypedEvent("excludedToolsChange", []); } - for (const [, controller] of this.taskInputAbortControllers) { - controller.abort(new Error("Connection ended")); - } - this.taskInputAbortControllers.clear(); // Restore the configured opt-in rather than carrying a mid-session // `setModernLogLevel` override into the next connection — and rather than // leaving it `undefined` after a `disconnect()` cleared it, which silently @@ -2160,195 +1746,166 @@ export class InspectorClient extends InspectorClientEventTarget { if (this.status === "connected") { return; } + this.status = "connecting"; + this.dispatchTypedEvent("statusChange", this.status); + try { + // Start from a clean session — see `resetSessionState` for why this is + // start-clean rather than relying on `disconnect()`. + this.resetSessionState(); + // Settle UI requests from any previous session before installing fresh + // receiver handlers. Binding close in resetSessionState aborts receiver + // callbacks; this sweep also covers ordinary peer requests. + this.clearAndAnnouncePendingPeerRequests(); + this.rejectPendingRawWireRequests("Connection ended"); + await this.closeTaskSessionBestEffort(); - // Start from a clean session — see `resetSessionState` for why this is - // start-clean rather than relying on `disconnect()`. - this.resetSessionState(); - // The two collections `resetSessionState` excludes as "settled on the way - // out", swept here as well — because one route out settles nothing. An - // `onerror` without an `onclose` only flips status to `"error"`: it runs - // neither teardown path, and it leaves `baseTransport` cached, so a - // `connect()` on this same instance reuses a *live* transport. That is the - // route the subscription-stream close exists for, and it strands these two - // the same way. The peer queue is the sharper of them — the web - // pending-request modal is derived from its length with no status gate, so - // it outlives the session, and a user answering it later would write - // *their* answer for the previous session's request id onto the new - // connection, arbitrarily far past the re-handshake. Note what the sweep - // does instead is emit a *cancel* for that same id, right here: still the - // settle-don't-discard rule, and this is the earliest moment available: - // the old connection is still the one on the wire here, and stays so at - // least until the conditional `dropCachedTransport()` below — which on a - // stdio server never runs at all, so the same transport carries straight - // through the re-handshake. - // - // Both helpers are idempotent (one guards on a non-empty queue, the other - // clears its map and re-rejecting a settled promise is a no-op), so these - // are no-ops on the routes that already ran them; and anything still - // pending here belongs to a session that is, by definition, no longer - // connected. - // - // Must stay *after* `resetSessionState()`, which reads as independent of it - // but is not: cancelling a task-augmented peer request settles it - // synchronously into the record callback, which ends in - // `upsertReceiverTask`. That is a no-op only because `clearReceiverTasks()` - // just emptied the map — hoisted above the reset, it would instead emit a - // `notifications/tasks/status` for the outgoing session's task, onto the - // transport this connect is about to reuse, moments before the reset drops - // the record anyway. - this.clearAndAnnouncePendingPeerRequests(); - this.rejectPendingRawWireRequests("Connection ended"); - - const oauthManager = this.oauthManager; - if ( - this.baseTransport && - this.isHttpOAuthConfig() && - oauthManager && - !this.transportHasAuthProvider && - !oauthManager.isEnterpriseManaged() && - (await oauthManager.isOAuthAuthorized()) - ) { - await this.dropCachedTransport(); - } - - // Create transport (single place for create / wrap / attach). - if (!this.baseTransport) { - const transportOptions: CreateTransportOptions = { - fetchFn: this.fetchFn, - pipeStderr: this.pipeStderr, - onStderr: (entry: StderrLogEntry) => { - this.dispatchStderrLog(entry); - }, - onFetchRequest: (entry: FetchRequestEntryBase) => { - this.dispatchFetchRequest({ ...entry, category: "transport" }); - this.noteTransportStream(entry); - }, - onFetchResponseBody: (id: string, body: string) => { - this.dispatchFetchRequestBodyUpdate(id, body); - }, - onFetchStreamUpdate: (id: string, stream: FetchStreamState) => { - this.dispatchFetchRequestStreamUpdate(id, stream); - }, - ...(this.serverSettings && { settings: this.serverSettings }), - }; - if (this.isHttpOAuthConfig() && oauthManager) { - // Record every 401/403 the transport sees, whatever else happens to - // it. The legacy first-authorization path deliberately runs with no - // authProvider and no challenge interception (see below), so the SDK - // raises a headerless `UnauthorizedError` and the client calls - // `authenticate()` with nothing in hand — this is the only place the - // challenge's RFC 9728 `resource_metadata` still exists (#2071). - const manager = oauthManager; - transportOptions.onAuthChallengeObserved = (challenge) => { - manager.noteObservedAuthChallenge(challenge); - }; - if (oauthManager.isEnterpriseManaged()) { - await oauthManager.trySilentEnterpriseManagedAuth(); - const provider = await oauthManager.createOAuthProviderForTransport(); - const tokens = await provider.tokens(); - if (!tokens?.access_token) { - const err = new Error( - "Unauthorized: EMA resource access token unavailable", - ) as Error & { status?: number; code?: number }; - err.status = 401; - err.code = 401; - throw err; - } - transportOptions.authProvider = provider; - } else if (await oauthManager.isOAuthAuthorized()) { - // Without stored tokens, omit authProvider so connect() surfaces a plain - // 401 instead of the SDK opening a browser before the app callback - // server is listening (TUI/CLI run authenticate() explicitly). - transportOptions.authProvider = - await oauthManager.createOAuthProviderForTransport(); - } - } + const oauthManager = this.oauthManager; if ( - this.directAuthRecovery && - this.directAuthRecoveryActive !== false && + this.baseTransport && this.isHttpOAuthConfig() && oauthManager && - // No stored tokens means no authProvider (see above), and then a 401 on - // the era-negotiation probe reaches the SDK as a raw `SdkHttpError`. - // The probe's classifier ignores the HTTP status — it only looks for a - // JSON-RPC error body — so it verdicts "not a modern server", and pin - // ("modern") mode rethrows that as ERA_NEGOTIATION_FAILED with the 401 - // discarded entirely: no status, not even a cause. Intercepting makes - // the 401 a typed AuthChallengeError, which survives the probe as - // `data.cause` for `findNestedAuthError` to recover (#1805). - // - // WORKAROUND (#1807, upstream modelcontextprotocol/typescript-sdk#2561): - // remove this clause once the SDK classifies a probe 401/403 as - // auth-required. `findNestedAuthError` is the permanent fix; the - // `|| this.probesProtocolEra()` clause below exists only to compensate - // for that upstream gap and should be deleted with it. - // - // Known, accepted side effect of turning intercept on with no stored - // tokens: `parseAuthChallengeFromResponse` treats 403 as a challenge - // too, so a probe answered 403 for a *non-auth* reason (a gateway - // rejecting the unknown `server/discover` method, say) now starts OAuth - // discovery instead of letting "auto" fall back to the legacy - // `initialize`. The outcome is a surfaced `oauthError`, not a hang, and - // it goes away with this clause. - (transportOptions.authProvider || this.probesProtocolEra()) + !this.transportHasAuthProvider && + !oauthManager.isEnterpriseManaged() && + (await oauthManager.isOAuthAuthorized()) ) { - transportOptions.interceptAuthChallenges = true; + await this.dropCachedTransport(); } - this.transportHasAuthProvider = !!transportOptions.authProvider; - const { transport: baseTransport } = this.transportClientFactory( - this.transportConfig, - transportOptions, - ); - this.baseTransport = baseTransport; - // What the factory was handed, not the live value: `transportOptions` - // was built before the OAuth awaits above, and a settings save landing - // during them would otherwise be reported as sent when it was not. - this.transportSettings = transportOptions.settings; - if (this.directAuthRecovery) { - this.directAuthRecoveryActive = !( - baseTransport instanceof RemoteClientTransport + + // Create transport (single place for create / wrap / attach). + if (!this.baseTransport) { + const transportOptions: CreateTransportOptions = { + fetchFn: this.fetchFn, + pipeStderr: this.pipeStderr, + onStderr: (entry: StderrLogEntry) => { + this.dispatchStderrLog(entry); + }, + onFetchRequest: (entry: FetchRequestEntryBase) => { + this.dispatchFetchRequest({ ...entry, category: "transport" }); + this.noteTransportStream(entry); + }, + onFetchResponseBody: (id: string, body: string) => { + this.dispatchFetchRequestBodyUpdate(id, body); + }, + onFetchStreamUpdate: (id: string, stream: FetchStreamState) => { + this.dispatchFetchRequestStreamUpdate(id, stream); + }, + ...(this.serverSettings && { settings: this.serverSettings }), + }; + if (this.isHttpOAuthConfig() && oauthManager) { + // Record every 401/403 the transport sees, whatever else happens to + // it. The legacy first-authorization path deliberately runs with no + // authProvider and no challenge interception (see below), so the SDK + // raises a headerless `UnauthorizedError` and the client calls + // `authenticate()` with nothing in hand — this is the only place the + // challenge's RFC 9728 `resource_metadata` still exists (#2071). + const manager = oauthManager; + transportOptions.onAuthChallengeObserved = (challenge) => { + manager.noteObservedAuthChallenge(challenge); + }; + if (oauthManager.isEnterpriseManaged()) { + await oauthManager.trySilentEnterpriseManagedAuth(); + const provider = + await oauthManager.createOAuthProviderForTransport(); + const tokens = await provider.tokens(); + if (!tokens?.access_token) { + const err = new Error( + "Unauthorized: EMA resource access token unavailable", + ) as Error & { status?: number; code?: number }; + err.status = 401; + err.code = 401; + throw err; + } + transportOptions.authProvider = provider; + } else if (await oauthManager.isOAuthAuthorized()) { + // Without stored tokens, omit authProvider so connect() surfaces a plain + // 401 instead of the SDK opening a browser before the app callback + // server is listening (TUI/CLI run authenticate() explicitly). + transportOptions.authProvider = + await oauthManager.createOAuthProviderForTransport(); + } + } + if ( + this.directAuthRecovery && + this.directAuthRecoveryActive !== false && + this.isHttpOAuthConfig() && + oauthManager && + // No stored tokens means no authProvider (see above), and then a 401 on + // the era-negotiation probe reaches the SDK as a raw `SdkHttpError`. + // The probe's classifier ignores the HTTP status — it only looks for a + // JSON-RPC error body — so it verdicts "not a modern server", and pin + // ("modern") mode rethrows that as ERA_NEGOTIATION_FAILED with the 401 + // discarded entirely: no status, not even a cause. Intercepting makes + // the 401 a typed AuthChallengeError, which survives the probe as + // `data.cause` for `findNestedAuthError` to recover (#1805). + // + // WORKAROUND (#1807, upstream modelcontextprotocol/typescript-sdk#2561): + // remove this clause once the SDK classifies a probe 401/403 as + // auth-required. `findNestedAuthError` is the permanent fix; the + // `|| this.probesProtocolEra()` clause below exists only to compensate + // for that upstream gap and should be deleted with it. + // + // Known, accepted side effect of turning intercept on with no stored + // tokens: `parseAuthChallengeFromResponse` treats 403 as a challenge + // too, so a probe answered 403 for a *non-auth* reason (a gateway + // rejecting the unknown `server/discover` method, say) now starts OAuth + // discovery instead of letting "auto" fall back to the legacy + // `initialize`. The outcome is a surfaced `oauthError`, not a hang, and + // it goes away with this clause. + (transportOptions.authProvider || this.probesProtocolEra()) + ) { + transportOptions.interceptAuthChallenges = true; + } + this.transportHasAuthProvider = !!transportOptions.authProvider; + const { transport: baseTransport } = this.transportClientFactory( + this.transportConfig, + transportOptions, ); + this.baseTransport = baseTransport; + // What the factory was handed, not the live value: `transportOptions` + // was built before the OAuth awaits above, and a settings save landing + // during them would otherwise be reported as sent when it was not. + this.transportSettings = transportOptions.settings; + if (this.directAuthRecovery) { + this.directAuthRecoveryActive = !( + baseTransport instanceof RemoteClientTransport + ); + } + if ( + baseTransport instanceof RemoteClientTransport && + oauthManager && + this.isHttpOAuthConfig() + ) { + baseTransport.setAuthRecovery({ + handleAuthChallenge: (challenge, options) => + oauthManager.handleAuthChallenge(challenge, options), + pushAuthState: () => this.pushRemoteAuthState(), + }); + baseTransport.setOnAuthChallenge((challenge) => { + void this.handleAmbientAuthChallenge(challenge); + }); + } + const messageTracking = this.createMessageTrackingCallbacks(); + this.transport = new MessageTrackingTransport( + baseTransport, + messageTracking, + { + rawRequestChannel: { + consume: (message) => this.consumeRawWireResponse(message), + }, + }, + ); + this.attachTransportListeners(this.baseTransport); } - if ( - baseTransport instanceof RemoteClientTransport && - oauthManager && - this.isHttpOAuthConfig() - ) { - baseTransport.setAuthRecovery({ - handleAuthChallenge: (challenge, options) => - oauthManager.handleAuthChallenge(challenge, options), - pushAuthState: () => this.pushRemoteAuthState(), - }); - baseTransport.setOnAuthChallenge((challenge) => { - void this.handleAmbientAuthChallenge(challenge); - }); - } - const messageTracking = this.createMessageTrackingCallbacks(); - this.transport = new MessageTrackingTransport( - baseTransport, - messageTracking, - { - rewriteIncomingResult: (message) => - this.rewriteModernTaskResult(message), - consumeIncomingResponse: (message) => - this.consumeRawWireResponse(message), - }, - ); - this.attachTransportListeners(this.baseTransport); - } - - if (!this.transport) { - throw new Error("Transport not initialized"); - } - try { - this.status = "connecting"; - this.dispatchTypedEvent("statusChange", this.status); + if (!this.transport) { + throw new Error("Transport not initialized"); + } // Register the handlers for server→client requests and the // capability-independent notifications before the handshake — see // `registerPeerRequestHandlers` for why the ordering is load-bearing. this.registerPeerRequestHandlers(); + this.bindReceiverTasks(); this.registerPeerNotificationHandlers(); // Connect-time timeout from per-server settings, defaulting to @@ -2452,6 +2009,7 @@ export class InspectorClient extends InspectorClientEventTarget { // #1395). If "connect" fired first, that gate would read undefined // capabilities and wipe tools/prompts/resources to empty on every connect. await this.fetchServerInfo(); + await this.attachTaskSession(); // Set initial logging level if configured and server supports it. // @@ -2621,6 +2179,7 @@ export class InspectorClient extends InspectorClientEventTarget { this.status = "error"; this.dispatchTypedEvent("statusChange", this.status); } + await this.closeTaskSessionBestEffort(); if (this.baseTransport && !this.transportHasAuthProvider) { await this.dropCachedTransport(); } @@ -2682,6 +2241,7 @@ export class InspectorClient extends InspectorClientEventTarget { await new Promise((r) => setTimeout(r, 10)); } } + await this.closeTaskSessionBestEffort(); try { await this.client.close(); } catch { @@ -2718,7 +2278,6 @@ export class InspectorClient extends InspectorClientEventTarget { // stream (best-effort — the transport is already going away) and bump the // generation so any in-flight re-listen/reconnect bails (#1630). this.resetSubscriptionStream(); - this.cancelledTaskIds.clear(); // Settle any pending raw-wire (modern tasks/*) requests so their callers // don't hang past teardown. Rejected outright on every disconnect: the // drain above polls the SDK's own response-handler map, which never holds @@ -2726,16 +2285,11 @@ export class InspectorClient extends InspectorClientEventTarget { // opt-in anyway — every production caller leaves `safeDisconnectTimeout` at // 0, so nothing is drained for anyone. this.rejectPendingRawWireRequests("Disconnected"); - // Abort any task paused at input_required so its poll loop unwinds. - for (const [, controller] of this.taskInputAbortControllers) { - controller.abort(new Error("Disconnected")); - } - this.taskInputAbortControllers.clear(); // Abort any in-flight ordinary tool call so its promise settles instead of // hanging past teardown; drop the controller reference either way. this.activeToolCallAbortController?.abort("Disconnected"); this.activeToolCallAbortController = undefined; - this.clearReceiverTasks(); + this.closeTaskReceiver(); this.capabilities = undefined; this.serverInfo = undefined; this.instructions = undefined; @@ -2831,6 +2385,10 @@ export class InspectorClient extends InspectorClientEventTarget { }; } + /** Authoritative generation-neutral capabilities for requester task behavior. */ + getTaskSessionCapabilities(): TaskCapabilities | undefined { + return this.taskSession?.capabilities; + } /** * True when the connection is modern (2026-07-28) AND the server advertised * the `io.modelcontextprotocol/tasks` extension (SEP-2663) in its @@ -2841,35 +2399,7 @@ export class InspectorClient extends InspectorClientEventTarget { * task store gate on the extension rather than the legacy capability. */ isTasksExtensionNegotiated(): boolean { - return ( - this.isModernEra() && - this.capabilities?.extensions?.[TASKS_EXTENSION_KEY] !== undefined - ); - } - - /** - * Build the full modern (2026-07-28) per-request envelope for a RAW tasks/* - * request. The SDK's codec normally stamps this envelope, but raw requests - * bypass the codec, and the modern server rejects a request whose - * `MCP-Protocol-Version` header names 2026-07-28 but omits the required - * envelope `_meta` keys (`protocolVersion`, `clientInfo`, plus - * `clientCapabilities` carrying the tasks extension). We reproduce it here. - */ - private withModernTaskEnvelope( - params: Record<string, unknown>, - ): Record<string, unknown> { - const clientCapabilities = { - ...this.clientCapabilities, - // Force-stamp the tasks extension regardless of what the client - // advertised at construction: the raw `tasks/*` channel requires it, and - // a user may disable general tasks advertisement via `advertisedExtensions` - // (#1738). So this stamp is load-bearing, not a redundant re-add. - extensions: { - ...this.clientCapabilities.extensions, - [TASKS_EXTENSION_KEY]: {}, - }, - }; - return this.withModernEnvelope(params, clientCapabilities); + return isTasksExtensionNegotiated(this.protocolEra, this.capabilities); } /** @@ -2911,247 +2441,77 @@ export class InspectorClient extends InspectorClientEventTarget { }; } - /** - * Transport-level rewrite of a modern (SEP-2663) `CreateTaskResult` - * (`resultType: "task"`) — the one task frame the SDK v2 codec rejects (tasks - * were removed, so the codec knows only `complete`/`input_required`). The true - * frame is already logged by `trackResponse`; here we hand the SDK a benign - * `CallToolResult` that carries the real `DetailedTask` under - * {@link MODERN_TASK_HANDLE_META}, where {@link pollTaskToolCall} reads it to - * drive the poll. Any other message passes through untouched. - */ - private rewriteModernTaskResult( - message: JSONRPCResultResponse, - ): JSONRPCMessage { - if (!isModernCreateTaskResult(message.result)) { - return message; - } - const task = message.result as ModernDetailedTask; - return { - ...message, - result: { - resultType: "complete", - content: [{ type: "text", text: `Modern task ${task.taskId} created` }], - _meta: { [MODERN_TASK_HANDLE_META]: task }, - }, - }; - } - - /** - * Send an extension method the SDK v2 era gate refuses to route — the modern - * `tasks/get` / `tasks/update` / `tasks/cancel`, which are spec-method names - * absent from the 2026-07-28 era, so `client.request` throws - * `MethodNotSupportedByProtocolVersion` before anything reaches the wire. - * - * We mint a string JSON-RPC id (the SDK only mints numeric ids, so ours never - * collide), send the raw frame straight through the transport (which still - * logs it via `trackRequest`, so the Protocol/Network tabs see it), and await - * the matching response — captured and consumed by the transport's - * consume-response hook so it never confuses the SDK Client. The response is - * validated with the caller's explicit schema. - */ private async rawWireRequest<T>( method: string, params: Record<string, unknown>, resultSchema: { parse: (value: unknown) => T }, + options: { + readonly signal?: AbortSignal; + readonly timeoutMs?: number; + } = {}, ): Promise<T> { - const transport = this.transport; - if (!transport) { - throw new Error("Client is not connected"); + let response: JsonRpcResponse; + try { + response = await this.rawWire.dispatch( + toJsonValue({ method, params }), + { signal: options.signal }, + options.timeoutMs, + ); + } catch (error) { + // This caller is the Inspector's own, not ext-tasks, so nothing above it + // unwraps the port-level DispatchError: hand back the annotated timeout + // (#2318) or the transport's own failure, as an SDK request would. + throw unwrapTaskDispatchError(error); } - const id = `inspector-ext-${(this.rawWireRequestCounter += 1)}`; - // `params` is an arbitrary caller-supplied record; the SDK types request - // params with a specific optional `_meta` shape it can't satisfy, so widen - // it with a single structural cast. Typing `message` as `JSONRPCRequest` - // (a `JSONRPCMessage` member) then needs no further cast. - const message: JSONRPCRequest = { - jsonrpc: "2.0", - id, - method, - params: params as JSONRPCRequest["params"], - }; - const timeoutMs = this.requestTimeout ?? 30_000; - const raw = await new Promise<unknown>((resolve, reject) => { - const timer = setTimeout(() => { - this.pendingRawWireRequests.delete(id); - // The same error, with the same annotation, as an SDK request that - // times out: this path bypasses `Protocol.request`, so it builds the - // SDK's own timeout shape and runs it through the decorator's - // annotation by hand — a raw-wire caller sees one kind of timeout, - // not two (#2318). - reject( - annotateRequestTimeout( - new SdkError(SdkErrorCode.RequestTimeout, "Request timed out", { - timeout: timeoutMs, - }), - method, - this.getConnectionDiagnostics(), - ), - ); - }, timeoutMs); - this.pendingRawWireRequests.set(id, { resolve, reject, timer }); - transport.send(message).catch((err: unknown) => { - const pending = this.pendingRawWireRequests.get(id); - if (pending) { - clearTimeout(pending.timer); - this.pendingRawWireRequests.delete(id); - } - // The browser's remote transport awaits the response inside `send`, - // so its relay wait can expire here first, as the SDK's timeout - // shape; annotate it exactly as the local timer above does (a - // non-timeout error passes through untouched). - reject( - annotateRequestTimeout( - err instanceof Error ? err : new Error(String(err)), - method, - this.getConnectionDiagnostics(), - ), - ); - }); - }); - return resultSchema.parse(raw); + if (response.kind === "error") { + throw new ProtocolError( + response.error.code, + response.error.message, + response.error.data, + ); + } + return resultSchema.parse(response.result); } - /** - * Transport consume-response hook: resolve/reject a pending - * {@link rawWireRequest} when its response arrives, and report it as consumed - * (so the transport does not forward it to the SDK Client, which never sent - * it). Returns false for any id we don't own, leaving normal SDK traffic - * untouched. - */ private consumeRawWireResponse( message: JSONRPCResultResponse | JSONRPCErrorResponse, ): boolean { - const id = String((message as { id?: unknown }).id); - const pending = this.pendingRawWireRequests.get(id); - if (!pending) { - return false; - } - this.pendingRawWireRequests.delete(id); - clearTimeout(pending.timer); - if ("error" in message) { - const err = (message as JSONRPCErrorResponse).error; - pending.reject(new Error(err?.message ?? `Request ${id} failed`)); - } else { - pending.resolve((message as JSONRPCResultResponse).result); - } - return true; + return this.rawWire.consume(message); } - /** - * Reject and clear all pending raw-wire requests — on every route out that - * can hold one, and at the top of `connect()` for the route in that settles - * nothing (see the comment there). - */ private rejectPendingRawWireRequests(reason: string): void { - for (const [, pending] of this.pendingRawWireRequests) { - clearTimeout(pending.timer); - pending.reject(new Error(reason)); - } - this.pendingRawWireRequests.clear(); + this.rawWire.rejectAll(reason); } - /** - * Get requestor task status by taskId (tasks we created on the server) - * @param taskId Task identifier - * @returns Task status - */ - async getRequestorTask(taskId: string): Promise<Task> { - if (!this.client) { - throw new Error("Client is not connected"); - } - // Modern (SEP-2663): `tasks/get` returns a `DetailedTask` (ttlMs/pollIntervalMs, - // inlined result/error/inputRequests) — a different wire shape than the - // deprecated SDK schema. Parse with the explicit modern schema and normalize - // onto the internal Task shape, stamping the extension client capability. - if (this.isTasksExtensionNegotiated()) { - const modern = await this.rawWireRequest( - "tasks/get", - this.withModernTaskEnvelope({ taskId }), - ModernGetTaskResultSchema, - ); - const task = normalizeModernTask(modern); - this.dispatchTypedEvent("requestorTaskUpdated", { - taskId: task.taskId, - task, - }); - return task; - } - // Legacy (2025-11-25): SDK v2 removed `client.experimental.tasks.*`; drive - // the `tasks/get` wire method directly with its deprecated-but-importable - // result schema. `GetTaskResult` is the flattened task object. - const task = (await this.client.request( - { method: "tasks/get", params: { taskId } }, - GetTaskResultSchema, - this.getRequestOptions(), - )) as Task; - - // Dispatch client-origin event (taskStatusChange is server-only) + /** Fetch a task created by this client and publish its latest state. */ + async getRequestorTask(taskId: string): Promise<InspectorTask> { + const view = await this.runTaskSessionOperation((session) => + session.task(extTaskId(taskId)).snapshot(), + ); + const task = toInspectorTask(view); this.dispatchTypedEvent("requestorTaskUpdated", { taskId: task.taskId, - task: task, + task, }); return task; } - /** - * Get requestor task result by taskId (tasks we created on the server) - * @param taskId Task identifier - * @returns Task result - */ + /** Fetch the terminal result of a task created by this client. */ async getRequestorTaskResult(taskId: string): Promise<CallToolResult> { - if (!this.client) { - throw new Error("Client is not connected"); - } - // `tasks/result` returns the task's stored payload; for a task-augmented - // tool call that payload is a CallToolResult, so validate with - // CallToolResultSchema (replacing the removed experimental helper). - return await this.client.request( - { method: "tasks/result", params: { taskId } }, - CallToolResultSchema, - this.getRequestOptions(), + const outcome = await this.runTaskSessionOperation((session) => + session.task(extTaskId(taskId)).result({ + resultCodec: taskToolResultCodec, + }), ); + return this.unwrapTaskOutcome(outcome); } - /** - * Cancel a running requestor task (task we created on the server) - * @param taskId Task identifier - * @returns Cancel result - */ + /** Cancel a running task created by this client. */ async cancelRequestorTask(taskId: string): Promise<void> { - if (!this.client) { - throw new Error("Client is not connected"); - } - // Mark before awaiting: cancelling unblocks the in-flight callToolStream, - // whose error message may arrive before this resolves — the stream's error - // path reads this set to label the task "cancelled" rather than "failed". - this.cancelledTaskIds.add(taskId); - // If the task is paused at `input_required` (its poll loop blocked on the - // pending-request modal), abort it so the modal closes and the poll observes - // the cancellation — otherwise the user is stuck answering a modal that a - // non-advancing server would keep re-showing. - const inputAbort = this.taskInputAbortControllers.get(taskId); - if (inputAbort) { - inputAbort.abort(new Error(`Task ${taskId} cancelled by user`)); - } - // Modern `tasks/cancel` is a raw-wire request (the SDK era gate blocks the - // spec-method name on 2026-07-28); legacy uses the SDK path + deprecated - // schema. - if (this.isTasksExtensionNegotiated()) { - await this.rawWireRequest( - "tasks/cancel", - this.withModernTaskEnvelope({ taskId }), - ModernCancelTaskResultSchema, - ); - } else { - await this.client.request( - { method: "tasks/cancel", params: { taskId } }, - CancelTaskResultSchema, - this.getRequestOptions(), - ); - } - - // Dispatch event + await this.runTaskSessionOperation((session) => + session.cancelTask(extTaskId(taskId)), + ); + this.cancelPendingTaskInput(taskId); this.dispatchTypedEvent("taskCancelled", { taskId }); } @@ -3162,21 +2522,13 @@ export class InspectorClient extends InspectorClientEventTarget { * observable status advances on a subsequent `tasks/get` poll (the update is * eventually consistent). Modern-only — legacy tasks surface input through the * server→client request channel, not `tasks/update`. - * - * @param taskId Task identifier - * @param inputResponses Responses keyed by the server's `inputRequests` ids */ async updateRequestorTask( taskId: string, inputResponses: Record<string, unknown>, ): Promise<void> { - if (!this.client) { - throw new Error("Client is not connected"); - } - await this.rawWireRequest( - "tasks/update", - this.withModernTaskEnvelope({ taskId, inputResponses }), - ModernUpdateTaskResultSchema, + await this.runTaskSessionOperation((session) => + session.task(extTaskId(taskId)).updateJson(inputResponses), ); } @@ -3212,28 +2564,53 @@ export class InspectorClient extends InspectorClientEventTarget { } /** - * List all requestor tasks with optional pagination (tasks we created on the server) - * @param cursor Optional pagination cursor - * @returns List of tasks with optional next cursor - */ + * Set the ambient abort signal that {@link getRequestOptions} folds into every + * request made while it is set, and return a disposer that restores the prior + * value (#1783). The daemon wraps a single serialized command in this so a + * caller disconnect cancels whatever request is in flight — for any method, + * not just tool calls — and the per-connection rpc queue frees instead of + * wedging behind a request the server never answers. + * + * It composes with a per-call `signal` via {@link combineAbortSignals}, so a + * tool call still runs its own {@link cancelToolCall} flow; this only adds + * coverage for the methods that pass no signal of their own. Pass `undefined` + * to clear. Nested scopes restore correctly because the disposer puts back + * exactly the value it replaced. + */ + setAmbientRequestSignal(signal: AbortSignal | undefined): () => void { + const previous = this.ambientRequestSignal; + this.ambientRequestSignal = signal; + return () => { + this.ambientRequestSignal = previous; + }; + } + + /** List server-held tasks created by this client. */ async listRequestorTasks( cursor?: string, - ): Promise<{ tasks: Task[]; nextCursor?: string }> { - if (!this.client) { - throw new Error("Client is not connected"); - } - const result = await this.client.request( - { - method: "tasks/list", - // `!== undefined`, not truthiness: a cursor is opaque and `""` is a - // legal value a server may hand back. Dropping it asks for page one - // again, so a caller walking pages would loop on the first page. - params: cursor !== undefined ? { cursor } : {}, - }, - ListTasksResultSchema, - this.getRequestOptions(), + ): Promise<{ tasks: InspectorTask[]; nextCursor?: string }> { + const result = await this.runTaskSessionOperation((session) => + session.listTasks(cursor), ); - return { tasks: result.tasks as Task[], nextCursor: result.nextCursor }; + return { + tasks: result.tasks.map((task) => toInspectorTask(task)), + nextCursor: result.nextCursor, + }; + } + + /** Run a task operation through auth recovery against the current session. */ + private async runTaskSessionOperation<T>( + operation: (session: TaskEnabledSession) => Promise<T>, + ): Promise<T> { + try { + return await this.withDirectAuthRecovery(() => { + const session = this.taskSession; + if (!session) throw new Error("Client is not connected"); + return operation(session); + }); + } catch (error) { + throw unwrapTaskDispatchError(error); + } } /** @@ -3435,6 +2812,7 @@ export class InspectorClient extends InspectorClientEventTarget { * On legacy connections a server never returns `input_required`, so the first * response is always complete and this is a single `client.request` call. */ + private async requestWithInputRequired<TSchema extends StandardSchemaV1>( method: "tools/call" | "prompts/get" | "resources/read", params: Record<string, unknown>, @@ -4063,6 +3441,9 @@ export class InspectorClient extends InspectorClientEventTarget { // goes straight through the transport (still logged for the Protocol / // Network tabs) and only the caller's schema is applied. Legacy keeps the // ordinary SDK path, which honors request options and `_meta` for us. + const requestOptions = this.getRequestOptions( + this.progressTokenOf(metadata), + ); const page = await this.invokeMcpClient( () => this.isModernEra() @@ -4070,11 +3451,15 @@ export class InspectorClient extends InspectorClientEventTarget { method, this.withModernEnvelope(params), pageSchema, + { + signal: requestOptions.signal, + timeoutMs: requestOptions.timeout, + }, ) : this.client!.request( { method, params }, pageSchema, - this.getRequestOptions(this.progressTokenOf(metadata)), + requestOptions, ), { method }, ); @@ -4257,11 +3642,11 @@ export class InspectorClient extends InspectorClientEventTarget { ListToolsResultSchema, "tools", ); - // Through `invokeMcpClient`, like the strict `listTools` above and like - // `salvageList`'s walk: this re-fetch can meet an auth challenge of its - // own, and outside that wrapper the recovery never runs. Its failure is - // then swallowed by `listAllTools`'s best-effort scan catch, leaving a - // stale excluded-tools set and no sign of why. + // Apply direct auth recovery and route modern pages below the SDK codec + // while legacy pages retain ordinary SDK request semantics. + const requestOptions = this.getRequestOptions( + this.progressTokenOf(metadata), + ); const page = await this.invokeMcpClient( () => this.isModernEra() @@ -4269,11 +3654,15 @@ export class InspectorClient extends InspectorClientEventTarget { "tools/list", this.withModernEnvelope(params), pageSchema, + { + signal: requestOptions.signal, + timeoutMs: requestOptions.timeout, + }, ) : this.client!.request( { method: "tools/list", params }, pageSchema, - this.getRequestOptions(this.progressTokenOf(metadata)), + requestOptions, ), { method: "tools/list" }, ); @@ -4434,6 +3823,9 @@ export class InspectorClient extends InspectorClientEventTarget { try { return await this.attemptToolCall(request, abortController.signal); } catch (error) { + const operationError = unwrapTaskDispatchError(error); + if (operationError instanceof ToolCallCancelledError) + throw operationError; // The controller was aborted. A deliberate `cancelToolCall()` (matched // by reason) means the SDK already sent `notifications/cancelled` if the // abort landed during a `client.request` leg — so surface a clean @@ -4451,7 +3843,7 @@ export class InspectorClient extends InspectorClientEventTarget { ) { throw new ToolCallCancelledError(tool.name); } - const urlElicitations = getUrlElicitationsFromError(error); + const urlElicitations = getUrlElicitationsFromError(operationError); if ( urlElicitations && urlElicitations.length > 0 && @@ -4516,9 +3908,11 @@ export class InspectorClient extends InspectorClientEventTarget { args, generalMetadata, toolSpecificMetadata, - error instanceof Error ? error.message : String(error), + operationError instanceof Error + ? operationError.message + : String(operationError), ); - throw error; + throw operationError; } } } @@ -4543,30 +3937,125 @@ export class InspectorClient extends InspectorClientEventTarget { return { ...args, ...convertToolParameters(tool, stringArgs) }; } - /** - * SEP-2243: mirror `x-mcp-header`-annotated arguments into `Mcp-Param-*` - * headers on a modern connection. The SDK only does this inside - * `client.callTool()` (and skips it in the browser), but we route - * `tools/call` through `client.request()` for manual MRTR driving (#1704), so - * we mirror ourselves. `Protocol.request` forwards `headers` (preserved - * across MRTR retry legs) to the transport, and the remote transport relays - * them to the backend's upstream send — issued server-side, where the browser - * skip doesn't apply. No-op on legacy/stdio (no annotations). - * - * Applied by BOTH `tools/call` entry points: a plain call - * ({@link attemptToolCall}) and a task-augmented one - * ({@link callToolStream}) — a strict modern server rejects either with - * `-32020` when the mirrored header is missing. - */ + /** Add modern x-mcp-* parameter mirrors without dropping caller headers. */ private applyMirroredParamHeaders( - tool: Tool, - convertedArgs: Record<string, JsonValue>, requestOptions: RequestOptions, + tool: Tool, + args: Record<string, JsonValue>, ): void { - if (this.protocolEra !== "modern") return; - const paramHeaders = mcpParamHeadersForTool(tool, convertedArgs); - if (Object.keys(paramHeaders).length === 0) return; - requestOptions.headers = { ...requestOptions.headers, ...paramHeaders }; + if (!this.isModernEra()) return; + const mirroredHeaders = mcpParamHeadersForTool(tool, args); + if (Object.keys(mirroredHeaders).length === 0) return; + requestOptions.headers = { + ...requestOptions.headers, + ...mirroredHeaders, + }; + } + + /** + * Return only the headers the ext-tasks raw call supports. The package/raw + * channel owns its fixed request timeout, and task progress is observed from + * transport events; ordinary SDK calls continue to use getRequestOptions(). + */ + private mirroredTaskParamHeaders( + tool: Tool, + args: Record<string, JsonValue>, + ): Readonly<Record<string, string>> | undefined { + if (!this.isModernEra()) return undefined; + const headers = mcpParamHeadersForTool(tool, args); + return Object.keys(headers).length === 0 ? undefined : headers; + } + + private async callTaskToolAndSettle( + tool: Tool, + args: Record<string, JsonValue>, + metadata: RequestMetadata | undefined, + preference: "allow" | "prefer", + retentionMs: number | undefined, + signal?: AbortSignal, + progressToken?: ProgressToken, + recovery?: { reference?: SerializedTaskReference }, + ): Promise<CallToolResult> { + const session = this.taskSession; + if (!session) throw new Error("Client is not connected"); + const headers = this.mirroredTaskParamHeaders(tool, args); + let lastTask: InspectorTask | undefined; + let outcomeEmitted = false; + if (progressToken !== undefined) this.taskProgress.acquire(progressToken); + try { + // On an auth-recovery rerun, resume the task the first attempt created + // (its reference is captured below), because repeating tools/call would + // start a duplicate task on the server. + const execution = recovery?.reference + ? await session.resumeTask<CallToolResult>(recovery.reference, { + resultCodec: taskToolResultCodec, + declaration: toolDeclarationFromMcpTool(tool), + signal, + }) + : await session.callTool<CallToolResult>( + tool.name, + toJsonValue(args) as Readonly<Record<string, TasksJsonValue>>, + { + resultCodec: taskToolResultCodec, + declaration: toolDeclarationFromMcpTool(tool), + signal, + requestTimeoutMs: this.requestTimeout, + // Forwarded so legacy calls routed through the SDK adapter keep + // the inactivity-timeout semantics of the old direct + // client.request path (modern raw dispatch re-arms on its own). + resetTimeoutOnProgress: this.resetTimeoutOnProgress, + task: { preference, retentionMs }, + ...(metadata === undefined + ? {} + : { + metadata: toJsonValue(metadata) as Readonly< + Record<string, TasksJsonValue> + >, + }), + ...(headers === undefined ? {} : { headers }), + }, + ); + if (recovery !== undefined && execution.kind === "task") { + recovery.reference = execution.serializeReference(); + } + const settlement = await execution.settle({ + signal, + onEvent: (event) => { + const emitted = this.emitTaskExecutionEvent(event, progressToken); + // A dispatched outcome event already carried the terminal error, + // so the catch below must not re-emit the same failure. + if (event.type === "outcome" && emitted !== undefined) + outcomeEmitted = true; + lastTask = emitted ?? lastTask; + }, + }); + if (settlement.outcome.status === "cancelled") { + // Synthesize the terminal "cancelled" update when none was observed, + // because the cancel ack ends the local task lifetime immediately — + // often before the server publishes a cancelled snapshot — and the + // UI must land on the true state without a refresh (#1455). + if (settlement.outcome.task === undefined && lastTask !== undefined) { + const task: InspectorTask = { ...lastTask, status: "cancelled" }; + const detail = { taskId: task.taskId, task }; + this.dispatchTypedEvent("toolCallTaskUpdated", detail); + this.dispatchTypedEvent("requestorTaskUpdated", detail); + } + throw new ToolCallCancelledError(tool.name); + } + return this.unwrapTaskOutcome(settlement.outcome); + } catch (error) { + const operationError = unwrapTaskDispatchError(error); + if ( + !(operationError instanceof ToolCallCancelledError) && + !outcomeEmitted + ) { + this.emitTaskError(lastTask, operationError); + } + throw operationError; + } finally { + if (progressToken !== undefined) + this.taskProgress.release(progressToken, lastTask?.taskId); + } } /** @@ -4587,116 +4076,82 @@ export class InspectorClient extends InspectorClientEventTarget { taskOptions, options, } = request; - const client = this.client; - if (!client) { - throw new Error("Client is not connected"); - } + if (!this.client) throw new Error("Client is not connected"); const convertedArgs = this.convertStringToolArgs(tool, args); - - // Merge general metadata with tool-specific metadata; tool-specific wins. - const callMetadata: RequestMetadata | undefined = + const callMetadata = generalMetadata || toolSpecificMetadata ? { ...(generalMetadata || {}), ...(toolSpecificMetadata || {}) } : undefined; - - const timestamp = new Date(); - // Fold in this client's defaultMetadata so server-wide _meta reaches - // the wire even when the caller passed nothing. const metadata = this.mergeMeta(callMetadata); + let invocationMetadata = metadata; + const timestamp = new Date(); - const callParams: { - name: string; - arguments: Record<string, JsonValue>; - _meta?: RequestMetadata; - task?: { ttl: number }; - } = { - name: tool.name, - arguments: convertedArgs, - _meta: metadata, - }; - if (taskOptions?.ttl != null) { - callParams.task = { ttl: taskOptions.ttl }; + let result: CallToolResult; + if (taskOptions === undefined && !this.isModernEra()) { + const params = { + name: tool.name, + arguments: convertedArgs, + ...(metadata ? { _meta: metadata } : {}), + }; + const requestOptions = this.getRequestOptions( + this.progressTokenOf(metadata), + signal, + ); + this.applyMirroredParamHeaders(requestOptions, tool, convertedArgs); + result = await this.invokeMcpClient( + () => + this.requestWithInputRequired( + "tools/call", + params, + CallToolResultSchema, + requestOptions, + ), + { method: "tools/call", toolName: tool.name }, + ); + } else { + const progressToken = this.progress + ? (this.progressTokenOf(metadata) ?? crypto.randomUUID()) + : undefined; + invocationMetadata = + progressToken === undefined + ? metadata + : { ...(metadata ?? {}), progressToken }; + // Shared across recovery reruns because a rerun must resume the task the + // first attempt already created rather than start a duplicate. + const recovery: { reference?: SerializedTaskReference } = {}; + result = await this.withDirectAuthRecovery( + () => + this.callTaskToolAndSettle( + tool, + convertedArgs, + invocationMetadata, + taskOptions === undefined ? "allow" : "prefer", + taskOptions?.ttl, + signal, + progressToken, + recovery, + ), + { method: "tools/call", toolName: tool.name }, + ); } - const requestOptions = this.getRequestOptions( - this.progressTokenOf(metadata), - signal, - ); - this.applyMirroredParamHeaders(tool, convertedArgs, requestOptions); - // Route through the MRTR driver (`requestWithInputRequired`) so a modern - // `input_required` result pauses at the pending-request UI and retries with - // the user's answer (#1704). Both eras use `client.request` with - // `CallToolResultSchema`; on legacy this is a single round. We deliberately - // do NOT use `client.callTool` (which would auto-fulfil / reject on an - // `input_required` result) — its only extra behavior over `request` is - // structuredContent output validation, which we already re-implement below - // via `validateToolOutput`. MCP Apps passthrough (skipOutputValidation) - // simply skips that check; both paths yield a CallToolResult once the - // driver returns a complete (non-`input_required`) result. - const rawResult = await this.invokeMcpClient( - () => - this.requestWithInputRequired( - "tools/call", - callParams, - CallToolResultSchema, - requestOptions, - ), - { method: "tools/call", toolName: tool.name }, - ); - - // Unsolicited modern task handle (SEP-2663): on a modern connection the - // server may answer ANY `tools/call` with a task rather than a result. The - // transport rewrote that frame into a `CallToolResult` carrying the real - // `DetailedTask` in `_meta`; poll it to completion here (the run-as-task - // path does the same via `callToolStream`) so the ordinary call resolves to - // the task's final result and the Tasks tab tracks it. - const taskHandle = (rawResult as CallToolResult)._meta?.[ - MODERN_TASK_HANDLE_META - ] as ModernDetailedTask | undefined; - const result = taskHandle - ? await this.pollModernTaskToTermination(taskHandle) - : rawResult; - - // Output-schema validation. SDK v2's `callTool` relaxed some checks (e.g. it - // no longer rejects a structuredContent with undeclared properties against a - // strict `additionalProperties: false` schema), so we run our own Ajv check - // to preserve the Inspector's v1 behavior: - // - default path: strict — a schema violation rejects the call (matching - // what a strict host would do), so the caller sees the error. - // - skipOutputValidation (MCP Apps passthrough): non-fatal — surface it as - // an advisory so a schema-violating-but-real result still reaches the app. const outputValidationError = this.validateToolOutput(tool, result); if (outputValidationError && !options?.skipOutputValidation) { - // Match the prior contract: on v1 a strict output-schema violation - // surfaced as the SDK's typed `McpError`/`ProtocolError` (code - // InvalidParams), not a bare Error — so downstream code that branches on - // `instanceof ProtocolError` / `error.code` keeps working. throw new ProtocolError( ProtocolErrorCode.InvalidParams, outputValidationError, ); } - const invocation: ToolCallInvocation = { toolName: tool.name, params: args, result, timestamp, success: true, - metadata, + metadata: invocationMetadata, outputValidationError, }; - - this.dispatchTypedEvent("toolCallResultChange", { - toolName: tool.name, - params: args, - result: invocation.result, - timestamp, - success: true, - metadata, - outputValidationError, - }); - + this.dispatchTypedEvent("toolCallResultChange", invocation); return invocation; } @@ -4787,401 +4242,187 @@ export class InspectorClient extends InspectorClientEventTarget { return validateToolOutput(this.outputValidator, tool, result); } - /** - * When a modern (SEP-2663) task is `input_required`, fulfil its embedded - * `inputRequests` through the pending-request UI and submit them via - * `tasks/update`. No-op for any other status. Shared by the streaming - * ({@link pollTaskToolCall}) and ordinary ({@link pollModernTaskToTermination}) - * poll loops so the input handling lives in one place. - * - * `priorRounds` is the count of `input_required` rounds already handled for - * this task; the return value is the updated count. A non-conformant server - * that keeps returning `input_required` without ever completing would - * otherwise re-prompt the user on every poll forever, so we bound it with the - * same {@link MRTR_MAX_ROUNDS} cap the MRTR driver uses. - */ - private async submitModernTaskInput( - detailed: ModernDetailedTask, - task: Task, - priorRounds: number, - signal?: AbortSignal, - ): Promise<number> { - if (task.status !== "input_required") { - return priorRounds; - } - const rounds = priorRounds + 1; - if (rounds > InspectorClient.MRTR_MAX_ROUNDS) { - throw new Error( - `Modern task "${task.taskId}" exceeded ${InspectorClient.MRTR_MAX_ROUNDS} input_required rounds without completing.`, - ); + private cancelPendingTaskInput(taskId: string): void { + for (const request of [ + ...this.pendingElicitations, + ...this.pendingSamples, + ]) { + if (request.taskId === taskId) request.cancel(); } - const inputResponses = await this.fulfilInputRequests( - this.tagInputRequestsWithTask(readInputRequests(detailed), task.taskId), - signal, - "task-input-required", + this.pendingElicitations = this.pendingElicitations.filter( + (request) => request.taskId !== taskId, ); - /* v8 ignore next 3 -- a conformant `input_required` task always carries - `inputRequests`, so `fulfilInputRequests` returns a (possibly empty) - object here, never undefined; the guard is defensive. */ - if (inputResponses) { - await this.updateRequestorTask(task.taskId, inputResponses); - } - return rounds; + this.pendingSamples = this.pendingSamples.filter( + (request) => request.taskId !== taskId, + ); + this.dispatchTypedEvent( + "pendingElicitationsChange", + this.pendingElicitations, + ); + this.dispatchTypedEvent("pendingSamplesChange", this.pendingSamples); } - /** - * Stamp `_meta[RELATED_TASK_META_KEY]` with the owning task id on each embedded - * request of a modern task's `inputRequests`. The pending-request UI reads that - * id (via `ElicitationCreateMessage.taskId`) so its Cancel control can cancel - * the TASK — not just answer the request — when a task is paused at - * `input_required`. - */ - private tagInputRequestsWithTask( - inputRequests: InputRequests | undefined, - taskId: string, - ): InputRequests | undefined { - /* v8 ignore next -- only called for an input_required task, which always - carries inputRequests; the undefined passthrough is defensive. */ - if (!inputRequests) return inputRequests; - const tagged: Record<string, unknown> = {}; - for (const [key, req] of Object.entries(inputRequests)) { - const request = req as { params?: { _meta?: Record<string, unknown> } }; - tagged[key] = { - ...request, - params: { - ...request.params, - _meta: { - ...request.params?._meta, - [RELATED_TASK_META_KEY]: { taskId }, - }, + /** Attach ext-tasks to the negotiated SDK session without taking over tool discovery. */ + private async attachTaskSession(): Promise<void> { + const client = this.client; + if (!client) return; + const endpointId = await taskSessionEndpointId( + this.transportConfig, + this.clientInfo, + ); + await this.closeTaskSession(); + // Re-check session ownership because disconnect() (or a transport crash) + // can overtake either await above: it closes the current task session + // while this method is suspended, and installing a new session on the + // torn-down client would leave extension-owned callbacks and state alive + // until a later reconnect/disconnect. `disconnecting` covers a teardown + // that claimed ownership but has not yet settled the status. + if ( + this.client !== client || + this.disconnecting || + this.status !== "connected" + ) + return; + this.taskSession = createTaskSessionFromClient(client, { + endpointId, + // ext-tasks is unbounded by default; the Inspector opts into the same + // runaway budget its own raw MRTR path enforces (MRTR_MAX_ROUNDS). + maxInputRounds: InspectorClient.MRTR_MAX_ROUNDS, + rawDispatch: this.dispatchTaskRequest, + v2RequestFraming: { + protocolVersion: this.protocolVersion!, + clientInfo: toJsonValue(this.clientInfo) as Readonly< + Record<string, TasksJsonValue> + >, + clientCapabilities: toJsonValue(this.clientCapabilities) as Readonly< + Record<string, TasksJsonValue> + >, + }, + // declaration. A no-op recovery fallback prevents duplicate tools/list traffic. + tools: { currentTool: () => undefined }, + onInputRequest: createApplicationInputHandler({ + elicitation: async (request, context) => { + const result = await this.enqueuePendingElicitation( + { + method: "elicitation/create", + params: request.params, + } as ElicitRequest, + taskInputOrigin(context.delivery), + context.signal, + ); + return { + ...jsonObject(result), + action: result.action, + ...(result.content === undefined + ? {} + : { + // Narrowing cast (subset of JsonValue): the elicitation UI + // produces schema-constrained scalars/string arrays, and + // the ext-tasks wire schema re-validates on send. + content: jsonObject(result.content) as Readonly< + Record<string, ApplicationElicitContentValue> + >, + }), + }; }, - }; - } - return tagged as InputRequests; - } - - /** - * Terminal outcome for a modern task: the inlined `CallToolResult` for a - * `completed` task (SEP-2663 removed the blocking `tasks/result`), or a - * `ProtocolError` for `failed` / `cancelled`. Shared so both poll loops agree - * on the result/error shape. - */ - private modernTaskTerminalOutcome( - task: Task, - detailed: ModernDetailedTask, - ): - | { type: "result"; result: CallToolResult } - | { type: "error"; error: ProtocolError } { - if (task.status === "completed") { - /* v8 ignore next -- a conformant `completed` task always inlines its - `result`; the `{ content: [] }` fallback is defensive. */ - return { - type: "result", - result: (detailed.result ?? { content: [] }) as CallToolResult, - }; - } - return { - type: "error", - error: new ProtocolError( - ProtocolErrorCode.InternalError, - task.statusMessage ?? `Task ${task.status}`, - ), - }; + sampling: async (request, context) => { + const result = await this.enqueuePendingSample( + { + method: "sampling/createMessage", + params: request.params, + } as CreateMessageRequest, + taskInputOrigin(context.delivery), + context.signal, + ); + return { + ...jsonObject(result), + model: result.model, + role: result.role, + // Narrowing cast (subset of JsonValue): the SDK sampling result + // carries protocol content blocks, and the ext-tasks wire schema + // re-validates on send — same shape as the elicitation cast above. + content: toJsonValue(result.content) as + | ApplicationSamplingContentBlock + | readonly ApplicationSamplingContentBlock[], + }; + }, + roots: async () => ({ roots: this.applicationRoots() }), + }), + onError: (error) => + this.logger.error({ error }, "ext-tasks background error"), + }); } /** - * Poll cadence for a task: the server-advertised `pollInterval` when - * positive, else the default. Shared by every task poll loop (both eras). + * Project configured roots to the ext-tasks handler shape. Mapped + * field-by-field because ApplicationRoot requires a typed string uri, + * which a JSON-record projection would erase. */ - private taskPollInterval(task: Task): number { - const advertised = task.pollInterval; - if (typeof advertised !== "number") return DEFAULT_TASK_POLL_INTERVAL_MS; - // A spec-conformant server never advertises a non-positive interval; the - // `> 0` guard is defensive against a malformed value. - /* v8 ignore next -- non-positive pollInterval is unreachable from a conformant server. */ - return advertised > 0 ? advertised : DEFAULT_TASK_POLL_INTERVAL_MS; + private applicationRoots(): readonly ApplicationRoot[] { + return ( + this.roots?.map((root) => ({ + uri: root.uri, + ...(root.name === undefined ? {} : { name: root.name }), + ...(root._meta === undefined ? {} : { _meta: jsonObject(root._meta) }), + })) ?? [] + ); } - /** - * Register a per-task abort controller (keyed by taskId) whose signal gates - * the task's `input_required` pending request, and return the signal plus a - * `release` cleanup. {@link cancelRequestorTask} aborts it to unblock a task - * paused at the pending-request modal. - */ - private registerTaskInputAbort(taskId: string): { - signal: AbortSignal; - release: () => void; - } { - const controller = new AbortController(); - this.taskInputAbortControllers.set(taskId, controller); - return { - signal: controller.signal, - release: () => { - // Only delete our own entry — tool calls are serial, so a second task - // never replaces this id's controller mid-poll; the guard is defensive. - /* v8 ignore next */ - if (this.taskInputAbortControllers.get(taskId) === controller) { - this.taskInputAbortControllers.delete(taskId); - } - }, - }; + /** Release extension-owned state without ever leaving the SDK adapter installed. */ + private async closeTaskSession(): Promise<void> { + const session = this.taskSession; + this.taskSession = null; + await session?.close(); } - /** - * Drive a modern (SEP-2663) task to a terminal state from a seed - * `DetailedTask`, dispatching task events so the Tasks tab and toasts track - * it, and return the completed task's inlined `CallToolResult` (or throw on - * `failed` / `cancelled`). Used by the ORDINARY `callTool` path when a server - * returns an unsolicited task handle (the run-as-task streaming path drives - * the equivalent loop inline in {@link pollTaskToolCall}). `input_required` - * rounds are answered through the pending-request UI and submitted via - * `tasks/update`. - */ - private async pollModernTaskToTermination( - seed: ModernDetailedTask, - ): Promise<CallToolResult> { - let detailed = seed; - let task = normalizeModernTask(detailed); - const emit = (t: Task): void => { - this.dispatchTypedEvent("toolCallTaskUpdated", { - taskId: t.taskId, - task: t, - }); - this.dispatchTypedEvent("requestorTaskUpdated", { - taskId: t.taskId, - task: t, - }); - }; - emit(task); - const { signal: inputSignal, release } = this.registerTaskInputAbort( - task.taskId, - ); + private async closeTaskSessionBestEffort(): Promise<void> { try { - let inputRounds = 0; - while (!InspectorClient.isTerminalTaskStatus(task.status)) { - inputRounds = await this.submitModernTaskInput( - detailed, - task, - inputRounds, - inputSignal, - ); - await new Promise((resolve) => - setTimeout(resolve, this.taskPollInterval(task)), - ); - detailed = await this.rawWireRequest( - "tasks/get", - this.withModernTaskEnvelope({ taskId: task.taskId }), - ModernGetTaskResultSchema, - ); - task = normalizeModernTask(detailed); - emit(task); - } - } finally { - release(); - } - const outcome = this.modernTaskTerminalOutcome(task, detailed); - if (outcome.type === "error") { - throw outcome.error; + await this.closeTaskSession(); + } catch (error) { + this.logger.warn({ error }, "Failed to close ext-tasks session"); } - return outcome.result; } - /** - * Poll a task-augmented tool call to completion. Replaces the removed - * `client.experimental.tasks.callToolStream` helper: it sends the - * task-augmented `tools/call` (the server responds with a task handle, i.e. a - * `CreateTaskResult`), then polls `tasks/get` until the task reaches a - * terminal status, yielding the same `taskCreated | taskStatus | result | - * error` message shapes the caller's `for await` loop consumes — so all the - * downstream event dispatch and terminal-state handling stays unchanged. - */ - private async *pollTaskToolCall( - params: CallToolRequest["params"], - requestOptions: RequestOptions, - ): AsyncGenerator< - | { type: "taskCreated"; task: Task } - | { type: "taskStatus"; task: Task } - | { type: "result"; result: CallToolResult } - | { type: "error"; error: ProtocolError } - > { - if (!this.client) { - throw new Error("Client is not connected"); - } - const client = this.client; - // The server streams `notifications/progress` for a task AFTER the - // task-augmented `tools/call` has already returned its `{ task }` handle. But - // SDK v2 deletes a request's progress subscription the moment that request - // resolves, so those later ticks would be dropped. Capture the subscription - // id the SDK registers for this request (the only new key in the private - // `_progressHandlers` map) so we can keep the caller's `onprogress` alive - // through the poll and clean it up when the task terminates. - // SDK gap: `Client` exposes no public API to keep a progress subscription - // alive across a resolved request (or to subscribe to progress by token), so - // we reach the private `_progressHandlers` map through a narrowed cast. A - // public "durable progress subscription" hook would remove this cast. - const progressHandlers = ( - client as unknown as { - _progressHandlers: Map<number, ProgressCallback>; - } - )._progressHandlers; - const keysBeforeRequest = new Set(progressHandlers.keys()); - // Create the task-augmented tool call. A task-capable server returns a task - // handle (`CreateTaskResult` = `{ task }`), but a server that completes - // synchronously (or for which the tool forbids/ignores task augmentation) - // may return an immediate `CallToolResult` instead — accept either with a - // union schema and branch on the presence of `task`. - // - // NOTE: the LEGACY task path does NOT opt into `allowInputRequired` (MRTR - // over legacy tasks is out of scope for #1704). The MODERN path (SEP-2663) - // instead surfaces a task's `input_required` through `tasks/get`'s - // `inputRequests` and answers via `tasks/update` (handled in the poll loop - // below), reusing the same pending-request UI. - const modernTasks = this.isTasksExtensionNegotiated(); - const requestPromise = client.request( - { - // On modern the SDK codec stamps the tasks-extension client capability - // into the request envelope (advertised at construction), so a server - // may answer with a `CreateTaskResult` — no per-call `_meta` needed. - method: "tools/call", - params, - }, - // Modern: the SDK codec can't decode a `resultType: "task"` result, so the - // transport rewrote it to a `CallToolResult` carrying the task handle in - // `_meta` — parse as a CallToolResult and read the handle below. Legacy: - // accept a `{ task }` handle or an immediate result. - modernTasks - ? CallToolResultSchema - : CreateTaskResultSchema.or(CallToolResultSchema), - requestOptions, - ); - // The SDK registers the progress handler synchronously while constructing - // the request promise (before this await), so the new key is present now. - // ASSUMES SERIAL CONSTRUCTION: `find` takes the first key not present in the - // pre-request snapshot, which is unambiguous only because no OTHER request - // registers a progress handler between the snapshot and this request's - // synchronous registration. Tool calls are user-driven and serial, so that - // holds today; if concurrent task-augmented calls are ever constructed in - // the same microtask window, two subscription ids could cross-wire and this - // must move to an SDK-supported correlation (see the delete-when-native note - // on `installReceiverTaskResponseBypass`). - const progressSubscriptionId = requestOptions.onprogress - ? [...progressHandlers.keys()].find((k) => !keysBeforeRequest.has(k)) - : undefined; - const created = await requestPromise; - - if (modernTasks) { - // Modern (SEP-2663): a task-creating `tools/call` came back as a - // `resultType: "task"` frame the SDK can't decode, so the transport - // rewrote it to a `CallToolResult` carrying the real `DetailedTask` under - // MODERN_TASK_HANDLE_META. A synchronous completion has no such handle — - // yield that `CallToolResult` directly. - const handle = (created as CallToolResult)._meta?.[ - MODERN_TASK_HANDLE_META - ] as ModernDetailedTask | undefined; - if (!handle) { - yield { type: "result", result: created as CallToolResult }; - return; - } - let detailed = handle; - let task = normalizeModernTask(detailed); - yield { type: "taskCreated", task }; - if (progressSubscriptionId != null && requestOptions.onprogress) { - progressHandlers.set(progressSubscriptionId, requestOptions.onprogress); - } - const { signal: inputSignal, release } = this.registerTaskInputAbort( - task.taskId, - ); - let inputRounds = 0; - try { - while (!InspectorClient.isTerminalTaskStatus(task.status)) { - // `input_required`: fulfil the embedded server→client requests through - // the same pending-request UI the MRTR path uses, then submit them via - // `tasks/update`. The update is eventually consistent — the task's - // status advances on a following `tasks/get`, so keep polling - // (bounded by MRTR_MAX_ROUNDS against a server that never advances). - // `inputSignal` fires if the task is cancelled while paused here. - inputRounds = await this.submitModernTaskInput( - detailed, - task, - inputRounds, - inputSignal, - ); - await new Promise((resolve) => - setTimeout(resolve, this.taskPollInterval(task)), - ); - detailed = await this.rawWireRequest( - "tasks/get", - this.withModernTaskEnvelope({ taskId: task.taskId }), - ModernGetTaskResultSchema, - ); - task = normalizeModernTask(detailed); - yield { type: "taskStatus", task }; - } - } finally { - release(); - if (progressSubscriptionId != null) { - progressHandlers.delete(progressSubscriptionId); - } - } - // Modern removes the blocking `tasks/result`: a completed task inlines its - // CallToolResult; failed/cancelled surface as an error. - yield this.modernTaskTerminalOutcome(task, detailed); - return; - } - - if (!("task" in created) || created.task == null) { - // Immediate result — no task was created; yield it directly. - yield { type: "result", result: created as CallToolResult }; - return; - } - let task = created.task as Task; - yield { type: "taskCreated", task }; + private emitTaskExecutionEvent( + event: TaskExecutionEvent<CallToolResult>, + progressToken?: ProgressToken, + ): InspectorTask | undefined { + const view = taskViewFromExecutionEvent(event); + if (view === undefined) return undefined; + const task = toInspectorTask(view); + if (progressToken !== undefined) + this.taskProgress.correlate(progressToken, task.taskId); + const outcomeDetail = + event.type !== "outcome" || event.outcome.status === "cancelled" + ? {} + : event.outcome.status === "completed" + ? { result: event.outcome.result } + : { error: toProtocolError(event.outcome.error) }; + const detail = { taskId: task.taskId, task, ...outcomeDetail }; + this.dispatchTypedEvent("toolCallTaskUpdated", detail); + this.dispatchTypedEvent("requestorTaskUpdated", detail); + return task; + } - // Revive the (now-deleted) progress subscription for the poll so task- - // execution progress ticks reach the caller's `onprogress`. - if (progressSubscriptionId != null && requestOptions.onprogress) { - progressHandlers.set(progressSubscriptionId, requestOptions.onprogress); - } + private unwrapTaskOutcome( + outcome: Parameters<typeof resultFromTaskOutcome<CallToolResult>>[0], + ): CallToolResult { try { - // Poll `tasks/get` until the task reaches a terminal status. Honour the - // server-advertised `pollInterval` when present, else the default cadence. - while (!InspectorClient.isTerminalTaskStatus(task.status)) { - await new Promise((resolve) => - setTimeout(resolve, this.taskPollInterval(task)), - ); - task = (await client.request( - { method: "tasks/get", params: { taskId: task.taskId } }, - GetTaskResultSchema, - this.getRequestOptions(), - )) as Task; - yield { type: "taskStatus", task }; - } - } finally { - if (progressSubscriptionId != null) { - progressHandlers.delete(progressSubscriptionId); - } + return resultFromTaskOutcome(outcome); + } catch (error) { + throw toProtocolError(error); } + } - if (task.status === "completed") { - const result = await client.request( - { method: "tasks/result", params: { taskId: task.taskId } }, - CallToolResultSchema, - this.getRequestOptions(), - ); - yield { type: "result", result }; - } else { - // failed | cancelled — surface as an error the caller's loop labels as - // "cancelled" (via cancelledTaskIds) or "failed". Carry a ProtocolError so - // the `error` payload matches the event map's type (the SDK helper this - // replaces also yielded a protocol-error-shaped value). - yield { - type: "error", - error: new ProtocolError( - ProtocolErrorCode.InternalError, - task.statusMessage ?? `Task ${task.status}`, - ), - }; - } + private emitTaskError( + lastTask: InspectorTask | undefined, + reason: unknown, + ): void { + if (!lastTask) return; + const error = toProtocolError(reason); + const detail = { taskId: lastTask.taskId, task: lastTask, error }; + this.dispatchTypedEvent("toolCallTaskUpdated", detail); + this.dispatchTypedEvent("requestorTaskUpdated", detail); } /** @@ -5200,232 +4441,72 @@ export class InspectorClient extends InspectorClientEventTarget { generalMetadata?: RequestMetadata, toolSpecificMetadata?: RequestMetadata, taskOptions?: { ttl?: number }, + options?: { skipOutputValidation?: boolean }, ): Promise<ToolCallInvocation> { - if (!this.client) { - throw new Error("Client is not connected"); - } + const convertedArgs = this.convertStringToolArgs(tool, args); + const callMetadata = + generalMetadata || toolSpecificMetadata + ? { ...(generalMetadata || {}), ...(toolSpecificMetadata || {}) } + : undefined; + const metadata = this.mergeMeta(callMetadata); + const progressToken = this.progress + ? (this.progressTokenOf(metadata) ?? crypto.randomUUID()) + : undefined; + const taskCallMetadata = + progressToken === undefined + ? metadata + : { ...(metadata ?? {}), progressToken }; + const timestamp = new Date(); try { - const convertedArgs = this.convertStringToolArgs(tool, args); - - // Merge general metadata with tool-specific metadata; tool-specific wins. - const callMetadata: RequestMetadata | undefined = - generalMetadata || toolSpecificMetadata - ? { ...(generalMetadata || {}), ...(toolSpecificMetadata || {}) } - : undefined; - - const timestamp = new Date(); - const metadata = this.mergeMeta(callMetadata); - - // Call the streaming API - const streamParams: Record<string, unknown> = { - name: tool.name, - arguments: convertedArgs, - }; - if (metadata) { - streamParams._meta = metadata; - } - if (taskOptions?.ttl != null) { - streamParams.task = { ttl: taskOptions.ttl }; - } - - let finalResult: CallToolResult | undefined; - let taskId: string | undefined; - let error: Error | undefined; - - // Correlate progress → task. getRequestOptions already wires onprogress to - // dispatch the generic progressNotification (keyed by the caller's - // progressToken). Wrap it so each tick that arrives after the task is - // created also dispatches requestorTaskProgress tagged with the taskId - // this stream owns — the only place that mapping is known. Ticks before - // taskCreated (rare) just fall through to the generic event. - // - // Gate on `this.progress`, mirroring getRequestOptions: when progress is - // globally disabled there's no inner handler to wrap, and we must not - // attach one here either — doing so would request a progress token (and - // emit requestorTaskProgress) for task calls only, bypassing the toggle - // that governs every other call path. - const requestOptions = this.getRequestOptions( - this.progressTokenOf(metadata), - ); - // The task-augmented `tools/call` needs the same SEP-2243 mirroring as the - // plain one — a strict modern server rejects it with -32020 otherwise. - this.applyMirroredParamHeaders(tool, convertedArgs, requestOptions); - if (this.progress) { - const innerOnProgress = requestOptions.onprogress; - requestOptions.onprogress = (progress: Progress) => { - innerOnProgress?.(progress); - if (taskId) { - this.dispatchTypedEvent("requestorTaskProgress", { - taskId, - progress, - }); - } - }; - } - - const stream = this.pollTaskToolCall( - streamParams as CallToolRequest["params"], - requestOptions, + // Shared across recovery reruns because a rerun must resume the task the + // first attempt already created rather than start a duplicate. + const recovery: { reference?: SerializedTaskReference } = {}; + const result = await this.withDirectAuthRecovery( + () => + this.callTaskToolAndSettle( + tool, + convertedArgs, + taskCallMetadata, + "prefer", + taskOptions?.ttl, + undefined, + progressToken, + recovery, + ), + { method: "tools/call", toolName: tool.name }, ); - - // Iterate through the async generator - for await (const message of stream) { - switch (message.type) { - case "taskCreated": - taskId = message.task.taskId; - this.dispatchTypedEvent("toolCallTaskUpdated", { - taskId: message.task.taskId, - task: message.task, - }); - this.dispatchTypedEvent("requestorTaskUpdated", { - taskId: message.task.taskId, - task: message.task, - }); - break; - - case "taskStatus": - if (!taskId) { - taskId = message.task.taskId; - } - this.dispatchTypedEvent("toolCallTaskUpdated", { - taskId: message.task.taskId, - task: message.task, - }); - this.dispatchTypedEvent("requestorTaskUpdated", { - taskId: message.task.taskId, - task: message.task, - }); - break; - - case "result": - finalResult = message.result as CallToolResult; - if (taskId) { - const completedTask: TaskWithOptionalCreatedAt = { - taskId, - ttl: null, - status: "completed", - statusMessage: "Task completed" as string, - lastUpdatedAt: new Date().toISOString(), - }; - this.dispatchTypedEvent("toolCallTaskUpdated", { - taskId, - task: completedTask, - result: finalResult, - }); - this.dispatchTypedEvent("requestorTaskUpdated", { - taskId, - task: completedTask, - result: finalResult, - }); - } - break; - - case "error": { - const errorMessage = - message.error.message || "Task execution failed"; - error = new Error(errorMessage); - if (taskId) { - // A user-cancelled task surfaces here as a generic error; report - // it as "cancelled" (not "failed") so the UI lands on the true - // terminal state immediately, matching what a refresh would show - // (#1455). - const cancelled = this.cancelledTaskIds.has(taskId); - // Consume the marker — task ids are single-use, so this keeps the - // set from growing across a long session of cancellations (the - // disconnect-clear stays the backstop for cancels whose task - // completed before the cancel landed and never hit this path). - this.cancelledTaskIds.delete(taskId); - const terminalTask: TaskWithOptionalCreatedAt = { - taskId, - ttl: null, - status: cancelled ? "cancelled" : "failed", - statusMessage: cancelled - ? "Client cancelled task execution." - : errorMessage, - lastUpdatedAt: new Date().toISOString(), - }; - this.dispatchTypedEvent("toolCallTaskUpdated", { - taskId, - task: terminalTask, - error: message.error, - }); - this.dispatchTypedEvent("requestorTaskUpdated", { - taskId, - task: terminalTask, - error: message.error, - }); - } - break; - } - } - } - - // If we got an error, throw it - if (error) { - throw error; - } - - // If we didn't get a result, something went wrong - // This can happen if the task completed but result wasn't in the stream - // Try to get it from the task result endpoint - if (!finalResult && taskId) { - try { - finalResult = await this.client.request( - { method: "tasks/result", params: { taskId } }, - CallToolResultSchema, - this.getRequestOptions(), // no metadata for fallback - ); - } catch (resultError) { - throw new Error( - `Tool call did not return a result: ${resultError instanceof Error ? resultError.message : String(resultError)}`, - { cause: resultError }, - ); - } - } - if (!finalResult) { - throw new Error("Tool call did not return a result"); + const outputValidationError = this.validateToolOutput(tool, result); + if (outputValidationError && !options?.skipOutputValidation) { + throw new ProtocolError( + ProtocolErrorCode.InvalidParams, + outputValidationError, + ); } - const invocation: ToolCallInvocation = { toolName: tool.name, params: args, - result: finalResult, + result, timestamp, success: true, - metadata, + metadata: taskCallMetadata, + outputValidationError, }; - - this.dispatchTypedEvent("toolCallResultChange", { - toolName: tool.name, - params: args, - result: invocation.result, - timestamp, - success: true, - metadata, - }); - + this.dispatchTypedEvent("toolCallResultChange", invocation); return invocation; } catch (error) { - // Merge general metadata with tool-specific metadata for error case - const callMetadata: RequestMetadata | undefined = - generalMetadata || toolSpecificMetadata - ? { ...(generalMetadata || {}), ...(toolSpecificMetadata || {}) } - : undefined; - - const timestamp = new Date(); - const metadata = this.mergeMeta(callMetadata); - - this.dispatchTypedEvent("toolCallResultChange", { - toolName: tool.name, - params: args, - result: null, - timestamp, - success: false, - error: error instanceof Error ? error.message : String(error), - metadata, - }); - - throw error; + const operationError = unwrapTaskDispatchError(error); + if (!(operationError instanceof ToolCallCancelledError)) { + this.dispatchFailedToolCall( + tool, + args, + generalMetadata, + toolSpecificMetadata, + operationError instanceof Error + ? operationError.message + : String(operationError), + ); + } + throw operationError; } } diff --git a/core/mcp/inspectorClientEventTarget.ts b/core/mcp/inspectorClientEventTarget.ts index f9ce6c8c74..e57d30a5a9 100644 --- a/core/mcp/inspectorClientEventTarget.ts +++ b/core/mcp/inspectorClientEventTarget.ts @@ -27,6 +27,7 @@ import type { ResourceSubscriptionStreamState, ExcludedTool, RequestMetadata, + InspectorTask, } from "./types.js"; import type { MalformedListItem } from "./listSalvage.js"; import type { ConnectionDiagnostics } from "./connectionDiagnostics.js"; @@ -37,7 +38,6 @@ import type { Root, Progress, ProgressToken, - Task, CallToolResult, ProtocolError, ProtocolEra, @@ -49,8 +49,8 @@ import type { JsonValue } from "../json/jsonUtils.js"; import type { OAuthTokens } from "@modelcontextprotocol/client"; import type { AuthChallenge } from "../auth/challenge.js"; -/** Task with createdAt optional so we can emit synthetic tasks (e.g. on result/error) that omit it. */ -export type TaskWithOptionalCreatedAt = Omit<Task, "createdAt"> & { +/** Task update shape retained for state-store compatibility. */ +export type TaskWithOptionalCreatedAt = Omit<InspectorTask, "createdAt"> & { createdAt?: string; }; @@ -164,7 +164,7 @@ export interface InspectorClientEventMap { resourceSubscriptionStreamChange: ResourceSubscriptionStreamState; // Task events /** Fired only from server notification notifications/tasks/status. */ - taskStatusChange: { taskId: string; task: Task }; + taskStatusChange: { taskId: string; task: InspectorTask }; /** Fired from callToolStream for each task update. */ toolCallTaskUpdated: { taskId: string; @@ -187,7 +187,7 @@ export interface InspectorClientEventMap { * event carries only the caller's progressToken, not the taskId). */ requestorTaskProgress: { taskId: string; progress: Progress }; - tasksChange: Task[]; + tasksChange: InspectorTask[]; // Signal events (no payload) connect: void; disconnect: void; diff --git a/core/mcp/inspectorClientProtocol.ts b/core/mcp/inspectorClientProtocol.ts index 8eefd758f9..2aeb860631 100644 --- a/core/mcp/inspectorClientProtocol.ts +++ b/core/mcp/inspectorClientProtocol.ts @@ -22,6 +22,7 @@ import type { ResourceSubscriptionStreamState, ExcludedTool, RequestMetadata, + InspectorTask, AppRendererClient, } from "./types.js"; import type { @@ -35,7 +36,6 @@ import type { Resource, ResourceTemplateType as ResourceTemplate, ServerCapabilities, - Task, Tool, } from "@modelcontextprotocol/client"; import type { JsonValue } from "../json/jsonUtils.js"; @@ -45,6 +45,7 @@ import type { InspectorClientEventTarget } from "./inspectorClientEventTarget.js import type { SkillEntry } from "./skillsSchemas.js"; import type { SkillsExtensionSupport } from "./skills.js"; import type { DirectoryReadResult } from "./skillsSchemas.js"; +import type { TaskCapabilities } from "@modelcontextprotocol/ext-tasks/client"; import type { SamplingCreateMessage } from "./samplingCreateMessage.js"; import type { ElicitationCreateMessage } from "./elicitationCreateMessage.js"; @@ -112,15 +113,17 @@ export interface InspectorClientProtocol extends InspectorClientEventTarget { ): Promise<{ resourceTemplates: ResourceTemplate[]; nextCursor?: string }>; listRequestorTasks( cursor?: string, - ): Promise<{ tasks: Task[]; nextCursor?: string }>; - /** Poll one requestor task's current status (era-aware: modern `DetailedTask` - * via `tasks/get`, or the legacy flattened task). Dispatches - * `requestorTaskUpdated`. Used by the modern task store's refresh (no - * `tasks/list`). */ - getRequestorTask(taskId: string): Promise<Task>; - /** True when a modern (2026-07-28) connection negotiated the - * `io.modelcontextprotocol/tasks` extension (SEP-2663). Gates the Tasks tab - * and the modern task store's poll-based refresh. */ + ): Promise<{ tasks: InspectorTask[]; nextCursor?: string }>; + /** Poll one requestor task's current status through the neutral task façade. */ + getRequestorTask(taskId: string): Promise<InspectorTask>; + /** + * Authoritative generation-neutral capabilities for requester task + * behavior. Required (not optional) because the managed task store gates + * refresh routing on `inventory` — an unimplemented method would silently + * fall back to `tasks/list` against sessions that reject it. + */ + getTaskSessionCapabilities(): TaskCapabilities | undefined; + /** Compatibility predicate for consumers that distinguish known-handle tasks. */ isTasksExtensionNegotiated(): boolean; /** The Skills extension (SEP-2640) the server declared, or `undefined`. diff --git a/core/mcp/messageTrackingTransport.ts b/core/mcp/messageTrackingTransport.ts index c7645dcc86..e84eff0086 100644 --- a/core/mcp/messageTrackingTransport.ts +++ b/core/mcp/messageTrackingTransport.ts @@ -39,35 +39,13 @@ export interface MessageTrackingCallbacks { trackSendFailure?: (message: JSONRPCRequest, error: unknown) => void; } -/** - * Optional rewrite of an incoming response BEFORE it reaches the SDK's codec. - * Used for extension result shapes the SDK v2 codec would reject outright — e.g. - * a modern (SEP-2663) `resultType: "task"` result, which the codec has no - * knowledge of (tasks were removed from the SDK). The ORIGINAL message is still - * what `trackResponse` logs (so the Protocol/Network tabs show the true wire); - * only the copy handed to the SDK is rewritten. Return the message unchanged to - * pass it through untouched. - */ -export type IncomingResultRewriter = ( - message: JSONRPCResultResponse, -) => JSONRPCMessage; - -/** - * Optional consumer for an incoming response the SDK Client did not originate — - * used for the raw-wire channel that drives extension methods the SDK v2 era - * gate refuses to send (e.g. modern `tasks/get`/`tasks/update`/`tasks/cancel`, - * which are spec-method names absent from the 2026-07-28 era). When this returns - * `true` the response is treated as fully handled and is NOT forwarded to the - * SDK Client (which has no pending request for it). The response is still logged - * by `trackResponse` first, so the Protocol/Network tabs see the true frame. - */ -export type IncomingResponseConsumer = ( - message: JSONRPCResultResponse | JSONRPCErrorResponse, -) => boolean; +/** Narrow raw-channel surface used to consume below-SDK responses. */ +export interface IncomingRawRequestChannel { + consume(message: JSONRPCResultResponse | JSONRPCErrorResponse): boolean; +} export interface MessageTrackingHooks { - rewriteIncomingResult?: IncomingResultRewriter; - consumeIncomingResponse?: IncomingResponseConsumer; + rawRequestChannel?: IncomingRawRequestChannel; } // Transport wrapper that intercepts all messages for tracking @@ -186,25 +164,12 @@ export class MessageTrackingTransport implements Transport { // Consume a response to a raw-wire request the SDK never sent (e.g. // a modern `tasks/get`); handled entirely by the caller, not the SDK. if ( - this.hooks.consumeIncomingResponse?.( + this.hooks.rawRequestChannel?.consume( message as JSONRPCResultResponse | JSONRPCErrorResponse, ) ) { return; } - // Rewrite a result the SDK codec can't decode (e.g. a modern - // `resultType: "task"` handle) AFTER logging the true wire, so the - // SDK receives a shape it accepts while the Protocol/Network tabs - // still show the real frame. - if (this.hooks.rewriteIncomingResult && "result" in message) { - const rewritten = this.hooks.rewriteIncomingResult( - message as JSONRPCResultResponse, - ); - if (rewritten !== message) { - handler(rewritten as T, extra); - return; - } - } } else if ("method" in message) { // This is a request coming from the server this.callbacks.trackRequest?.(message as JSONRPCRequest, "server"); diff --git a/core/mcp/modernTaskSchemas.ts b/core/mcp/modernTaskSchemas.ts deleted file mode 100644 index 8857175083..0000000000 --- a/core/mcp/modernTaskSchemas.ts +++ /dev/null @@ -1,132 +0,0 @@ -/** - * Modern (2026-07-28) task extension wire schemas — SEP-2663 - * (`io.modelcontextprotocol/tasks`). - * - * SDK v2 removed all built-in tasks support: the `Task` / `GetTaskResultSchema` - * / `CreateTaskResultSchema` it still exports are the **deprecated 2025-11-25** - * vocabulary (`ttl` / `pollInterval`, blocking `tasks/result`, `tasks/list`). - * The redesigned extension is a different wire shape — `ttlMs` / `pollIntervalMs`, - * a polymorphic `DetailedTask` that inlines `result` / `error` / `inputRequests` - * by status, `tasks/get` polling, a new `tasks/update`, no `tasks/list`, and no - * blocking `tasks/result`. There is no SDK schema for it, so the Inspector drives - * modern `tasks/*` as raw requests with these explicit schemas (the "explicit- - * schema raw-request form" the SDK docs prescribe). - * - * Schemas are intentionally permissive (`looseObject`) so an unknown wire field - * (e.g. a future status-specific member) passes through rather than failing the - * parse — the Inspector is a debugging tool and should surface, not reject. - */ - -import { z } from "zod/v4"; -import type { InputRequests, Task } from "@modelcontextprotocol/client"; - -/** SEP-2133 extension identifier for the redesigned Tasks extension (SEP-2663). */ -export const TASKS_EXTENSION_KEY = "io.modelcontextprotocol/tasks"; - -/** The modern protocol revision, used as the raw-request envelope's - * `protocolVersion` when the negotiated version isn't otherwise available. */ -export const MODERN_PROTOCOL_VERSION = "2026-07-28"; - -/** The `_meta` value stamped on modern task-eligible requests to declare the - * client supports the tasks extension (per-request capability, SEP-2663). */ -export const TASKS_EXTENSION_CLIENT_CAPABILITY = { - extensions: { [TASKS_EXTENSION_KEY]: {} }, -} as const; - -/** - * `_meta` key under which the transport-level rewriter stashes a modern task - * handle. SDK v2's codec rejects a `resultType: "task"` result outright (tasks - * were removed), so a task-creating `tools/call` response is rewritten to a - * benign `CallToolResult` carrying the real `DetailedTask` here, where the task - * poll driver reads it. See `MessageTrackingTransport`'s rewrite hook. - */ -export const MODERN_TASK_HANDLE_META = - "io.modelcontextprotocol/inspector/modernTaskHandle"; - -/** True when a decoded wire result is a modern `CreateTaskResult` - * (`resultType: "task"`) — the frame the SDK codec cannot handle. */ -export function isModernCreateTaskResult(result: unknown): boolean { - return ( - typeof result === "object" && - result !== null && - (result as { resultType?: unknown }).resultType === "task" && - typeof (result as { taskId?: unknown }).taskId === "string" - ); -} - -const ModernTaskStatusSchema = z.enum([ - "working", - "input_required", - "completed", - "failed", - "cancelled", -]); - -/** - * `DetailedTask` (SEP-2663): the modern task shape returned by `tasks/get` and - * carried by a `CreateTaskResult`. Status-specific members (`result`, `error`, - * `inputRequests`) are optional here because a single loose schema stands in for - * the wire union `Working | InputRequired | Completed | Failed | Cancelled`. - */ -export const ModernDetailedTaskSchema = z.looseObject({ - taskId: z.string(), - status: ModernTaskStatusSchema, - statusMessage: z.string().optional(), - createdAt: z.string(), - lastUpdatedAt: z.string(), - ttlMs: z.number().nullable().optional(), - pollIntervalMs: z.number().optional(), - /** Present on `completed`: the original request's result (e.g. CallToolResult). */ - result: z.record(z.string(), z.unknown()).optional(), - /** Present on `failed`: the JSON-RPC error that ended the task. */ - error: z.record(z.string(), z.unknown()).optional(), - /** Present on `input_required`: embedded server→client requests, keyed by id. */ - inputRequests: z.record(z.string(), z.unknown()).optional(), -}); - -export type ModernDetailedTask = z.infer<typeof ModernDetailedTaskSchema>; - -/** `GetTaskResult = Result & DetailedTask`. Same fields we need as the task itself. */ -export const ModernGetTaskResultSchema = ModernDetailedTaskSchema; - -/** `CreateTaskResult = Result & Task` (`resultType: "task"`). The seed task state. */ -export const ModernCreateTaskResultSchema = ModernDetailedTaskSchema; - -/** `UpdateTaskResult` — an empty acknowledgement (`resultType: "complete"`). */ -export const ModernUpdateTaskResultSchema = z.looseObject({}); - -/** `CancelTaskResult` — modern cancel acks with an empty/loose result. */ -export const ModernCancelTaskResultSchema = z.looseObject({}); - -/** - * Normalize a modern `DetailedTask` onto the internal (SDK 2025-11-25) `Task` - * shape the state store, events, and `TaskCard` consume: `ttlMs` → `ttl`, - * `pollIntervalMs` → `pollInterval`. The status-specific members - * (`result` / `error` / `inputRequests`) ride along structurally so the poll - * driver can read them; they are not part of the `Task` type but are harmless - * extra properties on the object (the card renders the full task JSON). - */ -export function normalizeModernTask(modern: ModernDetailedTask): Task { - const { ttlMs, pollIntervalMs, ...rest } = modern; - const normalized: Record<string, unknown> = { ...rest }; - // The internal Task requires `ttl: number | null`; map the modern `ttlMs` - // (which is itself `number | null`) straight across, defaulting to null. - normalized.ttl = ttlMs ?? null; - if (pollIntervalMs != null) normalized.pollInterval = pollIntervalMs; - // The loose modern schema is a structural superset of the internal Task - // (taskId/status/statusMessage/createdAt/lastUpdatedAt present; ttl/pollInterval - // mapped above). No SDK schema relates the two nominal types, so a single - // narrowing cast bridges the structurally-identical shape. - return normalized as unknown as Task; -} - -/** Read the embedded `inputRequests` map off a modern task, typed for - * {@link fulfilInputRequests}. The loose parse yields `unknown` values; the - * per-request `fulfilEmbeddedInputRequest` switch validates each by method. */ -export function readInputRequests( - modern: ModernDetailedTask, -): InputRequests | undefined { - // Structural bridge: the loose record parse cannot express the InputRequests - // union, but fulfilEmbeddedInputRequest validates each entry by its `method`. - return modern.inputRequests as InputRequests | undefined; -} diff --git a/core/mcp/node/config.ts b/core/mcp/node/config.ts index 3f23913f0d..09ca1a2a78 100644 --- a/core/mcp/node/config.ts +++ b/core/mcp/node/config.ts @@ -9,8 +9,10 @@ import type { StreamableHttpServerConfig, } from "../types.js"; import { isProtocolEra, normalizeServerType } from "../serverList.js"; +import { isSkillCatalogLimit } from "../skills.js"; import type { ServerProtocolEra } from "../types.js"; import { toRecord } from "../../json/jsonUtils.js"; +import { getOwnEntry } from "../../storage/own-entry.js"; /** * Options object passed to resolveServerConfigs by runners (parsed from argv). @@ -101,6 +103,31 @@ export function parseProtocolEra(value: string): ServerProtocolEra { return value; } +/** + * Build the Commander coerce for a skills-catalog budget flag + * (`--skill-catalog-max-skills` / `--skill-catalog-max-bytes`), so the CLI and + * TUI can set the per-server budget the web exposes in Server Settings (#2420). + * Accepts only a plain run of decimal digits denoting a positive safe integer — + * the values {@link isSkillCatalogLimit} keeps when it reads `mcp.json` — so + * `0`, `-1`, `1.5`, `1e3` and `0x10` are rejected up front, rather than falling + * back to the default the way a hand-edited file value does. `flag` names the + * option in the error. Pure function; no Commander dependency. + */ +export function skillCatalogLimitParser( + flag: string, +): (value: string) => number { + return (value: string): number => { + const trimmed = value.trim(); + const n = /^\d+$/.test(trimmed) ? Number(trimmed) : NaN; + if (!isSkillCatalogLimit(n)) { + throw new Error( + `Invalid ${flag}: ${value}. Expected a positive integer written as plain decimal digits, at most ${Number.MAX_SAFE_INTEGER}.`, + ); + } + return n; + }; +} + /** On-disk contents of a freshly seeded empty catalog (pretty-printed). */ const EMPTY_CATALOG_CONTENT = `${JSON.stringify({ mcpServers: {} }, null, 2)}\n`; @@ -174,6 +201,11 @@ function loadMcpServersConfig( /** * Loads a single server config from an MCP config file by name. * Delegates to loadMcpServersConfig (file existence and type normalization are done there). + * + * The lookup is an own-property read (`getOwnEntry`): `serverName` is user + * input, and a bare `mcpServers[serverName]` resolves an absent `constructor` + * or `toString` to the inherited `Object.prototype` member, which then fails + * downstream with an unrelated error instead of "not found" (#2537). */ function loadServerFromConfig( configPath: string, @@ -181,13 +213,14 @@ function loadServerFromConfig( writable: boolean, ): MCPServerConfig { const config = loadMcpServersConfig(configPath, writable); - if (!config.mcpServers[serverName]) { + const server = getOwnEntry(config.mcpServers, serverName); + if (!server) { const available = Object.keys(config.mcpServers).join(", "); throw new Error( `Server '${serverName}' not found in config file. Available servers: ${available}`, ); } - return config.mcpServers[serverName]; + return server; } /** Build one MCPServerConfig from ad-hoc options (no config file). */ diff --git a/core/mcp/node/index.ts b/core/mcp/node/index.ts index f8fa91fcf1..339cbf6487 100644 --- a/core/mcp/node/index.ts +++ b/core/mcp/node/index.ts @@ -2,6 +2,7 @@ export { parseKeyValuePair, parseHeaderPair, parseProtocolEra, + skillCatalogLimitParser, withDefaultCatalogPath, resolveServerConfigs, getNamedServerConfigs, diff --git a/core/mcp/node/servers.ts b/core/mcp/node/servers.ts index 8b7a2e348e..356c8e1f82 100644 --- a/core/mcp/node/servers.ts +++ b/core/mcp/node/servers.ts @@ -22,6 +22,7 @@ import { type ServerConfigOptions, } from "./config.js"; import { rehydrateMcpConfigFromKeychain } from "./server-secrets.js"; +import { getOwnEntry } from "../../storage/own-entry.js"; /** * A server resolved from a catalog/config file or an ad-hoc target, paired with @@ -42,6 +43,10 @@ export type ServerLoadOptions = ServerConfigOptions & { /** `--protocol-era`: overrides the file's `protocolEra` (or the legacy * default) the way `--header` overrides its headers (#2208). */ protocolEra?: ServerProtocolEra; + /** `--skill-catalog-max-skills` / `--skill-catalog-max-bytes`: override the + * file's per-server skills catalog budget the same way (#2420). */ + skillCatalogMaxSkills?: number; + skillCatalogMaxBytes?: number; /** Test injection; defaults to the selected store (see * `secret-store-selection.ts`) for catalog/config loads. */ secretStore?: SecretStore; @@ -81,14 +86,17 @@ function defaultServerSettings(): InspectorServerSettings { } /** - * Overlay `--header` / `--protocol-era` onto the settings lifted from the file - * (or onto nothing, for an ad-hoc target). Only those two fields are overridden - * — timeouts, OAuth, and the rest of the file's settings are preserved. Returns - * `base` untouched when neither flag was given. + * Overlay `--header` / `--protocol-era` / `--skill-catalog-max-*` onto the + * settings lifted from the file (or onto nothing, for an ad-hoc target). Only + * those fields are overridden — timeouts, OAuth, and the rest of the file's + * settings are preserved. Returns `base` untouched when no such flag was given. */ function mergeSettings( base: InspectorServerSettings | undefined, - overrides: Pick<ServerLoadOptions, "headers" | "protocolEra">, + overrides: Pick< + ServerLoadOptions, + "headers" | "protocolEra" | "skillCatalogMaxSkills" | "skillCatalogMaxBytes" + >, ): InspectorServerSettings | undefined { let settings = base; const fromHeaders = headersToServerSettings(overrides.headers); @@ -103,6 +111,18 @@ function mergeSettings( protocolEra: overrides.protocolEra, }; } + if (overrides.skillCatalogMaxSkills !== undefined) { + settings = { + ...(settings ?? defaultServerSettings()), + skillCatalogMaxSkills: overrides.skillCatalogMaxSkills, + }; + } + if (overrides.skillCatalogMaxBytes !== undefined) { + settings = { + ...(settings ?? defaultServerSettings()), + skillCatalogMaxBytes: overrides.skillCatalogMaxBytes, + }; + } return settings; } @@ -144,10 +164,10 @@ export async function loadServerEntries( env: serverOptions.env, cwd: serverOptions.cwd, }), - // Deliberate broadcast: a single `--header` set (and `--protocol-era`) - // is merged into EVERY server in the catalog/config (fine for the - // common single-server case; for multi-server files, prefer per-server - // settings in the file itself). + // Deliberate broadcast: a single `--header` set (and `--protocol-era`, + // `--skill-catalog-max-*`) is merged into EVERY server in the + // catalog/config (fine for the common single-server case; for + // multi-server files, prefer per-server settings in the file itself). settings: mergeSettings(entry.settings, serverOptions), }; } @@ -177,6 +197,10 @@ export async function loadServerEntries( * in config file") because `entries` may come from a file *or* a single ad-hoc * target — e.g. `--server foo` alongside a positional command resolves to just * `{ default }`, where "in config file" would be misleading. + * + * The lookup is an own-property read (`getOwnEntry`), so `--server constructor` + * (or any other inherited `Object.prototype` name) absent from `entries` + * reports "not found" rather than returning the inherited member (#2537). */ export function selectServerEntry( entries: Record<string, ResolvedServer>, @@ -184,7 +208,7 @@ export function selectServerEntry( ): ResolvedServer { const names = Object.keys(entries); if (serverName) { - const entry = entries[serverName]; + const entry = getOwnEntry(entries, serverName); if (!entry) { throw new Error( `Server '${serverName}' not found. Available servers: ${names.join(", ")}`, diff --git a/core/mcp/resultFile.ts b/core/mcp/resultFile.ts new file mode 100644 index 0000000000..e924902f90 --- /dev/null +++ b/core/mcp/resultFile.ts @@ -0,0 +1,123 @@ +/** + * Rendering an MCP result for a file on disk — the pure half of "save this + * result to a file", shared so every client writes the same bytes for the same + * result (#2571). + * + * The two encodings are the ones designed for the CLI's planned `--output` / + * `--output-format raw|json` (#2431), whose issue said the rendering should + * move here once a second client needed it. The TUI's `w` keybinding on its + * tool result view (#2571) is, for now, the only in-tree consumer; #2431 is + * expected to render through this module when it lands rather than carry its + * own copy. Only the rendering lives in `core/`: the write itself is + * client-owned, because each client reports a failed write its own way (a CLI + * exit code, a TUI status line) and the web client downloads rather than writes. + * + * - `json` is the whole result, pretty-printed with two-space indentation and a + * trailing newline. + * - `raw` is the result's payload as a consumer would want it on disk: the text + * of its text-bearing blocks joined by newlines, or — when it carries no text + * and exactly one binary block — that block's decoded bytes, so an image or + * audio result saves as a playable file rather than as base64 inside JSON. + * + * Browser-safe on purpose (no `Buffer`): `core/mcp` is consumed by the web + * client too, so binary payloads decode through `atob` into a `Uint8Array`, + * which Node's `fs.writeFile` accepts as readily as a `Buffer`. + */ + +export type ResultFileFormat = "raw" | "json"; + +export const RESULT_FILE_FORMATS: readonly ResultFileFormat[] = ["raw", "json"]; + +/** The two result shapes `raw` knows how to extract a payload from. */ +export type ResultFileMethod = "tools/call" | "resources/read"; + +/** + * `raw` was asked of a result that has no single raw form: no text and either + * no binary block or several of them. Callers add their own remedy (the CLI + * points at `--output-format json`; the TUI offers the json format). + */ +export class RawResultUnavailableError extends Error { + readonly method: ResultFileMethod; + readonly binaryCount: number; + + constructor(method: ResultFileMethod, binaryCount: number) { + super( + binaryCount === 0 + ? `The ${method} result has no text or binary content to write as raw.` + : `The ${method} result has ${binaryCount} binary blocks and no text, so it has no single raw form.`, + ); + this.name = "RawResultUnavailableError"; + this.method = method; + this.binaryCount = binaryCount; + } +} + +type Payload = { text: string } | { binary: string }; + +function asRecord(value: unknown): Record<string, unknown> | undefined { + return value !== null && typeof value === "object" && !Array.isArray(value) + ? (value as Record<string, unknown>) + : undefined; +} + +/** + * One payload per block: its text, or its base64 binary. A `tools/call` block + * carries text on `text`, binary on `data` (image/audio), and an embedded + * resource nests either under `resource`; a `resources/read` entry carries + * `text` or `blob` directly. A `resource_link` carries neither and is skipped. + */ +function blockPayload(block: unknown): Payload | undefined { + const b = asRecord(block); + if (!b) return undefined; + const nested = asRecord(b.resource); + if (nested) return blockPayload(nested); + if (typeof b.text === "string") return { text: b.text }; + if (typeof b.data === "string") return { binary: b.data }; + if (typeof b.blob === "string") return { binary: b.blob }; + return undefined; +} + +function decodeBase64(value: string): Uint8Array { + const binary = atob(value); + const bytes = new Uint8Array(binary.length); + for (let i = 0; i < binary.length; i++) bytes[i] = binary.charCodeAt(i); + return bytes; +} + +/** + * The `raw` form of a result: its texts joined by newlines, or the decoded + * bytes of its single binary block when it has no text at all. Throws + * {@link RawResultUnavailableError} when neither applies. + */ +export function renderRawResult( + result: unknown, + method: ResultFileMethod, +): string | Uint8Array { + const record = asRecord(result); + const list = method === "resources/read" ? record?.contents : record?.content; + const payloads = (Array.isArray(list) ? list : []) + .map(blockPayload) + .filter((p): p is Payload => p !== undefined); + const texts = payloads.flatMap((p) => ("text" in p ? [p.text] : [])); + if (texts.length > 0) return texts.join("\n"); + const binaries = payloads.flatMap((p) => ("binary" in p ? [p.binary] : [])); + if (binaries.length === 1) return decodeBase64(binaries[0]!); + throw new RawResultUnavailableError(method, binaries.length); +} + +/** Render a result in the requested file format. */ +export function renderResultForFile( + result: unknown, + method: ResultFileMethod, + format: ResultFileFormat, +): string | Uint8Array { + if (format === "raw") return renderRawResult(result, method); + return JSON.stringify(result, null, 2) + "\n"; +} + +/** Byte length of rendered output, for a "wrote N bytes" confirmation. */ +export function renderedByteLength(data: string | Uint8Array): number { + return typeof data === "string" + ? new TextEncoder().encode(data).length + : data.length; +} diff --git a/core/mcp/serverList.ts b/core/mcp/serverList.ts index 9814008ed5..1849c2ff4e 100644 --- a/core/mcp/serverList.ts +++ b/core/mcp/serverList.ts @@ -7,6 +7,7 @@ import { DEFAULT_CONNECTION_TIMEOUT_MS, + DEFAULT_ELICIT_CAPABILITY, DEFAULT_MAX_FETCH_REQUESTS, DEFAULT_MODERN_LOG_LEVEL, DEFAULT_PROTOCOL_ERA, @@ -20,6 +21,7 @@ import { } from "./skills.js"; import type { Root } from "@modelcontextprotocol/client"; import type { + ElicitCapabilityMode, InspectorServerSettings, RequestMetadata, MCPConfig, @@ -52,6 +54,28 @@ const VALID_PROTOCOL_ERAS: ReadonlySet<ServerProtocolEra> = new Set([ "modern", ]); +const VALID_ELICIT_CAPABILITIES: ReadonlySet<ElicitCapabilityMode> = new Set([ + "off", + "url", + "form", + "both", +]); + +/** + * Runtime guard for the `elicitCapability` literal, mirroring + * {@link isProtocolEra}: a hand-edited `mcp.json` read directly by the CLI/TUI + * can carry any string, and an unknown value should read back as the default + * rather than propagate to `createSessionClient`. + */ +export function isElicitCapability( + value: unknown, +): value is ElicitCapabilityMode { + return ( + typeof value === "string" && + VALID_ELICIT_CAPABILITIES.has(value as ElicitCapabilityMode) + ); +} + /** * Runtime guard for the `protocolEra` literal. `StoredMCPServer` types the * field as `ServerProtocolEra`, but a hand-edited `mcp.json` read directly by @@ -157,6 +181,7 @@ type StoredInspectorFields = Pick< | "headers" | "metadata" | "protocolEra" + | "elicitCapability" | "modernLogLevel" | "connectionTimeout" | "requestTimeout" @@ -535,6 +560,7 @@ export function storedFieldsToInspectorSettings( stored.oauth !== undefined || stored.roots !== undefined || stored.protocolEra !== undefined || + stored.elicitCapability !== undefined || stored.modernLogLevel !== undefined || stored.env !== undefined || stored.cwd !== undefined; @@ -590,6 +616,12 @@ export function storedFieldsToInspectorSettings( if (isProtocolEra(stored.protocolEra)) { settings.protocolEra = stored.protocolEra; } + // Like `protocolEra`: absent reads back as the default elicitation + // capability (`"both"`), and an unknown literal from a hand-edited file is + // dropped rather than surfaced. + if (isElicitCapability(stored.elicitCapability)) { + settings.elicitCapability = stored.elicitCapability; + } // Like `protocolEra`: absent reads back as the default modern log level (the // form defaults via `?? DEFAULT_MODERN_LOG_LEVEL`), and an unknown literal from // a hand-edited file is dropped rather than surfaced. @@ -749,6 +781,17 @@ export function inspectorSettingsToStoredFields( out.protocolEra = settings.protocolEra; } + // Persist only when it differs from the default elicitation capability; + // absent reads back as DEFAULT_ELICIT_CAPABILITY, so writing the default + // would inject the field into hand-edited files that never had it and break + // byte-stable round-trips. + if ( + settings.elicitCapability !== undefined && + settings.elicitCapability !== DEFAULT_ELICIT_CAPABILITY + ) { + out.elicitCapability = settings.elicitCapability; + } + // Persist only when it differs from the default modern log level; absent reads // back as DEFAULT_MODERN_LOG_LEVEL, so writing the default would inject the // field into files that never set it and break byte-stable round-trips. @@ -854,6 +897,7 @@ const INSPECTOR_FIELD_KEY_MAP = { headers: true, metadata: true, protocolEra: true, + elicitCapability: true, modernLogLevel: true, connectionTimeout: true, requestTimeout: true, diff --git a/core/mcp/skillsVerification.ts b/core/mcp/skillsVerification.ts index e497970c61..bb91e7f1f0 100644 --- a/core/mcp/skillsVerification.ts +++ b/core/mcp/skillsVerification.ts @@ -386,8 +386,18 @@ export async function verifySkills( // running total would CROSS the limit, so a conforming skill (≤ 16 MiB in // total, by definition) is never truncated. const manifest = withinBudget ? boundedManifest(declared) : []; + // ⚠️ The suggested command carries `--verify`. It named only + // `--method skills/get --uri`, which on its own prints the skill and + // checks nothing — so the escape hatch this message advertises returned + // no verdict at all (#2428). `skills-verify-cli.test.ts` runs the command + // back through the CLI. + // + // ⚠️ The URI is a `<uri>` placeholder, never `entry.uri` spliced in. The + // URI is server-controlled and shell metacharacters are legal in one, so + // a command built from it is a command a server wrote for the user to + // paste (Copilot). The report already carries the URI in its own `uri`. let incomplete = !withinBudget - ? `Not read: this run already reached its catalog budget of ${budget.maxSkills} skills / ${budget.maxBytes} bytes (raise it in the server's Skills settings). Nothing about this skill's files has been checked — verify it on its own with \`--method skills/get --uri\` to get a verdict.` + ? `Not read: this run already reached its catalog budget of ${budget.maxSkills} skills / ${budget.maxBytes} bytes (raise it in the server's Skills settings). Nothing about this skill's files has been checked — verify it on its own with \`--method skills/get --uri <uri> --verify\`, where <uri> is this report's uri, to get a verdict.` : manifest.length < declared.length ? `Only ${manifest.length} of ${declared.length} manifest entries were read: the skill exceeds the ${SKILL_MAX_RESOURCE_ENTRIES}-entry / ${SKILL_MAX_TOTAL_BYTES}-byte interoperability limits, so the rest were not fetched and cannot be reported on.` : undefined; diff --git a/core/mcp/state/managedRequestorTasksState.ts b/core/mcp/state/managedRequestorTasksState.ts index 8b83585c4b..d8a8499061 100644 --- a/core/mcp/state/managedRequestorTasksState.ts +++ b/core/mcp/state/managedRequestorTasksState.ts @@ -55,11 +55,19 @@ export class ManagedRequestorTasksState extends TypedEventTarget<ManagedRequesto constructor(client: InspectorClientProtocol) { super(); this.client = client; + // An event listener cannot await, so the automatic refreshes own their + // rejection here: a failed `tasks/list` (server error, auth challenge, the + // page cap) would otherwise escape as an unhandled rejection, which kills a + // Node host such as the TUI. The list keeps its last committed value; a + // caller-initiated `refresh()` still rejects to its caller. + const refreshQuietly = (): void => { + this.refresh().catch(() => {}); + }; const onConnect = (): void => { - void this.refresh(); + refreshQuietly(); }; const onTasksListChanged = (): void => { - void this.refresh(); + refreshQuietly(); }; const onStatusChange = (): void => { if (isTerminalStatus(this.client?.getStatus())) { @@ -144,21 +152,20 @@ export class ManagedRequestorTasksState extends TypedEventTarget<ManagedRequesto if (!client || client.getStatus() !== "connected") { return this.getTasks(); } - // Modern era (SEP-2663): the `io.modelcontextprotocol/tasks` extension has - // NO `tasks/list` — task handles are durable and client-held, arriving via - // task-augmented tool calls (server-directed `CreateTaskResult`) or - // unsolicited handles. "Refresh" therefore re-polls the tasks already known - // to this store via `tasks/get`; a task the server has since dropped (TTL) - // simply keeps its last-seen state. Gate on the extension being negotiated. - if (client.isTasksExtensionNegotiated()) { + // Routing is gated on the task session's authoritative `inventory` + // capability rather than era/extension predicates: a legacy Tasks session + // without `tasks/list` (inventory "known-handles") must re-poll known + // handles like a modern one — sending it through `listRequestorTasks()` + // would be rejected by ext-tasks as unsupported server inventory. + const inventory = + client.getTaskSessionCapabilities()?.inventory ?? "unsupported"; + if (inventory === "known-handles") { return this.refreshModern(client); } - // Legacy era: gate on the server's `tasks` capability — calling tasks/list - // against a server that doesn't advertise it returns -32601 "Method not - // found", which then surfaces in the console for every connect against any - // server that doesn't implement task tracking. Empty list is the right - // semantics for "this server doesn't support tasks." - if (!client.getCapabilities()?.tasks) { + // No task support at all: empty list is the right semantics for "this + // server doesn't support tasks", and calling tasks/list against it + // returns -32601 "Method not found" in the console on every connect. + if (inventory !== "server-list") { this.tasks = []; this.dispatchTypedEvent("tasksChange", this.tasks); return this.getTasks(); diff --git a/core/mcp/types.ts b/core/mcp/types.ts index 2e89d2f7b0..944caa4cb7 100644 --- a/core/mcp/types.ts +++ b/core/mcp/types.ts @@ -22,6 +22,7 @@ import type { import type { Client } from "@modelcontextprotocol/client"; import type { OAuthClientProvider } from "@modelcontextprotocol/client"; import type { Transport } from "@modelcontextprotocol/client"; +import type { TaskView } from "@modelcontextprotocol/ext-tasks/client"; import type { InspectorLogger } from "../logging/logger.js"; import type { AppElicitationRenderer } from "./appElicitation.js"; import type { @@ -40,6 +41,25 @@ import type { import type { OAuthStorage } from "../auth/storage.js"; import type { AuthChallenge } from "../auth/challenge.js"; +/** + * Requester task shape rendered by Inspector surfaces: a narrow normalized + * overlay of the ext-tasks `TaskView`. Shared fields (including the + * generation-neutral status union) are picked from the SDK type so the two + * contracts cannot drift; the overlay exists only because Inspector state + * stores require concrete timestamps (`TaskView` leaves them optional) and + * an unbranded task id. + */ +export interface InspectorTask extends Pick< + TaskView, + "status" | "statusMessage" | "ttl" | "pollInterval" +> { + taskId: string; + createdAt: string; + lastUpdatedAt: string; + /** Original generation-specific task payload, preserved without type claims. */ + raw?: Readonly<Record<string, unknown>>; +} + // Stdio transport config export interface StdioServerConfig { // Optional: stdio is the implicit default when `type` is absent. A @@ -125,6 +145,13 @@ export type StoredMCPServer = MCPServerConfig & { * (`"legacy"`). (#1626) */ protocolEra?: ServerProtocolEra; + /** + * Elicitation capability this client advertises to this server + * (`"off" | "url" | "form" | "both"`). Inspector-specific (no analog in the + * broader mcp.json ecosystem). Omitted on disk when it equals the default + * (`"both"`). Currently consumed by mcpdo only. (#1783) + */ + elicitCapability?: ElicitCapabilityMode; /** * Modern-era per-request log level stamped by default (`"off"` or one of the * eight logging levels). Inspector-specific. Omitted on disk when it equals @@ -730,6 +757,15 @@ export type ServerProtocolEra = "legacy" | "auto" | "modern"; /** The default per-server protocol era when none is configured. */ export const DEFAULT_PROTOCOL_ERA: ServerProtocolEra = "legacy"; +/** + * Elicitation capability mode a client advertises to a server for one + * connection — see {@link InspectorServerSettings.elicitCapability}. + */ +export type ElicitCapabilityMode = "off" | "url" | "form" | "both"; + +/** The default elicitation capability mode when none is configured. */ +export const DEFAULT_ELICIT_CAPABILITY: ElicitCapabilityMode = "both"; + /** * Per-server modern (2026-07-28) per-request log level (#1629). `logging/setLevel` * is gone on the modern era; instead the client opts into logs by stamping @@ -988,6 +1024,22 @@ export interface InspectorServerSettings { * omitted when it equals the default, keeping the file diff minimal. */ protocolEra?: ServerProtocolEra; + /** + * Elicitation capability this client advertises to the server for this + * connection: `"off"` (no `capabilities.elicitation` at all — the server + * sees a client that can't do elicitation and can fall back to whatever + * it does when the capability is absent, e.g. proceeding with defaults or + * failing its own way, rather than getting a guaranteed decline/cancel), + * `"url"` (URL-mode only), `"form"` (form-mode only), or `"both"`. Optional + * so a bare settings node reads back without one; absence means {@link + * DEFAULT_ELICIT_CAPABILITY} (`"both"`). Persisted on disk as + * `elicitCapability` and omitted when it equals the default. Currently + * consumed by mcpdo only (#1783) — a connect-time, sticky-per-connection + * choice rather than a per-call one, since a daemon-managed session can be + * reused by several later callers (interactive and scripted) over its + * lifetime. + */ + elicitCapability?: ElicitCapabilityMode; /** * Modern-era per-request log level stamped by default on this server's * connections (#1629). One of the eight logging levels, or `"off"` to not opt diff --git a/core/mcp/uriTemplate.ts b/core/mcp/uriTemplate.ts index a536dde397..b61dbfeb84 100644 --- a/core/mcp/uriTemplate.ts +++ b/core/mcp/uriTemplate.ts @@ -10,7 +10,7 @@ * * The **CLI is deliberately not a consumer**: it has no template form, and its * `resources/read` passes the already-expanded `--uri` straight to - * `readResource` (see `clients/cli/src/handlers/run-method.ts`). Nothing here + * `readResource` (see `core/cli/handlers/run-method.ts`). Nothing here * runs for it. * * ## Why this is not simply `new UriTemplate(t).expand(v)` diff --git a/clients/cli/src/open-url.ts b/core/node/openUrl.ts similarity index 85% rename from clients/cli/src/open-url.ts rename to core/node/openUrl.ts index 2e9319780f..4d3eaee6eb 100644 --- a/clients/cli/src/open-url.ts +++ b/core/node/openUrl.ts @@ -1,3 +1,14 @@ +/** + * The one browser-opener wrapper every Node client uses: the CLI's OAuth + * navigation, the TUI's OAuth navigation, and the web backend's `autoOpen` + * (#2533). It started life in the CLI (#2410, #2531); the TUI and the web + * launcher had their own bare `open(...)` calls with the same spawn-failure + * crash, so it lives here to keep three copies from drifting apart again. + * + * **Node-only** — `open` spawns a child process; never import it from browser + * code. + */ + import type { ChildProcess } from "node:child_process"; import open from "open"; diff --git a/docs/ai-software-factory.md b/docs/ai-software-factory.md index 1007ace158..99e4a3d675 100644 --- a/docs/ai-software-factory.md +++ b/docs/ai-software-factory.md @@ -159,7 +159,9 @@ rules in `AGENTS.md` cover: - **Branch names carry the target version first** — `v2/fix/2071-…`, `v1/fix/…` — cut from the matching `*/main`. -- **DCO signoff is a hard merge gate** (`git commit -s`). +- **DCO signoff is checked on every v2 PR** (`git commit -s`) by a repo-owned + workflow (`.github/workflows/dco.yml`), which replaced the suspended probot + DCO app. - **UI changes require before/after screenshots**, staged in a gitignored `pr-screenshots/` folder. - **A code review is requested and answered per-thread**, with a PR-level diff --git a/docs/cli-smoke-testing.md b/docs/cli-smoke-testing.md index 62039c82c3..4d833a388a 100644 --- a/docs/cli-smoke-testing.md +++ b/docs/cli-smoke-testing.md @@ -357,6 +357,12 @@ npx @modelcontextprotocol/inspector --cli --server-url "$SERVER_URL" --list-stor # → {"oauthStatePath":"/tmp/tmp.XXXX/oauth.json","storedServerUrls":[]} ``` +The secret store needs no equivalent isolation: each state file's store +entries are scoped by a namespace stamped into the file itself, so a smoke +run's tokens and the developer's real keychain entries for the same server +URL never share a slot (see [Where secrets are +stored](./secret-storage.md)). + For a server that genuinely needs a credential in CI, prefer a static header over OAuth entirely — `--header 'Authorization: Bearer <token>'`, with the token from your CI secret store. And **do not** put a credential in the URL: the CLI diff --git a/docs/docker.md b/docs/docker.md index 83b66bcbae..8115d84b7a 100644 --- a/docs/docker.md +++ b/docs/docker.md @@ -103,5 +103,5 @@ docker run --rm -u 0 --entrypoint chown \ -R node:node /data ``` -The image defaults to `--web` bound to `0.0.0.0:6274` with browser auto-open disabled; override the args to run another mode (`docker run --rm ghcr.io/modelcontextprotocol/inspector --cli …`). Pass `-e MCP_INSPECTOR_API_TOKEN=…` to set a known token (otherwise one is generated and printed in the logs), or `-e DANGEROUSLY_OMIT_AUTH=true` to disable auth. Binding `0.0.0.0` (all network interfaces) is refused by default outside a container — it exposes the process-spawning backend to the local network — so the image opts in explicitly with `DANGEROUSLY_BIND_ALL_INTERFACES=true` (already set in the `Dockerfile`); a bare `HOST=0.0.0.0` without that flag exits with an error. If you **remap the published port** (`-p 127.0.0.1:8080:6274`), the browser's origin (`http://localhost:8080`) no longer matches the in-container port, so set `-e ALLOWED_ORIGINS=http://localhost:8080,http://127.0.0.1:8080` (or run `-e CLIENT_PORT=8080 -p 127.0.0.1:8080:8080`) or connects will 403. `ALLOWED_ORIGINS` **replaces** the default list rather than merging, so list every loopback form you'll browse from (see the [web README](../clients/web/README.md#host-binding--the-origin-allow-list)). The image runs as the non-root `node` user and has a `HEALTHCHECK` that probes the web UI at the address `HOST` binds (a wildcard such as `0.0.0.0` is probed on loopback), so pinning `-e HOST=` to one interface keeps it valid. `--cli` and `--tui` have no web server, so the probe detects those modes from the container's arguments and reports the container healthy while it runs, with no `--no-healthcheck` needed. +The image defaults to `--web` bound to `0.0.0.0:6274` with browser auto-open disabled; override the args to run another mode (`docker run --rm ghcr.io/modelcontextprotocol/inspector --cli …`). Pass `-e MCP_INSPECTOR_API_TOKEN=…` to set a known token (otherwise one is generated and printed in the logs), or `-e DANGEROUSLY_OMIT_AUTH=true` to disable auth. Binding `0.0.0.0` (all network interfaces) is refused by default outside a container — it exposes the process-spawning backend to the local network — so the image opts in explicitly with `DANGEROUSLY_BIND_ALL_INTERFACES=true` (already set in the `Dockerfile`); a bare `HOST=0.0.0.0` without that flag exits with an error. If you **remap the published port** (`-p 127.0.0.1:8080:6274`), the browser's origin (`http://localhost:8080`) no longer matches the in-container port, so set `-e ALLOWED_ORIGINS=http://localhost:8080,http://127.0.0.1:8080` (or run `-e CLIENT_PORT=8080 -p 127.0.0.1:8080:8080`) or connects will 403. `ALLOWED_ORIGINS` **replaces** the default list rather than merging, so list every loopback form you'll browse from (see the [web README](../clients/web/README.md#host-binding--the-origin-allow-list)). The image runs as the non-root `node` user and has a `HEALTHCHECK` that probes the web UI at the address `HOST` binds (a wildcard such as `0.0.0.0` is probed on loopback), so pinning `-e HOST=` to one interface keeps it valid. `--cli` and `--tui` have no web server, so the probe detects those modes from the container's arguments and reports the container healthy while it runs, with no `--no-healthcheck` needed. An external orchestrator — a Kubernetes liveness/readiness probe, a Compose `healthcheck` — can probe `GET /healthz` on the web port: it needs no token and returns only `{"status":"ok"}` (see [Health check](../clients/web/README.md#health-check)). diff --git a/docs/environment-variables.md b/docs/environment-variables.md index 0acc315c16..392abcbfdd 100644 --- a/docs/environment-variables.md +++ b/docs/environment-variables.md @@ -2,7 +2,7 @@ Every environment variable that changes how the Inspector behaves at runtime, in one place: the Inspector's own variables, plus the standard system and Node variables it reads (`HOME`, the proxy variables, the Node TLS variables). Set them in the shell that launches `mcp-inspector` (or with `-e` for the [Docker image](./docker.md)). -The **Read by** column names the client whose process reads the variable: **web** is the Node backend that `--web` starts (the browser itself reads no environment), **CLI** and **TUI** are those clients, and **launcher** is the `mcp-inspector` bin that picks one of them. A variable read in shared `core/` code is marked with every client that reaches it. +The **Read by** column names the client whose process reads the variable: **web** is the Node backend that `--web` starts (the browser itself reads no environment), **CLI**, **TUI** and **mcpdo** (the connection CLI, and the daemon it starts) are those clients, and **launcher** is the `mcp-inspector` bin that picks one of them. A variable read in shared `core/` code is marked with every client that reaches it. ⚠️ **Unset a variable rather than setting it to an empty string.** The two are not interchangeable: `HOST=""` is read as an all-interfaces bind and refused, and an empty path variable such as `MCP_STORAGE_DIR=` or `MCP_INSPECTOR_LOG_DIR=` can resolve relative to the working directory instead of falling back to the default. A row says so explicitly where an empty value is treated as unset. @@ -12,22 +12,22 @@ A `~` in a default below means the home directory as described under [Home direc These guard the web backend, which spawns processes on request. Read [Host binding and the origin allow-list](../clients/web/README.md#host-binding--the-origin-allow-list) before widening any of them. -| Variable | Read by | Default | Effect | -| --------------------------------- | -------- | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `MCP_INSPECTOR_API_TOKEN` | web, CLI | a random token per launch | Bearer token guarding every `/api/*` route (`x-mcp-remote-auth: Bearer <token>`). Set it to use a known token instead of the generated one printed in the launch banner. The CLI reads it only to fill the `autoConnect` parameter of the deep link it emits. | -| `MCP_PROXY_AUTH_TOKEN` | web | — | **Deprecated** v1 name for `MCP_INSPECTOR_API_TOKEN`, used only when the new name is unset. | -| `DANGEROUSLY_OMIT_AUTH` | web | unset | Disables the API token entirely when set to `true` or `1` (trimmed, case-insensitive). Any other value — including `false`, `0` and empty — keeps auth on. | -| `HOST` | web, CLI | `127.0.0.1` | Address the web server binds. An all-interfaces host (`0.0.0.0`, `::`, an empty string, and equivalent spellings) is **refused** unless `DANGEROUSLY_BIND_ALL_INTERFACES` is enabled. The CLI reads it only to build its deep link. | -| `DANGEROUSLY_BIND_ALL_INTERFACES` | web | off | Opts in to an all-interfaces `HOST`. Only `true` or `1` (case-insensitive) enable it, so `false` reads as off. The Docker image sets it. | +| Variable | Read by | Default | Effect | +| --------------------------------- | -------- | ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `MCP_INSPECTOR_API_TOKEN` | web, CLI | a random token per launch | Bearer token guarding every `/api/*` route (`x-mcp-remote-auth: Bearer <token>`). Set it to use a known token instead of the generated one printed in the launch banner. The CLI reads it only to fill the `autoConnect` parameter of the deep link it emits. | +| `MCP_PROXY_AUTH_TOKEN` | web | — | **Deprecated** v1 name for `MCP_INSPECTOR_API_TOKEN`, used only when the new name is unset. | +| `DANGEROUSLY_OMIT_AUTH` | web | unset | Disables the API token entirely when set to `true` or `1` (trimmed, case-insensitive). Any other value — including `false`, `0` and empty — keeps auth on. | +| `HOST` | web, CLI | `127.0.0.1` | Address the web server binds. An all-interfaces host (`0.0.0.0`, `::`, an empty string, and equivalent spellings) is **refused** unless `DANGEROUSLY_BIND_ALL_INTERFACES` is enabled. The CLI reads it only to build its deep link. | +| `DANGEROUSLY_BIND_ALL_INTERFACES` | web | off | Opts in to an all-interfaces `HOST`. Only `true` or `1` (case-insensitive) enable it, so `false` reads as off. The Docker image sets it. | | `ALLOWED_ORIGINS` | web | derived from `HOST` | Comma-separated origins allowed to call the API. Unset, the list follows `HOST` at `CLIENT_PORT`: the loopback origins for a loopback host, the loopback origins plus `http://0.0.0.0` and `http://[::]` for an all-interfaces bind, and otherwise only the configured host's own origin (so binding a LAN address does **not** also allow `localhost`). **Replaces** the default list rather than adding to it, so list every form you browse from. Each entry must include the scheme (`http://localhost:6274`). The same list is the MCP Apps sandbox proxy's embedder allow-list (its `frame-ancestors` header and its referrer check), so a public Inspector origin must be listed here for the Apps tab to render. | ## Ports -| Variable | Read by | Default | Effect | -| --------------------- | -------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Variable | Read by | Default | Effect | +| --------------------- | -------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `CLIENT_PORT` | web, CLI | `6274` | Web UI port. Must be a **fixed** integer in 1–65535: `0` (an OS-assigned port) is rejected at startup, because the origin allow-list and the MCP Apps sandbox CSP are derived from it. | -| `MCP_SANDBOX_PORT` | web, CLI | `6275` | Port of the MCP Apps sandbox server, 0–65535. `0` asks the OS for a free port. An invalid value is ignored with a warning. | -| `SERVER_PORT` | web | — | v1's proxy port, now only a fallback for the sandbox port: used whenever `MCP_SANDBOX_PORT` does not yield a valid port — unset, empty, **or invalid**. | +| `MCP_SANDBOX_PORT` | web, CLI | `6275` | Port of the MCP Apps sandbox server, 0–65535. `0` asks the OS for a free port. An invalid value is ignored with a warning. | +| `SERVER_PORT` | web | — | v1's proxy port, now only a fallback for the sandbox port: used whenever `MCP_SANDBOX_PORT` does not yield a valid port — unset, empty, **or invalid**. | | `MCP_APP_ORIGIN_PORT` | web, CLI | `6278` | Port of the dedicated app-origin server, used only by an MCP App whose UI resource declares `_meta.ui.domain`; 0–65535, where `0` asks the OS for a free port. An invalid value is ignored with a warning. Pin it if your app's backend allowlists that origin. | The sandbox port resolves as: a valid `MCP_SANDBOX_PORT`, else a valid `SERVER_PORT`, else `6275`. The CLI reads `CLIENT_PORT`, `MCP_SANDBOX_PORT`, `MCP_APP_ORIGIN_PORT` and `HOST` only to build the deep link and port list it hands to a web session; it binds none of them. It normalizes `HOST` (an all-interfaces host becomes `localhost`, any other host is canonicalized), but it validates none of the three **port** variables: any non-empty port value is copied into the URLs and port-forwarding command as-is, with no range check, no warning, and no `SERVER_PORT` fallback, so a malformed value produces a broken hand-off rather than an error. @@ -36,38 +36,48 @@ The sandbox port resolves as: a valid `MCP_SANDBOX_PORT`, else a valid `SERVER_P The sandbox and app-origin servers advertise a URL built from their own bind (`http://localhost:6275/sandbox` under a wildcard bind). Behind a reverse proxy or an ingress, where the browser reaches them at a public hostname, set the address the browser should use instead. Neither changes what is bound: the process still listens on the port above, and routing the public address to it is the proxy's job. See [Behind a reverse proxy](../clients/web/README.md#host-binding--the-origin-allow-list). -| Variable | Read by | Default | Effect | -| ----------------------------- | ------- | --------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `MCP_SANDBOX_FULL_ADDRESS` | web | `http://<bind host>:<sandbox port>/sandbox` | Public URL of the MCP Apps sandbox proxy, returned as `sandboxUrl` by `GET /api/config` and printed in the banner (e.g. `https://inspector-sandbox.example.com/sandbox`). A bare origin gets `/sandbox` appended. An empty value counts as unset. | -| `MCP_APP_ORIGIN_FULL_ADDRESS` | web | `http://<bind host>:<app-origin port>` | Public origin that `_meta.ui.domain` app documents are published under (e.g. `https://inspector-apps.example.com`). **Origin only** — a path is refused, since documents are served at `<origin>/app-document/<id>`. An empty value counts as unset. | +| Variable | Read by | Default | Effect | +| ----------------------------- | ------- | ------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `MCP_SANDBOX_FULL_ADDRESS` | web | `http://<bind host>:<sandbox port>/sandbox` | Public URL of the MCP Apps sandbox proxy, returned as `sandboxUrl` by `GET /api/config` and printed in the banner (e.g. `https://inspector-sandbox.example.com/sandbox`). A bare origin gets `/sandbox` appended. An empty value counts as unset. | +| `MCP_APP_ORIGIN_FULL_ADDRESS` | web | `http://<bind host>:<app-origin port>` | Public origin that `_meta.ui.domain` app documents are published under (e.g. `https://inspector-apps.example.com`). **Origin only** — a path is refused, since documents are served at `<origin>/app-document/<id>`. An empty value counts as unset. | Both are **refused** — ignored with a warning, keeping the bind-derived address — when the value is not an absolute `http(s)` URL, carries credentials, a query string, a fragment or a wildcard, or is a bracketed IPv6 literal. ⚠️ **Each needs its own origin.** The MCP Apps spec requires the sandbox origin to differ from the Inspector's, so a sandbox address sharing an `ALLOWED_ORIGINS` origin is refused (`https://inspector.example.com/sandbox` behind the same hostname as the UI is the common case), as is an app origin equal to the Inspector's or the sandbox's. Neither is used unless its listener is on a fixed port — not when the port is `0`, and not when the pinned port was taken at startup (or collided with another Inspector port) and the server fell back to an OS-assigned one — since the proxy has no stable port to route it to. A plain-`http` address while `ALLOWED_ORIGINS` lists an `https` origin is used, but warned about: the browser blocks it as mixed content. ## Behavior -| Variable | Read by | Default | Effect | -| ------------------------ | ------------- | -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `MCP_AUTO_OPEN_ENABLED` | web, CLI | see effect | Whether a browser is opened for you. `false` never opens one; `true` always does. Unset (or any other value), the web client opens the UI at launch, and the CLI opens the OAuth authorization page only when stderr is a TTY. In the CLI, `true` also lets interactive OAuth **start** when neither stdin nor stderr is a TTY; otherwise that case fails with an auth-required error pointing at `--stored-auth-only`. **The TUI does not read it** and always opens the OAuth page. | -| `MCP_CATALOG_PATH` | web, CLI, TUI | `~/.mcp-inspector/mcp.json` | Default writable catalog, used when no `--catalog` is given. The CLI honors it only when no ad-hoc target (positional command, `--server-url`, or `--transport`) is given. See [MCP server configuration](./mcp-server-configuration.md). | -| `MCP_OAUTH_CALLBACK_URL` | CLI, TUI | `http://127.0.0.1:6276/oauth/callback` | Loopback redirect URL for the CLI/TUI OAuth flow. `--callback-url` takes precedence. | -| `NO_COLOR` | CLI | unset | Any non-empty value disables ANSI styling in the CLI's human-readable output. An empty `NO_COLOR=` counts as unset. | +| Variable | Read by | Default | Effect | +| ------------------------ | -------------------- | -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `MCP_AUTO_OPEN_ENABLED` | web, CLI, mcpdo | see effect | Whether a browser is opened for you. `false` never opens one; `true` always does. Unset (or any other value), the web client opens the UI at launch, and the CLI opens the OAuth authorization page only when stderr is a TTY. In the CLI, `true` also lets interactive OAuth **start** when neither stdin nor stderr is a TTY; otherwise that case fails with an auth-required error pointing at `--stored-auth-only`. mcpdo differs there: with no TTY on stdin or stderr and no `--stored-auth-only`, `connect` exits 0 with a pending connection and an `authUrl` to relay, and `true` keeps the blocking interactive flow instead. **The TUI does not read it** and always opens the OAuth page. | +| `MCP_CATALOG_PATH` | web, CLI, TUI, mcpdo | `~/.mcp-inspector/mcp.json` | Default writable catalog, used when no `--catalog` is given. The CLI honors it only when no ad-hoc target (positional command, `--server-url`, or `--transport`) is given. See [MCP server configuration](./mcp-server-configuration.md). | +| `MCP_OAUTH_CALLBACK_URL` | CLI, TUI, mcpdo | `http://127.0.0.1:6276/oauth/callback` | Loopback redirect URL for the CLI, TUI and mcpdo OAuth flow. In the CLI and TUI, `--callback-url` takes precedence; mcpdo has no such flag and reads only this variable. | +| `NO_COLOR` | CLI, mcpdo | unset | Any non-empty value disables ANSI styling in the CLI's human-readable output. An empty `NO_COLOR=` counts as unset. | ## Storage and state -| Variable | Read by | Default | Effect | -| -------------------------------- | ------------- | -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `MCP_STORAGE_DIR` | web, CLI, TUI | `~/.mcp-inspector/storage` | Storage directory. Relocates the OAuth state file (`oauth.json`) and the secrets file (`secrets.json`) for every client. For the **web** backend it also relocates `client.json`; the CLI and TUI find `client.json` through `MCP_CLIENT_CONFIG_PATH` instead. | -| `MCP_INSPECTOR_OAUTH_STATE_PATH` | CLI, TUI | `~/.mcp-inspector/storage/oauth.json` | Names the OAuth state file outright. Lookup order: this variable, then `<MCP_STORAGE_DIR>/oauth.json`, then `~/.mcp-inspector/storage/oauth.json`. ⚠️ Setting `MCP_STORAGE_DIR` alone does not isolate a CLI or TUI run if this variable is also exported. **The web backend does not read it** — it always uses `<MCP_STORAGE_DIR>/oauth.json`. | -| `MCP_CLIENT_CONFIG_PATH` | CLI, TUI | `~/.mcp-inspector/storage/client.json` | Install-level client config (CIMD, enterprise IdP). `--client-config` takes precedence. | +| Variable | Read by | Default | Effect | +| -------------------------------- | -------------------- | -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `MCP_STORAGE_DIR` | web, CLI, TUI, mcpdo | `~/.mcp-inspector/storage` | Storage directory. Relocates the OAuth state file (`oauth.json`) and the secrets file (`secrets.json`) for every client. For the **web** backend it also relocates `client.json`; the CLI and TUI find `client.json` through `MCP_CLIENT_CONFIG_PATH` instead. | +| `MCP_INSPECTOR_OAUTH_STATE_PATH` | CLI, TUI, mcpdo | `~/.mcp-inspector/storage/oauth.json` | Names the OAuth state file outright. Lookup order: this variable, then `<MCP_STORAGE_DIR>/oauth.json`, then `~/.mcp-inspector/storage/oauth.json`. ⚠️ Setting `MCP_STORAGE_DIR` alone does not isolate a CLI or TUI run if this variable is also exported. **The web backend does not read it** — it always uses `<MCP_STORAGE_DIR>/oauth.json`. Each state file keeps its own secret-store entries — they are scoped by a namespace stamped into the file — so per-profile state paths stay isolated even on a shared keychain (see [Where secrets are stored](./secret-storage.md)). | +| `MCP_CLIENT_CONFIG_PATH` | CLI, TUI, mcpdo | `~/.mcp-inspector/storage/client.json` | Install-level client config (CIMD, enterprise IdP). In the CLI and TUI, `--client-config` takes precedence; mcpdo has no such flag and reads only this variable. | ### Home directory Every default above that starts with `~` is built from the home directory the process sees, not from the OS account database: -| Variable | Read by | Effect | -| ------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| `HOME` | web, CLI, TUI | Base of `~/.mcp-inspector`: the default catalog, storage directory (`oauth.json`, `client.json`), secrets file, and TUI log directory. | -| `USERPROFILE` | web, CLI, TUI | Used in place of `HOME` when `HOME` is unset or empty — the normal case on Windows. ⚠️ If **neither** is set, as under some service managers, those defaults resolve against the **current working directory** instead. Set `HOME` or the specific path variables above. | +| Variable | Read by | Effect | +| ------------- | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `HOME` | web, CLI, TUI, mcpdo | Base of `~/.mcp-inspector`: the default catalog, storage directory (`oauth.json`, `client.json`), secrets file, and TUI log directory. | +| `USERPROFILE` | web, CLI, TUI, mcpdo | Used in place of `HOME` when `HOME` is unset or empty — the normal case on Windows. ⚠️ If **neither** is set, as under some service managers, those defaults resolve against the **current working directory** instead. Set `HOME` or the specific path variables above. The mcpdo daemon directory is the exception: with neither set, it falls back to the OS account's home directory (`os.homedir()`), not the working directory. | + +## mcpdo connection daemon + +mcpdo runs its connections in a local daemon (`mcpdod`) that it starts on first use. These select which daemon a command talks to, and how a command picks its connection. + +| Variable | Read by | Default | Effect | +| ------------------------------ | ------- | ------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `MCP_INSPECTOR_DAEMON_DIR` | mcpdo | `MCP_STORAGE_DIR` if set, else `~/.mcp-inspector` | Directory that holds the daemon's socket (`mcpdod.sock`), lock and published token. Takes precedence over `MCP_STORAGE_DIR`. `eval "$(mcpdo private)"` sets it to a fresh `0700` directory under the system temp directory, which gives that shell a daemon of its own. | +| `MCP_INSPECTOR_DAEMON_TOKEN` | mcpdo | unset | Bearer token every request to the daemon must carry. Unset, the daemon generates one at startup and publishes it to `mcpdod.token` in the daemon directory for same-user clients to read, so the shared daemon is still authenticated. `mcpdo private` sets it alongside `MCP_INSPECTOR_DAEMON_DIR`. Neither is a boundary against other processes running as your user (see the [mcpdo README](../clients/mcpdo/README.md)). | +| `MCP_ALLOW_DEFAULT_CONNECTION` | mcpdo | unset | When stdin is not a terminal, commands that act on a connection require an explicit `@name` or `--connection`. Set it to `1` to let them fall back to the most recently used connection, as they do interactively. Any other value keeps the requirement. | ## Secret store @@ -76,31 +86,31 @@ Where the Inspector's secrets (OAuth client secrets, the enterprise IdP client s > [!WARNING] > On a host with no OS keychain (Linux without libsecret or a Secret Service, headless or SSH sessions, Termux), the Inspector **automatically** stores secrets in a file that is **plaintext** unless `MCP_INSPECTOR_SECRET_KEY_FILE` or `MCP_INSPECTOR_SECRET_KEY` is set. See [the warning in Where secrets are stored](./secret-storage.md#how-the-store-is-chosen). -| Variable | Read by | Default | Effect | -| ---------------------------- | ------------- | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `MCP_INSPECTOR_SECRET_STORE` | web, CLI, TUI | probe the OS keychain | `keyring`, `file`, or `memory` (case-insensitive) picks the store outright and skips the probe. An empty or whitespace-only value counts as unset and silently runs automatic selection; any other value is ignored with a warning and also falls back to automatic selection. | -| `MCP_INSPECTOR_SECRET_FILE` | web, CLI, TUI | `~/.mcp-inspector/secrets.json` | Path of the file store. Lookup order: this variable, then `secrets.json` in `MCP_STORAGE_DIR` when that is set, then `~/.mcp-inspector/secrets.json`. ⚠️ The default sits **beside** the storage directory, not inside it. | -| `MCP_INSPECTOR_SECRET_KEY` | web, CLI, TUI | unset (file is plaintext, `0600`) | Passphrase that encrypts the file store; an empty or whitespace-only value counts as unset. Use a generated, high-entropy value. ⚠️ Changing or losing it makes the existing file unreadable; see [Where secrets are stored](./secret-storage.md#encryption) before rotating it. | -| `MCP_INSPECTOR_SECRET_KEY_FILE` | web, CLI, TUI | unset | Path of a file holding the passphrase; trailing line breaks are removed. Use this for Docker or Compose secrets, so the key stays out of the environment. Setting it together with a non-blank `MCP_INSPECTOR_SECRET_KEY` is an error, and so is setting it to an empty value. ⚠️ If the file is missing, unreadable or empty, or is the secrets file itself, the file store refuses to read or write rather than fall back to plaintext. | -| `MCP_INSPECTOR_PERSIST_TOKENS` | web, CLI, TUI | `all` | Which **acquired OAuth tokens** are persisted to the secret store: `all` (access + refresh tokens and IdP session tokens), `access` (access and ID tokens, but no refresh tokens), or `none` (no acquired tokens outlive the process; expect to re-authorize each run). Applies on write only — already-persisted tokens still load, and the next save under a stricter policy removes them from the store. Client secrets are registration credentials, not acquired tokens, and are persisted regardless. An empty value counts as unset; any other value is ignored with a warning and treated as `all`. | +| Variable | Read by | Default | Effect | +| ------------------------------- | -------------------- | --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `MCP_INSPECTOR_SECRET_STORE` | web, CLI, TUI, mcpdo | probe the OS keychain | `keyring`, `file`, or `memory` (case-insensitive) picks the store outright and skips the probe. An empty or whitespace-only value counts as unset and silently runs automatic selection; any other value is ignored with a warning and also falls back to automatic selection. | +| `MCP_INSPECTOR_SECRET_FILE` | web, CLI, TUI, mcpdo | `~/.mcp-inspector/secrets.json` | Path of the file store. Lookup order: this variable, then `secrets.json` in `MCP_STORAGE_DIR` when that is set, then `~/.mcp-inspector/secrets.json`. ⚠️ The default sits **beside** the storage directory, not inside it. | +| `MCP_INSPECTOR_SECRET_KEY` | web, CLI, TUI, mcpdo | unset (file is plaintext, `0600`) | Passphrase that encrypts the file store; an empty or whitespace-only value counts as unset. Use a generated, high-entropy value. ⚠️ Changing or losing it makes the existing file unreadable; see [Where secrets are stored](./secret-storage.md#encryption) before rotating it. | +| `MCP_INSPECTOR_SECRET_KEY_FILE` | web, CLI, TUI, mcpdo | unset | Path of a file holding the passphrase; trailing line breaks are removed. Use this for Docker or Compose secrets, so the key stays out of the environment. Setting it together with a non-blank `MCP_INSPECTOR_SECRET_KEY` is an error, and so is setting it to an empty value. ⚠️ If the file is missing, unreadable or empty, or is the secrets file itself, the file store refuses to read or write rather than fall back to plaintext. | +| `MCP_INSPECTOR_PERSIST_TOKENS` | web, CLI, TUI, mcpdo | `all` | Which **acquired OAuth tokens** are persisted to the secret store: `all` (access + refresh tokens and IdP session tokens), `access` (access and ID tokens, but no refresh tokens), or `none` (no acquired tokens outlive the process; expect to re-authorize each run). Applies on write only — already-persisted tokens still load, and the next save under a stricter policy removes them from the store. Client secrets are registration credentials, not acquired tokens, and are persisted regardless. An empty value counts as unset; any other value is ignored with a warning and treated as `all`. | When no store is configured, the choice also depends on whether the Inspector is running in a container, which it detects from `KUBERNETES_SERVICE_HOST` (or Docker's and Podman's marker files). That variable is set by the orchestrator, not by you. ## Logging and debugging -| Variable | Read by | Default | Effect | -| ----------------------- | ------------- | ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Variable | Read by | Default | Effect | +| ----------------------- | ------------- | ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `MCP_DEBUG` / `DEBUG` | launcher, TUI | off | Prints the full stack trace when the process exits on an error — the launcher for any `--web`/`--tui` failure, and the standalone TUI entry point for a startup failure. Any value other than empty, `0` or `false` (case-insensitive) turns it on. | -| `MCP_LOG_FILE` | web | unset (no log) | Appends the web backend's structured (pino, JSON lines) log to this file, creating its directory if needed. An empty value counts as unset. | -| `MCP_INSPECTOR_LOG_DIR` | TUI | `~/.mcp-inspector` | Directory of the TUI's `auth.log`. The TUI logs to a file so its output does not corrupt the terminal UI. | -| `LOG_LEVEL` | TUI | `info` | Level of the TUI's `auth.log`: one of `trace`, `debug`, `info`, `warn`, `error`, `fatal`, `silent`. An empty value is not replaced by `info`. | +| `MCP_LOG_FILE` | web | unset (no log) | Appends the web backend's structured (pino, JSON lines) log to this file, creating its directory if needed. An empty value counts as unset. | +| `MCP_INSPECTOR_LOG_DIR` | TUI | `~/.mcp-inspector` | Directory of the TUI's `auth.log`. The TUI logs to a file so its output does not corrupt the terminal UI. | +| `LOG_LEVEL` | TUI | `info` | Level of the TUI's `auth.log`: one of `trace`, `debug`, `info`, `warn`, `error`, `fatal`, `silent`. An empty value is not replaced by `info`. | ## Outbound proxy -| Variable | Read by | Default | Effect | -| ---------------------------- | ------------- | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Variable | Read by | Default | Effect | +| ---------------------------- | ------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `HTTPS_PROXY` / `HTTP_PROXY` | web, CLI, TUI | unset | Route connections to remote HTTP/SSE servers, including OAuth discovery and token requests, through a proxy. Lowercase forms are honored too. See [HTTP proxy support](../clients/cli/README.md#http-proxy-support). | -| `NO_PROXY` | web, CLI, TUI | unset | Hosts exempted from the proxy. | +| `NO_PROXY` | web, CLI, TUI | unset | Hosts exempted from the proxy. | ## Node.js variables diff --git a/docs/mcp-server-configuration.md b/docs/mcp-server-configuration.md index 3a9c5add14..f4918bfbe7 100644 --- a/docs/mcp-server-configuration.md +++ b/docs/mcp-server-configuration.md @@ -126,6 +126,7 @@ Without a `--` on the line the target is only the leading run of **non-dash** to | `-e <KEY=VALUE>` | Environment variable for a stdio server; repeatable | | | `--header "Name: Value"` | HTTP header for an HTTP/SSE server; repeatable | On web, requires an ad-hoc HTTP/SSE server | | `--protocol-era <era>` | `legacy`, `auto`, or `modern` — the era to negotiate | Sets [`protocolEra`](#inspector-specific-per-server-fields) without a file, on any transport. CLI/TUI also override a file's `protocolEra` with it; web, like `--header`, requires an ad-hoc target | +| `--skill-catalog-max-skills <n>` / `--skill-catalog-max-bytes <n>` | Skills catalog budget for a verification run; positive integer | **CLI/TUI only.** Set or override [`skillCatalogMaxSkills` / `skillCatalogMaxBytes`](#inspector-specific-per-server-fields) for every loaded server; web sets them per server in Server Settings instead | | `[target...]` | Positional command or URL for one ad-hoc server | | **`MCP_CATALOG_PATH` and ad-hoc targets differ by client.** The **CLI** ignores the env var when an ad-hoc target is given (a positional command, `--server-url`, or `--transport`), so a shell that exports it can still run one-off ad-hoc invocations without tripping the catalog/ad-hoc conflict. **Web and TUI read it unconditionally** — with it exported, an ad-hoc invocation such as `mcp-inspector --tui node build/index.js` is rejected as `--catalog cannot be combined with an ad-hoc server URL/command`. Unset the variable for that invocation on those two surfaces. @@ -190,7 +191,7 @@ These have no analog in the broader `mcp.json` ecosystem. Each is **omitted on w | `suppressNotificationStream` | `false` | Streamable HTTP, legacy era only: don't open the standalone `GET` notification stream. Server→client messages not carried on a request's own response stream won't arrive. `Last-Event-ID` resumption `GET`s still go out, and modern-era connections are unaffected (they never open this stream). A diagnostic and escape hatch for a server that times out every request after `initialize` because it cannot serve a second concurrent request ([#2317](https://github.com/modelcontextprotocol/inspector/issues/2317)) | | `advertisedExtensions` | — | Per-extension overrides for what the Inspector declares in `capabilities.extensions` | | `maxFetchRequests` | `1000` | Network-log retention for this server (`DEFAULT_MAX_FETCH_REQUESTS`); `0` means unlimited | -| `skillCatalogMaxSkills` | `256` | The maximum number of skills whose files are read in one verification run (`SKILL_MAX_CATALOG_SKILLS`) — the CLI's `--verify` and the TUI Skills pane. Positive integer; there is no unlimited value | +| `skillCatalogMaxSkills` | `256` | The maximum number of skills whose files are read in one verification run (`SKILL_MAX_CATALOG_SKILLS`) — the CLI's `--verify` and the TUI Skills pane's `v` (verify all); verifying a single skill always reads it in full. Positive integer; there is no unlimited value | | `skillCatalogMaxBytes` | `67108864` | The maximum number of bytes read across all skills in one verification run (`SKILL_MAX_CATALOG_BYTES`, 64 MiB). Positive integer | | `oauth` | — | `{ clientId, clientSecret, scopes, requestRefreshToken, revokeOnClear, authorizationParams, authorizationUrl, tokenUrl, enterpriseManaged, onInsufficientScope }` | @@ -228,6 +229,8 @@ The setting suppresses the grant declaration and the SDK's scope augmentation > > **Clear stored OAuth state** (Server Settings → Authorization) clears the Inspector's **local** copies — the tokens and the client information — and, where the authorization server supports it, revokes the grant there too (see `oauth.revokeOnClear` below), which settles the first. It does **not** touch the registration. For that, what happens next depends on how the client was obtained. A **dynamically registered** client is registered afresh on the next connect, and the new registration declares only `authorization_code`; the old one still exists at the AS, unused. A **preconfigured `oauth.clientId`** is reused as-is, so changing what that client declares is done at the authorization server, not here. +In standard (non-enterprise-managed) OAuth, a **preconfigured `oauth.clientId`** also needs the Inspector's **redirect URI** registered at the authorization server before the first connect (a dynamically registered client sends it itself). Under **enterprise-managed authorization** `oauth.clientId` is the Resource AS client, used only to redeem the token, and the authorization request carrying the redirect URI goes to the enterprise IdP — so register the URI on the IdP client configured in Client Settings instead. In the web client it is `<origin>/oauth/callback` — `http://localhost:6274/oauth/callback` by default — and it follows the address bar, so `localhost` vs `127.0.0.1`, a non-default port or a non-local host each change it. Server Settings → OAuth Settings shows the exact value, read-only with a copy button ([#2524](https://github.com/modelcontextprotocol/inspector/issues/2524)). The CLI and TUI use `--callback-url` / `MCP_OAUTH_CALLBACK_URL` instead (see [environment variables](./environment-variables.md)). + `oauth.revokeOnClear` (default `true`) controls whether clearing this server's stored OAuth state also **revokes the grant at the authorization server**, per [RFC 7009](https://datatracker.ietf.org/doc/html/rfc7009). Uncheck **Revoke tokens on clear** in Server Settings → Authorization to turn it off; only `false` is written to disk, so a server that never touched the setting keeps a minimal entry ([#2144](https://github.com/modelcontextprotocol/inspector/issues/2144)). Without it, clearing is silent from the authorization server's point of view: the Inspector deletes its local copy and the access token — and the refresh token, which is long-lived by design — stay valid there until they expire on their own. A day of connect/disconnect iteration leaves the AS holding a pile of grants for sessions that ended hours ago, and nothing in the Inspector can see or end them. RFC 7009 §1 describes this exact case; the clear is that moment. diff --git a/docs/quality-gate.md b/docs/quality-gate.md index 67e62c1be2..7168a962b0 100644 --- a/docs/quality-gate.md +++ b/docs/quality-gate.md @@ -14,7 +14,9 @@ Each client self-validates from its own folder; the root scripts chain them. The | **GitHub CI** (`.github/workflows/main.yml`) | Automatically, on every push | `npm install`, then `validate`, `verify:skills:cli`, `verify:build-gate`, `verify:bundle-externals`, `smoke` (which includes `smoke:web:chromium`), `test:storybook` — plus `coverage` in a parallel job ([#2159](https://github.com/modelcontextprotocol/inspector/issues/2159)) | | **The local gate** (`npm run local:gate`) | By hand, before you push | Every check above (the install is yours to run; `local:validate` stands in for `validate`, see below), **plus** the Firefox engine pass (`smoke:web:firefox`) | -The local gate runs **every check** CI runs, and is not a mirror. One of its steps has no GitHub CI counterpart: +One more CI check runs outside that table: **`.github/workflows/dco.yml`** fails on any commit that is not signed off ([#2566](https://github.com/modelcontextprotocol/inspector/issues/2566), [#2616](https://github.com/modelcontextprotocol/inspector/issues/2616)). It has two jobs. `DCO` runs on every *pull request with a `v2/**` base*, so `v2/main` and stacked v2 PRs, over the PR's own commits; v1 PRs and milestone PRs into `main` are out of its scope. `DCO (v2/main push)` re-checks every push that lands on `v2/main`, as a backstop for anything that merged without a passing PR check. The PR job runs on `pull_request`, so it is live as soon as the workflow is on `v2/main`. The trade-off is that a PR could edit its own check, and the push job could not catch that either, since a push run reads the workflow and the script from the pushed revision. It is accepted because PRs are maintainer-only, such an edit shows in the PR's diff, and the check exists to catch a *forgotten* signoff, not a forged one. The local gate runs the same script before a push as its `local:dco` stage (below), over `origin/v2/main..HEAD`. It replaced the probot DCO app, whose check vanished unnoticed when the app was suspended because it was never required. The replacement gates merges into `v2/main` as a **required** status check, set by the `v2/main - DCO` repository ruleset ([#2621](https://github.com/modelcontextprotocol/inspector/issues/2621)) because a workflow cannot declare itself required. The ruleset pins the check to GitHub Actions, so a hand-posted `DCO` status cannot satisfy it, and blocks deleting or force-pushing `v2/main`, which the push job's `before..after` range assumes never happens. + +The local gate runs **every check** `main.yml` runs, and is not a mirror. One of its steps has no GitHub CI counterpart: | Local-only step | Why it is local-only | | --- | --- | @@ -46,8 +48,9 @@ That is the readable half, and prose rots. The enforced half is `scripts/lib/wor | `npm run verify:dep-lockstep` | Guards the "one version per install-crossing dependency" invariant (#1896). v2 is not a workspace, so a client's test project compiles the shared first-party TypeScript — `core/`, `test-servers/src`, and the root-owned `vitest.shared.mts`, all of which resolve their dependencies from the **root** install — alongside the client's own sources, putting the same package in one `tsc` program twice. At the same version that's harmless; skewed, TypeScript must relate two structurally-distinct copies of every type, which for a recursive-generic surface is exponential (zod `4.3.6` vs `4.4.3` exhausted the 4GB tsc heap in `clients/web`). Derives its candidate set from **what actually enters each program** (#1965) — every client tsconfig project listed with `tsc --listFilesOnly` via the shared `scripts/lib/tsc-program.mjs`, each resolved `node_modules` file mapped to its owning install, keeping the packages that reach one program from two installs (a package whose declarations arrive only through another package's `.d.ts`, as `@modelcontextprotocol/sdk`'s do, is invisible to a scan of first-party imports). Prices each copy from the lockfile entry for the exact install path the program resolved, compares only the installs that met in one program, and **fails deny-by-default** on any disagreement not in the annotated `TOLERATED_SKEW` allowlist — empty today — with an allowlisted package tolerated only *within a major version*. A **second tier** (#2226) runs alongside it, asking the weaker but broader question the `AGENTS.md` rule actually states: does a package this repo *declares* anywhere resolve to two versions across our installs at all? Its candidate set is every name in any install's `dependencies`/`devDependencies`/`optionalDependencies` — unioned across the root and all four clients, so a copy declared by only one of them still counts — that **more than one install holds a top-level copy of** (17 packages today). Nested copies are excluded: one exists because some dependency asked for a different version, so it is that dependency's range to govern, not ours. Neither tier subsumes the other — the program tier sees a copy no manifest names (`@modelcontextprotocol/sdk`, arriving through another package's `.d.ts`), while the declared tier sees a **transitive** copy no program loads (cli's `@types/node`, hoisted via `@types/express` — the case that motivated it), two **clients** disagreeing with no root copy involved (`@types/react`, web against tui), and the peer shadows `eslint`/`typescript`/`vitest` that never enter a program. Same deny-by-default and same within-a-major rule, against its own `TOLERATED_DECLARED_SKEW` — also empty. Two limits: it reads **lockfiles**, so an uncommitted hand-installed copy is invisible, and it compares only declared names, so a purely transitive package no manifest names stays the first tier's business. Runs in `validate`. | `npm run verify:test-timeouts` | Guards the wall-clock budgets the test gates run under (#2323). The class it encodes against is a budget **nobody chose**: three of the six Vitest projects ran on Vitest's own `testTimeout: 5000` and five on its `hookTimeout`/`teardownTimeout: 10000`, sized for an idle machine rather than the one this team works on — three or four concurrent agent sessions in separate worktrees, each free to run a full `local:gate`, on eight logical cores. A correct, deterministic test cut off by such a budget fails a gate its diff did not break, which is the same "channel nobody trusts" failure [Lint has no warning tier](../AGENTS.md#lint-has-no-warning-tier) describes from the other direction; #2292, #1942 and #1742 were each an instance, found one site at a time. Deliberately **not** cited: #2278 (a missing condition wait around a geometry read) and #2250 (a real race in a test's own timing) were fixed by making the test wait for the right thing, and presenting a race fix as evidence for a larger ceiling would argue against #1596. It asks **Vitest itself** to resolve each of the four configs and reads the number a test actually gets, rather than checking that a key is absent from a config block — which would pass just as happily on a config that had stopped being loaded. A **seventh project** with no row, or a config file it does not discover (it knows all twelve filenames Vitest accepts, not just the two this repo uses), is an error rather than a silent skip. It also checks that every project loads `vitest.setup.shared.mts`. ⚠️ **Everything it reads comes from a resolved project, never from source text** — the two rules no config can report are asserted at runtime instead: `retry` by `vitest.setup.shared.mts` (which reads the value Vitest resolved, so a per-test option, a `describe` option, a project setting and a `--retry` flag are one check) and Testing Library's `asyncUtilTimeout` by `clients/web/src/test/asyncUtilTimeout.test.ts` (which reads what the project's own `waitFor`s use). Both began as source scanning in #2334 and the review found a new valid spelling missed in five consecutive rounds, so the boundary is deliberate: if a rule cannot be answered by asking the tool, assert it at runtime rather than reading the source for it. Observed **red** against the pre-#2323 config before it was trusted. Runs in `validate`; its decision logic has a sibling `verify-test-timeouts.test.mjs` under `test:scripts`. | `npm run verify:action-pins` | Guards the credentialed-job pinning rule ([#2484](https://github.com/modelcontextprotocol/inspector/issues/2484)). Actions stay on moving major tags by default (#2235), but a tag is mutable, so whoever can move one replaces the code running next to a credential. For every job under `.github/workflows/` that holds `id-token: write` / `packages: write` / `write-all` (its own `permissions:` or the workflow's), is handed any secret other than `GITHUB_TOKEN`, or uploads an artifact such a job downloads and `needs`, it fails on any `uses:` that is not a 40-hex commit SHA followed by an exact `# vX.Y.Z` comment. The comment is what keeps the pin in the monthly `dependency-refresh` sweep, which ranks a SHA pin by it to the patch. Offline, so it cannot check that the SHA matches the comment — resolve both from one tag lookup. | -| `npm run local:gate` | **Mandatory pre-push command.** `local:validate` → `verify:skills:cli` → `coverage` → `verify:build-gate` → `verify:bundle-externals` → `smoke` → `smoke:web:firefox` → `local:storybook`. Every GitHub CI check plus one local-only one — see [Two tiers](#two-tiers-github-ci-and-the-local-gate). Named `local:` rather than `ci` on purpose (#2146); there is no `npm run ci` alias. **Runs under a machine-wide lease** (`scripts/gate-lease.mjs`, [#2339](https://github.com/modelcontextprotocol/inspector/issues/2339)): a second `local:gate` started in another worktree waits for the first to finish instead of running alongside it, and queued gates start in the order they arrived ([#2473](https://github.com/modelcontextprotocol/inspector/issues/2473)). Measured on `aa56551b`, a quiet gate took 257s; two started together had one **fail** at 279s on a smoke port collision (`smoke:web:chromium` binds 6298, and so did the other gate's) and the survivor take 338s — so overlapping gates are red by construction, not merely slow, and back to back both finish green in ~2x257s. The wait prints the holder's pid and worktree; a holder that dies without releasing is taken over after 30s (`proper-lockfile` stale detection) — unless its lock directory cannot be removed, in which case the wait runs to its 45-minute cap and names the path; `INSPECTOR_SKIP_GATE_LEASE=1` bypasses it. The stages themselves are `local:gate:stages`, which the wrapper runs verbatim. Why it is a queue rather than a load wait or a worker cap, and what it does and does not change about capacity, is [Multi-agent testing](#multi-agent-testing-the-gate-lease). | +| `npm run local:gate` | **Mandatory pre-push command.** `local:validate` → `local:dco` → `verify:skills:cli` → `coverage` → `verify:build-gate` → `verify:bundle-externals` → `smoke` → `smoke:web:firefox` → `local:storybook`. Every GitHub CI check plus one local-only one — see [Two tiers](#two-tiers-github-ci-and-the-local-gate). Named `local:` rather than `ci` on purpose (#2146); there is no `npm run ci` alias. **Runs under a machine-wide lease** (`scripts/gate-lease.mjs`, [#2339](https://github.com/modelcontextprotocol/inspector/issues/2339)): a second `local:gate` started in another worktree waits for the first to finish instead of running alongside it, and queued gates start in the order they arrived ([#2473](https://github.com/modelcontextprotocol/inspector/issues/2473)). Measured on `aa56551b`, a quiet gate took 257s; two started together had one **fail** at 279s on a smoke port collision (`smoke:web:chromium` binds 6298, and so did the other gate's) and the survivor take 338s — so overlapping gates are red by construction, not merely slow, and back to back both finish green in ~2x257s. The wait prints the holder's pid and worktree; a holder that dies without releasing is taken over after 30s (`proper-lockfile` stale detection) — unless its lock directory cannot be removed, in which case the wait runs to its 45-minute cap and names the path; `INSPECTOR_SKIP_GATE_LEASE=1` bypasses it. The stages themselves are `local:gate:stages`, which the wrapper runs verbatim. Why it is a queue rather than a load wait or a worker cap, and what it does and does not change about capacity, is [Multi-agent testing](#multi-agent-testing-the-gate-lease). | | `npm run local:validate` | The gate's first stage: `validate` with each client's `test` leg removed (#2341) — the same `validate:guards` and `validate:core`, then every client's `check` — `format:check` + `lint` + `typecheck`, plus `build` for web, tui and launcher, which name one explicitly (cli's only validate-time build was `test`'s `pretest` hook, so it goes with the `test` leg and `coverage:cli` builds the binary once, later) — each client's `validate` being `check && test`. The suites run once, under `coverage`, instead of bare here and instrumented there. In the `local:` namespace so the workflow guard keeps it out of CI, where the bare pass is free. | +| `npm run local:dco` | The gate's second stage ([#2616](https://github.com/modelcontextprotocol/inspector/issues/2616)): `scripts/dco-check.mjs` over `origin/v2/main..HEAD`, excluding anything already on `origin/main` (`--exclude`; on a milestone-merge branch, cut from `main`, the bare range would be `main`'s own released and largely unsigned history), so a commit missing its `Signed-off-by:` trailer fails before it is pushed rather than on the PR's `DCO` check. The same script `.github/workflows/dco.yml` runs on every v2 PR and on every push to `v2/main`. Local-only: it reads `origin/v2/main`, which CI's shallow push checkout does not have, and `workflow-gate.test.mjs` pins it out of `validate`. A stale `origin/v2/main` only widens the range to commits that are already merged. | | `npm run local:storybook` | The gate's last stage: from `clients/web`, `npx playwright install chromium` and then `test:storybook` — the Storybook play functions, run headless by Vitest's browser project (CI runs `test:storybook` directly after its own Playwright install step, which is why this wrapper is in the `local:` namespace). Since [#2340](https://github.com/modelcontextprotocol/inspector/issues/2340) (PR [#2342](https://github.com/modelcontextprotocol/inspector/pull/2342)) the project's `optimizeDeps` carries `force: true` (`getStorybookOptimizeDeps` in `clients/web/server/vite-base-config.ts`), so Vite discards its dep pre-bundle cache and re-scans every story entry before the browser opens — the cold path CI takes on every PR. Without it, Vite keys that cache on the lockfile and the config, never on what the story graph imports, so after a story gains an import of an already-installed package the "valid" cache re-runs the optimizer mid-run and sends a `full-reload` the Vitest tester iframes do not act on; every story an already-loaded iframe renders from then on fails with `Failed to fetch dynamically imported module`, and the next run is green because the cache has caught up. Measured: a stale cache went 3/3 red → 3/3 green with `force` → 3/3 red reverted, load refuted as a cause (0 of 6 under load 8–12), cost ~0.8s per run against a ~25s stage. | | `npm run pack:verify` | Publish smoke — see [Publishing](./publishing.md). | diff --git a/docs/secret-storage.md b/docs/secret-storage.md index 6c03505b46..d4da053cb1 100644 --- a/docs/secret-storage.md +++ b/docs/secret-storage.md @@ -6,18 +6,18 @@ The Inspector keeps a few values out of `mcp.json`, `client.json` and `oauth.jso These values are stored as secrets: -| Value | Saved from | -| --------------------------------- | -------------------------------------- | -| A server's OAuth client secret | The server's OAuth settings | -| The enterprise IdP client secret | Client Settings (install-level) | -| Each stdio server's `env:` value | A stdio server's environment variables | -| Acquired OAuth tokens (access, refresh, ID) | Completing an OAuth flow | -| IdP session tokens (enterprise-managed auth) | Completing an IdP login | -| Dynamically registered client secrets | DCR during an OAuth flow | +| Value | Saved from | +| -------------------------------------------- | -------------------------------------- | +| A server's OAuth client secret | The server's OAuth settings | +| The enterprise IdP client secret | Client Settings (install-level) | +| Each stdio server's `env:` value | A stdio server's environment variables | +| Acquired OAuth tokens (access, refresh, ID) | Completing an OAuth flow | +| IdP session tokens (enterprise-managed auth) | Completing an IdP login | +| Dynamically registered client secrets | DCR during an OAuth flow | They are kept out of `mcp.json` so that sharing, committing or syncing the file does not leak credentials (#1356). When the Inspector saves an entry to a durable store, it leaves each `env` key in `mcp.json` with an empty value and omits the client secret; the real values live in the store. `headers` are **not** moved: they are saved in `mcp.json` exactly as written, so a header that carries a credential stays in the file. [MCP server configuration](./mcp-server-configuration.md) describes what that means for other tools reading the same file. -Acquired tokens follow the same rule for the OAuth state file: `oauth.json` keeps only non-secret state (flow bookkeeping, discovered metadata, public client ids), and the tokens and client secrets it used to hold live in the secret store. A pre-existing `oauth.json` that still carries plaintext tokens is migrated on first read — the tokens move into the store and the file is rewritten without them — but only when the store is durable; under the `memory` store the file is left as-is, since it is still the only durable copy. A token entry so malformed it cannot be a credential (for example, a hand-edited value of the wrong type) is left in the file rather than migrated, so it stays visible and clearable. `MCP_INSPECTOR_PERSIST_TOKENS=all|access|none` controls which acquired tokens are persisted at all (see [Environment variables](./environment-variables.md#secret-store)). +Acquired tokens follow the same rule for the OAuth state file: `oauth.json` keeps only non-secret state (flow bookkeeping, discovered metadata, public client ids), and the tokens and client secrets it used to hold live in the secret store. Each state file's store entries are scoped by a `secretsNamespace` UUID stamped into the file on its first write, so two state files (for example, per-profile `MCP_INSPECTOR_OAUTH_STATE_PATH` values) that authorize against the same server keep separate entries even in a shared store such as the OS keychain. A pre-namespace file is adopted transparently on its first save — its existing un-namespaced entries move under the new namespace, so the adopting file keeps its credentials. That migration deletes its sources, so it runs only under the real cross-process file lock; on a box where the lock cannot be taken, the save fails with a retryable error instead of racing a concurrent adopter — if you cannot make the lock directory writable there, the escape is to clear the file's stored OAuth state and re-authorize, since clearing does not migrate and a fresh state file mints its namespace without the lock. The namespace stamp is also one-way across versions: an Inspector older than the namespace (≤ 2.9.x) resolves only un-namespaced entries, so downgrading after adoption logs you out, and its next save strips the stamp — alternating old and new versions on one state file therefore re-adopts under a fresh namespace each time. The earlier namespace's entries are not stranded by this: each state file keeps a sidecar ledger, `<state file>.namespaces.json`, recording every namespace it has used and the servers and IdP issuers given store entries under each, and the next re-adoption (or removal of the state file) purges every recorded namespace the file no longer carries — on every store backend, the OS keychain included, since the ledger names each entry and nothing has to be enumerated. The same purge cleans up after a state file deleted by hand, at the next save to that path. Entries are recorded per store backend (the keychain, or a particular `secrets.json`), and only a run using that backend purges them — so a run that falls back to a different store leaves the keychain's record for the next keychain run rather than "purging" an empty store and forgetting it. A keychain run also purges what was recorded against a `secrets.json`, since moving that file into the keychain once one is available copies its entries across; and when a `secrets.json` run purges its own record, the record passes to the keychain rather than being forgotten, so those copies are still cleaned up the next time the keychain is in use. Like adoption's migration, the purge runs only under the real cross-process file lock; a namespace whose purge fails stays recorded and is retried on the next save, with a warning. One limit remains: if the ledger is deleted along with its state file, the entries it recorded cannot be found on the keyring backend, which cannot list its entries; on the `file` backend they remain visible in `secrets.json` (ids beginning `oauth+<namespace>+`) and can be removed by hand. If more than one pre-namespace state file was sharing a server's un-namespaced entry, the first one to adopt takes it with it, and each remaining pre-namespace profile re-authorizes that server once — its copy was already being overwritten by every other profile's saves, which is the bug the namespace fixes. Removing a profile that is still pre-namespace likewise purges the shared un-namespaced entries, as removal always has: the file being deleted is the store's only index of them, so leaving them would strand credentials nothing could find or clear again. A pre-existing `oauth.json` that still carries plaintext tokens is migrated on first read — the tokens move into the store and the file is rewritten without them — but only when the store is durable; under the `memory` store the file is left as-is, since it is still the only durable copy. A token entry so malformed it cannot be a credential (for example, a hand-edited value of the wrong type) is left in the file rather than migrated, so it stays visible and clearable. `MCP_INSPECTOR_PERSIST_TOKENS=all|access|none` controls which acquired tokens are persisted at all (see [Environment variables](./environment-variables.md#secret-store)). ## How the store is chosen @@ -31,15 +31,15 @@ Each process picks one store, once, the first time it needs it: the web backend The probe requires the keychain to both **read** and **enumerate** entries. On Linux the two can come from different providers: single entries can be served from the kernel keyring (keyutils) without a Secret Service, but enumeration — which deleting a server's credentials depends on — needs the Secret Service itself. A host with only the kernel keyring therefore falls back exactly like a host with no keychain at all. Earlier releases probed reads only and selected the keychain on such hosts; anything they stored lives in kernel memory only (it never survives a reboot) and is not read by the fallback store. To keep using those entries for the remainder of the boot session, set `MCP_INSPECTOR_SECRET_STORE=keyring`. -| Where you run it | Store | Secrets survive a restart? | -| ----------------------------------------------------------------------- | ----------------------------------- | -------------------------- | -| Desktop macOS or Windows, or Linux with a Secret Service running | OS keychain | Yes | -| Linux without libsecret or a Secret Service | File (`secrets.json`, mode `0600`) | Yes | -| Headless server or SSH session with no D-Bus session | File | Yes | -| Android/Termux | File | Yes | -| Container with **no volume** on the secrets directory | Memory | No, this session only | -| Container **with** a volume on the secrets directory | File | Yes | -| Any of the above with `MCP_INSPECTOR_SECRET_STORE` set | The store you named | Not with `memory`; with `file` in a container, only if the file is on a volume | +| Where you run it | Store | Secrets survive a restart? | +| ---------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------------------------------------ | +| Desktop macOS or Windows, or Linux with a Secret Service running | OS keychain | Yes | +| Linux without libsecret or a Secret Service | File (`secrets.json`, mode `0600`) | Yes | +| Headless server or SSH session with no D-Bus session | File | Yes | +| Android/Termux | File | Yes | +| Container with **no volume** on the secrets directory | Memory | No, this session only | +| Container **with** a volume on the secrets directory | File | Yes | +| Any of the above with `MCP_INSPECTOR_SECRET_STORE` set | The store you named | Not with `memory`; with `file` in a container, only if the file is on a volume | > [!WARNING] > **With no keychain, secrets go to a plaintext file, and you did not have to ask for it.** On a host where the keychain probe fails (Linux without libsecret or a running Secret Service such as GNOME Keyring or KWallet, a headless server or SSH session with no D-Bus session, or Android/Termux), the Inspector falls back **automatically** to `~/.mcp-inspector/secrets.json`. Unless you supply a key, that file is **unencrypted**. Mode `0600` keeps out other non-root users, but not root, not backups or copies of your home directory, and not any program running as you. The only signs are a warning on stderr when the store is selected and the footer in the web settings dialogs. @@ -58,7 +58,7 @@ The choice is made once per process. Installing a keychain while the Inspector i ### The memory store -`memory` keeps secrets for this process only; nothing is written anywhere and they are gone when it exits. Because it is not durable, the Inspector does **not** remove plaintext values that are already in `mcp.json`, `client.json` or `oauth.json` while it is active: in that case the file on disk is still the durable copy — including pre-existing OAuth tokens, which keep working across runs until a save changes or removes them. New or changed values are still kept out of the file, so a token *acquired or refreshed* under the memory store lasts this session only and needs a re-auth next run. +`memory` keeps secrets for this process only; nothing is written anywhere and they are gone when it exits. Because it is not durable, the Inspector does **not** remove plaintext values that are already in `mcp.json`, `client.json` or `oauth.json` while it is active: in that case the file on disk is still the durable copy — including pre-existing OAuth tokens, which keep working across runs until a save changes or removes them. New or changed values are still kept out of the file, so a token _acquired or refreshed_ under the memory store lasts this session only and needs a re-auth next run. ## The file store @@ -141,12 +141,12 @@ A successful move prints a message naming the file it removed. The same hand-off ## Changing the store -| To | Set | -| ------------------------------------------ | -------------------------------------------------------------------- | -| Always use the keychain | `MCP_INSPECTOR_SECRET_STORE=keyring` | -| Use a file even though a keychain exists | `MCP_INSPECTOR_SECRET_STORE=file` | -| Never write secrets to disk | `MCP_INSPECTOR_SECRET_STORE=memory` | -| Put the file somewhere else | `MCP_INSPECTOR_SECRET_FILE=/path/to/secrets.json`, or `MCP_STORAGE_DIR` | -| Encrypt the file | `MCP_INSPECTOR_SECRET_KEY_FILE=/path/to/key-file` (preferred), or `MCP_INSPECTOR_SECRET_KEY=<generated passphrase>` | +| To | Set | +| ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | +| Always use the keychain | `MCP_INSPECTOR_SECRET_STORE=keyring` | +| Use a file even though a keychain exists | `MCP_INSPECTOR_SECRET_STORE=file` | +| Never write secrets to disk | `MCP_INSPECTOR_SECRET_STORE=memory` | +| Put the file somewhere else | `MCP_INSPECTOR_SECRET_FILE=/path/to/secrets.json`, or `MCP_STORAGE_DIR` | +| Encrypt the file | `MCP_INSPECTOR_SECRET_KEY_FILE=/path/to/key-file` (preferred), or `MCP_INSPECTOR_SECRET_KEY=<generated passphrase>` | Every variable is also listed in [Environment variables](./environment-variables.md#secret-store). diff --git a/docs/test-servers.md b/docs/test-servers.md index 17fde6cef2..aea4f12022 100644 --- a/docs/test-servers.md +++ b/docs/test-servers.md @@ -6,6 +6,7 @@ The catalogue below is the reference. For how to build and run one, use the `/te - **In-process** — import the factories (`createTestServerHttp`, `createEchoTool`, …) and run the server inside the test's event loop (used by the HTTP integration paths). - **As a subprocess** — `test-servers/build/test-server-stdio.js` is spawned as a real stdio child (used by the CLI smoke and stdio integration tests). + Started with `--crashable` (`getCrashableTestMcpServerCommand()`), it also serves a `crash_server` tool that exits the process at a point the test picks, for exercising the client's mid-session crash handling. That tool is deliberately absent from the presets: it calls `process.exit`, which in-process would end the test runner. Configure a server declaratively with a JSON config (see `test-servers/configs/*.json`) selecting presets, then load it via `--config`. Because the servers are spawned as real subprocesses, the build output must exist first: @@ -735,12 +736,12 @@ The first resource-subscription listen is acknowledged **so the badge is reachab **Legacy** (`tasks-legacy-http.json`) advertises `capabilities.tasks` (`tasks: { list, cancel }`) with the `simple_task` / `progress_task` / `elicitation_task` presets. Run one of those tools with **Run as task** on, and the **Tasks** tab lists it (populated via `tasks/list`), polls `tasks/get`, fetches the payload with the blocking `tasks/result`, and cancels with `tasks/cancel`. -**Modern** (`tasks-modern-http.json`) sets `transport.modern: true` and `tasksExtension: true`, advertising the `io.modelcontextprotocol/tasks` extension (SEP-2663) and serving `modern_task` / `modern_input_task`. The **Tasks** tab is gated on the negotiated extension, not `capabilities.tasks`. +**Modern** (`tasks-modern-http.json`) sets `transport.modern: true` and `tasksExtension: true`, advertising the `io.modelcontextprotocol/tasks` extension (SEP-2663) and serving `modern_task` / `modern_input_task`. Connect with **Server Settings → Options → Protocol Era = Modern**; on the default Legacy era the handshake is `initialize` and no task is ever created. The **Tasks** tab is gated on the negotiated extension, not `capabilities.tasks`. -- Run `modern_task` as a task — the `tools/call` returns a `CreateTaskResult` (`resultType: "task"`, visible in the Protocol/Network tabs), the client polls **`tasks/get`** (no `tasks/list`), and the completed task inlines its result (no blocking `tasks/result`). +- Run `modern_task` — the tools declare no `taskSupport`, so there is no **Run as task** switch; the server answers the plain call with a task anyway. The `tools/call` returns a `CreateTaskResult` (`resultType: "task"`, visible in the Protocol/Network tabs), the client polls **`tasks/get`** (no `tasks/list`), and the completed task inlines its result (no blocking `tasks/result`). - Run `modern_input_task` — the task moves to `input_required`, surfacing an embedded elicitation through the pending-request modal. Answering it sends **`tasks/update`** with the `inputResponses`, and the next poll completes. -SDK v2 removed all tasks support **and** era-gates the `tasks/*` spec methods out of the modern era on both sides. So the Inspector drives the extension itself — the `resultType: "task"` frame is rewritten at the transport into a `CallToolResult` carrying the handle, and `tasks/get` / `update` / `cancel` ride a raw-wire request channel with the full modern envelope. The test server serves `tasks/*` from an Express interceptor ahead of the SDK handler, since the SDK's modern leg would answer them `-32601`. +SDK v2 removed all tasks support **and** era-gates the `tasks/*` spec methods out of the modern era on both sides. So the Inspector drives the extension through `@modelcontextprotocol/ext-tasks` (#2316): on the modern era every `tools/call` and `tasks/*` request goes out on the raw-wire channel the package requires as `rawDispatch` (`core/extension/tasks/rawWireChannel.ts`), below the SDK codec, and the package handles the `resultType: "task"` result, the `tasks/get` polling and `tasks/update` itself. The test server serves `tasks/*` from an Express interceptor ahead of the SDK handler, since the SDK's modern leg would answer them `-32601`. The Tasks tab's **Refresh** re-polls the handles already known to the client — modern has no server-side task list. diff --git a/package-lock.json b/package-lock.json index 32bd0f1af3..7b5c7f073d 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "@modelcontextprotocol/inspector", - "version": "2.9.0", + "version": "2.10.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "@modelcontextprotocol/inspector", - "version": "2.9.0", + "version": "2.10.0", "hasInstallScript": true, "license": "SEE LICENSE IN LICENSE", "dependencies": { @@ -14,6 +14,7 @@ "@modelcontextprotocol/client": "2.2.0", "@modelcontextprotocol/core": "2.2.0", "@modelcontextprotocol/ext-apps": "^2.0.0", + "@modelcontextprotocol/ext-tasks": "^0.2.2", "@modelcontextprotocol/server": "2.2.0", "@napi-rs/keyring": "^1.3.0", "@vitejs/plugin-react": "^6.0.0", @@ -33,7 +34,8 @@ "zod": "^4.4.3" }, "bin": { - "mcp-inspector": "clients/launcher/build/index.js" + "mcp-inspector": "clients/launcher/build/index.js", + "mcpdo": "clients/mcpdo/build/mcp-bin.js" }, "devDependencies": { "@eslint/js": "^10.0.1", @@ -689,6 +691,22 @@ } } }, + "node_modules/@modelcontextprotocol/ext-tasks": { + "version": "0.2.2", + "resolved": "https://registry.npmjs.org/@modelcontextprotocol/ext-tasks/-/ext-tasks-0.2.2.tgz", + "integrity": "sha512-qpyxUGaameAdtFF09NVPhe82xo70kkED7Kel5c+jAfPTUAh/V5xzWjd9iRBIZFNhK+r9egGIFglDZh0kPbXeoA==", + "dependencies": { + "zod": "^4.5.4" + }, + "peerDependencies": { + "@modelcontextprotocol/client": "^2.0.0" + }, + "peerDependenciesMeta": { + "@modelcontextprotocol/client": { + "optional": true + } + } + }, "node_modules/@modelcontextprotocol/server": { "version": "2.2.0", "resolved": "https://registry.npmjs.org/@modelcontextprotocol/server/-/server-2.2.0.tgz", @@ -1072,9 +1090,6 @@ "cpu": [ "arm64" ], - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1091,9 +1106,6 @@ "cpu": [ "arm64" ], - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -1110,9 +1122,6 @@ "cpu": [ "ppc64" ], - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1129,9 +1138,6 @@ "cpu": [ "s390x" ], - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1148,9 +1154,6 @@ "cpu": [ "x64" ], - "libc": [ - "glibc" - ], "license": "MIT", "optional": true, "os": [ @@ -1167,9 +1170,6 @@ "cpu": [ "x64" ], - "libc": [ - "musl" - ], "license": "MIT", "optional": true, "os": [ @@ -3675,9 +3675,6 @@ "cpu": [ "arm64" ], - "libc": [ - "glibc" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -3698,9 +3695,6 @@ "cpu": [ "arm64" ], - "libc": [ - "musl" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -3721,9 +3715,6 @@ "cpu": [ "x64" ], - "libc": [ - "glibc" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -3744,9 +3735,6 @@ "cpu": [ "x64" ], - "libc": [ - "musl" - ], "license": "MPL-2.0", "optional": true, "os": [ @@ -4355,9 +4343,9 @@ } }, "node_modules/proxy-addr": { - "version": "2.0.7", - "resolved": "https://registry.npmjs.org/proxy-addr/-/proxy-addr-2.0.7.tgz", - "integrity": "sha512-llQsMLSUDUPT44jdrU/O37qlnifitDP+ZwrmmZcoSKyLKvtZxpyV0n2/bD/N4tBAAZ/gJEdZU7KMraoK1+XYAg==", + "version": "2.0.8", + "resolved": "https://registry.npmjs.org/proxy-addr/-/proxy-addr-2.0.8.tgz", + "integrity": "sha512-5nnx0yGyVUcY6t9RnWcARWtwT9F1D8O9rt08htPvnd49W1IgZtmLkhu9WfMzQj1cFxjHIO6connUNVW5k7AVyQ==", "dev": true, "license": "MIT", "dependencies": { @@ -4366,6 +4354,10 @@ }, "engines": { "node": ">= 0.10" + }, + "funding": { + "type": "opencollective", + "url": "https://opencollective.com/express" } }, "node_modules/punycode": { @@ -4794,9 +4786,9 @@ } }, "node_modules/source-map-js": { - "version": "1.2.1", - "resolved": "https://registry.npmjs.org/source-map-js/-/source-map-js-1.2.1.tgz", - "integrity": "sha512-UXWMKhLOwVKb728IUtQPXxfYU+usdybtUrK/8uGE8CQMvrhOpwvzDBwj0QhSL7MQc7vIsISBG8VQ8+IDQxpfQA==", + "version": "1.2.2", + "resolved": "https://registry.npmjs.org/source-map-js/-/source-map-js-1.2.2.tgz", + "integrity": "sha512-KGj/8Y43x35aZVDtt+J4mK1hoLGHULMYfSkODJNQjNDC3oW1PqPoxMwo0pLUsWM/UEGzON/NxeHywEfNXNP3Vw==", "license": "BSD-3-Clause", "engines": { "node": ">=0.10.0" @@ -5537,9 +5529,9 @@ "license": "MIT" }, "node_modules/zod": { - "version": "4.4.3", - "resolved": "https://registry.npmjs.org/zod/-/zod-4.4.3.tgz", - "integrity": "sha512-ytENFjIJFl2UwYglde2jchW2Hwm4GJFLDiSXWdTrJQBIN9Fcyp7n4DhxJEiWNAJMV1/BqWfW/kkg71UDcHJyTQ==", + "version": "4.6.5", + "resolved": "https://registry.npmjs.org/zod/-/zod-4.6.5.tgz", + "integrity": "sha512-v5l/aFXZQeai4awLbOpSoHecE9UiMrnfx75tEXLjNonXVARxQ5mOeipTjROUchszUNCqnE+hqAMujRsRHsut2Q==", "license": "MIT", "funding": { "url": "https://github.com/sponsors/colinhacks" diff --git a/package.json b/package.json index d3b99e0f8f..51931a9565 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@modelcontextprotocol/inspector", - "version": "2.9.0", + "version": "2.10.0", "description": "The Model Context Protocol Inspector", "keywords": [ "MCP", @@ -18,7 +18,8 @@ "author": "The MCP Maintainers and Community", "type": "module", "bin": { - "mcp-inspector": "./clients/launcher/build/index.js" + "mcp-inspector": "./clients/launcher/build/index.js", + "mcpdo": "./clients/mcpdo/build/mcp-bin.js" }, "files": [ "clients/launcher/build", @@ -26,21 +27,26 @@ "clients/web/dist", "clients/web/static", "clients/cli/build", + "clients/mcpdo/build", "clients/tui/build", - "scripts/install-clients.mjs" + "scripts/install-clients.mjs", + "skills/mcpdo" ], "scripts": { "web": "node clients/launcher/build/index.js --web", "build:web:runner": "cd clients/web && npm run build:runner", "web:dev": "npm run build:web:runner && node clients/launcher/build/index.js --web --dev", - "build": "npm run build:web && npm run build:cli && npm run build:tui && npm run build:launcher", + "build": "npm run build:web && npm run build:cli && npm run build:mcpdo && npm run build:tui && npm run build:launcher", "build:cli": "cd clients/cli && npm run build", + "build:mcpdo": "cd clients/mcpdo && npm run build", + "build:mcpdo:dev": "cd clients/mcpdo && npm run build:dev", "build:tui": "cd clients/tui && npm run build", "build:web": "cd clients/web && npm run build", "build:launcher": "cd clients/launcher && npm run build", "local:gate": "node scripts/gate-lease.mjs npm run local:gate:stages", - "local:gate:stages": "npm run local:validate && npm run verify:skills:cli && npm run coverage && npm run verify:build-gate && npm run verify:bundle-externals && npm run smoke && npm run smoke:web:firefox && npm run local:storybook", - "local:validate": "npm run validate:guards && npm run validate:core && cd clients/web && npm run check && cd ../cli && npm run check && cd ../tui && npm run check && cd ../launcher && npm run check", + "local:gate:stages": "npm run local:validate && npm run local:dco && npm run verify:skills:cli && npm run coverage && npm run verify:build-gate && npm run verify:bundle-externals && npm run smoke && npm run smoke:web:firefox && npm run local:storybook", + "local:validate": "npm run validate:guards && npm run validate:core && cd clients/web && npm run check && cd ../cli && npm run check && cd ../mcpdo && npm run check && cd ../tui && npm run check && cd ../launcher && npm run check", + "local:dco": "node scripts/dco-check.mjs --base origin/v2/main --exclude origin/main", "local:storybook": "cd clients/web && npx playwright install chromium && npm run test:storybook", "verify:build-gate": "node scripts/verify-build-gate.mjs", "verify:bundle-externals": "node scripts/verify-bundle-externals.mjs", @@ -48,7 +54,25 @@ "verify:skills": "node scripts/verify-skills.mjs", "verify:skills:cli": "node scripts/verify-skills-cli.mjs", "test:scripts": "node --test \"scripts/**/*.test.mjs\"", - "validate": "npm run validate:guards && npm run validate:core && npm run validate:web && npm run validate:cli && npm run validate:tui && npm run validate:launcher", + "pr:review-request": "node scripts/pr-review-request.mjs", + "pr:review-wait": "node scripts/pr-review-wait.mjs", + "pr:review-fetch": "node scripts/pr-review-fetch.mjs", + "pr:link": "node scripts/pr-link-issue.mjs", + "dco:check": "node scripts/dco-check.mjs", + "pr:upload": "node scripts/pr-upload-screenshot.mjs", + "board:status": "node scripts/board-card-status.mjs", + "board:add": "node scripts/board-card-add.mjs", + "board:delete": "node scripts/board-card-delete.mjs", + "board:find-draft": "node scripts/board-draft-find.mjs", + "board:snapshot": "node scripts/board-snapshot.mjs", + "board:sweep": "node scripts/board-sweep.mjs", + "board:audit": "node scripts/board-audit.mjs", + "board:recover": "node scripts/board-recover.mjs", + "advisory:fork": "node scripts/advisory-fork.mjs", + "release:notes": "node scripts/release-notes.mjs", + "release:tag": "node scripts/release-tag.mjs", + "action:resolve-pin": "node scripts/action-pin-resolve.mjs", + "validate": "npm run validate:guards && npm run validate:core && npm run validate:web && npm run validate:cli && npm run validate:mcpdo && npm run validate:tui && npm run validate:launcher", "validate:guards": "npm run verify:install-fresh && npm run verify:format-coverage && npm run verify:skills && npm run verify:typecheck-coverage && npm run verify:dep-lockstep && npm run verify:test-timeouts && npm run verify:action-pins && npm run test:scripts", "verify:format-coverage": "node scripts/verify-format-coverage.mjs", "verify:dep-lockstep": "node scripts/verify-dep-lockstep.mjs", @@ -64,18 +88,21 @@ "format:check:scripts": "prettier --check \"scripts/**/*.{ts,tsx,mts,cts,js,jsx,mjs,cjs}\"", "format:shared": "prettier --write \"test-servers/src/**/*.{ts,tsx,mts,cts}\" vitest.shared.mts vitest.setup.shared.mts eslint.config.js", "format:check:shared": "prettier --check \"test-servers/src/**/*.{ts,tsx,mts,cts}\" vitest.shared.mts vitest.setup.shared.mts eslint.config.js", - "format": "npm run format:core && npm run format:scripts && npm run format:shared && cd clients/web && npm run format && cd ../cli && npm run format && cd ../tui && npm run format && cd ../launcher && npm run format", + "format": "npm run format:core && npm run format:scripts && npm run format:shared && cd clients/web && npm run format && cd ../cli && npm run format && cd ../mcpdo && npm run format && cd ../tui && npm run format && cd ../launcher && npm run format", "validate:cli": "cd clients/cli && npm run validate", + "validate:mcpdo": "cd clients/mcpdo && npm run validate", "validate:tui": "cd clients/tui && npm run validate", "validate:web": "cd clients/web && npm run validate", "validate:launcher": "cd clients/launcher && npm run validate", - "coverage": "npm run coverage:web && npm run coverage:cli && npm run coverage:tui && npm run coverage:launcher", + "coverage": "npm run coverage:web && npm run coverage:cli && npm run coverage:mcpdo && npm run coverage:tui && npm run coverage:launcher", "coverage:cli": "cd clients/cli && npm run test:coverage", + "coverage:mcpdo": "cd clients/mcpdo && npm run test:coverage", "coverage:tui": "cd clients/tui && npm run test:coverage", "coverage:web": "cd clients/web && npm run test:coverage", "coverage:launcher": "cd clients/launcher && npm run test:coverage", - "smoke": "npm run smoke:launcher && npm run smoke:cli && npm run smoke:tui && npm run smoke:web && npm run smoke:web:chromium && npm run smoke:web:tabs", + "smoke": "npm run smoke:launcher && npm run smoke:cli && npm run smoke:mcpdo && npm run smoke:tui && npm run smoke:web && npm run smoke:web:chromium && npm run smoke:web:tabs", "smoke:cli": "node scripts/smoke-cli.mjs", + "smoke:mcpdo": "node scripts/smoke-mcpdo.mjs", "smoke:tui": "node scripts/smoke-tui.mjs", "smoke:web": "node scripts/smoke-web.mjs", "smoke:web:browser": "node scripts/install-smoke-browser.mjs && node scripts/smoke-web-browser.mjs", @@ -90,13 +117,15 @@ "pack:verify": "node scripts/install-smoke-browser.mjs chromium && node scripts/pack-and-verify.mjs", "prepack": "npm run build", "postinstall": "node scripts/install-clients.mjs", - "skills:eval": "node scripts/skill-eval.mjs" + "skills:eval": "node scripts/skill-eval.mjs", + "skills:eval:mcpdo": "node scripts/skill-eval-mcpdo.mjs" }, "dependencies": { "@hono/node-server": "^2.0.12", "@modelcontextprotocol/client": "2.2.0", "@modelcontextprotocol/core": "2.2.0", "@modelcontextprotocol/ext-apps": "^2.0.0", + "@modelcontextprotocol/ext-tasks": "^0.2.2", "@modelcontextprotocol/server": "2.2.0", "@napi-rs/keyring": "^1.3.0", "@vitejs/plugin-react": "^6.0.0", diff --git a/scripts/action-pin-resolve.mjs b/scripts/action-pin-resolve.mjs new file mode 100644 index 0000000000..e1fa8d6037 --- /dev/null +++ b/scripts/action-pin-resolve.mjs @@ -0,0 +1,94 @@ +#!/usr/bin/env node +// Resolve an action's SHA pin and its exact-version comment from ONE tag +// lookup (#2558) — `npm run action:resolve-pin -- --repo actions/checkout +// --tag v5`. The resolver block the pre-push-gate skill previously +// transcribed inline. +// +// `verify:action-pins` is offline: it checks a credentialed job's `uses:` +// lines are `SHA # vX.Y.Z` pins but cannot check the SHA and the comment +// agree. This script is what makes them agree by construction — the SHA and +// the exact release both come from the same tag listing, so the comment the +// monthly sweep ranks by can never drift from the commit actually pinned. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { ghPaginatedList } from "./lib/gh.mjs"; + +const EXACT_TAG = /^v\d+\.\d+\.\d+$/; + +export function parsePinArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { repo: { type: "string" }, tag: { type: "string" } }, + }); + if (!values.repo || !/^[\w.-]+\/[\w.-]+$/.test(values.repo)) { + throw new Error( + `--repo must be owner/name, got ${values.repo ?? "nothing"}`, + ); + } + // The tag names a release line (a vN moving tag, or an exact vX.Y.Z) — + // its major is what the exact-version lookup is restricted to below. + // ONLY those two forms: a partial tag like v5.1 is an undocumented minor + // moving tag the exact-tag preservation above cannot reason about. + const major = /^v(\d+)(?:\.\d+\.\d+)?$/.exec(values.tag ?? "")?.[1]; + if (major === undefined) { + throw new Error( + `--tag must be a vN moving tag or exact vX.Y.Z, got ${values.tag ?? "nothing"}`, + ); + } + return { repo: values.repo, tag: values.tag, major: Number(major) }; +} + +/** + * The highest exact vX.Y.Z tag pointing at `sha` WITHIN the requested major + * — a commit can carry exact tags from several majors (a lagging line + * re-released from the same tree), and the comment must identify the release + * line that was asked for, not the numerically highest. + */ +export function exactVersionFor(tags, sha, major) { + const exact = tags + .filter( + (tag) => + EXACT_TAG.test(tag?.name ?? "") && + tag?.commit?.sha === sha && + Number(tag.name.slice(1).split(".")[0]) === major, + ) + .map((tag) => tag.name) + .sort((a, b) => { + const pa = a.slice(1).split(".").map(Number); + const pb = b.slice(1).split(".").map(Number); + return pa[0] - pb[0] || pa[1] - pb[1] || pa[2] - pb[2]; + }); + return exact.at(-1); +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { repo, tag, major } = parsePinArgs(argv); + + // ONE listing resolves both the SHA and the exact version. Resolving the + // moving tag through /commits/<tag> in a separate request would race an + // upstream retag between the two calls — the printed comment could name + // the previous release while the SHA pins the new one. + const tags = ghPaginatedList(spawn, `repos/${repo}/tags?per_page=100`); + const sha = tags.find((candidate) => candidate?.name === tag)?.commit?.sha; + if (!sha) { + throw new Error( + `could not resolve ${repo}@${tag} to a commit — no such tag`, + ); + } + const version = exactVersionFor(tags, sha, major); + if (!version) { + throw new Error( + `no exact v${major}.Y.Z tag in ${repo} points at ${sha} — pin by hand from the release page`, + ); + } + // An exact requested tag IS the version — re-deriving it from the SHA + // could mislabel it when several exact tags share one commit (v5.1.2 + // re-released unchanged as v5.2.0): the comment must name what was asked + // for, not the highest tag that happens to sit on the same tree. + console.log(`uses: ${repo}@${sha} # ${EXACT_TAG.test(tag) ? tag : version}`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/action-pin-resolve.test.mjs b/scripts/action-pin-resolve.test.mjs new file mode 100644 index 0000000000..e393778834 --- /dev/null +++ b/scripts/action-pin-resolve.test.mjs @@ -0,0 +1,117 @@ +// Tests for scripts/action-pin-resolve.mjs (#2558) — the SHA and the exact +// version come from the SAME tag listing (the offline guard cannot check they +// agree), numeric semver selection, and the no-exact-tag refusal. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { exactVersionFor, main, parsePinArgs } from "./action-pin-resolve.mjs"; + +test("parsePinArgs validates the repo slug and the tag shape", () => { + assert.deepEqual( + parsePinArgs(["--repo", "actions/checkout", "--tag", "v5"]), + { + repo: "actions/checkout", + tag: "v5", + major: 5, + }, + ); + assert.equal(parsePinArgs(["--repo", "a/b", "--tag", "v5.1.2"]).major, 5); + assert.throws( + () => parsePinArgs(["--repo", "checkout", "--tag", "v5"]), + /owner\/name/, + ); + assert.throws(() => parsePinArgs(["--repo", "a/b"]), /--tag/); + assert.throws( + () => parsePinArgs(["--repo", "a/b", "--tag", "main"]), + /vN moving tag/, + ); + // Only the two documented forms — a partial tag such as v5.1 is an + // undocumented minor moving tag and must be rejected, not resolved. + assert.throws( + () => parsePinArgs(["--repo", "a/b", "--tag", "v5.1"]), + /vN moving tag/, + ); +}); + +const SHA = "deadbeef"; +const tag = (name, sha = SHA) => ({ name, commit: { sha } }); + +test("exactVersionFor picks the highest exact tag on the SHA, numerically", () => { + // v5.10.0 > v5.9.1 numerically though not lexically. + assert.equal( + exactVersionFor( + [tag("v5"), tag("v5.9.1"), tag("v5.10.0"), tag("v4.9.9", "other")], + SHA, + 5, + ), + "v5.10.0", + ); + assert.equal(exactVersionFor([tag("v5")], SHA, 5), undefined); +}); + +test("exactVersionFor stays within the requested major", () => { + // One commit can carry exact tags from several majors (a lagging line + // re-released from the same tree) — the comment must name the line asked + // for, not the numerically highest. + const tags = [tag("v5.2.0"), tag("v6.0.0")]; + assert.equal(exactVersionFor(tags, SHA, 5), "v5.2.0"); + assert.equal(exactVersionFor(tags, SHA, 6), "v6.0.0"); + assert.equal(exactVersionFor(tags, SHA, 4), undefined); +}); + +function spawnScript({ tags }) { + return (cmd, args) => { + const joined = args.join(" "); + if (!joined.includes("/tags")) { + // The SHA must come from the same /tags listing as the version — a + // separate /commits/<tag> request would race an upstream retag. + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify([tags]), stderr: "" }; + }; +} + +test("main prints the uses: line with SHA and matching exact version", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + main( + ["--repo", "actions/checkout", "--tag", "v5"], + spawnScript({ tags: [tag("v5"), tag("v5.0.1")] }), + ); + assert.deepEqual(lines, [`uses: actions/checkout@${SHA} # v5.0.1`]); +}); + +test("main preserves an exact requested tag over a same-SHA higher tag", (t) => { + // v5.1.2 re-released unchanged as v5.2.0 shares its commit; asking for + // v5.1.2 must print v5.1.2, not the highest tag on the same tree. + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + main( + ["--repo", "actions/checkout", "--tag", "v5.1.2"], + spawnScript({ tags: [tag("v5"), tag("v5.1.2"), tag("v5.2.0")] }), + ); + assert.deepEqual(lines, [`uses: actions/checkout@${SHA} # v5.1.2`]); +}); + +test("main throws when the requested tag is not in the listing", () => { + assert.throws( + () => + main( + ["--repo", "actions/checkout", "--tag", "v9"], + spawnScript({ tags: [tag("v5"), tag("v5.0.1")] }), + ), + /could not resolve actions\/checkout@v9 .* no such tag/, + ); +}); + +test("main throws when no exact tag in the requested major points at the SHA", () => { + assert.throws( + () => + main( + ["--repo", "actions/checkout", "--tag", "v5"], + spawnScript({ tags: [tag("v5"), tag("v6.0.0")] }), + ), + /no exact v5\.Y\.Z tag/, + ); +}); diff --git a/scripts/advisory-fork.mjs b/scripts/advisory-fork.mjs new file mode 100644 index 0000000000..917bd3b9e4 --- /dev/null +++ b/scripts/advisory-fork.mjs @@ -0,0 +1,78 @@ +#!/usr/bin/env node +// Look up — or create only when absent — a security advisory's private fork +// (#2558) — `npm run advisory:fork -- --ghsa GHSA-xxxx-yyyy-zzzz [--create]`. +// The fork block the security-advisory skill previously transcribed inline. +// +// The ordering is the entire point of scripting this: the POST **creates** a +// private fork as a side effect, and an accidentally created fork needs the +// `delete_repo` scope to remove. So the advisory is READ FIRST, an existing +// fork is printed and the POST never runs, and creating one at all requires +// the explicit `--create` flag — a bare lookup can never mutate anything. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { REPO_SLUG, ghJson } from "./lib/gh.mjs"; + +const GHSA_PATTERN = + /^GHSA-[23456789cfghjmpqrvwx]{4}-[23456789cfghjmpqrvwx]{4}-[23456789cfghjmpqrvwx]{4}$/; + +export function parseForkArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { ghsa: { type: "string" }, create: { type: "boolean" } }, + }); + if (!values.ghsa || !GHSA_PATTERN.test(values.ghsa)) { + throw new Error( + `--ghsa must be a full GHSA id (GHSA-xxxx-yyyy-zzzz), got ${values.ghsa ?? "nothing"}`, + ); + } + return { ghsa: values.ghsa, create: values.create === true }; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { ghsa, create } = parseForkArgs(argv); + + // Read first — the create endpoint is never touched when a fork exists. + const advisory = ghJson(spawn, [ + "api", + `repos/${REPO_SLUG}/security-advisories/${ghsa}`, + ]); + const existing = advisory?.private_fork?.full_name; + if (existing) { + console.log(`fork: ${existing} (existing)`); + return; + } + if (!create) { + console.log(`fork: none — re-run with --create to make one for ${ghsa}`); + return; + } + + const fork = ghJson(spawn, [ + "api", + "-X", + "POST", + `repos/${REPO_SLUG}/security-advisories/${ghsa}/forks`, + ]); + // The 202 response's shape has varied (a repository object vs the advisory + // with private_fork nested) — accept either, and never report failure for + // a POST that succeeded: the fork now exists whatever the payload said. + const created = fork?.full_name ?? fork?.private_fork?.full_name; + if (created) { + console.log(`fork: ${created} (created)`); + return; + } + // Creation is asynchronous — re-read the advisory for the name. + const after = ghJson(spawn, [ + "api", + `repos/${REPO_SLUG}/security-advisories/${ghsa}`, + ])?.private_fork?.full_name; + console.log( + after + ? `fork: ${after} (created)` + : `fork: created, name pending — re-run without --create to confirm`, + ); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/advisory-fork.test.mjs b/scripts/advisory-fork.test.mjs new file mode 100644 index 0000000000..2bb5459807 --- /dev/null +++ b/scripts/advisory-fork.test.mjs @@ -0,0 +1,109 @@ +// Tests for scripts/advisory-fork.mjs (#2558) — the read-first ordering (the +// POST creates a fork as a side effect and needs delete_repo scope to undo) +// and the explicit --create gate. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { main, parseForkArgs } from "./advisory-fork.mjs"; + +const GHSA = "GHSA-2345-cfgh-jmpq"; + +test("parseForkArgs requires a full GHSA id and defaults create off", () => { + assert.deepEqual(parseForkArgs(["--ghsa", GHSA]), { + ghsa: GHSA, + create: false, + }); + assert.equal(parseForkArgs(["--ghsa", GHSA, "--create"]).create, true); + assert.throws(() => parseForkArgs(["--ghsa", "nope"]), /full GHSA id/); +}); + +function spawnScript({ fork = null, postPayload, afterFork } = {}) { + const calls = []; + let posted = false; + const spawn = (cmd, args) => { + calls.push(args); + const joined = args.join(" "); + let payload; + if (joined.includes("POST")) { + posted = true; + payload = postPayload ?? { + full_name: "modelcontextprotocol/inspector-ghsa-fork", + }; + } else if (joined.includes("security-advisories")) { + payload = { private_fork: posted ? (afterFork ?? fork) : fork }; + } else { + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("an existing fork is printed and the POST never runs", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript({ fork: { full_name: "mcp/fork-x" } }); + main(["--ghsa", GHSA, "--create"], spawn); + assert.deepEqual(lines, ["fork: mcp/fork-x (existing)"]); + assert.equal( + spawn.calls.some((args) => args.includes("POST")), + false, + ); +}); + +test("without --create a missing fork is reported, not created", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript(); + main(["--ghsa", GHSA], spawn); + assert.match(lines[0], /fork: none — re-run with --create/); + assert.equal( + spawn.calls.some((args) => args.includes("POST")), + false, + ); +}); + +test("with --create a missing fork is created and printed", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + main(["--ghsa", GHSA, "--create"], spawnScript()); + assert.deepEqual(lines, [ + "fork: modelcontextprotocol/inspector-ghsa-fork (created)", + ]); +}); + +test("a POST answering with the advisory shape still reports the fork", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + main( + ["--ghsa", GHSA, "--create"], + spawnScript({ + postPayload: { private_fork: { full_name: "mcp/nested-fork" } }, + }), + ); + assert.deepEqual(lines, ["fork: mcp/nested-fork (created)"]); +}); + +test("a nameless POST response re-reads the advisory, and never fails a fork that was made", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + // Re-read finds the name: + main( + ["--ghsa", GHSA, "--create"], + spawnScript({ + postPayload: {}, + afterFork: { full_name: "mcp/async-fork" }, + }), + ); + // Re-read still pending — reported as created, not as a failure, since the + // POST succeeded and the fork now exists: + main( + ["--ghsa", GHSA, "--create"], + spawnScript({ postPayload: {}, afterFork: null }), + ); + assert.deepEqual(lines, [ + "fork: mcp/async-fork (created)", + "fork: created, name pending — re-run without --create to confirm", + ]); +}); diff --git a/scripts/board-audit.mjs b/scripts/board-audit.mjs new file mode 100644 index 0000000000..379797b6a1 --- /dev/null +++ b/scripts/board-audit.mjs @@ -0,0 +1,202 @@ +#!/usr/bin/env node +// Audit both project boards against their invariants (#2558) — `npm run +// board:audit`. The audit pass of the issue-triage skill, which previously +// transcribed this as a ~70-line mktemp + jq block that had to be reproduced +// verbatim every run. The invariants themselves are AGENTS.md rules +// ("Issue-driven Work Style") and the issue-triage skill's audit table; this +// script is their mechanical form. +// +// Every dump is trusted only when complete: both board listings go through +// `itemListComplete`, and the issue dump is refused at its own --limit, so a +// silent truncation reads as an error instead of as a clean audit (#2451). +// +// Output is one line per check — `<count>\t<check>\t<first few offenders>` — +// and the exit code is non-zero when any check has offenders, so "the board +// is clean" is scriptable. + +import { spawnSync } from "node:child_process"; +import { REPO_SLUG, ghJson } from "./lib/gh.mjs"; +import { itemListComplete } from "./lib/board.mjs"; + +const ISSUE_LIMIT = 2000; +const V2_BOARD = 28; +const V1_BOARD = 11; +const TYPE_LABELS = [ + "bug", + "enhancement", + "documentation", + "chore", + "question", +]; +const SHOW = 10; + +/** Keep this repo's items — and drafts, whose repository is null. */ +export function own(items) { + return items.filter( + (item) => + item?.content?.repository == null || + item.content.repository === REPO_SLUG, + ); +} + +const isIssue = (item) => item?.content?.type === "Issue"; +const isGhsaDraft = (item) => + item?.content?.type === "DraftIssue" && + (item.content.title ?? "").startsWith("[GHSA-"); +const num = (item) => `#${item.content.number}`; + +/** + * Run every invariant over complete dumps. `issues` is the all-state issue + * dump; `v2`/`v1` are each board's own() items. Returns + * [{ check, offenders }] with offenders as display strings. + */ +export function auditChecks(issues, v2, v1) { + const byNumber = new Map(issues.map((issue) => [issue.number, issue])); + const issueOf = (item) => byNumber.get(item.content.number); + const open = (item) => issueOf(item)?.state === "OPEN"; + const closed = (item) => issueOf(item)?.state === "CLOSED"; + const labels = (item) => + (issueOf(item)?.labels ?? []).map((label) => label.name); + const milestone = (item) => issueOf(item)?.milestone != null; + + const v2Issues = v2.filter(isIssue); + const v1Issues = v1.filter(isIssue); + const v2Numbers = new Set(v2Issues.map((item) => item.content.number)); + + const openIssues = issues.filter((issue) => issue.state === "OPEN"); + const labelCount = (issue, names) => + (issue.labels ?? []).filter((label) => names.includes(label.name)).length; + + return [ + { + check: "double-boarded (a card on both #28 and #11)", + offenders: v1Issues + .filter((item) => v2Numbers.has(item.content.number)) + .map(num), + }, + { + check: "non-issue card (only [GHSA-…] drafts on #28 are allowed)", + offenders: [ + ...v2.filter((item) => !isIssue(item) && !isGhsaDraft(item)), + ...v1.filter((item) => !isIssue(item)), + ].map((item) => item.content?.title ?? item.id), + }, + { + check: "GHSA draft missing Status or Priority (#28)", + offenders: v2 + .filter(isGhsaDraft) + .filter((item) => item.status == null || item.priority == null) + .map((item) => item.content.title), + }, + { + check: "card with no Status", + offenders: [...v2, ...v1] + .filter((item) => item.status == null && !isGhsaDraft(item)) + .map((item) => + isIssue(item) ? num(item) : (item.content?.title ?? item.id), + ), + }, + { + check: "Incoming but milestoned (#28)", + offenders: v2Issues + .filter((item) => item.status === "Incoming" && milestone(item)) + .map(num), + }, + { + check: "past Incoming but no milestone (open, #28)", + offenders: v2Issues + .filter( + (item) => + item.status != null && + item.status !== "Incoming" && + item.status !== "Done" && + open(item) && + !milestone(item), + ) + .map(num), + }, + { + check: "v1-labeled issue on #28 (open)", + offenders: v2Issues + .filter((item) => open(item) && labels(item).includes("v1")) + .map(num), + }, + { + check: "v2-labeled issue on #11 (open)", + offenders: v1Issues + .filter((item) => open(item) && labels(item).includes("v2")) + .map(num), + }, + { + check: "open issue without exactly one version label (v1/v2)", + offenders: openIssues + .filter((issue) => labelCount(issue, ["v1", "v2"]) !== 1) + .map((issue) => `#${issue.number}`), + }, + { + check: `open issue without exactly one type label (${TYPE_LABELS.join("/")})`, + offenders: openIssues + .filter((issue) => labelCount(issue, TYPE_LABELS) !== 1) + .map((issue) => `#${issue.number}`), + }, + { + check: "open #28 card with no Priority", + offenders: v2Issues + .filter((item) => open(item) && item.priority == null) + .map(num), + }, + { + check: "closed-unshipped issue still carded (delete the card)", + offenders: [...v2Issues, ...v1Issues] + .filter( + (item) => closed(item) && issueOf(item)?.stateReason !== "COMPLETED", + ) + .map(num), + }, + { + check: "open issue carded Done", + offenders: [...v2Issues, ...v1Issues] + .filter((item) => open(item) && item.status === "Done") + .map(num), + }, + ]; +} + +export function main(_argv = process.argv.slice(2), spawn = spawnSync) { + const issues = ghJson(spawn, [ + "issue", + "list", + "--repo", + REPO_SLUG, + "--state", + "all", + "--limit", + String(ISSUE_LIMIT), + "--json", + "number,state,stateReason,labels,milestone", + ]); + if (issues.length >= ISSUE_LIMIT) { + throw new Error( + `issue listing hit --limit ${ISSUE_LIMIT} — raise it and re-run`, + ); + } + + const v2 = own(itemListComplete(spawn, V2_BOARD).items); + const v1 = own(itemListComplete(spawn, V1_BOARD).items); + + let dirty = false; + for (const { check, offenders } of auditChecks(issues, v2, v1)) { + const shown = offenders.slice(0, SHOW).join(" "); + console.log(`${offenders.length}\t${check}${shown ? `\t${shown}` : ""}`); + if (offenders.length > 0) { + dirty = true; + } + } + if (dirty) { + process.exitCode = 1; + } +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/board-audit.test.mjs b/scripts/board-audit.test.mjs new file mode 100644 index 0000000000..04357ec486 --- /dev/null +++ b/scripts/board-audit.test.mjs @@ -0,0 +1,165 @@ +// Tests for scripts/board-audit.mjs (#2558) — the own() repository filter +// (null keeps drafts — load-bearing), each invariant over a crafted fixture, +// and main()'s output/exit contract. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { auditChecks, main, own } from "./board-audit.mjs"; + +const SLUG = "modelcontextprotocol/inspector"; + +test("own keeps this repo's issues AND drafts (repository null)", () => { + const ours = { content: { type: "Issue", number: 1, repository: SLUG } }; + const draft = { content: { type: "DraftIssue", title: "[GHSA-…]" } }; + const foreign = { content: { type: "Issue", number: 9, repository: "o/r" } }; + assert.deepEqual(own([ours, draft, foreign]), [ours, draft]); +}); + +const issue = (number, over = {}) => ({ + number, + state: "OPEN", + stateReason: null, + labels: [{ name: "v2" }, { name: "bug" }], + milestone: { title: "2.5.0" }, + ...over, +}); +const card = (number, status, priority = "Medium") => ({ + id: `PVTI_${number}`, + status, + priority, + content: { type: "Issue", number, repository: SLUG }, +}); + +function checksByName(issues, v2, v1) { + return new Map( + auditChecks(issues, v2, v1).map(({ check, offenders }) => [ + check.split(" (")[0], + offenders, + ]), + ); +} + +test("a clean board produces no offenders anywhere", () => { + const issues = [ + issue(1), + issue(2, { labels: [{ name: "v1" }, { name: "chore" }], milestone: null }), + ]; + const checks = checksByName( + issues, + [card(1, "Todo")], + [{ ...card(2, "Todo"), priority: undefined }], + ); + for (const [name, offenders] of checks) { + assert.deepEqual(offenders, [], `check "${name}" should be clean`); + } +}); + +test("each invariant catches its own defect", () => { + const ghsaDraft = { + id: "PVTI_draft", + status: null, + priority: null, + content: { type: "DraftIssue", title: "[GHSA-2345-cfgh-jmpq] - x" }, + }; + const plainDraft = { + id: "PVTI_plain", + status: "Todo", + priority: "Low", + content: { type: "DraftIssue", title: "remember to tidy up" }, + }; + const issues = [ + issue(1), // double-boarded below + issue(2, { milestone: { title: "2.5.0" } }), // Incoming but milestoned + issue(3, { milestone: null }), // past Incoming, no milestone + issue(4, { labels: [{ name: "v1" }, { name: "bug" }] }), // v1 on #28 + issue(5, { labels: [{ name: "v2" }, { name: "bug" }] }), // v2 on #11 + issue(6, { labels: [{ name: "bug" }] }), // no version label + issue(7, { labels: [{ name: "v2" }] }), // no type label + issue(8, { priorityless: true }), // open, no Priority (via card below) + issue(9, { state: "CLOSED", stateReason: "NOT_PLANNED" }), // closed unshipped + issue(10), // open but Done + ]; + const v2 = [ + card(1, "Todo"), + card(2, "Incoming"), + card(3, "In Progress"), + card(4, "Todo"), + card(6, "Todo"), + card(7, "Todo"), + { ...card(8, "Todo"), priority: null }, + card(9, "Todo"), + card(10, "Done"), + ghsaDraft, + plainDraft, + ]; + const v1 = [card(1, "Todo"), card(5, "Todo")]; + + const checks = checksByName(issues, v2, v1); + assert.deepEqual(checks.get("double-boarded"), ["#1"]); + assert.deepEqual(checks.get("non-issue card"), ["remember to tidy up"]); + assert.deepEqual(checks.get("GHSA draft missing Status or Priority"), [ + "[GHSA-2345-cfgh-jmpq] - x", + ]); + assert.deepEqual(checks.get("Incoming but milestoned"), ["#2"]); + assert.deepEqual(checks.get("past Incoming but no milestone"), ["#3"]); + assert.deepEqual(checks.get("v1-labeled issue on #28"), ["#4"]); + // #1 (double-boarded, v2-labeled) legitimately also trips the #11 check. + assert.deepEqual(checks.get("v2-labeled issue on #11"), ["#1", "#5"]); + assert.deepEqual(checks.get("open issue without exactly one version label"), [ + "#6", + ]); + assert.deepEqual(checks.get("open issue without exactly one type label"), [ + "#7", + ]); + assert.deepEqual(checks.get("open #28 card with no Priority"), ["#8"]); + assert.deepEqual(checks.get("closed-unshipped issue still carded"), ["#9"]); + assert.deepEqual(checks.get("open issue carded Done"), ["#10"]); +}); + +test("a GHSA draft with no Status is not double-counted as 'no Status'", () => { + const ghsaDraft = { + id: "PVTI_draft", + status: null, + priority: null, + content: { type: "DraftIssue", title: "[GHSA-2345-cfgh-jmpq] - x" }, + }; + const checks = checksByName([], [ghsaDraft], []); + assert.deepEqual(checks.get("card with no Status"), []); + assert.equal(checks.get("GHSA draft missing Status or Priority").length, 1); +}); + +/** A spawn answering the issue dump and both board dumps. */ +function spawnScript({ issues, v2 = [], v1 = [] }) { + return (cmd, args) => { + const joined = args.join(" "); + let payload; + if (joined.includes("issue list")) { + payload = issues; + } else if (joined.includes("item-list")) { + const items = args[2] === "28" ? v2 : v1; + payload = { items, totalCount: items.length }; + } else { + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; +} + +test("main prints one line per check and exits non-zero when dirty", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const before = process.exitCode; + main([], spawnScript({ issues: [issue(1)], v2: [], v1: [] })); + // 13 checks, one line each; #1 is open+milestoned but unboarded is NOT an + // audit concern (that is the sweep's), so the only hits are none → clean? + // No: issue 1 is not carded anywhere, so every per-card check is clean and + // the per-issue label checks pass — the audit is clean and exits 0. + assert.equal(lines.length, 13); + assert.equal(process.exitCode, before); + + lines.length = 0; + main([], spawnScript({ issues: [issue(1)], v2: [card(1, "Done")], v1: [] })); + assert.ok(lines.some((line) => line.startsWith("1\topen issue carded Done"))); + assert.equal(process.exitCode, 1); + process.exitCode = before; +}); diff --git a/scripts/board-card-add.mjs b/scripts/board-card-add.mjs new file mode 100644 index 0000000000..cfc0612bd8 --- /dev/null +++ b/scripts/board-card-add.mjs @@ -0,0 +1,209 @@ +#!/usr/bin/env node +// Add an issue's card to a board and set its fields (#2558) — `npm run +// board:add -- --issue <N> --status Todo [--priority Medium] [--board 28]`. +// The add-card recipe board-ops previously transcribed inline (and +// issue-create step 4 points at). +// +// Same properties as `board-card-status.mjs`: every id resolved by name at +// run time, and the Status VERIFIED by reading it back — `card: …` prints +// only on a confirmed match. On board #28 `--priority` is REQUIRED — every +// v2 board item has a Priority (AGENTS.md), so an add without one would mint +// a card `board:audit` immediately flags. Board #11 has no Priority field, +// so there the flag is refused by name resolution instead. +// +// The command is IDEMPOTENT rather than rolling back on failure. `item-add` +// is itself idempotent — two invocations (or a concurrent add) can both +// receive the same item id — so this script can never prove it created the +// card it holds an id for, and deleting on failure could destroy a card a +// concurrent operation owns. Instead: an existing card whose requested +// fields are unset or already match is (re)configured in place — which is +// also what lets a re-run finish a card a failed edit left partial — and an +// existing card holding a CONTRADICTING value is refused before any write. +// +// The issue is also resolved as an ISSUE before the first write: issue and +// PR numbers share one namespace, `/issues/<PR number>` redirects to the PR, +// and `item-add` accepts it — which would board a PR (forbidden) and fail +// only at the later verify, leaving the bad card behind. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { OWNER, REPO, ghJson, requirePositiveInt } from "./lib/gh.mjs"; +import { + DEFAULT_BOARD, + boardFields, + editItemField, + fieldOption, + findCard, + projectId as resolveProjectId, + requireSupportedBoard, +} from "./lib/board.mjs"; + +export function parseAddArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { + issue: { type: "string" }, + status: { type: "string" }, + priority: { type: "string" }, + board: { type: "string" }, + }, + }); + if (!values.status) { + throw new Error("--status is required (e.g. --status Todo)"); + } + const board = requireSupportedBoard( + values.board === undefined + ? DEFAULT_BOARD + : requirePositiveInt(values.board, "--board"), + ); + if (board === DEFAULT_BOARD && values.priority === undefined) { + throw new Error( + "--priority is required on board #28 — every v2 board item has a Priority (derive it with the issue-triage rubric); only board #11, which has no Priority field, omits it", + ); + } + return { + issue: requirePositiveInt(values.issue, "--issue"), + status: values.status, + priority: values.priority, + board, + }; +} + +/** + * Look the issue's card up and refuse when a requested field holds a + * DIFFERENT value — the real duplicate-add mistake (a wrong issue number, an + * issue already moving through the board), where proceeding would overwrite + * board state someone else set. Returns the card's item id when it exists + * with unset/matching fields (safe to configure in place), else undefined. + */ +function assertNoConflicts(spawn, issue, project, board, status, priority) { + const existing = findCard(spawn, issue, project); + if (!existing?.id) { + return undefined; + } + const conflicts = []; + const statusNow = existing.fieldValueByName?.name; + if (statusNow != null && statusNow !== status) { + conflicts.push(`Status "${statusNow}"`); + } + if (priority !== undefined) { + const priorityNow = findCard(spawn, issue, project, "Priority") + ?.fieldValueByName?.name; + if (priorityNow != null && priorityNow !== priority) { + conflicts.push(`Priority "${priorityNow}"`); + } + } + if (conflicts.length > 0) { + throw new Error( + `#${issue} already has a card on board #${board} reading ${conflicts.join(" and ")} — not adding a duplicate or overwriting a configured card; move it with board:status`, + ); + } + return existing.id; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { issue, status, priority, board } = parseAddArgs(argv); + + const project = resolveProjectId(spawn, board); + const fields = boardFields(spawn, board); + // Resolve EVERY option before the first write, so a bad name cannot leave + // a half-configured card behind. + const statusIds = fieldOption(fields, "Status", status); + const priorityIds = + priority === undefined + ? undefined + : fieldOption(fields, "Priority", priority); + + // Resolve the number as an ISSUE before the first write — findCard queries + // repository.issue(number:), so a PR number (same namespace, and item-add + // would accept its URL) fails here instead of boarding a forbidden PR card. + // An existing card with unset/matching fields is configured in place: that + // is a re-run finishing this command's own earlier failure, or a harmless + // repeat, and treating it as an error would make every failure after + // item-add unrecoverable (see the header — rollback deletion is unsound). + let itemId = assertNoConflicts( + spawn, + issue, + project, + board, + status, + priority, + ); + if (itemId === undefined) { + const addedId = ghJson(spawn, [ + "project", + "item-add", + String(board), + "--owner", + OWNER, + "--url", + `https://github.com/${OWNER}/${REPO}/issues/${issue}`, + "--format", + "json", + ]).id; + if (!addedId) { + throw new Error(`item-add returned no id for #${issue}`); + } + // item-add is idempotent, so the empty pre-check does not prove this id + // names a NEW, unconfigured card — a concurrent invocation may have + // added AND configured it between the check and the add. Re-read it and + // re-apply the same conflict checks before the first field write; the + // added id is used only if the card is not yet visible to the lookup. + itemId = + assertNoConflicts(spawn, issue, project, board, status, priority) ?? + addedId; + } + + try { + editItemField( + spawn, + project, + itemId, + statusIds.fieldId, + statusIds.optionId, + ); + if (priorityIds) { + editItemField( + spawn, + project, + itemId, + priorityIds.fieldId, + priorityIds.optionId, + ); + } + + // Verify by reading each set field back — never report an unconfirmed add. + const after = findCard(spawn, issue, project); + const now = after?.fieldValueByName?.name ?? "(none)"; + if (now !== status) { + throw new Error(`card reads "${now}" after the add, not "${status}"`); + } + if (priority !== undefined) { + const priorityNow = + findCard(spawn, issue, project, "Priority")?.fieldValueByName?.name ?? + "(none)"; + if (priorityNow !== priority) { + throw new Error( + `card Priority reads "${priorityNow}" after the add, not "${priority}"`, + ); + } + } + console.log( + `card: ${now}${priority === undefined ? "" : ` / ${priority}`} (board #${board})`, + ); + } catch (cause) { + const message = cause instanceof Error ? cause.message : String(cause); + // NO rollback deletion: item-add is idempotent, so this invocation may + // hold the id of a card a concurrent operation created or has since + // configured — deleting it would destroy their work. The card is left + // in place and a re-run finishes it (or refuses, naming the conflict). + throw new Error( + `${message} — card ${itemId} on board #${board} may be partially configured; re-run this command to finish it, or remove it with board:delete`, + { cause }, + ); + } +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/board-card-add.test.mjs b/scripts/board-card-add.test.mjs new file mode 100644 index 0000000000..1d61fa1aeb --- /dev/null +++ b/scripts/board-card-add.test.mjs @@ -0,0 +1,282 @@ +// Tests for scripts/board-card-add.mjs (#2558) — argv validation, resolving +// every option BEFORE the first write, and the add-then-verify orchestration +// for Status and the optional Priority. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { main, parseAddArgs } from "./board-card-add.mjs"; + +test("parseAddArgs validates and defaults", () => { + assert.deepEqual( + parseAddArgs(["--issue", "7", "--status", "Todo", "--priority", "High"]), + { + issue: 7, + status: "Todo", + priority: "High", + board: 28, + }, + ); + // Board #11 has no Priority field, so only there may --priority be absent. + assert.equal( + parseAddArgs(["--issue", "7", "--status", "Todo", "--board", "11"]) + .priority, + undefined, + ); + assert.throws(() => parseAddArgs(["--issue", "7"]), /--status/); + // Only the two live boards may be mutated — a typo naming another real + // org project would otherwise add the card there. + assert.throws( + () => parseAddArgs(["--issue", "7", "--status", "Todo", "--board", "12"]), + /must be 28 \(v2\) or 11 \(v1\)/, + ); + // Every v2 board item has a Priority — an add without one is refused. + assert.throws( + () => parseAddArgs(["--issue", "7", "--status", "Todo"]), + /--priority is required on board #28/, + ); +}); + +const FIELDS = [ + { + id: "F_status", + name: "Status", + options: [{ id: "opt_todo", name: "Todo" }], + }, + { + id: "F_priority", + name: "Priority", + options: [{ id: "opt_med", name: "Medium" }], + }, +]; + +/** + * A spawn for main()'s flow: view, field-list, the pre-add issue lookup + * (empty until item-add unless `preexisting`), item-add, item-edits, then + * verify lookups. A graphql read of a field returns `before[field]` until an + * item-edit touches that field, and `after[field]` from then on — which is + * how a test shows a preexisting partial card being finished in place. + * `issueLookup:"pr"` makes every graphql lookup fail the way a PR number + * does. + */ +function spawnScript({ + fields = FIELDS, + after = { Status: "Todo", Priority: "Medium" }, + before, + preexisting = false, + issueLookup = "issue", + priorityEditFails = false, +} = {}) { + const calls = []; + const edited = new Set(); + let added = false; + const spawn = (cmd, args) => { + calls.push(args); + const joined = args.join(" "); + let payload; + if (joined.includes("project view")) { + payload = { id: "PVT_x" }; + } else if (joined.includes("field-list")) { + payload = { fields }; + } else if (joined.includes("item-add")) { + added = true; + payload = { id: "PVTI_new" }; + } else if (joined.includes("item-edit")) { + if (priorityEditFails && joined.includes("opt_med")) { + return { status: 1, stdout: "", stderr: "priority edit boom" }; + } + if (joined.includes("F_status")) edited.add("Status"); + if (joined.includes("F_priority")) edited.add("Priority"); + return { status: 0, stdout: "{}", stderr: "" }; + } else if (joined.includes("graphql")) { + if (issueLookup === "pr") { + return { + status: 1, + stdout: "", + stderr: "Could not resolve to an Issue with the number of 7.", + }; + } + const field = /fieldValueByName\(name:"(\w+)"\)/.exec(joined)[1]; + const values = before && !edited.has(field) ? before : after; + payload = { + data: { + repository: { + issue: { + projectItems: { + nodes: + added || preexisting + ? [ + { + id: "PVTI_new", + project: { id: "PVT_x" }, + fieldValueByName: values[field] + ? { name: values[field] } + : null, + }, + ] + : [], + }, + }, + }, + }, + }; + } else { + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("main adds, sets both fields, verifies, and prints the card line", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript(); + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn); + assert.deepEqual(lines, ["card: Todo / Medium (board #28)"]); + const edits = spawn.calls.filter((args) => args.includes("item-edit")); + assert.equal(edits.length, 2); + assert.ok(edits[0].includes("opt_todo")); + assert.ok(edits[1].includes("opt_med")); +}); + +test("main sets only Status on board #11, which has no Priority", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript({ after: { Status: "Todo" } }); + main(["--issue", "7", "--status", "Todo", "--board", "11"], spawn); + assert.deepEqual(lines, ["card: Todo (board #11)"]); + assert.equal( + spawn.calls.filter((args) => args.includes("item-edit")).length, + 1, + ); +}); + +test("a PR number fails the pre-add issue lookup, before item-add", () => { + // Issue and PR numbers share one namespace and item-add accepts a PR URL — + // the issue(number:) lookup must refuse it before the first write. + const spawn = spawnScript({ issueLookup: "pr" }); + assert.throws( + () => + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn), + /Could not resolve to an Issue/, + ); + assert.equal( + spawn.calls.some((args) => args.includes("item-add")), + false, + ); +}); + +test("a card holding a contradicting value is refused, before any write", () => { + // The real duplicate-add mistake: the issue is already moving through the + // board. Overwriting its Status would destroy board state someone set. + const spawn = spawnScript({ + preexisting: true, + before: { Status: "In Progress", Priority: "Medium" }, + }); + assert.throws( + () => + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn), + /already has a card on board #28 reading Status "In Progress"/, + ); + assert.equal( + spawn.calls.some( + (args) => args.includes("item-add") || args.includes("item-edit"), + ), + false, + ); +}); + +test("a matching preexisting card is reconfigured idempotently, no item-add", (t) => { + // item-add is idempotent upstream, so the command is too: a re-run after + // success (or a concurrent add that won the race) confirms and reports. + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript({ preexisting: true }); + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn); + assert.deepEqual(lines, ["card: Todo / Medium (board #28)"]); + assert.equal( + spawn.calls.some((args) => args.includes("item-add")), + false, + ); +}); + +test("a partially configured card is finished in place by a re-run", (t) => { + // The recovery path rollback used to foreclose: a failed Priority edit + // left Status set and Priority unset; the re-run completes it. + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript({ + preexisting: true, + before: { Status: "Todo" }, + }); + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn); + assert.deepEqual(lines, ["card: Todo / Medium (board #28)"]); + assert.equal( + spawn.calls.some((args) => args.includes("item-add")), + false, + ); + assert.equal( + spawn.calls.filter((args) => args.includes("item-edit")).length, + 2, + ); +}); + +test("a bad option name fails BEFORE item-add, leaving nothing half-made", () => { + const spawn = spawnScript({ fields: [FIELDS[0]] }); + assert.throws( + () => + main(["--issue", "7", "--status", "Todo", "--priority", "High"], spawn), + /no "Priority"/, + ); + assert.equal( + spawn.calls.some((args) => args.includes("item-add")), + false, + ); +}); + +test("main refuses to report an unconfirmed add", () => { + // before:{} keeps the post-add conflict recheck clean (unset fields), so + // the failure is the verify's own: the edit "succeeded" but did not take. + const spawn = spawnScript({ before: {}, after: { Status: "Incoming" } }); + assert.throws( + () => + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn), + /reads "Incoming"/, + ); +}); + +test("a card configured concurrently between pre-check and add is not overwritten", () => { + // item-add is idempotent: an empty pre-check does not prove the returned + // id names a new card. The post-add recheck must apply the same conflict + // rules before the first field write. + const spawn = spawnScript({ + before: { Status: "In Progress", Priority: "Medium" }, + }); + assert.throws( + () => + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn), + /already has a card on board #28 reading Status "In Progress"/, + ); + assert.ok(spawn.calls.some((args) => args.includes("item-add"))); + assert.equal( + spawn.calls.some((args) => args.includes("item-edit")), + false, + ); +}); + +test("a failure after item-add leaves the card and names the re-run", () => { + // NEVER a rollback delete: item-add is idempotent, so this invocation + // cannot prove it created the card — deleting could destroy a concurrent + // operation's card. The error says how to finish instead. + const spawn = spawnScript({ priorityEditFails: true }); + assert.throws( + () => + main(["--issue", "7", "--status", "Todo", "--priority", "Medium"], spawn), + /priority edit boom.*PVTI_new.*may be partially configured.*re-run/s, + ); + assert.equal( + spawn.calls.some((args) => args.includes("item-delete")), + false, + ); +}); diff --git a/scripts/board-card-delete.mjs b/scripts/board-card-delete.mjs new file mode 100644 index 0000000000..69b7eb90ad --- /dev/null +++ b/scripts/board-card-delete.mjs @@ -0,0 +1,137 @@ +#!/usr/bin/env node +// Delete an issue's board card (#2558) — `npm run board:delete -- --issue <N> +// [--board 28] [--reason duplicate|not-planned]`. The delete recipe board-ops +// previously transcribed inline, including the issue-side ITEM_ID lookup it +// depended on. +// +// "Done means the work shipped": an issue closed as duplicate / won't fix / +// not planned / obsolete shipped nothing, so its card is DELETED, not parked +// in Done. Deleting the card touches the board only — the issue keeps its +// labels and comments and stays searchable forever. +// +// `--reason` also closes the issue with the matching machine-readable state +// reason. `duplicate` cannot be set through `gh issue close --reason` (it +// accepts only completed / not planned), so the close goes through the API — +// the same PATCH the skill documented. +// +// An absent card is ALWAYS an error unless `--allow-missing-card` says +// otherwise. The flag exists for one case: a prior run deleted the card and +// then failed the close transiently, so the retry finds no card and still +// needs to PATCH. Implicit retry-on-absence was rejected in review — a wrong +// `--board`, or a mistyped issue number, would close an issue while its real +// card survives. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { OWNER, REPO_SLUG, gh, requirePositiveInt } from "./lib/gh.mjs"; +import { + DEFAULT_BOARD, + findCard, + projectId as resolveProjectId, + requireSupportedBoard, +} from "./lib/board.mjs"; + +const CLOSE_REASONS = { duplicate: "duplicate", "not-planned": "not_planned" }; + +export function parseDeleteArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { + issue: { type: "string" }, + board: { type: "string" }, + reason: { type: "string" }, + "allow-missing-card": { type: "boolean" }, + }, + }); + if ( + values.reason !== undefined && + !Object.hasOwn(CLOSE_REASONS, values.reason) + ) { + throw new Error( + `--reason must be one of: ${Object.keys(CLOSE_REASONS).join(", ")}`, + ); + } + if (values["allow-missing-card"] === true && values.reason === undefined) { + throw new Error( + "--allow-missing-card is only meaningful with --reason — without a close there is nothing left to do when the card is absent", + ); + } + return { + issue: requirePositiveInt(values.issue, "--issue"), + board: requireSupportedBoard( + values.board === undefined + ? DEFAULT_BOARD + : requirePositiveInt(values.board, "--board"), + ), + reason: values.reason, + allowMissingCard: values["allow-missing-card"] === true, + }; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { issue, board, reason, allowMissingCard } = parseDeleteArgs(argv); + + const project = resolveProjectId(spawn, board); + const card = findCard(spawn, issue, project); + if (!card?.id) { + // Absent card: fail unless the caller explicitly declared this a retry + // of a delete-succeeded/close-failed run. Closing implicitly on --reason + // alone would close the issue on a wrong --board or a typo'd number + // while its real card survives. + if (!allowMissingCard) { + throw new Error( + `#${issue} has no card on board #${board} — nothing deleted` + + (reason + ? "; if a prior run already deleted it, re-run with --allow-missing-card to close anyway" + : ""), + ); + } + console.log( + `no card: #${issue} on board #${board} — closing anyway (--allow-missing-card)`, + ); + } else { + const del = gh(spawn, [ + "project", + "item-delete", + String(board), + "--owner", + OWNER, + "--id", + card.id, + "--format", + "json", + ]); + if (del.status !== 0) { + throw new Error(`item-delete failed: ${(del.stderr ?? "").trim()}`); + } + + // Verify by looking the card up again — never report an unconfirmed delete. + if (findCard(spawn, issue, project)?.id) { + throw new Error( + `#${issue} still has a card on board #${board} after delete`, + ); + } + console.log(`deleted: card for #${issue} on board #${board}`); + } + + if (reason) { + const close = gh(spawn, [ + "api", + `repos/${REPO_SLUG}/issues/${issue}`, + "-X", + "PATCH", + "-f", + "state=closed", + "-f", + `state_reason=${CLOSE_REASONS[reason]}`, + ]); + if (close.status !== 0) { + throw new Error(`close failed: ${(close.stderr ?? "").trim()}`); + } + console.log(`closed: #${issue} (${CLOSE_REASONS[reason]})`); + } +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/board-card-delete.test.mjs b/scripts/board-card-delete.test.mjs new file mode 100644 index 0000000000..f37315c03e --- /dev/null +++ b/scripts/board-card-delete.test.mjs @@ -0,0 +1,146 @@ +// Tests for scripts/board-card-delete.mjs (#2558) — argv validation, the +// delete-then-verify orchestration, and the optional API close with the +// machine-readable reason `gh issue close` cannot set. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { main, parseDeleteArgs } from "./board-card-delete.mjs"; + +test("parseDeleteArgs validates the reason vocabulary", () => { + assert.deepEqual(parseDeleteArgs(["--issue", "7"]), { + issue: 7, + board: 28, + reason: undefined, + allowMissingCard: false, + }); + assert.equal( + parseDeleteArgs(["--issue", "7", "--reason", "duplicate"]).reason, + "duplicate", + ); + assert.throws( + () => parseDeleteArgs(["--issue", "7", "--reason", "wontfix"]), + /duplicate, not-planned/, + ); + // Own-property check: an inherited name must not pass as a close reason. + assert.throws( + () => parseDeleteArgs(["--issue", "7", "--reason", "toString"]), + /duplicate, not-planned/, + ); + // The retry flag is only meaningful with a close to retry. + assert.throws( + () => parseDeleteArgs(["--issue", "7", "--allow-missing-card"]), + /only meaningful with --reason/, + ); +}); + +/** A spawn whose card lookup returns a card until the delete, then none. */ +function spawnScript({ card = true, stillThere = false } = {}) { + const calls = []; + let deleted = false; + const spawn = (cmd, args) => { + calls.push(args); + const joined = args.join(" "); + let payload; + if (joined.includes("project view")) { + payload = { id: "PVT_x" }; + } else if (joined.includes("graphql")) { + const present = card && (!deleted || stillThere); + payload = { + data: { + repository: { + issue: { + projectItems: { + nodes: present + ? [{ id: "PVTI_x", project: { id: "PVT_x" } }] + : [], + }, + }, + }, + }, + }; + } else if (joined.includes("item-delete")) { + deleted = true; + return { status: 0, stdout: "{}", stderr: "" }; + } else if (joined.includes("PATCH")) { + return { status: 0, stdout: "{}", stderr: "" }; + } else { + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("main deletes, verifies the card is gone, and prints", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript(); + main(["--issue", "7"], spawn); + assert.deepEqual(lines, ["deleted: card for #7 on board #28"]); + assert.equal( + spawn.calls.some((args) => args.includes("PATCH")), + false, + ); +}); + +test("main closes with the API reason when --reason is given", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript(); + main(["--issue", "7", "--reason", "duplicate"], spawn); + assert.deepEqual(lines, [ + "deleted: card for #7 on board #28", + "closed: #7 (duplicate)", + ]); + const patch = spawn.calls.find((args) => args.includes("PATCH")); + assert.ok(patch.includes("state_reason=duplicate")); +}); + +test("main throws when there is no card to delete", () => { + assert.throws( + () => main(["--issue", "7"], spawnScript({ card: false })), + /no card on board #28/, + ); +}); + +test("--reason alone still refuses an absent card", () => { + // A wrong --board or a typo'd issue number must not close the issue while + // its real card survives — closing past an absent card is opt-in. + assert.throws( + () => + main( + ["--issue", "7", "--reason", "not-planned"], + spawnScript({ card: false }), + ), + /no card on board #28.*--allow-missing-card/s, + ); +}); + +test("an explicit --allow-missing-card retry closes past the absent card", (t) => { + // A prior run may have deleted the card and failed the PATCH transiently — + // the declared retry must not stop at "no card" with the issue still open. + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript({ card: false }); + main( + ["--issue", "7", "--reason", "not-planned", "--allow-missing-card"], + spawn, + ); + assert.deepEqual(lines, [ + "no card: #7 on board #28 — closing anyway (--allow-missing-card)", + "closed: #7 (not_planned)", + ]); + assert.equal( + spawn.calls.some((args) => args.includes("item-delete")), + false, + ); +}); + +test("main refuses to report an unconfirmed delete", () => { + assert.throws( + () => main(["--issue", "7"], spawnScript({ stillThere: true })), + /still has a card/, + ); +}); diff --git a/scripts/board-card-status.mjs b/scripts/board-card-status.mjs new file mode 100644 index 0000000000..160aa2db64 --- /dev/null +++ b/scripts/board-card-status.mjs @@ -0,0 +1,91 @@ +#!/usr/bin/env node +// Move an issue's board card to a Status (#2558) — `npm run board:status -- +// --issue <N> --status "In Review" [--board 28]`. The card-move block the +// pr-flow skill (steps 1 and 6) and board-ops previously transcribed inline — +// the most fragile of the four, because a half-pasted block silently no-ops. +// +// Three properties carried over from the inline block, all load-bearing: +// +// - Every id is resolved BY NAME at run time. Single-select option ids are +// regenerated whenever a field's option list is edited (board-ops' hazard), +// so no option id is hardcoded here and a recreated option still resolves. +// - The card is found FROM THE ISSUE (`projectItems`), selected by the +// board's node id — never from `gh project item-list`, whose `--limit` +// truncates silently past the board's size. +// - The move is VERIFIED by reading the Status back; `card: <Status>` on +// stdout is printed only on a confirmed match, so "the step is done only +// when it prints `card: …`" is enforced by the exit code too. +// +// `--board 11` works unchanged for a v1 issue — same column names, every id +// resolved by name against that board. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { requirePositiveInt } from "./lib/gh.mjs"; +import { + DEFAULT_BOARD, + boardFields, + cardOnProject, + editItemField, + fieldOption, + findCard, + projectId as resolveProjectId, + requireSupportedBoard, +} from "./lib/board.mjs"; + +export { DEFAULT_BOARD, cardOnProject, fieldOption }; + +export function parseStatusArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { + issue: { type: "string" }, + status: { type: "string" }, + board: { type: "string" }, + }, + }); + if (!values.status) { + throw new Error("--status is required (e.g. --status 'In Review')"); + } + return { + issue: requirePositiveInt(values.issue, "--issue"), + status: values.status, + board: requireSupportedBoard( + values.board === undefined + ? DEFAULT_BOARD + : requirePositiveInt(values.board, "--board"), + ), + }; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { issue, status, board } = parseStatusArgs(argv); + + const project = resolveProjectId(spawn, board); + const { fieldId, optionId } = fieldOption( + boardFields(spawn, board), + "Status", + status, + ); + + const card = findCard(spawn, issue, project); + if (!card?.id) { + throw new Error( + `#${issue} has no card on board #${board} — board it first (/issue-create step 4)`, + ); + } + + editItemField(spawn, project, card.id, fieldId, optionId); + + // Verify by reading the Status back — never report an unconfirmed move. + const after = findCard(spawn, issue, project); + const now = after?.fieldValueByName?.name ?? "(none)"; + if (now !== status) { + throw new Error(`card reads "${now}" after the edit, not "${status}"`); + } + console.log(`card: ${now}`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/board-card-status.test.mjs b/scripts/board-card-status.test.mjs new file mode 100644 index 0000000000..5214a61e69 --- /dev/null +++ b/scripts/board-card-status.test.mjs @@ -0,0 +1,148 @@ +// Tests for scripts/board-card-status.mjs (#2558) — name-based id resolution, +// issue-side card lookup, and `main()`'s edit-then-verify orchestration +// through an injected spawn. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + cardOnProject, + fieldOption, + main, + parseStatusArgs, +} from "./board-card-status.mjs"; + +const FIELDS = [ + { + id: "F_status", + name: "Status", + options: [ + { id: "opt_todo", name: "Todo" }, + { id: "opt_rev", name: "In Review" }, + ], + }, +]; + +test("fieldOption resolves field and option ids by name", () => { + assert.deepEqual(fieldOption(FIELDS, "Status", "In Review"), { + fieldId: "F_status", + optionId: "opt_rev", + }); + assert.throws(() => fieldOption(FIELDS, "Priority", "High"), /no "Priority"/); + // The error names the options that DO exist, so a renamed column is obvious. + assert.throws( + () => fieldOption(FIELDS, "Status", "Review"), + /Todo, In Review/, + ); +}); + +const cardResponse = (statusName, projectId = "PVT_x") => ({ + data: { + repository: { + issue: { + projectItems: { + nodes: [ + { id: "PVTI_other", project: { id: "PVT_other" } }, + { + id: "PVTI_ours", + project: { id: projectId }, + fieldValueByName: statusName ? { name: statusName } : null, + }, + ], + }, + }, + }, + }, +}); + +test("cardOnProject selects by project node id and throws on a bad shape", () => { + assert.equal(cardOnProject(cardResponse("Todo"), "PVT_x").id, "PVTI_ours"); + assert.equal(cardOnProject(cardResponse("Todo"), "PVT_absent"), undefined); + assert.throws(() => cardOnProject({ data: {} }, "PVT_x"), /unexpected/); +}); + +test("parseStatusArgs validates and defaults the board to 28", () => { + assert.deepEqual(parseStatusArgs(["--issue", "7", "--status", "Todo"]), { + issue: 7, + status: "Todo", + board: 28, + }); + assert.equal( + parseStatusArgs(["--issue", "7", "--status", "Todo", "--board", "11"]) + .board, + 11, + ); + assert.throws(() => parseStatusArgs(["--issue", "7"]), /--status/); +}); + +/** + * A spawn answering main()'s five calls in order: project view, field-list, + * card lookup, item-edit, verify lookup. `after` is the Status the verify + * read returns. + */ +function spawnScript({ + after = "In Review", + card = true, + editStatus = 0, +} = {}) { + const calls = []; + let lookups = 0; + const spawn = (cmd, args) => { + calls.push(args); + const joined = args.join(" "); + let payload; + if (joined.includes("project view")) { + payload = { id: "PVT_x" }; + } else if (joined.includes("field-list")) { + payload = { fields: FIELDS }; + } else if (joined.includes("graphql")) { + lookups += 1; + payload = card + ? cardResponse(lookups === 1 ? "Todo" : after) + : { data: { repository: { issue: { projectItems: { nodes: [] } } } } }; + } else if (joined.includes("item-edit")) { + return { status: editStatus, stdout: "{}", stderr: "edit refused" }; + } else { + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("main edits the card and prints card: <Status> only after verifying", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript(); + main(["--issue", "7", "--status", "In Review"], spawn); + assert.deepEqual(lines, ["card: In Review"]); + + const edit = spawn.calls.find((args) => args.includes("item-edit")); + assert.ok(edit.includes("PVTI_ours")); + assert.ok(edit.includes("F_status")); + assert.ok(edit.includes("opt_rev")); +}); + +test("main throws when the issue has no card on the board", () => { + const spawn = spawnScript({ card: false }); + assert.throws( + () => main(["--issue", "7", "--status", "In Review"], spawn), + /no card on board #28/, + ); +}); + +test("main throws when item-edit fails", () => { + const spawn = spawnScript({ editStatus: 1 }); + assert.throws( + () => main(["--issue", "7", "--status", "In Review"], spawn), + /edit refused/, + ); +}); + +test("main refuses to report an unconfirmed move", () => { + const spawn = spawnScript({ after: "Todo" }); + assert.throws( + () => main(["--issue", "7", "--status", "In Review"], spawn), + /reads "Todo"/, + ); +}); diff --git a/scripts/board-draft-find.mjs b/scripts/board-draft-find.mjs new file mode 100644 index 0000000000..cc99e16657 --- /dev/null +++ b/scripts/board-draft-find.mjs @@ -0,0 +1,69 @@ +#!/usr/bin/env node +// Find a GHSA advisory draft card by title (#2558) — `npm run +// board:find-draft -- --ghsa GHSA-xxxx-yyyy-zzzz [--board 28]`. The +// title-lookup block board-ops previously transcribed inline. +// +// A draft card has no repository and no issue number, so the issue-side +// lookup the other board scripts use cannot find one — it is matched by the +// bracketed GHSA id prefix in its title, never by words from the summary (a +// summary is free text and two advisories can share one). The listing is +// trusted only when complete (`itemListComplete`), so a truncated dump reads +// as an error rather than as "no draft card". + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { + DEFAULT_BOARD, + itemListComplete, + requireSupportedBoard, +} from "./lib/board.mjs"; +import { requirePositiveInt } from "./lib/gh.mjs"; + +const GHSA_PATTERN = + /^GHSA-[23456789cfghjmpqrvwx]{4}-[23456789cfghjmpqrvwx]{4}-[23456789cfghjmpqrvwx]{4}$/; + +export function parseFindDraftArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { ghsa: { type: "string" }, board: { type: "string" } }, + }); + if (!values.ghsa || !GHSA_PATTERN.test(values.ghsa)) { + throw new Error( + `--ghsa must be a full GHSA id (GHSA-xxxx-yyyy-zzzz), got ${values.ghsa ?? "nothing"}`, + ); + } + return { + ghsa: values.ghsa, + board: requireSupportedBoard( + values.board === undefined + ? DEFAULT_BOARD + : requirePositiveInt(values.board, "--board"), + ), + }; +} + +/** Draft items whose title carries the bracketed GHSA id prefix. */ +export function draftsFor(items, ghsa) { + return items.filter( + (item) => + item?.content?.type === "DraftIssue" && + (item.content.title ?? "").startsWith(`[${ghsa}]`), + ); +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { ghsa, board } = parseFindDraftArgs(argv); + + const { items } = itemListComplete(spawn, board); + const drafts = draftsFor(items, ghsa); + if (drafts.length === 0) { + throw new Error(`no draft card titled [${ghsa}] on #${board}`); + } + for (const draft of drafts) { + console.log(`ITEM=${draft.id} ${draft.content.title}`); + } +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/board-draft-find.test.mjs b/scripts/board-draft-find.test.mjs new file mode 100644 index 0000000000..545ff329be --- /dev/null +++ b/scripts/board-draft-find.test.mjs @@ -0,0 +1,70 @@ +// Tests for scripts/board-draft-find.mjs (#2558) — GHSA id validation, the +// bracketed-prefix title match, and the complete-listing dependency. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { draftsFor, main, parseFindDraftArgs } from "./board-draft-find.mjs"; + +const GHSA = "GHSA-2345-cfgh-jmpq"; + +test("parseFindDraftArgs requires a full GHSA id", () => { + assert.deepEqual(parseFindDraftArgs(["--ghsa", GHSA]), { + ghsa: GHSA, + board: 28, + }); + assert.throws( + () => parseFindDraftArgs(["--ghsa", "GHSA-123"]), + /full GHSA id/, + ); + assert.throws(() => parseFindDraftArgs([]), /full GHSA id/); +}); + +const ITEMS = [ + { id: "PVTI_issue", content: { type: "Issue", number: 7 } }, + { + id: "PVTI_draft", + content: { type: "DraftIssue", title: `[${GHSA}] - proxy SSRF` }, + }, + { + id: "PVTI_other", + content: { type: "DraftIssue", title: "[GHSA-aaaa-bbbb-cccc] - other" }, + }, +]; + +test("draftsFor matches the bracketed id prefix on drafts only", () => { + assert.deepEqual( + draftsFor(ITEMS, GHSA).map((item) => item.id), + ["PVTI_draft"], + ); + assert.deepEqual(draftsFor(ITEMS, "GHSA-9999-9999-9999"), []); +}); + +const spawnScript = (items) => () => ({ + status: 0, + stdout: JSON.stringify({ items, totalCount: items.length }), + stderr: "", +}); + +test("main prints ITEM=<id> <title> for each match", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + main(["--ghsa", GHSA], spawnScript(ITEMS)); + assert.deepEqual(lines, [`ITEM=PVTI_draft [${GHSA}] - proxy SSRF`]); +}); + +test("main throws when no draft card matches", () => { + assert.throws( + () => main(["--ghsa", GHSA], spawnScript([ITEMS[0]])), + new RegExp(`no draft card titled \\[${GHSA}\\]`), + ); +}); + +test("main refuses a truncated listing rather than reporting absence", () => { + const spawn = () => ({ + status: 0, + stdout: JSON.stringify({ items: ITEMS, totalCount: 500 }), + stderr: "", + }); + assert.throws(() => main(["--ghsa", GHSA], spawn), /INCOMPLETE/); +}); diff --git a/scripts/board-recover.mjs b/scripts/board-recover.mjs new file mode 100644 index 0000000000..6a7ac46349 --- /dev/null +++ b/scripts/board-recover.mjs @@ -0,0 +1,346 @@ +#!/usr/bin/env node +// Recover from a deleted single-select board option (#2558) — the two +// mechanical phases of board-ops' recovery recipe. Step 2 of that recipe — +// recreating the option while echoing every surviving option's id — is +// deliberately NOT scripted: it edits the field schema, the very operation +// the option-deletion hazard is about, and stays a human act in the web UI. +// +// npm run board:recover -- --phase diff --snapshot <path> [--field Status] +// npm run board:recover -- --phase reapply --lost <path> --option-id <id> +// +// `diff` dumps the broken board (complete or refused) and compares it against +// the snapshot: a card is LOST only when it is null now AND held a value in +// the snapshot — a card already blank in the snapshot, or added since, is not +// recovery's to touch. The lost set is written to lost-ids.json BESIDE the +// snapshot only when every lost card held the SAME snapshot value, since +// reapply assigns one option id to all of them; a mixed grouping is printed +// and refused. The file records the board, field and held value alongside the +// ids, so `reapply` can verify that --option-id is actually the recreated +// option for that value — any other valid option id on the field is refused +// rather than silently rewriting every lost card to the wrong value. Reapply +// also re-reads each card's field immediately before editing it (a whole-board +// preflight cannot hold across a paced loop) and aborts on any card that is +// gone or no longer blank, paced to stay under the API's abuse limits. + +import { spawnSync } from "node:child_process"; +import { readFileSync, rmSync, writeFileSync } from "node:fs"; +import { dirname, join } from "node:path"; +import { parseArgs } from "node:util"; +import { setTimeout as delay } from "node:timers/promises"; +import { requirePositiveInt } from "./lib/gh.mjs"; +import { + assertOutsideRepo, + boardFields, + editItemField, + itemFieldValue, + itemListComplete, + projectId as resolveProjectId, + requireSupportedBoard, +} from "./lib/board.mjs"; + +const EDIT_PACING_MS = 400; + +export function parseRecoverArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { + phase: { type: "string" }, + snapshot: { type: "string" }, + lost: { type: "string" }, + "option-id": { type: "string" }, + field: { type: "string" }, + board: { type: "string" }, + }, + }); + const board = + values.board === undefined + ? undefined + : requireSupportedBoard(requirePositiveInt(values.board, "--board")); + if (values.phase === "diff") { + if (!values.snapshot) { + throw new Error("--phase diff needs --snapshot <path>"); + } + // board stays undefined when not given: diff takes it from the snapshot + // (board:snapshot records it), and an explicit flag may only CONFIRM what + // the snapshot records — a mismatch is refused in main. + return { + phase: "diff", + snapshot: values.snapshot, + field: values.field ?? "Status", + board, + }; + } + if (values.phase === "reapply") { + if (!values.lost || !values["option-id"]) { + throw new Error( + "--phase reapply needs --lost <path> and --option-id <id>", + ); + } + // board and field stay undefined when not given: reapply takes both from + // the lost file (written by diff), and an explicit flag may only CONFIRM + // what the file records — a mismatch is refused in main. + return { + phase: "reapply", + lost: values.lost, + optionId: values["option-id"], + field: values.field, + board, + }; + } + throw new Error('--phase must be "diff" or "reapply"'); +} + +/** item-list exposes each single-select field under its lowercased name. */ +const fieldKey = (field) => field.toLowerCase(); + +/** + * Cards null in the broken dump that held a value in the snapshot, grouped by + * that value: [{ value, count, ids }]. A card null in both dumps was blank + * before the deletion, and a card absent from the snapshot was added after it + * — neither is recovery's to overwrite, so both are excluded. + */ +export function lostGrouping(snapshotItems, brokenItems, field) { + const key = fieldKey(field); + const held = new Map(); + for (const item of snapshotItems) { + if (item[key] != null) { + held.set(item.id, item[key]); + } + } + const groups = new Map(); + for (const item of brokenItems) { + if (item[key] != null || !held.has(item.id)) { + continue; + } + const value = held.get(item.id); + const group = groups.get(value) ?? { value, count: 0, ids: [] }; + group.count += 1; + group.ids.push(item.id); + groups.set(value, group); + } + return [...groups.values()]; +} + +export async function main( + argv = process.argv.slice(2), + spawn = spawnSync, + sleep = delay, +) { + const parsed = parseRecoverArgs(argv); + + if (parsed.phase === "diff") { + // lost-ids.json is derived BESIDE the snapshot, so a --snapshot inside + // the worktree would write private board item ids one `git add -A` from + // a PR — the same refusal board:snapshot applies to its --dir. + const lostDir = dirname(parsed.snapshot); + assertOutsideRepo(lostDir, process.cwd()); + // A previous diff's lost-ids.json must not survive this run: a diff that + // fails or ends with nothing recoverable would otherwise leave stale ids + // for reapply to consume. Remove it before anything can fail. + const lostPath = join(lostDir, "lost-ids.json"); + rmSync(lostPath, { force: true }); + const snapshot = JSON.parse(readFileSync(parsed.snapshot, "utf8")); + // A truncated snapshot cannot be diffed against: the cards it omits + // would be excluded from the lost set and reapply would then report + // success while leaving them orphaned. The snapshot carries its own + // totalCount (board:snapshot writes the verified dump whole), so refuse + // one whose items fall short of it — or one missing either key. + if ( + !Array.isArray(snapshot.items) || + typeof snapshot.totalCount !== "number" || + snapshot.items.length !== snapshot.totalCount + ) { + throw new Error( + `${parsed.snapshot} is not a complete board snapshot ` + + `(${snapshot.items?.length ?? "?"} items of totalCount ` + + `${snapshot.totalCount ?? "?"}) — retake it with board:snapshot`, + ); + } + // The snapshot proves which board it belongs to (board:snapshot records + // it): diffing a board-11 snapshot against board 28's dump would match + // nothing and confidently print "lost: 0 cards". An explicit --board may + // only confirm it; a snapshot predating the recorded key (hand-taken + // with gh directly) cannot prove its board, so there the flag is + // REQUIRED rather than defaulted — the silent wrong-board diff is the + // exact failure this check exists for. + let board; + if (snapshot.board !== undefined) { + if (!Number.isInteger(snapshot.board)) { + throw new Error( + `${parsed.snapshot} records a non-numeric board (${snapshot.board})`, + ); + } + requireSupportedBoard(snapshot.board, `${parsed.snapshot}'s board`); + if (parsed.board !== undefined && parsed.board !== snapshot.board) { + throw new Error( + `--board ${parsed.board} does not match the snapshot's board #${snapshot.board}`, + ); + } + board = snapshot.board; + } else { + if (parsed.board === undefined) { + throw new Error( + `${parsed.snapshot} records no board — pass --board naming the board it was taken from`, + ); + } + board = parsed.board; + } + // The field must exist on the board — a typo (--field Priorty) would + // otherwise read every card's value as undefined, match nothing, and + // falsely report no lost cards. + if ( + !boardFields(spawn, board).some( + (candidate) => candidate.name === parsed.field, + ) + ) { + throw new Error(`board #${board} has no "${parsed.field}" field`); + } + // itemListComplete refuses a truncated dump, so lost-ids.json is written + // only from a complete picture of the broken board. + const broken = itemListComplete(spawn, board); + const groups = lostGrouping(snapshot.items, broken.items, parsed.field); + for (const { value, count } of groups) { + console.log(`was ${value}: ${count}`); + } + if (groups.length === 0) { + console.log("lost: 0 cards — nothing to recover"); + return; + } + // Reapply assigns ONE option id to every lost card, so the lost set is + // only actionable when it held a single value. A mixed grouping means + // something besides the option deletion blanked cards — refuse it. + if (groups.length > 1) { + process.exitCode = 1; + console.error( + `lost cards held ${groups.length} different values — one option id ` + + `cannot restore them all; not writing ${lostPath}`, + ); + return; + } + const lostIds = groups[0].ids; + // The artifact records what the lost cards HELD, not just their ids, so + // reapply can refuse an --option-id that is valid on the field but is not + // the recreated option for this value. + const payload = { + board, + field: parsed.field, + value: groups[0].value, + ids: lostIds, + }; + // Same protections as the snapshot itself: the ids are private board + // data, so owner-only, and exclusive so a file planted between the + // removal above and this write is refused rather than followed. + writeFileSync(lostPath, JSON.stringify(payload, null, 2), { + mode: 0o600, + flag: "wx", + }); + console.log( + `lost: ${lostIds.length} cards (all "${groups[0].value}") → ${lostPath}`, + ); + return; + } + + const recorded = JSON.parse(readFileSync(parsed.lost, "utf8")); + if ( + recorded === null || + typeof recorded !== "object" || + Array.isArray(recorded) || + !Number.isInteger(recorded.board) || + typeof recorded.field !== "string" || + typeof recorded.value !== "string" || + !Array.isArray(recorded.ids) || + recorded.ids.length === 0 || + recorded.ids.some((id) => typeof id !== "string") + ) { + throw new Error( + `${parsed.lost} is not a lost file written by --phase diff ` + + `({ board, field, value, ids })`, + ); + } + // Same boundary the diff phase applies to a snapshot's recorded board: a + // hand-edited lost file must not point this mutation loop at an arbitrary + // org project. + requireSupportedBoard(recorded.board, `${parsed.lost}'s board`); + // The file is authoritative for board and field; an explicit flag may only + // confirm it. A silent override would let reapply run against a different + // board or field than the one diff actually measured. + if (parsed.board !== undefined && parsed.board !== recorded.board) { + throw new Error( + `--board ${parsed.board} does not match the lost file's board #${recorded.board}`, + ); + } + if (parsed.field !== undefined && parsed.field !== recorded.field) { + throw new Error( + `--field ${parsed.field} does not match the lost file's field "${recorded.field}"`, + ); + } + const lostIds = recorded.ids; + const project = resolveProjectId(spawn, recorded.board); + const field = boardFields(spawn, recorded.board).find( + (candidate) => candidate.name === recorded.field, + ); + if (!field) { + throw new Error( + `board #${recorded.board} has no "${recorded.field}" field`, + ); + } + // --option-id must be the RECREATED option for the value the lost cards + // held — any other valid option id on the field would succeed and silently + // rewrite every lost card to the wrong value. + const option = (field.options ?? []).find( + (candidate) => candidate.id === parsed.optionId, + ); + if (!option) { + const known = (field.options ?? []) + .map((o) => `"${o.name}" (${o.id})`) + .join(", "); + throw new Error( + `"${recorded.field}" has no option with id ${parsed.optionId} (has: ${known})`, + ); + } + if (option.name !== recorded.value) { + throw new Error( + `--option-id ${parsed.optionId} is "${option.name}" but the lost cards ` + + `held "${recorded.value}" — pass the recreated "${recorded.value}" option's id`, + ); + } + // Fail fast, before the FIRST edit, when the list is already stale — someone + // may have legitimately set one of these cards while the option was being + // recreated, or deleted one. This makes a stale-at-start run all-or-nothing; + // the per-card read in the loop below is what holds at mutation time. + const current = itemListComplete(spawn, recorded.board); + const byId = new Map(current.items.map((item) => [item.id, item])); + const key = fieldKey(recorded.field); + const stale = lostIds.filter( + (id) => !byId.has(id) || byId.get(id)[key] != null, + ); + if (stale.length > 0) { + throw new Error( + `${stale.length} of ${lostIds.length} lost cards are gone or no longer ` + + `blank (${stale.slice(0, 5).join(", ")}${stale.length > 5 ? ", …" : ""}) ` + + `— the lost list is stale; re-run --phase diff and retry`, + ); + } + let applied = 0; + for (const id of lostIds) { + // Re-read THIS card immediately before its edit: with hundreds of cards + // and 400 ms pacing, the preflight above goes stale mid-loop, and a card + // someone set during the run must not be overwritten. + const now = itemFieldValue(spawn, id, recorded.field); + if (!now.exists || now.value !== null) { + throw new Error( + `card ${id} is ${now.exists ? `no longer blank ("${now.value}")` : "gone"} — ` + + `the lost list went stale mid-run (${applied} of ${lostIds.length} ` + + `reapplied); re-run --phase diff and retry with the remainder`, + ); + } + editItemField(spawn, project, id, field.id, parsed.optionId); + applied += 1; + await sleep(EDIT_PACING_MS); + } + console.log(`reapplied: ${lostIds.length} cards → option ${parsed.optionId}`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + await main(); +} diff --git a/scripts/board-recover.test.mjs b/scripts/board-recover.test.mjs new file mode 100644 index 0000000000..006b5711ca --- /dev/null +++ b/scripts/board-recover.test.mjs @@ -0,0 +1,591 @@ +// Tests for scripts/board-recover.mjs (#2558) — the diff phase's +// complete-dump refusal and snapshot grouping (the safety check that the +// orphaned set is exactly the cards that held the deleted option), and the +// reapply phase's paced re-application. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + mkdtempSync, + existsSync, + readFileSync, + statSync, + writeFileSync, +} from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { lostGrouping, main, parseRecoverArgs } from "./board-recover.mjs"; + +test("parseRecoverArgs demands each phase's own inputs", () => { + assert.equal( + parseRecoverArgs(["--phase", "diff", "--snapshot", "/tmp/s.json"]).phase, + "diff", + ); + assert.throws(() => parseRecoverArgs(["--phase", "diff"]), /--snapshot/); + assert.throws( + () => parseRecoverArgs(["--phase", "reapply", "--lost", "x"]), + /--option-id/, + ); + assert.throws( + () => parseRecoverArgs(["--phase", "undo"]), + /"diff" or "reapply"/, + ); +}); + +test("lostGrouping keeps only cards that lost a snapshot value", () => { + const snapshot = [ + { id: "a", status: "Done" }, + { id: "b", status: "Done" }, + { id: "c", status: null }, // blank before the deletion — not lost + { id: "d", status: "Todo" }, // untouched + ]; + const broken = [ + { id: "a", status: null }, + { id: "b", status: null }, + { id: "c", status: null }, + { id: "d", status: "Todo" }, + { id: "e", status: null }, // added after the snapshot — not lost + ]; + assert.deepEqual(lostGrouping(snapshot, broken, "Status"), [ + { value: "Done", count: 2, ids: ["a", "b"] }, + ]); +}); + +const dumpSpawn = + (items, totalCount = items.length) => + (cmd, args) => { + if (args.join(" ").includes("field-list")) { + return { + status: 0, + stdout: JSON.stringify({ + fields: [{ id: "F_status", name: "Status" }], + }), + stderr: "", + }; + } + return { + status: 0, + stdout: JSON.stringify({ items, totalCount }), + stderr: "", + }; + }; + +/** A snapshot file as board:snapshot writes it — board recorded, complete. */ +const snapshotFile = (items, board = 28) => + JSON.stringify({ board, items, totalCount: items.length }); + +test("diff writes lost-ids.json beside the snapshot and prints the grouping", async (t) => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "board-28-snapshot.json"); + writeFileSync( + snapshotPath, + snapshotFile([ + { id: "a", status: "Done" }, + { id: "b", status: "Todo" }, + ]), + ); + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + await main( + ["--phase", "diff", "--snapshot", snapshotPath], + dumpSpawn([ + { id: "a", status: null }, // orphaned + { id: "b", status: "Todo" }, + ]), + ); + const lostPath = join(dir, "lost-ids.json"); + assert.deepEqual(JSON.parse(readFileSync(lostPath, "utf8")), { + board: 28, + field: "Status", + value: "Done", + ids: ["a"], + }); + // Same protections as the snapshot: private ids, owner-only, exclusive. + assert.equal(statSync(lostPath).mode & 0o777, 0o600); + assert.deepEqual(lines, [ + "was Done: 1", + `lost: 1 cards (all "Done") → ${lostPath}`, + ]); +}); + +test("diff refuses to write lost-ids.json when lost cards held mixed values", async (t) => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "board-28-snapshot.json"); + writeFileSync( + snapshotPath, + snapshotFile([ + { id: "a", status: "Done" }, + { id: "b", status: "Todo" }, + ]), + ); + const lines = []; + const errors = []; + t.mock.method(console, "log", (line) => lines.push(line)); + t.mock.method(console, "error", (line) => errors.push(line)); + const previousExitCode = process.exitCode; + try { + await main( + ["--phase", "diff", "--snapshot", snapshotPath], + dumpSpawn([ + { id: "a", status: null }, + { id: "b", status: null }, + ]), + ); + assert.equal(process.exitCode, 1); + } finally { + process.exitCode = previousExitCode; + } + assert.ok(!existsSync(join(dir, "lost-ids.json"))); + assert.deepEqual(lines, ["was Done: 1", "was Todo: 1"]); + assert.match(errors.join("\n"), /2 different values/); +}); + +test("diff reports nothing to recover when no card lost a value", async (t) => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "board-28-snapshot.json"); + writeFileSync(snapshotPath, snapshotFile([{ id: "a", status: "Done" }])); + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + await main( + ["--phase", "diff", "--snapshot", snapshotPath], + dumpSpawn([{ id: "a", status: "Done" }]), + ); + assert.ok(!existsSync(join(dir, "lost-ids.json"))); + assert.deepEqual(lines, ["lost: 0 cards — nothing to recover"]); +}); + +test("diff refuses a truncated broken-board dump", async () => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "s.json"); + writeFileSync(snapshotPath, snapshotFile([])); + await assert.rejects( + main( + ["--phase", "diff", "--snapshot", snapshotPath], + dumpSpawn([{ id: "a", status: null }], 500), + ), + /INCOMPLETE/, + ); +}); + +test("diff refuses a truncated snapshot by its own totalCount", async () => { + // A snapshot whose items fall short of its totalCount omits cards whose + // lost values can never be recovered from it — reapply would then report + // success while leaving them orphaned. The same shape check refuses a + // file with no items array or no totalCount at all. + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "s.json"); + writeFileSync( + snapshotPath, + JSON.stringify({ items: [{ id: "a", status: "Done" }], totalCount: 500 }), + ); + await assert.rejects( + main(["--phase", "diff", "--snapshot", snapshotPath], () => + assert.fail("nothing should be spawned"), + ), + /not a complete board snapshot \(1 items of totalCount 500\)/, + ); + writeFileSync(snapshotPath, JSON.stringify({ items: [] })); + await assert.rejects( + main(["--phase", "diff", "--snapshot", snapshotPath], () => + assert.fail("nothing should be spawned"), + ), + /not a complete board snapshot/, + ); +}); + +test("diff removes a stale lost-ids.json before doing anything", async (t) => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "board-28-snapshot.json"); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, JSON.stringify(["stale"])); + + // A failing diff (unreadable snapshot) must not preserve the stale file… + await assert.rejects( + main(["--phase", "diff", "--snapshot", snapshotPath], () => + assert.fail("nothing should be spawned"), + ), + ); + assert.ok(!existsSync(lostPath)); + + // …and neither does a diff that finds nothing to recover. + writeFileSync(lostPath, JSON.stringify(["stale"])); + writeFileSync(snapshotPath, snapshotFile([{ id: "a", status: "Done" }])); + t.mock.method(console, "log", () => {}); + await main( + ["--phase", "diff", "--snapshot", snapshotPath], + dumpSpawn([{ id: "a", status: "Done" }]), + ); + assert.ok(!existsSync(lostPath)); +}); + +test("diff refuses an explicit --board that contradicts the snapshot", async () => { + // A board-11 snapshot diffed against board 28's dump would match nothing + // and confidently print "lost: 0 cards". + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "s.json"); + writeFileSync(snapshotPath, snapshotFile([{ id: "a", status: "Done" }], 11)); + await assert.rejects( + main(["--phase", "diff", "--snapshot", snapshotPath, "--board", "28"], () => + assert.fail("nothing should be spawned"), + ), + /--board 28 does not match the snapshot's board #11/, + ); +}); + +test("diff takes its board from the snapshot and resolves fields on it", async (t) => { + // No --board flag: the snapshot's recorded board (11) decides which board + // is dumped and which board the lost file names. + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "s.json"); + writeFileSync(snapshotPath, snapshotFile([{ id: "a", status: "Done" }], 11)); + const boardsAsked = []; + const spawn = (cmd, args) => { + const joined = args.join(" "); + if (joined.includes("field-list") || joined.includes("item-list")) { + boardsAsked.push(args[2]); + } + if (joined.includes("field-list")) { + return { + status: 0, + stdout: JSON.stringify({ fields: [{ id: "F", name: "Status" }] }), + stderr: "", + }; + } + return { + status: 0, + stdout: JSON.stringify({ + items: [{ id: "a", status: null }], + totalCount: 1, + }), + stderr: "", + }; + }; + t.mock.method(console, "log", () => {}); + await main(["--phase", "diff", "--snapshot", snapshotPath], spawn); + assert.deepEqual(boardsAsked, ["11", "11"]); + assert.equal( + JSON.parse(readFileSync(join(dir, "lost-ids.json"), "utf8")).board, + 11, + ); +}); + +test("diff refuses a field the board does not have", async () => { + // A typo (--field Priorty) would read every card's value as undefined, + // match nothing, and falsely report no lost cards. + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "s.json"); + writeFileSync(snapshotPath, snapshotFile([{ id: "a", status: "Done" }])); + await assert.rejects( + main( + ["--phase", "diff", "--snapshot", snapshotPath, "--field", "Priorty"], + dumpSpawn([{ id: "a", status: null }]), + ), + /board #28 has no "Priorty" field/, + ); + assert.ok(!existsSync(join(dir, "lost-ids.json"))); +}); + +test("diff requires an explicit --board for a snapshot with no board key", async (t) => { + // A hand-taken `gh project item-list` dump predates the recorded key and + // cannot prove its board — defaulting would diff a #11 snapshot against + // #28 and confidently report "lost: 0 cards". + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "s.json"); + writeFileSync( + snapshotPath, + JSON.stringify({ items: [{ id: "a", status: "Done" }], totalCount: 1 }), + ); + await assert.rejects( + main(["--phase", "diff", "--snapshot", snapshotPath], () => + assert.fail("nothing should be spawned"), + ), + /records no board — pass --board/, + ); + // With the flag the legacy snapshot still diffs, against the named board. + t.mock.method(console, "log", () => {}); + await main( + ["--phase", "diff", "--snapshot", snapshotPath, "--board", "28"], + dumpSpawn([{ id: "a", status: null }]), + ); + assert.equal( + JSON.parse(readFileSync(join(dir, "lost-ids.json"), "utf8")).board, + 28, + ); +}); + +test("diff refuses a snapshot recording an unsupported board", async () => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const snapshotPath = join(dir, "s.json"); + writeFileSync(snapshotPath, snapshotFile([{ id: "a", status: "Done" }], 12)); + await assert.rejects( + main(["--phase", "diff", "--snapshot", snapshotPath], () => + assert.fail("nothing should be spawned"), + ), + /must be 28 \(v2\) or 11 \(v1\)/, + ); +}); + +test("diff refuses a snapshot path inside the repo", async () => { + // lost-ids.json is derived beside the snapshot — written in the worktree + // it is private board data one `git add -A` from a PR. + await assert.rejects( + main( + [ + "--phase", + "diff", + "--snapshot", + join(process.cwd(), "board-28-snapshot.json"), + ], + () => assert.fail("nothing should be spawned"), + ), + /private/, + ); +}); + +/** A valid lost file as --phase diff writes it. */ +const lostFile = (ids, value = "Done") => + JSON.stringify({ board: 28, field: "Status", value, ids }); + +/** + * A reapply spawn: `items` is the current board (preflight dump AND the + * per-card reads), `options` the Status field's option list. `onEdit` + * collects item-edit args when provided; absent, an edit is a failure. + */ +const reapplySpawn = + (items, { options = [{ id: "opt_new", name: "Done" }], onEdit } = {}) => + (cmd, args) => { + const joined = args.join(" "); + if (joined.includes("project view")) { + return { status: 0, stdout: JSON.stringify({ id: "PVT_x" }), stderr: "" }; + } + if (joined.includes("field-list")) { + return { + status: 0, + stdout: JSON.stringify({ + fields: [{ id: "F_status", name: "Status", options }], + }), + stderr: "", + }; + } + if (joined.includes("item-list")) { + return { + status: 0, + stdout: JSON.stringify({ items, totalCount: items.length }), + stderr: "", + }; + } + if (joined.includes("api graphql")) { + // The per-card read immediately before an edit. + const id = args.find((arg) => arg.startsWith("id=")).slice(3); + const item = items.find((candidate) => candidate.id === id); + const node = + item === undefined + ? null + : { + fieldValueByName: + item.status == null ? null : { name: item.status }, + }; + return { + status: 0, + stdout: JSON.stringify({ data: { node } }), + stderr: "", + }; + } + if (joined.includes("item-edit")) { + assert.ok(onEdit, `no edit may run here: ${joined}`); + onEdit(args); + return { status: 0, stdout: "{}", stderr: "" }; + } + assert.fail(`unexpected gh call: ${joined}`); + }; + +test("reapply edits each lost card with pacing and reports the count", async (t) => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, lostFile(["a", "b"])); + + const edits = []; + const spawn = reapplySpawn( + [ + { id: "a", status: null }, + { id: "b", status: null }, + ], + { onEdit: (args) => edits.push(args) }, + ); + const sleeps = []; + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + await main( + ["--phase", "reapply", "--lost", lostPath, "--option-id", "opt_new"], + spawn, + async (ms) => sleeps.push(ms), + ); + assert.equal(edits.length, 2); + assert.ok(edits.every((args) => args.includes("opt_new"))); + assert.deepEqual(sleeps, [400, 400]); + assert.deepEqual(lines, ["reapplied: 2 cards → option opt_new"]); +}); + +test("reapply refuses a lost file that is not diff's own format", async () => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + // The pre-#2559 bare-array format records no field/value, so --option-id + // could not be verified against what the cards held — refused outright. + writeFileSync(lostPath, JSON.stringify(["a", "b"])); + await assert.rejects( + main(["--phase", "reapply", "--lost", lostPath, "--option-id", "x"], () => { + assert.fail("nothing should be spawned"); + }), + /not a lost file written by --phase diff/, + ); +}); + +test("reapply refuses a lost file recording an unsupported board", async () => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync( + lostPath, + JSON.stringify({ board: 12, field: "Status", value: "Done", ids: ["a"] }), + ); + await assert.rejects( + main(["--phase", "reapply", "--lost", lostPath, "--option-id", "x"], () => { + assert.fail("nothing should be spawned"); + }), + /must be 28 \(v2\) or 11 \(v1\)/, + ); +}); + +test("reapply refuses an explicit flag that contradicts the lost file", async () => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, lostFile(["a"])); + await assert.rejects( + main( + [ + "--phase", + "reapply", + "--lost", + lostPath, + "--option-id", + "x", + "--board", + "11", + ], + () => assert.fail("nothing should be spawned"), + ), + /--board 11 does not match the lost file's board #28/, + ); + await assert.rejects( + main( + [ + "--phase", + "reapply", + "--lost", + lostPath, + "--option-id", + "x", + "--field", + "Priority", + ], + () => assert.fail("nothing should be spawned"), + ), + /--field Priority does not match the lost file's field "Status"/, + ); +}); + +test("reapply refuses an option id the field does not have", async () => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, lostFile(["a"])); + await assert.rejects( + main( + ["--phase", "reapply", "--lost", lostPath, "--option-id", "bogus"], + reapplySpawn([{ id: "a", status: null }]), + ), + /"Status" has no option with id bogus \(has: "Done" \(opt_new\)\)/, + ); +}); + +test("reapply refuses an option whose name is not the recorded value", async () => { + // A valid option id from the SAME field that is not the recreated option + // would silently rewrite every lost card to the wrong value. + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, lostFile(["a"])); + await assert.rejects( + main( + ["--phase", "reapply", "--lost", lostPath, "--option-id", "opt_todo"], + reapplySpawn([{ id: "a", status: null }], { + options: [ + { id: "opt_new", name: "Done" }, + { id: "opt_todo", name: "Todo" }, + ], + }), + ), + /is "Todo" but the lost cards held "Done"/, + ); +}); + +test("reapply refuses a lost card that is no longer blank", async () => { + // Someone legitimately set the card between diff and reapply — the stale + // list must not overwrite that newer value. Caught by the preflight, + // before any edit. + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, lostFile(["a", "b"])); + await assert.rejects( + main( + ["--phase", "reapply", "--lost", lostPath, "--option-id", "opt_new"], + reapplySpawn([ + { id: "a", status: null }, + { id: "b", status: "In Progress" }, + ]), + ), + /no longer blank.*stale.*re-run --phase diff/s, + ); +}); + +test("reapply refuses a lost card that no longer exists", async () => { + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, lostFile(["a", "gone"])); + await assert.rejects( + main( + ["--phase", "reapply", "--lost", lostPath, "--option-id", "opt_new"], + reapplySpawn([{ id: "a", status: null }]), + ), + /gone or no longer blank.*gone/s, + ); +}); + +test("reapply aborts mid-loop when a card is set during the run", async (t) => { + // The preflight passes (both cards blank at the start), then card b is set + // while card a's edit is pacing — the per-card read must catch it and no + // edit may land on b. + const dir = mkdtempSync(join(tmpdir(), "board-recover-test-")); + const lostPath = join(dir, "lost-ids.json"); + writeFileSync(lostPath, lostFile(["a", "b"])); + const items = [ + { id: "a", status: null }, + { id: "b", status: null }, + ]; + const edits = []; + const spawn = reapplySpawn(items, { + onEdit: (args) => { + edits.push(args); + // Simulate the concurrent maintainer: after a's edit, b gets a value. + items[1].status = "Todo"; + }, + }); + t.mock.method(console, "log", () => {}); + await assert.rejects( + main( + ["--phase", "reapply", "--lost", lostPath, "--option-id", "opt_new"], + spawn, + async () => {}, + ), + /card b is no longer blank \("Todo"\).*went stale mid-run \(1 of 2 reapplied\)/s, + ); + assert.equal(edits.length, 1); + assert.ok(edits[0].includes("a")); +}); diff --git a/scripts/board-snapshot.mjs b/scripts/board-snapshot.mjs new file mode 100644 index 0000000000..76ec29b1dd --- /dev/null +++ b/scripts/board-snapshot.mjs @@ -0,0 +1,77 @@ +#!/usr/bin/env node +// Snapshot a board before touching a field's options (#2558) — `npm run +// board:snapshot [-- --board 28] [--dir <path>]`. The snapshot block +// board-ops previously transcribed inline — one command that is the +// difference between a five-minute restore and reconstructing ~200 statuses +// by inference. +// +// Two refusals carried over from the inline block, both load-bearing: +// +// - A truncated snapshot cannot restore the cards it dropped, so the dump is +// trusted only when complete (`itemListComplete`) and nothing is written +// otherwise. +// - The boards are private, so a snapshot is a full dump of item ids and +// every card's Status and Priority. Written inside the working tree it is +// one `git add -A` away from being published in a PR, so the default is a +// fresh temp dir and a `--dir` under the current working directory is +// refused. + +import { spawnSync } from "node:child_process"; +import { mkdtempSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { parseArgs } from "node:util"; +import { + DEFAULT_BOARD, + assertOutsideRepo, + itemListComplete, + requireSupportedBoard, +} from "./lib/board.mjs"; +import { requirePositiveInt } from "./lib/gh.mjs"; + +export function parseSnapshotArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { board: { type: "string" }, dir: { type: "string" } }, + }); + return { + board: requireSupportedBoard( + values.board === undefined + ? DEFAULT_BOARD + : requirePositiveInt(values.board, "--board"), + ), + dir: values.dir, + }; +} + +/** + * The canonical-path outside-repo refusal lives in lib/board.mjs + * (`canonical` / `assertOutsideRepo`) — board-recover.mjs applies the same + * check to the directory it derives lost-ids.json into. + */ + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { board, dir } = parseSnapshotArgs(argv); + + const target = dir ?? mkdtempSync(join(tmpdir(), "board-")); + assertOutsideRepo(target, process.cwd()); + + // itemListComplete throws on a truncated dump, so nothing partial is written. + // The board number is recorded IN the snapshot so board-recover's diff can + // prove the snapshot belongs to the board it is diffing against instead of + // trusting a default or a flag. + const dump = { board, ...itemListComplete(spawn, board) }; + const path = join(target, `board-${board}-snapshot.json`); + // The dump is private: 0600 so a shared --dir (e.g. /tmp) never leaves it + // world-readable, and "wx" so an existing file (or a symlink planted at the + // path) is refused rather than followed or overwritten. + writeFileSync(path, JSON.stringify(dump, null, 2), { + mode: 0o600, + flag: "wx", + }); + console.log(`snapshot: ${path} (${dump.totalCount} items)`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/board-snapshot.test.mjs b/scripts/board-snapshot.test.mjs new file mode 100644 index 0000000000..3c9ee0c98e --- /dev/null +++ b/scripts/board-snapshot.test.mjs @@ -0,0 +1,85 @@ +// Tests for scripts/board-snapshot.mjs (#2558) — the two refusals: nothing +// is written from a truncated dump, and nothing is ever written inside the +// repo (the boards are private; a snapshot in the worktree is one +// `git add -A` from a PR). Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + existsSync, + mkdirSync, + mkdtempSync, + readFileSync, + readdirSync, + statSync, + symlinkSync, +} from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { main, parseSnapshotArgs } from "./board-snapshot.mjs"; +import { assertOutsideRepo } from "./lib/board.mjs"; +test("parseSnapshotArgs defaults the board and passes --dir through", () => { + assert.deepEqual(parseSnapshotArgs([]), { board: 28, dir: undefined }); + assert.deepEqual(parseSnapshotArgs(["--board", "11", "--dir", "/tmp/x"]), { + board: 11, + dir: "/tmp/x", + }); +}); + +test("assertOutsideRepo refuses the repo root and anything under it", () => { + assert.throws(() => assertOutsideRepo("/repo", "/repo"), /private/); + assert.throws(() => assertOutsideRepo("/repo/sub", "/repo"), /private/); + // A sibling whose name shares the prefix is fine. + assertOutsideRepo("/repo-sibling", "/repo"); + assertOutsideRepo("/elsewhere", "/repo"); +}); + +test("assertOutsideRepo resolves symlinks — a link into the repo is refused", () => { + const outside = mkdtempSync(join(tmpdir(), "board-snapshot-test-")); + const repo = join(outside, "repo"); + const linkToRepo = join(outside, "to-repo"); + mkdirSync(repo); + symlinkSync(repo, linkToRepo); + assert.throws(() => assertOutsideRepo(linkToRepo, repo), /private/); + // And the repo root reached through its own symlink still anchors the check. + assert.throws( + () => assertOutsideRepo(join(repo, "sub"), linkToRepo), + /private/, + ); +}); + +const ITEMS = [{ id: "PVTI_a", status: "Todo" }]; +const spawnScript = (totalCount) => () => ({ + status: 0, + stdout: JSON.stringify({ items: ITEMS, totalCount }), + stderr: "", +}); + +test("main writes the verified dump and prints its path", (t) => { + const dir = mkdtempSync(join(tmpdir(), "board-snapshot-test-")); + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + main(["--dir", dir], spawnScript(1)); + const path = join(dir, "board-28-snapshot.json"); + assert.deepEqual(lines, [`snapshot: ${path} (1 items)`]); + const written = JSON.parse(readFileSync(path, "utf8")); + assert.deepEqual(written.items, ITEMS); + // The board is recorded in the snapshot so board-recover's diff can prove + // the snapshot belongs to the board it diffs against. + assert.equal(written.board, 28); + // The dump is private: owner-only, and an existing file is never reused. + assert.equal(statSync(path).mode & 0o777, 0o600); + assert.throws(() => main(["--dir", dir], spawnScript(1)), /EEXIST/); +}); + +test("main writes nothing from a truncated dump", () => { + const dir = mkdtempSync(join(tmpdir(), "board-snapshot-test-")); + assert.throws(() => main(["--dir", dir], spawnScript(500)), /INCOMPLETE/); + assert.deepEqual(readdirSync(dir), []); +}); + +test("main refuses a --dir inside the working directory", () => { + const inside = join(process.cwd(), "pr-screenshots"); + assert.throws(() => main(["--dir", inside], spawnScript(1)), /private/); + assert.equal(existsSync(join(inside, "board-28-snapshot.json")), false); +}); diff --git a/scripts/board-sweep.mjs b/scripts/board-sweep.mjs new file mode 100644 index 0000000000..0599c40bb7 --- /dev/null +++ b/scripts/board-sweep.mjs @@ -0,0 +1,82 @@ +#!/usr/bin/env node +// Sweep for unboarded open issues (#2558) — `npm run board:sweep`. Pass 1 of +// the issue-triage skill, which previously transcribed this as a mktemp + +// dual-dump + jq-union block. +// +// Diffs the open issues against BOTH boards — diffing against #28 alone +// reports a v1 issue correctly carded on #11 as unboarded, which is how a +// past sweep double-boarded one (#1929). Both dumps are trusted only when +// complete (`itemListComplete`), because a silently truncated listing makes a +// carded issue read as unboarded and get double-carded. Prints each unboarded +// issue with its destination (milestoned already → Todo, else → Incoming) and +// exits non-zero when any exist, so "the sweep is clean" is checkable. + +import { spawnSync } from "node:child_process"; +import { REPO_SLUG, ghJson } from "./lib/gh.mjs"; +import { itemListComplete } from "./lib/board.mjs"; + +export const BOARDS = [28, 11]; +const ISSUE_LIMIT = 2000; + +/** Issue numbers carded on a board, filtered to this repo's real issues. */ +export function boardedNumbers(items) { + return items + .filter( + (item) => + item?.content?.type === "Issue" && + item.content.repository === REPO_SLUG, + ) + .map((item) => item.content.number); +} + +/** The unboarded open issues, each with the destination column it should get. */ +export function unboarded(openIssues, boarded) { + const carded = new Set(boarded); + return openIssues + .filter((issue) => !carded.has(issue.number)) + .map((issue) => ({ + number: issue.number, + destination: issue.milestone + ? `Todo (has milestone ${issue.milestone.title})` + : "Incoming", + })); +} + +export function main(_argv = process.argv.slice(2), spawn = spawnSync) { + const open = ghJson(spawn, [ + "issue", + "list", + "--repo", + REPO_SLUG, + "--state", + "open", + "--limit", + String(ISSUE_LIMIT), + "--json", + "number,milestone", + ]); + // `gh issue list` reports no total, so a listing AT the limit is treated as + // possibly truncated — the same refusal the board dumps get. + if (open.length >= ISSUE_LIMIT) { + throw new Error( + `open-issue listing hit --limit ${ISSUE_LIMIT} — raise it and re-run`, + ); + } + + const boarded = BOARDS.flatMap((board) => + boardedNumbers(itemListComplete(spawn, board).items), + ); + + const missing = unboarded(open, boarded); + for (const issue of missing) { + console.log(`#${issue.number}\t→ ${issue.destination}`); + } + console.log(`unboarded: ${missing.length}`); + if (missing.length > 0) { + process.exitCode = 1; + } +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/board-sweep.test.mjs b/scripts/board-sweep.test.mjs new file mode 100644 index 0000000000..fd39e15d7e --- /dev/null +++ b/scripts/board-sweep.test.mjs @@ -0,0 +1,106 @@ +// Tests for scripts/board-sweep.mjs (#2558) — the union across BOTH boards +// (diffing against #28 alone double-boards a correctly-carded v1 issue), the +// Todo/Incoming destination split, and the truncation refusals. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { boardedNumbers, main, unboarded } from "./board-sweep.mjs"; + +const SLUG = "modelcontextprotocol/inspector"; + +test("boardedNumbers keeps only this repo's real issues", () => { + assert.deepEqual( + boardedNumbers([ + { content: { type: "Issue", number: 1, repository: SLUG } }, + { content: { type: "Issue", number: 2, repository: "other/repo" } }, + { content: { type: "DraftIssue", title: "[GHSA-…]" } }, + ]), + [1], + ); +}); + +test("unboarded splits destinations by milestone", () => { + assert.deepEqual( + unboarded( + [ + { number: 1, milestone: { title: "2.5.0" } }, + { number: 2, milestone: null }, + { number: 3, milestone: null }, + ], + [3], + ), + [ + { number: 1, destination: "Todo (has milestone 2.5.0)" }, + { number: 2, destination: "Incoming" }, + ], + ); +}); + +/** A spawn answering the issue list and both board dumps. */ +function spawnScript({ open, boards }) { + return (cmd, args) => { + const joined = args.join(" "); + let payload; + if (joined.includes("issue list")) { + payload = open; + } else if (joined.includes("item-list")) { + const board = args[2]; + const items = boards[board] ?? []; + payload = { items, totalCount: items.length }; + } else { + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; +} + +const carded = (number) => ({ + content: { type: "Issue", number, repository: SLUG }, +}); + +test("main unions both boards and reports only truly unboarded issues", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const before = process.exitCode; + main( + [], + spawnScript({ + open: [ + { number: 1, milestone: null }, // carded on #28 + { number: 2, milestone: null }, // carded on #11 — NOT unboarded + { number: 3, milestone: { title: "2.5.0" } }, // unboarded + ], + boards: { 28: [carded(1)], 11: [carded(2)] }, + }), + ); + assert.deepEqual(lines, ["#3\t→ Todo (has milestone 2.5.0)", "unboarded: 1"]); + assert.equal(process.exitCode, 1); + process.exitCode = before; +}); + +test("main exits clean when every open issue is carded", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const before = process.exitCode; + main( + [], + spawnScript({ + open: [{ number: 1, milestone: null }], + boards: { 28: [carded(1)], 11: [] }, + }), + ); + assert.deepEqual(lines, ["unboarded: 0"]); + assert.equal(process.exitCode, before); +}); + +test("main refuses an issue listing at its own limit", () => { + const open = Array.from({ length: 2000 }, (_, i) => ({ + number: i + 1, + milestone: null, + })); + assert.throws( + () => main([], spawnScript({ open, boards: {} })), + /hit --limit 2000/, + ); +}); diff --git a/scripts/dco-check.mjs b/scripts/dco-check.mjs new file mode 100644 index 0000000000..f9135dc56f --- /dev/null +++ b/scripts/dco-check.mjs @@ -0,0 +1,240 @@ +#!/usr/bin/env node +// DCO signoff check (#2566) — `npm run dco:check -- --base <rev> [--head <rev>] +// [--exclude <rev>…]`. +// +// Replaces the probot DCO app, which was suspended and whose check simply +// stopped appearing after #1981 (2026-08-12). Nothing failed when it vanished, +// because it was never a required check — so this is a check the repo owns, +// run by `.github/workflows/dco.yml` on every v2 PR — any `v2/**` base, so +// stacked PRs too — and, as a backstop, on every push that lands on `v2/main` +// (#2616; v1 and milestone PRs into `main` are out of its scope). The PR job +// is meant to be made REQUIRED so a future outage blocks merges instead of +// passing silently. +// +// The rule is the app's: every commit in `base..head` must carry a +// `Signed-off-by: Name <email>` line whose name AND email match the commit's +// author or its committer (one identity — a name from one and an email from +// the other is not a match). Names and emails compare case-insensitively, +// after trimming. One relaxation: for a commit GitHub itself committed (a +// squash merge), an author EMAIL match is enough — see `webFlowAuthorMatch`. +// The app's two exemptions are kept: +// +// - merge commits (more than one parent), which certify nothing new; and +// - bot-authored commits — an author email of the GitHub noreply shape +// `<id>+<login>[bot]@users.noreply.github.com`. +// +// ⚠️ The bot exemption reads an email any committer can set, so it is not a +// security control — but neither is the trailer: `Signed-off-by:` is +// self-asserted text either way. The check exists to catch a FORGOTTEN +// signoff, which is the failure that actually happens, not a forged one. +// +// The range is read from local git, not the API, so the workflow needs only +// `contents: read` and a full-history checkout, and the same command works +// before pushing. `base..head` is exactly the commit list GitHub shows on a +// PR: everything reachable from head that the base branch does not already +// contain. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; + +// Records are NUL-separated (`git log -z`): `git commit` refuses a NUL in a +// message, so no body can split a record. Fields use \x1f, and the message +// is the LAST field and is re-joined, so a body that happens to contain \x1f +// still parses whole. +const FS = "\x1f"; +const RS = "\0"; +const LOG_FORMAT = ["%H", "%P", "%an", "%ae", "%cn", "%ce", "%B"].join("%x1f"); + +const SIGNOFF = /^\s*Signed-off-by:\s*(.+?)\s*<([^<>]+)>\s*$/gim; +const BOT_EMAIL = /^\d+\+[^@\s]+\[bot\]@users\.noreply\.github\.com$/i; + +export function parseDcoArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { + base: { type: "string" }, + head: { type: "string" }, + exclude: { type: "string", multiple: true }, + }, + }); + if (!values.base) { + throw new Error("--base <rev> is required (e.g. origin/v2/main)"); + } + return { + base: values.base, + head: values.head ?? "HEAD", + exclude: values.exclude ?? [], + }; +} + +/** Split `git log -z --format=<LOG_FORMAT>` output into commit records. */ +export function parseLog(stdout) { + return stdout + .split(RS) + .filter((record) => record.trim() !== "") + .map((record) => { + const [sha, parents, an, ae, cn, ce, ...body] = record.split(FS); + return { + sha, + parents: parents.split(" ").filter(Boolean), + author: { name: an, email: ae }, + committer: { name: cn, email: ce }, + message: body.join(FS), + }; + }); +} + +/** Every `Signed-off-by:` identity in a commit message. */ +export function signoffs(message) { + return [...message.matchAll(SIGNOFF)].map(([, name, email]) => ({ + name, + email, + })); +} + +const norm = (value) => value.trim().toLowerCase(); +const sameIdentity = (a, b) => + norm(a.name) === norm(b.name) && norm(a.email) === norm(b.email); + +// The identity GitHub commits as when it creates a commit itself — a squash +// merge, a rebase merge, a web edit. Matched as a whole identity, name and +// email, like every other comparison here: git metadata is user-set either +// way, but a commit that merely borrows the email is not a web-flow commit. +const GITHUB_WEB_FLOW = { name: "GitHub", email: "noreply@github.com" }; + +/** + * GitHub writes a squash-merge commit's author NAME from the merger's GitHub + * profile ("Cliff Hall"), while the squashed trailers carry their git + * `user.name` ("cliffhall") — same person, same email, different name. Four + * such squash merges sit on `v2/main` (#2213, #2322, #2379, #2383), and the + * push job (#2616) sees every one of them. So for a commit GitHub itself + * committed, an author EMAIL match is accepted. A signoff is still required, + * and every other commit still needs name and email from one identity. + */ +const webFlowAuthorMatch = (sig, commit) => + sameIdentity(commit.committer, GITHUB_WEB_FLOW) && + norm(sig.email) === norm(commit.author.email); + +/** + * Why this commit fails the check, or `null` when it passes or is exempt. + * Exempt commits return `null` as well — they are reported separately by + * `classify`. + */ +export function failureReason(commit) { + const found = signoffs(commit.message); + if (found.length === 0) return "no Signed-off-by trailer"; + const matches = found.some( + (sig) => + sameIdentity(sig, commit.author) || + sameIdentity(sig, commit.committer) || + webFlowAuthorMatch(sig, commit), + ); + if (matches) return null; + const listed = found.map((sig) => `${sig.name} <${sig.email}>`).join(", "); + return ( + `Signed-off-by (${listed}) matches neither the author ` + + `(${commit.author.name} <${commit.author.email}>) nor the committer ` + + `(${commit.committer.name} <${commit.committer.email}>)` + ); +} + +/** `"merge"`, `"bot"`, or `null` when the commit must be signed off. */ +export function exemption(commit) { + if (commit.parents.length > 1) return "merge"; + if (BOT_EMAIL.test(commit.author.email.trim())) return "bot"; + return null; +} + +/** Partition commits into checked/exempt and collect the failures. */ +export function classify(commits) { + const failures = []; + let exempt = 0; + for (const commit of commits) { + if (exemption(commit)) { + exempt++; + continue; + } + const reason = failureReason(commit); + if (reason) failures.push({ commit, reason }); + } + return { checked: commits.length - exempt, exempt, failures }; +} + +/** + * `--exclude <rev>` drops everything reachable from `rev` as well, on top of + * `base` (#2616). `local:dco` excludes `origin/main`: a milestone-merge branch + * is cut from `main`, so `origin/v2/main..HEAD` there is `main`'s own history + * — 117 already-released commits, the v1 era and the 2.0.0 candidates among + * them, almost none signed — and the gate the release runs on that branch + * would fail on all of it. A commit already on `main` has shipped, so it says + * nothing about the change being checked. On a feature branch the exclusion + * is a no-op, since it holds nothing `main` has that `v2/main` lacks. + */ +function describeRange(base, head, exclude) { + const range = `${base}..${head}`; + return exclude.length ? `${range} (excluding ${exclude.join(", ")})` : range; +} + +function readRange(spawn, base, head, exclude) { + const result = spawn( + "git", + [ + "log", + "-z", + `--format=${LOG_FORMAT}`, + `${base}..${head}`, + ...exclude.map((rev) => `^${rev}`), + ], + { encoding: "utf8", maxBuffer: 64 * 1024 * 1024 }, + ); + if (result.error) throw result.error; + if (result.status !== 0) { + throw new Error( + `git log ${describeRange(base, head, exclude)} failed: ${(result.stderr ?? "").trim()}`, + ); + } + return parseLog(result.stdout); +} + +/** Returns the process exit code: 0 when every commit passes, 1 otherwise. */ +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { base, head, exclude } = parseDcoArgs(argv); + const commits = readRange(spawn, base, head, exclude); + const range = describeRange(base, head, exclude); + const { checked, exempt, failures } = classify(commits); + + if (failures.length === 0) { + console.log( + `dco: OK — ${checked} commit(s) signed off in ${range}` + + (exempt ? ` (${exempt} merge/bot commit(s) exempt)` : ""), + ); + return 0; + } + + console.error( + `dco: FAIL — ${failures.length} of ${checked} commit(s) in ${range} lack a matching signoff:\n`, + ); + for (const { commit, reason } of failures) { + const subject = commit.message.split("\n", 1)[0]; + console.error(` ${commit.sha.slice(0, 12)} ${subject}\n ${reason}`); + } + console.error( + // `--rebase-merges` keeps any merge commit (and a conflict resolution + // that lives only in it) instead of flattening the branch; the merges + // themselves stay unsigned, which is fine since they are exempt. + `\nRepair (sole author, nobody building on the branch):\n` + + ` git rebase --rebase-merges --signoff ${base}\n` + + ` git push --force-with-lease\n` + + `Prevent it next time with \`git commit -s\`.`, + ); + return 1; +} + +if (import.meta.url === `file://${process.argv[1]}`) { + try { + process.exitCode = main(); + } catch (error) { + console.error(`dco: ${error.message}`); + process.exitCode = 2; + } +} diff --git a/scripts/dco-check.test.mjs b/scripts/dco-check.test.mjs new file mode 100644 index 0000000000..510b5e8851 --- /dev/null +++ b/scripts/dco-check.test.mjs @@ -0,0 +1,334 @@ +// Tests for scripts/dco-check.mjs (#2566) — the signoff rule one case per +// clause (identity match, the two exemptions, the failure report), plus a +// run of `main()` against a real throwaway git repository, since the log +// format and its parser only mean anything together. Run via +// `npm run test:scripts`. + +import { after, test } from "node:test"; +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import { mkdtempSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { + classify, + exemption, + failureReason, + main, + parseDcoArgs, + signoffs, +} from "./dco-check.mjs"; + +const ADA = { name: "Ada Lovelace", email: "ada@example.com" }; +const BOB = { name: "Bob Builder", email: "bob@example.com" }; + +function commit({ + author = ADA, + committer = author, + parents = ["p1"], + message = "subject\n", +} = {}) { + return { sha: "a".repeat(40), parents, author, committer, message }; +} + +const signed = (who, subject = "subject") => + `${subject}\n\nbody\n\nSigned-off-by: ${who.name} <${who.email}>\n`; + +test("parseDcoArgs requires --base and defaults --head to HEAD", () => { + assert.deepEqual(parseDcoArgs(["--base", "origin/v2/main"]), { + base: "origin/v2/main", + head: "HEAD", + exclude: [], + }); + assert.deepEqual(parseDcoArgs(["--base", "a", "--head", "b"]), { + base: "a", + head: "b", + exclude: [], + }); + assert.deepEqual( + parseDcoArgs(["--base", "a", "--exclude", "x", "--exclude", "y"]).exclude, + ["x", "y"], + ); + assert.throws(() => parseDcoArgs([]), /--base/); +}); + +test("signoffs reads every trailer, tolerating case and spacing", () => { + assert.deepEqual( + signoffs( + "x\n\nsigned-off-by: Ada Lovelace <ada@example.com> \nSigned-off-by: Bob Builder <bob@example.com>\n", + ), + [ADA, BOB], + ); + assert.deepEqual(signoffs("x\n\nSigned-off-by: no email here\n"), []); + // Mid-line text is not a trailer. + assert.deepEqual( + signoffs("mentions Signed-off-by: Ada <ada@example.com> inline\n"), + [], + ); +}); + +test("a signoff matching the author passes", () => { + assert.equal(failureReason(commit({ message: signed(ADA) })), null); +}); + +test("a signoff matching the committer passes", () => { + assert.equal( + failureReason( + commit({ author: BOB, committer: ADA, message: signed(ADA) }), + ), + null, + ); +}); + +test("matching ignores case and surrounding whitespace", () => { + assert.equal( + failureReason( + commit({ + message: signed({ name: "ada lovelace", email: "ADA@Example.com" }), + }), + ), + null, + ); +}); + +test("a missing trailer fails", () => { + assert.equal(failureReason(commit()), "no Signed-off-by trailer"); +}); + +test("a trailer for someone else fails and names every identity", () => { + const reason = failureReason(commit({ message: signed(BOB) })); + assert.match(reason, /Bob Builder <bob@example.com>/); + assert.match(reason, /author \(Ada Lovelace <ada@example.com>\)/); +}); + +test("name from one identity and email from the other is not a match", () => { + const reason = failureReason( + commit({ + author: ADA, + committer: BOB, + message: signed({ name: ADA.name, email: BOB.email }), + }), + ); + assert.notEqual(reason, null); +}); + +// A GitHub squash merge: author name from the GitHub profile, committer +// GitHub's web-flow identity, trailers carrying the git `user.name` (#2616). +const GITHUB = { name: "GitHub", email: "noreply@github.com" }; +const ADA_PROFILE = { name: "Ada L.", email: ADA.email }; +const ADA_GIT = { name: "ada", email: ADA.email }; + +test("a GitHub-committed squash passes on an author email match", () => { + assert.equal( + failureReason( + commit({ + author: ADA_PROFILE, + committer: GITHUB, + message: signed(ADA_GIT), + }), + ), + null, + ); +}); + +test("the email-only match applies to GitHub-committed commits alone", () => { + assert.notEqual( + failureReason( + commit({ author: ADA_PROFILE, committer: BOB, message: signed(ADA_GIT) }), + ), + null, + ); +}); + +test("the web-flow identity is matched whole, not by its email alone", () => { + assert.notEqual( + failureReason( + commit({ + author: ADA_PROFILE, + committer: { name: "Mallory", email: GITHUB.email }, + message: signed(ADA_GIT), + }), + ), + null, + ); +}); + +test("a GitHub-committed commit still needs the AUTHOR's email", () => { + assert.notEqual( + failureReason( + commit({ author: ADA_PROFILE, committer: GITHUB, message: signed(BOB) }), + ), + null, + ); + assert.equal( + failureReason(commit({ author: ADA_PROFILE, committer: GITHUB })), + "no Signed-off-by trailer", + ); +}); + +test("exemption: merge commits and noreply bot authors only", () => { + assert.equal(exemption(commit({ parents: ["p1", "p2"] })), "merge"); + assert.equal( + exemption( + commit({ + author: { + name: "github-actions[bot]", + email: "41898282+github-actions[bot]@users.noreply.github.com", + }, + }), + ), + "bot", + ); + // A noreply address that is not a bot's is not exempt. + assert.equal( + exemption( + commit({ + author: { name: "Ada", email: "123+ada@users.noreply.github.com" }, + }), + ), + null, + ); + // A root commit (no parents) is checked like any other. + assert.equal(exemption(commit({ parents: [] })), null); +}); + +test("classify: one unsigned commit fails the range — no partial credit", () => { + const result = classify([ + commit({ message: signed(ADA) }), + commit({ message: "unsigned\n" }), + commit({ parents: ["p1", "p2"] }), + ]); + assert.equal(result.checked, 2); + assert.equal(result.exempt, 1); + assert.equal(result.failures.length, 1); + assert.equal(result.failures[0].reason, "no Signed-off-by trailer"); +}); + +// --- main() against a real repository ------------------------------------- + +const repos = []; +after(() => { + for (const dir of repos) rmSync(dir, { recursive: true, force: true }); +}); + +function git(cwd, args, env = {}) { + const result = spawnSync("git", args, { + cwd, + encoding: "utf8", + env: { ...process.env, ...env }, + }); + assert.equal(result.status, 0, result.stderr); + return result.stdout.trim(); +} + +/** A repo whose `base` branch holds one unsigned commit (outside any range). */ +function makeRepo() { + const dir = mkdtempSync(path.join(tmpdir(), "dco-check-")); + repos.push(dir); + git(dir, ["init", "-q", "-b", "base"]); + git(dir, ["config", "user.name", ADA.name]); + git(dir, ["config", "user.email", ADA.email]); + git(dir, ["config", "commit.gpgsign", "false"]); + git(dir, ["commit", "-q", "--allow-empty", "-m", "base, unsigned"]); + git(dir, ["checkout", "-q", "-b", "feature"]); + return dir; +} + +const spawnIn = (cwd) => (cmd, args, opts) => + spawnSync(cmd, args, { ...opts, cwd }); + +function runMain(dir, extra = []) { + const out = []; + const err = []; + const log = console.log; + const error = console.error; + console.log = (line) => out.push(line); + console.error = (line) => err.push(line); + try { + const code = main(["--base", "base", ...extra], spawnIn(dir)); + return { code, out: out.join("\n"), err: err.join("\n") }; + } finally { + console.log = log; + console.error = error; + } +} + +test("main passes a range of signed commits and ignores the base's history", () => { + const dir = makeRepo(); + git(dir, ["commit", "-q", "-s", "--allow-empty", "-m", "one"]); + git(dir, ["commit", "-q", "-s", "--allow-empty", "-m", "two\n\nbody"]); + const { code, out } = runMain(dir); + assert.equal(code, 0); + assert.match(out, /dco: OK — 2 commit\(s\) signed off in base\.\.HEAD/); +}); + +test("main --exclude drops a released branch's history from the range", () => { + // The milestone-merge shape (#2616): a branch cut from `main`, whose own + // unsigned history must not fail a gate run against `v2/main`. + const dir = makeRepo(); + git(dir, ["commit", "-q", "--allow-empty", "-m", "released, unsigned"]); + git(dir, ["branch", "released"]); + git(dir, ["commit", "-q", "-s", "--allow-empty", "-m", "new, signed"]); + assert.equal(runMain(dir).code, 1, "without --exclude the old commit fails"); + const { code, out } = runMain(dir, ["--exclude", "released"]); + assert.equal(code, 0); + assert.match( + out, + /1 commit\(s\) signed off in base\.\.HEAD \(excluding released\)/, + ); +}); + +test("main fails on one unsigned commit and prints the repair", () => { + const dir = makeRepo(); + git(dir, ["commit", "-q", "-s", "--allow-empty", "-m", "signed"]); + git(dir, ["commit", "-q", "--allow-empty", "-m", "forgot the signoff"]); + const { code, err } = runMain(dir); + assert.equal(code, 1); + assert.match(err, /1 of 2 commit\(s\)/); + assert.match(err, /forgot the signoff\n {4}no Signed-off-by trailer/); + assert.match(err, /git rebase --rebase-merges --signoff base/); + assert.match(err, /git push --force-with-lease/); +}); + +test("main exempts a merge commit and a bot-authored commit", () => { + const dir = makeRepo(); + git(dir, ["commit", "-q", "-s", "--allow-empty", "-m", "signed"]); + git(dir, ["checkout", "-q", "-b", "side", "base"]); + git(dir, ["commit", "-q", "-s", "--allow-empty", "-m", "side, signed"]); + git(dir, ["checkout", "-q", "feature"]); + git(dir, ["merge", "-q", "--no-ff", "--no-edit", "side"]); + git(dir, ["commit", "-q", "--allow-empty", "-m", "bot, unsigned"], { + GIT_AUTHOR_NAME: "github-actions[bot]", + GIT_AUTHOR_EMAIL: "41898282+github-actions[bot]@users.noreply.github.com", + }); + const { code, out } = runMain(dir); + assert.equal(code, 0); + assert.match( + out, + /2 commit\(s\) signed off .* \(2 merge\/bot commit\(s\) exempt\)/, + ); +}); + +test("main parses a message containing the old separators intact", () => { + const dir = makeRepo(); + git(dir, [ + "commit", + "-q", + "-s", + "--allow-empty", + "-m", + "odd \x1e record \x1f field bytes\n\nbody \x1e\x1f too", + ]); + git(dir, ["commit", "-q", "-s", "--allow-empty", "-m", "next"]); + const { code, out } = runMain(dir); + assert.equal(code, 0); + assert.match(out, /2 commit\(s\) signed off/); +}); + +test("main throws on a revision git cannot resolve", () => { + const dir = makeRepo(); + assert.throws( + () => main(["--base", "no-such-ref"], spawnIn(dir)), + /git log no-such-ref\.\.HEAD failed/, + ); +}); diff --git a/scripts/dependabot-alerts.mjs b/scripts/dependabot-alerts.mjs index 26f7f53900..3e8d621a0f 100644 --- a/scripts/dependabot-alerts.mjs +++ b/scripts/dependabot-alerts.mjs @@ -56,6 +56,7 @@ import { spawnSync } from "node:child_process"; import { readFileSync } from "node:fs"; import semver from "semver"; +import { escapeTableCell as cell } from "./lib/markdown-cell.mjs"; /** Board #28 (v2). The project and field node ids are stable; option ids are not. */ export const PROJECT_ID = "PVT_kwDOCt2Azc4BJVxt"; @@ -431,9 +432,6 @@ export function buildIssueTitle(group) { return `chore(deps): bump \`${group.package}\` to \`${group.fixedIn}\` in \`${group.manifestPath}\` (${n} ${n === 1 ? "advisory" : "advisories"})`; } -/** Escape a value going into a Markdown table cell. */ -const cell = (value) => String(value).replace(/\|/g, "\\|"); - const PLACEMENT_DOC = "https://github.com/modelcontextprotocol/inspector/blob/v2/main/AGENTS.md#dependency-placement"; @@ -675,8 +673,7 @@ export function buildNewAdvisoryComment(group, added) { const rows = group.advisories .filter((a) => added.includes(a.ghsa)) .map( - (a) => - `| [${a.ghsa}](${a.url}) | ${a.severity} | ${a.summary.replace(/\|/g, "\\|")} |`, + (a) => `| [${a.ghsa}](${a.url}) | ${a.severity} | ${cell(a.summary)} |`, ) .join("\n"); return [ diff --git a/scripts/dependency-refresh.mjs b/scripts/dependency-refresh.mjs index 79a4bf4103..735212ec0b 100644 --- a/scripts/dependency-refresh.mjs +++ b/scripts/dependency-refresh.mjs @@ -52,6 +52,7 @@ export const INSTALLS = [ { dir: "clients/cli", label: "clients/cli" }, { dir: "clients/tui", label: "clients/tui" }, { dir: "clients/launcher", label: "clients/launcher" }, + { dir: "clients/mcpdo", label: "clients/mcpdo" }, ]; /** Where the `uses:` refs this sweep checks live, relative to the repo root. */ diff --git a/scripts/dependency-refresh.test.mjs b/scripts/dependency-refresh.test.mjs index a70b151e9c..951df72445 100644 --- a/scripts/dependency-refresh.test.mjs +++ b/scripts/dependency-refresh.test.mjs @@ -305,6 +305,22 @@ test("main throws when npm outdated exits with an undocumented status", () => { assert.equal(ghCall(spawn, "create"), undefined); }); +test("INSTALLS enrolls the root and every client install", () => { + // A client absent here is silently skipped by the monthly sweep — its + // client-only devDependencies would never show up in `npm outdated`. + assert.deepEqual( + INSTALLS.map((i) => i.dir), + [ + ".", + "clients/web", + "clients/cli", + "clients/tui", + "clients/launcher", + "clients/mcpdo", + ], + ); +}); + test("main sweeps every install and files one milestoned issue", () => { const spawn = fakeSpawn({ outdated: { diff --git a/scripts/install-clients.mjs b/scripts/install-clients.mjs index 94867b0a45..815ef2da27 100644 --- a/scripts/install-clients.mjs +++ b/scripts/install-clients.mjs @@ -27,7 +27,7 @@ import { dirname, join, resolve, sep } from "node:path"; import { fileURLToPath } from "node:url"; const repoRoot = resolve(dirname(fileURLToPath(import.meta.url)), ".."); -const CLIENTS = ["web", "cli", "tui", "launcher"]; +const CLIENTS = ["web", "cli", "mcpdo", "tui", "launcher"]; if (process.env.INSPECTOR_SKIP_CLIENT_INSTALL) { console.log( diff --git a/scripts/lib/board.mjs b/scripts/lib/board.mjs new file mode 100644 index 0000000000..1fa97188b2 --- /dev/null +++ b/scripts/lib/board.mjs @@ -0,0 +1,225 @@ +// Shared project-board plumbing for the maintainer-workflow scripts (#2558): +// `board-card-status.mjs`, `board-card-add.mjs`, `board-card-delete.mjs`, +// `board-draft-find.mjs`, `board-snapshot.mjs`, `board-sweep.mjs`, +// `board-audit.mjs` and `board-recover.mjs`. The invariants the board-ops and +// issue-triage skills could previously only state in prose next to their +// inline blocks are enforced here once, under test: +// +// - Every id is resolved BY NAME at run time. Single-select option ids are +// regenerated whenever a field's option list is edited (board-ops' hazard), +// so nothing here hardcodes one. +// - An issue's card is found FROM THE ISSUE (`projectItems`), selected by the +// board's node id — project numbers are per-owner, and the issue-side +// lookup has no exposure to board size. +// - A whole-board listing is trusted only when COMPLETE: `gh project +// item-list --limit N` truncates silently past N, so `itemListComplete` +// compares `.items | length` against `.totalCount` and throws rather than +// letting a truncated dump read as a smaller board (#2451's defect class). +// - A private dump NEVER lands inside the repo: `assertOutsideRepo` resolves +// symlinks before comparing, so a linked path cannot smuggle board data +// into the worktree (one `git add -A` from a PR). + +import { realpathSync } from "node:fs"; +import { basename, dirname, join, resolve, sep } from "node:path"; +import { OWNER, REPO, gh, ghGraphql, ghJson } from "./gh.mjs"; + +/** + * The canonical form of a path: symlinks resolved. A path that does not + * exist yet canonicalizes its deepest existing ancestor and re-joins the + * rest, so a planned subdirectory still anchors to the real tree. + * `resolve()` alone would let a symlinked path (e.g. /tmp/to-repo → the + * worktree) place a private dump inside the repo. + */ +export function canonical(path, realpath = realpathSync) { + const full = resolve(path); + try { + return realpath(full); + } catch { + const parent = dirname(full); + if (parent === full) { + return full; + } + return join(canonical(parent, realpath), basename(full)); + } +} + +/** Throw when `dir` is inside `cwd` — a private dump never lands in the worktree. */ +export function assertOutsideRepo(dir, cwd, realpath = realpathSync) { + const target = canonical(dir, realpath); + const root = canonical(cwd, realpath); + if (target === root || target.startsWith(root + sep)) { + throw new Error( + `refusing to write private board data inside the repo (${target}) — ` + + `the boards are private; use a directory outside ${root}`, + ); + } +} + +export const DEFAULT_BOARD = 28; + +/** + * The only boards these scripts may touch — v2's #28 and v1's #11 + * (AGENTS.md's two live boards). Project numbers are per-owner and the org + * holds other real projects, so an unrestricted `--board` typo would add or + * edit cards on an unrelated board — and any non-28 number would also bypass + * board-card-add's #28-only Priority requirement. + */ +export const SUPPORTED_BOARDS = [DEFAULT_BOARD, 11]; + +/** Refuse a board number outside SUPPORTED_BOARDS; returns it otherwise. */ +export function requireSupportedBoard(board, what = "--board") { + if (!SUPPORTED_BOARDS.includes(board)) { + throw new Error( + `${what} must be 28 (v2) or 11 (v1) — refusing to touch board #${board}`, + ); + } + return board; +} + +/** Default `--limit` for whole-board dumps — headroom over the board's size. */ +export const ITEM_LIST_LIMIT = 2000; + +/** Resolve a board's project node id by number. */ +export function projectId(spawn, board) { + const id = ghJson(spawn, [ + "project", + "view", + String(board), + "--owner", + OWNER, + "--format", + "json", + ]).id; + if (!id) { + throw new Error(`could not resolve project id for board #${board}`); + } + return id; +} + +/** A board's fields (names, ids, options), for name-based resolution. */ +export function boardFields(spawn, board) { + return ( + ghJson(spawn, [ + "project", + "field-list", + String(board), + "--owner", + OWNER, + "--format", + "json", + ]).fields ?? [] + ); +} + +/** Resolve a single-select field's id and one option's id, both by name. */ +export function fieldOption(fields, fieldName, optionName) { + const field = fields.find((candidate) => candidate.name === fieldName); + if (!field) { + throw new Error(`board has no "${fieldName}" field`); + } + const option = (field.options ?? []).find( + (candidate) => candidate.name === optionName, + ); + if (!option) { + const known = (field.options ?? []).map((o) => o.name).join(", "); + throw new Error( + `"${fieldName}" has no option "${optionName}" (has: ${known})`, + ); + } + return { fieldId: field.id, optionId: option.id }; +} + +/** The issue-side card query, parameterized by the field to read back. */ +export function cardQuery(fieldName) { + return `query($n:Int!){repository(owner:"${OWNER}",name:"${REPO}"){issue(number:$n){projectItems(first:100){nodes{id project{id} fieldValueByName(name:"${fieldName}"){... on ProjectV2ItemFieldSingleSelectValue{name}}}}}}}`; +} + +/** The issue's card on the given project, from a `projectItems` response. */ +export function cardOnProject(response, project) { + const nodes = response?.data?.repository?.issue?.projectItems?.nodes; + if (!Array.isArray(nodes)) { + throw new Error( + `unexpected projectItems response shape: ${JSON.stringify(response)}`, + ); + } + return nodes.find((node) => node?.project?.id === project); +} + +/** Find an issue's card on a project, reading back one field's value. */ +export function findCard(spawn, issue, project, fieldName = "Status") { + return cardOnProject( + ghGraphql(spawn, cardQuery(fieldName), { n: issue }), + project, + ); +} + +/** + * A whole-board dump, trusted only when complete. Returns the parsed + * `{ items, totalCount }` object so callers that persist it (the snapshot) + * write exactly what was verified. + */ +export function itemListComplete(spawn, board, limit = ITEM_LIST_LIMIT) { + const dump = ghJson(spawn, [ + "project", + "item-list", + String(board), + "--owner", + OWNER, + "--format", + "json", + "--limit", + String(limit), + ]); + if ( + !Array.isArray(dump.items) || + typeof dump.totalCount !== "number" || + dump.items.length !== dump.totalCount + ) { + throw new Error( + `board #${board} listing INCOMPLETE or malformed ` + + `(${dump.items?.length ?? "?"} of ${dump.totalCount ?? "?"}) — raise --limit`, + ); + } + return dump; +} + +/** + * Read one card's single-select field value at this moment, by item node id. + * Returns `{ exists, value }`: `exists` false when the item no longer + * resolves (a deleted card), `value` null when the field is blank. This is + * the per-card read board-recover's reapply makes immediately before each + * edit — a whole-board preflight cannot hold across a paced loop. + */ +export function itemFieldValue(spawn, itemId, fieldName) { + const response = ghGraphql( + spawn, + `query($id:ID!){node(id:$id){... on ProjectV2Item{fieldValueByName(name:"${fieldName}"){... on ProjectV2ItemFieldSingleSelectValue{name}}}}}`, + { id: itemId }, + ); + const node = response?.data?.node; + if (node == null) { + return { exists: false, value: null }; + } + return { exists: true, value: node.fieldValueByName?.name ?? null }; +} + +/** Edit one single-select field on a card, throwing on a non-zero exit. */ +export function editItemField(spawn, project, itemId, fieldId, optionId) { + const edit = gh(spawn, [ + "project", + "item-edit", + "--project-id", + project, + "--id", + itemId, + "--field-id", + fieldId, + "--single-select-option-id", + optionId, + "--format", + "json", + ]); + if (edit.status !== 0) { + throw new Error(`item-edit failed: ${(edit.stderr ?? "").trim()}`); + } +} diff --git a/scripts/lib/board.test.mjs b/scripts/lib/board.test.mjs new file mode 100644 index 0000000000..d732f42fa7 --- /dev/null +++ b/scripts/lib/board.test.mjs @@ -0,0 +1,117 @@ +// Tests for scripts/lib/board.mjs (#2558) — the shared board plumbing: id +// resolution by name, the issue-side card lookup, and the complete-listing +// guard (#2451's silent-truncation class). Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + boardFields, + cardQuery, + editItemField, + findCard, + itemListComplete, + projectId, + requireSupportedBoard, +} from "./board.mjs"; + +const ok = (payload) => ({ + status: 0, + stdout: JSON.stringify(payload), + stderr: "", +}); + +test("requireSupportedBoard accepts only the two live boards", () => { + assert.equal(requireSupportedBoard(28), 28); + assert.equal(requireSupportedBoard(11), 11); + // Project numbers are per-owner and the org holds other real projects — + // a typo must not mutate (or read as empty) an unrelated board. + assert.throws( + () => requireSupportedBoard(12), + /must be 28 \(v2\) or 11 \(v1\).*board #12/, + ); + assert.throws( + () => requireSupportedBoard(12, "the snapshot's board"), + /the snapshot's board must be/, + ); +}); + +test("projectId resolves by board number and throws on a missing id", () => { + assert.equal( + projectId((cmd, args) => { + assert.equal(cmd, "gh"); + assert.ok(args.includes("view") && args.includes("11")); + return ok({ id: "PVT_v1" }); + }, 11), + "PVT_v1", + ); + assert.throws(() => projectId(() => ok({}), 28), /could not resolve/); +}); + +test("boardFields returns the fields list, defaulting to empty", () => { + assert.deepEqual( + boardFields(() => ok({ fields: [{ name: "Status" }] }), 28), + [{ name: "Status" }], + ); + assert.deepEqual( + boardFields(() => ok({}), 28), + [], + ); +}); + +test("cardQuery parameterizes the field read back", () => { + assert.match(cardQuery("Priority"), /fieldValueByName\(name:"Priority"\)/); +}); + +test("findCard selects the card by project node id", () => { + const spawn = () => + ok({ + data: { + repository: { + issue: { + projectItems: { + nodes: [ + { id: "PVTI_a", project: { id: "PVT_other" } }, + { + id: "PVTI_b", + project: { id: "PVT_ours" }, + fieldValueByName: { name: "Todo" }, + }, + ], + }, + }, + }, + }, + }); + assert.equal(findCard(spawn, 7, "PVT_ours").id, "PVTI_b"); + assert.equal(findCard(spawn, 7, "PVT_absent"), undefined); +}); + +test("itemListComplete returns a complete dump and throws on truncation", () => { + const items = [{ id: "a" }, { id: "b" }]; + assert.deepEqual( + itemListComplete(() => ok({ items, totalCount: 2 }), 28).items, + items, + ); + // Truncated: more items exist than the dump holds. + assert.throws( + () => itemListComplete(() => ok({ items, totalCount: 500 }), 28), + /INCOMPLETE.*2 of 500/, + ); + // Malformed: a failed call's output has neither key. + assert.throws(() => itemListComplete(() => ok({}), 28), /INCOMPLETE/); +}); + +test("editItemField throws on a non-zero exit with the stderr", () => { + assert.throws( + () => + editItemField( + () => ({ status: 1, stdout: "", stderr: "nope" }), + "PVT_x", + "PVTI_x", + "F_x", + "opt_x", + ), + /item-edit failed: nope/, + ); + editItemField(() => ok({}), "PVT_x", "PVTI_x", "F_x", "opt_x"); +}); diff --git a/scripts/lib/gh.mjs b/scripts/lib/gh.mjs new file mode 100644 index 0000000000..84752059af --- /dev/null +++ b/scripts/lib/gh.mjs @@ -0,0 +1,107 @@ +// Shared `gh` invocation helpers for the maintainer-workflow scripts (#2558) +// — the pr:*, board:*, advisory:* and action:* helpers. Those scripts replace +// command blocks that the skills previously transcribed inline, so the +// failure modes the skills could only warn about in prose are handled here +// once, under test. +// +// Every function takes its spawn function as a parameter (callers default it +// to `spawnSync`), the same injectability pattern `dependabot-alerts.mjs` and +// `dependency-refresh.mjs` set, so orchestration is testable without starting +// a `gh` process. Auth is `gh`'s own — nothing here sees a token. + +export const OWNER = "modelcontextprotocol"; +export const REPO = "inspector"; +export const REPO_SLUG = `${OWNER}/${REPO}`; + +/** + * The Copilot code-review bot's exact REST login. Matched exactly, never by + * prefix — on a public PR a user whose login merely starts with the bot's + * name could otherwise pass as "the Copilot review" (satisfying a wait's + * expected count, or being printed as the round to act on). + */ +export const COPILOT_REVIEWER_LOGIN = "copilot-pull-request-reviewer[bot]"; + +/** + * `spawnSync`'s default `maxBuffer` is 1 MiB, which a whole-board or + * all-issue JSON listing can exceed — board #28 already passes 500 items and + * each item carries its content — killing the call with ENOBUFS before + * `itemListComplete` ever gets to validate the dump. 64 MiB keeps the bound + * explicit while leaving complete dumps ample headroom. + */ +export const GH_MAX_BUFFER = 64 * 1024 * 1024; + +/** + * Run `gh` with the given args. Throws only on spawn failure (gh not + * installed); a non-zero exit is the caller's to interpret via the result. + */ +export function gh(spawn, args) { + const result = spawn("gh", args, { + encoding: "utf8", + maxBuffer: GH_MAX_BUFFER, + }); + if (result.error) { + throw result.error; + } + return result; +} + +/** + * Run `gh` and parse its stdout as JSON, throwing on a non-zero exit with the + * stderr in the message. An API/auth failure must throw rather than read as an + * empty result — the pr-flow skill's wait loop documents why: an error + * swallowed into a zero makes a background wait retry blind forever. + */ +export function ghJson(spawn, args) { + const result = gh(spawn, args); + if (result.status !== 0) { + throw new Error( + `gh ${args.join(" ")} failed (${result.status}): ${(result.stderr ?? "").trim()}`, + ); + } + return JSON.parse(result.stdout); +} + +/** + * Fetch a paginated REST list endpoint completely. + * + * ⚠️ `--slurp` is load-bearing (same note as `dependabot-alerts.mjs`): without + * it `gh api --paginate` concatenates one JSON array per page into invalid + * JSON. With it the output is an array of pages, flattened here. + */ +export function ghPaginatedList(spawn, path) { + const pages = ghJson(spawn, ["api", "--paginate", "--slurp", path]); + if (!Array.isArray(pages) || pages.some((page) => !Array.isArray(page))) { + throw new Error(`unexpected non-list response from ${path}`); + } + return pages.flat(); +} + +/** + * Run a GraphQL query/mutation. `fields` values are passed with `-F` for + * numbers (typed as Int) and `-f` for strings, matching how the skills' inline + * blocks passed them. + */ +export function ghGraphql(spawn, query, fields = {}) { + const args = ["api", "graphql"]; + for (const [name, value] of Object.entries(fields)) { + args.push( + typeof value === "number" ? "-F" : "-f", + `${name}=${String(value)}`, + ); + } + args.push("-f", `query=${query}`); + return ghJson(spawn, args); +} + +/** + * Parse a `--flag`-style value as a positive integer, throwing a usage-shaped + * error naming the flag — shared by every script's argv validation. + */ +export function requirePositiveInt(value, flag) { + if (value === undefined || !/^[1-9][0-9]*$/.test(value)) { + throw new Error( + `${flag} must be a positive integer (got ${value ?? "nothing"})`, + ); + } + return Number(value); +} diff --git a/scripts/lib/gh.test.mjs b/scripts/lib/gh.test.mjs new file mode 100644 index 0000000000..5aead316c1 --- /dev/null +++ b/scripts/lib/gh.test.mjs @@ -0,0 +1,101 @@ +// Tests for scripts/lib/gh.mjs (#2558) — the shared `gh` invocation helpers +// under the maintainer-workflow scripts. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + gh, + ghJson, + ghGraphql, + ghPaginatedList, + requirePositiveInt, + GH_MAX_BUFFER, +} from "./gh.mjs"; + +function spawnReturning(result) { + const calls = []; + const spawn = (cmd, args, opts) => { + calls.push({ cmd, args, opts }); + return result; + }; + spawn.calls = calls; + return spawn; +} + +test("gh passes args through and returns the raw result", () => { + const spawn = spawnReturning({ status: 0, stdout: "x", stderr: "" }); + const result = gh(spawn, ["api", "whatever"]); + assert.equal(result.stdout, "x"); + assert.deepEqual(spawn.calls[0].args, ["api", "whatever"]); + assert.equal(spawn.calls[0].cmd, "gh"); +}); + +test("gh raises maxBuffer above spawnSync's 1 MiB default", () => { + // A whole-board or all-issue listing can exceed 1 MiB, and ENOBUFS would + // kill the call before itemListComplete could validate the dump. + const spawn = spawnReturning({ status: 0, stdout: "x", stderr: "" }); + gh(spawn, ["api", "whatever"]); + assert.equal(spawn.calls[0].opts.maxBuffer, GH_MAX_BUFFER); + assert.ok(GH_MAX_BUFFER > 1024 * 1024); +}); + +test("gh throws on a spawn-level error (gh not installed)", () => { + const spawn = spawnReturning({ error: new Error("ENOENT") }); + assert.throws(() => gh(spawn, ["api"]), /ENOENT/); +}); + +test("ghJson parses stdout and throws on non-zero exit with stderr", () => { + const ok = spawnReturning({ status: 0, stdout: '{"a":1}', stderr: "" }); + assert.deepEqual(ghJson(ok, ["api", "x"]), { a: 1 }); + + const bad = spawnReturning({ status: 1, stdout: "", stderr: "auth broke" }); + assert.throws(() => ghJson(bad, ["api", "x"]), /auth broke/); +}); + +test("ghPaginatedList slurps and flattens pages", () => { + const spawn = spawnReturning({ + status: 0, + stdout: "[[1,2],[3]]", + stderr: "", + }); + assert.deepEqual(ghPaginatedList(spawn, "repos/o/r/things"), [1, 2, 3]); + assert.deepEqual(spawn.calls[0].args, [ + "api", + "--paginate", + "--slurp", + "repos/o/r/things", + ]); +}); + +test("ghPaginatedList rejects a non-list response", () => { + const notList = spawnReturning({ status: 0, stdout: '{"x":1}', stderr: "" }); + assert.throws(() => ghPaginatedList(notList, "p"), /non-list/); + const mixedPages = spawnReturning({ + status: 0, + stdout: "[[1],2]", + stderr: "", + }); + assert.throws(() => ghPaginatedList(mixedPages, "p"), /non-list/); +}); + +test("ghGraphql passes numbers with -F and strings with -f", () => { + const spawn = spawnReturning({ status: 0, stdout: "{}", stderr: "" }); + ghGraphql(spawn, "query($n:Int!){x}", { n: 7, s: "abc" }); + assert.deepEqual(spawn.calls[0].args, [ + "api", + "graphql", + "-F", + "n=7", + "-f", + "s=abc", + "-f", + "query=query($n:Int!){x}", + ]); +}); + +test("requirePositiveInt accepts positives and rejects everything else", () => { + assert.equal(requirePositiveInt("42", "--pr"), 42); + for (const bad of [undefined, "0", "-1", "1.5", "abc", "", "07"]) { + assert.throws(() => requirePositiveInt(bad, "--pr"), /--pr/); + } +}); diff --git a/scripts/lib/markdown-cell.mjs b/scripts/lib/markdown-cell.mjs new file mode 100644 index 0000000000..c97d2d0a2d --- /dev/null +++ b/scripts/lib/markdown-cell.mjs @@ -0,0 +1,19 @@ +/** + * Escape a value for a Markdown table cell (#2546). + * + * The issue-filing sweeps (`dependabot-alerts.mjs`, `sdk-watch.mjs`) build + * table rows from text they do not control: advisory summaries and upstream + * release data. A `|` in that text ends the cell early, so it is escaped as + * `\|`. That alone is incomplete: a value that itself ends in a backslash, + * such as `abc\`, would become `abc\\|`, where the backslash pair cancels and + * the pipe is live again, breaking the row. So backslashes are escaped + * **first**, then pipes (CodeQL js/incomplete-sanitization, alerts #74–#76). + * + * One helper for both scripts, so the order cannot drift between copies. + * + * @param {unknown} value + * @returns {string} + */ +export function escapeTableCell(value) { + return String(value).replace(/\\/g, "\\\\").replace(/\|/g, "\\|"); +} diff --git a/scripts/lib/markdown-cell.test.mjs b/scripts/lib/markdown-cell.test.mjs new file mode 100644 index 0000000000..c9b760dc99 --- /dev/null +++ b/scripts/lib/markdown-cell.test.mjs @@ -0,0 +1,31 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { escapeTableCell } from "./markdown-cell.mjs"; + +test("leaves plain text alone", () => { + assert.equal(escapeTableCell("plain text"), "plain text"); +}); + +test("escapes a pipe so it cannot end the cell", () => { + assert.equal(escapeTableCell("a|b"), "a\\|b"); +}); + +test("escapes a trailing backslash before a following pipe can pair with it", () => { + // `abc\` + an escaped pipe must not read as `abc\\|` (a live pipe). + assert.equal(escapeTableCell("abc\\"), "abc\\\\"); + assert.equal(escapeTableCell("abc\\|d"), "abc\\\\\\|d"); +}); + +test("coerces non-strings", () => { + assert.equal(escapeTableCell(42), "42"); + assert.equal(escapeTableCell(null), "null"); +}); + +test("every pipe in the output is escaped by an odd run of backslashes", () => { + for (const input of ["a|b", "a\\|b", "a\\\\|b", "\\", "|", "x\\"]) { + const out = escapeTableCell(input); + for (const m of out.matchAll(/(\\*)\|/g)) { + assert.equal(m[1].length % 2, 1, `${JSON.stringify(input)} → ${out}`); + } + } +}); diff --git a/scripts/lib/mcpdo-eval-matchers.mjs b/scripts/lib/mcpdo-eval-matchers.mjs new file mode 100644 index 0000000000..8558d22842 --- /dev/null +++ b/scripts/lib/mcpdo-eval-matchers.mjs @@ -0,0 +1,853 @@ +// Structured matchers for mcpdo behavior evals (`skills:eval:mcpdo`). +// +// A behavior case asserts WHAT an agent did with mcpdo, from the transcript +// records the eval shim (`mcpdo-eval-shim.mjs`) writes: one JSON line per +// invocation, `{ argv, exit, start, end, events }`. Everything here works on +// parsed argv ARRAYS and captured output — never on re-parsed shell strings, +// which would re-implement the shell's quoting rules badly and drift from +// what actually ran. +// +// Matcher semantics are deliberately loose where the CLI is generous. mcpdo +// accepts the same tool call as `tools/call get_sum a:=2 b:=3`, +// `tools/call --tool-name get_sum --tool-arg a=2 b=3`, or +// `tools/call get_sum '{"a":2,"b":3}'`, with the connection as `@name`, as +// `--connection name`, or implicit via MRU. A case should pass for every +// correct spelling and fail for a wrong tool, wrong argument, or wrong +// server — so parsing normalizes all spellings into one shape before +// matching, and values compare canonically (`"2"` matches `2`). + +/** + * Global and per-command flags that consume exactly one following token. + * + * Deliberately a fixed list rather than a "flags eat the next token" + * heuristic: boolean flags like `--task` would otherwise swallow the tool + * name. An unknown value-flag degrades softly — its value shows up as a + * stray positional, which no matcher field reads. + */ +const VALUE_FLAGS = new Set([ + "--format", + "--connection", + "--conn", + "--catalog", + "--config", + "--tool-name", + "--tool-args-json", + "--uri", + "--transport", + "--cwd", + "--connect-timeout", + "--era", + "--elicit", + "-e", +]); + +/** + * Variadic flags (`<pairs...>`): consume following `key=value` tokens until + * one stops looking like a pair. + */ +const VARIADIC_FLAGS = new Set(["--metadata", "--tool-arg", "--tool-metadata"]); + +/** JSON.parse when the text parses, the raw string otherwise. */ +function looseParse(text) { + try { + return JSON.parse(text); + } catch { + return text; + } +} + +/** + * Parse one mcpdo invocation's argv (everything after `mcpdo`) into the + * fields matchers assert on. + * + * @param {string[]} argv + * @returns {{ cmd: string | null, connection: string | null, tool: string | + * null, args: Record<string, unknown>, positionals: string[] }} + */ +export function parseMcpdoArgv(argv) { + let connection = null; + let toolNameFlag = null; + const args = {}; + const positionals = []; + + for (let i = 0; i < argv.length; i++) { + let token = argv[i]; + // `--flag=value` form: split so the flag matches the sets below the same + // as the space-separated form; otherwise `--connection=x` would parse as + // an unknown boolean flag and silently drop the value. + let inline = null; + if (token.startsWith("--")) { + const eq = token.indexOf("="); + if (eq !== -1) { + inline = token.slice(eq + 1); + token = token.slice(0, eq); + } + } + if (VARIADIC_FLAGS.has(token)) { + const pairs = inline !== null ? [inline] : []; + while ( + i + 1 < argv.length && + !argv[i + 1].startsWith("-") && + argv[i + 1].includes("=") + ) { + pairs.push(argv[++i]); + } + if (token === "--tool-arg") { + for (const pair of pairs) { + const eq = pair.indexOf("="); + args[pair.slice(0, eq)] = looseParse(pair.slice(eq + 1)); + } + } + continue; + } + if (VALUE_FLAGS.has(token)) { + const value = inline ?? argv[++i]; + if (token === "--connection" || token === "--conn") connection = value; + if (token === "--tool-name") toolNameFlag = value; + if (token === "--tool-args-json") { + const parsed = looseParse(value); + if (parsed !== null && typeof parsed === "object") { + Object.assign(args, parsed); + } + } + continue; + } + if (token.startsWith("--")) continue; // boolean flag + if (token.startsWith("@") && token.length > 1) { + // Positional `@name` selects the connection wherever it appears. + if (connection === null) connection = token.slice(1); + continue; + } + positionals.push(token); + } + + let cmd = positionals[0] ?? null; + let rest = positionals.slice(1); + if (cmd === "daemon" && rest.length > 0) { + // `daemon stop` etc. — the subcommand is part of what a case asserts. + cmd = `daemon ${rest[0]}`; + rest = rest.slice(1); + } + + let tool = toolNameFlag; + for (const token of rest) { + if (token.includes(":=")) { + const sep = token.indexOf(":="); + args[token.slice(0, sep)] = looseParse(token.slice(sep + 2)); + continue; + } + if (token.startsWith("{")) { + const parsed = looseParse(token); + if (parsed !== null && typeof parsed === "object") { + Object.assign(args, parsed); + } + continue; + } + // First bare positional after the command: the tool (or prompt, or + // task id — the field is generic on purpose; `tool` is just its most + // common reading). + if (tool === null) tool = token; + } + + return { cmd, connection, tool, args, positionals }; +} + +/** Recursively: parse JSON-looking strings, sort object keys. */ +function normalize(value) { + const v = typeof value === "string" ? looseParse(value) : value; + if (v !== null && typeof v === "object") { + if (Array.isArray(v)) return v.map(normalize); + return Object.fromEntries( + Object.keys(v) + .sort() + .map((k) => [k, normalize(v[k])]), + ); + } + return v; +} + +/** Stable stringify: objects with sorted keys, so shape compares by value. */ +function canonical(value) { + return JSON.stringify(normalize(value)); +} + +/** + * Loose value equality: `2`, `"2"`, and a JSON string `"2"` all match, and + * objects compare deeply with key order ignored. Argument values arrive as + * strings from `--tool-arg a=2` and as numbers from `a:=2`; a case should + * not care which spelling the agent picked. + */ +export function valuesMatch(expected, actual) { + return canonical(expected) === canonical(actual); +} + +/** Concatenated output for one stream of a transcript record. */ +export function streamText(record, stream) { + return (record.events ?? []) + .filter((e) => e.stream === stream) + .map((e) => e.data) + .join(""); +} + +/** Read `a.b.c` out of a parsed JSON value. */ +function readPath(value, dotted) { + let cur = value; + for (const key of dotted.split(".")) { + if (cur === null || typeof cur !== "object") return undefined; + cur = cur[key]; + } + return cur; +} + +/** + * Per-stream views of a transcript with a mapping back to GLOBAL event + * order. + * + * Two problems solved at once. A pipe does not preserve write boundaries, so + * a phase's pattern may span two recorded chunks — matching must run over + * each stream's concatenated text, not per event. But interactive ordering + * ("the stdin answer came after the stdout prompt") is BETWEEN streams, so + * every character also needs a position on the one shared timeline; segments + * carry that mapping. + * + * @param {object} record One shim transcript record. + * @returns {Map<string, { text: string, segments: { streamStart: number, + * globalStart: number, len: number }[] }>} + */ +export function buildTimeline(record) { + const streams = new Map(); + let global = 0; + for (const event of record.events ?? []) { + const data = String(event.data ?? ""); + let entry = streams.get(event.stream); + if (!entry) { + entry = { text: "", segments: [] }; + streams.set(event.stream, entry); + } + entry.segments.push({ + streamStart: entry.text.length, + globalStart: global, + len: data.length, + }); + entry.text += data; + global += data.length; + } + return streams; +} + +/** Global timeline position of a stream-local offset. */ +function globalPos(entry, streamOffset) { + for (const seg of entry.segments) { + if (streamOffset < seg.streamStart + seg.len) { + return seg.globalStart + (streamOffset - seg.streamStart); + } + } + return Number.MAX_SAFE_INTEGER; +} + +/** + * Match ordered phases against one invocation's interleaved transcript. + * + * Each phase is `{ stream, match }`: a regex that must appear on that stream + * strictly AFTER (on the global timeline) where the previous phase matched. + * This is what turns the shim's event capture into assertions like "stdout + * showed the auth URL before exit" or "stdin answered only after the prompt + * appeared". + * + * @param {{ stream: string, match: string }[]} phases + * @param {object} record One shim transcript record. + * @returns {string | null} `null` on match, else the first failure reason. + */ +export function matchPhases(phases, record) { + const streams = buildTimeline(record); + let cursor = -1; + for (const [i, phase] of phases.entries()) { + const entry = streams.get(phase.stream); + if (!entry) { + return `phase ${i} /${phase.match}/: no ${phase.stream} data recorded`; + } + const re = new RegExp(phase.match, "g"); + let found = -1; + for (const m of entry.text.matchAll(re)) { + const end = globalPos(entry, m.index + Math.max(m[0].length, 1) - 1); + if (end > cursor) { + found = end; + break; + } + } + if (found === -1) { + const anywhere = new RegExp(phase.match).test(entry.text); + return anywhere + ? `phase ${i} /${phase.match}/ matched ${phase.stream} only BEFORE phase ${i - 1}` + : `phase ${i} /${phase.match}/ not found on ${phase.stream}`; + } + cursor = found; + } + return null; +} + +/** + * Match one transcript record against one matcher. + * + * @param {object} matcher See `validateBehaviorCase` for the shape. + * @param {object} record One shim transcript record. + * @returns {string | null} `null` on match, else the first mismatch reason. + */ +export function matchCall(matcher, record) { + const parsed = parseMcpdoArgv(record.argv ?? []); + if (parsed.cmd !== matcher.cmd) { + return `cmd is \`${parsed.cmd}\`, expected \`${matcher.cmd}\``; + } + if (matcher.connection !== undefined) { + // `connect` names its connection positionally (`connect test-stdio`), + // which the generic parse files under `tool` — for this command the + // target IS the connection being established, so match either spelling. + const conn = + parsed.connection ?? (matcher.cmd === "connect" ? parsed.tool : null); + // An absent connection is accepted: the hermetic env holds exactly one + // entry, so the implicit MRU can only be the right one. An EXPLICIT + // wrong connection is the bug this field exists to catch. + if (conn !== null && conn !== matcher.connection) { + return `connection is \`${conn}\`, expected \`${matcher.connection}\``; + } + } + if (matcher.tool !== undefined && parsed.tool !== matcher.tool) { + return `tool is \`${parsed.tool}\`, expected \`${matcher.tool}\``; + } + for (const [key, expected] of Object.entries(matcher.args ?? {})) { + if (!(key in parsed.args)) { + return `arg \`${key}\` missing (args: ${JSON.stringify(parsed.args)})`; + } + if (!valuesMatch(expected, parsed.args[key])) { + return `arg \`${key}\` is ${JSON.stringify(parsed.args[key])}, expected ${JSON.stringify(expected)}`; + } + } + // A failing invocation must not satisfy a matcher unless the case says so: + // `connect` that exited 1 did not connect. + const wantExit = matcher.exit ?? 0; + if (record.exit !== wantExit) { + return `exit is ${record.exit}, expected ${wantExit}`; + } + const stdout = streamText(record, "stdout"); + if (matcher.result !== undefined) { + let parsedOut; + try { + parsedOut = JSON.parse(stdout); + } catch { + // Distinct diagnostic on purpose: the call may have been right while + // the case asserted JSON against human text output. + return `\`result\` asserted but stdout is not JSON (use stdoutMatch for text output)`; + } + for (const [path, expected] of Object.entries(matcher.result)) { + const actual = readPath(parsedOut, path); + if (!valuesMatch(expected, actual)) { + return `result path \`${path}\` is ${JSON.stringify(actual)}, expected ${JSON.stringify(expected)}`; + } + } + } + if (matcher.stdoutMatch !== undefined) { + if (!new RegExp(matcher.stdoutMatch).test(stdout)) { + return `stdout does not match /${matcher.stdoutMatch}/`; + } + } + if (matcher.phases !== undefined) { + const reason = matchPhases(matcher.phases, record); + if (reason !== null) return reason; + } + return null; +} + +/** + * Check the transcript contains the expected calls as an ordered + * subsequence — other calls in between are fine (exploring with + * `tools/list` first is correct behavior, not noise). + * + * @param {object[]} expectCalls + * @param {object[]} records + * @returns {{ ok: boolean, failures: string[] }} + */ +export function evalExpectCalls(expectCalls, records) { + const failures = []; + let cursor = 0; + for (const [i, matcher] of expectCalls.entries()) { + let matched = -1; + const nearMisses = []; + for (let j = cursor; j < records.length; j++) { + const reason = matchCall(matcher, records[j]); + if (reason === null) { + matched = j; + break; + } + // Same command, wrong details: that is the interesting diagnostic. + if (parseMcpdoArgv(records[j].argv ?? []).cmd === matcher.cmd) { + nearMisses.push(reason); + } + } + if (matched === -1) { + const detail = + nearMisses.length > 0 + ? ` (near miss: ${nearMisses[0]})` + : records.length === 0 + ? " (no mcpdo invocations recorded)" + : ""; + failures.push( + `expectCalls[${i}] \`${matcher.cmd}\` not satisfied${detail}`, + ); + } else { + cursor = matched + 1; + } + } + return { ok: failures.length === 0, failures }; +} + +const MATCHER_KEYS = new Set([ + "cmd", + "connection", + "tool", + "args", + "exit", + "result", + "stdoutMatch", + "phases", +]); + +const PHASE_STREAMS = new Set(["stdin", "stdout", "stderr"]); + +/** + * Validate one behavior case. Local to the mcpdo eval on purpose — the + * shared `skill-manifest.mjs` schema describes trigger cases for every + * skill, while `expectCalls` is this harness's private contract. + * + * Unknown matcher keys are errors, not ignored: a typoed `tooll` would + * otherwise silently assert nothing and report a hit. + * + * @param {object} c + * @param {number} i Case index, for error messages. + * @returns {string[]} Errors; empty when valid. + */ +export function validateBehaviorCase(c, i) { + const errors = []; + if (typeof c.prompt !== "string" || c.prompt.trim() === "") { + errors.push(`behavior case ${i}: \`prompt\` must be a non-empty string`); + } + if (!Array.isArray(c.expectCalls) || c.expectCalls.length === 0) { + errors.push( + `behavior case ${i}: \`expectCalls\` must be a non-empty array`, + ); + return errors; + } + c.expectCalls.forEach((m, j) => { + const at = `behavior case ${i} expectCalls[${j}]`; + if (m === null || typeof m !== "object" || Array.isArray(m)) { + errors.push(`${at}: must be an object`); + return; + } + if (typeof m.cmd !== "string" || m.cmd.trim() === "") { + errors.push(`${at}: \`cmd\` is required`); + } + for (const key of Object.keys(m)) { + if (!MATCHER_KEYS.has(key)) { + errors.push(`${at}: unknown key \`${key}\``); + } + } + for (const key of ["connection", "tool", "stdoutMatch"]) { + if (m[key] !== undefined && typeof m[key] !== "string") { + errors.push(`${at}: \`${key}\` must be a string`); + } + } + for (const key of ["args", "result"]) { + if ( + m[key] !== undefined && + (m[key] === null || typeof m[key] !== "object" || Array.isArray(m[key])) + ) { + errors.push(`${at}: \`${key}\` must be an object`); + } + } + if (m.exit !== undefined && !Number.isInteger(m.exit)) { + errors.push(`${at}: \`exit\` must be an integer`); + } + if (m.stdoutMatch !== undefined && typeof m.stdoutMatch === "string") { + try { + new RegExp(m.stdoutMatch); + } catch { + errors.push(`${at}: \`stdoutMatch\` is not a valid regex`); + } + } + if (m.phases !== undefined) { + if (!Array.isArray(m.phases) || m.phases.length === 0) { + errors.push(`${at}: \`phases\` must be a non-empty array`); + } else { + m.phases.forEach((p, k) => { + if (p === null || typeof p !== "object" || Array.isArray(p)) { + errors.push(`${at}: phases[${k}] must be an object`); + return; + } + for (const key of Object.keys(p)) { + if (key !== "stream" && key !== "match") { + errors.push(`${at}: phases[${k}] unknown key \`${key}\``); + } + } + if (!PHASE_STREAMS.has(p.stream)) { + errors.push( + `${at}: phases[${k}].stream must be one of ${[...PHASE_STREAMS].join(", ")}`, + ); + } + if (typeof p.match !== "string") { + errors.push(`${at}: phases[${k}].match must be a string`); + } else { + try { + new RegExp(p.match); + } catch { + errors.push(`${at}: phases[${k}].match is not a valid regex`); + } + } + }); + } + } + }); + if (c.autoConsent !== undefined && typeof c.autoConsent !== "boolean") { + errors.push(`behavior case ${i}: \`autoConsent\` must be a boolean`); + } + if ( + c.consentAfterReply !== undefined && + typeof c.consentAfterReply !== "boolean" + ) { + errors.push(`behavior case ${i}: \`consentAfterReply\` must be a boolean`); + } + if ( + c.consentDelayMs !== undefined && + (!Number.isInteger(c.consentDelayMs) || c.consentDelayMs < 0) + ) { + errors.push( + `behavior case ${i}: \`consentDelayMs\` must be a non-negative integer`, + ); + } + for (const key of ["expectReply", "expectReplyFollowUp"]) { + if (c[key] === undefined) continue; + if (typeof c[key] !== "string" || c[key].trim() === "") { + errors.push(`behavior case ${i}: \`${key}\` must be a non-empty string`); + } else { + try { + new RegExp(c[key]); + } catch { + errors.push(`behavior case ${i}: \`${key}\` is not a valid regex`); + } + } + } + // `followUp` makes a case MULTI-TURN: turn 1 runs the prompt, the harness + // completes the simulated sign-in, then turn 2 resumes the same agent + // session with this text. Only meaningful alongside `autoConsent` (nothing + // completes the sign-in otherwise) and currently claude-only (session + // resume), but neither is enforced here — the harness skips a multi-turn + // case for a non-claude agent, and a missing clicker just fails the case. + if (c.followUp !== undefined) { + if (typeof c.followUp !== "string" || c.followUp.trim() === "") { + errors.push( + `behavior case ${i}: \`followUp\` must be a non-empty string`, + ); + } + if (c.expectReplyFollowUp === undefined) { + errors.push( + `behavior case ${i}: \`followUp\` requires \`expectReplyFollowUp\` ` + + `(a multi-turn case must assert what the resumed turn told the user)`, + ); + } + } else if (c.expectReplyFollowUp !== undefined) { + errors.push( + `behavior case ${i}: \`expectReplyFollowUp\` requires \`followUp\``, + ); + } + // `expectLastCallTurn1` asserts the agent's LAST mcpdo command before the + // follow-up turn (all of a single-turn case) was this one — the gate that + // proves a pure-connect case STOPPED after relaying the sign-in URL instead + // of polling. For a case with a task command to run, use + // `rejectCompletedTurn1` instead (last-command is brittle under chaining). + if (c.expectLastCallTurn1 !== undefined) { + if ( + typeof c.expectLastCallTurn1 !== "string" || + c.expectLastCallTurn1.trim() === "" + ) { + errors.push( + `behavior case ${i}: \`expectLastCallTurn1\` must be a non-empty string`, + ); + } + } + // `rejectCompletedTurn1` names the task command (e.g. `tools/list`) whose + // SUCCESS (exit 0) before the follow-up means the agent completed the task + // in turn 1 instead of stopping to let the user sign in — a poll or a + // genuine completion. See `succeededBefore` for why this replaced the + // brittle "last command was connect" check for the list-tools case. + if (c.rejectCompletedTurn1 !== undefined) { + if ( + typeof c.rejectCompletedTurn1 !== "string" || + c.rejectCompletedTurn1.trim() === "" + ) { + errors.push( + `behavior case ${i}: \`rejectCompletedTurn1\` must be a non-empty string`, + ); + } + } + errors.push(...validateCaseServers(c, i)); + return errors; +} + +/** + * Claude's session id, read from its `-p --output-format stream-json` NDJSON + * (`rawLogPath`). Claude stamps the same `session_id` on its init system + * event and its final result event; `--resume <id>` continues that session, + * which is what makes the OAuth behavior cases multi-turn. Returns the last + * one seen (robust to a torn final line), or null when absent — a non-claude + * stream, or one never captured. + * + * @param {string} raw NDJSON text as captured via `rawLogPath`. + * @returns {string | null} + */ +export function claudeSessionId(raw) { + let id = null; + for (const line of raw.split("\n")) { + if (!line.includes("session_id")) continue; + let evt; + try { + evt = JSON.parse(line); + } catch { + continue; + } + if (typeof evt?.session_id === "string" && evt.session_id !== "") { + id = evt.session_id; + } + } + return id; +} + +/** + * The `cmd` of the last mcpdo invocation that STARTED before `boundaryTs` + * (epoch-ms), by start time — not array position, because a `connect` that + * relays a URL exits fast while an earlier failed `tools/list` may still be + * recorded around it. The multi-turn OAuth cases use this with the + * follow-up's send time as the boundary to assert the agent's final first-turn + * move was `connect` (it stopped to relay the URL) rather than a poll/retry. + * A single-turn case passes `Infinity` to cover every record. Returns null + * when nothing qualifies. + * + * @param {Array<{argv?: string[], start?: number}>} records + * @param {number} boundaryTs + * @returns {string | null} + */ +export function lastCommandBefore(records, boundaryTs) { + let last = null; + let lastStart = -Infinity; + for (const r of records) { + if (typeof r.start !== "number" || r.start >= boundaryTs) continue; + if (r.start >= lastStart) { + lastStart = r.start; + last = parseMcpdoArgv(r.argv ?? []).cmd; + } + } + return last; +} + +/** + * Did any mcpdo invocation of `cmd` SUCCEED (exit 0) before `boundaryTs`? + * + * This is the behavioral completion signal the multi-turn OAuth cases gate on, + * and it is deliberately NOT `lastCommandBefore`. An agent that relays the + * sign-in URL and then waits still frequently runs the task command + * optimistically in the same turn — e.g. `connect && tools/list`, chained in a + * single shell line — because `connect` exits 0 even when sign-in is pending. + * That chained `tools/list` runs against a not-yet-connected server, exits + * NON-zero (`auth_required`), and the agent ignores it and waits: correct + * behavior, but `lastCommandBefore` reads the trailing `tools/list` and wrongly + * flags it. What actually distinguishes "waited" from "barged through" is + * whether the task command SUCCEEDED in turn 1 — a poll-until-connected or a + * genuine completion exits 0; an optimistic chained attempt does not. + * + * @param {Array<{argv?: string[], start?: number, exit?: number}>} records + * @param {number} boundaryTs + * @param {string} cmd + * @returns {boolean} + */ +export function succeededBefore(records, boundaryTs, cmd) { + for (const r of records) { + if (typeof r.start !== "number" || r.start >= boundaryTs) continue; + if (r.exit !== 0) continue; + if (parseMcpdoArgv(r.argv ?? []).cmd === cmd) return true; + } + return false; +} + +/** + * The text an agent actually showed its user, from the raw agent-session + * NDJSON stream (`agent-session.ndjson`). + * + * This is deliberately a different surface from the shim transcript: a case's + * `expectCalls` assert what the agent RAN, while `expectReply` asserts what it + * SAID — the two can diverge exactly when the agent reads something in a + * command's output (a sign-in URL) and fails to relay it to the user, which + * no transcript matcher can see. + * + * Handles both agents' event shapes: Claude's `-p --output-format stream-json` + * (`{type:"assistant", message:{content:[{type:"text", text}]}}`) and the + * Copilot CLI's `--output-format json` (`{type:"assistant.message", + * data:{content: "<text>"}}`). Malformed lines are CLI noise, not + * observations, same as the collectors in skill-eval.mjs. + * + * What counts as "shown to the user" is host-specific, so it is a parameter + * rather than a rule: Claude Code renders `thinking` blocks whose signature + * is narration as ordinary foreground text — a sign-in URL relayed there + * reached the user (observed live) — while the Copilot CLI shows its + * reasoning in a dimmed font users routinely skip. The per-agent defaults + * live in {@link REPLY_TEXT_OPTIONS}; pass `options` to override in a test + * or when tuning what a given host actually surfaces. + * + * @param {string} raw NDJSON text as captured via `rawLogPath`. + * @param {object} [options] + * @param {boolean} [options.includeThinking] Also count Claude `thinking` + * blocks as user-visible text. + * @returns {string} Every user-visible assistant text block, joined with + * newlines. + */ +export function assistantReplyText(raw, options = {}) { + const { includeThinking = false } = options; + const texts = []; + for (const line of raw.split("\n")) { + if (!line.trim()) continue; + let evt; + try { + evt = JSON.parse(line); + } catch { + continue; + } + if (evt?.type === "assistant") { + for (const block of evt.message?.content ?? []) { + if (block?.type === "text" && typeof block.text === "string") { + texts.push(block.text); + } else if ( + includeThinking && + block?.type === "thinking" && + typeof block.thinking === "string" + ) { + texts.push(block.thinking); + } + } + } else if ( + evt?.type === "assistant.message" && + typeof evt.data?.content === "string" + ) { + texts.push(evt.data.content); + } + } + return texts.join("\n"); +} + +/** + * Per-agent defaults for {@link assistantReplyText}: what each host's UI + * actually puts in front of the user. Claude Code displays narration-style + * thinking as foreground text; the Copilot CLI dims its reasoning, so only + * proper assistant messages count there. + */ +export const REPLY_TEXT_OPTIONS = { + claude: { includeThinking: true }, + copilot: { includeThinking: false }, +}; + +/** + * Validate a behavior case's server declaration: optional `server` (single + * spec under the default catalog name) or `servers` (name→spec map), never + * both. Names become catalog entry names — the agent sees them via + * `servers/list`, so they are scenario content. + * + * @param {object} c A behavior case. + * @param {number} i Case index, for error messages. + * @returns {string[]} + */ +export function validateCaseServers(c, i) { + if (c.server !== undefined && c.servers !== undefined) { + return [ + `behavior case ${i}: \`server\` and \`servers\` are mutually exclusive`, + ]; + } + if (c.servers !== undefined) { + const at = `behavior case ${i} \`servers\``; + if ( + c.servers === null || + typeof c.servers !== "object" || + Array.isArray(c.servers) + ) { + return [`${at}: must be a name→spec object`]; + } + const names = Object.keys(c.servers); + if (names.length === 0) return [`${at}: must not be empty`]; + const errors = []; + for (const name of names) { + if (!/^[A-Za-z0-9_.-]+$/.test(name)) { + errors.push(`${at}: \`${name}\` is not a valid catalog entry name`); + continue; + } + if (c.servers[name] === undefined) { + errors.push( + `${at}.${name}: spec must not be undefined (omit \`servers\` for the default server)`, + ); + continue; + } + errors.push( + ...validateServerSpec(c.servers[name], i, `\`servers\`.${name}`), + ); + } + return errors; + } + return validateServerSpec(c.server, i); +} + +/** + * Validate one server spec: either `{ "url": "<http(s) endpoint>" }` for a + * server the harness does not manage, or the test-servers declarative + * config-file shape (serverInfo + preset refs). A composed spec with no + * transport (or stdio) is served through the eval's stdio launcher; with + * `transport.type: "streamable-http"` it is started in-process + * (`TestServerHttp`) and the catalog entry points at its URL — OAuth via the + * spec's `oauth` block rides on the same instance. Only the discriminating + * structure is checked here — preset names and capability switches are the + * framework's contract, validated by `resolveConfig` when the server starts. + * + * @param {object | undefined} server + * @param {number} i Case index, for error messages. + * @param {string} [label] Field label for error messages. + * @returns {string[]} + */ +export function validateServerSpec(server, i, label = "`server`") { + if (server === undefined) return []; + const at = `behavior case ${i} ${label}`; + if (server === null || typeof server !== "object" || Array.isArray(server)) { + return [`${at}: must be an object`]; + } + if ("url" in server) { + const errors = []; + if (typeof server.url !== "string" || !/^https?:\/\//.test(server.url)) { + errors.push(`${at}: \`url\` must be an http(s) URL`); + } + for (const key of Object.keys(server)) { + if (key !== "url") { + errors.push(`${at}: \`url\` form takes no other keys (got \`${key}\`)`); + } + } + return errors; + } + if ( + typeof server.serverInfo?.name !== "string" || + typeof server.serverInfo?.version !== "string" + ) { + return [ + `${at}: composed form needs \`serverInfo\` with \`name\` and \`version\` (or use the \`url\` form)`, + ]; + } + if ( + server.transport !== undefined && + server.transport?.type !== "stdio" && + server.transport?.type !== "streamable-http" + ) { + return [ + `${at}: composed transport must be "stdio" (default) or "streamable-http" — sse fixtures are not supported by the harness`, + ]; + } + return []; +} diff --git a/scripts/lib/mcpdo-eval-matchers.test.mjs b/scripts/lib/mcpdo-eval-matchers.test.mjs new file mode 100644 index 0000000000..1678ac33b6 --- /dev/null +++ b/scripts/lib/mcpdo-eval-matchers.test.mjs @@ -0,0 +1,728 @@ +// Tests for the behavior-eval matcher library: every spelling mcpdo accepts +// for the same call must normalize to the same parse, and a matcher must +// fail for the reasons a case exists to catch (wrong tool, wrong argument, +// wrong server, nonzero exit) with a reason a human can act on. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + assistantReplyText, + REPLY_TEXT_OPTIONS, + parseMcpdoArgv, + valuesMatch, + matchCall, + matchPhases, + evalExpectCalls, + streamText, + validateBehaviorCase, + validateCaseServers, + validateServerSpec, + claudeSessionId, + lastCommandBefore, + succeededBefore, +} from "./mcpdo-eval-matchers.mjs"; + +const record = (argv, { exit = 0, stdout = "", stderr = "" } = {}) => ({ + argv, + exit, + start: 1, + end: 2, + events: [ + ...(stdout ? [{ t: 1, stream: "stdout", data: stdout }] : []), + ...(stderr ? [{ t: 1, stream: "stderr", data: stderr }] : []), + ], +}); + +test("parseMcpdoArgv: every tools/call spelling normalizes the same", () => { + const spellings = [ + ["tools/call", "get_sum", "a:=2", "b:=3", "--connection", "test-stdio"], + ["tools/call", "--conn", "test-stdio", "get_sum", "a:=2", "b:=3"], + ["@test-stdio", "tools/call", "get_sum", '{"a":2,"b":3}'], + [ + "tools/call", + "--tool-name", + "get_sum", + "--tool-arg", + "a=2", + "b=3", + "--connection", + "test-stdio", + ], + [ + "tools/call", + "get_sum", + "--tool-args-json", + '{"a":2,"b":3}', + "--connection", + "test-stdio", + ], + [ + "tools/call", + "--tool-name=get_sum", + "--tool-arg=a=2", + "b=3", + "--connection=test-stdio", + ], + ]; + for (const argv of spellings) { + const p = parseMcpdoArgv(argv); + assert.equal(p.cmd, "tools/call", argv.join(" ")); + assert.equal(p.connection, "test-stdio", argv.join(" ")); + assert.equal(p.tool, "get_sum", argv.join(" ")); + assert.ok(valuesMatch(2, p.args.a), argv.join(" ")); + assert.ok(valuesMatch(3, p.args.b), argv.join(" ")); + } +}); + +test("parseMcpdoArgv: global flags do not become positionals", () => { + const p = parseMcpdoArgv([ + "--format", + "json", + "--plain", + "tools/list", + "--connection", + "x", + ]); + assert.equal(p.cmd, "tools/list"); + assert.equal(p.connection, "x"); + assert.equal(p.tool, null); +}); + +test("parseMcpdoArgv: daemon subcommand folds into cmd", () => { + assert.equal(parseMcpdoArgv(["daemon", "stop"]).cmd, "daemon stop"); + assert.equal(parseMcpdoArgv(["daemon", "status"]).cmd, "daemon status"); +}); + +test("parseMcpdoArgv: boolean flag does not swallow the tool name", () => { + const p = parseMcpdoArgv(["tools/call", "--task", "slow_tool", "n:=1"]); + assert.equal(p.tool, "slow_tool"); + assert.ok(valuesMatch(1, p.args.n)); +}); + +test("valuesMatch: canonical across strings, numbers, and key order", () => { + assert.ok(valuesMatch(2, "2")); + assert.ok(valuesMatch("hello", "hello")); + assert.ok(valuesMatch({ b: 1, a: "2" }, { a: 2, b: 1 })); + assert.ok(!valuesMatch(2, 3)); + assert.ok(!valuesMatch("2", "2x")); +}); + +test("matchCall: full match on the add case", () => { + const r = record( + ["tools/call", "get_sum", "a:=2", "b:=3", "--connection", "test-stdio"], + { stdout: '{"result":5}' }, + ); + assert.equal( + matchCall( + { + cmd: "tools/call", + connection: "test-stdio", + tool: "get_sum", + args: { a: 2, b: 3 }, + result: { result: 5 }, + }, + r, + ), + null, + ); +}); + +test("matchCall: implicit MRU connection is accepted, explicit wrong one is not", () => { + const m = { cmd: "tools/list", connection: "test-stdio" }; + assert.equal(matchCall(m, record(["tools/list"])), null); + const wrong = matchCall(m, record(["tools/list", "--connection", "other"])); + assert.match(wrong, /connection is `other`/); +}); + +test("matchCall: connect names its connection positionally", () => { + const m = { cmd: "connect", connection: "test-stdio" }; + assert.equal(matchCall(m, record(["connect", "test-stdio"])), null); + assert.equal(matchCall(m, record(["connect", "@test-stdio"])), null); + assert.equal( + matchCall(m, record(["connect", "--connection", "test-stdio"])), + null, + ); + assert.match( + matchCall(m, record(["connect", "other-entry"])), + /connection is `other-entry`/, + ); +}); + +test("matchCall: wrong tool, wrong arg, missing arg, nonzero exit", () => { + const base = { + cmd: "tools/call", + tool: "get_sum", + args: { a: 2, b: 3 }, + }; + assert.match( + matchCall(base, record(["tools/call", "echo", "a:=2", "b:=3"])), + /tool is `echo`/, + ); + assert.match( + matchCall(base, record(["tools/call", "get_sum", "a:=2", "b:=4"])), + /arg `b` is 4/, + ); + assert.match( + matchCall(base, record(["tools/call", "get_sum", "a:=2"])), + /arg `b` missing/, + ); + assert.match( + matchCall( + base, + record(["tools/call", "get_sum", "a:=2", "b:=3"], { exit: 1 }), + ), + /exit is 1/, + ); +}); + +test("matchCall: result against text output is a distinct diagnostic", () => { + const r = record(["tools/call", "get_sum", "a:=2", "b:=3"], { + stdout: "result: 5\n", + }); + assert.match( + matchCall({ cmd: "tools/call", result: { result: 5 } }, r), + /stdout is not JSON/, + ); + assert.equal(matchCall({ cmd: "tools/call", stdoutMatch: "5" }, r), null); +}); + +test("matchCall: result reads dotted paths", () => { + const r = record(["tools/list"], { + stdout: '{"tools":[{"name":"get_sum"}]}', + }); + assert.equal( + matchCall({ cmd: "tools/list", result: { "tools.0.name": "get_sum" } }, r), + null, + ); +}); + +test("evalExpectCalls: ordered subsequence with unrelated calls between", () => { + const records = [ + record(["daemon", "status"]), + record(["connect", "test-stdio"]), + record(["tools/list"]), + record(["tools/call", "get_sum", "a:=2", "b:=3"]), + ]; + const { ok } = evalExpectCalls( + [ + { cmd: "connect", connection: "test-stdio" }, + { cmd: "tools/call", tool: "get_sum", args: { a: 2, b: 3 } }, + ], + records, + ); + assert.ok(ok); +}); + +test("evalExpectCalls: order violations and misses carry diagnostics", () => { + const records = [ + record(["tools/call", "get_sum", "a:=2", "b:=4"]), + record(["connect", "test-stdio"]), + ]; + const out = evalExpectCalls( + [ + { cmd: "connect", connection: "test-stdio" }, + { cmd: "tools/call", args: { b: 3 } }, + ], + records, + ); + assert.ok(!out.ok); + // connect matched (index 1), so the tools/call must come after it — the + // earlier wrong call does not count and there is no later one. + assert.equal(out.failures.length, 1); + assert.match(out.failures[0], /expectCalls\[1\]/); +}); + +test("evalExpectCalls: empty transcript says so", () => { + const out = evalExpectCalls([{ cmd: "connect" }], []); + assert.match(out.failures[0], /no mcpdo invocations recorded/); +}); + +test("streamText concatenates one stream in order", () => { + const r = { + events: [ + { t: 1, stream: "stdout", data: "a" }, + { t: 2, stream: "stderr", data: "X" }, + { t: 3, stream: "stdout", data: "b" }, + ], + }; + assert.equal(streamText(r, "stdout"), "ab"); + assert.equal(streamText(r, "stderr"), "X"); +}); + +test("validateBehaviorCase: accepts the real shape", () => { + assert.deepEqual( + validateBehaviorCase( + { + kind: "behavior", + prompt: "Add 2 and 3", + expectCalls: [ + { + cmd: "tools/call", + connection: "test-stdio", + tool: "get_sum", + args: { a: 2, b: 3 }, + stdoutMatch: "5", + }, + ], + }, + 0, + ), + [], + ); +}); + +test("validateBehaviorCase: catches typos, bad types, bad regex", () => { + const errs = validateBehaviorCase( + { + prompt: "", + expectCalls: [ + { cmd: "", tooll: "x" }, + { cmd: "ok", args: [], exit: "0", stdoutMatch: "(" }, + "nope", + ], + }, + 3, + ); + assert.ok(errs.some((e) => /`prompt`/.test(e))); + assert.ok(errs.some((e) => /`cmd` is required/.test(e))); + assert.ok(errs.some((e) => /unknown key `tooll`/.test(e))); + assert.ok(errs.some((e) => /`args` must be an object/.test(e))); + assert.ok(errs.some((e) => /`exit` must be an integer/.test(e))); + assert.ok(errs.some((e) => /not a valid regex/.test(e))); + assert.ok(errs.some((e) => /must be an object/.test(e))); +}); + +test("validateBehaviorCase: autoConsent must be a boolean", () => { + const base = { prompt: "p", expectCalls: [{ cmd: "connect" }] }; + assert.deepEqual(validateBehaviorCase({ ...base, autoConsent: true }, 0), []); + assert.ok( + validateBehaviorCase({ ...base, autoConsent: "yes" }, 0).some((e) => + /`autoConsent` must be a boolean/.test(e), + ), + ); +}); + +test("validateBehaviorCase: consentDelayMs and expectReply", () => { + const base = { prompt: "p", expectCalls: [{ cmd: "connect" }] }; + assert.deepEqual( + validateBehaviorCase( + { ...base, consentDelayMs: 20000, expectReply: "oauth/authorize\\?" }, + 0, + ), + [], + ); + assert.ok( + validateBehaviorCase({ ...base, consentDelayMs: -1 }, 0).some((e) => + /`consentDelayMs` must be a non-negative integer/.test(e), + ), + ); + assert.ok( + validateBehaviorCase({ ...base, consentDelayMs: 1.5 }, 0).some((e) => + /`consentDelayMs` must be a non-negative integer/.test(e), + ), + ); + assert.ok( + validateBehaviorCase({ ...base, expectReply: "" }, 0).some((e) => + /`expectReply` must be a non-empty string/.test(e), + ), + ); + assert.ok( + validateBehaviorCase({ ...base, expectReply: "(" }, 0).some((e) => + /`expectReply` is not a valid regex/.test(e), + ), + ); + assert.deepEqual( + validateBehaviorCase({ ...base, consentAfterReply: true }, 0), + [], + ); + assert.ok( + validateBehaviorCase({ ...base, consentAfterReply: "yes" }, 0).some((e) => + /`consentAfterReply` must be a boolean/.test(e), + ), + ); +}); + +test("validateBehaviorCase: multi-turn fields (followUp, expectReplyFollowUp, expectLastCallTurn1)", () => { + const base = { prompt: "p", expectCalls: [{ cmd: "connect" }] }; + // A valid multi-turn case: followUp paired with expectReplyFollowUp. + assert.deepEqual( + validateBehaviorCase( + { + ...base, + followUp: "continue", + expectReplyFollowUp: "add", + expectLastCallTurn1: "connect", + }, + 0, + ), + [], + ); + // followUp without expectReplyFollowUp is an error. + assert.ok( + validateBehaviorCase({ ...base, followUp: "go" }, 0).some((e) => + /`followUp` requires `expectReplyFollowUp`/.test(e), + ), + ); + // expectReplyFollowUp without followUp is an error. + assert.ok( + validateBehaviorCase({ ...base, expectReplyFollowUp: "add" }, 0).some((e) => + /`expectReplyFollowUp` requires `followUp`/.test(e), + ), + ); + // Empty / bad-regex / wrong-type checks. + assert.ok( + validateBehaviorCase( + { ...base, followUp: "", expectReplyFollowUp: "add" }, + 0, + ).some((e) => /`followUp` must be a non-empty string/.test(e)), + ); + assert.ok( + validateBehaviorCase( + { ...base, followUp: "go", expectReplyFollowUp: "(" }, + 0, + ).some((e) => /`expectReplyFollowUp` is not a valid regex/.test(e)), + ); + assert.ok( + validateBehaviorCase({ ...base, expectLastCallTurn1: "" }, 0).some((e) => + /`expectLastCallTurn1` must be a non-empty string/.test(e), + ), + ); + // expectLastCallTurn1 is valid on its own (a single-turn relay-and-stop gate). + assert.deepEqual( + validateBehaviorCase({ ...base, expectLastCallTurn1: "connect" }, 0), + [], + ); + // rejectCompletedTurn1 (the robust list-tools gate) has the same shape rules. + assert.ok( + validateBehaviorCase({ ...base, rejectCompletedTurn1: "" }, 0).some((e) => + /`rejectCompletedTurn1` must be a non-empty string/.test(e), + ), + ); + assert.deepEqual( + validateBehaviorCase({ ...base, rejectCompletedTurn1: "tools/list" }, 0), + [], + ); +}); + +test("claudeSessionId: last session_id wins, null when absent", () => { + const raw = + JSON.stringify({ type: "system", subtype: "init", session_id: "s-1" }) + + "\n" + + "not json\n" + + JSON.stringify({ type: "assistant", message: { content: [] } }) + + "\n" + + JSON.stringify({ type: "result", session_id: "s-2" }) + + "\n"; + assert.equal(claudeSessionId(raw), "s-2"); + assert.equal(claudeSessionId(""), null); + assert.equal(claudeSessionId(JSON.stringify({ type: "assistant" })), null); + // A torn final line (truncated mid-write) must not lose an earlier id. + assert.equal( + claudeSessionId( + JSON.stringify({ session_id: "s-1" }) + '\n{"session_id":"s-', + ), + "s-1", + ); +}); + +test("lastCommandBefore: final command by start time, honoring the boundary", () => { + const rec = (cmd, start) => ({ argv: [cmd], start }); + const records = [ + rec("tools/list", 10), // failed pre-connect probe + rec("connect", 20), // relayed the URL, then stopped + rec("tools/list", 40), // turn 2, after the follow-up + ]; + // Before the follow-up (boundary 30): connect was the last thing. + assert.equal(lastCommandBefore(records, 30), "connect"); + // Infinity covers every record (a single-turn case). + assert.equal(lastCommandBefore(records, Infinity), "tools/list"); + // A later-exiting but earlier-starting record does not displace connect. + assert.equal( + lastCommandBefore( + [rec("connect", 20), { argv: ["connections/show"], start: 15 }], + 30, + ), + "connect", + ); + // Records with no numeric start are ignored; nothing qualifying is null. + assert.equal(lastCommandBefore([{ argv: ["connect"] }], 30), null); + assert.equal(lastCommandBefore([], 30), null); +}); + +test("succeededBefore: exit-0 task command before the boundary", () => { + const rec = (cmd, start, exit) => ({ argv: [cmd], start, exit }); + // A chained `connect && tools/list` where tools/list errors (auth_required) + // then succeeds in turn 2: NOT completed in turn 1. + const chained = [ + rec("connect", 20, 0), + rec("tools/list", 22, 3), // optimistic, errored — agent waited + rec("tools/list", 40, 0), // turn 2, after sign-in + ]; + assert.equal(succeededBefore(chained, 30, "tools/list"), false); + // A poll that ran tools/list to success in turn 1: completed without waiting. + const polled = [rec("connect", 20, 0), rec("tools/list", 25, 0)]; + assert.equal(succeededBefore(polled, 30, "tools/list"), true); + // Argv carries a connection prefix (`@name tools/list`) — still matched. + assert.equal( + succeededBefore( + [{ argv: ["@secure-add", "tools/list"], start: 25, exit: 0 }], + 30, + "tools/list", + ), + true, + ); + // Records with no numeric start, or at/after the boundary, are ignored. + assert.equal( + succeededBefore([{ argv: ["tools/list"], exit: 0 }], 30, "tools/list"), + false, + ); + assert.equal( + succeededBefore([rec("tools/list", 40, 0)], 30, "tools/list"), + false, + ); +}); + +test("assistantReplyText: reads both agents' event shapes, skips noise", () => { + const claude = + JSON.stringify({ + type: "assistant", + message: { + content: [ + { type: "text", text: "open http://127.0.0.1:1/oauth/authorize?x=1" }, + { type: "tool_use", name: "Bash", input: {} }, + { + type: "thinking", + thinking: "narrated aside with a secret-ish url", + }, + ], + }, + }) + "\n"; + const copilot = + JSON.stringify({ + type: "assistant.message", + data: { content: "then sign in", toolRequests: [] }, + }) + "\n"; + const noise = + "not json\n" + + JSON.stringify({ type: "tool_result", content: "secret url" }) + + "\n" + + JSON.stringify({ type: "assistant.message_delta", data: { content: 5 } }) + + "\n"; + const text = assistantReplyText(claude + noise + copilot); + assert.equal( + text, + "open http://127.0.0.1:1/oauth/authorize?x=1\nthen sign in", + ); + // Tool results are what the agent SAW, not what it said — they must never + // satisfy a reply matcher. + assert.ok(!text.includes("secret url")); + // Thinking blocks count only when the host displays them (Claude Code + // narration) — opt-in via includeThinking, default off. + assert.ok(!text.includes("narrated aside")); + const withThinking = assistantReplyText(claude + noise + copilot, { + includeThinking: true, + }); + assert.ok(withThinking.includes("narrated aside")); + assert.deepEqual(REPLY_TEXT_OPTIONS.claude, { includeThinking: true }); + assert.deepEqual(REPLY_TEXT_OPTIONS.copilot, { includeThinking: false }); + assert.equal(assistantReplyText(""), ""); +}); + +test("matchPhases: interleaved prompt/answer/result ordering", () => { + const r = { + argv: ["tools/call", "collect"], + exit: 0, + events: [ + { t: 1, stream: "stdout", data: "Enter your na" }, + { t: 2, stream: "stdout", data: "me: " }, // pattern spans chunks + { t: 3, stream: "stdin", data: "Ada\n" }, + { t: 4, stream: "stdout", data: '{"ok":true}\n' }, + ], + }; + assert.equal( + matchPhases( + [ + { stream: "stdout", match: "Enter your name" }, + { stream: "stdin", match: "Ada" }, + { stream: "stdout", match: '"ok"' }, + ], + r, + ), + null, + ); + // The answer cannot come before the prompt. + assert.match( + matchPhases( + [ + { stream: "stdin", match: "Ada" }, + { stream: "stdout", match: "Enter your name" }, + ], + r, + ), + /matched stdout only BEFORE/, + ); + assert.match( + matchPhases([{ stream: "stderr", match: "x" }], r), + /no stderr data recorded/, + ); + assert.match( + matchPhases([{ stream: "stdout", match: "missing" }], r), + /not found on stdout/, + ); +}); + +test("matchCall: phases participate in a full matcher", () => { + const r = record(["connect", "test-stdio"], { + stdout: "Visit https://idp.example/auth to continue\nConnection ready\n", + }); + assert.equal( + matchCall( + { + cmd: "connect", + phases: [ + { stream: "stdout", match: "https://idp\\.example/auth" }, + { stream: "stdout", match: "Connection ready" }, + ], + }, + r, + ), + null, + ); +}); + +test("validateBehaviorCase: phases schema", () => { + const errs = validateBehaviorCase( + { + prompt: "p", + expectCalls: [ + { + cmd: "connect", + phases: [ + { stream: "socket", match: "x" }, + { stream: "stdout", match: "(", extra: 1 }, + "nope", + ], + }, + { cmd: "ok", phases: [] }, + ], + }, + 0, + ); + assert.ok(errs.some((e) => /phases\[0\]\.stream must be one of/.test(e))); + assert.ok( + errs.some((e) => /phases\[1\]\.match is not a valid regex/.test(e)), + ); + assert.ok(errs.some((e) => /phases\[1\] unknown key `extra`/.test(e))); + assert.ok(errs.some((e) => /phases\[2\] must be an object/.test(e))); + assert.ok(errs.some((e) => /`phases` must be a non-empty array/.test(e))); +}); + +test("validateServerSpec: url form, composed form, and rejects", () => { + assert.deepEqual(validateServerSpec(undefined, 0), []); + assert.deepEqual( + validateServerSpec({ url: "http://127.0.0.1:3999/mcp" }, 0), + [], + ); + assert.deepEqual( + validateServerSpec( + { + serverInfo: { name: "composed", version: "1.0.0" }, + tools: [{ preset: "add" }], + }, + 0, + ), + [], + ); + assert.ok( + validateServerSpec({ url: "ftp://x" }, 0).some((e) => + /http\(s\) URL/.test(e), + ), + ); + assert.ok( + validateServerSpec({ url: "http://x", tools: [{ preset: "add" }] }, 0).some( + (e) => /no other keys/.test(e), + ), + ); + assert.ok( + validateServerSpec({ tools: [{ preset: "add" }] }, 0).some((e) => + /needs `serverInfo`/.test(e), + ), + ); + assert.deepEqual( + validateServerSpec( + { + serverInfo: { name: "protected-api", version: "1.0.0" }, + transport: { type: "streamable-http" }, + oauth: { enabled: true, mode: "combined" }, + }, + 0, + ), + [], + "http composed form (incl. oauth) is valid", + ); + assert.ok( + validateServerSpec( + { + serverInfo: { name: "c", version: "1" }, + transport: { type: "sse" }, + }, + 0, + ).some((e) => /sse fixtures are not supported/.test(e)), + ); + assert.ok(validateServerSpec([], 0).some((e) => /must be an object/.test(e))); +}); + +test("validateCaseServers: map form, exclusivity, and per-entry labels", () => { + const spec = { serverInfo: { name: "s", version: "1" } }; + assert.deepEqual( + validateCaseServers( + { servers: { calendar: spec, "weather-api": { url: "http://x/mcp" } } }, + 0, + ), + [], + ); + assert.ok( + validateCaseServers({ server: spec, servers: { a: spec } }, 0).some((e) => + /mutually exclusive/.test(e), + ), + ); + assert.ok( + validateCaseServers({ servers: {} }, 0).some((e) => + /must not be empty/.test(e), + ), + ); + assert.ok( + validateCaseServers({ servers: ["x"] }, 0).some((e) => + /name→spec object/.test(e), + ), + ); + assert.ok( + validateCaseServers({ servers: { "bad name!": spec } }, 0).some((e) => + /not a valid catalog entry name/.test(e), + ), + ); + assert.ok( + validateCaseServers({ servers: { a: undefined } }, 0).some((e) => + /must not be undefined/.test(e), + ), + ); + // Nested spec errors carry the entry name. + assert.ok( + validateCaseServers({ servers: { alpha: { url: "ftp://x" } } }, 3).some( + (e) => /behavior case 3 `servers`\.alpha/.test(e), + ), + ); +}); + +test("validateBehaviorCase: server field is validated through the case", () => { + const errs = validateBehaviorCase( + { prompt: "p", expectCalls: [{ cmd: "connect" }], server: { url: "nope" } }, + 2, + ); + assert.ok(errs.some((e) => /behavior case 2 `server`/.test(e))); +}); + +test("validateBehaviorCase: empty expectCalls is an error", () => { + const errs = validateBehaviorCase({ prompt: "p", expectCalls: [] }, 0); + assert.ok(errs.some((e) => /non-empty array/.test(e))); +}); diff --git a/scripts/lib/mcpdo-eval-server-launcher.mjs b/scripts/lib/mcpdo-eval-server-launcher.mjs new file mode 100644 index 0000000000..596c033252 --- /dev/null +++ b/scripts/lib/mcpdo-eval-server-launcher.mjs @@ -0,0 +1,57 @@ +#!/usr/bin/env node +// Stdio entrypoint for a COMPOSED test server, owned by the mcpdo behavior +// eval (`skills:eval:mcpdo`). +// +// A behavior case may carry a `server` field — the test-servers declarative +// config-file shape (serverInfo + preset refs + capability switches; see +// test-servers/src/load-config.ts). The harness writes it to disk and points +// the sample's catalog entry at this launcher, so each case talks to exactly +// the server it needs: an eliciting tool, task tools, subscriptions, +// whatever the preset registry can compose. +// +// Deliberately NOT an extension of `test-server-stdio.js`: that entrypoint +// is a purpose-built default composition and stays that way. This launcher +// is the composition path the framework already exposes — `loadConfig` → +// `resolveConfig` → `TestServerStdio` (whose constructor takes any +// ServerConfig; only its standalone main hard-wires the default). Built +// output is imported, same as the eval's use of the default server: the +// eval measures what a user would run, and `scripts/` cannot import the +// workspace package by name anyway. +// +// Usage: mcpdo-eval-server-launcher.mjs <config.(json|yaml)> + +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const ROOT = path.resolve( + path.dirname(fileURLToPath(import.meta.url)), + "..", + "..", +); + +async function main() { + const configPath = process.argv[2]; + if (!configPath) { + process.stderr.write( + "mcpdo-eval-server-launcher: usage: mcpdo-eval-server-launcher.mjs <config.(json|yaml)>\n", + ); + process.exit(2); + } + const { loadConfig, resolveConfig, TestServerStdio } = await import( + path.join(ROOT, "test-servers", "build", "index.js") + ); + const loaded = loadConfig(path.resolve(configPath)); + if (loaded.transport?.type !== "stdio") { + throw new Error( + `config transport.type must be "stdio" (got ${JSON.stringify(loaded.transport?.type)})`, + ); + } + const config = resolveConfig(loaded); + await new TestServerStdio(config).start(); + // Stdio transport holds the process open; exit is the client closing us. +} + +main().catch((err) => { + process.stderr.write(`mcpdo-eval-server-launcher: ${err?.message ?? err}\n`); + process.exit(1); +}); diff --git a/scripts/lib/mcpdo-eval-shim.mjs b/scripts/lib/mcpdo-eval-shim.mjs new file mode 100644 index 0000000000..3143c1a2d1 --- /dev/null +++ b/scripts/lib/mcpdo-eval-shim.mjs @@ -0,0 +1,139 @@ +#!/usr/bin/env node +// Transparent recording wrapper around the real mcpdo, for behavior evals. +// +// The behavior eval (`skills:eval:mcpdo`) needs to know what an agent's +// shell actually asked mcpdo to do and what mcpdo answered — WITHOUT parsing +// shell strings out of the agent's event stream (quoting rules, chained +// commands and subshells make that a reimplementation of `sh`). So the +// harness puts a `mcpdo` shim first on PATH; the shim execs this file, which +// spawns the REAL CLI and tees all three stdio streams through untouched +// while recording a timestamped transcript. +// +// One JSON line per invocation, appended to `$MCPDO_EVAL_LOG`: +// +// { argv, exit, start, end, events: [{ t, stream, data }] } +// +// `events` interleaves stdin/stdout/stderr in arrival order, so an +// interactive invocation (elicitation prompts answered over stdin, two-phase +// OAuth output on one blocked pipe) is captured as faithfully as an atomic +// one — for the simple case the transcript degenerates to a single stdout +// event. Output is recorded VERBATIM: the shim never injects `--format json` +// or any other flag, because which format the agent asked for is part of +// what is being measured. +// +// The record is written on child exit with a single appendFileSync — an +// O_APPEND write of one line, atomic enough for concurrent invocations +// within a sample (POSIX; the eval is POSIX-only, stated where it runs). + +import { spawn } from "node:child_process"; +import { appendFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +/** + * Run the real CLI, teeing stdio and recording the transcript. + * + * Injectable so tests can drive it with a fixture child and in-memory + * streams; the module main wires the real process. + * + * @param {object} opts + * @param {string} opts.realBin Path to the real mcp-bin.js. + * @param {string[]} opts.argv What the agent passed after `mcpdo`. + * @param {string} opts.logPath Transcript destination (NDJSON, appended). + * @param {NodeJS.ReadableStream} opts.stdin + * @param {NodeJS.WritableStream} opts.stdout + * @param {NodeJS.WritableStream} opts.stderr + * @param {typeof spawn} [opts.spawnFn] + * @param {typeof appendFileSync} [opts.appendFn] + * @returns {Promise<number>} The child's exit code. + */ +export function runShim({ + realBin, + argv, + logPath, + stdin, + stdout, + stderr, + spawnFn = spawn, + appendFn = appendFileSync, +}) { + return new Promise((resolve, reject) => { + const record = { + argv, + exit: null, + start: Date.now(), + end: null, + events: [], + }; + const child = spawnFn(process.execPath, [realBin, ...argv], { + stdio: ["pipe", "pipe", "pipe"], + }); + const tap = (stream) => (chunk) => { + record.events.push({ + t: Date.now(), + stream, + data: chunk.toString("utf8"), + }); + }; + stdin.on("data", (chunk) => { + tap("stdin")(chunk); + child.stdin.write(chunk); + }); + stdin.on("end", () => child.stdin.end()); + // The child may exit without reading piped stdin; that EPIPE is its + // business, not a shim failure. + child.stdin.on("error", () => {}); + child.stdout.on("data", (chunk) => { + tap("stdout")(chunk); + stdout.write(chunk); + }); + child.stderr.on("data", (chunk) => { + tap("stderr")(chunk); + stderr.write(chunk); + }); + child.on("error", reject); + child.on("close", (code) => { + record.exit = code ?? 1; + record.end = Date.now(); + appendFn(logPath, JSON.stringify(record) + "\n"); + resolve(record.exit); + }); + }); +} + +const isMain = + process.argv[1] && + path.resolve(process.argv[1]) === fileURLToPath(import.meta.url); + +if (isMain) { + const realBin = process.env.MCPDO_EVAL_REAL; + const logPath = process.env.MCPDO_EVAL_LOG; + if (!realBin || !logPath) { + process.stderr.write( + "mcpdo-eval-shim: MCPDO_EVAL_REAL and MCPDO_EVAL_LOG must be set\n", + ); + process.exit(2); + } + runShim({ + realBin, + argv: process.argv.slice(2), + logPath, + stdin: process.stdin, + stdout: process.stdout, + stderr: process.stderr, + }).then( + // `process.exitCode` + natural exit, not `process.exit()`: exit() + // discards queued stdout/stderr writes, truncating large JSON results + // on macOS pipes. Destroying stdin releases the last open handle so + // the process drains its writes and exits on its own. + (code) => { + process.exitCode = code; + process.stdin.destroy(); + }, + (err) => { + process.stderr.write(`mcpdo-eval-shim: ${err?.message ?? err}\n`); + process.exitCode = 2; + process.stdin.destroy(); + }, + ); +} diff --git a/scripts/lib/mcpdo-eval-shim.test.mjs b/scripts/lib/mcpdo-eval-shim.test.mjs new file mode 100644 index 0000000000..d7ff66cea0 --- /dev/null +++ b/scripts/lib/mcpdo-eval-shim.test.mjs @@ -0,0 +1,90 @@ +// The shim's whole job is fidelity: same argv to the real binary, same bytes +// on the same streams, same exit code — with a transcript on the side. So +// the test drives it end-to-end against a fixture "real CLI" and asserts on +// all four at once. A unit test of the internals would pass while the tee +// dropped a stream (Copilot). + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { spawn } from "node:child_process"; +import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const SHIM = path.join( + path.dirname(fileURLToPath(import.meta.url)), + "mcpdo-eval-shim.mjs", +); + +// Echoes argv on stdout, a marker + upper-cased stdin on stderr, exits 3. +const FIXTURE = ` +let input = ""; +process.stdin.on("data", (c) => (input += c)); +process.stdin.on("end", () => { + process.stdout.write("argv:" + process.argv.slice(2).join(",") + "\\n"); + process.stderr.write("err:" + input.toUpperCase()); + process.exit(3); +}); +`; + +function runShimProcess(args, { stdinText = "", env = {} } = {}) { + return new Promise((resolve, reject) => { + const child = spawn(process.execPath, [SHIM, ...args], { + env: { ...process.env, ...env }, + stdio: ["pipe", "pipe", "pipe"], + }); + let stdout = ""; + let stderr = ""; + child.stdout.on("data", (c) => (stdout += c)); + child.stderr.on("data", (c) => (stderr += c)); + child.on("error", reject); + child.on("close", (code) => resolve({ code, stdout, stderr })); + child.stdin.end(stdinText); + }); +} + +test("shim tees stdio verbatim, mirrors exit, and records the transcript", async () => { + const dir = mkdtempSync(path.join(os.tmpdir(), "mcpdo-shim-test-")); + try { + const realBin = path.join(dir, "fixture.mjs"); + const logPath = path.join(dir, "log.ndjson"); + writeFileSync(realBin, FIXTURE); + + const out = await runShimProcess(["tools/call", "get_sum", "a:=2"], { + stdinText: "hi", + env: { MCPDO_EVAL_REAL: realBin, MCPDO_EVAL_LOG: logPath }, + }); + assert.equal(out.code, 3); + assert.equal(out.stdout, "argv:tools/call,get_sum,a:=2\n"); + assert.equal(out.stderr, "err:HI"); + + const records = readFileSync(logPath, "utf8") + .trim() + .split("\n") + .map((l) => JSON.parse(l)); + assert.equal(records.length, 1); + const r = records[0]; + assert.deepEqual(r.argv, ["tools/call", "get_sum", "a:=2"]); + assert.equal(r.exit, 3); + assert.ok(r.start <= r.end); + const byStream = (s) => + r.events + .filter((e) => e.stream === s) + .map((e) => e.data) + .join(""); + assert.equal(byStream("stdin"), "hi"); + assert.equal(byStream("stdout"), "argv:tools/call,get_sum,a:=2\n"); + assert.equal(byStream("stderr"), "err:HI"); + } finally { + rmSync(dir, { recursive: true, force: true }); + } +}); + +test("shim refuses to run without its env contract", async () => { + const out = await runShimProcess(["anything"], { + env: { MCPDO_EVAL_REAL: "", MCPDO_EVAL_LOG: "" }, + }); + assert.equal(out.code, 2); + assert.match(out.stderr, /MCPDO_EVAL_REAL and MCPDO_EVAL_LOG/); +}); diff --git a/scripts/lib/workflow-gate.test.mjs b/scripts/lib/workflow-gate.test.mjs index ae5911fc7a..dbc04bd814 100644 --- a/scripts/lib/workflow-gate.test.mjs +++ b/scripts/lib/workflow-gate.test.mjs @@ -614,6 +614,25 @@ describe("the gate's name", () => { assert.match(scripts["local:gate:stages"], /local:storybook/); }); + it("checks DCO signoffs before pushing, and only locally (#2616)", () => { + // The pre-push half of the DCO check: `dco.yml` reports on the PR, this + // catches the unsigned commit before it is pushed. It reads + // `origin/v2/main`, which CI's shallow push checkout does not have — so it + // must stay unreachable from `validate`, which CI runs. + assert.equal( + scripts["local:dco"], + "node scripts/dco-check.mjs --base origin/v2/main --exclude origin/main", + ); + assert.ok( + scriptChainRuns(scripts, "local:gate", "local:dco"), + "local:gate must run local:dco", + ); + assert.ok( + !reachableScripts(scripts, "validate").has("local:dco"), + "validate (CI) must not reach local:dco", + ); + }); + it("runs its stages under the lease wrapper, and nothing else (#2339)", () => { // The wrapper serializes gates across worktrees. It must be the ONLY thing // `local:gate` does — a stage placed beside it would run outside the @@ -639,7 +658,7 @@ describe("the gate's name", () => { // that keep it honest: the gate no longer reaches a client's bare `test`, // it still reaches every non-test check `validate` reaches, and `validate` // itself (CI's inner loop) is untouched. - const clients = ["web", "cli", "tui", "launcher"]; + const clients = ["web", "cli", "mcpdo", "tui", "launcher"]; const clientScripts = Object.fromEntries( clients.map((c) => [ c, @@ -693,7 +712,7 @@ describe("the gate's name", () => { for (const name of inner) if ( name !== "validate" && - !/^validate:(web|cli|tui|launcher)$/.test(name) + !/^validate:(web|cli|mcpdo|tui|launcher)$/.test(name) ) assert.ok(gate.has(name), `local:validate must reach ${name}`); }); diff --git a/scripts/pack-and-verify.mjs b/scripts/pack-and-verify.mjs index 04538775c5..7e88ef79b8 100644 --- a/scripts/pack-and-verify.mjs +++ b/scripts/pack-and-verify.mjs @@ -391,6 +391,109 @@ try { fail(`\`--cli … tools/list\` missing expected "echo" tool`); } + // 4b². The tarball also ships the `mcpdo` connection-CLI bin (#1783). Verify + // the installed shim resolves and runs: `--help` (dispatch/build + // resolution) plus a daemon-free command (`servers/list` against the + // same catalog — no daemon spawn, no MCP connection), so a wrong bin + // path or an incompletely packed mcpdo build fails the gate. + step("verifying installed `mcpdo` (--help, daemon-free servers/list)..."); + const mcpdoBin = join( + work, + "node_modules", + ".bin", + process.platform === "win32" ? "mcpdo.cmd" : "mcpdo", + ); + if (!existsSync(mcpdoBin)) { + fail(`installed \`mcpdo\` bin not found at ${mcpdoBin}`); + } + const runMcpdo = (args, extraEnv = {}) => { + const r = spawnSync(shellArgs([mcpdoBin])[0], shellArgs(args), { + cwd: work, + encoding: "utf8", + env: { ...process.env, ...extraEnv }, + shell: WIN_SHELL, + }); + return { status: r.status, output: `${r.stdout ?? ""}${r.stderr ?? ""}` }; + }; + const mcpdoHelp = runMcpdo(["--help"]); + if (mcpdoHelp.status !== 0 || !mcpdoHelp.output.includes("Usage: mcpdo")) { + fail( + `\`mcpdo --help\` exited ${mcpdoHelp.status} or missing usage banner\n` + + mcpdoHelp.output.slice(0, 800), + ); + } + // Point the daemon dir into the throwaway consumer so the check never sees + // (or touches) a real daemon on the host. + const mcpdoServers = runMcpdo( + ["servers/list", "--catalog", catalogPath, "--plain"], + { MCP_INSPECTOR_DAEMON_DIR: join(work, "mcpdo-daemon") }, + ); + if (mcpdoServers.status !== 0 || !mcpdoServers.output.includes("test")) { + fail( + `\`mcpdo servers/list\` exited ${mcpdoServers.status} or missing "test" entry\n` + + mcpdoServers.output.slice(0, 800), + ); + } + + // 4b³. Daemon lifecycle from the installed package: `connect` must locate + // and spawn the separately shipped `build/mcpdod.js` — the daemon-free + // checks above pass even when that artifact is missing or mislocated, + // yet every connection command would fail at startup. Connect against + // the same stdio fixture, verify the connection is listed, then tear + // everything down (stop the daemon before any fail() so nothing + // outlives the check). + step("verifying installed `mcpdo` daemon flow (connect/list/disconnect)..."); + const mcpdoDaemonEnv = { + MCP_INSPECTOR_DAEMON_DIR: join(work, "mcpdo-daemon"), + }; + const stopMcpdoDaemon = () => runMcpdo(["daemon", "stop"], mcpdoDaemonEnv); + const failMcpdoDaemonFlow = (message) => { + stopMcpdoDaemon(); + fail(message); + }; + const mcpdoConnect = runMcpdo( + ["connect", "test", "--catalog", catalogPath, "--plain"], + mcpdoDaemonEnv, + ); + if (mcpdoConnect.status !== 0 || !mcpdoConnect.output.includes("test")) { + failMcpdoDaemonFlow( + `\`mcpdo connect test\` exited ${mcpdoConnect.status} — the packaged ` + + `daemon (build/mcpdod.js) likely failed to start\n` + + mcpdoConnect.output.slice(0, 800), + ); + } + const mcpdoConnections = runMcpdo( + ["connections/list", "--plain"], + mcpdoDaemonEnv, + ); + if ( + mcpdoConnections.status !== 0 || + !mcpdoConnections.output.includes("test") + ) { + failMcpdoDaemonFlow( + `\`mcpdo connections/list\` exited ${mcpdoConnections.status} or missing ` + + `the "test" connection\n` + + mcpdoConnections.output.slice(0, 800), + ); + } + const mcpdoDisconnect = runMcpdo( + ["disconnect", "test", "--plain"], + mcpdoDaemonEnv, + ); + if (mcpdoDisconnect.status !== 0) { + failMcpdoDaemonFlow( + `\`mcpdo disconnect test\` exited ${mcpdoDisconnect.status}\n` + + mcpdoDisconnect.output.slice(0, 800), + ); + } + const mcpdoStop = stopMcpdoDaemon(); + if (mcpdoStop.status !== 0) { + fail( + `\`mcpdo daemon stop\` exited ${mcpdoStop.status}\n` + + mcpdoStop.output.slice(0, 800), + ); + } + // 4c. Prod `--web` boot from the installed package — THE critical packaging // path: the runner must locate and serve the shipped `dist` (not rebuild // it) and inject the auth token. Run non-blocking and poll `/`. diff --git a/scripts/pr-link-issue.mjs b/scripts/pr-link-issue.mjs new file mode 100644 index 0000000000..d4f9159d3f --- /dev/null +++ b/scripts/pr-link-issue.mjs @@ -0,0 +1,82 @@ +#!/usr/bin/env node +// Link a PR to the issue it closes (#2558) — `npm run pr:link -- --pr <N> +// --issue <M>`. Step 6 of the pr-flow skill, which previously transcribed +// this as a three-call GraphQL block. +// +// Closing keywords only auto-link for PRs targeting the DEFAULT branch, and +// v2 PRs target `v2/main` — so `Closes #N` there is only a cross-reference. +// The `addCloseIssueReferences` mutation adds the manual closing reference +// (the UI's Development sidebar), which is what puts the PR in the card's +// Linked pull requests field. The link is VERIFIED by reading the PR's +// `closingIssuesReferences` back; `linked: …` prints only on a confirmed +// match, so an unconfirmed mutation fails loudly instead of leaving a card +// with no PR. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { + OWNER, + REPO, + REPO_SLUG, + ghGraphql, + ghJson, + requirePositiveInt, +} from "./lib/gh.mjs"; + +const ISSUE_ID_QUERY = `query($n:Int!){repository(owner:"${OWNER}",name:"${REPO}"){issue(number:$n){id}}}`; +const LINK_MUTATION = `mutation($i:ID!,$p:[ID!]!){addCloseIssueReferences(input:{issueId:$i, pullRequestIds:$p}){clientMutationId}}`; +// `first:100` is the connection's maximum page — with `first:10` a PR already +// linked to ten issues would verify the wrong page and report a successful +// mutation as a failure. +const VERIFY_QUERY = `query($n:Int!){repository(owner:"${OWNER}",name:"${REPO}"){pullRequest(number:$n){closingIssuesReferences(first:100){nodes{number}}}}}`; + +export function parseLinkArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { pr: { type: "string" }, issue: { type: "string" } }, + }); + return { + pr: requirePositiveInt(values.pr, "--pr"), + issue: requirePositiveInt(values.issue, "--issue"), + }; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { pr, issue } = parseLinkArgs(argv); + + const issueId = ghGraphql(spawn, ISSUE_ID_QUERY, { n: issue })?.data + ?.repository?.issue?.id; + if (!issueId) { + throw new Error(`could not resolve issue #${issue}`); + } + const prId = ghJson(spawn, [ + "pr", + "view", + String(pr), + "--repo", + REPO_SLUG, + "--json", + "id", + ]).id; + if (!prId) { + throw new Error(`could not resolve PR #${pr}`); + } + + ghGraphql(spawn, LINK_MUTATION, { i: issueId, p: prId }); + + // Verify: the PR must now list the issue — never report an unconfirmed link. + const linked = ( + ghGraphql(spawn, VERIFY_QUERY, { n: pr })?.data?.repository?.pullRequest + ?.closingIssuesReferences?.nodes ?? [] + ).map((node) => node.number); + if (!linked.includes(issue)) { + throw new Error( + `PR #${pr} closingIssuesReferences reads [${linked.join(", ")}] — #${issue} is not in it`, + ); + } + console.log(`linked: PR #${pr} closes #${issue}`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/pr-link-issue.test.mjs b/scripts/pr-link-issue.test.mjs new file mode 100644 index 0000000000..92b05a1e85 --- /dev/null +++ b/scripts/pr-link-issue.test.mjs @@ -0,0 +1,85 @@ +// Tests for scripts/pr-link-issue.mjs (#2558) — argv validation and the +// link-then-verify orchestration: `linked:` prints only when the PR's +// closingIssuesReferences actually lists the issue. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { main, parseLinkArgs } from "./pr-link-issue.mjs"; + +test("parseLinkArgs requires both numbers", () => { + assert.deepEqual(parseLinkArgs(["--pr", "2559", "--issue", "2558"]), { + pr: 2559, + issue: 2558, + }); + assert.throws(() => parseLinkArgs(["--pr", "2559"]), /--issue/); + assert.throws(() => parseLinkArgs(["--issue", "0", "--pr", "1"]), /--issue/); +}); + +/** A spawn answering: issue-id query, pr view, mutation, verify query. */ +function spawnScript({ linked = [2558], issueId = "I_issue" } = {}) { + const calls = []; + const spawn = (cmd, args) => { + calls.push(args); + const joined = args.join(" "); + let payload; + if (joined.includes("issue(number:$n){id}")) { + payload = { + data: { repository: { issue: issueId ? { id: issueId } : null } }, + }; + } else if (joined.includes("pr view")) { + payload = { id: "PR_id" }; + } else if (joined.includes("addCloseIssueReferences")) { + payload = { + data: { addCloseIssueReferences: { clientMutationId: null } }, + }; + } else if (joined.includes("closingIssuesReferences")) { + payload = { + data: { + repository: { + pullRequest: { + closingIssuesReferences: { + nodes: linked.map((number) => ({ number })), + }, + }, + }, + }, + }; + } else { + assert.fail(`unexpected gh call: ${joined}`); + } + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("main links and prints only after verifying the reference", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnScript(); + main(["--pr", "2559", "--issue", "2558"], spawn); + assert.deepEqual(lines, ["linked: PR #2559 closes #2558"]); + // The mutation got both node ids. + const mutation = spawn.calls.find((args) => + args.some((arg) => arg.includes("addCloseIssueReferences")), + ); + assert.ok(mutation.some((arg) => arg === "i=I_issue")); + assert.ok(mutation.some((arg) => arg === "p=PR_id")); +}); + +test("main throws when the issue cannot be resolved", () => { + assert.throws( + () => + main(["--pr", "2559", "--issue", "9999"], spawnScript({ issueId: null })), + /could not resolve issue #9999/, + ); +}); + +test("main refuses to report an unconfirmed link", () => { + assert.throws( + () => + main(["--pr", "2559", "--issue", "2558"], spawnScript({ linked: [17] })), + /reads \[17\]/, + ); +}); diff --git a/scripts/pr-review-fetch.mjs b/scripts/pr-review-fetch.mjs new file mode 100644 index 0000000000..16209db853 --- /dev/null +++ b/scripts/pr-review-fetch.mjs @@ -0,0 +1,105 @@ +#!/usr/bin/env node +// Fetch a Copilot review round's content (#2558) — `npm run pr:review-fetch +// -- --pr <N> [--review <REVIEW_ID>]`. Step 8 of the pr-flow skill, which +// previously transcribed this as a paginated fetch + jq block rebuilt every +// round. +// +// Without `--review` it resolves the LATEST Copilot review by `submitted_at`; +// with it, the named review is selected from the same listing, so both paths +// print the same header and body. Comments are fetched by REVIEW id — the +// unpaginated /reviews listing hides later rounds behind your own replies — +// and paginated completely, because a round you only half fetch is a round +// you only half answer. +// +// The review BODY is printed in full: the headline sentence and the +// "Suppressed comments" block live there, and a zero-comment round can still +// name a real bug in either (pr-flow 7c's definition of "clean" reads all +// three channels). + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { + REPO_SLUG, + COPILOT_REVIEWER_LOGIN, + ghPaginatedList, + requirePositiveInt, +} from "./lib/gh.mjs"; + +/** The latest Copilot-posted review in a flattened listing, or undefined. */ +export function latestCopilotReview(reviews) { + return reviews + .filter((review) => review?.user?.login === COPILOT_REVIEWER_LOGIN) + .sort((a, b) => + String(a.submitted_at ?? "").localeCompare(String(b.submitted_at ?? "")), + ) + .at(-1); +} + +/** One review comment, shaped for replying into its thread by id. */ +export function formatComment(comment) { + const line = comment.line ?? comment.original_line ?? "?"; + return `COMMENT=${comment.id} ${comment.path}:${line}\n${comment.body}`; +} + +export function parseFetchArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { pr: { type: "string" }, review: { type: "string" } }, + }); + return { + pr: requirePositiveInt(values.pr, "--pr"), + review: + values.review === undefined + ? undefined + : requirePositiveInt(values.review, "--review"), + }; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { pr, review } = parseFetchArgs(argv); + + const reviews = ghPaginatedList( + spawn, + `repos/${REPO_SLUG}/pulls/${pr}/reviews`, + ); + // Both paths resolve a full review object, so the header and body — where + // the headline and "Suppressed comments" findings live — print either way. + const selected = + review === undefined + ? latestCopilotReview(reviews) + : reviews.find((candidate) => candidate.id === review); + if (!selected) { + throw new Error( + review === undefined + ? `PR #${pr} has no Copilot review` + : `PR #${pr} has no review ${review}`, + ); + } + // The contract is "fetch a Copilot round" on both paths — a mistyped + // --review id that lands on a human review must fail, not be classified + // as the round to act on. + if (selected.user?.login !== COPILOT_REVIEWER_LOGIN) { + throw new Error( + `review ${selected.id} was posted by "${selected.user?.login ?? "(unknown)"}", not ${COPILOT_REVIEWER_LOGIN}`, + ); + } + const header = `REVIEW=${selected.id} SUBMITTED=${selected.submitted_at}\n${selected.body}`; + + const comments = ghPaginatedList( + spawn, + `repos/${REPO_SLUG}/pulls/${pr}/reviews/${selected.id}/comments`, + ); + + console.log(header); + console.log(`\n--- ${comments.length} inline comment(s) ---`); + for (const comment of comments) { + console.log(`\n${formatComment(comment)}`); + } + console.log( + `\nreply per thread: gh api repos/${REPO_SLUG}/pulls/${pr}/comments/<COMMENT>/replies -f body='…'`, + ); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/pr-review-fetch.test.mjs b/scripts/pr-review-fetch.test.mjs new file mode 100644 index 0000000000..77cee333ee --- /dev/null +++ b/scripts/pr-review-fetch.test.mjs @@ -0,0 +1,140 @@ +// Tests for scripts/pr-review-fetch.mjs (#2558) — round resolution, +// formatting, and `main()` through an injected spawn. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + formatComment, + latestCopilotReview, + main, + parseFetchArgs, +} from "./pr-review-fetch.mjs"; + +const COPILOT = "copilot-pull-request-reviewer[bot]"; +const review = (id, login, submitted) => ({ + id, + user: { login }, + submitted_at: submitted, + body: `body of ${id}`, +}); + +test("latestCopilotReview picks the newest Copilot review by submitted_at", () => { + const latest = latestCopilotReview([ + review(1, COPILOT, "2026-01-02T00:00:00Z"), + review(2, "alice", "2026-01-05T00:00:00Z"), + review(3, COPILOT, "2026-01-03T00:00:00Z"), + ]); + assert.equal(latest.id, 3); + assert.equal( + latestCopilotReview([review(2, "alice", "2026-01-05")]), + undefined, + ); +}); + +test("latestCopilotReview matches the bot login exactly, not by prefix", () => { + // A public-PR user whose login merely starts with the bot's name must not + // be read as "the Copilot round". + const latest = latestCopilotReview([ + review(1, COPILOT, "2026-01-01T00:00:00Z"), + review(2, "copilot-pull-request-reviewer", "2026-01-02T00:00:00Z"), + review(3, "copilot-pull-request-reviewer-fake", "2026-01-03T00:00:00Z"), + ]); + assert.equal(latest.id, 1); +}); + +test("formatComment names the thread id and falls back to original_line", () => { + assert.equal( + formatComment({ id: 7, path: "a.ts", line: 12, body: "b" }), + "COMMENT=7 a.ts:12\nb", + ); + assert.match( + formatComment({ id: 7, path: "a.ts", original_line: 4, body: "b" }), + /a\.ts:4/, + ); + assert.match(formatComment({ id: 7, path: "a.ts", body: "b" }), /a\.ts:\?/); +}); + +test("parseFetchArgs validates --pr and optional --review", () => { + assert.deepEqual(parseFetchArgs(["--pr", "3"]), { pr: 3, review: undefined }); + assert.equal(parseFetchArgs(["--pr", "3", "--review", "9"]).review, 9); + assert.throws(() => parseFetchArgs([]), /--pr/); +}); + +function spawnFor({ reviews, comments }) { + const calls = []; + const spawn = (cmd, args) => { + calls.push(args); + const path = args.at(-1); + const payload = path.endsWith("/reviews") ? reviews : comments; + assert.ok(payload, `unexpected gh call: ${args.join(" ")}`); + return { status: 0, stdout: JSON.stringify([payload]), stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("main resolves the latest round and prints header + every comment", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnFor({ + reviews: [review(5, COPILOT, "2026-01-01T00:00:00Z")], + comments: [ + { id: 11, path: "x.ts", line: 2, body: "first" }, + { id: 12, path: "y.ts", line: 8, body: "second" }, + ], + }); + main(["--pr", "4"], spawn); + const out = lines.join("\n"); + assert.match(out, /REVIEW=5 SUBMITTED=2026-01-01T00:00:00Z/); + assert.match(out, /body of 5/); + assert.match(out, /2 inline comment/); + assert.match(out, /COMMENT=11 x\.ts:2/); + assert.match(out, /COMMENT=12 y\.ts:8/); + // Comments were fetched by REVIEW id, not from the reviews listing. + assert.ok(spawn.calls[1].at(-1).includes("/reviews/5/comments")); +}); + +test("main with --review prints the named review's full header and body", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = spawnFor({ + reviews: [ + review(77, COPILOT, "2026-01-01T00:00:00Z"), + review(99, COPILOT, "2026-01-02T00:00:00Z"), + ], + comments: [], + }); + main(["--pr", "4", "--review", "77"], spawn); + assert.ok(spawn.calls[1].at(-1).includes("/reviews/77/comments")); + const out = lines.join("\n"); + assert.match(out, /REVIEW=77 SUBMITTED=2026-01-01T00:00:00Z/); + assert.match(out, /body of 77/); +}); + +test("main with --review throws when the review does not exist", () => { + const spawn = spawnFor({ + reviews: [review(99, COPILOT, "2026-01-02T00:00:00Z")], + }); + assert.throws( + () => main(["--pr", "4", "--review", "77"], spawn), + /no review 77/, + ); +}); + +test("main with --review refuses a review the bot did not post", () => { + // The contract is "fetch a Copilot round" — a mistyped id landing on a + // human review must fail, not be printed as the round to act on. + const spawn = spawnFor({ + reviews: [review(77, "alice", "2026-01-01T00:00:00Z")], + }); + assert.throws( + () => main(["--pr", "4", "--review", "77"], spawn), + /posted by "alice"/, + ); +}); + +test("main throws when the PR has no Copilot review", () => { + const spawn = spawnFor({ reviews: [review(2, "alice", "2026-01-01")] }); + assert.throws(() => main(["--pr", "4"], spawn), /no Copilot review/); +}); diff --git a/scripts/pr-review-request.mjs b/scripts/pr-review-request.mjs new file mode 100644 index 0000000000..edfceeadfb --- /dev/null +++ b/scripts/pr-review-request.mjs @@ -0,0 +1,63 @@ +#!/usr/bin/env node +// Request a Copilot review round on a PR (#2558) — `npm run pr:review-request +// -- --pr <N>`. Step 7a of the pr-flow skill, which previously transcribed +// this as an inline GraphQL block rebuilt from scratch on every round. +// +// Only the GraphQL `requestReviews` mutation with the Copilot BOT id works: +// REST, `gh pr edit --add-reviewer`, `userIds`, and `copilot-swe-agent` all +// fail or silently drop the request. `union:true` adds to any existing +// reviewers instead of replacing them. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { OWNER, REPO, ghGraphql, requirePositiveInt } from "./lib/gh.mjs"; + +/** + * The `copilot-pull-request-reviewer` bot's node id — last verified + * 2026-10-01. If the mutation starts rejecting it, re-resolve via the PR's + * `suggestedReviewers`/`reviewRequests` connections or the web UI's reviewer + * picker network calls; no plain lookup-by-login API exists for bots. + */ +export const COPILOT_BOT_ID = "BOT_kgDOCnlnWA"; + +/** Parse argv (everything after `node script.mjs`) into a validated PR number. */ +export function parseRequestArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { pr: { type: "string" } }, + }); + return { pr: requirePositiveInt(values.pr, "--pr") }; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { pr } = parseRequestArgs(argv); + + const prId = ghGraphql( + spawn, + `query($n:Int!){repository(owner:"${OWNER}",name:"${REPO}"){pullRequest(number:$n){id}}}`, + { n: pr }, + ).data?.repository?.pullRequest?.id; + if (!prId) { + throw new Error(`PR #${pr} not found in ${OWNER}/${REPO}`); + } + + const result = ghGraphql( + spawn, + "mutation($pr:ID!,$bot:[ID!]!){requestReviews(input:{pullRequestId:$pr, botIds:$bot, union:true}){pullRequest{id}}}", + { pr: prId, bot: COPILOT_BOT_ID }, + ); + if (!result.data?.requestReviews?.pullRequest?.id) { + throw new Error( + `requestReviews returned no pullRequest — response: ${JSON.stringify(result)}`, + ); + } + + console.log(`requested: Copilot review on PR #${pr}`); + console.log( + `next: npm run pr:review-wait -- --pr ${pr} --expected <review count to reach>`, + ); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/pr-review-request.test.mjs b/scripts/pr-review-request.test.mjs new file mode 100644 index 0000000000..3fcc529b40 --- /dev/null +++ b/scripts/pr-review-request.test.mjs @@ -0,0 +1,66 @@ +// Tests for scripts/pr-review-request.mjs (#2558) — argv validation and +// `main()` through an injected spawn, so no `gh` process is started. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + COPILOT_BOT_ID, + main, + parseRequestArgs, +} from "./pr-review-request.mjs"; + +/** A spawn that answers each `gh` call from a queue of JSON payloads. */ +function spawnQueue(payloads) { + const calls = []; + const spawn = (cmd, args) => { + calls.push(args); + const payload = payloads.shift(); + assert.ok( + payload !== undefined, + `unexpected extra gh call: ${args.join(" ")}`, + ); + return { status: 0, stdout: JSON.stringify(payload), stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("parseRequestArgs requires a positive --pr", () => { + assert.deepEqual(parseRequestArgs(["--pr", "2556"]), { pr: 2556 }); + assert.throws(() => parseRequestArgs([]), /--pr/); + assert.throws(() => parseRequestArgs(["--pr", "zero"]), /--pr/); +}); + +test("main resolves the PR id, requests the Copilot bot, and confirms", (t) => { + const spawn = spawnQueue([ + { data: { repository: { pullRequest: { id: "PR_abc" } } } }, + { data: { requestReviews: { pullRequest: { id: "PR_abc" } } } }, + ]); + const log = t.mock.method(console, "log", () => {}); + main(["--pr", "9"], spawn); + + // The mutation call carries the bot id and the resolved PR id. + const mutation = spawn.calls[1]; + assert.ok(mutation.includes(`bot=${COPILOT_BOT_ID}`)); + assert.ok(mutation.includes("pr=PR_abc")); + assert.match(mutation.join(" "), /requestReviews/); + assert.match( + log.mock.calls[0].arguments[0], + /requested: Copilot review on PR #9/, + ); +}); + +test("main throws when the PR does not exist", () => { + const spawn = spawnQueue([{ data: { repository: { pullRequest: null } } }]); + assert.throws(() => main(["--pr", "9"], spawn), /not found/); +}); + +test("main throws when the mutation returns no pullRequest", (t) => { + t.mock.method(console, "log", () => {}); + const spawn = spawnQueue([ + { data: { repository: { pullRequest: { id: "PR_abc" } } } }, + { data: { requestReviews: { pullRequest: null } } }, + ]); + assert.throws(() => main(["--pr", "9"], spawn), /no pullRequest/); +}); diff --git a/scripts/pr-review-wait.mjs b/scripts/pr-review-wait.mjs new file mode 100644 index 0000000000..72ed8c1da4 --- /dev/null +++ b/scripts/pr-review-wait.mjs @@ -0,0 +1,154 @@ +#!/usr/bin/env node +// Wait for a Copilot review round to resolve (#2558) — `npm run +// pr:review-wait -- --pr <N> --expected <K>`. Step 7b of the pr-flow skill, +// which previously transcribed this as a ~30-line bash loop rebuilt from +// scratch on every round. +// +// A round ends one of two ways: Copilot POSTS a review, or its pending +// request DISAPPEARS without one (the session failed, or occasionally it has +// nothing to say). Waiting only for the review hangs forever on the second +// case, so this watches both, plus a hard deadline. The outcome is the last +// stdout line, exactly one of: +// +// ROUND=posted — the Copilot review count reached --expected +// ROUND=ended-without-review — the pending request cleared with no review +// ROUND=timed-out — the deadline passed with the request pending +// +// `--expected` is the review COUNT to reach, not a delta: earlier rounds' +// reviews are still on the PR, so round two waits for a count of 2. All three +// outcomes exit 0 — they are answers, and the caller's decision table (pr-flow +// 7c) owns what each means. A gh/API/parse failure instead THROWS and exits +// non-zero: an error swallowed into a zero count would make a background wait +// retry blind forever, which is the exact failure the skill's inline loop +// spent half its lines defending against. +// +// Intended to run in the background (`npm run pr:review-wait … &` or an async +// shell) — per AGENTS.md "Waiting on long-running work", remote state the +// harness cannot observe is polled inside ONE backgrounded process, never one +// check per turn. On ROUND=posted, give the inline comments a further ~60s +// before fetching; they lag the review body (pr-flow step 8). + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { + OWNER, + REPO, + REPO_SLUG, + COPILOT_REVIEWER_LOGIN, + ghGraphql, + ghPaginatedList, + requirePositiveInt, +} from "./lib/gh.mjs"; + +/** Remote-API polling floor per AGENTS.md "Waiting on long-running work". */ +export const POLL_INTERVAL_MS = 30_000; +/** Rounds normally land in 2–10 minutes; 25 is the skill's historical cap. */ +export const DEFAULT_TIMEOUT_MINUTES = 25; + +/** Count the Copilot-posted reviews in a full (flattened) review listing. */ +export function copilotReviewCount(reviews) { + return reviews.filter( + (review) => review?.user?.login === COPILOT_REVIEWER_LOGIN, + ).length; +} + +/** + * The spellings under which a pending Copilot review request appears — + * matched exactly, never by substring, so another requested reviewer whose + * login merely contains "copilot" cannot keep the waiter polling after the + * real request resolved without a review. + */ +export const PENDING_COPILOT_LOGINS = new Set([ + COPILOT_REVIEWER_LOGIN, // REST login + "copilot-pull-request-reviewer", // GraphQL Bot login (no [bot] suffix) + "Copilot", // the reviewer slug the request UI/API uses +]); + +/** Count pending Copilot review requests in the GraphQL response. */ +export function pendingCopilotRequests(response) { + const nodes = response?.data?.repository?.pullRequest?.reviewRequests?.nodes; + if (!Array.isArray(nodes)) { + throw new Error( + `unexpected reviewRequests response shape: ${JSON.stringify(response)}`, + ); + } + return nodes.filter((node) => + PENDING_COPILOT_LOGINS.has(node?.requestedReviewer?.login ?? ""), + ).length; +} + +/** + * Poll until the round resolves; returns the ROUND outcome string. Deps are + * injectable for tests: `spawn` (gh), `sleep(ms)`, `now()` in ms. + */ +export async function waitForRound({ pr, expected, timeoutMinutes }, deps) { + const { spawn, sleep, now } = deps; + const deadline = now() + timeoutMinutes * 60_000; + + const count = () => + copilotReviewCount( + ghPaginatedList(spawn, `repos/${REPO_SLUG}/pulls/${pr}/reviews`), + ); + // first:100 is the connection maximum — a PR's requested-reviewer list is + // short, but a first page sized below the max could leave Copilot unseen + // on a PR with many requestees, reading as ended-without-review while the + // request is still active. + const pending = () => + pendingCopilotRequests( + ghGraphql( + spawn, + `query($n:Int!){repository(owner:"${OWNER}",name:"${REPO}"){pullRequest(number:$n){reviewRequests(first:100){nodes{requestedReviewer{... on Bot{login} ... on User{login}}}}}}}`, + { n: pr }, + ), + ); + + for (;;) { + if (count() >= expected) { + return "posted"; + } + if (pending() === 0) { + // The request can clear a beat before its review becomes visible. + await sleep(POLL_INTERVAL_MS); + return count() >= expected ? "posted" : "ended-without-review"; + } + if (now() >= deadline) { + return "timed-out"; + } + await sleep(POLL_INTERVAL_MS); + } +} + +export function parseWaitArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { + pr: { type: "string" }, + expected: { type: "string" }, + "timeout-minutes": { type: "string" }, + }, + }); + return { + pr: requirePositiveInt(values.pr, "--pr"), + expected: requirePositiveInt(values.expected, "--expected"), + timeoutMinutes: + values["timeout-minutes"] === undefined + ? DEFAULT_TIMEOUT_MINUTES + : requirePositiveInt(values["timeout-minutes"], "--timeout-minutes"), + }; +} + +export async function main( + argv = process.argv.slice(2), + deps = { + spawn: spawnSync, + sleep: (ms) => new Promise((resolve) => setTimeout(resolve, ms)), + now: Date.now, + }, +) { + const outcome = await waitForRound(parseWaitArgs(argv), deps); + console.log(`ROUND=${outcome}`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + await main(); +} diff --git a/scripts/pr-review-wait.test.mjs b/scripts/pr-review-wait.test.mjs new file mode 100644 index 0000000000..01f39e3f7e --- /dev/null +++ b/scripts/pr-review-wait.test.mjs @@ -0,0 +1,189 @@ +// Tests for scripts/pr-review-wait.mjs (#2558) — the pure counters and the +// polling orchestration in `waitForRound()`, driven through injected +// spawn/sleep/now so nothing real is polled and no time passes. Run via +// `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + POLL_INTERVAL_MS, + copilotReviewCount, + main, + parseWaitArgs, + pendingCopilotRequests, + waitForRound, +} from "./pr-review-wait.mjs"; + +const review = (login) => ({ user: { login } }); +const COPILOT = "copilot-pull-request-reviewer[bot]"; + +test("copilotReviewCount counts only Copilot reviews", () => { + assert.equal( + copilotReviewCount([review(COPILOT), review("alice"), review(COPILOT), {}]), + 2, + ); + // Exact login, not a prefix — a lookalike account must not satisfy a wait. + assert.equal( + copilotReviewCount([ + review("copilot-pull-request-reviewer"), + review("copilot-pull-request-reviewer-fake"), + ]), + 0, + ); +}); + +test("pendingCopilotRequests matches known logins exactly, throws on bad shape", () => { + const resp = (nodes) => ({ + data: { repository: { pullRequest: { reviewRequests: { nodes } } } }, + }); + assert.equal( + pendingCopilotRequests( + resp([ + { requestedReviewer: { login: "Copilot" } }, + { requestedReviewer: { login: "copilot-pull-request-reviewer" } }, + { requestedReviewer: { login: COPILOT } }, + { requestedReviewer: { login: "alice" } }, + { requestedReviewer: null }, + ]), + ), + 3, + ); + // Exact logins only — another reviewer whose login merely contains + // "copilot" must not keep the waiter polling to timeout. + assert.equal( + pendingCopilotRequests( + resp([ + { requestedReviewer: { login: "my-copilot-team" } }, + { requestedReviewer: { login: "copilot" } }, + ]), + ), + 0, + ); + assert.throws(() => pendingCopilotRequests({ data: {} }), /unexpected/); +}); + +/** + * Drives waitForRound with scripted per-poll state. `states` is consumed one + * entry per reviews-or-pending fetch pair: each entry holds the review logins + * and the pending reviewer logins the API "returns" at that point in time. + */ +function fakeDeps(timeline) { + let slept = 0; + const sleeps = []; + const spawn = (cmd, args) => { + const state = timeline[0]; + assert.ok(state, `gh call after timeline exhausted: ${args.join(" ")}`); + if (args.includes("--paginate")) { + return { + status: 0, + stdout: JSON.stringify([state.reviews.map(review)]), + stderr: "", + }; + } + // The pending lookup must request the connection maximum — a smaller + // first page could leave Copilot unseen among many requestees and read + // as ended-without-review while the request is still active. + const query = args.find((arg) => arg.startsWith("query=")); + assert.match(query, /reviewRequests\(first:100\)/); + return { + status: 0, + stdout: JSON.stringify({ + data: { + repository: { + pullRequest: { + reviewRequests: { + nodes: state.pending.map((login) => ({ + requestedReviewer: { login }, + })), + }, + }, + }, + }, + }), + stderr: "", + }; + }; + const sleep = (ms) => { + sleeps.push(ms); + slept += ms; + timeline.shift(); // time advances: next poll sees the next state + return Promise.resolve(); + }; + const now = () => slept; + return { spawn, sleep, now, sleeps }; +} + +const args = { pr: 1, expected: 2, timeoutMinutes: 25 }; + +test("posted immediately when the count is already reached", async () => { + const deps = fakeDeps([{ reviews: [COPILOT, COPILOT], pending: [] }]); + assert.equal(await waitForRound(args, deps), "posted"); + assert.deepEqual(deps.sleeps, []); +}); + +test("request cleared: recounts once after a grace sleep — posted", async () => { + const deps = fakeDeps([ + { reviews: [COPILOT], pending: [] }, + { reviews: [COPILOT, COPILOT], pending: [] }, + ]); + assert.equal(await waitForRound(args, deps), "posted"); + assert.deepEqual(deps.sleeps, [POLL_INTERVAL_MS]); +}); + +test("request cleared and no review arrives — ended-without-review", async () => { + const deps = fakeDeps([ + { reviews: [COPILOT], pending: [] }, + { reviews: [COPILOT], pending: [] }, + ]); + assert.equal(await waitForRound(args, deps), "ended-without-review"); +}); + +test("still pending: polls until the review lands — posted", async () => { + const deps = fakeDeps([ + { reviews: [COPILOT], pending: ["Copilot"] }, + { reviews: [COPILOT], pending: ["Copilot"] }, + { reviews: [COPILOT, COPILOT], pending: [] }, + ]); + assert.equal(await waitForRound(args, deps), "posted"); + assert.deepEqual(deps.sleeps, [POLL_INTERVAL_MS, POLL_INTERVAL_MS]); +}); + +test("deadline passes while the request is still pending — timed-out", async () => { + const deps = fakeDeps([{ reviews: [COPILOT], pending: ["Copilot"] }]); + assert.equal( + await waitForRound({ ...args, timeoutMinutes: 0 }, deps), + "timed-out", + ); +}); + +test("a gh failure throws instead of reading as a zero count", async () => { + const spawn = () => ({ status: 1, stdout: "", stderr: "rate limited" }); + await assert.rejects( + waitForRound(args, { spawn, sleep: () => {}, now: () => 0 }), + /rate limited/, + ); +}); + +test("parseWaitArgs validates and defaults the timeout", () => { + assert.deepEqual(parseWaitArgs(["--pr", "5", "--expected", "2"]), { + pr: 5, + expected: 2, + timeoutMinutes: 25, + }); + assert.equal( + parseWaitArgs(["--pr", "5", "--expected", "1", "--timeout-minutes", "3"]) + .timeoutMinutes, + 3, + ); + assert.throws(() => parseWaitArgs(["--pr", "5"]), /--expected/); +}); + +test("main prints the ROUND contract line", async (t) => { + const log = t.mock.method(console, "log", () => {}); + const deps = fakeDeps([ + { reviews: [COPILOT], pending: [] }, + { reviews: [COPILOT], pending: [] }, + ]); + await main(["--pr", "1", "--expected", "1"], deps); + assert.equal(log.mock.calls[0].arguments[0], "ROUND=posted"); +}); diff --git a/scripts/pr-upload-screenshot.mjs b/scripts/pr-upload-screenshot.mjs new file mode 100644 index 0000000000..715438ea44 --- /dev/null +++ b/scripts/pr-upload-screenshot.mjs @@ -0,0 +1,95 @@ +#!/usr/bin/env node +// Upload a screenshot to GitHub's user-attachments store (#2558) — `npm run +// pr:upload -- --file pr-screenshots/foo.png [--name shown-name.png]`. The +// upload block the pr-flow skill previously transcribed inline; the returned +// URL is what gets embedded in the PR body or comment. +// +// Two things the inline block learned the hard way, preserved here: the +// upload parameters go in the QUERY STRING — a JSON body fails with "Invalid +// name for request" — and the body is the file's raw bytes. The token comes +// from `gh auth token` and travels only in the Authorization header, never +// in an argv where other processes could read it. + +import { spawnSync } from "node:child_process"; +import { readFileSync } from "node:fs"; +import { basename, extname } from "node:path"; +import { parseArgs } from "node:util"; +import { REPO_SLUG, gh, ghJson } from "./lib/gh.mjs"; + +const CONTENT_TYPES = { + ".png": "image/png", + ".jpg": "image/jpeg", + ".jpeg": "image/jpeg", + ".gif": "image/gif", + ".webp": "image/webp", + ".mp4": "video/mp4", + ".mov": "video/quicktime", +}; + +export function parseUploadArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { file: { type: "string" }, name: { type: "string" } }, + }); + if (!values.file) { + throw new Error("--file <path> is required"); + } + const name = values.name ?? basename(values.file); + const contentType = CONTENT_TYPES[extname(name).toLowerCase()]; + if (!contentType) { + throw new Error( + `unsupported extension on "${name}" — known: ${Object.keys(CONTENT_TYPES).join(", ")}`, + ); + } + return { file: values.file, name, contentType }; +} + +export async function main( + argv = process.argv.slice(2), + spawn = spawnSync, + fetchFn = globalThis.fetch, +) { + const { file, name, contentType } = parseUploadArgs(argv); + + const token = gh(spawn, ["auth", "token"]); + if (token.status !== 0 || !(token.stdout ?? "").trim()) { + throw new Error("`gh auth token` returned no token — run `gh auth login`"); + } + const repositoryId = ghJson(spawn, ["api", `repos/${REPO_SLUG}`]).id; + if (!repositoryId) { + throw new Error(`could not resolve repository id for ${REPO_SLUG}`); + } + + const url = new URL("https://uploads.github.com/user-attachments/assets"); + url.searchParams.set("repository_id", String(repositoryId)); + url.searchParams.set("name", name); + url.searchParams.set("content_type", contentType); + + const response = await fetchFn(url, { + method: "POST", + headers: { + Authorization: `token ${token.stdout.trim()}`, + "Content-Type": contentType, + }, + body: readFileSync(file), + }); + const json = await response.json(); + if (!response.ok) { + throw new Error( + `upload failed (${response.status}): ${JSON.stringify(json)}`, + ); + } + // The attachment URL's field name has varied — but a success with neither + // is unusable output, not a hosted URL; reject it rather than print it. + const hosted = json.url ?? json.href; + if (!hosted) { + throw new Error( + `upload succeeded but returned no url/href: ${JSON.stringify(json)}`, + ); + } + console.log(hosted); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + await main(); +} diff --git a/scripts/pr-upload-screenshot.test.mjs b/scripts/pr-upload-screenshot.test.mjs new file mode 100644 index 0000000000..f5a38c5c9e --- /dev/null +++ b/scripts/pr-upload-screenshot.test.mjs @@ -0,0 +1,102 @@ +// Tests for scripts/pr-upload-screenshot.mjs (#2558) — the content-type map, +// the query-string parameter placement (a JSON body fails upstream), the raw +// bytes body, and the header-only token. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdtempSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { main, parseUploadArgs } from "./pr-upload-screenshot.mjs"; + +test("parseUploadArgs maps extensions and honors --name", () => { + assert.deepEqual(parseUploadArgs(["--file", "shots/connect.png"]), { + file: "shots/connect.png", + name: "connect.png", + contentType: "image/png", + }); + assert.equal( + parseUploadArgs(["--file", "a.png", "--name", "demo.mp4"]).contentType, + "video/mp4", + ); + assert.throws(() => parseUploadArgs(["--file", "notes.txt"]), /unsupported/); + assert.throws(() => parseUploadArgs([]), /--file/); +}); + +function spawnScript() { + return (cmd, args) => { + const joined = args.join(" "); + if (joined === "auth token") { + return { status: 0, stdout: "gho_secret\n", stderr: "" }; + } + if (joined.includes("repos/")) { + return { status: 0, stdout: JSON.stringify({ id: 4242 }), stderr: "" }; + } + assert.fail(`unexpected gh call: ${joined}`); + }; +} + +test("main uploads raw bytes with query-string params and prints the URL", async (t) => { + const dir = mkdtempSync(join(tmpdir(), "pr-upload-test-")); + const file = join(dir, "proof.png"); + writeFileSync(file, Buffer.from([1, 2, 3])); + + let request; + const fetchFn = async (url, init) => { + request = { url: new URL(url), init }; + return { + ok: true, + status: 201, + json: async () => ({ url: "https://x/a.png" }), + }; + }; + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + await main(["--file", file], spawnScript(), fetchFn); + + assert.deepEqual(lines, ["https://x/a.png"]); + assert.equal(request.url.searchParams.get("repository_id"), "4242"); + assert.equal(request.url.searchParams.get("name"), "proof.png"); + assert.equal(request.url.searchParams.get("content_type"), "image/png"); + assert.equal(request.init.headers.Authorization, "token gho_secret"); + assert.deepEqual([...request.init.body], [1, 2, 3]); +}); + +test("main throws on an upload rejection with the response body", async () => { + const dir = mkdtempSync(join(tmpdir(), "pr-upload-test-")); + const file = join(dir, "proof.png"); + writeFileSync(file, Buffer.from([1])); + await assert.rejects( + main(["--file", file], spawnScript(), async () => ({ + ok: false, + status: 422, + json: async () => ({ message: "Invalid name for request" }), + })), + /upload failed \(422\).*Invalid name/, + ); +}); + +test("main rejects a success response with no hosted URL", async () => { + const dir = mkdtempSync(join(tmpdir(), "pr-upload-test-")); + const file = join(dir, "proof.png"); + writeFileSync(file, Buffer.from([1])); + await assert.rejects( + main(["--file", file], spawnScript(), async () => ({ + ok: true, + status: 201, + json: async () => ({ id: 99 }), + })), + /no url\/href/, + ); +}); + +test("main throws when gh has no token", async () => { + await assert.rejects( + main(["--file", "x.png"], (cmd, args) => + args.join(" ") === "auth token" + ? { status: 1, stdout: "", stderr: "not logged in" } + : assert.fail("nothing else should run"), + ), + /no token/, + ); +}); diff --git a/scripts/release-notes.mjs b/scripts/release-notes.mjs new file mode 100644 index 0000000000..28db4f32e1 --- /dev/null +++ b/scripts/release-notes.mjs @@ -0,0 +1,376 @@ +#!/usr/bin/env node +// Assemble a release's notes and create the GitHub Release (#2550) — +// `npm run release:notes -- --merge-branch <b> --ledger-url <u> [...]`. +// The recipe the release skill's step 3a previously transcribed inline. +// +// The notes are four parts, in this order: GitHub's generated "What's +// Changed" list (the same `releases/generate-notes` API the UI button uses), +// the smoke-ledger line, a `## Known issue(s)` section from what the +// maintainer passes in (a judgment call, so never generated), and a +// `## Thanks for helping us improve` section crediting the community members +// whose issues the listed PRs close. GitHub adds everyone `@`-mentioned in a +// release body to its Contributors strip, so that section is what puts the +// reporters there. +// +// FAIL FAST. Every `gh`/`git` call is checked and a failure throws: a +// permission lookup that errors must never read as "community" (that could +// credit a maintainer), and a rate-limited PR or issue lookup must never be +// silently skipped (that would publish a partial Thanks list). A permission +// value outside the known set throws for the same reason. +// +// The default run is a PREVIEW that prints the notes and creates nothing. +// Publishing the Release triggers `package` → `publish` (npm) and the GHCR +// image, so it is never implicit: `--draft` creates a draft for review and +// `--publish` publishes. Both pass the bare `x.y.z` tag and `--target main`, +// and both refuse unless that version is the one on origin/main. + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; +import { versionFrom } from "./release-tag.mjs"; + +export const REPO = "modelcontextprotocol/inspector"; +const [OWNER, NAME] = REPO.split("/"); + +// What `collaborators/{user}/permission` reports in its `permission` field. +// A public repo reports `read` for anyone without a role, so a community +// member is `read` (or `triage`/`none`), and a maintainer is anyone above it. +const MAINTAINER_PERMISSIONS = new Set(["admin", "maintain", "write"]); +const COMMUNITY_PERMISSIONS = new Set(["triage", "read", "none"]); + +export const THANKS_LEAD_IN = + "This release addresses issues reported by these community members. Thank you for taking the time to file them:"; + +export function parseNotesArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { + "merge-branch": { type: "string" }, + "ledger-url": { type: "string" }, + "known-issue": { type: "string", multiple: true }, + version: { type: "string" }, + "previous-tag": { type: "string" }, + draft: { type: "boolean" }, + publish: { type: "boolean" }, + }, + }); + const mergeBranch = values["merge-branch"]; + const ledgerUrl = values["ledger-url"]; + if (!mergeBranch || !ledgerUrl) { + throw new Error("--merge-branch and --ledger-url are both required"); + } + if (!/^https:\/\/\S+$/.test(ledgerUrl)) { + throw new Error(`--ledger-url "${ledgerUrl}" is not an https URL`); + } + for (const [flag, tag] of [ + ["--version", values.version], + ["--previous-tag", values["previous-tag"]], + ]) { + if (tag !== undefined && !isStableTag(tag)) { + throw new Error(`${flag} "${tag}" is not a bare x.y.z`); + } + } + if (values.draft && values.publish) { + throw new Error("--draft and --publish are mutually exclusive"); + } + // A Release created here always starts from the previous stable tag; an + // override (a typo, a stale value) would publish the wrong range of + // changes and credits, so it is a preview-only knob like --version. + if ((values.draft || values.publish) && values["previous-tag"]) { + throw new Error("--previous-tag is preview-only"); + } + return { + mergeBranch, + ledgerUrl, + knownIssues: values["known-issue"] ?? [], + version: values.version, + previousTag: values["previous-tag"], + mode: values.publish ? "publish" : values.draft ? "draft" : "preview", + }; +} + +function run(spawn, cmd, args, input) { + const result = spawn(cmd, args, { encoding: "utf8", input }); + if (result.error) { + throw result.error; + } + if (result.status !== 0) { + throw new Error( + `${cmd} ${args.join(" ")} failed: ${(result.stderr ?? "").trim()}`, + ); + } + return (result.stdout ?? "").trim(); +} + +function graphql(spawn, query, variables) { + const args = ["api", "graphql", "-f", `query=${query}`]; + for (const [key, value] of Object.entries(variables)) { + // -F types a number as Int; -f keeps a cursor a String. + args.push(typeof value === "number" ? "-F" : "-f", `${key}=${value}`); + } + const response = JSON.parse(run(spawn, "gh", args)); + if (response.errors?.length) { + throw new Error(`gh api graphql: ${JSON.stringify(response.errors)}`); + } + return response.data; +} + +export function isStableTag(tag) { + return /^\d+\.\d+\.\d+$/.test(tag); +} + +function compareVersions(a, b) { + const pa = a.split(".").map(Number); + const pb = b.split(".").map(Number); + for (let i = 0; i < 3; i++) { + if (pa[i] !== pb[i]) return pa[i] - pb[i]; + } + return 0; +} + +/** + * The highest STABLE tag below `version`. Whole-name match: this repo also + * has 2.0.0-rc.N, x.y.z-hotfix, x.y.z-amended and v2-alpha-1 tags, and + * picking one of those as the previous tag would drop changes from the list. + */ +export function previousStableTag(tags, version) { + const below = tags + .filter(isStableTag) + .filter((tag) => compareVersions(tag, version) < 0) + .sort(compareVersions); + if (below.length === 0) { + throw new Error(`no stable x.y.z tag below ${version}`); + } + return below[below.length - 1]; +} + +/** Every PR of this repo the generated list links to, ascending. */ +export function pullNumbersFrom(generated) { + const pattern = new RegExp(`https://github\\.com/${REPO}/pull/(\\d+)`, "g"); + return [...new Set([...generated.matchAll(pattern)].map((m) => +m[1]))].sort( + (a, b) => a - b, + ); +} + +/** + * Issue numbers a PR body closes by keyword — GitHub's nine closing keywords, + * with its optional colon. A bare `#N` only, so a cross-repo `owner/repo#N` + * is not mistaken for one of ours. + */ +export function closingKeywordIssues(body) { + const pattern = /\b(?:close[sd]?|fix(?:e[sd])?|resolve[sd]?):?\s+#(\d+)\b/gi; + return [...proseOf(body ?? "").matchAll(pattern)].map((m) => +m[1]); +} + +/** + * The body with every place GitHub ignores a closing keyword blanked out: + * HTML comments (PR templates leave `<!-- Closes #… -->` behind), fenced + * code, inline code and blockquotes. Four-space indented code is left + * alone: telling it from an indented list continuation needs a full + * Markdown parser, and masking a real `Closes` line would lose a credit. A keyword quoted in any of those does + * not close the issue, so it must not credit the issue's author either. + */ +export function proseOf(body) { + return ( + body + .replace(/<!--[\s\S]*?(?:-->|$)/g, " ") + // A closing fence is the opening fence's character, at least as many, + // and nothing after it but spaces or tabs; anything else is content. + .replace( + /^ {0,3}((`|~)\2{2,})[^\n]*\n[\s\S]*?(?:^ {0,3}\1\2*[ \t]*$|(?![\s\S]))/gm, + " ", + ) + .replace(/(`+)[\s\S]*?\1/g, " ") + .replace(/^ {0,3}>.*$/gm, " ") + ); +} + +const CLOSING_QUERY = `query($n:Int!,$after:String){repository(owner:"${OWNER}",name:"${NAME}"){pullRequest(number:$n){body closingIssuesReferences(first:100,after:$after){pageInfo{hasNextPage endCursor} nodes{number repository{nameWithOwner}}}}}}`; + +/** Every issue of this repo one PR closes: manual links (paginated) + body keywords. */ +export function issuesClosedBy(spawn, pr) { + const numbers = new Set(); + let body; + let after; + for (;;) { + const variables = after === undefined ? { n: pr } : { n: pr, after }; + const pull = graphql(spawn, CLOSING_QUERY, variables).repository + .pullRequest; + if (!pull) { + throw new Error(`PR #${pr} not found`); + } + body ??= pull.body; + const refs = pull.closingIssuesReferences; + for (const node of refs.nodes) { + if (node.repository.nameWithOwner === REPO) numbers.add(node.number); + } + if (!refs.pageInfo.hasNextPage) break; + after = refs.pageInfo.endCursor; + } + for (const n of closingKeywordIssues(body)) numbers.add(n); + return numbers; +} + +const AUTHOR_QUERY = `query($n:Int!){repository(owner:"${OWNER}",name:"${NAME}"){issueOrPullRequest(number:$n){__typename ... on Issue{author{login __typename}}}}}`; + +/** + * The login to credit for issue `n`, or null when there is nobody to credit: + * the number is a PR rather than an issue, or the author is a bot (the + * SDK-watch and Dependabot sweeps) or a deleted account. + */ +export function creditableAuthor(spawn, n) { + const node = graphql(spawn, AUTHOR_QUERY, { n }).repository + .issueOrPullRequest; + if (!node) { + throw new Error(`#${n} not found`); + } + if (node.__typename !== "Issue") return null; + return node.author?.__typename === "User" ? node.author.login : null; +} + +/** True for a maintainer, false for a community member; throws otherwise. */ +export function isMaintainer(spawn, login) { + const permission = run(spawn, "gh", [ + "api", + `repos/${REPO}/collaborators/${login}/permission`, + "--jq", + ".permission", + ]); + if (MAINTAINER_PERMISSIONS.has(permission)) return true; + if (COMMUNITY_PERMISSIONS.has(permission)) return false; + throw new Error(`unexpected permission "${permission}" for @${login}`); +} + +/** Community reporter → the issues of theirs the listed PRs close. */ +export function collectReporters(spawn, pulls) { + const issues = new Set(); + for (const pr of pulls) { + for (const n of issuesClosedBy(spawn, pr)) issues.add(n); + } + const reporters = new Map(); + const maintainer = new Map(); + for (const n of [...issues].sort((a, b) => a - b)) { + const login = creditableAuthor(spawn, n); + if (login === null) continue; + if (!maintainer.has(login)) { + maintainer.set(login, isMaintainer(spawn, login)); + } + if (maintainer.get(login)) continue; + if (!reporters.has(login)) reporters.set(login, []); + reporters.get(login).push(n); + } + return reporters; +} + +/** One line per person, most issues first, then by name; "" when nobody. */ +export function formatThanks(reporters) { + if (reporters.size === 0) return ""; + const lines = [...reporters] + .sort( + ([a, ia], [b, ib]) => + ib.length - ia.length || + a.localeCompare(b, "en", { sensitivity: "base" }), + ) + .map( + ([login, nums]) => `* @${login} (${nums.map((n) => `#${n}`).join(", ")})`, + ); + return `## Thanks for helping us improve\n\n${THANKS_LEAD_IN}\n\n${lines.join("\n")}`; +} + +export function formatKnownIssues(knownIssues) { + if (knownIssues.length === 0) return ""; + const heading = knownIssues.length === 1 ? "Known issue" : "Known issues"; + return `## ${heading}\n\n${knownIssues.join("\n\n")}`; +} + +export function assembleNotes({ + generated, + mergeBranch, + ledgerUrl, + knownIssues, + thanks, +}) { + const ledger = `**Smoke test ledger for milestone branch**: [${mergeBranch}](${ledgerUrl})`; + const sections = [formatKnownIssues(knownIssues), thanks].filter(Boolean); + return [`${generated.trimEnd()}\n${ledger}`, ...sections].join("\n\n") + "\n"; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const args = parseNotesArgs(argv); + + // Tags first: a refspec-less fetch rewrites FETCH_HEAD, so the main fetch + // has to come last for FETCH_HEAD to be main (see release-tag.mjs on why + // FETCH_HEAD and not the origin/main tracking ref). + run(spawn, "git", ["fetch", "origin", "--tags"]); + run(spawn, "git", ["fetch", "origin", "main"]); + const sha = run(spawn, "git", ["rev-parse", "FETCH_HEAD"]); + const mainVersion = versionFrom( + run(spawn, "git", ["show", `${sha}:package.json`]), + ); + const version = args.version ?? mainVersion; + if (args.mode !== "preview" && version !== mainVersion) { + // --target main would attach this version's Release to another + // version's tree; regenerating an older release's notes is preview-only. + throw new Error( + `--${args.mode} creates ${version} on main, but origin/main is ${mainVersion}`, + ); + } + const previousTag = + args.previousTag ?? + previousStableTag(run(spawn, "git", ["tag", "-l"]).split("\n"), version); + console.error(`release notes: ${previousTag} → ${version}`); + + const generated = run(spawn, "gh", [ + "api", + `repos/${REPO}/releases/generate-notes`, + "-f", + `tag_name=${version}`, + "-f", + "target_commitish=main", + "-f", + `previous_tag_name=${previousTag}`, + "--jq", + ".body", + ]); + const pulls = pullNumbersFrom(generated); + console.error(`crediting reporters of issues closed by ${pulls.length} PRs`); + const notes = assembleNotes({ + generated, + mergeBranch: args.mergeBranch, + ledgerUrl: args.ledgerUrl, + knownIssues: args.knownIssues, + thanks: formatThanks(collectReporters(spawn, pulls)), + }); + + if (args.mode === "preview") { + process.stdout.write(notes); + console.error( + "preview only — nothing created; re-run with --draft (or --publish)", + ); + return notes; + } + const create = [ + "release", + "create", + version, + "--repo", + REPO, + "--target", + "main", + "--title", + version, + "--notes-file", + "-", + args.mode === "draft" ? "--draft" : "--latest", + ]; + const url = run(spawn, "gh", create, notes); + console.error( + args.mode === "draft" + ? `draft created: ${url} — review it, then publish (this triggers npm + GHCR)` + : `published: ${url}`, + ); + return notes; +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/release-notes.test.mjs b/scripts/release-notes.test.mjs new file mode 100644 index 0000000000..356cabbceb --- /dev/null +++ b/scripts/release-notes.test.mjs @@ -0,0 +1,536 @@ +// Tests for scripts/release-notes.mjs (#2550) — the issue-author mapping and +// its exclusions (maintainers by permission, bots, PRs, other repos), the +// paginated closing references, fail-fast on every API error, the note +// layout 2.9.0 shipped with, and that nothing is created without an explicit +// --draft / --publish. Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + REPO, + THANKS_LEAD_IN, + assembleNotes, + closingKeywordIssues, + formatKnownIssues, + formatThanks, + main, + parseNotesArgs, + previousStableTag, + pullNumbersFrom, +} from "./release-notes.mjs"; + +const BASE = [ + "--merge-branch", + "v2/chore/mm", + "--ledger-url", + "https://l.example/x", +]; +const SHA = "abc123"; +const PULL = (n) => `https://github.com/${REPO}/pull/${n}`; + +test("parseNotesArgs defaults to a preview and validates its inputs", () => { + assert.deepEqual(parseNotesArgs(BASE), { + mergeBranch: "v2/chore/mm", + ledgerUrl: "https://l.example/x", + knownIssues: [], + version: undefined, + previousTag: undefined, + mode: "preview", + }); + assert.equal(parseNotesArgs([...BASE, "--draft"]).mode, "draft"); + assert.equal(parseNotesArgs([...BASE, "--publish"]).mode, "publish"); + assert.deepEqual( + parseNotesArgs([...BASE, "--known-issue", "a", "--known-issue", "b"]) + .knownIssues, + ["a", "b"], + ); + assert.throws(() => parseNotesArgs([]), /both required/); + assert.throws( + () => parseNotesArgs(["--merge-branch", "b", "--ledger-url", "ftp://x"]), + /not an https URL/, + ); + assert.throws( + () => parseNotesArgs([...BASE, "--draft", "--publish"]), + /mutually exclusive/, + ); + assert.throws( + () => parseNotesArgs([...BASE, "--version", "v2.9.0"]), + /--version "v2.9.0" is not a bare x.y.z/, + ); + assert.throws( + () => parseNotesArgs([...BASE, "--previous-tag", "2.0.0-rc.1"]), + /--previous-tag/, + ); + for (const mode of ["--draft", "--publish"]) { + assert.throws( + () => parseNotesArgs([...BASE, "--previous-tag", "2.8.0", mode]), + /--previous-tag is preview-only/, + ); + } + assert.equal( + parseNotesArgs([...BASE, "--previous-tag", "2.8.0"]).previousTag, + "2.8.0", + ); +}); + +test("previousStableTag skips RC, hotfix and v-prefixed tags", () => { + const tags = [ + "2.8.0", + "2.9.0", + "2.10.0-rc.1", + "2.9.1-hotfix", + "v2-alpha-1", + "2.10.0", + "", + ]; + assert.equal(previousStableTag(tags, "2.10.0"), "2.9.0"); + assert.equal(previousStableTag(tags, "2.9.0"), "2.8.0"); + assert.equal(previousStableTag(tags, "2.11.0"), "2.10.0"); + assert.throws(() => previousStableTag(tags, "1.0.0"), /no stable/); +}); + +test("pullNumbersFrom takes this repo's PR links only, deduplicated", () => { + const generated = [ + `* a by @x in ${PULL(12)}`, + `* b by @y in ${PULL(3)}`, + `* @y made their first contribution in ${PULL(3)}`, + "* c in https://github.com/other/repo/pull/7", + ].join("\n"); + assert.deepEqual(pullNumbersFrom(generated), [3, 12]); +}); + +test("closingKeywordIssues reads every closing keyword, never a cross-repo ref", () => { + const body = [ + "Closes #1", + "fixes: #2, Resolved #3", + "close #4 and FIXED #5", + "Refs #6", + "Closes other/repo#7", + "prefixes #8", + ].join("\n"); + assert.deepEqual(closingKeywordIssues(body), [1, 2, 3, 4, 5]); + assert.deepEqual(closingKeywordIssues(null), []); +}); + +test("closingKeywordIssues ignores keywords GitHub ignores: code, quotes, comments", () => { + const body = [ + "Closes #1", + "<!-- Closes #2 -->", + "<!--", + "Fixes #3", + "-->", + "Documents `Closes #4` and ``fixes #5``.", + "```sh", + "Closes #6", + "```", + "~~~", + "resolves #7", + "~~~", + "> Closes #8", + " > fixes #9", + "Resolves #10", + ].join("\n"); + assert.deepEqual(closingKeywordIssues(body), [1, 10]); + // An unterminated fence or comment swallows the rest, as GitHub renders it. + assert.deepEqual(closingKeywordIssues("Closes #1\n```\nCloses #2"), [1]); + assert.deepEqual(closingKeywordIssues("Closes #1\n<!-- Closes #2"), [1]); + // A "fence" with text after it does not close the block; a longer one does. + assert.deepEqual( + closingKeywordIssues("```\n```sh\nCloses #2\n`````\nCloses #3"), + [3], + ); + assert.deepEqual( + closingKeywordIssues("~~~\n```\nCloses #2\n~~~ \t\nFixes #3"), + [3], + ); +}); + +test("formatThanks orders by issue count, then name case-insensitively", () => { + const reporters = new Map([ + ["zed", [5]], + ["Amy", [9]], + ["many", [1, 2, 3]], + ["bob", [4]], + ]); + assert.equal( + formatThanks(reporters), + [ + "## Thanks for helping us improve", + "", + THANKS_LEAD_IN, + "", + "* @many (#1, #2, #3)", + "* @Amy (#9)", + "* @bob (#4)", + "* @zed (#5)", + ].join("\n"), + ); + assert.equal(formatThanks(new Map()), ""); +}); + +test("formatKnownIssues pluralizes and is omitted when empty", () => { + assert.equal(formatKnownIssues([]), ""); + assert.equal(formatKnownIssues(["one"]), "## Known issue\n\none"); + assert.equal(formatKnownIssues(["a", "b"]), "## Known issues\n\na\n\nb"); +}); + +test("assembleNotes lays the parts out as 2.9.0 shipped them", () => { + const notes = assembleNotes({ + generated: "## What's Changed\n* x\n\n**Full Changelog**: c\n", + mergeBranch: "mm", + ledgerUrl: "https://l", + knownIssues: ["K"], + thanks: "## Thanks for helping us improve\n\nT", + }); + assert.equal( + notes, + [ + "## What's Changed", + "* x", + "", + "**Full Changelog**: c", + "**Smoke test ledger for milestone branch**: [mm](https://l)", + "", + "## Known issue", + "", + "K", + "", + "## Thanks for helping us improve", + "", + "T", + "", + ].join("\n"), + ); + assert.equal( + assembleNotes({ + generated: "G", + mergeBranch: "mm", + ledgerUrl: "https://l", + knownIssues: [], + thanks: "", + }), + "G\n**Smoke test ledger for milestone branch**: [mm](https://l)\n", + ); +}); + +// A fake `gh`/`git` over a small repo model. `failOn` makes the first call +// whose joined argv includes it fail, the way a rate limit or a 404 would. +function world({ + mainVersion = "2.10.0", + tags = ["2.8.0", "2.9.0", "2.10.0-rc.1"], + pulls = {}, + issues = {}, + perms = {}, + failOn, +} = {}) { + const calls = []; + const generated = [ + "## What's Changed", + ...Object.keys(pulls).map((n) => `* change by @dev in ${PULL(n)}`), + "", + "**Full Changelog**: https://example/compare", + ].join("\n"); + const ok = (stdout) => ({ status: 0, stdout, stderr: "" }); + const spawn = (cmd, args, opts) => { + calls.push({ cmd, args, input: opts?.input }); + const joined = args.join(" "); + if (failOn && joined.includes(failOn)) { + return { status: 1, stdout: "", stderr: "API rate limit exceeded" }; + } + if (cmd === "git") { + if (joined === "rev-parse FETCH_HEAD") return ok(`${SHA}\n`); + if (joined === `show ${SHA}:package.json`) + return ok(JSON.stringify({ version: mainVersion })); + if (joined === "tag -l") return ok(tags.join("\n")); + return ok(""); + } + assert.equal(cmd, "gh"); + if (args[1].endsWith("/releases/generate-notes")) return ok(generated); + if (args[0] === "release") return ok("https://github.com/r/releases/1"); + if (args[1] === "graphql") { + const vars = Object.fromEntries( + args + .filter((_, i) => args[i - 1] === "-F" || args[i - 1] === "-f") + .map((kv) => kv.split(/=(.*)/s).slice(0, 2)), + ); + const n = Number(vars.n); + if (vars.query.includes("closingIssuesReferences")) { + const pr = pulls[n]; + const pages = pr.pages ?? [pr.closing ?? []]; + const index = vars.after ? Number(vars.after) : 0; + const hasNextPage = index + 1 < pages.length; + return ok( + JSON.stringify({ + data: { + repository: { + pullRequest: { + body: pr.body ?? "", + closingIssuesReferences: { + pageInfo: { + hasNextPage, + endCursor: hasNextPage ? String(index + 1) : null, + }, + nodes: pages[index].map((ref) => + typeof ref === "number" + ? { number: ref, repository: { nameWithOwner: REPO } } + : ref, + ), + }, + }, + }, + }, + }), + ); + } + return ok( + JSON.stringify({ + data: { repository: { issueOrPullRequest: issues[n] ?? null } }, + }), + ); + } + const login = args[1].match(/collaborators\/([^/]+)\/permission/)[1]; + return ok(perms[login]); + }; + spawn.calls = calls; + return spawn; +} + +const user = (login) => ({ + __typename: "Issue", + author: { login, __typename: "User" }, +}); + +function quiet(t) { + const out = []; + t.mock.method(console, "error", () => {}); + t.mock.method(process.stdout, "write", (chunk) => { + out.push(chunk); + return true; + }); + return out; +} + +test("credits community reporters only — no maintainers, bots, PRs or other repos", (t) => { + const out = quiet(t); + const spawn = world({ + pulls: { + 10: { body: "Closes #1\nFixes #2", closing: [3] }, + 11: { + body: "Resolves #4, closes #5, closes #6, closes #1", + closing: [{ number: 7, repository: { nameWithOwner: "other/repo" } }], + }, + }, + issues: { + 1: user("reporter"), + 2: user("maint"), + 3: user("reporter"), + 4: { __typename: "Issue", author: { login: "bot", __typename: "Bot" } }, + 5: { __typename: "PullRequest" }, + 6: { __typename: "Issue", author: null }, + }, + perms: { reporter: "read", maint: "write" }, + }); + const notes = main(BASE, spawn); + assert.match(notes, /\* @reporter \(#1, #3\)\n$/); + assert.doesNotMatch(notes, /@maint|@bot|#5|#7/); + assert.equal(out.join(""), notes); + // Each person's permission is looked up once, however many issues they filed. + const permissionCalls = spawn.calls.filter((c) => + c.args[1]?.includes("/permission"), + ); + assert.deepEqual( + permissionCalls.map((c) => c.args[1].split("/")[4]), + ["reporter", "maint"], + ); + // The preview creates nothing. + assert.equal( + spawn.calls.some((c) => c.args[0] === "release"), + false, + ); +}); + +test("generate-notes is asked for the previous STABLE tag to main", (t) => { + quiet(t); + const spawn = world(); + main(BASE, spawn); + const call = spawn.calls.find((c) => + c.args[1]?.endsWith("/releases/generate-notes"), + ); + assert.deepEqual(call.args.slice(2), [ + "-f", + "tag_name=2.10.0", + "-f", + "target_commitish=main", + "-f", + "previous_tag_name=2.9.0", + "--jq", + ".body", + ]); + // The tags fetch precedes the main fetch, so FETCH_HEAD is main. + const fetches = spawn.calls.filter((c) => c.args[0] === "fetch"); + assert.deepEqual( + fetches.map((c) => c.args.join(" ")), + ["fetch origin --tags", "fetch origin main"], + ); +}); + +test("an explicit --version/--previous-tag regenerates an older release's notes", (t) => { + quiet(t); + const spawn = world({ mainVersion: "2.10.0" }); + main([...BASE, "--version", "2.8.0", "--previous-tag", "2.7.0"], spawn); + const call = spawn.calls.find((c) => + c.args[1]?.endsWith("/releases/generate-notes"), + ); + assert.ok(call.args.includes("tag_name=2.8.0")); + assert.ok(call.args.includes("previous_tag_name=2.7.0")); + assert.equal( + spawn.calls.some((c) => c.args.join(" ") === "tag -l"), + false, + ); +}); + +test("the Thanks section is omitted when no community reporter remains", (t) => { + quiet(t); + const notes = main( + BASE, + world({ + pulls: { 10: { body: "Closes #1" } }, + issues: { 1: user("maint") }, + perms: { maint: "admin" }, + }), + ); + assert.doesNotMatch(notes, /Thanks/); +}); + +test("closing references are followed across every page", (t) => { + quiet(t); + const spawn = world({ + pulls: { 10: { pages: [[1], [2], [3]] } }, + issues: { 1: user("a"), 2: user("b"), 3: user("c") }, + perms: { a: "read", b: "triage", c: "none" }, + }); + const notes = main(BASE, spawn); + assert.match(notes, /@a \(#1\)\n\* @b \(#2\)\n\* @c \(#3\)/); + const cursors = spawn.calls + .filter((c) => c.args.some((a) => a.includes("closingIssuesReferences"))) + .map((c) => c.args.find((a) => a.startsWith("after=")) ?? null); + assert.deepEqual(cursors, [null, "after=1", "after=2"]); +}); + +for (const [what, failOn] of [ + ["a permission lookup", "/permission"], + ["a closing-references lookup", "closingIssuesReferences"], + ["an issue-author lookup", "issueOrPullRequest"], + ["generate-notes", "generate-notes"], + ["the fetch", "fetch origin main"], +]) { + test(`a failed ${what} aborts — never a partial Thanks list`, (t) => { + quiet(t); + const spawn = world({ + pulls: { 10: { body: "Closes #1" } }, + issues: { 1: user("maybe-maint") }, + perms: { "maybe-maint": "read" }, + failOn, + }); + assert.throws( + () => main([...BASE, "--publish"], spawn), + /rate limit exceeded/, + ); + assert.equal( + spawn.calls.some((c) => c.args[0] === "release"), + false, + ); + }); +} + +test("an unknown permission value aborts rather than reading as community", (t) => { + quiet(t); + assert.throws( + () => + main( + BASE, + world({ + pulls: { 10: { body: "Closes #1" } }, + issues: { 1: user("who") }, + perms: { who: "" }, + }), + ), + /unexpected permission "" for @who/, + ); +}); + +test("GraphQL errors and missing nodes abort", (t) => { + quiet(t); + const missing = world({ pulls: { 10: { body: "Closes #99" } } }); + assert.throws(() => main(BASE, missing), /#99 not found/); + + const errors = world({ pulls: { 10: {} } }); + const inner = errors; + const spawn = (cmd, args, opts) => + args[1] === "graphql" + ? { status: 0, stdout: '{"errors":[{"message":"boom"}]}', stderr: "" } + : inner(cmd, args, opts); + assert.throws(() => main(BASE, spawn), /boom/); + + const noPull = (cmd, args, opts) => + args[1] === "graphql" + ? { + status: 0, + stdout: '{"data":{"repository":{"pullRequest":null}}}', + stderr: "", + } + : inner(cmd, args, opts); + assert.throws(() => main(BASE, noPull), /PR #10 not found/); +}); + +test("a spawn error is rethrown", () => { + const boom = new Error("ENOENT"); + assert.throws( + () => main(BASE, () => ({ error: boom })), + (e) => e === boom, + ); +}); + +test("--draft creates a draft at main under the bare tag, notes on stdin", (t) => { + quiet(t); + const spawn = world(); + const notes = main([...BASE, "--draft"], spawn); + const create = spawn.calls.find((c) => c.args[0] === "release"); + assert.deepEqual(create.args, [ + "release", + "create", + "2.10.0", + "--repo", + REPO, + "--target", + "main", + "--title", + "2.10.0", + "--notes-file", + "-", + "--draft", + ]); + assert.equal(create.input, notes); +}); + +test("--publish publishes as latest", (t) => { + quiet(t); + const spawn = world(); + main([...BASE, "--publish"], spawn); + const create = spawn.calls.find((c) => c.args[0] === "release"); + assert.equal(create.args.at(-1), "--latest"); + assert.equal(create.args.includes("--draft"), false); +}); + +test("creating a version other than origin/main's is refused", (t) => { + quiet(t); + const spawn = world({ mainVersion: "2.10.0" }); + assert.throws( + () => main([...BASE, "--version", "2.9.0", "--draft"], spawn), + /--draft creates 2\.9\.0 on main, but origin\/main is 2\.10\.0/, + ); + assert.equal( + spawn.calls.some((c) => c.cmd === "gh"), + false, + ); +}); diff --git a/scripts/release-tag.mjs b/scripts/release-tag.mjs new file mode 100644 index 0000000000..59a5464b6e --- /dev/null +++ b/scripts/release-tag.mjs @@ -0,0 +1,82 @@ +#!/usr/bin/env node +// Tag a release on origin/main (#2558) — `npm run release:tag [-- --push]`. +// The tag block the release skill previously transcribed inline. +// +// The hazard this exists for: the tag must point at `origin/main`'s commit, +// NEVER at a local HEAD — a local branch that is ahead or behind tags the +// wrong tree. So the SHA and the version both come from `origin/main` after +// an explicit fetch, and the default run is a DRY RUN that prints what would +// be tagged; `--push` is the human gate on the decision, not on the +// composition. +// +// The tag is the BARE `x.y.z` — this repo's release tags carry no `v` prefix +// (the release skill's rule; npm's own default would have minted `vx.y.z`). + +import { spawnSync } from "node:child_process"; +import { parseArgs } from "node:util"; + +export function parseTagArgs(argv) { + const { values } = parseArgs({ + args: argv, + options: { push: { type: "boolean" } }, + }); + return { push: values.push === true }; +} + +function git(spawn, args) { + const result = spawn("git", args, { encoding: "utf8" }); + if (result.error) { + throw result.error; + } + if (result.status !== 0) { + throw new Error( + `git ${args.join(" ")} failed: ${(result.stderr ?? "").trim()}`, + ); + } + return (result.stdout ?? "").trim(); +} + +/** The version from a raw package.json, validated as x.y.z. */ +export function versionFrom(packageJson) { + const version = JSON.parse(packageJson).version; + if (typeof version !== "string" || !/^\d+\.\d+\.\d+$/.test(version)) { + throw new Error( + `origin/main package.json version "${version}" is not x.y.z`, + ); + } + return version; +} + +export function main(argv = process.argv.slice(2), spawn = spawnSync) { + const { push } = parseTagArgs(argv); + + git(spawn, ["fetch", "origin", "main"]); + // Read the SHA from FETCH_HEAD — what that fetch literally just retrieved. + // `rev-parse origin/main` would read the tracking ref, which a source-only + // `fetch origin main` updates only opportunistically — a stale tracking + // ref could be tagged while the output claims an explicit fetch made it + // current. Resolve the SHA once and read everything else FROM that SHA: + // worktrees share refs, so a fetch elsewhere can move things between + // commands — reading package.json off a ref name could pair commit A's + // version with commit B's tag target, the exact mismatch this helper + // exists to prevent. + const sha = git(spawn, ["rev-parse", "FETCH_HEAD"]); + const version = versionFrom(git(spawn, ["show", `${sha}:package.json`])); + const tag = version; // bare x.y.z — no v prefix on this repo's release tags + + if (!push) { + console.log( + `would tag: ${tag} → ${sha} (origin/main) — re-run with --push`, + ); + return; + } + // Push the ref directly from the SHA — no local tag is created, so a + // failed push leaves nothing behind and a retry starts clean (a local + // `git tag` first would make the retry fail with "already exists"). + git(spawn, ["push", "origin", `${sha}:refs/tags/${tag}`]); + console.log(`tagged: ${tag} → ${sha}`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main(); +} diff --git a/scripts/release-tag.test.mjs b/scripts/release-tag.test.mjs new file mode 100644 index 0000000000..865426eaf6 --- /dev/null +++ b/scripts/release-tag.test.mjs @@ -0,0 +1,76 @@ +// Tests for scripts/release-tag.mjs (#2558) — both the version and the SHA +// come from origin/main after an explicit fetch (never a local HEAD), the +// default run is a dry run, --push tags that exact SHA, and the tag is the +// bare x.y.z (no v prefix). Run via `npm run test:scripts`. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { main, parseTagArgs, versionFrom } from "./release-tag.mjs"; + +test("parseTagArgs defaults to a dry run", () => { + assert.deepEqual(parseTagArgs([]), { push: false }); + assert.deepEqual(parseTagArgs(["--push"]), { push: true }); +}); + +test("versionFrom validates the shape", () => { + assert.equal(versionFrom('{"version":"2.4.1"}'), "2.4.1"); + assert.throws(() => versionFrom('{"version":"2.4"}'), /not x\.y\.z/); + assert.throws(() => versionFrom("{}"), /not x\.y\.z/); +}); + +const SHA = "abc123def456"; + +function gitSpawn() { + const calls = []; + const spawn = (cmd, args) => { + assert.equal(cmd, "git"); + calls.push(args); + const joined = args.join(" "); + let stdout = ""; + if (joined === `show ${SHA}:package.json`) { + stdout = '{"version":"2.4.1"}'; + } else if (joined === "rev-parse FETCH_HEAD") { + stdout = `${SHA}\n`; + } + return { status: 0, stdout, stderr: "" }; + }; + spawn.calls = calls; + return spawn; +} + +test("the default run fetches, prints what it would tag, and tags nothing", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = gitSpawn(); + main([], spawn); + assert.deepEqual(lines, [ + `would tag: 2.4.1 → ${SHA} (origin/main) — re-run with --push`, + ]); + assert.deepEqual(spawn.calls[0], ["fetch", "origin", "main"]); + assert.equal( + spawn.calls.some((args) => args[0] === "tag" || args[0] === "push"), + false, + ); +}); + +test("--push pushes origin/main's SHA as the tag ref — no local tag", (t) => { + const lines = []; + t.mock.method(console, "log", (line) => lines.push(line)); + const spawn = gitSpawn(); + main(["--push"], spawn); + assert.deepEqual(lines, [`tagged: 2.4.1 → ${SHA}`]); + // The ref is pushed directly from the SHA, so a failed push leaves no + // local tag behind and the retry starts clean. + assert.ok(!spawn.calls.some((args) => args[0] === "tag")); + assert.deepEqual( + spawn.calls.find((args) => args[0] === "push"), + ["push", "origin", `${SHA}:refs/tags/2.4.1`], + ); +}); + +test("a failing git call throws with its stderr", () => { + assert.throws( + () => main([], () => ({ status: 128, stdout: "", stderr: "no remote" })), + /git fetch origin main failed: no remote/, + ); +}); diff --git a/scripts/sdk-watch.mjs b/scripts/sdk-watch.mjs index fe46833ceb..ae7d68515c 100644 --- a/scripts/sdk-watch.mjs +++ b/scripts/sdk-watch.mjs @@ -18,12 +18,12 @@ // Four things shape the design, each verified against this repo before it was // written: // -// 1. **Two upstreams, not one.** `client`/`core`/`server`/`server-legacy` all +// 1. **One issue per upstream.** `client`/`core`/`server`/`server-legacy` all // ship from `modelcontextprotocol/typescript-sdk` and release in lockstep; -// `ext-apps` ships from its own repo on its own cadence. Treating them as -// one group would file an issue naming a version that only some of the -// packages have, so `SDK_GROUPS` keeps them separate and each gets its own -// issue and its own marker. +// `ext-apps` and `ext-tasks` each ship from their own repo on their own +// cadence. Treating them as one group would file an issue naming a +// version that only some of the packages have, so `SDK_GROUPS` keeps them +// separate and each gets its own issue and its own marker. // 2. **Compare the INSTALLED version, not the declared range.** #1063 phrases // the check as "is the current version > than the one we have in our // package.json", which is exact today only because the four SDK packages @@ -34,7 +34,7 @@ // whether the fix is a manifest edit or a lockfile refresh — but the // comparison is against the lockfile. // 3. **A new SDK package must not be watched silently by nobody.** The group -// table is a hardcoded list, so a fifth `@modelcontextprotocol/*` package +// table is a hardcoded list, so another `@modelcontextprotocol/*` package // added to the root manifest would never be checked and nothing would say // so. `assertEveryPackageWatched` turns that into a loud failure instead — // the sweep goes red rather than reporting a clean night over a package it @@ -67,6 +67,7 @@ import { spawnSync } from "node:child_process"; import { appendFileSync, readFileSync } from "node:fs"; import semver from "semver"; +import { escapeTableCell as cell } from "./lib/markdown-cell.mjs"; /** The branch this repo ships from, and whose manifests are read. */ export const TARGET_BRANCH = "v2/main"; @@ -76,8 +77,8 @@ export const TARGET_BRANCH = "v2/main"; * * Split by REPOSITORY rather than by npm scope: the four `typescript-sdk` * packages are cut from one release and always share a version, so one issue - * covers the whole bump, while `ext-apps` moves independently and would - * otherwise drag three unrelated packages into its title. + * covers the whole bump, while `ext-apps` and `ext-tasks` each move + * independently and would otherwise drag unrelated packages into their titles. */ export const SDK_GROUPS = [ { @@ -97,6 +98,12 @@ export const SDK_GROUPS = [ repo: "modelcontextprotocol/ext-apps", packages: ["@modelcontextprotocol/ext-apps"], }, + { + key: "ext-tasks", + label: "MCP Tasks extension SDK", + repo: "modelcontextprotocol/ext-tasks", + packages: ["@modelcontextprotocol/ext-tasks"], + }, ]; /** Every package under this prefix is in scope for the watch. */ @@ -334,7 +341,7 @@ export function parseSupersededMarker(body) { /** * Fail loudly when the root manifest declares an SDK package no group watches. * - * The group table is hardcoded, so an added fifth package would be checked by + * The group table is hardcoded, so a newly added package would be checked by * nobody and the sweep would still print a clean result — a silent blind spot * in the one mechanism that exists to remove a silent blind spot. Throwing * turns "we forgot to add it here" into a red run on the next night. @@ -439,8 +446,6 @@ export function buildIssueTitle(state) { return `chore(deps): upgrade the ${state.group.label} to ${state.target}`; } -const cell = (value) => String(value).replace(/\|/g, "\\|"); - /** * Does adopting `target` require editing the root manifest, or only the lockfile? * @@ -529,7 +534,7 @@ export function buildIssueBody(state) { "### Upgrade checklist", "", ...manifestChecklist(rows, target), - "- [ ] Re-check the bundler `external` lists (`clients/{cli,tui}/tsup.config.ts`, `clients/web/tsup.runner.config.ts`) if the release adds or renames an entry point; `npm run verify:bundle-externals` enforces this against the built output.", + "- [ ] Re-check the bundler `external` lists (`clients/{cli,mcpdo,tui}/tsup.config.ts`, `clients/web/tsup.runner.config.ts`) if the release adds or renames an entry point; `npm run verify:bundle-externals` enforces this against the built output.", "- [ ] `npm run format`, then `npm run local:gate`.", "", "An automated review of what actually changed upstream — and which parts of this app it touches — is posted as a comment below.", diff --git a/scripts/sdk-watch.test.mjs b/scripts/sdk-watch.test.mjs index 3ec86af7d1..b1543bb15a 100644 --- a/scripts/sdk-watch.test.mjs +++ b/scripts/sdk-watch.test.mjs @@ -44,6 +44,7 @@ import { const SDK = SDK_GROUPS[0]; const EXT = SDK_GROUPS[1]; +const TASKS = SDK_GROUPS[2]; /** Every version triple current, so a group is behind only where a test says so. */ function currentVersions(overrides = {}) { @@ -997,6 +998,30 @@ test("main files one issue per upstream when both groups are behind", () => { assert.equal(filed[1].from, "1.7.5", "from is the installed version"); }); +test("main watches ext-tasks as its own upstream group", () => { + // ext-tasks ships from its own repo, so it needs its own group entry + // rather than being folded into a sibling's. + const spawn = fakeSpawn({ latest: latestAt(TASKS, "0.2.0") }); + const output = outputFile(); + writeFileSync(output, ""); + + main("o/r", spawn, { + readFile: fakeReadFile({ + declared: { "@modelcontextprotocol/ext-tasks": "0.1.0" }, + installed: { "@modelcontextprotocol/ext-tasks": "0.1.0" }, + }), + output, + }); + + const filed = readFiled(output); + assert.deepEqual( + filed.map((f) => f.repo), + ["modelcontextprotocol/ext-tasks"], + ); + assert.equal(filed[0].from, "0.1.0"); + assert.equal(filed[0].to, "0.2.0"); +}); + /** An open issue this sweep already filed for `target`. */ function existingIssue(number, target, group = SDK) { return { diff --git a/scripts/skill-eval-mcpdo.mjs b/scripts/skill-eval-mcpdo.mjs new file mode 100644 index 0000000000..7c54b36312 --- /dev/null +++ b/scripts/skill-eval-mcpdo.mjs @@ -0,0 +1,1087 @@ +#!/usr/bin/env node +// Trigger eval for the PRODUCT skill `skills/mcpdo` (the mcpdo connection CLI). +// +// `skills:eval` measures the dev-workflow skills in `.claude/skills`. The +// mcpdo skill is a different animal: it SHIPS in the npm package and is +// installed into a *user's* skills directory, so measuring it inside this +// repo's checkout would measure the wrong environment twice over — +// +// - agents discover project skills from `<cwd>/.claude/skills`, so a run +// with `cwd` at the repo root cannot see `skills/mcpdo` at all, and +// - the repo checkout loads AGENTS.md and ten dev skills, none of which a +// user of mcpdo has in context. +// +// So each sample runs in a SANDBOX project: a fresh directory outside the +// repo containing only `.claude/skills/mcpdo` (a copy of the shipped skill, +// evals excluded). That is the closest headless approximation of "a user with +// the mcpdo skill installed asks their agent something". +// +// The sandbox lives under `~/.cache`, not `os.tmpdir()`, for two reasons that +// are both about flags `runPrompt` already passes: the Copilot run carries +// `--disallow-temp-dir`, and a sandbox *inside* the repo would be walked up +// past — the agent would resolve the repo as the project root and load the +// dev skills instead of the sandbox's. +// +// What a trigger case proves — and does not. A hit means only "the agent +// loaded the mcpdo skill for this prompt". The case set therefore skews +// toward UNPROMPTED-RECOGNITION prompts (no mention of mcpdo: "what MCP +// servers am I connected to?") where the decision to reach for the skill is +// the entire measurement, plus negatives that guard against a description +// broadened into firing on every prompt containing "server" or "tools". A +// task-framed "use mcpdo to ..." prompt is near-guaranteed to fire and earns +// one smoke case, no more. Whether the agent then runs the RIGHT mcpdo +// commands is a separate (behavior) measurement this script does not make. +// +// Like `skills:eval`, this is NOT part of `validate`, `local:gate`, or CI: +// it spends metered model calls and the measurement is a hit rate over +// samples. Run it before and after editing `skills/mcpdo/SKILL.md`, and +// compare the same cases. +// +// BEHAVIOR cases (`"kind": "behavior"` in evals.json) go one layer deeper +// than trigger cases: did the agent run the RIGHT mcpdo commands? Each +// sample gets a hermetic environment — a private daemon binding, a throwaway +// storage dir, and a catalog holding exactly one entry (`test-stdio`, this +// repo's stdio test server) — and a recording `mcpdo` shim first on PATH +// that tees stdio through the real build while appending a transcript per +// invocation (lib/mcpdo-eval-shim.mjs). Scoring matches the transcript +// against the case's `expectCalls` (lib/mcpdo-eval-matchers.mjs): an ordered +// subsequence of structured matchers over parsed argv, exit codes, and +// captured output — never over shell strings from the agent's event stream. +// The shell containment is command-scoped approval, probed live on both +// CLIs: Claude `--allowedTools "Bash(mcpdo *)"`, Copilot `--allow-tool +// 'shell(mcpdo:*)'`; anything else effectful is auto-denied fast in headless +// mode and the run continues. (Copilot's `--available-tools` is deliberately +// NOT used here: its availability names differ from its pattern names, and +// filtering the shell tool out entirely makes the model fabricate command +// output — measured, not hypothesized.) Behavior cases are POSIX-only: the +// shim is installed as a `#!/bin/sh` wrapper. +// +// Usage: +// npm run skills:eval:mcpdo +// RUNS=5 THRESHOLD=0.8 npm run skills:eval:mcpdo +// AGENT=copilot npm run skills:eval:mcpdo +// BEHAVIOR_RUNS=4 BEHAVIOR_THRESHOLD=0.75 npm run skills:eval:mcpdo + +import { spawn } from "node:child_process"; +import { + chmodSync, + cpSync, + existsSync, + mkdirSync, + readFileSync, + renameSync, + rmSync, + writeFileSync, +} from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { randomBytes, randomUUID } from "node:crypto"; +import { fileURLToPath } from "node:url"; +import { AGENTS, formatReport, runPrompt } from "./skill-eval.mjs"; +import { parseSkill, validateEvalCases } from "./lib/skill-manifest.mjs"; +import { + assistantReplyText, + REPLY_TEXT_OPTIONS, + evalExpectCalls, + streamText, + validateBehaviorCase, + claudeSessionId, + lastCommandBefore, + succeededBefore, +} from "./lib/mcpdo-eval-matchers.mjs"; + +const ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), ".."); +const SKILL_DIR = path.join(ROOT, "skills", "mcpdo"); +// The cases live OUTSIDE the skill directory, deliberately: `skills/mcpdo` is +// a published artifact (npm `files`, installed into user skill dirs), and the +// skill IS the directory — eval data inside it would ship to every user and +// pollute what an agent may read. Dev-only measurement data belongs with the +// package that owns the skill, not in its payload. (The dev-workflow skills +// under `.claude/skills` keep evals inline because those directories never +// leave the repo.) +const EVALS_FILE = path.join(ROOT, "clients", "mcpdo", "evals", "evals.json"); +const SKILL_NAME = "mcpdo"; + +// Behavior-eval fixtures: the real CLI build the shim wraps, the shim +// itself, and the stdio test server the private catalog points at. Builds, +// not sources — the eval measures what a user would run. +const REAL_BIN = path.join(ROOT, "clients", "mcpdo", "build", "mcp-bin.js"); +const SHIM_SRC = path.join(ROOT, "scripts", "lib", "mcpdo-eval-shim.mjs"); +const TEST_SERVER_BIN = path.join( + ROOT, + "test-servers", + "build", + "test-server-stdio.js", +); +const SERVER_LAUNCHER = path.join( + ROOT, + "scripts", + "lib", + "mcpdo-eval-server-launcher.mjs", +); + +const THRESHOLD = Number(process.env.THRESHOLD ?? 0.8); +const RUNS = Number(process.env.RUNS ?? 3); +const CONCURRENCY = Number(process.env.CONCURRENCY ?? 4); +// `all` runs every known agent in sequence; a single agent name narrows the +// run (e.g. AGENT=claude for cheaper iteration). +const AGENT = process.env.AGENT ?? "all"; +// Behavior knobs are separate from the trigger ones: a multi-turn agentic +// run costs an order of magnitude more than a one-turn trigger sample, and +// its hit rate is honestly lower — 0.5 strict to start, tightened as the +// skill improves rather than loosened to pass. +const BEHAVIOR_RUNS = Number(process.env.BEHAVIOR_RUNS ?? RUNS); +const BEHAVIOR_THRESHOLD = Number(process.env.BEHAVIOR_THRESHOLD ?? 0.5); +const BEHAVIOR_TURNS = Number(process.env.BEHAVIOR_TURNS ?? 14); +// Substring filter over case prompts (both sections) for cheap iteration on +// one case: `CASE_MATCH="add 2 and 3" npm run skills:eval:mcpdo`. A filtered +// run is a dev probe, not a measurement — the summary says so when active. +const CASE_MATCH = process.env.CASE_MATCH ?? ""; + +/** + * Build the sandbox project one sample set runs in. + * + * The skill is COPIED rather than symlinked: a symlink would resolve back + * inside the repo, and whether an agent's project-root detection follows it + * is exactly the kind of version-dependent behavior a measurement should not + * sit on. + * + * @returns {string} The sandbox directory. + */ +export function makeSandbox() { + const dir = path.join( + os.homedir(), + ".cache", + "mcpdo-skill-eval", + `${process.pid}-${randomBytes(4).toString("hex")}`, + ); + const dest = path.join(dir, ".claude", "skills", SKILL_NAME); + mkdirSync(dest, { recursive: true }); + cpSync(SKILL_DIR, dest, { recursive: true }); + // A README so a human finding a leaked sandbox knows what it was. + writeFileSync( + path.join(dir, "README.md"), + "Scratch project for `npm run skills:eval:mcpdo`; safe to delete.\n", + ); + return dir; +} + +/** + * Load and validate the committed cases, partitioned by kind. + * + * Trigger cases go through the shared `validateEvalCases` schema; behavior + * cases through this harness's own `validateBehaviorCase` — `expectCalls` is + * a private contract of this runner, and the shared schema should not grow + * fields only one skill's eval understands. + * + * @returns {{ trigger: object[], behavior: object[] }} + */ +export function loadCases() { + const skill = parseSkill( + SKILL_NAME, + readFileSync(path.join(SKILL_DIR, "SKILL.md"), "utf8"), + ); + if (skill.errors.length > 0) { + throw new Error(`skills/mcpdo/SKILL.md does not parse: ${skill.errors[0]}`); + } + if (!skill.modelInvoked) { + throw new Error( + "skills/mcpdo is not model-invoked; a trigger eval of it measures nothing", + ); + } + const all = JSON.parse(readFileSync(EVALS_FILE, "utf8")); + if (!Array.isArray(all)) { + throw new Error("clients/mcpdo/evals/evals.json must be an array"); + } + // `kind` is an explicit discriminator, required on every case: a defaulted + // kind would let a typo ("behaviour") silently demote a behavior case to a + // trigger case and fail with an unrelated schema error. + const KINDS = new Set(["trigger", "behavior"]); + const kindErrors = all.flatMap((c, i) => + KINDS.has(c?.kind) + ? [] + : [ + `case ${i}: \`kind\` must be one of ${[...KINDS].join(", ")} (got ${JSON.stringify(c?.kind)})`, + ], + ); + if (kindErrors.length > 0) { + throw new Error(`clients/mcpdo/evals/evals.json: ${kindErrors.join("; ")}`); + } + const trigger = all.filter((c) => c.kind === "trigger"); + const behavior = all.filter((c) => c.kind === "behavior"); + const errors = [ + ...validateEvalCases(SKILL_NAME, trigger, new Set([SKILL_NAME])), + ...behavior.flatMap((c, i) => validateBehaviorCase(c, i)), + ]; + if (errors.length > 0) { + throw new Error(`clients/mcpdo/evals/evals.json: ${errors.join("; ")}`); + } + return { trigger, behavior }; +} + +/** + * Agent arguments for a BEHAVIOR run: the trigger policy plus shell, scoped + * to mcpdo by command-level approval. + * + * SECURITY NOTE — this scoping is NOT a boundary. The approval patterns are + * prefix matches (`mcpdo x && anything` passes), and `mcpdo connect` itself + * launches arbitrary stdio commands by design. What the patterns do is keep a + * COOPERATING model from drifting into unrelated shell work (probed live: + * unapproved effectful commands fail fast and the run continues, so the + * denials cost turns, not hangs). The agent's command environment is + * minimized separately (see agentEnv in skill-eval.mjs); a runner who wants + * hard isolation from a misbehaving model should run this suite inside an + * OS-level sandbox of their choice (container, VM, dedicated user). + * + * Passed to `runPrompt` as `agentArgsFn` — a replacement, because the + * trigger policy's `--deny-tool shell` / `--disallowedTools Bash` cannot be + * retracted by appending. + * + * @param {string} agent + * @param {number} maxTurns + * @param {string | null} [resumeSessionId] The session id for a multi-turn + * OAuth behavior case. For claude it is the prior turn's captured id, added + * as `--resume` on the second turn only. For copilot it is an id the harness + * minted, added as `--session-id` on both turns (which that flag sets then + * resumes). Null for a single-turn case, where both agents start fresh. + * @returns {string[]} + */ +export function behaviorAgentArgs(agent, maxTurns, resumeSessionId = null) { + if (agent === "copilot") { + return [ + "--output-format", + "json", + // One flag both sets and resumes: on turn 1 it mints the session under + // the id the harness chose, on turn 2 the same id continues it (verified + // against the live CLI). Unlike claude there is nothing to parse out of + // the stream — the harness owns the id — so it is passed on BOTH turns. + ...(resumeSessionId ? ["--session-id", resumeSessionId] : []), + // No `--available-tools`: its availability names differ from the + // approval-pattern names (the shell tool is `bash` in events but + // `shell(...)` in patterns), and naming it wrong silently removes the + // tool — after which the model FABRICATES command output. Everything + // unapproved is auto-denied in headless mode (drift reduction, not + // containment — see the function doc). + "--allow-tool", + "view,glob,grep,skill", + "--allow-tool", + "shell(mcpdo:*)", + "--deny-tool", + "write", + "--deny-tool", + "url", + "--disable-builtin-mcps", + "--disallow-temp-dir", + "--no-ask-user", + "--no-auto-update", + ]; + } + if (agent !== "claude") throw new Error(`unknown agent \`${agent}\``); + return [ + "-p", + // Continue the prior turn's session when resuming; reads the new prompt + // from stdin exactly as a fresh `-p` run does. + ...(resumeSessionId ? ["--resume", resumeSessionId] : []), + "--output-format", + "stream-json", + "--verbose", + "--max-turns", + String(maxTurns), + "--tools", + "Read,Glob,Grep,Skill,Bash", + "--allowedTools", + "Read,Glob,Grep,Skill,Bash(mcpdo *)", + "--disallowedTools", + "Write,Edit,NotebookEdit,Task,Agent,SlashCommand,WebFetch,WebSearch,KillShell", + "--strict-mcp-config", + ]; +} + +/** + * Normalize a behavior case's server declaration to a name→spec map. + * `servers` wins (validation forbids both); `server` is sugar for a single + * entry under the default name; neither means the default composition. + * + * @param {object} c A behavior case. + * @returns {Record<string, object | undefined>} + */ +export function caseServers(c) { + if (c.servers !== undefined) return c.servers; + return { "test-stdio": c.server }; +} + +/** + * Build one behavior sample's hermetic mcpdo world, next to (not inside) its + * sandbox so the agent's cwd stays clean. + * + * Private daemon binding (own dir + minted token — the same isolation + * `mcpdo private` gives a shell), throwaway storage, a catalog holding + * exactly the case's servers, and a bin dir whose `mcpdo` is the recording + * shim. + * + * No `MCP_ALLOW_DEFAULT_CONNECTION`: agents run non-TTY, so the mcpdo + * itself refuses implicit-MRU targeting (`requireExplicitConnection`) — + * every successful targeting call in a transcript names its connection, + * which is what keeps `connection` matchers decidable even with several + * catalog entries. + * + * Per-entry spec forms (see `validateServerSpec`): undefined → the stdio + * test server in its DEFAULT composition; `{url}` → an already-running HTTP + * fixture; composed with stdio (or no) transport → config on disk, served + * through the eval's stdio launcher; composed with streamable-http + * transport → started IN-PROCESS (`TestServerHttp`) and the entry points at + * its URL. In-process because nothing forces a process boundary for HTTP + * (the daemon only spawns stdio commands), the fixture can't pollute the + * transcript (the shim records only mcpdo invocations), and teardown is a + * direct `stop()`. OAuth rides on the same instance (`oauth` in the spec). + * + * @param {string} sandbox The sample's sandbox dir (from `makeSandbox`). + * @param {Record<string, object | undefined>} [servers] Name→spec map (from + * `caseServers`). + * @returns {Promise<{ env: Record<string, string>, logPath: string, + * teardown: () => Promise<void> }>} + */ +export async function makeBehaviorEnv( + sandbox, + servers = { "test-stdio": undefined }, +) { + const envDir = `${sandbox}-env`; + const daemonDir = path.join(envDir, "daemon"); + const storageDir = path.join(envDir, "storage"); + const binDir = path.join(envDir, "bin"); + const logPath = path.join(envDir, "mcpdo-transcript.ndjson"); + const catalogPath = path.join(envDir, "catalog.json"); + mkdirSync(daemonDir, { recursive: true, mode: 0o700 }); + mkdirSync(storageDir, { recursive: true }); + mkdirSync(binDir, { recursive: true }); + /** In-process HTTP fixtures to stop at teardown. */ + const httpServers = []; + const entries = {}; + try { + for (const [name, spec] of Object.entries(servers)) { + if (spec === undefined) { + entries[name] = { + type: "stdio", + command: process.execPath, + args: [TEST_SERVER_BIN], + }; + } else if ("url" in spec) { + entries[name] = { type: "streamable-http", url: spec.url }; + } else if (spec.transport?.type === "streamable-http") { + const serverConfigPath = path.join( + envDir, + `server-config-${name}.json`, + ); + writeFileSync(serverConfigPath, JSON.stringify(spec, null, 2)); + const { loadConfig, resolveConfig, TestServerHttp } = await import( + path.join(ROOT, "test-servers", "build", "index.js") + ); + const server = new TestServerHttp( + resolveConfig(loadConfig(serverConfigPath)), + ); + await server.start(); + httpServers.push(server); + entries[name] = { type: "streamable-http", url: server.url }; + } else { + const serverConfigPath = path.join( + envDir, + `server-config-${name}.json`, + ); + // The launcher is always a stdio child; the case spec needn't say so + // (and validateServerSpec rejects a spec that says otherwise). + writeFileSync( + serverConfigPath, + JSON.stringify({ transport: { type: "stdio" }, ...spec }, null, 2), + ); + entries[name] = { + type: "stdio", + command: process.execPath, + args: [SERVER_LAUNCHER, serverConfigPath], + }; + } + } + } catch (err) { + for (const s of httpServers) await s.stop().catch(() => {}); + rmSync(envDir, { recursive: true, force: true }); + throw err; + } + writeFileSync(catalogPath, JSON.stringify({ mcpServers: entries }, null, 2)); + const shimBin = path.join(binDir, "mcpdo"); + writeFileSync( + shimBin, + `#!/bin/sh\nexec ${JSON.stringify(process.execPath)} ${JSON.stringify(SHIM_SRC)} "$@"\n`, + ); + chmodSync(shimBin, 0o755); + const env = { + PATH: `${binDir}${path.delimiter}${process.env.PATH ?? ""}`, + MCP_INSPECTOR_DAEMON_DIR: daemonDir, + MCP_INSPECTOR_DAEMON_TOKEN: randomBytes(32).toString("base64url"), + MCP_STORAGE_DIR: storageDir, + MCP_CATALOG_PATH: catalogPath, + // Ephemeral OAuth callback port per sample: the default is a fixed port + // (6276), which concurrent samples' detached auth helpers fight over — + // the loser's consent redirect lands on the winner's helper and sign-in + // silently never completes (measured: intermittent secure-add misses + // with endless `authorized: false` polls). + MCP_OAUTH_CALLBACK_URL: "http://127.0.0.1:0/oauth/callback", + MCPDO_EVAL_REAL: REAL_BIN, + MCPDO_EVAL_LOG: logPath, + }; + const teardown = async ({ keepEnvDir = false } = {}) => { + // Daemon first (it may hold connections into the HTTP fixtures), then + // the fixtures. Direct spawn of the real build, not the shim: teardown + // must not appear in the transcript, and must work even if the shim is + // broken. Async spawn, NOT spawnSync — the in-process fixtures share + // this event loop, and a synchronous wait would deadlock any daemon + // shutdown that talks to them (measured: the sync variant stalled). + await new Promise((resolve) => { + const child = spawn(process.execPath, [REAL_BIN, "daemon", "stop"], { + env: { ...process.env, ...env }, + stdio: "ignore", + }); + const timer = setTimeout(() => child.kill("SIGKILL"), 15000); + child.on("exit", () => { + clearTimeout(timer); + resolve(); + }); + child.on("error", () => { + clearTimeout(timer); + resolve(); + }); + }); + for (const s of httpServers) await s.stop().catch(() => {}); + if (!keepEnvDir) rmSync(envDir, { recursive: true, force: true }); + }; + return { env, logPath, envDir, teardown }; +} + +/** Parse the shim transcript; tolerate a torn final line, never silent-drop. */ +export function readTranscript(logPath) { + if (!existsSync(logPath)) return []; + const lines = readFileSync(logPath, "utf8").split("\n"); + const lastNonEmpty = lines.findLastIndex((l) => l.trim() !== ""); + const records = []; + for (let i = 0; i <= lastNonEmpty; i++) { + if (lines[i].trim() === "") continue; + try { + records.push(JSON.parse(lines[i])); + } catch { + // Only the final record can legitimately be malformed (a write torn + // by a kill); anything earlier is corruption the scorer must not + // silently misread as agent behavior. + if (i === lastNonEmpty) break; + throw new Error( + `${logPath}: malformed transcript record at line ${i + 1}`, + ); + } + } + return records; +} + +/** + * Simulate the human side of a headless OAuth flow: visit the sign-in link + * the agent was handed and approve the consent page. + * + * Non-TTY `connect` against an auth-requiring server exits 0 with the + * authorize URL in its output while a detached helper holds the loopback + * callback. In a real session the agent relays that URL and a human clicks + * it; here the harness is the human. GET shows the composable AS's consent + * page; POSTing the same params approves it; following the redirect delivers + * code+state to the helper's loopback listener, which finishes the token + * exchange — after which the agent's next call revives the connection. + * + * @param {string} authUrl The full /oauth/authorize URL from the transcript. + */ +export async function clickConsent(authUrl) { + let res = await fetch(authUrl, { redirect: "manual" }); + if (res.status === 200) { + const u = new URL(authUrl); + res = await fetch(`${u.origin}${u.pathname}`, { + method: "POST", + headers: { "Content-Type": "application/x-www-form-urlencoded" }, + body: new URLSearchParams(u.searchParams), + redirect: "manual", + }); + } + if (res.status !== 301 && res.status !== 302) { + throw new Error( + `consent POST: expected redirect, got ${res.status}: ${(await res.text()).slice(0, 200)}`, + ); + } + const location = res.headers.get("location"); + if (!location) throw new Error("consent redirect missing Location header"); + // The loopback callback owned by the detached auth helper. + const cb = await fetch(new URL(location, authUrl)); + await cb.text().catch(() => {}); +} + +const AUTHORIZE_URL_RE = + /https?:\/\/[^\s"'<>\\]+\/oauth\/authorize\?[^\s"'<>\\]*/g; + +/** + * Watch a behavior sample's transcript for authorize URLs and auto-approve + * each one once (`autoConsent` cases). Polling the transcript, not the live + * streams: the shim appends a record when an invocation exits, and non-TTY + * connect exits as soon as it prints the URL, so the link shows up while + * the agent is still mid-session. Click failures are logged, not thrown — + * the case then fails on its own matchers, with this as the diagnostic. + * + * @param {string} logPath The sample's transcript path. + * @param {object} [opts] + * @param {number} [opts.delayMs] How long a discovered URL sits unclicked. The + * simulated human signing in instantly would let the agent skip the relay + * entirely (its first poll already shows the connection up), so an + * `expectReply` case holds the click back long enough that the agent has + * to tell the user about the link, exactly as in a real session. + * @param {(() => string) | null} [opts.replyText] When set, a URL is clicked + * only once it appears verbatim in the text this getter returns — the + * agent's own user-facing replies. This makes consent *causal* on the + * relay: the simulated human can only open a link the agent actually + * showed them, so a session where the agent hoards the URL never + * authenticates and fails on `expectCalls` too, exactly as a real user is + * stranded. + * @returns {{ stop: () => void, clickedCount: () => number }} + */ +export function startConsentClicker(logPath, opts = {}) { + const { delayMs = 0, replyText = null } = opts; + const clicked = new Set(); + const firstSeen = new Map(); + let inFlight = false; + let succeeded = 0; + const timer = setInterval(() => { + if (inFlight) return; + const now = Date.now(); + const relayed = replyText === null ? null : replyText(); + const urls = new Set(); + for (const record of readTranscript(logPath)) { + // Scan joined per-stream text, not individual chunks: a long authorize + // URL (e.g. inside pretty-printed `--format json` output) can be split + // across stream chunks, and a per-chunk scan then sees only fragments + // (measured: copilot misses where sign-in silently never happened). + for (const stream of ["stdout", "stderr"]) { + for (const url of streamText(record, stream).match(AUTHORIZE_URL_RE) ?? + []) { + if (clicked.has(url)) continue; + if (relayed !== null && !relayed.includes(url)) continue; + if (!firstSeen.has(url)) firstSeen.set(url, now); + if (now - firstSeen.get(url) >= delayMs) urls.add(url); + } + } + } + if (urls.size === 0) return; + for (const url of urls) clicked.add(url); + inFlight = true; + (async () => { + for (const url of urls) { + try { + await clickConsent(url); + succeeded++; + console.error(` autoConsent: clicked ${new URL(url).pathname}`); + } catch (err) { + console.error(` autoConsent: click failed — ${err.message}`); + } + } + })().finally(() => { + inFlight = false; + }); + }, 250); + return { + stop: () => clearInterval(timer), + // How many consent clicks have COMPLETED (not merely been scheduled) — the + // multi-turn flow waits on this to know the simulated sign-in happened + // before it resumes the session. + clickedCount: () => succeeded, + }; +} + +/** + * One out-of-band `connections/show @name --format json` against the sample's + * daemon, read back as parsed JSON. Spawned as the REAL bin directly (not + * through the PATH shim), so it never appears in the scored transcript — + * `expectCalls` sees only what the AGENT ran. + * + * @param {NodeJS.ProcessEnv} env The sample's hermetic env (binds the daemon). + * @param {string} name Connection/catalog name. + * @returns {Promise<Record<string, unknown> | null>} Parsed result, or null + * on any failure (daemon down, unknown connection, unparseable output). + */ +function daemonShowJson(env, name) { + return new Promise((resolve) => { + const child = spawn( + process.execPath, + [REAL_BIN, "connections/show", `@${name}`, "--format", "json"], + { env, stdio: ["ignore", "pipe", "ignore"] }, + ); + let out = ""; + child.stdout.on("data", (d) => { + out += d; + }); + child.on("error", () => resolve(null)); + child.on("close", () => { + try { + resolve(JSON.parse(out)); + } catch { + resolve(null); + } + }); + }); +} + +/** + * Wait for a pending OAuth connection to finish signing in, driving the + * completion itself. `connections/show` revives a pending entry once tokens + * are on disk (see the daemon's show handler), so polling it both WAITS for + * the simulated human's consent click to land AND performs the revive the + * agent's next op would — leaving the connection live for the resumed turn. + * Ready ⇔ the result no longer carries `authUrl`/`pendingAuth`. Returns false + * on timeout (the agent likely never relayed the URL, so nothing was clicked); + * the case then fails on its own matchers, with the timeout as the signal. + * + * @param {NodeJS.ProcessEnv} env The sample's hermetic env. + * @param {string} name Connection/catalog name. + * @param {number} [timeoutMs] + * @returns {Promise<boolean>} + */ +async function waitForConnectionReady(env, name, timeoutMs = 20000) { + const deadline = Date.now() + timeoutMs; + for (;;) { + const result = await daemonShowJson(env, name); + if (result && !result.authUrl && result.pendingAuth !== true) return true; + if (Date.now() >= deadline) return false; + await new Promise((r) => setTimeout(r, 400)); + } +} + +/** + * Where a failed sample's artifacts are preserved for diagnosis: the shim + * transcript, the raw agent stream, the composed catalog/server configs, and + * the case itself. Never auto-cleaned — delete by hand when done. + */ +export const FAILURES_DIR = path.join( + os.homedir(), + ".cache", + "mcpdo-skill-eval", + "failures", +); + +/** + * Run one behavior sample: fresh sandbox + hermetic env, one agent session, + * transcript scored against the case's `expectCalls`. On a miss the sample's + * env dir (transcript, agent stream, configs, case) is preserved under + * {@link FAILURES_DIR} and its path returned as `artifactsDir`. + * + * @param {object} c A behavior case. + * @param {string} agent One of `AGENTS`. + * @returns {Promise<{ hit: boolean, failures: string[], calls: number }>} + */ +async function runBehaviorSample(c, agent) { + const sandbox = makeSandbox(); + let ready; + try { + ready = await makeBehaviorEnv(sandbox, caseServers(c)); + } catch (err) { + // makeBehaviorEnv cleans up its own envDir on failure; the sandbox + // predates it and is ours to reclaim. + rmSync(sandbox, { recursive: true, force: true }); + throw err; + } + const { env, logPath, envDir, teardown } = ready; + const rawLogPath = path.join(envDir, "agent-session.ndjson"); + // The reply gate reads the live raw agent stream, appended per chunk by + // runPrompt — so mid-session the getter sees every reply sent so far. + const replyText = () => + existsSync(rawLogPath) + ? assistantReplyText( + readFileSync(rawLogPath, "utf8"), + REPLY_TEXT_OPTIONS[agent], + ) + : ""; + const clicker = + c.autoConsent === true + ? startConsentClicker(logPath, { + delayMs: c.consentDelayMs ?? 0, + replyText: c.consentAfterReply === true ? replyText : null, + }) + : null; + let keepEnvDir = false; + let artifactsDir = null; + // A multi-turn case (`followUp`) runs the prompt, completes the simulated + // sign-in out-of-band, then resumes the SAME claude session with the + // follow-up text — modelling a real session where the agent relays the URL, + // ends its turn, and continues once the user says they signed in. Turn 1's + // user-facing reply and the follow-up's send time are captured so the + // scorer can separate "shown the URL in turn 1" and "stopped after connect" + // from the resumed turn's work. + const isMultiTurn = + typeof c.followUp === "string" && c.followUp.trim() !== ""; + const name = Object.keys(caseServers(c))[0]; + // Copilot's session id is the harness's to choose — minted up front and set + // on turn 1 via `--session-id` so turn 2 can resume it. Claude instead + // stamps its own id, which turn 1 emits and we read back below; so turn 1 + // passes no id for claude. + const copilotSessionId = + isMultiTurn && agent === "copilot" ? randomUUID() : null; + let turn1Reply = ""; + let followUpAt = Number.POSITIVE_INFINITY; + try { + await runPrompt(c.prompt, { + cwd: sandbox, + agent, + maxTurns: BEHAVIOR_TURNS, + env, + agentArgsFn: behaviorAgentArgs, + rawLogPath, + resumeSessionId: copilotSessionId, + }); + if (isMultiTurn) { + turn1Reply = replyText(); + // Drive + wait for the simulated sign-in the clicker performs, so the + // resumed turn finds the connection live (as `connect` promised the + // user it would "complete automatically"). + const ready = await waitForConnectionReady(env, name); + if (!ready) { + console.error( + ` multi-turn: ${name} never finished signing in before the follow-up`, + ); + } + // Copilot uses the id we minted; claude's is read back from its stream. + const sessionId = + agent === "copilot" + ? copilotSessionId + : existsSync(rawLogPath) + ? claudeSessionId(readFileSync(rawLogPath, "utf8")) + : null; + followUpAt = Date.now(); + if (sessionId) { + await runPrompt(c.followUp, { + cwd: sandbox, + agent, + maxTurns: BEHAVIOR_TURNS, + env, + agentArgsFn: behaviorAgentArgs, + rawLogPath, + resumeSessionId: sessionId, + }); + } else { + console.error( + ` multi-turn: no ${agent} session id captured — cannot resume`, + ); + } + } + const records = readTranscript(logPath); + const { ok: callsOk, failures } = evalExpectCalls(c.expectCalls, records); + let ok = callsOk; + if (c.expectReply !== undefined) { + // What the agent SAID, not what it ran: the raw agent stream is the + // only record of the user-facing reply (see assistantReplyText). For a + // multi-turn case this is the FIRST turn's reply — the URL must reach + // the user before they sign in, not after. + const reply = isMultiTurn ? turn1Reply : replyText(); + if (!new RegExp(c.expectReply).test(reply)) { + ok = false; + failures.push( + `expectReply: no assistant text matched /${c.expectReply}/ ` + + `(the agent never showed it to the user)`, + ); + } + } + if (c.expectLastCallTurn1 !== undefined) { + // Pure-connect cases (no task command to run): the agent's final + // first-turn mcpdo command must be `connect` — it relayed the URL and + // stopped, with nothing legitimate to chain after it. + const last = lastCommandBefore(records, followUpAt); + if (last !== c.expectLastCallTurn1) { + ok = false; + failures.push( + `expectLastCallTurn1: turn 1's final mcpdo command was ` + + `\`${last ?? "(none)"}\`, expected \`${c.expectLastCallTurn1}\` ` + + `(did it poll/retry after showing the URL instead of ending the turn?)`, + ); + } + } + if (c.rejectCompletedTurn1 !== undefined) { + // Cases with a task command (e.g. `tools/list`): did the agent STOP to + // let the user sign in, or barrel through? The signal is whether the + // task command SUCCEEDED in turn 1 — a poll that ran until the + // connection came up, or a genuine completion, exits 0; an optimistic + // `connect && tools/list` chain exits non-zero (auth_required) and the + // agent correctly waits. Last-command cannot tell those apart; this can. + if (succeededBefore(records, followUpAt, c.rejectCompletedTurn1)) { + ok = false; + failures.push( + `rejectCompletedTurn1: \`${c.rejectCompletedTurn1}\` succeeded ` + + `(exit 0) in turn 1 — the agent completed the task instead of ` + + `waiting for the user to sign in (polled/retried, or never stopped).`, + ); + } + } + if (c.expectReplyFollowUp !== undefined) { + // What the RESUMED turn told the user — turn-2 text only, found by + // stripping the turn-1 prefix the full reply is built on. + const full = replyText(); + const afterText = full.startsWith(turn1Reply) + ? full.slice(turn1Reply.length) + : full; + if (!new RegExp(c.expectReplyFollowUp).test(afterText)) { + ok = false; + failures.push( + `expectReplyFollowUp: no assistant text after the follow-up ` + + `matched /${c.expectReplyFollowUp}/ ` + + `(it never reported the result once access was granted)`, + ); + } + } + // Compact transcript for miss diagnostics: what the agent actually ran + // and what it got back — the eval's equivalent of a stack trace. + const transcript = ok + ? [] + : records.map((r) => ({ + argv: r.argv, + exit: r.exit, + start: r.start, + end: r.end, + out: streamText(r, "stdout").slice(0, 300), + err: streamText(r, "stderr").slice(0, 300), + })); + if (!ok || process.env.MCPDO_EVAL_KEEP === "1") { + keepEnvDir = true; + writeFileSync( + path.join(envDir, "case.json"), + JSON.stringify({ agent, case: c, failures }, null, 2), + ); + artifactsDir = path.join( + FAILURES_DIR, + `${new Date().toISOString().replace(/[:.]/g, "-")}-${agent}-${path.basename(envDir)}`, + ); + } + return { + hit: ok, + failures, + calls: records.length, + transcript, + artifactsDir, + }; + } finally { + clicker?.stop(); + await teardown({ keepEnvDir }); + if (keepEnvDir && artifactsDir !== null) { + mkdirSync(FAILURES_DIR, { recursive: true }); + renameSync(envDir, artifactsDir); + } + rmSync(sandbox, { recursive: true, force: true }); + } +} + +/** + * Run `fn` over `items` with at most `n` in flight. + * + * A rejecting item must not blow up the pool mid-run: sibling samples own + * live resources (spawned daemons, sandbox dirs) that only their own + * try/finally reclaims, so an immediate `Promise.all` rejection would exit + * the process before that cleanup runs. With `onError`, a rejection is + * mapped to a result and the run continues. Without it, the pool stops + * taking new items, lets in-flight siblings finish (and clean up), then + * rethrows the first error. Exported for tests. + * + * @param {Array<T>} items + * @param {number} n Max concurrency. + * @param {(item: T) => Promise<R>} fn + * @param {(item: T, err: unknown) => R} [onError] + * @returns {Promise<R[]>} + * @template T, R + */ +export async function pool(items, n, fn, onError) { + const out = new Array(items.length); + let firstError = null; + let i = 0; + await Promise.all( + Array.from({ length: Math.min(n, items.length) }, async () => { + while (firstError === null && i < items.length) { + const idx = i++; + try { + out[idx] = await fn(items[idx]); + } catch (err) { + if (onError) { + out[idx] = onError(items[idx], err); + } else { + firstError ??= err; + } + } + } + }), + ); + if (firstError !== null) throw firstError; + return out; +} + +async function main() { + if (AGENT !== "all" && !AGENTS.includes(AGENT)) { + console.error( + `skills:eval:mcpdo — unknown AGENT \`${AGENT}\`; known: all, ${AGENTS.join(", ")}`, + ); + process.exit(1); + } + const agents = AGENT === "all" ? [...AGENTS] : [AGENT]; + for (const [name, value] of [ + ["THRESHOLD", THRESHOLD], + ["BEHAVIOR_THRESHOLD", BEHAVIOR_THRESHOLD], + ]) { + if (!Number.isFinite(value) || value < 0 || value > 1) { + console.error( + `skills:eval:mcpdo — ${name} must be a number in [0, 1] (got ${process.env[name]}).`, + ); + process.exit(1); + } + } + const { trigger, behavior } = loadCases(); + const byMatch = (c) => c.prompt.includes(CASE_MATCH); + if (CASE_MATCH !== "") { + console.log( + `skills:eval:mcpdo — CASE_MATCH filter active (${JSON.stringify(CASE_MATCH)}); this is a dev probe, not a measurement`, + ); + } + let failed = 0; + for (const agent of agents) { + failed += await runTriggerSection( + trigger.filter(byMatch).map((c) => ({ ...c, from: SKILL_NAME })), + agent, + ); + failed += await runBehaviorSection(behavior.filter(byMatch), agent); + } + process.exit(failed > 0 ? 1 : 0); +} + +/** + * Trigger section: did the skill fire? One shared read-only sandbox. + * + * @returns {Promise<number>} Failed case count. + */ +async function runTriggerSection(cases, agent) { + if (cases.length === 0) return 0; + const sandbox = makeSandbox(); + console.log( + `skills:eval:mcpdo trigger — ${cases.length} cases x ${RUNS} runs, agent ${agent}, sandbox ${sandbox}`, + ); + const samples = cases.flatMap((c) => Array.from({ length: RUNS }, () => c)); + try { + const results = await pool(samples, CONCURRENCY, async (c) => { + const invoked = await runPrompt(c.prompt, { + cwd: sandbox, + agent, + maxTurns: 1, + }); + return { c, invoked }; + }); + // `ours` is just {mcpdo}: a negative case asserts THIS skill stayed + // quiet. The sandbox has no other project skill, but the contributor's + // ~/.claude skills are still visible to the run and are not ours to + // assert about. + const { lines, failed } = formatReport( + cases, + results, + new Set([SKILL_NAME]), + { + threshold: THRESHOLD, + chainThreshold: 0.5, + chainMaxTurns: 1, + agent, + }, + ); + for (const line of lines) console.log(line); + return failed; + } finally { + rmSync(sandbox, { recursive: true, force: true }); + } +} + +/** + * Behavior section: did the agent run the right mcpdo commands? One hermetic + * world per SAMPLE — samples must not share MRU or connection state. + * + * @returns {Promise<number>} Failed case count. + */ +async function runBehaviorSection(cases, agent) { + if (cases.length === 0) return 0; + if (process.platform === "win32") { + console.log( + "skills:eval:mcpdo behavior — skipped: the recording shim is POSIX-only", + ); + return 0; + } + if (!existsSync(REAL_BIN) || !existsSync(TEST_SERVER_BIN)) { + console.error( + "skills:eval:mcpdo behavior — missing builds; run `npm run build` first" + + ` (need ${path.relative(ROOT, REAL_BIN)} and ${path.relative(ROOT, TEST_SERVER_BIN)})`, + ); + return 1; + } + // Both agents can resume a prior session, so multi-turn cases run on all of + // them (claude via `--resume`, copilot via a minted `--session-id`). + console.log( + `skills:eval:mcpdo behavior — ${cases.length} cases x ${BEHAVIOR_RUNS} runs, agent ${agent}, budget ${BEHAVIOR_TURNS} turns`, + ); + const samples = cases.flatMap((c) => + Array.from({ length: BEHAVIOR_RUNS }, () => c), + ); + const results = await pool( + samples, + CONCURRENCY, + async (c) => ({ + c, + ...(await runBehaviorSample(c, agent)), + }), + // Infrastructure failure (spawn error, fixture died), not a model miss — + // scored as a miss with an explicit reason so the section finishes and + // sibling samples' daemons/sandboxes still get their teardown. + (c, err) => ({ + c, + hit: false, + calls: 0, + failures: [`sample error: ${err?.message ?? String(err)}`], + transcript: [], + artifactsDir: null, + }), + ); + let failed = 0; + for (const c of cases) { + const mine = results.filter((r) => r.c === c); + const hits = mine.filter((r) => r.hit).length; + const rate = hits / mine.length; + const pass = rate >= BEHAVIOR_THRESHOLD; + if (!pass) failed++; + console.log( + `${pass ? "PASS" : "FAIL"} behavior ${hits}/${mine.length} (need ${BEHAVIOR_THRESHOLD}) — ${c.prompt}`, + ); + for (const r of mine) { + if (!r.hit) { + console.log( + ` miss (${r.calls} mcpdo calls): ${r.failures.join("; ")}`, + ); + if (r.artifactsDir) console.log(` artifacts: ${r.artifactsDir}`); + const t0 = r.transcript?.[0]?.start; + for (const t of r.transcript ?? []) { + const at = + typeof t.start === "number" && typeof t0 === "number" + ? ` @${((t.start - t0) / 1000).toFixed(1)}s+${((t.end - t.start) / 1000).toFixed(1)}s` + : ""; + console.log( + ` $ mcpdo ${t.argv.join(" ")} -> ${t.exit}${at}` + + (t.out ? `\n out: ${t.out.replace(/\n/g, "\\n")}` : "") + + (t.err ? `\n err: ${t.err.replace(/\n/g, "\\n")}` : ""), + ); + } + } + } + } + return failed; +} + +if ( + process.argv[1] && + path.resolve(process.argv[1]) === fileURLToPath(import.meta.url) +) { + main().catch((e) => { + console.error(e.message ?? e); + process.exit(1); + }); +} diff --git a/scripts/skill-eval-mcpdo.test.mjs b/scripts/skill-eval-mcpdo.test.mjs new file mode 100644 index 0000000000..4df695193f --- /dev/null +++ b/scripts/skill-eval-mcpdo.test.mjs @@ -0,0 +1,675 @@ +// Tests for the mcpdo behavior eval's deterministic layer: the hermetic +// environment builder and the composed-server launcher. The model-dependent +// hit rates stay manual (`npm run skills:eval:mcpdo`); everything a hit +// depends on that is NOT a model decision is nailed down here. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { spawn } from "node:child_process"; +import { + existsSync, + mkdtempSync, + readFileSync, + rmSync, + statSync, + writeFileSync, +} from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; +import { + caseServers, + clickConsent, + makeBehaviorEnv, + readTranscript, + startConsentClicker, + behaviorAgentArgs, + loadCases, + pool, +} from "./skill-eval-mcpdo.mjs"; + +const ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), ".."); +const LAUNCHER = path.join( + ROOT, + "scripts", + "lib", + "mcpdo-eval-server-launcher.mjs", +); +const TEST_SERVERS_BUILD = path.join(ROOT, "test-servers", "build", "index.js"); + +const tempSandbox = () => + path.join(mkdtempSync(path.join(os.tmpdir(), "mcpdo-eval-test-")), "sandbox"); + +// makeBehaviorEnv builds `${sandbox}-env`; clean both. +const cleanup = (sandbox) => { + rmSync(path.dirname(sandbox), { recursive: true, force: true }); +}; + +test( + "makeBehaviorEnv: default world shape", + { skip: process.platform === "win32" }, + async () => { + const sandbox = tempSandbox(); + try { + const { env, logPath } = await makeBehaviorEnv(sandbox); + const envDir = `${sandbox}-env`; + + const catalog = JSON.parse( + readFileSync(path.join(envDir, "catalog.json"), "utf8"), + ); + assert.deepEqual(Object.keys(catalog.mcpServers), ["test-stdio"]); + assert.equal(catalog.mcpServers["test-stdio"].type, "stdio"); + assert.match( + catalog.mcpServers["test-stdio"].args.join(" "), + /test-server-stdio\.js/, + ); + + // The shim shadows any globally installed mcpdo. + const shim = path.join(envDir, "bin", "mcpdo"); + assert.ok(statSync(shim).mode & 0o100, "shim must be executable"); + assert.match(readFileSync(shim, "utf8"), /mcpdo-eval-shim\.mjs/); + assert.ok(env.PATH.startsWith(path.join(envDir, "bin") + path.delimiter)); + + // The private-daemon trio plus the shim contract. + assert.equal(env.MCP_INSPECTOR_DAEMON_DIR, path.join(envDir, "daemon")); + assert.ok(env.MCP_INSPECTOR_DAEMON_TOKEN.length >= 32); + assert.equal(env.MCP_STORAGE_DIR, path.join(envDir, "storage")); + assert.equal(env.MCP_CATALOG_PATH, path.join(envDir, "catalog.json")); + assert.equal(env.MCPDO_EVAL_LOG, logPath); + assert.ok(existsSync(env.MCPDO_EVAL_REAL) || true); // path shape only + } finally { + cleanup(sandbox); + } + }, +); + +test( + "makeBehaviorEnv: url server spec points the entry at the fixture", + { skip: process.platform === "win32" }, + async () => { + const sandbox = tempSandbox(); + try { + await makeBehaviorEnv( + sandbox, + caseServers({ server: { url: "http://127.0.0.1:3999/mcp" } }), + ); + const catalog = JSON.parse( + readFileSync(path.join(`${sandbox}-env`, "catalog.json"), "utf8"), + ); + assert.deepEqual(catalog.mcpServers["test-stdio"], { + type: "streamable-http", + url: "http://127.0.0.1:3999/mcp", + }); + } finally { + cleanup(sandbox); + } + }, +); + +test( + "makeBehaviorEnv: composed server spec is written and served via the launcher", + { skip: process.platform === "win32" }, + async () => { + const sandbox = tempSandbox(); + try { + const spec = { + serverInfo: { name: "composed-test", version: "1.0.0" }, + tools: [{ preset: "add" }], + }; + await makeBehaviorEnv(sandbox, caseServers({ server: spec })); + const envDir = `${sandbox}-env`; + const entry = JSON.parse( + readFileSync(path.join(envDir, "catalog.json"), "utf8"), + ).mcpServers["test-stdio"]; + assert.equal(entry.type, "stdio"); + assert.equal(entry.args[0], LAUNCHER); + assert.deepEqual( + JSON.parse(readFileSync(entry.args[1], "utf8")), + { transport: { type: "stdio" }, ...spec }, + "the config on disk is the case's spec plus the stdio transport", + ); + } finally { + cleanup(sandbox); + } + }, +); + +test("caseServers: normalizes server / servers / neither", () => { + const spec = { serverInfo: { name: "s", version: "1" } }; + assert.deepEqual(caseServers({}), { "test-stdio": undefined }); + assert.deepEqual(caseServers({ server: spec }), { "test-stdio": spec }); + assert.deepEqual(caseServers({ servers: { a: spec } }), { a: spec }); +}); + +test( + "makeBehaviorEnv: multiple servers get their own entries and config files", + { skip: process.platform === "win32" }, + async () => { + const sandbox = tempSandbox(); + try { + const calendar = { + serverInfo: { name: "calendar", version: "1.0.0" }, + tools: [{ preset: "add" }], + }; + await makeBehaviorEnv(sandbox, { + calendar, + "weather-api": { url: "http://127.0.0.1:3999/mcp" }, + }); + const envDir = `${sandbox}-env`; + const catalog = JSON.parse( + readFileSync(path.join(envDir, "catalog.json"), "utf8"), + ); + assert.deepEqual(Object.keys(catalog.mcpServers).sort(), [ + "calendar", + "weather-api", + ]); + const cal = catalog.mcpServers.calendar; + assert.equal(cal.args[0], LAUNCHER); + assert.match(cal.args[1], /server-config-calendar\.json$/); + assert.deepEqual(catalog.mcpServers["weather-api"], { + type: "streamable-http", + url: "http://127.0.0.1:3999/mcp", + }); + } finally { + cleanup(sandbox); + } + }, +); + +// In-process HTTP fixture: a composed spec with streamable-http transport +// must come up inside the harness, get a catalog entry pointing at its live +// URL, answer an MCP initialize over HTTP, and die at teardown. OAuth rides +// on the same instance, so its AS metadata endpoint is asserted too. +test( + "makeBehaviorEnv: http composed server runs in-process (with oauth metadata)", + { skip: !existsSync(TEST_SERVERS_BUILD) || process.platform === "win32" }, + async () => { + const sandbox = tempSandbox(); + let teardown; + try { + const env = await makeBehaviorEnv(sandbox, { + "protected-api": { + transport: { type: "streamable-http" }, + serverInfo: { name: "protected-api", version: "1.0.0" }, + tools: [{ preset: "add" }], + oauth: { enabled: true, mode: "combined", requireAuth: true }, + }, + }); + teardown = env.teardown; + const entry = JSON.parse( + readFileSync(path.join(`${sandbox}-env`, "catalog.json"), "utf8"), + ).mcpServers["protected-api"]; + assert.equal(entry.type, "streamable-http"); + assert.match(entry.url, /^http:\/\/localhost:\d+\/mcp$/); + + const origin = entry.url.replace(/\/mcp$/, ""); + const meta = await fetch( + `${origin}/.well-known/oauth-authorization-server`, + ); + assert.equal(meta.status, 200); + const asMeta = await meta.json(); + assert.equal(asMeta.issuer, origin); + assert.ok(asMeta.authorization_endpoint.startsWith(origin)); + + // Unauthenticated MCP request → the protected resource must challenge, + // not serve (401 + WWW-Authenticate), proving oauth guards the entry + // the catalog points at. + const res = await fetch(entry.url, { + method: "POST", + headers: { + "content-type": "application/json", + accept: "application/json, text/event-stream", + }, + body: JSON.stringify({ + jsonrpc: "2.0", + id: 1, + method: "initialize", + params: { + protocolVersion: "2025-06-18", + capabilities: {}, + clientInfo: { name: "eval-test", version: "0.0.0" }, + }, + }), + }); + assert.equal(res.status, 401); + assert.ok(res.headers.get("www-authenticate")); + + await teardown(); + teardown = undefined; + await assert.rejects( + fetch(`${origin}/.well-known/oauth-authorization-server`), + undefined, + "fixture must be gone after teardown", + ); + } finally { + if (teardown) await teardown(); + cleanup(sandbox); + } + }, +); + +test( + "makeBehaviorEnv: plain http composed server answers initialize", + { skip: !existsSync(TEST_SERVERS_BUILD) || process.platform === "win32" }, + async () => { + const sandbox = tempSandbox(); + let teardown; + try { + const env = await makeBehaviorEnv(sandbox, { + api: { + transport: { type: "streamable-http" }, + serverInfo: { name: "plain-api", version: "1.0.0" }, + tools: [{ preset: "add" }], + }, + }); + teardown = env.teardown; + const entry = JSON.parse( + readFileSync(path.join(`${sandbox}-env`, "catalog.json"), "utf8"), + ).mcpServers.api; + const res = await fetch(entry.url, { + method: "POST", + headers: { + "content-type": "application/json", + accept: "application/json, text/event-stream", + }, + body: JSON.stringify({ + jsonrpc: "2.0", + id: 1, + method: "initialize", + params: { + protocolVersion: "2025-06-18", + capabilities: {}, + clientInfo: { name: "eval-test", version: "0.0.0" }, + }, + }), + }); + assert.equal(res.status, 200); + assert.match(await res.text(), /"name":\s*"plain-api"/); + } finally { + if (teardown) await teardown(); + cleanup(sandbox); + } + }, +); + +test("readTranscript: missing file and torn tail line", () => { + const dir = mkdtempSync(path.join(os.tmpdir(), "mcpdo-eval-test-")); + try { + assert.deepEqual(readTranscript(path.join(dir, "absent.ndjson")), []); + const p = path.join(dir, "log.ndjson"); + const good = JSON.stringify({ argv: ["tools/list"], exit: 0, events: [] }); + writeFileSync(p, `${good}\n{"argv":["to`); + const records = readTranscript(p); + assert.equal(records.length, 1); + assert.deepEqual(records[0].argv, ["tools/list"]); + } finally { + rmSync(dir, { recursive: true, force: true }); + } +}); + +test("readTranscript: corruption before the final record throws", () => { + const dir = mkdtempSync(path.join(os.tmpdir(), "mcpdo-eval-test-")); + try { + const p = path.join(dir, "log.ndjson"); + const good = JSON.stringify({ argv: ["tools/list"], exit: 0, events: [] }); + // Only a torn FINAL line is a benign kill artifact; a malformed record + // with records after it is corruption the scorer must not misread as + // the agent never running that command. + writeFileSync(p, `${good}\n{"argv":["to\n${good}\n`); + assert.throws( + () => readTranscript(p), + /malformed transcript record at line 2/, + ); + } finally { + rmSync(dir, { recursive: true, force: true }); + } +}); + +test("consent clicker: finds an authorize URL split across stream chunks", async () => { + const { createServer } = await import("node:http"); + const hits = { callback: 0 }; + const server = createServer((req, res) => { + const u = new URL(req.url, "http://127.0.0.1"); + if (u.pathname === "/oauth/authorize" && req.method === "GET") { + res.writeHead(200, { "Content-Type": "text/html" }); + res.end("<form>consent</form>"); + } else if (u.pathname === "/oauth/authorize" && req.method === "POST") { + res.writeHead(302, { + Location: `http://127.0.0.1:${server.address().port}/cb?code=x&state=s`, + }); + res.end(); + } else if (u.pathname === "/cb") { + hits.callback++; + res.writeHead(200); + res.end("done"); + } else { + res.writeHead(404); + res.end(); + } + }); + await new Promise((r) => server.listen(0, "127.0.0.1", r)); + const port = server.address().port; + const authUrl = `http://127.0.0.1:${port}/oauth/authorize?client_id=c&state=s`; + const cut = authUrl.indexOf("authorize?") + 12; // mid-query split + const dir = mkdtempSync(path.join(os.tmpdir(), "mcpdo-eval-test-")); + const logPath = path.join(dir, "log.ndjson"); + writeFileSync( + logPath, + JSON.stringify({ + argv: ["connect", "secure"], + exit: 0, + events: [ + { + t: 1, + stream: "stdout", + data: `"authUrl": "${authUrl.slice(0, cut)}`, + }, + { t: 2, stream: "stdout", data: `${authUrl.slice(cut)}"` }, + ], + }) + "\n", + ); + const clicker = startConsentClicker(logPath); + try { + const deadline = Date.now() + 5000; + while (hits.callback === 0 && Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 50)); + } + assert.equal(hits.callback, 1); + } finally { + clicker.stop(); + server.close(); + rmSync(dir, { recursive: true, force: true }); + } +}); + +test("consent clicker: approves each authorize URL from the transcript once", async () => { + const { createServer } = await import("node:http"); + const hits = { get: 0, post: 0, callback: 0 }; + const server = createServer((req, res) => { + const u = new URL(req.url, "http://127.0.0.1"); + if (u.pathname === "/oauth/authorize" && req.method === "GET") { + hits.get++; + res.writeHead(200, { "Content-Type": "text/html" }); + res.end("<form>consent</form>"); + } else if (u.pathname === "/oauth/authorize" && req.method === "POST") { + hits.post++; + res.writeHead(302, { + Location: `http://127.0.0.1:${server.address().port}/cb?code=x&state=s`, + }); + res.end(); + } else if (u.pathname === "/cb") { + hits.callback++; + res.writeHead(200); + res.end("done"); + } else { + res.writeHead(404); + res.end(); + } + }); + await new Promise((r) => server.listen(0, "127.0.0.1", r)); + const port = server.address().port; + const authUrl = `http://127.0.0.1:${port}/oauth/authorize?client_id=c&state=s`; + const dir = mkdtempSync(path.join(os.tmpdir(), "mcpdo-eval-test-")); + const logPath = path.join(dir, "log.ndjson"); + const record = JSON.stringify({ + argv: ["connect", "secure"], + exit: 0, + events: [{ t: 1, stream: "stdout", data: `"authUrl": "${authUrl}"` }], + }); + writeFileSync(logPath, `${record}\n`); + const clicker = startConsentClicker(logPath); + try { + const deadline = Date.now() + 5000; + // Wait on clickedCount, not hits.callback: the /cb hit fires INSIDE + // clickConsent, and succeeded++ only runs after it returns — so waiting on + // the callback races the increment and reads clickedCount() as 0. + while (clicker.clickedCount() === 0 && Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 50)); + } + assert.equal(hits.get, 1); + assert.equal(hits.post, 1); + assert.equal(hits.callback, 1); + // clickedCount reflects COMPLETED clicks — the multi-turn wait keys on it. + assert.equal(clicker.clickedCount(), 1); + // Same URL appearing again (an agent re-printing it) is not re-clicked. + writeFileSync(logPath, `${record}\n${record}\n`); + await new Promise((r) => setTimeout(r, 700)); + assert.equal(hits.post, 1); + assert.equal(clicker.clickedCount(), 1); + } finally { + clicker.stop(); + server.close(); + rmSync(dir, { recursive: true, force: true }); + } +}); + +test("consent clicker: replyText gate holds the click until the agent relays", async () => { + const { createServer } = await import("node:http"); + const hits = { post: 0, callback: 0 }; + const server = createServer((req, res) => { + const u = new URL(req.url, "http://127.0.0.1"); + if (u.pathname === "/oauth/authorize" && req.method === "GET") { + res.writeHead(200, { "Content-Type": "text/html" }); + res.end("<form>consent</form>"); + } else if (u.pathname === "/oauth/authorize" && req.method === "POST") { + hits.post++; + res.writeHead(302, { + Location: `http://127.0.0.1:${server.address().port}/cb?code=x&state=s`, + }); + res.end(); + } else if (u.pathname === "/cb") { + hits.callback++; + res.writeHead(200); + res.end("done"); + } else { + res.writeHead(404); + res.end(); + } + }); + await new Promise((r) => server.listen(0, "127.0.0.1", r)); + const port = server.address().port; + const authUrl = `http://127.0.0.1:${port}/oauth/authorize?client_id=c&state=s`; + const dir = mkdtempSync(path.join(os.tmpdir(), "mcpdo-eval-test-")); + const logPath = path.join(dir, "log.ndjson"); + writeFileSync( + logPath, + JSON.stringify({ + argv: ["connect", "secure"], + exit: 0, + events: [{ t: 1, stream: "stdout", data: `"authUrl": "${authUrl}"` }], + }) + "\n", + ); + let reply = "working on it"; + const clicker = startConsentClicker(logPath, { replyText: () => reply }); + try { + // URL is in the transcript but not in any reply: the simulated human + // cannot open a link the agent never showed them. + await new Promise((r) => setTimeout(r, 700)); + assert.equal(hits.post, 0); + reply = `open this to sign in: ${authUrl}`; + const deadline = Date.now() + 5000; + while (hits.callback === 0 && Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 50)); + } + assert.equal(hits.post, 1); + assert.equal(hits.callback, 1); + } finally { + clicker.stop(); + server.close(); + rmSync(dir, { recursive: true, force: true }); + } +}); + +test("clickConsent: throws when the AS does not redirect", async () => { + const { createServer } = await import("node:http"); + const server = createServer((_req, res) => { + res.writeHead(400); + res.end("nope"); + }); + await new Promise((r) => server.listen(0, "127.0.0.1", r)); + const port = server.address().port; + try { + await assert.rejects( + clickConsent(`http://127.0.0.1:${port}/oauth/authorize?x=1`), + /expected redirect, got 400/, + ); + } finally { + server.close(); + } +}); + +test("loadCases: the committed evals file validates and partitions", () => { + const { trigger, behavior } = loadCases(); + assert.ok(trigger.length >= 5); + assert.ok(behavior.length >= 2); + assert.ok(trigger.every((c) => c.kind === "trigger")); + assert.ok(behavior.every((c) => c.kind === "behavior")); +}); + +test("behaviorAgentArgs: resume id becomes --resume (claude) / --session-id (copilot)", () => { + const fresh = behaviorAgentArgs("claude", 14); + assert.ok(!fresh.includes("--resume")); + const resumed = behaviorAgentArgs("claude", 14, "sess-abc"); + const i = resumed.indexOf("--resume"); + assert.ok(i >= 0, "resume flag present"); + assert.equal(resumed[i + 1], "sess-abc"); + // -p stays first; --resume rides right after it, before the output format. + assert.equal(resumed[0], "-p"); + assert.ok(resumed.indexOf("--resume") < resumed.indexOf("--output-format")); + // Copilot takes the id too, as `--session-id` (one flag that both sets the + // session on turn 1 and resumes it on turn 2). + const copFresh = behaviorAgentArgs("copilot", 14); + assert.ok(!copFresh.includes("--session-id")); + const copResumed = behaviorAgentArgs("copilot", 14, "sess-abc"); + const j = copResumed.indexOf("--session-id"); + assert.ok(j >= 0, "session-id flag present for copilot"); + assert.equal(copResumed[j + 1], "sess-abc"); +}); + +// End-to-end: the launcher must actually SERVE the composed config. One MCP +// handshake over newline-delimited JSON-RPC, then tools/list, asserting the +// composed tool set (and only it). Skipped when the test-servers build is +// absent — the eval itself requires builds too, and says so. +test( + "launcher serves a composed config over stdio", + { skip: !existsSync(TEST_SERVERS_BUILD) || process.platform === "win32" }, + async () => { + const dir = mkdtempSync(path.join(os.tmpdir(), "mcpdo-eval-test-")); + const configPath = path.join(dir, "server.json"); + writeFileSync( + configPath, + JSON.stringify({ + transport: { type: "stdio" }, + serverInfo: { name: "composed-test", version: "1.0.0" }, + tools: [{ preset: "add" }], + }), + ); + const child = spawn(process.execPath, [LAUNCHER, configPath], { + stdio: ["pipe", "pipe", "pipe"], + }); + try { + const responses = new Map(); + let buf = ""; + let notify; + const arrived = new Promise((r) => (notify = r)); + child.stdout.on("data", (chunk) => { + buf += chunk.toString(); + let nl; + while ((nl = buf.indexOf("\n")) !== -1) { + const line = buf.slice(0, nl); + buf = buf.slice(nl + 1); + if (line.trim() === "") continue; + const msg = JSON.parse(line); + if (msg.id !== undefined) { + responses.set(msg.id, msg); + notify(); + } + } + }); + const waitFor = async (id, ms = 10000) => { + const deadline = Date.now() + ms; + while (!responses.has(id)) { + if (Date.now() > deadline) { + throw new Error(`no response ${id}; stderr may explain`); + } + await new Promise((r) => setTimeout(r, 25)); + } + return responses.get(id); + }; + const send = (msg) => child.stdin.write(JSON.stringify(msg) + "\n"); + + send({ + jsonrpc: "2.0", + id: 1, + method: "initialize", + params: { + protocolVersion: "2025-06-18", + capabilities: {}, + clientInfo: { name: "eval-test", version: "0.0.0" }, + }, + }); + const init = await waitFor(1); + assert.equal(init.result.serverInfo.name, "composed-test"); + send({ jsonrpc: "2.0", method: "notifications/initialized" }); + send({ jsonrpc: "2.0", id: 2, method: "tools/list", params: {} }); + const tools = (await waitFor(2)).result.tools.map((t) => t.name); + assert.deepEqual(tools, ["add"]); + } finally { + child.kill("SIGTERM"); + rmSync(dir, { recursive: true, force: true }); + } + }, +); + +test("pool: onError maps a rejecting item and every sibling still runs", async () => { + const done = []; + const results = await pool( + [1, 2, 3, 4, 5], + 2, + async (n) => { + if (n === 3) throw new Error(`boom-${n}`); + done.push(n); + return `ok-${n}`; + }, + (n, err) => `mapped-${n}:${err.message}`, + ); + assert.deepEqual(results, [ + "ok-1", + "ok-2", + "mapped-3:boom-3", + "ok-4", + "ok-5", + ]); + assert.deepEqual(done.sort(), [1, 2, 4, 5]); +}); + +test("pool: without onError it lets in-flight cleanup finish, stops new work, rethrows", async () => { + const cleaned = []; + const started = []; + let releaseSlow; + const slow = new Promise((r) => { + releaseSlow = r; + }); + await assert.rejects( + pool([1, 2, 3, 4], 2, async (n) => { + started.push(n); + try { + if (n === 1) { + // Rejects while item 2 is still in flight; its finally must run + // before the pool settles (that's where real samples tear down + // their daemons/sandboxes). + throw new Error("infra"); + } + await slow; + return n; + } finally { + cleaned.push(n); + if (n === 1) setTimeout(releaseSlow, 20); + } + }), + /infra/, + ); + // Item 2's cleanup ran; items 3 and 4 were never started after the abort. + assert.deepEqual(cleaned.sort(), [1, 2]); + assert.deepEqual(started.sort(), [1, 2]); +}); diff --git a/scripts/skill-eval.mjs b/scripts/skill-eval.mjs index 735643d77a..615e7594d2 100755 --- a/scripts/skill-eval.mjs +++ b/scripts/skill-eval.mjs @@ -49,7 +49,13 @@ // the run has established it needs one — so they never share a column. import { spawn } from "node:child_process"; -import { readFileSync, existsSync, readdirSync, statSync } from "node:fs"; +import { + readFileSync, + existsSync, + readdirSync, + statSync, + appendFileSync, +} from "node:fs"; import path from "node:path"; import { fileURLToPath } from "node:url"; import { cliSpawnArgs, probeCliVersion } from "./lib/claude-cli.mjs"; @@ -651,6 +657,83 @@ export function stopLiveCopilotRuns(runs = liveCopilotRuns) { return n; } +/** + * Minimal environment for a spawned agent. The agent under test is a + * nondeterministic model with shell access — its actions are untrusted, so it + * must not inherit the developer's full environment (arbitrary exported + * credentials, cloud keys, tokens). Instead of spreading `process.env`, pick + * only what an agent CLI needs to run and authenticate: + * + * - process basics (PATH, HOME, TMPDIR, locale, terminal), + * - proxy configuration, and + * - the agent's OWN auth/config vars (ANTHROPIC_ and CLAUDE_ prefixes for + * claude; GITHUB_, GH_, and COPILOT_ prefixes for copilot) — the agent + * needs its credentials, but the other agent's (and everything else) + * stays out. + * + * HOME remains real because both CLIs keep auth state under it; env + * minimization limits what leaks into the model's command environment, not + * filesystem access. OS-level isolation (container, VM, dedicated user) is + * the runner's responsibility if they want a hard boundary. + * + * @param {string} agent `claude` or `copilot`. + * @param {NodeJS.ProcessEnv} [source] Injectable for tests. + * @returns {Record<string, string>} + */ +export function agentEnv(agent, source = process.env) { + const base = [ + "PATH", + "HOME", + "TMPDIR", + "TERM", + "SHELL", + "USER", + "LOGNAME", + "LANG", + "LC_ALL", + "HTTP_PROXY", + "HTTPS_PROXY", + "NO_PROXY", + "http_proxy", + "https_proxy", + "no_proxy", + "SSL_CERT_FILE", + "NODE_EXTRA_CA_CERTS", + // Windows equivalents; harmless no-ops elsewhere. + "SYSTEMROOT", + "APPDATA", + "LOCALAPPDATA", + "USERPROFILE", + "TEMP", + "TMP", + "PATHEXT", + "COMSPEC", + ]; + const prefixes = + agent === "copilot" + ? ["GITHUB_", "GH_", "COPILOT_", "XDG_"] + : // AWS_/GOOGLE_/CLOUD_ML_ carry Bedrock and Vertex credentials/region; + // claude routes through them when CLAUDE_CODE_USE_BEDROCK/VERTEX is + // set, and dropping them fails auth on those runs. + ["ANTHROPIC_", "CLAUDE_", "XDG_", "AWS_", "GOOGLE_", "CLOUD_ML_"]; + const env = {}; + // Windows environment keys keep mixed-case spellings (`Path`, `SystemRoot`, + // `ComSpec`), so the base-list match is case-insensitive; the original + // spelling is preserved in the returned env. Prefixes stay case-sensitive — + // the vendor vars they name are uppercase on every platform. + const baseUpper = base.map((k) => k.toUpperCase()); + for (const key of Object.keys(source)) { + if (source[key] === undefined) continue; + if ( + baseUpper.includes(key.toUpperCase()) || + prefixes.some((p) => key.startsWith(p)) + ) { + env[key] = source[key]; + } + } + return env; +} + /** * Drive one fresh session and return the payloads the `Skill` tool was called * with. @@ -674,6 +757,22 @@ export function runPrompt( maxTurns = 1, agent = "claude", killFn = killTree, + // Additive seams for the mcpdo BEHAVIOR eval (skill-eval-mcpdo.mjs): + // `env` merges over the minimal agent environment (the behavior eval puts + // a recording shim first on PATH and binds a private daemon), and + // `agentArgsFn` replaces the whole argument builder — replacement, not + // appending, because a policy that must allow shell cannot be reached by + // appending to one that denies it (`--deny-tool shell` has no inverse + // flag). Defaults preserve this file's read-only trigger policy exactly. + env = null, + agentArgsFn = agentArgs, + // Optional raw capture of the agent's stdout stream (NDJSON events) for + // post-mortem diagnosis of failed samples. + rawLogPath = null, + // Continue a prior session rather than starting fresh (multi-turn behavior + // eval). Passed to `agentArgsFn` as its third argument; the default + // `agentArgs` ignores it, and only the mcpdo behavior builder acts on it. + resumeSessionId = null, } = {}, ) { return new Promise((resolve, reject) => { @@ -685,9 +784,11 @@ export function runPrompt( // process table. const { command, args, options } = cliSpawnArgs( agent, - agentArgs(agent, maxTurns), + agentArgsFn(agent, maxTurns, resumeSessionId), { cwd, + // Never the full inherited environment: see agentEnv. + env: { ...agentEnv(agent), ...(env ?? {}) }, stdio: ["pipe", "pipe", "inherit"], // Its own process group, so `killTree` can reach the native binary the // wrapper starts. Windows has no groups; `taskkill /T` covers it. @@ -716,6 +817,13 @@ export function runPrompt( // stream rather than restarting at each read. let turnOffset = 0; p.stdout.on("data", (chunk) => { + if (rawLogPath !== null) { + try { + appendFileSync(rawLogPath, chunk); + } catch { + // Diagnostics only — never fail the run over the raw log. + } + } if (stopped) return; const parsed = collect(buf + chunk.toString(), turnOffset); buf = parsed.rest; diff --git a/scripts/skill-eval.test.mjs b/scripts/skill-eval.test.mjs index d9a38d0046..d41fab3113 100644 --- a/scripts/skill-eval.test.mjs +++ b/scripts/skill-eval.test.mjs @@ -17,6 +17,7 @@ import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; import path from "node:path"; import { + agentEnv, caseHit, chainHit, formatReport, @@ -1053,3 +1054,92 @@ test("a rejected Copilot run leaves the in-flight set", async () => { ); assert.equal(stopLiveCopilotRuns(), 0); }); + +test("agentEnv: agents get only process basics and their own credentials", () => { + const source = { + PATH: "/usr/bin", + HOME: "/Users/dev", + AWS_SECRET_ACCESS_KEY: "bedrock-key", + GOOGLE_APPLICATION_CREDENTIALS: "/creds.json", + CLOUD_ML_REGION: "us-east5", + NPM_TOKEN: "leak-me-not", + ANTHROPIC_API_KEY: "claude-key", + CLAUDE_CODE_FLAG: "1", + GH_TOKEN: "copilot-key", + GITHUB_TOKEN: "copilot-key-2", + COPILOT_MODEL: "m", + XDG_CONFIG_HOME: "/Users/dev/.config", + }; + const claude = agentEnv("claude", source); + assert.equal(claude.PATH, "/usr/bin"); + assert.equal(claude.HOME, "/Users/dev"); + assert.equal(claude.ANTHROPIC_API_KEY, "claude-key"); + assert.equal(claude.CLAUDE_CODE_FLAG, "1"); + assert.equal(claude.XDG_CONFIG_HOME, "/Users/dev/.config"); + // Bedrock/Vertex credentials are claude's own auth route + // (CLAUDE_CODE_USE_BEDROCK/VERTEX), so AWS_/GOOGLE_/CLOUD_ML_ pass through. + assert.equal(claude.AWS_SECRET_ACCESS_KEY, "bedrock-key"); + assert.equal(claude.GOOGLE_APPLICATION_CREDENTIALS, "/creds.json"); + assert.equal(claude.CLOUD_ML_REGION, "us-east5"); + assert.ok(!("NPM_TOKEN" in claude)); + assert.ok(!("GH_TOKEN" in claude)); + assert.ok(!("GITHUB_TOKEN" in claude)); + const copilot = agentEnv("copilot", source); + assert.equal(copilot.GH_TOKEN, "copilot-key"); + assert.equal(copilot.GITHUB_TOKEN, "copilot-key-2"); + assert.equal(copilot.COPILOT_MODEL, "m"); + assert.ok(!("ANTHROPIC_API_KEY" in copilot)); + assert.ok(!("AWS_SECRET_ACCESS_KEY" in copilot)); + assert.ok(!("GOOGLE_APPLICATION_CREDENTIALS" in copilot)); +}); + +test("agentEnv: mixed-case Windows base keys pass through with spelling intact", () => { + const source = { + Path: "C:\\Windows;C:\\Windows\\System32", + SystemRoot: "C:\\Windows", + ComSpec: "C:\\Windows\\System32\\cmd.exe", + AppData: "C:\\Users\\dev\\AppData\\Roaming", + // Mixed case only helps base-list names; prefixes stay case-sensitive. + Anthropic_Api_Key: "not-a-real-prefix-match", + }; + const env = agentEnv("claude", source); + assert.equal(env.Path, source.Path); + assert.equal(env.SystemRoot, source.SystemRoot); + assert.equal(env.ComSpec, source.ComSpec); + assert.equal(env.AppData, source.AppData); + assert.ok(!("Anthropic_Api_Key" in env)); +}); + +test("runPrompt: spawned agent env is minimal plus the caller's overlay", async () => { + process.env.SKILL_EVAL_TEST_SECRET = "leak-me-not"; + try { + let seen; + const spawnFn = (command, args, options) => { + seen = options.env; + const child = new EventEmitter(); + child.stdout = new EventEmitter(); + child.stdin = { end: () => {} }; + queueMicrotask(() => { + child.stdout.emit( + "data", + Buffer.from( + JSON.stringify({ type: "result", subtype: "success" }) + "\n", + ), + ); + child.emit("close", 0); + }); + return child; + }; + await runPrompt("p", { + agent: "claude", + spawnFn, + killFn: () => {}, + env: { MCPDO_EVAL_LOG: "/tmp/x" }, + }); + assert.ok(!("SKILL_EVAL_TEST_SECRET" in seen)); + assert.equal(seen.MCPDO_EVAL_LOG, "/tmp/x"); + assert.equal(seen.PATH, process.env.PATH); + } finally { + delete process.env.SKILL_EVAL_TEST_SECRET; + } +}); diff --git a/scripts/smoke-cli.mjs b/scripts/smoke-cli.mjs index c1c6e69b29..b88e837095 100644 --- a/scripts/smoke-cli.mjs +++ b/scripts/smoke-cli.mjs @@ -439,7 +439,7 @@ try { } let envelope; try { - // The documented envelope contract (clients/cli/src/error-handler.ts) is + // The documented envelope contract (core/cli/error-handler.ts) is // that the envelope is the *last* stderr line — `2>&1 | tail -1 | jq // .error` — because stderr legitimately carries diagnostics first (the // secret-store fallback banner on a host without a usable keychain, e.g. diff --git a/scripts/smoke-mcpdo.mjs b/scripts/smoke-mcpdo.mjs new file mode 100644 index 0000000000..1984401ca5 --- /dev/null +++ b/scripts/smoke-mcpdo.mjs @@ -0,0 +1,232 @@ +#!/usr/bin/env node +/** + * End-to-end smoke test for the experimental mcpdo daemon CLI + * (`clients/mcpdo`). The unit/integration suite covers the daemon and + * command surface piecewise; this script drives the BUILT binary the way an + * agent shell would — non-TTY, catalog-based — and asserts the headline + * lifecycle end to end: + * + * 1. `connect <entry>` resolves a catalog entry and spawns/uses the + * background daemon (auto-ensure path). + * 2. `@entry tools/call` round-trips a real stdio MCP server. + * 3. A tool that elicits (`submit_ticket`) parks: the CLI returns + * immediately with an `elicitation/respond <id>` handle instead of + * hanging (the skill's non-TTY contract). + * 4. `elicitation/respond <id> field:=value ...` resumes the parked call + * and the original tool result comes back. + * 5. `daemon status --format json` reports the connection and + * `stopping: false`. + * 6. `disconnect` + `daemon stop` tear down, and the daemon process + * actually exits (socket file released). + * + * Fully hermetic: private daemon dir/socket, storage dir, catalog, and + * daemon token under a temp dir — the developer's real mcpdo daemon (if + * any) is untouched. Exits non-zero on any mismatch. + * + * Expects `clients/mcpdo/build` to be built first (the validate / CI + * ordering guarantees this). The composed test server (`test-servers/build`) + * is rebuilt on every run — see `scripts/lib/ensure-test-servers.mjs`. + */ + +import { spawnSync } from "node:child_process"; +import { randomBytes } from "node:crypto"; +import { + existsSync, + mkdirSync, + mkdtempSync, + rmSync, + writeFileSync, +} from "node:fs"; +import { tmpdir } from "node:os"; +import { join, resolve } from "node:path"; +import { ensureTestServers } from "./lib/ensure-test-servers.mjs"; + +const repoRoot = resolve(import.meta.dirname, ".."); +const mcpdoBin = join(repoRoot, "clients", "mcpdo", "build", "mcp-bin.js"); +const serverLauncher = join( + repoRoot, + "scripts", + "lib", + "mcpdo-eval-server-launcher.mjs", +); + +function fail(message) { + console.error(`smoke:mcpdo FAILED — ${message}`); + process.exit(1); +} + +if (!existsSync(mcpdoBin)) { + fail(`missing build artifact ${mcpdoBin} — run \`npm run build\` first`); +} + +// Rebuilt on every run — presence is not freshness (#2111). The launcher +// imports `test-servers/build/index.js`, which the same tsc pass emits; +// `fixtures` is the named entry that pins the emit actually happened. +try { + ensureTestServers({ + repoRoot, + label: "smoke:mcpdo", + requires: ["fixtures"], + }); +} catch (e) { + fail(e.message); +} +const testServersIndex = join(repoRoot, "test-servers", "build", "index.js"); +if (!existsSync(testServersIndex)) { + fail(`test-servers build did not emit ${testServersIndex}`); +} + +// Hermetic sandbox: everything the daemon touches lives under here. +const sandbox = mkdtempSync(join(tmpdir(), "mcpdo-smoke-")); +const daemonDir = join(sandbox, "daemon"); +const storageDir = join(sandbox, "storage"); +mkdirSync(daemonDir); +mkdirSync(storageDir); + +// One composed stdio server: a plain tool (get_sum) for the basic call and +// an intrinsically-eliciting tool (submit_ticket) for the park round-trip. +const serverConfigPath = join(sandbox, "server-config.json"); +writeFileSync( + serverConfigPath, + JSON.stringify( + { + transport: { type: "stdio" }, + serverInfo: { name: "helpdesk", version: "1.0.0" }, + tools: [{ preset: "get_sum" }, { preset: "submit_ticket" }], + }, + null, + 2, + ), +); +const catalogPath = join(sandbox, "catalog.json"); +writeFileSync( + catalogPath, + JSON.stringify( + { + mcpServers: { + helpdesk: { + type: "stdio", + command: process.execPath, + args: [serverLauncher, serverConfigPath], + }, + }, + }, + null, + 2, + ), +); + +const SMOKE_ENV = { + MCP_INSPECTOR_DAEMON_DIR: daemonDir, + MCP_INSPECTOR_DAEMON_TOKEN: randomBytes(32).toString("base64url"), + MCP_STORAGE_DIR: storageDir, + MCP_CATALOG_PATH: catalogPath, + // Memory store keeps the smoke off the host keychain (no OAuth here, but + // the store is probed at startup) — mirrors smoke-cli.mjs. + MCP_INSPECTOR_SECRET_STORE: "memory", +}; + +/** Run one mcpdo invocation. Returns { status, stdout, stderr }. */ +function runMcpdo(args) { + const r = spawnSync(process.execPath, [mcpdoBin, ...args], { + cwd: repoRoot, + env: { ...process.env, ...SMOKE_ENV }, + encoding: "utf-8", + timeout: 60_000, + }); + if (r.error) fail(`mcpdo ${args.join(" ")} did not run: ${r.error.message}`); + return r; +} + +function step(name, args, { expectStatus = 0, match = [] } = {}) { + const r = runMcpdo(args); + if (r.status !== expectStatus) { + fail( + `${name}: expected exit ${expectStatus}, got ${r.status}\nstdout:\n${r.stdout}\nstderr:\n${r.stderr}`, + ); + } + for (const m of match) { + if (!m.test(r.stdout)) { + fail(`${name}: stdout did not match ${m}\nstdout:\n${r.stdout}`); + } + } + console.log(`smoke:mcpdo ok — ${name}`); + return r; +} + +const sleep = (ms) => + new Promise((resolveSleep) => setTimeout(resolveSleep, ms)); + +try { + // 1. Connect the catalog entry (auto-spawns the private daemon). + step("connect", ["connect", "helpdesk"], { + match: [/@helpdesk/], + }); + + // 2. Plain tool call round-trip. + step("tools/call", ["@helpdesk", "tools/call", "get_sum", "a:=2", "b:=3"], { + match: [/"result":\s*5/], + }); + + // 3. Eliciting tool parks instead of hanging; the CLI hands back a + // respond command with the elicitation id. + const parked = step( + "tools/call parks on elicitation", + ["@helpdesk", "tools/call", "submit_ticket", "summary:=Printer jammed"], + { match: [/Input required|elicitationPending/, /elicitation\/respond/] }, + ); + const idMatch = parked.stdout.match(/elicitation\/respond (\S+)/); + if (!idMatch) { + fail(`could not extract elicitation id from:\n${parked.stdout}`); + } + + // 4. Respond resumes the parked call; the tool's real result comes back. + step( + "elicitation/respond resumes the call", + [ + "elicitation/respond", + idMatch[1], + "contact_name:=Ada Lovelace", + "contact_email:=ada@example.com", + ], + { match: [/TCK-\d+/, /Ada Lovelace/] }, + ); + + // 5. Status sees the live connection and a non-stopping daemon. + const status = step("daemon status", [ + "--format", + "json", + "daemon", + "status", + ]); + const parsedStatus = JSON.parse(status.stdout); + if (parsedStatus.connections?.[0]?.name !== "helpdesk") { + fail(`daemon status missing helpdesk connection:\n${status.stdout}`); + } + if (parsedStatus.stopping !== false) { + fail(`daemon status should report stopping: false:\n${status.stdout}`); + } + + // 6. Teardown: disconnect, stop, and confirm the daemon really exits + // (socket removed once shutdown completes). + step("disconnect", ["disconnect", "helpdesk"]); + step("daemon stop", ["daemon", "stop"]); + const socketPath = join(daemonDir, "mcpdod.sock"); + const deadline = Date.now() + 10_000; + while (existsSync(socketPath)) { + if (Date.now() > deadline) + fail("daemon socket still present 10s after stop"); + await sleep(100); + } + console.log("smoke:mcpdo ok — daemon exited (socket released)"); + + console.log("smoke:mcpdo PASSED"); +} finally { + // Belt and braces: if a step failed mid-flight, don't leak the daemon. + spawnSync(process.execPath, [mcpdoBin, "daemon", "stop"], { + env: { ...process.env, ...SMOKE_ENV }, + encoding: "utf-8", + timeout: 30_000, + }); + rmSync(sandbox, { recursive: true, force: true }); +} diff --git a/scripts/verify-bundle-externals.mjs b/scripts/verify-bundle-externals.mjs index 8b2774f8e6..b9f2123e3e 100644 --- a/scripts/verify-bundle-externals.mjs +++ b/scripts/verify-bundle-externals.mjs @@ -34,8 +34,12 @@ const repoRoot = resolve(dirname(fileURLToPath(import.meta.url)), ".."); /** * Clients that ship a tsup bundle, with the config to read `external` from and - * the build directory to inspect. `clients/launcher` is plain `tsc` — it emits - * no bundle and inlines nothing — so it has nothing to check. + * the build directory to inspect. `entry` names the file whose presence + * proves a build actually ran; it defaults to `index.js` (what web/cli/tui + * each name their single tsup entry) and is overridden only when a client's + * tsup config uses a different entry name, like mcpdo's multi-entry `mcp-bin`. + * `clients/launcher` is plain `tsc` — it emits no bundle and inlines nothing — + * so it has nothing to check. */ export const BUNDLED_CLIENTS = [ { @@ -53,6 +57,12 @@ export const BUNDLED_CLIENTS = [ config: "clients/tui/tsup.config.ts", build: "clients/tui/build", }, + { + name: "mcpdo", + config: "clients/mcpdo/tsup.config.ts", + build: "clients/mcpdo/build", + entry: "mcp-bin.js", + }, ]; /** @@ -205,10 +215,11 @@ function main() { ); for (const client of BUNDLED_CLIENTS) { const buildDir = join(repoRoot, client.build); - const entry = join(buildDir, "index.js"); + const entryName = client.entry ?? "index.js"; + const entry = join(buildDir, entryName); if (!existsSync(entry)) { failures.push( - `${client.name}: ${client.build}/index.js is missing — run \`npm run build\` first.`, + `${client.name}: ${client.build}/${entryName} is missing — run \`npm run build\` first.`, ); continue; } diff --git a/scripts/verify-format-coverage.mjs b/scripts/verify-format-coverage.mjs index c7abd3fa78..f617cac5eb 100644 --- a/scripts/verify-format-coverage.mjs +++ b/scripts/verify-format-coverage.mjs @@ -51,6 +51,7 @@ const MANIFESTS = [ ".", "clients/web", "clients/cli", + "clients/mcpdo", "clients/tui", "clients/launcher", ]; diff --git a/scripts/verify-skills.mjs b/scripts/verify-skills.mjs index 891c6eda4a..5d4addfbd3 100755 --- a/scripts/verify-skills.mjs +++ b/scripts/verify-skills.mjs @@ -312,6 +312,26 @@ function main(argv = process.argv.slice(2)) { ); } + // Shipped skills (`skills/`, packaged with a client rather than loaded from + // `.claude/skills`) get the same frontmatter/structure validation — a + // truncated SKILL.md would otherwise ship silently — but no eval-case or + // listing-budget requirements: they are not part of this repo's own + // agent skill listing. + if (!override && existsSync(path.join(ROOT, "skills"))) { + const shippedDir = path.join(ROOT, "skills"); + for (const dir of skillDirs(shippedDir)) { + const file = path.join(shippedDir, dir, "SKILL.md"); + if (!existsSync(file)) { + failures.push(`skills/${dir}: no SKILL.md`); + continue; + } + const skill = parseSkill(dir, readFileSync(file, "utf8")); + for (const e of skill.errors) { + failures.push(`skills/${dir}/SKILL.md: ${e}`); + } + } + } + if (failures.length > 0) { console.error(`verify:skills — ${failures.length} problem(s):\n`); for (const f of failures) console.error(" " + f); diff --git a/scripts/verify-test-timeouts.mjs b/scripts/verify-test-timeouts.mjs index 59f0394ac8..3363ef9439 100644 --- a/scripts/verify-test-timeouts.mjs +++ b/scripts/verify-test-timeouts.mjs @@ -93,6 +93,7 @@ export const EXPECTED_PROJECTS = Object.freeze({ cli: EXPECTED_TIMEOUTS, tui: EXPECTED_TIMEOUTS, launcher: EXPECTED_TIMEOUTS, + mcpdo: EXPECTED_TIMEOUTS, }); /** @@ -108,6 +109,7 @@ export const CONFIG_ROOTS = Object.freeze([ { root: "clients/cli", projects: ["cli"] }, { root: "clients/tui", projects: ["tui"] }, { root: "clients/launcher", projects: ["launcher"] }, + { root: "clients/mcpdo", projects: ["mcpdo"] }, ]); /** diff --git a/scripts/verify-test-timeouts.test.mjs b/scripts/verify-test-timeouts.test.mjs index 8c5a727857..2f6add7390 100644 --- a/scripts/verify-test-timeouts.test.mjs +++ b/scripts/verify-test-timeouts.test.mjs @@ -123,6 +123,7 @@ test("a Vitest config this guard does not check is an error", () => { "clients/cli", "clients/tui", "clients/launcher", + "clients/mcpdo", "clients/desktop", ]); assert.equal(failures.length, 1); @@ -131,7 +132,7 @@ test("a Vitest config this guard does not check is an error", () => { test("a stale row naming a config that no longer exists is an error", () => { const failures = checkConfigRootCoverage(["clients/web", "clients/cli"]); - assert.equal(failures.length, 2); + assert.equal(failures.length, 3); for (const f of failures) assert.match(f, /has no Vitest config on disk/); }); diff --git a/skills/mcpdo/SKILL.md b/skills/mcpdo/SKILL.md new file mode 100644 index 0000000000..790be6226f --- /dev/null +++ b/skills/mcpdo/SKILL.md @@ -0,0 +1,165 @@ +--- +name: mcpdo +description: Access MCP (Model Context Protocol) servers and their tools, resources, and prompts through the mcpdo CLI — connections it holds extend your capabilities alongside any built-in MCP support. Use this skill for any question or task about MCP servers, connections, or tools (e.g. "what MCP servers am I connected to?", "what MCP tools do I have?"), and whenever a task requires inspecting, testing, or scripting against an MCP server (stdio or HTTP) rather than writing custom client code. +disable-model-invocation: false +--- + +# mcpdo — MCP Inspector connection CLI + +Connect to an MCP server once, then run many commands against that named +connection. + +```bash +mcpdo servers/list # catalog entries you can connect +mcpdo connect entry-name # connect a catalog entry +mcpdo connect https://example.com/mcp # ad-hoc HTTP/SSE target +mcpdo connect node server.js # ad-hoc stdio target + +mcpdo connections/list # open connections +mcpdo @entry-name tools/list +mcpdo @entry-name tools/call <toolName> arg:=value +mcpdo @entry-name resources/list +mcpdo @entry-name resources/read <uri> +mcpdo @entry-name prompts/list + +mcpdo --connection entry-name tools/list # flag form of @entry-name + +mcpdo disconnect entry-name +``` + +The canonical flow is `servers/list` → `connect <entry>` → `@entry <command>`. +Run `mcpdo help` or `mcpdo <command> --help` for the full, authoritative list of +commands and flags (auth management, elicitation controls, daemon control, and +more). + +## How to think about mcpdo + +- **Connections extend your toolset.** Treat every open connection's tools, + resources, and prompts as part of your available capabilities: when facing a + task (or before saying "I can't do that"), check `mcpdo connections/list` + and the connected servers' `tools/list`, and call those tools mid-task like + any built-in tool. +- **Answer capability questions with it.** "What MCP servers/tools do I have?" + includes mcpdo connections (`connections/list`) and catalog entries that are + connectable but not connected (`servers/list`). +- **Don't auto-connect.** Catalog entries don't describe capabilities, and + connecting can start an auth flow involving the human. Connect when the user + directs it, or when a task clearly needs it — and say so first. +- **Inspect via commands, not the filesystem.** Entries, connections, auth + state, and daemon state all have read commands; never list or read + `~/.mcp-inspector` directly. The one exception is _editing_ the catalog + (below). +- **Make it always-on (optional).** Skills load only on demand; for standing + awareness of mcpdo in a project, append `mcpdo agent-help --instructions` + output to the project's `CLAUDE.md`/`AGENTS.md`. (`mcpdo agent-help` + prints this guide; `--skill-path` prints the installable skill file's + path.) + +## The catalog + +- `servers/list` / `servers/show <name>` read the **catalog**: the writable + entry file at `~/.mcp-inspector/mcp.json` (standard `mcpServers` shape), + overridable per shell via `--catalog <path>` or `MCP_CATALOG_PATH`. + `servers/list` prints the resolved source path. +- `servers/*` shows entries on disk; `connections/*` shows live daemon state. + Two shells with different catalogs share the same connections. +- There are no CLI edit commands, by design: add or remove entries by editing + the catalog file directly. `mcpdo connect entry --config path/to/mcp.json` + instead connects an entry from a read-only foreign config file. + +## Conventions + +- `--format json` outputs JSON; the default, `--format text`, is + human-readable. +- Always qualify commands with `@name` or `--connection <name>` (shorthand + `--conn`) from an agent shell: with non-interactive (non-TTY) stdin, mcpdo + requires an explicit connection and errors without one. Omitting it falls + back to the most-recently-used connection only on an interactive TTY, or + anywhere when `MCP_ALLOW_DEFAULT_CONNECTION=1` is set. +- Connections persist across separate `mcpdo` invocations and **self-heal**: a + dropped transport (expired session, exited stdio child) transparently + re-dials on next use with stored credentials. Don't monitor or reconnect + manually; only an `auth_required` error needs action (re-run `connect`). + `mcpdo disconnect` ends one connection; `mcpdo daemon stop` resets + everything. +- The `[legacy]` / `[modern]` era tag on `connections/list` is the negotiated + protocol generation (`legacy` = classic `initialize` handshake — current and + fine, not deprecated). Informational only. + +## Auth + +- Auth is automatic at connect time and stored for reuse (`mcpdo auth/list` / + `mcpdo auth/clear`). `auth/list` shows each stored URL with the catalog or + connection names it is "known as" and a `● live` marker when a current + connection holds it, so the store entries line up with the servers you use. + `auth/clear` accepts either the store URL or one of those friendly names + (`mcpdo auth/clear hosted-everything`); a name that resolves to no URL (a + stdio server) or to more than one URL is rejected with guidance. EMA IdP + login records lead with their bare issuer URL and carry a trailing + `enterprise IdP login` marker, and clear by that bare issuer URL + (`mcpdo auth/clear https://idp.example.com`). Run `connect` as its **own + command** — never chained with `&&`, `;`, or a follow-on + `tools/list`/`connections/show`. A connect that needs sign-in exits 0 while + the connection is still unusable, so a chained command runs against a + not-yet-connected server, errors, and clutters the output you relay to the + user. Connect, check whether it returned `pendingAuth`, and only then decide + the next step. When a + browser sign-in is needed and stdin is non-TTY, + `connect` exits 0 immediately with `pendingAuth: true` and an `authUrl`. + The connection is **not** connected and **not** usable until the user signs + in — exit 0 is not success here. Show that `authUrl` to the user as literal + plain text as the **last thing in your reply**, ask them to open it, sign in, + and tell you when they're done, then **end your turn and wait**. Do not retry + the command, poll, sleep, or say you are connected: the user can't see your + message until the turn ends, so any further tool call in the same turn only + delays the link reaching them. Once the user says they have signed in, run + `connections/show @name` to complete the sign-in, then continue the original + task. (`connections/show @name` completes a finished sign-in itself and + reprints the `authUrl` if you need to show it again; `connections/list` stays + read-only but reports `pendingAuthSignedIn: true` — "signed in — completing on + next use" — once the user's part is done.) Never reconnect to fix a pending + sign-in. +- To force a fresh sign-in on an already-open connection, `mcpdo disconnect + <name> --clear-auth` (`-c`) tears it down and clears its stored tokens in one + step, so the next plain `connect` re-triggers the browser flow. Prefer it over + a separate `disconnect` + `auth/clear <url>` — it takes the connection name + (not the URL) and closes the window where the live connection keeps working on + in-memory tokens. `connect <name> --relogin` (`-r`) is the equivalent when you + are reconnecting anyway. +- Enterprise-managed auth (EMA) works the same way. `mcpdo auth/ema-login` + from a non-TTY shell exits 0 immediately with `pendingLogin: true` and an + `authUrl`. Show that `authUrl` to the user as literal plain text as the last + thing in your reply, ask them to sign in and tell you when they're done, then + end your turn and wait — do not poll while they are signing in. Once they + confirm, run `mcpdo auth/ema-status`; `loginState` reads `logged_in` once the + sign-in has completed in the background. After that, connects to EMA servers + mint + tokens silently with no further sign-in. Connecting to an EMA server + *without* a prior IdP login parks like any other pending sign-in, with the + IdP link as its `authUrl`. `mcpdo auth/ema-logout` clears local EMA state + only; when the IdP advertises an end-session endpoint the output includes a + URL to end the IdP browser session too — relay it to the user, who may + ignore it if they only meant to reset local state. +- Tokens live in the OS keychain when one is available. On keychain-less + hosts mcpdo falls back to the shared secrets file (never the in-memory + store — mcpdo is multi-process, so a per-process store can't carry a token + from the sign-in helper to the daemon). `MCP_INSPECTOR_SECRET_STORE` + (`keyring|file|memory`) still overrides explicitly. + +## Elicitations (server asks a question mid-call) + +- On an interactive TTY, mcpdo prompts inline. From an agent shell (non-TTY or + `--format json`), the **command returns immediately** (exit 0) with an + `elicitationPending` payload carrying the question, schema, and an + `elicitationId`; the underlying MCP **tool call stays parked** on the daemon + awaiting your response. Never wait on or time-box the mcpdo command itself — + it has already exited; the pending work lives daemon-side. +- Answer with `mcpdo elicitation/respond <elicitationId> field:=value ...` + (repeat if the server asks again), or end it with `--decline` or `--cancel`. + For URL-mode elicitations, relay the URL to the user, then confirm with + `elicitation/respond <id> --done` (or `--cancel`; URL mode has no decline). + The response returns the + final tool result. +- Parked calls expire after 10 minutes; one parked call per connection. Pass + `--elicit off` on `connect` to have well-behaved servers fall back to their + own defaults instead of asking. diff --git a/specification/v2_auth_ema.md b/specification/v2_auth_ema.md index 1ceaa33e59..2f5abfa9c2 100644 --- a/specification/v2_auth_ema.md +++ b/specification/v2_auth_ema.md @@ -300,7 +300,7 @@ Install-level IdP credentials are edited in **Client Settings**, separate from p The next connect to a cleared EMA server has no cached access token, so a **401** triggers leg 1 (IdP login) via `authenticate()`. - **Web OAuth store singleton:** web uses `getWebRemoteOAuthStorage()` (`clients/web/src/lib/remoteOAuthStorage.ts`) — memoized `RemoteOAuthStorage` backed by `/api/storage/oauth` — so Client Settings sign-out, EMA IdP session, connect, and per-server clear all mutate the same in-memory `OAuthStorageBase` view as the active `InspectorClient`. -**Future (not implemented):** explicit **Sign out** may additionally invoke the IdP's OIDC **end-session** / logout endpoint (RP-initiated logout) so the IdP SSO cookie is cleared — not just inspector-local IdP/resource token state. Today sign-out is **local-only** (clear `idpSessions`, tagged EMA resource entries, and leg-1 PKCE); the IdP may still treat the browser as signed in and skip the login prompt on the next authorize redirect until the IdP session expires or the user signs out at the IdP. +**End-session URL (implemented):** `clearEmaIdpSession` returns `{ endSessionUrl? }` — the IdP's OIDC **end-session** (RP-initiated logout) URL with `id_token_hint`, built from the discovery metadata cached at login under the leg-1 key (both the session's idToken and that cache are read *before* the clear destroys them; no network is involved, and a missing session, missing cache, or missing `end_session_endpoint` just omits the URL). For the cache to carry the field at all, IdP discovery at login fetches the **OpenID Connect Discovery** document (`/.well-known/openid-configuration`) first — `end_session_endpoint` is an OIDC-only field, and the SDK's RFC 8414-first discovery would cache an OAuth metadata document that legitimately omits it (`discoverIdpMetadata` falls back to SDK discovery when the OIDC document is unavailable; sessions cached before this change need one re-login to pick the field up). The clear itself stays local-only — the URL is for the caller to relay so the user can also end the IdP's browser SSO session. `mcpdo auth/ema-logout` prints it; web/TUI sign-out does not surface it yet, so there the IdP may still treat the browser as signed in and skip the login prompt on the next authorize redirect until the IdP session expires or the user signs out at the IdP. Auto-*invoking* the endpoint (with `post_logout_redirect_uri` confirmation) remains future work. **Copy / UX notes:** User-facing text explains enterprise-managed authorization in plain language (org IdP sign-in vs each server's OAuth login). It does **not** reference protocol jargon (`leg 1`, resource AS, etc.) or storage filenames in the form. Per-server EMA enablement and MCP-server OAuth credentials remain in **Server Settings** (see below). @@ -518,7 +518,7 @@ Design decisions for EMA are complete. Remaining work is the phased plan and che - [x] **Web shared OAuth store** — `RemoteOAuthStorage` / `oauth.json` for web + CLI + TUI parity (#1548); see §Shared file-backed OAuth state - [ ] Client profile persistence (migrate from `client.json`; may extend or replace web Client Settings) - [ ] Optional: adopt `@modelcontextprotocol/client` v2 Layer-2 helpers for legs 2–3 (replace `wire.ts`) or full v2 transport for EMA -- [ ] Optional: IdP **end-session** / RP-initiated logout on explicit **Sign out** (today sign-out clears inspector-local IdP session and tagged EMA resource tokens only; IdP browser SSO may remain active) +- [x] Partial: IdP **end-session** URL on explicit **Sign out** — `clearEmaIdpSession` returns the RP-initiated logout URL (`end_session_endpoint` + `id_token_hint`, from the login-time metadata cache); `mcpdo auth/ema-logout` prints it. Still open: web/TUI surfacing it, and optionally auto-invoking the endpoint with `post_logout_redirect_uri` confirmation --- diff --git a/specification/v2_auth_mid_session.md b/specification/v2_auth_mid_session.md index cf714040df..c022bedae6 100644 --- a/specification/v2_auth_mid_session.md +++ b/specification/v2_auth_mid_session.md @@ -141,7 +141,7 @@ RemoteClientTransport StreamableHTTPClientTransport | Web remote client | `core/mcp/remote/remoteClientTransport.ts`, `core/mcp/inspectorClient.ts` | | Web app | `clients/web/src/App.tsx`, `lib/oauthResume.ts`, `utils/pendingReauth.ts`, `lib/browserTabVisibility.ts`, `components/groups/StepUpAuthModal/` | | TUI | `clients/tui/src/App.tsx`, `utils/tuiOAuth.ts` | -| CLI | `clients/cli/src/cliOAuth.ts` | +| CLI | `core/cli/cliOAuth.ts` | | Runner OAuth (TUI/CLI) | `core/auth/node/runner-interactive-oauth.ts`, `oauth-callback-server.ts` | | Step-up test fixture | `test-servers/configs/oauth-step-up-demo.json`, `test-servers/src/test-server-oauth.ts` | @@ -393,7 +393,7 @@ CLI never spawns TUI/web for auth — completes locally or fails. | Reauth for affected server | `handleAuthRecoveryRequired()` switches server when needed | | Clear OAuth | Auth tab **S** | -### CLI UX (`clients/cli/src/cliOAuth.ts`) +### CLI UX (`core/cli/cliOAuth.ts`) | Situation | Behavior | | ------------------------------ | -------------------------------------------------------------------------------- | diff --git a/specification/v2_catalog_launch_config.md b/specification/v2_catalog_launch_config.md index 6fa0b9a7e1..6b23307276 100644 --- a/specification/v2_catalog_launch_config.md +++ b/specification/v2_catalog_launch_config.md @@ -512,7 +512,7 @@ G1, G4, and launcher details: [v2_cli_tui_launcher.md](v2_cli_tui_launcher.md). | [#1183](https://github.com/modelcontextprotocol/inspector/issues/1183) — auto-connect | Open | UC5 web ergonomics | | [#1348](https://github.com/modelcontextprotocol/inspector/issues/1348) — import from other clients | Open | UC2 web UI | | [#1435](https://github.com/modelcontextprotocol/inspector/issues/1435) — registry import | Open | UC2 registry path | -| [#1432](https://github.com/modelcontextprotocol/inspector/issues/1432) — CLI v2 | Open | Session CLI umbrella | +| [#1432](https://github.com/modelcontextprotocol/inspector/issues/1432) — CLI v2 | Closed (#1783) | Connection CLI umbrella — as-built: [v2_cli_v2.md](v2_cli_v2.md); `mcpdo` daemon CLI shipped in [#1783](https://github.com/modelcontextprotocol/inspector/pull/1783) | | [#1352](https://github.com/modelcontextprotocol/inspector/pull/1352) / [#1358](https://github.com/modelcontextprotocol/inspector/pull/1358) | Merged | Flat settings on disk | | [#1356](https://github.com/modelcontextprotocol/inspector/pull/1356) | Merged | Secrets in keychain | diff --git a/specification/v2_cli_tui_launcher.md b/specification/v2_cli_tui_launcher.md index fb42b296bd..804bd52937 100644 --- a/specification/v2_cli_tui_launcher.md +++ b/specification/v2_cli_tui_launcher.md @@ -6,7 +6,7 @@ ## Summary -v2 ships three non-web Inspector incarnations alongside the web client: a **one-shot CLI**, an **interactive TUI**, and a **launcher** that routes to web, CLI, or TUI from a single `mcp-inspector` binary. All three consume the same `core/` source as the web client via the `@inspector/core` path alias and run on the shared `InspectorClient` stack ported from v1.5/main. +v2 ships four non-web Inspector incarnations alongside the web client: a **one-shot CLI**, an **interactive TUI**, an experimental **connection CLI** (`mcpdo`, `clients/mcpdo/`), and a **launcher** that routes to web, CLI, or TUI from a single `mcp-inspector` binary. All four consume the same `core/` source as the web client via the `@inspector/core` path alias and run on the shared `InspectorClient` stack ported from v1.5/main. This document describes how those clients are built, wired, and tested today, and records known gaps. For catalog vs launch-time config semantics (`--config`, `--catalog`, import), see [Catalog and Launch Configuration](v2_catalog_launch_config.md). @@ -20,7 +20,7 @@ This document describes how those clients are built, wired, and tested today, an ## Non-goals -- **CLI v2 sessions** (connect once, many subcommands) — tracked separately in [#1432](https://github.com/modelcontextprotocol/inspector/issues/1432). +- **CLI v2 connections** (connect once, many subcommands) — as-built in [v2_cli_v2.md](v2_cli_v2.md) (`mcpdo` bin connection-first; `mcp-inspector --cli` stays one-shot); tracked by [#1432](https://github.com/modelcontextprotocol/inspector/issues/1432). - **npm workspaces** — v2 uses a fat root package plus per-client `package.json` for dev dependencies; the launcher resolves sibling `build/` outputs via relative paths, not workspace hoisting. - _Why not workspaces:_ `core/` is consumed by **bundling** — a Vite alias for the browser, tsup inlining for the Node clients — not by symlinked package resolution, so workspaces' main benefit (cross-package linking) does not apply. Each client also pins `react` / `@modelcontextprotocol/sdk` to its own `node_modules` (see `vitest.shared.mts`) to avoid dual-package-instance hazards, which hoisting works against. And the published `@modelcontextprotocol/inspector` is a single flat fat package that workspaces would complicate rather than simplify. - _Cost (from-source dev only):_ there is no hoisting, so each client keeps its own `node_modules`. A root `postinstall` (`scripts/install-clients.mjs`) cascades `npm install` into every client, so a single `npm install` at the repo root populates them all — re-run it after a pull that changes a client's dependencies. The cascade no-ops outside a source checkout (it exits early when running from `node_modules`, and the published tarball ships only each client's `build/`, no client `package.json`), so end users of the published package are unaffected. Set `INSPECTOR_SKIP_CLIENT_INSTALL=1` to skip the cascade (e.g. CI that installs each client itself). @@ -34,7 +34,8 @@ This document describes how those clients are built, wired, and tested today, an | Artifact | Path | Build | Published bin | | ---------- | ------------------------------- | ------------------------------------------------------ | -------------------------------------------------------- | | Launcher | `clients/launcher/` | `tsc` → `build/index.js` | Root `mcp-inspector` → `clients/launcher/build/index.js` | -| CLI | `clients/cli/` | `tsup` → `build/index.js` | `mcp-inspector-cli` (client package only) | +| CLI | `clients/cli/` | `tsup` → `build/index.js` | `mcp-inspector-cli` (client package only; one-shot) | +| mcpdo | `clients/mcpdo/` | `tsup` → `build/mcp-bin.js` + `build/mcpdod.js` | `mcpdo` (experimental; ships in the inspector package) | | TUI | `clients/tui/` | `tsup` → `build/index.js` | `mcp-inspector-tui` (client package only) | | Web runner | `clients/web/server/run-web.ts` | `tsup` (`build:runner`) → `clients/web/build/index.js` | `mcp-inspector-web` (client package only) | @@ -71,22 +72,22 @@ Root scripts `inspector`, `web`, and `web:dev` are thin wrappers around the laun ## Shared core consumption -All three clients import from `@inspector/core/...` (mapped to `../../core/` source). +All four clients import from `@inspector/core/...` (mapped to `../../core/` source). -| Concern | Web | CLI / TUI | Launcher | +| Concern | Web | CLI / mcpdo / TUI | Launcher | | -------------- | --------------------------------- | -------------------------------------------------------- | ------------------------- | | Dev typecheck | `tsconfig.app.json` paths | per-client `tsconfig.json` paths | `tsconfig.json` (no core) | | Runtime bundle | Vite alias | tsup `noExternal: [/^@inspector\/core/]` + esbuild alias | n/a | | Tests | Vitest projects in `clients/web/` | Vitest + `vitest.shared.mts` aliases | none | -`vitest.shared.mts` at repo root centralizes `@inspector/core` and test-server aliases plus bare-module pins (`react`, `pino`, SDK, etc.) so CLI/TUI Vitest configs stay aligned with web. +`vitest.shared.mts` at repo root centralizes `@inspector/core` and test-server aliases plus bare-module pins (`react`, `pino`, SDK, etc.) so CLI/mcpdo/TUI Vitest configs stay aligned with web. **Resolved design choices:** | Topic | Decision | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | | Core package | No separate `inspector-core` npm package; source-only `core/` | -| CLI/TUI build | tsup bundles `@inspector/core` into `build/index.js` | +| CLI/TUI build | tsup bundles `@inspector/core` into each client's `build/` (CLI/TUI `index.js`; mcpdo `mcp-bin.js` + `mcpdod.js`) | | Core tests | Not duplicated under cli/tui; web unit + integration suites cover `core/` | | Default config path | `loadServerEntries()` applies `withDefaultCatalogPath()` → `~/.mcp-inspector/mcp.json` when no `--catalog`/`--config` and no ad-hoc target | @@ -94,7 +95,7 @@ All three clients import from `@inspector/core/...` (mapped to `../../core/` sou ## CLI -**Model:** one-shot — each invocation connects, runs a single `--method`, prints JSON to stdout, disconnects, exits. Same surface as v1.5; session-oriented CLI v2 is future work ([#1432](https://github.com/modelcontextprotocol/inspector/issues/1432)). +**Model:** one-shot — each invocation connects, runs a single `--method`, prints a result to stdout, disconnects, exits. Same surface as v1.5. Connection-oriented CLI v2 (`mcpdo`) is documented as-built in [v2_cli_v2.md](v2_cli_v2.md) ([#1432](https://github.com/modelcontextprotocol/inspector/issues/1432)). **Entry:** `clients/cli/src/index.ts` exports `runCli(argv)`; `src/cli.ts` owns Commander parsing and `InspectorClient` orchestration. diff --git a/specification/v2_cli_v2.md b/specification/v2_cli_v2.md new file mode 100644 index 0000000000..78e4ddcee6 --- /dev/null +++ b/specification/v2_cli_v2.md @@ -0,0 +1,180 @@ +# Inspector CLI v2 (connection-oriented) + +### [Brief](README.md) | [V1 Problems](v1_problems.md) | [V2 Scope](v2_scope.md) | [V2 Tech Stack](v2_web_client.md) | [V2 UX](v2_ux.md) | [V2 Auth](v2_auth.md) | [V2 New Spec Impact](v2_new_spec_impact.md) + +#### [CLI, TUI, Launcher](v2_cli_tui_launcher.md) | CLI v2 | [Catalog / launch config](v2_catalog_launch_config.md) + +Documentation of the **experimental** connection-oriented Inspector CLI (`mcpdo`) and how it relates to the frozen one-shot path (`mcp-inspector --cli`). Tracked by [#1432](https://github.com/modelcontextprotocol/inspector/issues/1432). `mcpdo` is a separate client under `clients/mcpdo/`, shipped as the `mcpdo` bin in `@modelcontextprotocol/inspector` (experimental). + +**Related:** [CLI, TUI, and Launcher](v2_cli_tui_launcher.md), [Catalog and Launch Configuration](v2_catalog_launch_config.md), [Storage](v2_storage.md), [Auth](v2_auth.md), [`clients/mcpdo/README.md`](../clients/mcpdo/README.md), [`clients/cli/README.md`](../clients/cli/README.md) (one-shot). + +--- + +## Overview + +| | **One-shot** | **Connection** | +| ---------- | ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------- | +| Entrypoint | `mcp-inspector --cli` | `mcpdo` | +| Lifecycle | Connect → one `--method` → disconnect | Connect once → many subcommands → disconnect | +| Process | In-process only | Short-lived front-end + implicit connection daemon (IPC) | +| Package | `clients/cli` (ships with `@modelcontextprotocol/inspector`) | `clients/mcpdo` (experimental; ships the `mcpdo` bin with `@modelcontextprotocol/inspector`) | + +Both use `@inspector/core` `InspectorClient` and shared `core/cli/handlers/run-method.ts` (#2461). One-shot never starts the daemon. `mcpdo` does not accept `--method`. + +```bash +mcpdo servers/list --config mcp.json +mcpdo servers/show my-server --config mcp.json +mcpdo connect myserver --config mcp.json +mcpdo tools/list +mcpdo tools/call search query:=hello +mcpdo @other resources/list +mcpdo disconnect +``` + +Optional private daemon for one shell (`ssh-agent` style): + +```bash +eval "$(mcpdo private)" +mcpdo connect myserver --config mcp.json +mcpdo tools/list +``` + +--- + +## As-built + +### Entrypoints and layout + +| Piece | Location | +| -------------------- | ------------------------------------------------------------------------------------------------------------------------------ | +| One-shot | `clients/cli/src/cli.ts`, `index.ts`; interactive OAuth in `core/cli/cliOAuth.ts` (shared with mcpdo) | +| Connection front-end | `clients/mcpdo/src/connection/` (`mcp.ts`, `dispatch.ts`, `authorize.ts`, `format-*.ts`, `private-env.ts`) + `mcp-bin.ts` | +| Daemon | `clients/mcpdo/src/daemon/` → `clients/mcpdo/build/mcpdod.js` | +| Shared handlers | `core/cli/handlers/` (`run-method.ts`, `method-types.ts`, `servers-list.ts`, …); one-shot-only output in `clients/cli/src/handlers/` (`emit-result.ts`, …) | + +``` +mcp-inspector --cli … mcpdo … + │ │ + ▼ ▼ + clients/cli clients/mcpdo + cli.ts connection/mcp.ts + │ │ NDJSON IPC + │ daemon (build/mcpdod.js) + └──────────┬─────────────┘ + ▼ + clients/cli handlers/run-method.ts → InspectorClient +``` + +### One-shot (`mcp-inspector --cli`) + +Frozen automation contract. Each invocation: resolve server → connect → `runMethod` → print → disconnect. Never uses the connection daemon. + +| `--method` | Notes | +| ----------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | +| `initialize`, `tools/list`, `tools/call`, `resources/list`, `resources/read`, `resources/templates/list`, `prompts/list`, `prompts/get`, `logging/setLevel` | Core one-shot surface (`ONE_SHOT_METHODS`) | +| `servers/list`, `servers/show` | Catalog only (no MCP connect); `servers/show` needs `--server` | + +Anything else (e.g. `logging/tail`, `resources/subscribe`, `tasks/*`, `roots/*`) is a **usage error before connect** — one-shot must not hang on stream outcomes. + +**Output:** `--format text` = pretty JSON of bare result; `json` = `{ result[, appInfo] }` envelope. Exit codes `0`–`5` + stderr `ErrorEnvelope`. + +**Auth:** Interactive OAuth + mid-session recovery in-process (`cliOAuth.ts`); `--stored-auth-only`, `--use-stored-auth`, handoff flags. See [clients/cli/README.md](../clients/cli/README.md). + +### Connection CLI (`mcpdo`) + +#### Commands + +| Category | Commands | +| ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Catalog | `servers/list`, `servers/show <name>` | +| Connection | `connect` (`--relogin`), `disconnect`, `connections/list`, `connections/use` | +| Auth store | `auth/list`, `auth/clear` / `auth/clear --all` | +| Daemon | `private`, `daemon status`, `daemon stop` | +| MCP | `tools/list`, `tools/call`, `resources/*`, `prompts/*`, `logging/setLevel`, `logging/tail`, `tasks/*`, `roots/list`, `roots/set` (`initialize` is deliberately not registered — connection metadata comes from `connections/show`) | + +**Globals (before subcommand):** `--format text|json`, `--plain`, `--connection <name>`, `--catalog` / `--config`, `--stored-auth-only`. + +**Connection select:** leading `@name` and/or `--connection <name>`. Tool args: `key:=value`, inline JSON, or `--tool-arg` / `--tool-args-json`. + +**Connect forms:** catalog entry / `--server` / ad-hoc URL or command; optional `@name` to override connection name (default = entry id). + +#### Output + +| Flag | Behaviour | +| ------------------------- | ----------------------------------------------------------------------------------------------- | +| `--format text` (default) | Human-readable. On a TTY: ANSI color / bold / dim / OSC 8 links unless `--plain` or `NO_COLOR`. | +| `--format json` | Pretty-printed payload (**no** `{ result }` envelope; never ANSI). | +| Streams | Long-lived until Ctrl-C; human lines or pretty JSON events per `--format`. | + +#### Default connection (MRU) + +- Omit `@name` / `--connection` → MRU (TTY). +- Explicit `@name` / `--connection` always wins. +- Non-TTY: require explicit connection unless `MCP_ALLOW_DEFAULT_CONNECTION=1`. +- `connections/list`, `connections/use <name>`; `daemon status` / `connections/list` do **not** auto-spawn the daemon. + +#### Daemon + +**IPC ops:** `ping`, `connect`, `disconnect`, `connections/list`, `connections/use`, `connections/show`, `daemon/status`, `daemon/stop`, `rpc`, `stream`. + +- One `InspectorClient` per named connection; auto-spawn on first need; idle exit ~60s after last disconnect **or** after a connection-less spawn with no successful connect; `daemon stop` tears down immediately. +- Socket/lock mode `0600` (best-effort). Config (incl. secrets) over IPC after listen — not on daemon argv. +- Errors that are not already `CliExitCodeError` go through `classifyError` (exit-code parity with one-shot). + +| Context | Path | +| -------------------------- | ------------------------------------------------------------------------------------------------------------------------- | +| Shared default | `~/.mcp-inspector/mcpdod.sock` (+ `mcpdod.lock`, `mcpdod.token`, `mcpdod.log`) | +| `MCP_STORAGE_DIR` | Socket/lock under that dir (CI isolation; same family as `oauth.json`) | +| `MCP_INSPECTOR_DAEMON_DIR` | Wins over storage dir when set (spawn pin / private) | +| Private | `$TMPDIR/mcp-conn-<uid>/<id>/` (0700, short id — `sun_path` caps socket paths at 104 bytes on macOS) from `mcpdo private` | + +| Mode | Trust | +| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Shared (default)** | Auto-generated token, published to `mcpdod.token` (0600) in the daemon dir (0700). Same-UID peer that can read the dir can drive connections (intentional cross-terminal share); there is no unauthenticated request path. | +| **Private** | `eval "$(mcpdo private)"` exports `MCP_INSPECTOR_DAEMON_DIR` + `MCP_INSPECTOR_DAEMON_TOKEN`. Daemon requires the token on every request. OAuth store remains shared unless the user also sets `MCP_STORAGE_DIR`. Daemon starts lazily on first IPC. | + +#### Auth (connection) + +- Same `oauth.json` store as other Inspector clients. +- **Connect-time:** daemon connect → on `auth_required`, front-end `authorizeInFrontend()` (unless `--stored-auth-only`) → retry connect. +- **`--relogin`:** clear any stored OAuth for the server URL before connect; interactive login still runs only if auth is required afterward. No-op for stdio / targets with no URL-keyed store entry (do not reject — same semantics, nothing to clear). +- **Mid-session** step-up during `rpc` / `stream`: **not implemented** (see To-do). Use one-shot, or disconnect / re-auth / reconnect. +- Connection `connect` does not expose one-shot OAuth flags (`--client-id`, `--callback-url`, …); env / defaults / `MCP_OAUTH_CALLBACK_URL` only. + +#### One-shot ↔ connection mapping + +| One-shot | Connection | +| ---------------------------------------------------------- | ------------------------------------------------------------ | +| `… --catalog mcp.json --server s --method tools/list` | `mcpdo connect --catalog mcp.json s` then `mcpdo tools/list` | +| `… --method tools/call --tool-name X --tool-args-json '…'` | `mcpdo tools/call X key:=val` / `'{"…"}'` | +| `… --method servers/list` | `mcpdo servers/list` | +| `… --method servers/show --server <name>` | `mcpdo servers/show <name>` | + +### Testing + +| Client | Runner | Coverage | +| ------------------------------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | +| One-shot (`clients/cli`) | In-process `runCli()`; thin binary e2e | Per-file ≥90 on `clients/cli/src` + `core/cli`. Exclusion: `src/index.ts`. | +| Connection CLI (`clients/mcpdo`) | In-process `runMcp()`; daemon IPC + stream + private-token tests | Per-file ≥90 on `clients/mcpdo/src`. Exclusions: `mcp-bin.ts`, `daemon/run.ts` (bootstraps only). | + +Both are wired into root `validate` / `coverage`. + +--- + +## To-do + +| Item | Notes | +| ----------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Mid-session auth over IPC** | Challenge + step-up UX on the invoking `mcpdo` during `rpc`/`stream`. Connect-time only today. | +| **Per-socket request serialization** | Requests on one connection are handled as lines arrive (single line capped at 1 MiB); safe while clients use one request per connection. | +| **Shared `createCliInspectorClient`** | Daemon / authorize / one-shot construct clients separately. | +| **Split `registerRpcCommands`** | Large Commander switch in `connection/mcp.ts`. | +| **`mcpdo daemon run`** | Optional foreground debug (not a Commander subcommand; `build/mcpdod.js` works today). | +| **Launcher help polish** | Make `mcpdo` vs `--cli` unmistakable in launcher `--help` / docs. | +| **Connection `connect` OAuth flag parity** | One-shot has `--client-id` / `--callback-url` / handoff; connection authorize uses defaults / env only. | +| **Peer-cred / stronger private IPC** | Private mode uses bearer token; optional OS peer checks beyond that. | +| **Stream fan-out / `mcpdo attach`** | One consumer per stream invocation today. | +| **Sampling CLI** | Still TUI/web. mcpdo handles server-driven _elicitation_ (URL + form modes, `--elicit` capability override) since #1783; sampling remains unimplemented. Behavior as shipped: on an interactive TTY, mcpdo prompts inline; any non-TTY or `--format json` caller gets the elicitation **parked** — the command returns immediately with an `elicitationPending` payload and the caller answers later via `elicitation/respond <id>` (or `--decline` / `--cancel`). Nothing auto-declines. URL mode never auto-accepts: completion is only confirmed by an explicit answer. | +| **Ephemeral no-`connect` shortcuts on `mcpdo`** | Out of scope (keep two mental models). | +| **`MCP_SESSION` env** | Superseded by require-explicit-on-non-TTY + `MCP_ALLOW_DEFAULT_CONNECTION=1`. | +| **Human `--full` schema dumps** | Optional formatter polish. | diff --git a/test-servers/src/modern-tasks.ts b/test-servers/src/modern-tasks.ts index 15531921e1..e449d84dfe 100644 --- a/test-servers/src/modern-tasks.ts +++ b/test-servers/src/modern-tasks.ts @@ -60,17 +60,24 @@ interface ModernTaskEntry { inputSatisfied?: boolean; /** The `inputResponses` the client submitted, echoed back in the result. */ inputResponses?: Record<string, unknown>; + /** How many input rounds this task has already surfaced. Loops key their + * `inputRequests` by this so each round is a *distinct* request — the + * ext-tasks SDK fingerprints request keys and skips a recurring one, so a + * reused key would be read as the already-answered round and the loop would + * never advance past round 1. */ + inputRound: number; } /** The embedded elicitation an `input_required` modern task surfaces. Shaped as * a standalone `elicitation/create` request so the client's pending-request UI - * (reused from the MRTR path) renders it and returns an `ElicitResult`. */ -function confirmInputRequests(): Record<string, unknown> { + * (reused from the MRTR path) renders it and returns an `ElicitResult`. The + * request is keyed by the current round so successive rounds are distinct. */ +function confirmInputRequests(round: number): Record<string, unknown> { return { - confirm: { + [`confirm_${round}`]: { method: "elicitation/create", params: { - message: "Approve this task before it continues?", + message: `Approve step ${round + 1} of this task before it continues?`, requestedSchema: { type: "object", properties: { @@ -127,6 +134,7 @@ export class ModernTaskRuntime { lastUpdatedAt: now, args, pollsRemaining: SIMPLE_WORKING_POLLS, + inputRound: 0, }; this.tasks.set(entry.taskId, entry); return { @@ -152,6 +160,10 @@ export class ModernTaskRuntime { const entry = this.requireTask(taskId); entry.inputSatisfied = true; if (inputResponses) entry.inputResponses = inputResponses; + // A loop never completes: advance the round so its next poll surfaces a + // fresh, distinct input request (a reused key would be deduped by the SDK + // and the loop would wedge at round 1). + if (entry.kind === "loop") entry.inputRound += 1; this.touch(entry); return { resultType: "complete" }; } @@ -196,8 +208,10 @@ export class ModernTaskRuntime { entry.status = "completed"; } } else if (entry.kind === "loop") { - // Never advances: stays input_required every poll (ignores tasks/update), - // so a client is re-prompted indefinitely — exercises its round cap. + // Never completes: stays input_required every poll. Each answered round + // bumps inputRound (in updateTask) so the next poll surfaces a distinct + // input request — the client is genuinely re-prompted until its own + // round cap trips, rather than wedging on a deduped repeat key. entry.status = "input_required"; } else { // input task: request input, then complete once the client has answered. @@ -222,7 +236,7 @@ export class ModernTaskRuntime { base.pollIntervalMs = DEFAULT_POLL_INTERVAL_MS; } if (entry.status === "input_required") { - base.inputRequests = confirmInputRequests(); + base.inputRequests = confirmInputRequests(entry.inputRound); } if (entry.status === "completed") { base.result = this.completedResult(entry); diff --git a/test-servers/src/preset-registry.ts b/test-servers/src/preset-registry.ts index fe0f92f6ec..115ce0a909 100644 --- a/test-servers/src/preset-registry.ts +++ b/test-servers/src/preset-registry.ts @@ -24,6 +24,7 @@ import { createCollectSampleTool, createListRootsTool, createCollectFormElicitationTool, + createSubmitTicketTool, createMrtrTool, createMrtrMultiRoundTool, createMrtrRootsTool, @@ -149,6 +150,8 @@ function resolveToolPreset( return createListRootsTool(); case "collect_elicitation": return createCollectFormElicitationTool(); + case "submit_ticket": + return createSubmitTicketTool(); case "mrtr_confirm": return createMrtrTool(); case "mrtr_two_step": diff --git a/test-servers/src/test-server-fixtures.ts b/test-servers/src/test-server-fixtures.ts index 9db734abbb..3fe73a35c2 100644 --- a/test-servers/src/test-server-fixtures.ts +++ b/test-servers/src/test-server-fixtures.ts @@ -494,6 +494,83 @@ export function createCollectFormElicitationTool(): ToolDefinition { }; } +/** + * A tool whose elicitation is INTRINSIC: the elicited fields are deliberately + * absent from the input schema, so a caller cannot pre-supply them as + * arguments — the only way to a ticket number is to answer the mid-call + * `elicitation/create`. That is what makes it a fixture for measuring + * elicitation *behavior* (does an agent recognize the request, map known + * facts onto the requested schema, and use the result?), where + * `collect_elicitation` — whose message/schema are caller-composed — can + * only exercise the wire. Legacy-only, like `collect_elicitation`: it calls + * `server.elicitInput`, which errors on the 2026-07-28 leg. + */ +export function createSubmitTicketTool(): ToolDefinition { + return { + name: "submit_ticket", + description: + "File a helpdesk ticket with the given summary and return the ticket number. Contact details are collected separately when the ticket is filed.", + inputSchema: { + summary: z.string().describe("One-line summary of the issue"), + }, + handler: async ( + params: Record<string, unknown>, + context?: TestServerContext, + ): Promise<CallToolResult> => { + if (!context) { + throw new Error("Server context not available"); + } + const summary = params.summary as string; + const result = await context.server.server.elicitInput({ + message: "Who should we contact about this ticket?", + requestedSchema: { + type: "object", + properties: { + contact_name: { + type: "string", + title: "Contact name", + description: "Full name of the person to contact", + }, + contact_email: { + type: "string", + title: "Contact email", + description: "Email address for updates on this ticket", + }, + }, + required: ["contact_name", "contact_email"], + }, + }); + if (result.action !== "accept") { + return toToolResult( + `Ticket not filed: contact details ${result.action === "decline" ? "declined" : "cancelled"}.`, + ); + } + const content = (result.content ?? {}) as Record<string, unknown>; + // Content-derived id (djb2 over the full submission), so the fixture + // holds no state and distinct submissions get distinct numbers except + // for genuine hash collisions in the 4-digit space. + const payload = JSON.stringify([ + summary, + content.contact_name, + content.contact_email, + ]); + let hash = 5381; + for (let i = 0; i < payload.length; i++) { + hash = ((hash * 33) ^ payload.charCodeAt(i)) >>> 0; + } + const ticket = `TCK-${(1000 + (hash % 9000)).toString()}`; + return toToolResult( + JSON.stringify({ + ticket, + summary, + contact_name: content.contact_name, + contact_email: content.contact_email, + }), + ); + }, + }; +} + /** Canonical URI for {@link createAppElicitationResource}, referenced by {@link createAppElicitationTool}'s `_meta.ui.resourceUri`. */ export const APP_ELICITATION_URI = "ui://demo/choose-option.html"; diff --git a/test-servers/src/test-server-stdio.ts b/test-servers/src/test-server-stdio.ts index 3d115f7946..63eaebc64b 100644 --- a/test-servers/src/test-server-stdio.ts +++ b/test-servers/src/test-server-stdio.ts @@ -3,20 +3,132 @@ /** * Test MCP server for stdio transport testing * Can be used programmatically or run as a standalone executable + * + * Run with {@link CRASHABLE_FLAG} it also serves {@link createCrashServerTool}, + * which kills this process at a point the test chooses — the fixture for the + * Inspector's mid-session crash reconciliation (#2437). That tool lives here, + * and only reaches a server through this file's standalone entry, because it + * calls `process.exit`: wired into the in-process HTTP server it would end the + * test runner rather than the server. */ +import * as z from "zod/v4"; import { McpServer } from "@modelcontextprotocol/server"; import { StdioServerTransport } from "@modelcontextprotocol/server/stdio"; import { fileURLToPath } from "url"; import type { ServerConfig, ResourceDefinition, + ToolDefinition, } from "./test-server-fixtures.js"; import { getDefaultServerConfig, createMcpServer, + createCollectFormElicitationTool, + createCollectSampleTool, } from "./test-server-fixtures.js"; +/** argv flag that makes the standalone server serve {@link createCrashServerTool}. */ +export const CRASHABLE_FLAG = "--crashable"; + +/** Name of the tool {@link createCrashServerTool} registers. */ +export const CRASH_SERVER_TOOL_NAME = "crash_server"; + +/** + * Create a `crash_server` tool that ends this server's process, so a test can + * crash a session at a point it picks rather than relying on whatever path + * happens to drop the connection. + * + * - `respond: false` (the default) exits without answering, so the + * `tools/call` that triggered it — and anything else in flight, such as a + * pending elicitation or sampling request — is still outstanding when the + * process dies. + * - `respond: true` answers first and exits `delayMs` later, so the crash + * lands on an idle session with nothing in flight. + * - `stderr` is written before exiting, standing in for a real server's dying + * words. + * + * The exit waits for stderr and stdout to flush first: pipe writes are + * asynchronous on macOS, so exiting straight after one can drop it — the + * dying words, or the `respond: true` answer, whose loss would turn an idle + * crash back into a crash with the call in flight. + * + * ⚠️ Only for a server running in a process of its own (the standalone entry + * below, behind {@link CRASHABLE_FLAG}) — see the module header. + */ +export function createCrashServerTool(): ToolDefinition { + return { + name: CRASH_SERVER_TOOL_NAME, + description: + "Exit this server process (test fixture for mid-session crash handling)", + inputSchema: { + respond: z + .boolean() + .optional() + .describe("Answer this call before exiting (default false)"), + delayMs: z + .number() + .int() + .min(0) + .optional() + .describe("Milliseconds to wait before exiting (default 0)"), + stderr: z + .string() + .optional() + .describe("Message to write to stderr just before exiting"), + }, + handler: async (params: Record<string, unknown>) => { + const respond = params.respond === true; + const delayMs = typeof params.delayMs === "number" ? params.delayMs : 0; + const stderr = + typeof params.stderr === "string" ? params.stderr : undefined; + // An empty write's callback runs once every write queued before it on + // that stream has flushed — so on stdout it is a barrier behind the + // `respond: true` answer, which the SDK has written by the time the + // timer below fires. + const flushStdoutThenExit = () => + process.stdout.write("", () => process.exit(1)); + const exit = () => { + if (stderr === undefined) return flushStdoutThenExit(); + process.stderr.write(`${stderr}\n`, flushStdoutThenExit); + }; + if (respond) { + setTimeout(exit, delayMs); + return { + content: [ + { + type: "text", + text: `Exiting in ${delayMs}ms`, + }, + ], + }; + } + // Never settles: the process is gone before an answer could be sent. + return new Promise(() => { + setTimeout(exit, delayMs); + }); + }, + }; +} + +/** + * The standalone server's config for {@link CRASHABLE_FLAG}: the default + * config plus the crash tool, and the two peer-request tools a test pairs with + * it to have a sampling or elicitation request pending when the process dies. + */ +export function getCrashableServerConfig(): ServerConfig { + const config = getDefaultServerConfig(); + return { + ...config, + tools: [ + ...(config.tools ?? []), + createCollectFormElicitationTool(), + createCollectSampleTool(), + createCrashServerTool(), + ], + }; +} + export class TestServerStdio { private mcpServer: McpServer; private transport?: StdioServerTransport; @@ -108,7 +220,22 @@ export function getTestMcpServerCommand(): { command: string; args: string[] } { }; } -// If run as a standalone script, start with default config +/** + * {@link getTestMcpServerCommand} for the crashable variant: the same server + * started with {@link CRASHABLE_FLAG}, serving {@link getCrashableServerConfig}. + */ +export function getCrashableTestMcpServerCommand(): { + command: string; + args: string[]; +} { + return { + command: "node", + args: [getTestMcpServerPath(), CRASHABLE_FLAG], + }; +} + +// If run as a standalone script, start with the default config — or, with +// CRASHABLE_FLAG on the command line, the crashable one // Check if this file is being executed directly (not imported) const isMainModule = import.meta.url.endsWith(process.argv[1] || "") || @@ -116,7 +243,11 @@ const isMainModule = (process.argv[1]?.endsWith("test-server-stdio.js") ?? false); if (isMainModule) { - const server = new TestServerStdio(getDefaultServerConfig()); + const server = new TestServerStdio( + process.argv.includes(CRASHABLE_FLAG) + ? getCrashableServerConfig() + : getDefaultServerConfig(), + ); server .start() .then(() => {