diff --git a/README.md b/README.md index 545b905..9ae27e2 100644 --- a/README.md +++ b/README.md @@ -154,12 +154,31 @@ authenticates every request independently of anything the MCP client does: and a truncated hash, not the content. Set `NANOKVM_MCP_AUDIT_FULL=true` to log full text instead (useful for debugging, but keep the log file access-restricted if you do). -**Do not expose `NANOKVM_MCP_BIND` directly to the LAN or the internet.** If you need to -reach the sidecar from another machine, put it on a [Tailscale](https://tailscale.com) -tailnet instead — the stock NanoKVM firmware already supports installing Tailscale, so -this costs nothing and is the recommended exposure path. Bind the sidecar to the -tailscale interface address (or leave it on loopback and use `tailscale serve`/a local -SSH tunnel) rather than to `0.0.0.0`. +**Do not expose `NANOKVM_MCP_BIND` directly to the LAN or the internet.** To reach the +sidecar from another machine, leave the bind on loopback and forward a local port over +SSH — see [Durable access](#durable-access-a-launchd-supervised-tunnel). That keeps two +independent layers in front of a keystroke injector: an SSH credential to reach the port +at all, and the bearer token to use it. + +**A note on Tailscale, which this README previously recommended.** Putting the device on a +tailnet is sound reasoning about *exposure* — the LAN is the wrong place for this listener +— but running `tailscaled` **on the NanoKVM does not fit its memory budget**, and the +earlier claim here that it "costs nothing" was wrong. On this hardware `tailscaled` settles +around **40 MB resident** ([sipeed/NanoKVM#366](https://github.com/sipeed/NanoKVM/issues/366)), +against the **43 MB available** this project measures on-device. Upstream reports it +needing `GOMEMLIMIT=100`, `GOGC=50`, and a cron reboot every two days to stay up — and +still being OOM-killed at 78 MB — on a board with *more* free RAM than ours +([sipeed/NanoKVM#660](https://github.com/sipeed/NanoKVM/issues/660)). Tuning does not +recover it: roughly 86% of that heap is wireguard-go per-interface buffer pools, an +allocation floor rather than collectable garbage +([tailscale/tailscale#16258](https://github.com/tailscale/tailscale/issues/16258)). An OOM +here is not graceful — the kernel takes the largest resident process, which may be the +firmware's own video pipeline, so the KVM function dies along with your way back in. + +If you later need to reach the device from outside the LAN, run the Tailscale **subnet +router on another always-on host** on that LAN rather than on the NanoKVM. The device +spends no RAM, the tunnel below rides over the tailnet unchanged, and the MCP listener +still never leaves loopback. ## Connecting an MCP client @@ -170,7 +189,7 @@ Point your MCP client at the daemon's HTTP endpoint with the bearer token in the { "mcpServers": { "nanokvm": { - "url": "http://:8080/", + "url": "http://127.0.0.1:8080/", "headers": { "Authorization": "Bearer ${NANOKVM_MCP_TOKEN}" } @@ -179,18 +198,20 @@ Point your MCP client at the daemon's HTTP endpoint with the bearer token in the } ``` -Replace `` with the NanoKVM's Tailscale address (or `127.0.0.1` plus an SSH -tunnel/port-forward — see [Security](#security-model) for why the LAN address is -discouraged) and `` with the value of `NANOKVM_MCP_TOKEN`, or the token printed to -`/data/nanokvm-mcp/daemon.log` if you didn't set one. +The URL is `127.0.0.1` because the sidecar stays bound to loopback and you reach it +through an SSH port-forward — see [Security](#security-model) for why the LAN address is +discouraged, and [Durable access](#durable-access-a-launchd-supervised-tunnel) for a +forward that survives reboots. `NANOKVM_MCP_TOKEN` holds the device's token, or the one +printed to `/data/nanokvm-mcp/daemon.log` if you didn't set an explicit value. ### Claude Code over an SSH tunnel (tested end-to-end) This is the setup validated against real hardware. It uses an SSH master connection that carries the port-forward and authenticates once; every later `ssh`/`scp` rides the same socket without re-prompting. (The master exits after -an hour *idle*, so this is a per-session flow, not a permanent one — a durable -path is tracked in [#16](https://github.com/tylervick/nanokvm-mcp/issues/16).) +an hour *idle*, so this is a per-session flow, not a permanent one — for +day-to-day use, set up the supervised tunnel in +[Durable access](#durable-access-a-launchd-supervised-tunnel) instead.) **1. Open the master connection + tunnel** (one password prompt): @@ -257,6 +278,171 @@ has no file-based header option), so it's briefly visible to `ps` — fine on a single-user machine, but on a shared host prefer checking with curl's `@file` header form instead. +### Durable access: a launchd-supervised tunnel + +The master-socket flow above authenticates once and lapses after an hour idle — +right for a working session, wrong for day-to-day. For a forward that comes back +on its own after a network drop, a device reboot, or a laptop wake, hand a plain +tunnel to launchd. + +This changes nothing in the [security model](#security-model): the sidecar stays +bound to `127.0.0.1`, and reaching it still costs an SSH credential *plus* the +bearer token. What changes is only who restarts the tunnel. + +This is a LAN-scoped path — it needs the client to reach the device's SSH port. +See [Security](#security-model) for the off-LAN extension. + +**1. Install a key on the device.** An unattended reconnect can't answer a +password prompt, and the stock firmware ships no authorized keys. Check what the +device's dropbear supports first — ed25519 needs dropbear 2020.79 or newer: + +```sh +ssh -v root@ exit 2>&1 | grep 'remote software version' +``` + +Generate a dedicated key (use `-t rsa -b 3072` instead if that banner is older) +and append it, paying the password prompt one last time. The `restrict` prefix +confines the key to what the tunnel actually needs: + +```sh +ssh-keygen -t ed25519 -f ~/.ssh/nanokvm -C nanokvm-mcp -N '' +{ printf 'restrict,no-pty,command="/bin/false" '; cat ~/.ssh/nanokvm.pub; } | \ +ssh -o PubkeyAuthentication=no -o PreferredAuthentications=password root@ \ + 'mkdir -p /root/.ssh && chmod 700 /root/.ssh && cat >> /root/.ssh/authorized_keys \ + && chmod 600 /root/.ssh/authorized_keys' +``` + +**Understand what this key is before you install it.** `-N ''` gives it no +passphrase, because launchd has no one to ask for one — so `~/.ssh/nanokvm` is a +plaintext credential at rest, and whoever reads that file gets whatever the key +grants. Left unrestricted, that is a root shell on the device, not merely the +port-forward. Three things reduce the blast radius, in descending order of how +much they buy you: + +- **The `restrict,no-pty,command="/bin/false"` prefix above.** `restrict` denies + everything and re-enables nothing; `command="/bin/false"` forces a dead command + for any shell or exec request. `ssh -N` asks for neither, so the forward still + works while interactive use of the key does not. Dropbear parses `restrict` only + on newer builds — if yours ignores it the key still works, just unrestricted, and + if it rejects the line outright password auth still gets you in. Neither case + locks you out, but check which you got (below) rather than assuming. +- **Keep it out of your agent.** Don't `ssh-add` this key. The plist references it + by path, so the agent never needs it, and staying out of the agent is also what + keeps the dropbear auth-attempt flood from coming back. +- **A passphrase plus a keychain-backed agent** (`ssh-add --apple-use-keychain`) is + strictly safer at rest, but it defeats the point here: the whole reason for this + section is a tunnel that reconnects with nobody logged in. If your threat model + wants the passphrase, keep the per-session master-socket flow above instead and + accept the hourly re-auth. + +Some dropbear builds read `/etc/dropbear/authorized_keys` instead. Verify before +going further — a launchd job that can't authenticate just respawns forever. Test +the forward rather than a shell, since the forced command denies the latter by +design: + +```sh +ssh -i ~/.ssh/nanokvm -o IdentitiesOnly=yes -f -N -L 18080:127.0.0.1:8080 root@ \ + && nc -z 127.0.0.1 18080 && echo "key auth + forwarding OK" +``` + +If that prints nothing, retry with `-v` and look for `Authentication succeeded`. +Should a shell come back instead of the forced command failing, the restriction +options were ignored — the key works but is unrestricted, so treat `~/.ssh/nanokvm` +accordingly. Clean up the test forward when done (`pkill -f 18080:127.0.0.1:8080`). + +With a key in place, `-i ~/.ssh/nanokvm -o IdentitiesOnly=yes` *replaces* the +`PubkeyAuthentication=no` trio above rather than adding to it: dropbear is offered +exactly one key, so there is no agent flood to work around. + +The key survives firmware *app* updates (they touch only `/kvmapp`, `/root/old`, +and `/root/.kvmcache`) but not a rootfs reflash — redo this step after one. + +**2. Connect once interactively, then clear any ad-hoc master.** The first +connection pins the device's host key in `known_hosts`; the launchd job runs with +`StrictHostKeyChecking` at its default and will refuse an unknown host rather than +trust it blindly. Then release port 8080, or the supervised tunnel will fail its +forward and respawn in a loop: + +```sh +ssh -S /tmp/nkvm.sock -O exit root@ 2>/dev/null +``` + +**3. Write the agent** to `~/Library/LaunchAgents/com.nanokvm.tunnel.plist`. +`ProgramArguments` gets no shell expansion, so spell out your home directory and +the device address in full — `~` and `$HOME` will not work. Replace +`YOUR-USER` and `DEVICE-ADDRESS` before loading it; unlike the shell snippets +elsewhere in this README, the placeholders here can't use angle brackets, which +XML would read as tags: + +```xml + + + + + Label + com.nanokvm.tunnel + ProgramArguments + + /usr/bin/ssh + -N + -i /Users/YOUR-USER/.ssh/nanokvm + -o IdentitiesOnly=yes + -o ControlPath=none + -o ControlMaster=no + -o ExitOnForwardFailure=yes + -o ConnectTimeout=10 + -o ServerAliveInterval=30 + -o ServerAliveCountMax=3 + -L 8080:127.0.0.1:8080 + root@DEVICE-ADDRESS + + RunAtLoad + KeepAlive + ThrottleInterval 10 + StandardErrorPath /tmp/nanokvm-tunnel.log + + +``` + +Four details there are load-bearing, and each maps to a way this fails silently: + +- **No `-f`.** launchd supervises a *foreground* process. A self-backgrounding + `ssh` looks like an instant exit and gets respawned forever. +- **`ControlPath=none`, not just `ControlMaster=no`.** These are not the same + switch, and getting it wrong is the subtle one. Per `ssh_config(5)`, + `ControlMaster=no` is the *default* and is precisely the value that lets a + session **join** an existing master — it governs whether this `ssh` becomes a + master, not whether it reuses one. Only `ControlPath=none` disables sharing. If + your `~/.ssh/config` sets a `ControlPath`, the job would otherwise hand its + forward to that master and exit immediately, giving you a respawn loop and a + tunnel that dies whenever the master does. +- **`ExitOnForwardFailure=yes` with `ServerAliveInterval`/`ServerAliveCountMax`.** + Together these make `ssh` exit within ~90 s of the link going away instead of + holding a dead forward open. launchd can only restart something that exits. +- **`ConnectTimeout=10`.** Without it a connect to a sleeping or absent device sits + in the system TCP timeout for over a minute before launchd gets its exit code, + which stretches every reconnect attempt for no benefit on a LAN. + +**4. Load it and confirm:** + +```sh +launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.nanokvm.tunnel.plist +launchctl print gui/$(id -u)/com.nanokvm.tunnel | grep -E 'state|pid' +nc -z 127.0.0.1 8080 && echo "tunnel up" +``` + +`nc -z` confirms the forward; to check the MCP endpoint end-to-end, use the MCP +Inspector command above. To stop it, or to reload after editing the plist: + +```sh +launchctl bootout gui/$(id -u)/com.nanokvm.tunnel +``` + +With the tunnel supervised, the `claude mcp add` registration from step 3 above +needs no changes — it already points at `127.0.0.1:8080`, and now that address +answers without a manual re-auth first. + ## Tools 14 tools are registered by default (7 in read-only mode): diff --git a/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md b/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md index 1290100..c871114 100644 --- a/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md +++ b/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md @@ -130,7 +130,7 @@ upstream code carry a header noting origin. Laptop — Claude Code / Claude Desktop │ │ MCP streamable HTTP + bearer token - │ (over Tailscale tailnet, recommended) + │ (over an SSH port-forward to loopback) ▼ NanoKVM device (riscv64 / SG2002) ┌─────────────────────────────────────────────┐ @@ -246,13 +246,28 @@ On-device means no stdio, so the endpoint is network-reachable and **must** auth independently of tool annotations. This is transport security, not a guardrail. - Streamable HTTP via the official `modelcontextprotocol/go-sdk` (v1.6.1) -- Bearer token from `/etc/kvm/.nanokvm_mcp_token`, mode 0600, generated on first run +- ~~Bearer token from `/etc/kvm/.nanokvm_mcp_token`, mode 0600, generated on first run~~ + **Corrected 2026-07-30 (#16):** as built, the token comes from the `NANOKVM_MCP_TOKEN` + environment variable (`internal/config/config.go`), which the init script sources from + `/root/nanokvm-mcp/nanokvm-mcp.env`; if unset, one is generated per start and logged to + `/data/nanokvm-mcp/daemon.log`. No `/etc/kvm/.nanokvm_mcp_token` is ever read — the only + `/etc/kvm` token this project touches is the firmware's `.picoclaw_internal_token`. - Constant-time comparison; non-bearer requests get 401 - **Default bind `127.0.0.1:8080`.** Exposing it requires explicit configuration. A keystroke injector should not become LAN-reachable by default. -- The README documents Tailscale as the recommended exposure path. The stock firmware +- ~~The README documents Tailscale as the recommended exposure path. The stock firmware already supports installing it, so putting this on a tailnet instead of the LAN is the - largest available security win and costs nothing. + largest available security win and costs nothing.~~ **Superseded — see below.** + + **Corrected 2026-07-30 (#16).** The security reasoning holds — the LAN is the wrong + place for this listener — but "costs nothing" was wrong, and the recommendation has + changed to an SSH port-forward to the loopback bind. `tailscaled` on this hardware + settles around 40 MB resident ([sipeed/NanoKVM#366](https://github.com/sipeed/NanoKVM/issues/366)) + against the 43 MB available measured in Device recon above, and roughly 86% of that + heap is wireguard-go per-interface buffer pools rather than collectable garbage + ([tailscale/tailscale#16258](https://github.com/tailscale/tailscale/issues/16258)), so + `GOMEMLIMIT` does not recover it. Off-LAN reach, when it is needed, belongs on a + subnet router on another host — not on the device. ## Guardrails @@ -415,4 +430,4 @@ that case is documented, not solved. | Memory pressure on a thin ~18 MB headroom (43 MB available) | Measured 8.1 MB idle, 7.1 MB binary. `GOMEMLIMIT` in the init script; picoclaw path never decodes JPEG; `publicBackend` hard-caps resolution to keep decode under budget; enforced in CI and the device smoke test | | `publicBackend` screenshot spikes memory (no scaled decode in Go's `image/jpeg`) | Resolution cap plus `debug.FreeOSMemory()`; documented as a degraded path, and the reason picoclaw is preferred | | Session-lock contention with PicoClaw | Detected and surfaced as a clear tool error rather than a hang | -| Network-exposed keystroke injection | Loopback default bind, bearer auth, read-only mode, audit log, Tailscale guidance | +| Network-exposed keystroke injection | Loopback default bind, bearer auth, read-only mode, audit log, SSH port-forward guidance |