From b9d16e56f5e2e2932890330627efda57bd9b999f Mon Sep 17 00:00:00 2001 From: Tyler Vick Date: Thu, 30 Jul 2026 10:52:24 -0700 Subject: [PATCH 1/2] docs: durable SSH tunnel for day-to-day access (closes #16) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The ad-hoc master socket lapses after an hour idle. Document a launchd-supervised port-forward that reconnects on its own, and correct the Tailscale recommendation it replaces. Tailscale on the device was evaluated and rejected on memory. tailscaled settles around 40 MB resident on this hardware (sipeed/NanoKVM#366) against the 43 MB available measured in device recon; upstream needs GOMEMLIMIT=100, GOGC=50 and a two-day cron reboot to keep it up on a board with more free RAM than ours, and still saw an OOM kill at 78 MB (sipeed/NanoKVM#660). GOMEMLIMIT does not recover it — ~86% of that heap is wireguard-go per-interface buffer pools, an allocation floor rather than collectable garbage (tailscale/tailscale#16258). The README's "costs nothing" claim was wrong and is corrected here and in the design spec; the exposure reasoning behind it still holds. The tunnel leaves the security model unchanged: the bind stays on loopback, and reaching it costs an SSH credential plus the bearer token. Off-LAN reach, when it is wanted, belongs on a subnet router on another host rather than on the device. Documents the SSH key prerequisite unattended reconnect requires, which replaces PR #18's dropbear agent-flood workaround rather than adding to it, and the three plist details that otherwise fail silently: no -f, ControlMaster=no, and ExitOnForwardFailure with server-alive limits. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 166 ++++++++++++++++-- .../2026-07-22-nanokvm-mcp-sidecar-design.md | 14 +- 2 files changed, 165 insertions(+), 15 deletions(-) diff --git a/README.md b/README.md index 545b905..9cc83d1 100644 --- a/README.md +++ b/README.md @@ -154,12 +154,31 @@ authenticates every request independently of anything the MCP client does: and a truncated hash, not the content. Set `NANOKVM_MCP_AUDIT_FULL=true` to log full text instead (useful for debugging, but keep the log file access-restricted if you do). -**Do not expose `NANOKVM_MCP_BIND` directly to the LAN or the internet.** If you need to -reach the sidecar from another machine, put it on a [Tailscale](https://tailscale.com) -tailnet instead — the stock NanoKVM firmware already supports installing Tailscale, so -this costs nothing and is the recommended exposure path. Bind the sidecar to the -tailscale interface address (or leave it on loopback and use `tailscale serve`/a local -SSH tunnel) rather than to `0.0.0.0`. +**Do not expose `NANOKVM_MCP_BIND` directly to the LAN or the internet.** To reach the +sidecar from another machine, leave the bind on loopback and forward a local port over +SSH — see [Durable access](#durable-access-a-launchd-supervised-tunnel). That keeps two +independent layers in front of a keystroke injector: an SSH credential to reach the port +at all, and the bearer token to use it. + +**A note on Tailscale, which this README previously recommended.** Putting the device on a +tailnet is sound reasoning about *exposure* — the LAN is the wrong place for this listener +— but running `tailscaled` **on the NanoKVM does not fit its memory budget**, and the +earlier claim here that it "costs nothing" was wrong. On this hardware `tailscaled` settles +around **40 MB resident** ([sipeed/NanoKVM#366](https://github.com/sipeed/NanoKVM/issues/366)), +against the **43 MB available** this project measures on-device. Upstream reports it +needing `GOMEMLIMIT=100`, `GOGC=50`, and a cron reboot every two days to stay up — and +still being OOM-killed at 78 MB — on a board with *more* free RAM than ours +([sipeed/NanoKVM#660](https://github.com/sipeed/NanoKVM/issues/660)). Tuning does not +recover it: roughly 86% of that heap is wireguard-go per-interface buffer pools, an +allocation floor rather than collectable garbage +([tailscale/tailscale#16258](https://github.com/tailscale/tailscale/issues/16258)). An OOM +here is not graceful — the kernel takes the largest resident process, which may be the +firmware's own video pipeline, so the KVM function dies along with your way back in. + +If you later need to reach the device from outside the LAN, run the Tailscale **subnet +router on another always-on host** on that LAN rather than on the NanoKVM. The device +spends no RAM, the tunnel below rides over the tailnet unchanged, and the MCP listener +still never leaves loopback. ## Connecting an MCP client @@ -170,7 +189,7 @@ Point your MCP client at the daemon's HTTP endpoint with the bearer token in the { "mcpServers": { "nanokvm": { - "url": "http://:8080/", + "url": "http://127.0.0.1:8080/", "headers": { "Authorization": "Bearer ${NANOKVM_MCP_TOKEN}" } @@ -179,18 +198,20 @@ Point your MCP client at the daemon's HTTP endpoint with the bearer token in the } ``` -Replace `` with the NanoKVM's Tailscale address (or `127.0.0.1` plus an SSH -tunnel/port-forward — see [Security](#security-model) for why the LAN address is -discouraged) and `` with the value of `NANOKVM_MCP_TOKEN`, or the token printed to -`/data/nanokvm-mcp/daemon.log` if you didn't set one. +The URL is `127.0.0.1` because the sidecar stays bound to loopback and you reach it +through an SSH port-forward — see [Security](#security-model) for why the LAN address is +discouraged, and [Durable access](#durable-access-a-launchd-supervised-tunnel) for a +forward that survives reboots. `NANOKVM_MCP_TOKEN` holds the device's token, or the one +printed to `/data/nanokvm-mcp/daemon.log` if you didn't set an explicit value. ### Claude Code over an SSH tunnel (tested end-to-end) This is the setup validated against real hardware. It uses an SSH master connection that carries the port-forward and authenticates once; every later `ssh`/`scp` rides the same socket without re-prompting. (The master exits after -an hour *idle*, so this is a per-session flow, not a permanent one — a durable -path is tracked in [#16](https://github.com/tylervick/nanokvm-mcp/issues/16).) +an hour *idle*, so this is a per-session flow, not a permanent one — for +day-to-day use, set up the supervised tunnel in +[Durable access](#durable-access-a-launchd-supervised-tunnel) instead.) **1. Open the master connection + tunnel** (one password prompt): @@ -257,6 +278,125 @@ has no file-based header option), so it's briefly visible to `ps` — fine on a single-user machine, but on a shared host prefer checking with curl's `@file` header form instead. +### Durable access: a launchd-supervised tunnel + +The master-socket flow above authenticates once and lapses after an hour idle — +right for a working session, wrong for day-to-day. For a forward that comes back +on its own after a network drop, a device reboot, or a laptop wake, hand a plain +tunnel to launchd. + +This changes nothing in the [security model](#security-model): the sidecar stays +bound to `127.0.0.1`, and reaching it still costs an SSH credential *plus* the +bearer token. What changes is only who restarts the tunnel. + +This is a LAN-scoped path — it needs the client to reach the device's SSH port. +See [Security](#security-model) for the off-LAN extension. + +**1. Install a key on the device.** An unattended reconnect can't answer a +password prompt, and the stock firmware ships no authorized keys. Check what the +device's dropbear supports first — ed25519 needs dropbear 2020.79 or newer: + +```sh +ssh -v root@ exit 2>&1 | grep 'remote software version' +``` + +Generate a dedicated key (use `-t rsa -b 3072` instead if that banner is older) +and append it, paying the password prompt one last time: + +```sh +ssh-keygen -t ed25519 -f ~/.ssh/nanokvm -C nanokvm-mcp -N '' +ssh -o PubkeyAuthentication=no -o PreferredAuthentications=password root@ \ + 'mkdir -p /root/.ssh && chmod 700 /root/.ssh && cat >> /root/.ssh/authorized_keys \ + && chmod 600 /root/.ssh/authorized_keys' < ~/.ssh/nanokvm.pub +``` + +Some dropbear builds read `/etc/dropbear/authorized_keys` instead. Verify before +going further — a launchd job that can't authenticate just respawns forever: + +```sh +ssh -i ~/.ssh/nanokvm -o IdentitiesOnly=yes root@ true && echo "key auth OK" +``` + +With a key in place, `-i ~/.ssh/nanokvm -o IdentitiesOnly=yes` *replaces* the +`PubkeyAuthentication=no` trio above rather than adding to it: dropbear is offered +exactly one key, so there is no agent flood to work around. + +The key survives firmware *app* updates (they touch only `/kvmapp`, `/root/old`, +and `/root/.kvmcache`) but not a rootfs reflash — redo this step after one. + +**2. Connect once interactively, then clear any ad-hoc master.** The first +connection pins the device's host key in `known_hosts`; the launchd job runs with +`StrictHostKeyChecking` at its default and will refuse an unknown host rather than +trust it blindly. Then release port 8080, or the supervised tunnel will fail its +forward and respawn in a loop: + +```sh +ssh -S /tmp/nkvm.sock -O exit root@ 2>/dev/null +``` + +**3. Write the agent** to `~/Library/LaunchAgents/com.nanokvm.tunnel.plist`. +`ProgramArguments` gets no shell expansion, so spell out your home directory and +the device address in full — `~` and `$HOME` will not work: + +```xml + + + + + Label + com.nanokvm.tunnel + ProgramArguments + + /usr/bin/ssh + -N + -i /Users//.ssh/nanokvm + -o IdentitiesOnly=yes + -o ControlMaster=no + -o ExitOnForwardFailure=yes + -o ServerAliveInterval=30 + -o ServerAliveCountMax=3 + -L 8080:127.0.0.1:8080 + root@ + + RunAtLoad + KeepAlive + ThrottleInterval 10 + StandardErrorPath /tmp/nanokvm-tunnel.log + + +``` + +Three details there are load-bearing, and each maps to a way this fails silently: + +- **No `-f`.** launchd supervises a *foreground* process. A self-backgrounding + `ssh` looks like an instant exit and gets respawned forever. +- **`ControlMaster=no`.** If your `~/.ssh/config` sets `ControlMaster auto`, the + job would hand its forward to an existing master and exit immediately — the same + respawn loop, and a tunnel that dies whenever that master does. +- **`ExitOnForwardFailure=yes` with `ServerAliveInterval`/`ServerAliveCountMax`.** + Together these make `ssh` exit within ~90 s of the link going away instead of + holding a dead forward open. launchd can only restart something that exits. + +**4. Load it and confirm:** + +```sh +launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.nanokvm.tunnel.plist +launchctl print gui/$(id -u)/com.nanokvm.tunnel | grep -E 'state|pid' +nc -z 127.0.0.1 8080 && echo "tunnel up" +``` + +`nc -z` confirms the forward; to check the MCP endpoint end-to-end, use the MCP +Inspector command above. To stop it, or to reload after editing the plist: + +```sh +launchctl bootout gui/$(id -u)/com.nanokvm.tunnel +``` + +With the tunnel supervised, the `claude mcp add` registration from step 3 above +needs no changes — it already points at `127.0.0.1:8080`, and now that address +answers without a manual re-auth first. + ## Tools 14 tools are registered by default (7 in read-only mode): diff --git a/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md b/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md index 1290100..9f5ac0d 100644 --- a/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md +++ b/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md @@ -130,7 +130,7 @@ upstream code carry a header noting origin. Laptop — Claude Code / Claude Desktop │ │ MCP streamable HTTP + bearer token - │ (over Tailscale tailnet, recommended) + │ (over an SSH port-forward to loopback) ▼ NanoKVM device (riscv64 / SG2002) ┌─────────────────────────────────────────────┐ @@ -254,6 +254,16 @@ independently of tool annotations. This is transport security, not a guardrail. already supports installing it, so putting this on a tailnet instead of the LAN is the largest available security win and costs nothing. + **Corrected 2026-07-30 (#16).** The security reasoning holds — the LAN is the wrong + place for this listener — but "costs nothing" was wrong, and the recommendation has + changed to an SSH port-forward to the loopback bind. `tailscaled` on this hardware + settles around 40 MB resident ([sipeed/NanoKVM#366](https://github.com/sipeed/NanoKVM/issues/366)) + against the 43 MB available measured in Device recon above, and roughly 86% of that + heap is wireguard-go per-interface buffer pools rather than collectable garbage + ([tailscale/tailscale#16258](https://github.com/tailscale/tailscale/issues/16258)), so + `GOMEMLIMIT` does not recover it. Off-LAN reach, when it is needed, belongs on a + subnet router on another host — not on the device. + ## Guardrails - **Annotations** on all 14 tools, as above. @@ -415,4 +425,4 @@ that case is documented, not solved. | Memory pressure on a thin ~18 MB headroom (43 MB available) | Measured 8.1 MB idle, 7.1 MB binary. `GOMEMLIMIT` in the init script; picoclaw path never decodes JPEG; `publicBackend` hard-caps resolution to keep decode under budget; enforced in CI and the device smoke test | | `publicBackend` screenshot spikes memory (no scaled decode in Go's `image/jpeg`) | Resolution cap plus `debug.FreeOSMemory()`; documented as a degraded path, and the reason picoclaw is preferred | | Session-lock contention with PicoClaw | Detected and surfaced as a clear tool error rather than a hang | -| Network-exposed keystroke injection | Loopback default bind, bearer auth, read-only mode, audit log, Tailscale guidance | +| Network-exposed keystroke injection | Loopback default bind, bearer auth, read-only mode, audit log, SSH port-forward guidance | From 7e9c42ba81f93be922f70a5cbb4b6cf6d9a08339 Mon Sep 17 00:00:00 2001 From: Tyler Vick Date: Thu, 30 Jul 2026 11:26:32 -0700 Subject: [PATCH 2/2] =?UTF-8?q?docs:=20address=20review=20=E2=80=94=20Cont?= =?UTF-8?q?rolPath,=20key=20restrictions,=20plist=20validity?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ControlMaster=no does not prevent multiplexing. Per ssh_config(5) it is the default and is exactly the value that lets a session join an existing master; it governs whether this ssh becomes a master, not whether it reuses one. Only ControlPath=none disables sharing. The plist and the explanation for it were both wrong — a reader with ControlPath in their ssh config would have hit the respawn loop that bullet claimed to prevent. Add ConnectTimeout=10 so a connect to a sleeping device fails in 10s rather than sitting in the system TCP timeout before launchd sees an exit. The plist used / placeholders, which XML parses as elements, so the block as printed was not valid and was not what plutil validated earlier. Switched to YOUR-USER/DEVICE-ADDRESS and noted why the angle-bracket convention used elsewhere cannot apply inside XML. The block now lints verbatim. Document what the SSH key actually grants: -N '' leaves it unencrypted at rest, and an unrestricted key in root's authorized_keys is a root shell, not just a forward. Install it behind restrict,no-pty,command="/bin/false" so -N forwarding still works while interactive use does not, with the dropbear-version caveat and a verification that tests the forward rather than a shell. Note that neither failure mode locks you out, since password auth remains. Spec: strike through the superseded Tailscale sentence rather than only appending a correction under it, and correct the bearer-token source — it is the NANOKVM_MCP_TOKEN environment variable per internal/config/config.go, not /etc/kvm/.nanokvm_mcp_token, which nothing reads. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 68 ++++++++++++++++--- .../2026-07-22-nanokvm-mcp-sidecar-design.md | 11 ++- 2 files changed, 65 insertions(+), 14 deletions(-) diff --git a/README.md b/README.md index 9cc83d1..9ae27e2 100644 --- a/README.md +++ b/README.md @@ -301,22 +301,55 @@ ssh -v root@ exit 2>&1 | grep 'remote software version' ``` Generate a dedicated key (use `-t rsa -b 3072` instead if that banner is older) -and append it, paying the password prompt one last time: +and append it, paying the password prompt one last time. The `restrict` prefix +confines the key to what the tunnel actually needs: ```sh ssh-keygen -t ed25519 -f ~/.ssh/nanokvm -C nanokvm-mcp -N '' +{ printf 'restrict,no-pty,command="/bin/false" '; cat ~/.ssh/nanokvm.pub; } | \ ssh -o PubkeyAuthentication=no -o PreferredAuthentications=password root@ \ 'mkdir -p /root/.ssh && chmod 700 /root/.ssh && cat >> /root/.ssh/authorized_keys \ - && chmod 600 /root/.ssh/authorized_keys' < ~/.ssh/nanokvm.pub + && chmod 600 /root/.ssh/authorized_keys' ``` +**Understand what this key is before you install it.** `-N ''` gives it no +passphrase, because launchd has no one to ask for one — so `~/.ssh/nanokvm` is a +plaintext credential at rest, and whoever reads that file gets whatever the key +grants. Left unrestricted, that is a root shell on the device, not merely the +port-forward. Three things reduce the blast radius, in descending order of how +much they buy you: + +- **The `restrict,no-pty,command="/bin/false"` prefix above.** `restrict` denies + everything and re-enables nothing; `command="/bin/false"` forces a dead command + for any shell or exec request. `ssh -N` asks for neither, so the forward still + works while interactive use of the key does not. Dropbear parses `restrict` only + on newer builds — if yours ignores it the key still works, just unrestricted, and + if it rejects the line outright password auth still gets you in. Neither case + locks you out, but check which you got (below) rather than assuming. +- **Keep it out of your agent.** Don't `ssh-add` this key. The plist references it + by path, so the agent never needs it, and staying out of the agent is also what + keeps the dropbear auth-attempt flood from coming back. +- **A passphrase plus a keychain-backed agent** (`ssh-add --apple-use-keychain`) is + strictly safer at rest, but it defeats the point here: the whole reason for this + section is a tunnel that reconnects with nobody logged in. If your threat model + wants the passphrase, keep the per-session master-socket flow above instead and + accept the hourly re-auth. + Some dropbear builds read `/etc/dropbear/authorized_keys` instead. Verify before -going further — a launchd job that can't authenticate just respawns forever: +going further — a launchd job that can't authenticate just respawns forever. Test +the forward rather than a shell, since the forced command denies the latter by +design: ```sh -ssh -i ~/.ssh/nanokvm -o IdentitiesOnly=yes root@ true && echo "key auth OK" +ssh -i ~/.ssh/nanokvm -o IdentitiesOnly=yes -f -N -L 18080:127.0.0.1:8080 root@ \ + && nc -z 127.0.0.1 18080 && echo "key auth + forwarding OK" ``` +If that prints nothing, retry with `-v` and look for `Authentication succeeded`. +Should a shell come back instead of the forced command failing, the restriction +options were ignored — the key works but is unrestricted, so treat `~/.ssh/nanokvm` +accordingly. Clean up the test forward when done (`pkill -f 18080:127.0.0.1:8080`). + With a key in place, `-i ~/.ssh/nanokvm -o IdentitiesOnly=yes` *replaces* the `PubkeyAuthentication=no` trio above rather than adding to it: dropbear is offered exactly one key, so there is no agent flood to work around. @@ -336,7 +369,10 @@ ssh -S /tmp/nkvm.sock -O exit root@ 2>/dev/null **3. Write the agent** to `~/Library/LaunchAgents/com.nanokvm.tunnel.plist`. `ProgramArguments` gets no shell expansion, so spell out your home directory and -the device address in full — `~` and `$HOME` will not work: +the device address in full — `~` and `$HOME` will not work. Replace +`YOUR-USER` and `DEVICE-ADDRESS` before loading it; unlike the shell snippets +elsewhere in this README, the placeholders here can't use angle brackets, which +XML would read as tags: ```xml @@ -350,14 +386,16 @@ the device address in full — `~` and `$HOME` will not work: /usr/bin/ssh -N - -i /Users//.ssh/nanokvm + -i /Users/YOUR-USER/.ssh/nanokvm -o IdentitiesOnly=yes + -o ControlPath=none -o ControlMaster=no -o ExitOnForwardFailure=yes + -o ConnectTimeout=10 -o ServerAliveInterval=30 -o ServerAliveCountMax=3 -L 8080:127.0.0.1:8080 - root@ + root@DEVICE-ADDRESS RunAtLoad KeepAlive @@ -367,16 +405,24 @@ the device address in full — `~` and `$HOME` will not work: ``` -Three details there are load-bearing, and each maps to a way this fails silently: +Four details there are load-bearing, and each maps to a way this fails silently: - **No `-f`.** launchd supervises a *foreground* process. A self-backgrounding `ssh` looks like an instant exit and gets respawned forever. -- **`ControlMaster=no`.** If your `~/.ssh/config` sets `ControlMaster auto`, the - job would hand its forward to an existing master and exit immediately — the same - respawn loop, and a tunnel that dies whenever that master does. +- **`ControlPath=none`, not just `ControlMaster=no`.** These are not the same + switch, and getting it wrong is the subtle one. Per `ssh_config(5)`, + `ControlMaster=no` is the *default* and is precisely the value that lets a + session **join** an existing master — it governs whether this `ssh` becomes a + master, not whether it reuses one. Only `ControlPath=none` disables sharing. If + your `~/.ssh/config` sets a `ControlPath`, the job would otherwise hand its + forward to that master and exit immediately, giving you a respawn loop and a + tunnel that dies whenever the master does. - **`ExitOnForwardFailure=yes` with `ServerAliveInterval`/`ServerAliveCountMax`.** Together these make `ssh` exit within ~90 s of the link going away instead of holding a dead forward open. launchd can only restart something that exits. +- **`ConnectTimeout=10`.** Without it a connect to a sleeping or absent device sits + in the system TCP timeout for over a minute before launchd gets its exit code, + which stretches every reconnect attempt for no benefit on a LAN. **4. Load it and confirm:** diff --git a/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md b/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md index 9f5ac0d..c871114 100644 --- a/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md +++ b/docs/superpowers/specs/2026-07-22-nanokvm-mcp-sidecar-design.md @@ -246,13 +246,18 @@ On-device means no stdio, so the endpoint is network-reachable and **must** auth independently of tool annotations. This is transport security, not a guardrail. - Streamable HTTP via the official `modelcontextprotocol/go-sdk` (v1.6.1) -- Bearer token from `/etc/kvm/.nanokvm_mcp_token`, mode 0600, generated on first run +- ~~Bearer token from `/etc/kvm/.nanokvm_mcp_token`, mode 0600, generated on first run~~ + **Corrected 2026-07-30 (#16):** as built, the token comes from the `NANOKVM_MCP_TOKEN` + environment variable (`internal/config/config.go`), which the init script sources from + `/root/nanokvm-mcp/nanokvm-mcp.env`; if unset, one is generated per start and logged to + `/data/nanokvm-mcp/daemon.log`. No `/etc/kvm/.nanokvm_mcp_token` is ever read — the only + `/etc/kvm` token this project touches is the firmware's `.picoclaw_internal_token`. - Constant-time comparison; non-bearer requests get 401 - **Default bind `127.0.0.1:8080`.** Exposing it requires explicit configuration. A keystroke injector should not become LAN-reachable by default. -- The README documents Tailscale as the recommended exposure path. The stock firmware +- ~~The README documents Tailscale as the recommended exposure path. The stock firmware already supports installing it, so putting this on a tailnet instead of the LAN is the - largest available security win and costs nothing. + largest available security win and costs nothing.~~ **Superseded — see below.** **Corrected 2026-07-30 (#16).** The security reasoning holds — the LAN is the wrong place for this listener — but "costs nothing" was wrong, and the recommendation has