From 83fcd008357d14746bb3188d29e80f50b60ad647 Mon Sep 17 00:00:00 2001 From: Serge Gatezh <2880401+gatezh@users.noreply.github.com> Date: Wed, 9 Sep 2026 10:42:06 -0600 Subject: [PATCH] docs: move host-level runbooks to the wiki and record the split rule The two Docker disk documents added in #123 describe Docker Desktop's settings-store.json, the vscode volume and a Mac's free space. Nothing in this repo can invalidate them; they go stale when Docker Desktop changes. Keeping them in-tree meant they were reviewed on this repo's cadence for no benefit, and edited through branch -> PR -> CI at exactly the moment you least want that: during an incident, with a full disk. Moved to the wiki as "Docker Disk Maintenance" and "Incident: Docker Disk Exhaustion (2026-09-07)", with their relative cross-links rewritten and Sources footers citing #123. README keeps its entries, repointed. The rule that decided it is now in .claude/CLAUDE.md rather than left implicit: a document that goes stale when this repo's code changes belongs in the repo; one that goes stale when an external system changes, or that spans several images, belongs in the wiki. That is why per-image READMEs do not move -- a tag list has to change in the same commit as its Dockerfile, and splitting them guarantees drift. The wiki gains three consumer-facing pages in the same pass (Choosing an Image, Image Tags and Rebuild Policy, Troubleshooting), so README now points there for cross-image questions no single README owns. Tradeoff accepted: the moved content is no longer reviewed in a PR and no longer in this repo's git log. It has its own history in the wiki repo, and it has zero coupling to code here, so nothing that could drift does. --- .claude/CLAUDE.md | 20 +++ README.md | 14 +- docs/disk-usage.md | 171 ---------------------- docs/docker-maintenance-cheatsheet.md | 200 -------------------------- 4 files changed, 31 insertions(+), 374 deletions(-) delete mode 100644 docs/disk-usage.md delete mode 100644 docs/docker-maintenance-cheatsheet.md diff --git a/.claude/CLAUDE.md b/.claude/CLAUDE.md index 1b81825..95fe306 100644 --- a/.claude/CLAUDE.md +++ b/.claude/CLAUDE.md @@ -28,6 +28,26 @@ Dockerfiles for custom devcontainer images on GitHub Container Registry (ghcr.io - Standalone: `{primary-version}` only (e.g., `0.11.0`) - Multi-tool images: `{tool1}{version}-{tool2}{version}` (e.g., `bun1.3.9-hugo0.156.0`) +## Where Documentation Goes + +**If a document goes stale when this repo's code changes, it belongs in the repo. If it goes stale when an external system changes β€” Docker Desktop, GitHub settings, an upstream tool β€” or it spans several images, it belongs in the [wiki](https://github.com/gatezh/devcontainers/wiki).** + +| In the repo | | +|---|---| +| Per-image `README.md` | A tag list or tool table must change in the same commit as its Dockerfile | +| `.claude/CLAUDE.md`, `.claude/rules/*` | Conventions enforced in review | +| `.github/workflows/README.md` | Describes the workflows beside it | +| `docs/plans/`, `docs/superpowers/` | Planning and design artifacts | + +| In the wiki | | +|---|---| +| Cross-image guides | Belong to no single image | +| Host / Docker Desktop procedures | Track Docker, not this repo | +| GitHub settings runbooks | Configuration that can't be reviewed in a PR | +| Incident writeups | Operational history, not code | + +Never duplicate: wiki pages link to READMEs, READMEs link back. Every wiki page ends with a `## Sources` section citing the PRs/issues it came from and a `*Last verified:*` date. + ## Code Style - 2-space indentation in JSON/YAML diff --git a/README.md b/README.md index 25de221..c0d8dc7 100644 --- a/README.md +++ b/README.md @@ -2,6 +2,8 @@ This repository contains Dockerfiles for custom Docker images hosted on GitHub Container Registry (ghcr.io). +**New here?** The [wiki](https://github.com/gatezh/devcontainers/wiki) has a [guide to picking an image](https://github.com/gatezh/devcontainers/wiki/Choosing-an-Image) and explains [what `latest` means and when it moves](https://github.com/gatezh/devcontainers/wiki/Image-Tags-and-Rebuild-Policy). + ## πŸ“š Image Documentation ### Devcontainer Images @@ -16,10 +18,16 @@ This repository contains Dockerfiles for custom Docker images hosted on GitHub C - **[ralphex-fe](./ralphex-fe/README.md)** - Bun + Hugo Extended on ralphex base (standalone image) -### Maintenance +## πŸ“– Guides (wiki) + +Cross-image guides and host-level procedures live in the [wiki](https://github.com/gatezh/devcontainers/wiki), because they go stale when Docker or GitHub changes rather than when this repo does. -- **[Docker maintenance cheatsheet](./docs/docker-maintenance-cheatsheet.md)** - commands for "low disk space", a hung Docker, and safe cleanup -- **[Why Docker fills the disk](./docs/disk-usage.md)** - what actually grows, why, and the settings that prevent an outage +- **[Choosing an Image](https://github.com/gatezh/devcontainers/wiki/Choosing-an-Image)** - which of the six images you want +- **[Image Tags and Rebuild Policy](https://github.com/gatezh/devcontainers/wiki/Image-Tags-and-Rebuild-Policy)** - what `latest` means, which tags are immutable, when rebuilds happen +- **[Troubleshooting](https://github.com/gatezh/devcontainers/wiki/Troubleshooting)** - symptoms that span more than one image +- **[Docker Disk Maintenance](https://github.com/gatezh/devcontainers/wiki/Docker-Disk-Maintenance)** - commands for "low disk space", a hung Docker, and safe cleanup +- **[Incident: Docker Disk Exhaustion (2026-09-07)](https://github.com/gatezh/devcontainers/wiki/Incident-Docker-Disk-Exhaustion-2026-09-07)** - what actually grows, why, and the settings that prevent an outage +- **[Branch Protection and Renovate Auto-Merge](https://github.com/gatezh/devcontainers/wiki/Branch-Protection-and-Renovate-Auto-Merge)** - maintainer runbook ## Repository Structure diff --git a/docs/disk-usage.md b/docs/disk-usage.md deleted file mode 100644 index bfe705d..0000000 --- a/docs/disk-usage.md +++ /dev/null @@ -1,171 +0,0 @@ -# Why Docker fills the disk, and what actually fixes it - -Written after the 2026-09-07 incident: Docker Desktop's VM died with -`no space left on device`, taking every container with it. The Mac had 2.8 GB -free; `Docker.raw` had grown to 80 GB. This is the third occurrence. - -The point of this document is that **almost none of it was a Docker bug or a -misconfiguration in this repo**. It is normal, expected growth in caches that -nothing prunes, behind a limit that was never set. - ---- - -## The one setting that turns a nuisance into an outage - -Docker Desktop's `settings-store.json` had **no `DiskSizeMiB` key**, so the VM -disk defaulted to a **1 TB** virtual maximum: - -``` -ls -lh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw -# -rw-r--r-- 1 user staff 1.0T <- apparent (virtual max) -du -sh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw -# 80G <- actually allocated (sparse file) -``` - -With a 1 TB ceiling on a 460 GB Mac, Docker's effective limit is *the entire -machine*. So instead of Docker hitting its own wall and returning a normal -`no space left on device` to whatever was writing, it kept growing until macOS -ran out β€” and the VM itself died. - -**Fix: set a disk usage limit below your typical free space.** 64 GB was chosen -(steady-state need is ~25–30 GB). Docker Desktop β†’ Settings β†’ Resources β†’ -Advanced β†’ Disk usage limit. - -This is the single highest-value change. It does not stop the growth; it makes -the failure survivable and local to Docker. - ---- - -## Where the space actually goes - -Measured on 2026-09-08, inside a 53 GB `Docker.raw`: - -| Consumer | Size | Pruned by anything? | -|---|---|---| -| `vscode` volume (VS Code server + extension cache) | 23.5 GB | **No** | -| Named volumes (node_modules, pgdata, claude-config…) | 31 GB total | Only manually | -| Images | 9.5 GB | `docker image prune` | -| Build cache | 3.9 GB | builder GC (configured) | -| Container writable layers | 6.3 GB | `docker container prune` | - -### 1. The `vscode` volume β€” the big one - -Created automatically by the Dev Containers extension (nothing in this repo -mounts it) to cache the VS Code Server across container rebuilds. It contained: - -- **11.1 GB of server installs** β€” 17 of them. One directory per VS Code commit, - ~500 MB–1.3 GB each, going back ~9 months. Old ones are never deleted. -- **11 GB of `extensionsCache`** β€” of which `anthropic.claude-code` alone was - **84 cached versions = 6.7 GB**, and `openai.chatgpt` 26 versions = 3.5 GB. -- **819 MB of orphaned `.vsix` downloads** β€” UUID-named temp files left behind by - interrupted downloads. - -Two things make this worse here than for most people: - -- **AI extensions ship several releases per week.** A cache designed for - monthly-ish updates now takes ~30x the churn, and keeps every version. -- **Running both Alpine and glibc devcontainers doubles it.** Every VS Code - release is installed twice, as `linux-arm64` *and* `alpine-arm64`. - -### 2. Daily `:latest` pulls leave dangling images - -The images in this repo are rebuilt by CI daily, and each devcontainer's -`initializeCommand` runs `docker pull …:latest` on open. The previous ~2.5 GB -image becomes dangling locally and nothing removes it automatically β€” builder GC -only touches build cache, and the GHCR cleanup workflow only touches the remote -registry. Regular `docker image prune` is required. - -### 3. node_modules volumes are duplicated per variant - -Compose prefixes volume names with the project name, so the default and -`-sandbox` variants of a devcontainer get *separate* `node_modules` volumes even -though `${localWorkspaceFolderBasename}` was intended to share them. Opening both -variants stores dependencies twice. - -### 4. Long-running containers accumulate writable layers - -`mise install` / `bun install` caches land in the container's writable layer -rather than a volume. One long-lived container reached 13 GB on its own. - ---- - -## Two macOS behaviours that make diagnosis confusing - -**`df -h /` lies.** It reports the sealed read-only system snapshot. Real free -space is on the data volume: - -``` -df -h /System/Volumes/Data -``` - -**Freeing space inside Docker may not return it to the Mac.** APFS local Time -Machine snapshots are copy-on-write and pin the old blocks. After reclaiming -29 GB and watching `Docker.raw` shrink 80 β†’ 51 GB, host free space did not move -at all, because two same-day snapshots still referenced the old contents. They -expire on their own within ~24h and the space then returns. - -``` -tmutil listlocalsnapshots / # if space did not come back, look here first -``` - ---- - -## What to do about it β€” official tools only - -Ranked by value: - -1. **Set the Docker disk usage limit** (above). Do this regardless of everything - else. -2. **Delete the `vscode` volume periodically** β€” stop your devcontainers, then - `docker volume rm vscode`. The extension recreates it containing only the - current server. Reclaims ~18 GB; costs one server re-download per - devcontainer, once. -3. **`docker image prune -f`** on a regular basis β€” this is the one that offsets - the daily `:latest` pulls. -4. **VS Code's own commands**: `Dev Containers: Clean Up Dev Containers…` and - `Dev Containers: Clean Up Dev Volumes…`. - -See [`docker-maintenance-cheatsheet.md`](./docker-maintenance-cheatsheet.md) for -the exact commands. - -### There is no built-in cleanup to enable - -Verified against Microsoft's documentation and the extension manifest: VS Code -ships **no** automatic pruning of the server cache or `extensionsCache`, and no -setting to bound their size. There is nothing being "missed". - -The one relevant official switch is **`dev.containers.cacheVolume`** (default -`true`) β€” *"Controls whether a Docker volume should be used to cache the VS Code -server and extensions."* Setting it to `false` removes the shared volume -entirely, and the server then lives in each container's writable layer, which -`docker container prune` can reclaim. - -**Not recommended for this setup**: because these devcontainers pull `:latest` on -every open and CI rebuilds daily, containers are recreated often, and each -recreation would re-download the server. The shared cache is genuinely earning -its keep here β€” it just needs occasional emptying. - ---- - -## Things that look like solutions but are not - -- **`docker system prune -a --volumes`** β€” wipes database and credential volumes - (`*-pgdata`, `*-claude-config`, `*-fish-data`). Never run it here. -- **`docker volume prune`** β€” same problem; orphaned `*-claude-config` and - `*-fish-data` volumes are kept deliberately, they hold credentials and history. -- **Scheduled-cleanup containers** β€” Watchtower was archived in Dec 2025 and - `spotify/docker-gc` is unmaintained. -- **Shrinking the Docker disk to reclaim space** β€” reducing the limit recreates - the disk image and destroys all volumes. Back up first (see the runbook in - `~/docker-volume-backups/`). - ---- - -## Prevention checklist - -- [ ] Docker disk usage limit set (64 GB) -- [ ] `docker image prune -f` run periodically -- [ ] `vscode` volume emptied when it exceeds ~10 GB -- [ ] `"log-driver": "local"` in `~/.docker/daemon.json` β€” replaces unbounded - `json-file` container logs with rotating, compressed ones -- [ ] Retired project? `docker compose down -v` in its folder releases its volumes diff --git a/docs/docker-maintenance-cheatsheet.md b/docs/docker-maintenance-cheatsheet.md deleted file mode 100644 index 98f44f4..0000000 --- a/docs/docker-maintenance-cheatsheet.md +++ /dev/null @@ -1,200 +0,0 @@ -# Docker maintenance cheatsheet (macOS) - -Stock commands only β€” nothing custom to install. Background and the *why* live in -[`disk-usage.md`](./disk-usage.md). - ---- - -## "Low disk space" warning β€” do this first - -```bash -# 1. Real free space (df -h / reports the read-only system snapshot, ignore it) -df -h /System/Volumes/Data - -# 2. How big is Docker really? (ls shows the sparse virtual max, du shows actual) -du -sh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw - -# 3. What inside Docker is using it -docker system df -``` - -## Safe reclaim, in order of value - -```bash -docker image prune -f # dangling images only - safe, offsets daily :latest pulls -docker builder prune -f # build cache beyond the keep-storage in daemon.json -docker container prune -f # STOPPED containers only - check nothing was killed mid-work -``` - -`Docker.raw` auto-TRIMs after a prune, so the file shrinks on its own. No manual -`fstrim` needed on current Docker Desktop. - -## The big one: the `vscode` volume - -Usually the single largest item (23.5 GB when last measured). Nothing prunes it. - -```bash -docker system df -v | grep -i vscode # check its size first -``` - -**Closing VS Code is not enough, and neither is stopping the containers.** -`docker volume rm` blocks on any container that *references* the volume, running -or not β€” the containers must be removed. Expect this error otherwise: - -``` -Error response from daemon: remove vscode: volume is in use - [] -``` - -Full procedure: - -```bash -# 1. Which containers reference it -docker ps -a --filter volume=vscode --format '{{.Names}}\t{{.State}}' - -# 2. Stop any that are running. Note: devcontainers with -# restart: unless-stopped come back by themselves and stay up even with -# VS Code closed. -docker stop - -# 3. Remove them. NEVER add -v: that would delete the named volumes too, -# including *-claude-config (credentials) and *-pgdata (databases). -docker rm - -# 4. Now the volume will go -docker volume rm vscode -``` - -Safe to do because devcontainer source is bind-mounted from the host and state -lives in named volumes, both of which survive `docker rm`. What you lose is the -container writable layer (mise/bun caches) β€” rebuilt on next "Reopen in -Container", along with a one-time VS Code server re-download. - -Do this when the volume exceeds ~10 GB, roughly quarterly. - -## VS Code's own cleanup (Command Palette) - -- `Dev Containers: Clean Up Dev Containers…` -- `Dev Containers: Clean Up Dev Volumes…` - -## Never run these - -```bash -docker system prune -a --volumes # DESTROYS pgdata / claude-config / fish-data -docker volume prune -a # same - those volumes hold credentials and DBs -``` - -To release a *retired* project's volumes deliberately: `docker compose down -v` -in that project's folder. - ---- - -## Reclaimed space didn't show up in `df`? - -APFS local snapshots pin the freed blocks β€” copy-on-write means deleting data -inside `Docker.raw` returns nothing to the pool while a snapshot references it. - -```bash -tmutil listlocalsnapshots / # same-day snapshots are the usual cause -``` - -They expire on their own within ~24h and the space returns. To force it (this -is what macOS itself runs under pressure, `4` = urgency): - -```bash -tmutil thinlocalsnapshots / 30000000000 4 -``` - -Deleting local snapshots does **not** affect Time Machine backups on an external -or network disk. - ---- - -## Docker is hung β€” every command just sits there - -Symptom: `docker ps` never returns (rather than erroring). Usually means the -Linux VM died but the host-side backend still holds the socket. - -```bash -# Confirm it -grep -iE 'no space|GET /error' \ - ~/Library/Containers/com.docker.docker/Data/log/host/com.docker.backend.log | tail - -# Recover -osascript -e 'quit app "Docker Desktop"' -pgrep -f 'com\.docker\.back[e]nd' # note the PID, then: kill -9 -open -a Docker -until docker info >/dev/null 2>&1; do sleep 5; done; echo "daemon up" -``` - -Note the `back[e]nd` bracket trick: a plain `pkill -f "com.docker.backend"` -matches the killing shell's own command line and kills itself instead. - -**After a crash, `docker container prune -f` is not safe.** Containers that were -*running* are now `Exited (255)` and indistinguishable from long-idle ones. -Separate them by stop time before pruning: - -```bash -docker inspect -f '{{.State.Status}}|{{.State.FinishedAt}}|{{.Name}}' $(docker ps -aq) | sort -t'|' -k2 -``` - -Crash-killed containers all share the daemon-boot timestamp. - ---- - -## Settings worth having - -**Disk usage limit** β€” Docker Desktop β†’ Settings β†’ Resources β†’ Advanced. -Set it *below* your typical free space (64 GB here). Without it the default is a -1 TB virtual disk, i.e. Docker can consume the whole Mac. Reducing the limit -recreates the disk and destroys all volumes β€” back up first. - -**Log rotation** β€” `~/.docker/daemon.json`: - -```json -{ - "builder": { "gc": { "enabled": true, "defaultKeepStorage": "5GB" } }, - "log-driver": "local" -} -``` - -Every line a container writes to stdout/stderr is stored on disk, and the default -`json-file` driver **never rotates unless you tell it to** β€” its `max-size` -defaults to `-1` (unlimited) with `max-file: 1`. A chatty dev server or a -database logging every query grows that file forever, and it shows up nowhere in -`docker system df`. - -| | `json-file` (default) | `local` | -|---|---|---| -| Rotation | none (`max-size: -1`, `max-file: 1`) | 20 MB Γ— 5 files | -| Compression | no | yes, by default | -| Cap per container | unbounded | ~100 MB | - -`docker logs` works normally with `local`. Docker's docs do *not* state a -preference between the two drivers, so this is a better-defaults choice rather -than an official recommendation β€” the equally valid alternative is keeping -`json-file` and adding `"log-opts": {"max-size": "10m", "max-file": "3"}`, which -preserves compatibility if anything ever parses the raw JSON log files. With -`local`, don't: Docker warns those files "are designed to be exclusively -accessed by the Docker daemon". - -Requires a Docker restart, and applies only to **newly created** containers β€” so -the cheapest moment to change it is right after a reset, when none exist. - ---- - -## Volume backups - -Tarballs, manifest, and `backup.sh` / `restore.sh` live in -`~/docker-volume-backups/`. Restore verifies sha256 against `MANIFEST.txt` and -refuses to overwrite a volume that a running container has mounted. - -```bash -~/docker-volume-backups//restore.sh --verify # integrity check only -~/docker-volume-backups//restore.sh # restore missing volumes -~/docker-volume-backups//restore.sh --force # REPLACE existing ones -``` - -Worth backing up: `*-claude-config`, `*-fish-data`, `*-fish-history`, `*pgdata*`, -`*sql-data*`, `*storage-data*`. -Regenerable, don't bother: `vscode`, `*node-modules*`, `*playwright-browsers*`, -images, container layers.