diff --git a/README.md b/README.md index fff3e7c..25de221 100644 --- a/README.md +++ b/README.md @@ -16,6 +16,11 @@ This repository contains Dockerfiles for custom Docker images hosted on GitHub C - **[ralphex-fe](./ralphex-fe/README.md)** - Bun + Hugo Extended on ralphex base (standalone image) +### Maintenance + +- **[Docker maintenance cheatsheet](./docs/docker-maintenance-cheatsheet.md)** - commands for "low disk space", a hung Docker, and safe cleanup +- **[Why Docker fills the disk](./docs/disk-usage.md)** - what actually grows, why, and the settings that prevent an outage + ## Repository Structure Each subdirectory represents a Docker image project. Devcontainer images use the following structure: diff --git a/docs/disk-usage.md b/docs/disk-usage.md new file mode 100644 index 0000000..bfe705d --- /dev/null +++ b/docs/disk-usage.md @@ -0,0 +1,171 @@ +# Why Docker fills the disk, and what actually fixes it + +Written after the 2026-09-07 incident: Docker Desktop's VM died with +`no space left on device`, taking every container with it. The Mac had 2.8 GB +free; `Docker.raw` had grown to 80 GB. This is the third occurrence. + +The point of this document is that **almost none of it was a Docker bug or a +misconfiguration in this repo**. It is normal, expected growth in caches that +nothing prunes, behind a limit that was never set. + +--- + +## The one setting that turns a nuisance into an outage + +Docker Desktop's `settings-store.json` had **no `DiskSizeMiB` key**, so the VM +disk defaulted to a **1 TB** virtual maximum: + +``` +ls -lh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw +# -rw-r--r-- 1 user staff 1.0T <- apparent (virtual max) +du -sh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw +# 80G <- actually allocated (sparse file) +``` + +With a 1 TB ceiling on a 460 GB Mac, Docker's effective limit is *the entire +machine*. So instead of Docker hitting its own wall and returning a normal +`no space left on device` to whatever was writing, it kept growing until macOS +ran out — and the VM itself died. + +**Fix: set a disk usage limit below your typical free space.** 64 GB was chosen +(steady-state need is ~25–30 GB). Docker Desktop → Settings → Resources → +Advanced → Disk usage limit. + +This is the single highest-value change. It does not stop the growth; it makes +the failure survivable and local to Docker. + +--- + +## Where the space actually goes + +Measured on 2026-09-08, inside a 53 GB `Docker.raw`: + +| Consumer | Size | Pruned by anything? | +|---|---|---| +| `vscode` volume (VS Code server + extension cache) | 23.5 GB | **No** | +| Named volumes (node_modules, pgdata, claude-config…) | 31 GB total | Only manually | +| Images | 9.5 GB | `docker image prune` | +| Build cache | 3.9 GB | builder GC (configured) | +| Container writable layers | 6.3 GB | `docker container prune` | + +### 1. The `vscode` volume — the big one + +Created automatically by the Dev Containers extension (nothing in this repo +mounts it) to cache the VS Code Server across container rebuilds. It contained: + +- **11.1 GB of server installs** — 17 of them. One directory per VS Code commit, + ~500 MB–1.3 GB each, going back ~9 months. Old ones are never deleted. +- **11 GB of `extensionsCache`** — of which `anthropic.claude-code` alone was + **84 cached versions = 6.7 GB**, and `openai.chatgpt` 26 versions = 3.5 GB. +- **819 MB of orphaned `.vsix` downloads** — UUID-named temp files left behind by + interrupted downloads. + +Two things make this worse here than for most people: + +- **AI extensions ship several releases per week.** A cache designed for + monthly-ish updates now takes ~30x the churn, and keeps every version. +- **Running both Alpine and glibc devcontainers doubles it.** Every VS Code + release is installed twice, as `linux-arm64` *and* `alpine-arm64`. + +### 2. Daily `:latest` pulls leave dangling images + +The images in this repo are rebuilt by CI daily, and each devcontainer's +`initializeCommand` runs `docker pull …:latest` on open. The previous ~2.5 GB +image becomes dangling locally and nothing removes it automatically — builder GC +only touches build cache, and the GHCR cleanup workflow only touches the remote +registry. Regular `docker image prune` is required. + +### 3. node_modules volumes are duplicated per variant + +Compose prefixes volume names with the project name, so the default and +`-sandbox` variants of a devcontainer get *separate* `node_modules` volumes even +though `${localWorkspaceFolderBasename}` was intended to share them. Opening both +variants stores dependencies twice. + +### 4. Long-running containers accumulate writable layers + +`mise install` / `bun install` caches land in the container's writable layer +rather than a volume. One long-lived container reached 13 GB on its own. + +--- + +## Two macOS behaviours that make diagnosis confusing + +**`df -h /` lies.** It reports the sealed read-only system snapshot. Real free +space is on the data volume: + +``` +df -h /System/Volumes/Data +``` + +**Freeing space inside Docker may not return it to the Mac.** APFS local Time +Machine snapshots are copy-on-write and pin the old blocks. After reclaiming +29 GB and watching `Docker.raw` shrink 80 → 51 GB, host free space did not move +at all, because two same-day snapshots still referenced the old contents. They +expire on their own within ~24h and the space then returns. + +``` +tmutil listlocalsnapshots / # if space did not come back, look here first +``` + +--- + +## What to do about it — official tools only + +Ranked by value: + +1. **Set the Docker disk usage limit** (above). Do this regardless of everything + else. +2. **Delete the `vscode` volume periodically** — stop your devcontainers, then + `docker volume rm vscode`. The extension recreates it containing only the + current server. Reclaims ~18 GB; costs one server re-download per + devcontainer, once. +3. **`docker image prune -f`** on a regular basis — this is the one that offsets + the daily `:latest` pulls. +4. **VS Code's own commands**: `Dev Containers: Clean Up Dev Containers…` and + `Dev Containers: Clean Up Dev Volumes…`. + +See [`docker-maintenance-cheatsheet.md`](./docker-maintenance-cheatsheet.md) for +the exact commands. + +### There is no built-in cleanup to enable + +Verified against Microsoft's documentation and the extension manifest: VS Code +ships **no** automatic pruning of the server cache or `extensionsCache`, and no +setting to bound their size. There is nothing being "missed". + +The one relevant official switch is **`dev.containers.cacheVolume`** (default +`true`) — *"Controls whether a Docker volume should be used to cache the VS Code +server and extensions."* Setting it to `false` removes the shared volume +entirely, and the server then lives in each container's writable layer, which +`docker container prune` can reclaim. + +**Not recommended for this setup**: because these devcontainers pull `:latest` on +every open and CI rebuilds daily, containers are recreated often, and each +recreation would re-download the server. The shared cache is genuinely earning +its keep here — it just needs occasional emptying. + +--- + +## Things that look like solutions but are not + +- **`docker system prune -a --volumes`** — wipes database and credential volumes + (`*-pgdata`, `*-claude-config`, `*-fish-data`). Never run it here. +- **`docker volume prune`** — same problem; orphaned `*-claude-config` and + `*-fish-data` volumes are kept deliberately, they hold credentials and history. +- **Scheduled-cleanup containers** — Watchtower was archived in Dec 2025 and + `spotify/docker-gc` is unmaintained. +- **Shrinking the Docker disk to reclaim space** — reducing the limit recreates + the disk image and destroys all volumes. Back up first (see the runbook in + `~/docker-volume-backups/`). + +--- + +## Prevention checklist + +- [ ] Docker disk usage limit set (64 GB) +- [ ] `docker image prune -f` run periodically +- [ ] `vscode` volume emptied when it exceeds ~10 GB +- [ ] `"log-driver": "local"` in `~/.docker/daemon.json` — replaces unbounded + `json-file` container logs with rotating, compressed ones +- [ ] Retired project? `docker compose down -v` in its folder releases its volumes diff --git a/docs/docker-maintenance-cheatsheet.md b/docs/docker-maintenance-cheatsheet.md new file mode 100644 index 0000000..98f44f4 --- /dev/null +++ b/docs/docker-maintenance-cheatsheet.md @@ -0,0 +1,200 @@ +# Docker maintenance cheatsheet (macOS) + +Stock commands only — nothing custom to install. Background and the *why* live in +[`disk-usage.md`](./disk-usage.md). + +--- + +## "Low disk space" warning — do this first + +```bash +# 1. Real free space (df -h / reports the read-only system snapshot, ignore it) +df -h /System/Volumes/Data + +# 2. How big is Docker really? (ls shows the sparse virtual max, du shows actual) +du -sh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw + +# 3. What inside Docker is using it +docker system df +``` + +## Safe reclaim, in order of value + +```bash +docker image prune -f # dangling images only - safe, offsets daily :latest pulls +docker builder prune -f # build cache beyond the keep-storage in daemon.json +docker container prune -f # STOPPED containers only - check nothing was killed mid-work +``` + +`Docker.raw` auto-TRIMs after a prune, so the file shrinks on its own. No manual +`fstrim` needed on current Docker Desktop. + +## The big one: the `vscode` volume + +Usually the single largest item (23.5 GB when last measured). Nothing prunes it. + +```bash +docker system df -v | grep -i vscode # check its size first +``` + +**Closing VS Code is not enough, and neither is stopping the containers.** +`docker volume rm` blocks on any container that *references* the volume, running +or not — the containers must be removed. Expect this error otherwise: + +``` +Error response from daemon: remove vscode: volume is in use - [] +``` + +Full procedure: + +```bash +# 1. Which containers reference it +docker ps -a --filter volume=vscode --format '{{.Names}}\t{{.State}}' + +# 2. Stop any that are running. Note: devcontainers with +# restart: unless-stopped come back by themselves and stay up even with +# VS Code closed. +docker stop + +# 3. Remove them. NEVER add -v: that would delete the named volumes too, +# including *-claude-config (credentials) and *-pgdata (databases). +docker rm + +# 4. Now the volume will go +docker volume rm vscode +``` + +Safe to do because devcontainer source is bind-mounted from the host and state +lives in named volumes, both of which survive `docker rm`. What you lose is the +container writable layer (mise/bun caches) — rebuilt on next "Reopen in +Container", along with a one-time VS Code server re-download. + +Do this when the volume exceeds ~10 GB, roughly quarterly. + +## VS Code's own cleanup (Command Palette) + +- `Dev Containers: Clean Up Dev Containers…` +- `Dev Containers: Clean Up Dev Volumes…` + +## Never run these + +```bash +docker system prune -a --volumes # DESTROYS pgdata / claude-config / fish-data +docker volume prune -a # same - those volumes hold credentials and DBs +``` + +To release a *retired* project's volumes deliberately: `docker compose down -v` +in that project's folder. + +--- + +## Reclaimed space didn't show up in `df`? + +APFS local snapshots pin the freed blocks — copy-on-write means deleting data +inside `Docker.raw` returns nothing to the pool while a snapshot references it. + +```bash +tmutil listlocalsnapshots / # same-day snapshots are the usual cause +``` + +They expire on their own within ~24h and the space returns. To force it (this +is what macOS itself runs under pressure, `4` = urgency): + +```bash +tmutil thinlocalsnapshots / 30000000000 4 +``` + +Deleting local snapshots does **not** affect Time Machine backups on an external +or network disk. + +--- + +## Docker is hung — every command just sits there + +Symptom: `docker ps` never returns (rather than erroring). Usually means the +Linux VM died but the host-side backend still holds the socket. + +```bash +# Confirm it +grep -iE 'no space|GET /error' \ + ~/Library/Containers/com.docker.docker/Data/log/host/com.docker.backend.log | tail + +# Recover +osascript -e 'quit app "Docker Desktop"' +pgrep -f 'com\.docker\.back[e]nd' # note the PID, then: kill -9 +open -a Docker +until docker info >/dev/null 2>&1; do sleep 5; done; echo "daemon up" +``` + +Note the `back[e]nd` bracket trick: a plain `pkill -f "com.docker.backend"` +matches the killing shell's own command line and kills itself instead. + +**After a crash, `docker container prune -f` is not safe.** Containers that were +*running* are now `Exited (255)` and indistinguishable from long-idle ones. +Separate them by stop time before pruning: + +```bash +docker inspect -f '{{.State.Status}}|{{.State.FinishedAt}}|{{.Name}}' $(docker ps -aq) | sort -t'|' -k2 +``` + +Crash-killed containers all share the daemon-boot timestamp. + +--- + +## Settings worth having + +**Disk usage limit** — Docker Desktop → Settings → Resources → Advanced. +Set it *below* your typical free space (64 GB here). Without it the default is a +1 TB virtual disk, i.e. Docker can consume the whole Mac. Reducing the limit +recreates the disk and destroys all volumes — back up first. + +**Log rotation** — `~/.docker/daemon.json`: + +```json +{ + "builder": { "gc": { "enabled": true, "defaultKeepStorage": "5GB" } }, + "log-driver": "local" +} +``` + +Every line a container writes to stdout/stderr is stored on disk, and the default +`json-file` driver **never rotates unless you tell it to** — its `max-size` +defaults to `-1` (unlimited) with `max-file: 1`. A chatty dev server or a +database logging every query grows that file forever, and it shows up nowhere in +`docker system df`. + +| | `json-file` (default) | `local` | +|---|---|---| +| Rotation | none (`max-size: -1`, `max-file: 1`) | 20 MB × 5 files | +| Compression | no | yes, by default | +| Cap per container | unbounded | ~100 MB | + +`docker logs` works normally with `local`. Docker's docs do *not* state a +preference between the two drivers, so this is a better-defaults choice rather +than an official recommendation — the equally valid alternative is keeping +`json-file` and adding `"log-opts": {"max-size": "10m", "max-file": "3"}`, which +preserves compatibility if anything ever parses the raw JSON log files. With +`local`, don't: Docker warns those files "are designed to be exclusively +accessed by the Docker daemon". + +Requires a Docker restart, and applies only to **newly created** containers — so +the cheapest moment to change it is right after a reset, when none exist. + +--- + +## Volume backups + +Tarballs, manifest, and `backup.sh` / `restore.sh` live in +`~/docker-volume-backups/`. Restore verifies sha256 against `MANIFEST.txt` and +refuses to overwrite a volume that a running container has mounted. + +```bash +~/docker-volume-backups//restore.sh --verify # integrity check only +~/docker-volume-backups//restore.sh # restore missing volumes +~/docker-volume-backups//restore.sh --force # REPLACE existing ones +``` + +Worth backing up: `*-claude-config`, `*-fish-data`, `*-fish-history`, `*pgdata*`, +`*sql-data*`, `*storage-data*`. +Regenerable, don't bother: `vscode`, `*node-modules*`, `*playwright-browsers*`, +images, container layers.