From bebd8a99614917942d2166e01330e1509b357dcb Mon Sep 17 00:00:00 2001 From: Serge Gatezh <2880401+gatezh@users.noreply.github.com> Date: Tue, 8 Sep 2026 11:55:52 -0600 Subject: [PATCH 1/2] docs: add Docker disk usage runbook and maintenance cheatsheet MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Docker Desktop's VM died on 2026-09-07 with "no space left on device", taking every container with it — the third occurrence. The cause was not a bug or a misconfiguration in this repo, but unbounded caches behind a limit never set. - docs/disk-usage.md — why it happens. settings-store.json had no DiskSizeMiB, so the VM disk defaulted to a 1 TB virtual max; on a 460 GB Mac that makes Docker's effective limit the whole machine, so a full disk kills the VM instead of returning ENOSPC to the caller. Documents the growth sources specific to this setup: the Dev Containers shared `vscode` volume (23.5 GB — 17 server installs, plus 11 GB of extensionsCache of which anthropic.claude-code alone was 84 versions / 6.7 GB), daily :latest pulls leaving dangling images, node_modules volumes duplicated across the default and -sandbox variants, and writable-layer growth from in-container caches. Running both Alpine and glibc devcontainers doubles every server install. Also records the two macOS behaviours that make this hard to diagnose: `df -h /` reports the sealed read-only system snapshot, and APFS local snapshots pin freed blocks so reclaimed space does not return for ~24h. - docs/docker-maintenance-cheatsheet.md — what to run. Triage, safe reclaim order, emptying the vscode volume, recovering a hung daemon, the post-crash `container prune` caveat, and volume backup/restore. Stock Docker and VS Code tooling only. There is no built-in pruning to enable: the Dev Containers extension ships no automatic cleanup of the server cache and no setting to bound its size. The one relevant switch, dev.containers.cacheVolume, is documented along with why disabling it is the wrong trade-off here — these devcontainers pull :latest on open, so containers are recreated often and each recreation would re-download the server. --- README.md | 5 + docs/disk-usage.md | 171 ++++++++++++++++++++++++ docs/docker-maintenance-cheatsheet.md | 180 ++++++++++++++++++++++++++ 3 files changed, 356 insertions(+) create mode 100644 docs/disk-usage.md create mode 100644 docs/docker-maintenance-cheatsheet.md diff --git a/README.md b/README.md index fff3e7c..25de221 100644 --- a/README.md +++ b/README.md @@ -16,6 +16,11 @@ This repository contains Dockerfiles for custom Docker images hosted on GitHub C - **[ralphex-fe](./ralphex-fe/README.md)** - Bun + Hugo Extended on ralphex base (standalone image) +### Maintenance + +- **[Docker maintenance cheatsheet](./docs/docker-maintenance-cheatsheet.md)** - commands for "low disk space", a hung Docker, and safe cleanup +- **[Why Docker fills the disk](./docs/disk-usage.md)** - what actually grows, why, and the settings that prevent an outage + ## Repository Structure Each subdirectory represents a Docker image project. Devcontainer images use the following structure: diff --git a/docs/disk-usage.md b/docs/disk-usage.md new file mode 100644 index 0000000..bfe705d --- /dev/null +++ b/docs/disk-usage.md @@ -0,0 +1,171 @@ +# Why Docker fills the disk, and what actually fixes it + +Written after the 2026-09-07 incident: Docker Desktop's VM died with +`no space left on device`, taking every container with it. The Mac had 2.8 GB +free; `Docker.raw` had grown to 80 GB. This is the third occurrence. + +The point of this document is that **almost none of it was a Docker bug or a +misconfiguration in this repo**. It is normal, expected growth in caches that +nothing prunes, behind a limit that was never set. + +--- + +## The one setting that turns a nuisance into an outage + +Docker Desktop's `settings-store.json` had **no `DiskSizeMiB` key**, so the VM +disk defaulted to a **1 TB** virtual maximum: + +``` +ls -lh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw +# -rw-r--r-- 1 user staff 1.0T <- apparent (virtual max) +du -sh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw +# 80G <- actually allocated (sparse file) +``` + +With a 1 TB ceiling on a 460 GB Mac, Docker's effective limit is *the entire +machine*. So instead of Docker hitting its own wall and returning a normal +`no space left on device` to whatever was writing, it kept growing until macOS +ran out — and the VM itself died. + +**Fix: set a disk usage limit below your typical free space.** 64 GB was chosen +(steady-state need is ~25–30 GB). Docker Desktop → Settings → Resources → +Advanced → Disk usage limit. + +This is the single highest-value change. It does not stop the growth; it makes +the failure survivable and local to Docker. + +--- + +## Where the space actually goes + +Measured on 2026-09-08, inside a 53 GB `Docker.raw`: + +| Consumer | Size | Pruned by anything? | +|---|---|---| +| `vscode` volume (VS Code server + extension cache) | 23.5 GB | **No** | +| Named volumes (node_modules, pgdata, claude-config…) | 31 GB total | Only manually | +| Images | 9.5 GB | `docker image prune` | +| Build cache | 3.9 GB | builder GC (configured) | +| Container writable layers | 6.3 GB | `docker container prune` | + +### 1. The `vscode` volume — the big one + +Created automatically by the Dev Containers extension (nothing in this repo +mounts it) to cache the VS Code Server across container rebuilds. It contained: + +- **11.1 GB of server installs** — 17 of them. One directory per VS Code commit, + ~500 MB–1.3 GB each, going back ~9 months. Old ones are never deleted. +- **11 GB of `extensionsCache`** — of which `anthropic.claude-code` alone was + **84 cached versions = 6.7 GB**, and `openai.chatgpt` 26 versions = 3.5 GB. +- **819 MB of orphaned `.vsix` downloads** — UUID-named temp files left behind by + interrupted downloads. + +Two things make this worse here than for most people: + +- **AI extensions ship several releases per week.** A cache designed for + monthly-ish updates now takes ~30x the churn, and keeps every version. +- **Running both Alpine and glibc devcontainers doubles it.** Every VS Code + release is installed twice, as `linux-arm64` *and* `alpine-arm64`. + +### 2. Daily `:latest` pulls leave dangling images + +The images in this repo are rebuilt by CI daily, and each devcontainer's +`initializeCommand` runs `docker pull …:latest` on open. The previous ~2.5 GB +image becomes dangling locally and nothing removes it automatically — builder GC +only touches build cache, and the GHCR cleanup workflow only touches the remote +registry. Regular `docker image prune` is required. + +### 3. node_modules volumes are duplicated per variant + +Compose prefixes volume names with the project name, so the default and +`-sandbox` variants of a devcontainer get *separate* `node_modules` volumes even +though `${localWorkspaceFolderBasename}` was intended to share them. Opening both +variants stores dependencies twice. + +### 4. Long-running containers accumulate writable layers + +`mise install` / `bun install` caches land in the container's writable layer +rather than a volume. One long-lived container reached 13 GB on its own. + +--- + +## Two macOS behaviours that make diagnosis confusing + +**`df -h /` lies.** It reports the sealed read-only system snapshot. Real free +space is on the data volume: + +``` +df -h /System/Volumes/Data +``` + +**Freeing space inside Docker may not return it to the Mac.** APFS local Time +Machine snapshots are copy-on-write and pin the old blocks. After reclaiming +29 GB and watching `Docker.raw` shrink 80 → 51 GB, host free space did not move +at all, because two same-day snapshots still referenced the old contents. They +expire on their own within ~24h and the space then returns. + +``` +tmutil listlocalsnapshots / # if space did not come back, look here first +``` + +--- + +## What to do about it — official tools only + +Ranked by value: + +1. **Set the Docker disk usage limit** (above). Do this regardless of everything + else. +2. **Delete the `vscode` volume periodically** — stop your devcontainers, then + `docker volume rm vscode`. The extension recreates it containing only the + current server. Reclaims ~18 GB; costs one server re-download per + devcontainer, once. +3. **`docker image prune -f`** on a regular basis — this is the one that offsets + the daily `:latest` pulls. +4. **VS Code's own commands**: `Dev Containers: Clean Up Dev Containers…` and + `Dev Containers: Clean Up Dev Volumes…`. + +See [`docker-maintenance-cheatsheet.md`](./docker-maintenance-cheatsheet.md) for +the exact commands. + +### There is no built-in cleanup to enable + +Verified against Microsoft's documentation and the extension manifest: VS Code +ships **no** automatic pruning of the server cache or `extensionsCache`, and no +setting to bound their size. There is nothing being "missed". + +The one relevant official switch is **`dev.containers.cacheVolume`** (default +`true`) — *"Controls whether a Docker volume should be used to cache the VS Code +server and extensions."* Setting it to `false` removes the shared volume +entirely, and the server then lives in each container's writable layer, which +`docker container prune` can reclaim. + +**Not recommended for this setup**: because these devcontainers pull `:latest` on +every open and CI rebuilds daily, containers are recreated often, and each +recreation would re-download the server. The shared cache is genuinely earning +its keep here — it just needs occasional emptying. + +--- + +## Things that look like solutions but are not + +- **`docker system prune -a --volumes`** — wipes database and credential volumes + (`*-pgdata`, `*-claude-config`, `*-fish-data`). Never run it here. +- **`docker volume prune`** — same problem; orphaned `*-claude-config` and + `*-fish-data` volumes are kept deliberately, they hold credentials and history. +- **Scheduled-cleanup containers** — Watchtower was archived in Dec 2025 and + `spotify/docker-gc` is unmaintained. +- **Shrinking the Docker disk to reclaim space** — reducing the limit recreates + the disk image and destroys all volumes. Back up first (see the runbook in + `~/docker-volume-backups/`). + +--- + +## Prevention checklist + +- [ ] Docker disk usage limit set (64 GB) +- [ ] `docker image prune -f` run periodically +- [ ] `vscode` volume emptied when it exceeds ~10 GB +- [ ] `"log-driver": "local"` in `~/.docker/daemon.json` — replaces unbounded + `json-file` container logs with rotating, compressed ones +- [ ] Retired project? `docker compose down -v` in its folder releases its volumes diff --git a/docs/docker-maintenance-cheatsheet.md b/docs/docker-maintenance-cheatsheet.md new file mode 100644 index 0000000..c592a09 --- /dev/null +++ b/docs/docker-maintenance-cheatsheet.md @@ -0,0 +1,180 @@ +# Docker maintenance cheatsheet (macOS) + +Stock commands only — nothing custom to install. Background and the *why* live in +[`disk-usage.md`](./disk-usage.md). + +--- + +## "Low disk space" warning — do this first + +```bash +# 1. Real free space (df -h / reports the read-only system snapshot, ignore it) +df -h /System/Volumes/Data + +# 2. How big is Docker really? (ls shows the sparse virtual max, du shows actual) +du -sh ~/Library/Containers/com.docker.docker/Data/vms/0/data/Docker.raw + +# 3. What inside Docker is using it +docker system df +``` + +## Safe reclaim, in order of value + +```bash +docker image prune -f # dangling images only - safe, offsets daily :latest pulls +docker builder prune -f # build cache beyond the keep-storage in daemon.json +docker container prune -f # STOPPED containers only - check nothing was killed mid-work +``` + +`Docker.raw` auto-TRIMs after a prune, so the file shrinks on its own. No manual +`fstrim` needed on current Docker Desktop. + +## The big one: the `vscode` volume + +Usually the single largest item (23.5 GB when last measured). Nothing prunes it. + +```bash +docker system df -v | grep -i vscode # check its size first +``` + +**Closing VS Code is not enough, and neither is stopping the containers.** +`docker volume rm` blocks on any container that *references* the volume, running +or not — the containers must be removed. Expect this error otherwise: + +``` +Error response from daemon: remove vscode: volume is in use - [] +``` + +Full procedure: + +```bash +# 1. Which containers reference it +docker ps -a --filter volume=vscode --format '{{.Names}}\t{{.State}}' + +# 2. Stop any that are running. Note: devcontainers with +# restart: unless-stopped come back by themselves and stay up even with +# VS Code closed. +docker stop + +# 3. Remove them. NEVER add -v: that would delete the named volumes too, +# including *-claude-config (credentials) and *-pgdata (databases). +docker rm + +# 4. Now the volume will go +docker volume rm vscode +``` + +Safe to do because devcontainer source is bind-mounted from the host and state +lives in named volumes, both of which survive `docker rm`. What you lose is the +container writable layer (mise/bun caches) — rebuilt on next "Reopen in +Container", along with a one-time VS Code server re-download. + +Do this when the volume exceeds ~10 GB, roughly quarterly. + +## VS Code's own cleanup (Command Palette) + +- `Dev Containers: Clean Up Dev Containers…` +- `Dev Containers: Clean Up Dev Volumes…` + +## Never run these + +```bash +docker system prune -a --volumes # DESTROYS pgdata / claude-config / fish-data +docker volume prune -a # same - those volumes hold credentials and DBs +``` + +To release a *retired* project's volumes deliberately: `docker compose down -v` +in that project's folder. + +--- + +## Reclaimed space didn't show up in `df`? + +APFS local snapshots pin the freed blocks — copy-on-write means deleting data +inside `Docker.raw` returns nothing to the pool while a snapshot references it. + +```bash +tmutil listlocalsnapshots / # same-day snapshots are the usual cause +``` + +They expire on their own within ~24h and the space returns. To force it (this +is what macOS itself runs under pressure, `4` = urgency): + +```bash +tmutil thinlocalsnapshots / 30000000000 4 +``` + +Deleting local snapshots does **not** affect Time Machine backups on an external +or network disk. + +--- + +## Docker is hung — every command just sits there + +Symptom: `docker ps` never returns (rather than erroring). Usually means the +Linux VM died but the host-side backend still holds the socket. + +```bash +# Confirm it +grep -iE 'no space|GET /error' \ + ~/Library/Containers/com.docker.docker/Data/log/host/com.docker.backend.log | tail + +# Recover +osascript -e 'quit app "Docker Desktop"' +pgrep -f 'com\.docker\.back[e]nd' # note the PID, then: kill -9 +open -a Docker +until docker info >/dev/null 2>&1; do sleep 5; done; echo "daemon up" +``` + +Note the `back[e]nd` bracket trick: a plain `pkill -f "com.docker.backend"` +matches the killing shell's own command line and kills itself instead. + +**After a crash, `docker container prune -f` is not safe.** Containers that were +*running* are now `Exited (255)` and indistinguishable from long-idle ones. +Separate them by stop time before pruning: + +```bash +docker inspect -f '{{.State.Status}}|{{.State.FinishedAt}}|{{.Name}}' $(docker ps -aq) | sort -t'|' -k2 +``` + +Crash-killed containers all share the daemon-boot timestamp. + +--- + +## Settings worth having + +**Disk usage limit** — Docker Desktop → Settings → Resources → Advanced. +Set it *below* your typical free space (64 GB here). Without it the default is a +1 TB virtual disk, i.e. Docker can consume the whole Mac. Reducing the limit +recreates the disk and destroys all volumes — back up first. + +**Log rotation** — `~/.docker/daemon.json`: + +```json +{ + "builder": { "gc": { "enabled": true, "defaultKeepStorage": "5GB" } }, + "log-driver": "local" +} +``` + +`local` rotates and compresses; the default `json-file` grows unbounded. +Requires a Docker restart, and applies to newly created containers. + +--- + +## Volume backups + +Tarballs, manifest, and `backup.sh` / `restore.sh` live in +`~/docker-volume-backups/`. Restore verifies sha256 against `MANIFEST.txt` and +refuses to overwrite a volume that a running container has mounted. + +```bash +~/docker-volume-backups//restore.sh --verify # integrity check only +~/docker-volume-backups//restore.sh # restore missing volumes +~/docker-volume-backups//restore.sh --force # REPLACE existing ones +``` + +Worth backing up: `*-claude-config`, `*-fish-data`, `*-fish-history`, `*pgdata*`, +`*sql-data*`, `*storage-data*`. +Regenerable, don't bother: `vscode`, `*node-modules*`, `*playwright-browsers*`, +images, container layers. From 072889dc6aa54ba5d6f9ba99d5768da31fb598bb Mon Sep 17 00:00:00 2001 From: Serge Gatezh <2880401+gatezh@users.noreply.github.com> Date: Tue, 8 Sep 2026 12:19:36 -0600 Subject: [PATCH 2/2] docs: spell out what the local log driver actually changes The log-rotation section asserted that json-file "grows unbounded" without saying why or what local replaces it with, which is the kind of claim a reader has to go verify before acting on it. Per Docker's driver docs: json-file defaults to max-size -1 (unlimited) with max-file 1, so it does not rotate at all unless configured; local defaults to 20 MB x 5 files with compression on, capping a container at roughly 100 MB. Adds that comparison as a table, notes docker logs still works, and records the caveat that local's files are "designed to be exclusively accessed by the Docker daemon". Also stops overclaiming: Docker's docs state no preference between the two drivers, so this is presented as a better-defaults choice, with the json-file + log-opts alternative given for anyone who needs raw-JSON compatibility. Notes that the setting only affects newly created containers, which makes a post-reset moment the cheapest time to apply it. --- docs/docker-maintenance-cheatsheet.md | 24 ++++++++++++++++++++++-- 1 file changed, 22 insertions(+), 2 deletions(-) diff --git a/docs/docker-maintenance-cheatsheet.md b/docs/docker-maintenance-cheatsheet.md index c592a09..98f44f4 100644 --- a/docs/docker-maintenance-cheatsheet.md +++ b/docs/docker-maintenance-cheatsheet.md @@ -157,8 +157,28 @@ recreates the disk and destroys all volumes — back up first. } ``` -`local` rotates and compresses; the default `json-file` grows unbounded. -Requires a Docker restart, and applies to newly created containers. +Every line a container writes to stdout/stderr is stored on disk, and the default +`json-file` driver **never rotates unless you tell it to** — its `max-size` +defaults to `-1` (unlimited) with `max-file: 1`. A chatty dev server or a +database logging every query grows that file forever, and it shows up nowhere in +`docker system df`. + +| | `json-file` (default) | `local` | +|---|---|---| +| Rotation | none (`max-size: -1`, `max-file: 1`) | 20 MB × 5 files | +| Compression | no | yes, by default | +| Cap per container | unbounded | ~100 MB | + +`docker logs` works normally with `local`. Docker's docs do *not* state a +preference between the two drivers, so this is a better-defaults choice rather +than an official recommendation — the equally valid alternative is keeping +`json-file` and adding `"log-opts": {"max-size": "10m", "max-file": "3"}`, which +preserves compatibility if anything ever parses the raw JSON log files. With +`local`, don't: Docker warns those files "are designed to be exclusively +accessed by the Docker daemon". + +Requires a Docker restart, and applies only to **newly created** containers — so +the cheapest moment to change it is right after a reset, when none exist. ---