Repository navigation
docs: add Docker disk usage runbook and maintenance cheatsheet - #123
Merged
Merged
Conversation
Docker Desktop's VM died on 2026-09-07 with "no space left on device", taking every container with it — the third occurrence. The cause was not a bug or a misconfiguration in this repo, but unbounded caches behind a limit never set. - docs/disk-usage.md — why it happens. settings-store.json had no DiskSizeMiB, so the VM disk defaulted to a 1 TB virtual max; on a 460 GB Mac that makes Docker's effective limit the whole machine, so a full disk kills the VM instead of returning ENOSPC to the caller. Documents the growth sources specific to this setup: the Dev Containers shared `vscode` volume (23.5 GB — 17 server installs, plus 11 GB of extensionsCache of which anthropic.claude-code alone was 84 versions / 6.7 GB), daily :latest pulls leaving dangling images, node_modules volumes duplicated across the default and -sandbox variants, and writable-layer growth from in-container caches. Running both Alpine and glibc devcontainers doubles every server install. Also records the two macOS behaviours that make this hard to diagnose: `df -h /` reports the sealed read-only system snapshot, and APFS local snapshots pin freed blocks so reclaimed space does not return for ~24h. - docs/docker-maintenance-cheatsheet.md — what to run. Triage, safe reclaim order, emptying the vscode volume, recovering a hung daemon, the post-crash `container prune` caveat, and volume backup/restore. Stock Docker and VS Code tooling only. There is no built-in pruning to enable: the Dev Containers extension ships no automatic cleanup of the server cache and no setting to bound its size. The one relevant switch, dev.containers.cacheVolume, is documented along with why disabling it is the wrong trade-off here — these devcontainers pull :latest on open, so containers are recreated often and each recreation would re-download the server.
The log-rotation section asserted that json-file "grows unbounded" without saying why or what local replaces it with, which is the kind of claim a reader has to go verify before acting on it. Per Docker's driver docs: json-file defaults to max-size -1 (unlimited) with max-file 1, so it does not rotate at all unless configured; local defaults to 20 MB x 5 files with compression on, capping a container at roughly 100 MB. Adds that comparison as a table, notes docker logs still works, and records the caveat that local's files are "designed to be exclusively accessed by the Docker daemon". Also stops overclaiming: Docker's docs state no preference between the two drivers, so this is presented as a better-defaults choice, with the json-file + log-opts alternative given for anyone who needs raw-JSON compatibility. Notes that the setting only affects newly created containers, which makes a post-reset moment the cheapest time to apply it.
gatezh
added a commit
that referenced
this pull request
Sep 9, 2026
`automerge: true` has never merged a PR since it was added in #116. #121 sat open, green and CLEAN for 3.5 weeks; #124 was on the same path. Root cause is a race created by `platformAutomerge: false`. That setting means only a Renovate run can merge, and a run merges when it observes an already-green branch. But @anthropic-ai/claude-code ships ~2 releases/day while Renovate runs every 2-9 days, so every run found a newer version, force-pushed the branch (resetting CI to pending) and ended seconds later. The last run is typical: pushed at 19:56:39, run ended 19:56:45, first check went green at 19:56:53, last at 19:59:40 -- nobody was watching. The run that could merge is always the run that just invalidated CI. Note this is not fixable with `minimumReleaseAge`: at any threshold there are still newly-eligible versions by the next run, so the force-push repeats. Switch to `platformAutomerge: true` so GitHub's native auto-merge merges on green with no Renovate run involved. This is also the freshest option -- no version-age delay at all. Native auto-merge needs something to wait for, i.e. branch protection with a required check, and a required check that never runs blocks a PR forever. CI is currently path-filtered at the `on:` level, so a docs-only PR (#123 touches only README.md and docs/*.md) triggers no CI at all and would deadlock. So drop the paths filter and add one `CI complete` job aggregating the others, passing on success-or-skipped so path-filtered builds still don't block. Per-image builds are still gated by detect-changes; the always-on jobs are lint-only. Verified with actionlint (exit 0, no findings).
gatezh
added a commit
that referenced
this pull request
Sep 9, 2026
* fix(ci): make Renovate auto-merge actually fire `automerge: true` has never merged a PR since it was added in #116. #121 sat open, green and CLEAN for 3.5 weeks; #124 was on the same path. Root cause is a race created by `platformAutomerge: false`. That setting means only a Renovate run can merge, and a run merges when it observes an already-green branch. But @anthropic-ai/claude-code ships ~2 releases/day while Renovate runs every 2-9 days, so every run found a newer version, force-pushed the branch (resetting CI to pending) and ended seconds later. The last run is typical: pushed at 19:56:39, run ended 19:56:45, first check went green at 19:56:53, last at 19:59:40 -- nobody was watching. The run that could merge is always the run that just invalidated CI. Note this is not fixable with `minimumReleaseAge`: at any threshold there are still newly-eligible versions by the next run, so the force-push repeats. Switch to `platformAutomerge: true` so GitHub's native auto-merge merges on green with no Renovate run involved. This is also the freshest option -- no version-age delay at all. Native auto-merge needs something to wait for, i.e. branch protection with a required check, and a required check that never runs blocks a PR forever. CI is currently path-filtered at the `on:` level, so a docs-only PR (#123 touches only README.md and docs/*.md) triggers no CI at all and would deadlock. So drop the paths filter and add one `CI complete` job aggregating the others, passing on success-or-skipped so path-filtered builds still don't block. Per-image builds are still gated by detect-changes; the always-on jobs are lint-only. Verified with actionlint (exit 0, no findings). * fix(ci): drop redundant platformAutomerge, soak non-claude bumps, guard the gate Review follow-ups on this branch. platformAutomerge:true is Renovate's own default (renovate-schema.json: platformAutomerge.default = true), so the explicit setting was noise. Deleted it and kept only the part of the comment that is still load-bearing: what the config depends on being configured on the GitHub side. These bumps merge unreviewed and publish to ghcr.io, so a compromised upstream release would reach the published images with no human in the loop. Added minimumReleaseAge: '3 days' as a soak period, with a second packageRule clearing it for @anthropic-ai/claude-code, which is tracked at latest on purpose. internalChecksFilter defaults to 'strict', so a too-young version is never offered and the group PR simply carries whichever tools are eligible. ci-complete is about to become the only required check on master, gating unattended merges, so two hardening changes: - A new job added to this workflow but omitted from `needs` would fail while the gate stayed green. The first step now derives the job list from the workflow file with yq and fails if `needs` has drifted. - The failure message named a bare result ('failure') with no job attached. Iterating toJSON(needs) instead of join(needs.*.result) keeps the job ids, so the error now says which job failed and how. Also recorded why this job uses always() rather than !cancelled(): GitHub counts a skipped required check as passing, so !cancelled() would turn a cancelled run into a green gate. Verified: actionlint exit 0; renovate-config-validator "Config validated successfully"; both jq filters and the yq job-list extraction exercised locally against success/skipped/failure/cancelled fixtures.
gatezh
added a commit
that referenced
this pull request
Sep 9, 2026
…133) The two Docker disk documents added in #123 describe Docker Desktop's settings-store.json, the vscode volume and a Mac's free space. Nothing in this repo can invalidate them; they go stale when Docker Desktop changes. Keeping them in-tree meant they were reviewed on this repo's cadence for no benefit, and edited through branch -> PR -> CI at exactly the moment you least want that: during an incident, with a full disk. Moved to the wiki as "Docker Disk Maintenance" and "Incident: Docker Disk Exhaustion (2026-09-07)", with their relative cross-links rewritten and Sources footers citing #123. README keeps its entries, repointed. The rule that decided it is now in .claude/CLAUDE.md rather than left implicit: a document that goes stale when this repo's code changes belongs in the repo; one that goes stale when an external system changes, or that spans several images, belongs in the wiki. That is why per-image READMEs do not move -- a tag list has to change in the same commit as its Dockerfile, and splitting them guarantees drift. The wiki gains three consumer-facing pages in the same pass (Choosing an Image, Image Tags and Rebuild Policy, Troubleshooting), so README now points there for cross-image questions no single README owns. Tradeoff accepted: the moved content is no longer reviewed in a PR and no longer in this repo's git log. It has its own history in the wiki repo, and it has zero coupling to code here, so nothing that could drift does.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Docker Desktop's VM died on 2026-09-07 with
no space left on device, taking every container with it — the third occurrence of this. Recovery took a while mostly because the failure looks like something else: everydockercommand hangs rather than erroring, since the host-side backend keeps holding the socket after the Linux VM dies.None of it turned out to be a bug or a misconfiguration in this repo. It was normal growth in caches that nothing prunes, behind a disk limit that was never set. That is worth writing down, because the knowledge otherwise lives in one person's memory of one bad afternoon.
What's here
docs/disk-usage.md— why it happens:settings-store.jsonhad noDiskSizeMiB, so the VM disk defaulted to a 1 TB virtual max. On a 460 GB Mac that makes Docker's effective limit the entire machine — so instead of Docker hitting its own wall and returningENOSPC, it grew until macOS ran out and the VM died.vscodevolume at 23.5 GB (17 VS Code server installs, plus 11 GB ofextensionsCache—anthropic.claude-codealone was 84 cached versions / 6.7 GB), daily:latestpulls orphaning the previous image,node_modulesvolumes duplicated across the default and-sandboxvariants, and writable-layer growth from in-containermise/buncaches.linux-arm64andalpine-arm64.df -h /reports the sealed read-only system snapshot (real free space is on/System/Volumes/Data), and APFS local snapshots pin freed blocks — after reclaiming 29 GB and watchingDocker.rawshrink 80 → 51 GB, host free space did not move at all until the snapshots expired.docs/docker-maintenance-cheatsheet.md— what to run: triage, safe reclaim order, emptying thevscodevolume, recovering a hung daemon, and volume backup/restore.Both are linked from the README under a new "Maintenance" heading.
Notes
docker volume rm vscode, which reclaims the same space with one documented command.extensionsCache, and no setting to bound them. The one relevant switch isdev.containers.cacheVolume(defaulttrue), documented along with why disabling it is the wrong trade-off here — these devcontainers pull:lateston open, so containers are recreated often and each recreation would re-download the server..github/workflows/is triggered by markdown paths.