Skip to content

docs: add Docker disk usage runbook and maintenance cheatsheet - #123

Merged
gatezh merged 3 commits into
masterfrom
docs/docker-disk-maintenance
Sep 9, 2026
Merged

gatezh merged 3 commits into
masterfrom
docs/docker-disk-maintenance

Conversation

@gatezh

@gatezh gatezh commented Sep 8, 2026

Copy link
Copy Markdown
Owner

Why

Docker Desktop's VM died on 2026-09-07 with no space left on device, taking every container with it — the third occurrence of this. Recovery took a while mostly because the failure looks like something else: every docker command hangs rather than erroring, since the host-side backend keeps holding the socket after the Linux VM dies.

None of it turned out to be a bug or a misconfiguration in this repo. It was normal growth in caches that nothing prunes, behind a disk limit that was never set. That is worth writing down, because the knowledge otherwise lives in one person's memory of one bad afternoon.

What's here

docs/disk-usage.md — why it happens:

  • settings-store.json had no DiskSizeMiB, so the VM disk defaulted to a 1 TB virtual max. On a 460 GB Mac that makes Docker's effective limit the entire machine — so instead of Docker hitting its own wall and returning ENOSPC, it grew until macOS ran out and the VM died.
  • Where the space actually goes, measured: the Dev Containers shared vscode volume at 23.5 GB (17 VS Code server installs, plus 11 GB of extensionsCache — anthropic.claude-code alone was 84 cached versions / 6.7 GB), daily :latest pulls orphaning the previous image, node_modules volumes duplicated across the default and -sandbox variants, and writable-layer growth from in-container mise/bun caches.
  • Running both Alpine and glibc devcontainers doubles every VS Code release — each lands as linux-arm64 and alpine-arm64.
  • Two macOS behaviours that make diagnosis confusing: df -h / reports the sealed read-only system snapshot (real free space is on /System/Volumes/Data), and APFS local snapshots pin freed blocks — after reclaiming 29 GB and watching Docker.raw shrink 80 → 51 GB, host free space did not move at all until the snapshots expired.

docs/docker-maintenance-cheatsheet.md — what to run: triage, safe reclaim order, emptying the vscode volume, recovering a hung daemon, and volume backup/restore.

Both are linked from the README under a new "Maintenance" heading.

Notes

  • Stock tooling only. No scripts to install, nothing scheduled. An earlier draft of this was a custom cleanup script on a LaunchAgent; it was dropped in favour of docker volume rm vscode, which reclaims the same space with one documented command.
  • There is no built-in cleanup being missed. Verified against Microsoft's docs and the extension manifest: the Dev Containers extension ships no automatic pruning of the server cache or extensionsCache, and no setting to bound them. The one relevant switch is dev.containers.cacheVolume (default true), documented along with why disabling it is the wrong trade-off here — these devcontainers pull :latest on open, so containers are recreated often and each recreation would re-download the server.
  • Documents-only change; no workflow in .github/workflows/ is triggered by markdown paths.

Docker Desktop's VM died on 2026-09-07 with "no space left on device", taking
every container with it — the third occurrence. The cause was not a bug or a
misconfiguration in this repo, but unbounded caches behind a limit never set.

- docs/disk-usage.md — why it happens. settings-store.json had no DiskSizeMiB,
  so the VM disk defaulted to a 1 TB virtual max; on a 460 GB Mac that makes
  Docker's effective limit the whole machine, so a full disk kills the VM
  instead of returning ENOSPC to the caller. Documents the growth sources
  specific to this setup: the Dev Containers shared `vscode` volume (23.5 GB —
  17 server installs, plus 11 GB of extensionsCache of which
  anthropic.claude-code alone was 84 versions / 6.7 GB), daily :latest pulls
  leaving dangling images, node_modules volumes duplicated across the default
  and -sandbox variants, and writable-layer growth from in-container caches.
  Running both Alpine and glibc devcontainers doubles every server install.
  Also records the two macOS behaviours that make this hard to diagnose:
  `df -h /` reports the sealed read-only system snapshot, and APFS local
  snapshots pin freed blocks so reclaimed space does not return for ~24h.

- docs/docker-maintenance-cheatsheet.md — what to run. Triage, safe reclaim
  order, emptying the vscode volume, recovering a hung daemon, the post-crash
  `container prune` caveat, and volume backup/restore.

Stock Docker and VS Code tooling only. There is no built-in pruning to enable:
the Dev Containers extension ships no automatic cleanup of the server cache and
no setting to bound its size. The one relevant switch, dev.containers.cacheVolume,
is documented along with why disabling it is the wrong trade-off here — these
devcontainers pull :latest on open, so containers are recreated often and each
recreation would re-download the server.
The log-rotation section asserted that json-file "grows unbounded" without
saying why or what local replaces it with, which is the kind of claim a reader
has to go verify before acting on it.

Per Docker's driver docs: json-file defaults to max-size -1 (unlimited) with
max-file 1, so it does not rotate at all unless configured; local defaults to
20 MB x 5 files with compression on, capping a container at roughly 100 MB.
Adds that comparison as a table, notes docker logs still works, and records the
caveat that local's files are "designed to be exclusively accessed by the Docker
daemon".

Also stops overclaiming: Docker's docs state no preference between the two
drivers, so this is presented as a better-defaults choice, with the json-file +
log-opts alternative given for anyone who needs raw-JSON compatibility. Notes
that the setting only affects newly created containers, which makes a post-reset
moment the cheapest time to apply it.
gatezh added a commit that referenced this pull request Sep 9, 2026
`automerge: true` has never merged a PR since it was added in #116. #121 sat
open, green and CLEAN for 3.5 weeks; #124 was on the same path.

Root cause is a race created by `platformAutomerge: false`. That setting means
only a Renovate run can merge, and a run merges when it observes an
already-green branch. But @anthropic-ai/claude-code ships ~2 releases/day while
Renovate runs every 2-9 days, so every run found a newer version, force-pushed
the branch (resetting CI to pending) and ended seconds later. The last run is
typical: pushed at 19:56:39, run ended 19:56:45, first check went green at
19:56:53, last at 19:59:40 -- nobody was watching. The run that could merge is
always the run that just invalidated CI.

Note this is not fixable with `minimumReleaseAge`: at any threshold there are
still newly-eligible versions by the next run, so the force-push repeats.

Switch to `platformAutomerge: true` so GitHub's native auto-merge merges on
green with no Renovate run involved. This is also the freshest option -- no
version-age delay at all.

Native auto-merge needs something to wait for, i.e. branch protection with a
required check, and a required check that never runs blocks a PR forever. CI is
currently path-filtered at the `on:` level, so a docs-only PR (#123 touches only
README.md and docs/*.md) triggers no CI at all and would deadlock. So drop the
paths filter and add one `CI complete` job aggregating the others, passing on
success-or-skipped so path-filtered builds still don't block. Per-image builds
are still gated by detect-changes; the always-on jobs are lint-only.

Verified with actionlint (exit 0, no findings).
gatezh added a commit that referenced this pull request Sep 9, 2026
* fix(ci): make Renovate auto-merge actually fire

`automerge: true` has never merged a PR since it was added in #116. #121 sat
open, green and CLEAN for 3.5 weeks; #124 was on the same path.

Root cause is a race created by `platformAutomerge: false`. That setting means
only a Renovate run can merge, and a run merges when it observes an
already-green branch. But @anthropic-ai/claude-code ships ~2 releases/day while
Renovate runs every 2-9 days, so every run found a newer version, force-pushed
the branch (resetting CI to pending) and ended seconds later. The last run is
typical: pushed at 19:56:39, run ended 19:56:45, first check went green at
19:56:53, last at 19:59:40 -- nobody was watching. The run that could merge is
always the run that just invalidated CI.

Note this is not fixable with `minimumReleaseAge`: at any threshold there are
still newly-eligible versions by the next run, so the force-push repeats.

Switch to `platformAutomerge: true` so GitHub's native auto-merge merges on
green with no Renovate run involved. This is also the freshest option -- no
version-age delay at all.

Native auto-merge needs something to wait for, i.e. branch protection with a
required check, and a required check that never runs blocks a PR forever. CI is
currently path-filtered at the `on:` level, so a docs-only PR (#123 touches only
README.md and docs/*.md) triggers no CI at all and would deadlock. So drop the
paths filter and add one `CI complete` job aggregating the others, passing on
success-or-skipped so path-filtered builds still don't block. Per-image builds
are still gated by detect-changes; the always-on jobs are lint-only.

Verified with actionlint (exit 0, no findings).

* fix(ci): drop redundant platformAutomerge, soak non-claude bumps, guard the gate

Review follow-ups on this branch.

platformAutomerge:true is Renovate's own default (renovate-schema.json:
platformAutomerge.default = true), so the explicit setting was noise. Deleted
it and kept only the part of the comment that is still load-bearing: what the
config depends on being configured on the GitHub side.

These bumps merge unreviewed and publish to ghcr.io, so a compromised upstream
release would reach the published images with no human in the loop. Added
minimumReleaseAge: '3 days' as a soak period, with a second packageRule
clearing it for @anthropic-ai/claude-code, which is tracked at latest on
purpose. internalChecksFilter defaults to 'strict', so a too-young version is
never offered and the group PR simply carries whichever tools are eligible.

ci-complete is about to become the only required check on master, gating
unattended merges, so two hardening changes:

- A new job added to this workflow but omitted from `needs` would fail while
  the gate stayed green. The first step now derives the job list from the
  workflow file with yq and fails if `needs` has drifted.
- The failure message named a bare result ('failure') with no job attached.
  Iterating toJSON(needs) instead of join(needs.*.result) keeps the job ids,
  so the error now says which job failed and how.

Also recorded why this job uses always() rather than !cancelled(): GitHub
counts a skipped required check as passing, so !cancelled() would turn a
cancelled run into a green gate.

Verified: actionlint exit 0; renovate-config-validator "Config validated
successfully"; both jq filters and the yq job-list extraction exercised
locally against success/skipped/failure/cancelled fixtures.
@gatezh
gatezh merged commit 945219c into master Sep 9, 2026
10 checks passed
@gatezh
gatezh deleted the docs/docker-disk-maintenance branch September 9, 2026 16:26
gatezh added a commit that referenced this pull request Sep 9, 2026
…133)

The two Docker disk documents added in #123 describe Docker Desktop's
settings-store.json, the vscode volume and a Mac's free space. Nothing in this
repo can invalidate them; they go stale when Docker Desktop changes. Keeping
them in-tree meant they were reviewed on this repo's cadence for no benefit,
and edited through branch -> PR -> CI at exactly the moment you least want that:
during an incident, with a full disk.

Moved to the wiki as "Docker Disk Maintenance" and "Incident: Docker Disk
Exhaustion (2026-09-07)", with their relative cross-links rewritten and Sources
footers citing #123. README keeps its entries, repointed.

The rule that decided it is now in .claude/CLAUDE.md rather than left implicit:
a document that goes stale when this repo's code changes belongs in the repo;
one that goes stale when an external system changes, or that spans several
images, belongs in the wiki. That is why per-image READMEs do not move -- a tag
list has to change in the same commit as its Dockerfile, and splitting them
guarantees drift.

The wiki gains three consumer-facing pages in the same pass (Choosing an Image,
Image Tags and Rebuild Policy, Troubleshooting), so README now points there for
cross-image questions no single README owns.

Tradeoff accepted: the moved content is no longer reviewed in a PR and no longer
in this repo's git log. It has its own history in the wiki repo, and it has zero
coupling to code here, so nothing that could drift does.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant