diff --git a/active.md b/active.md index 538d0ce8..fd9a4176 100644 --- a/active.md +++ b/active.md @@ -51,42 +51,6 @@ - current verdict: Heart's last committed dashboard (2026-08-11T05:51Z, i.e. BEFORE the green run) reads STALE score 65, listing `no release validation for current source` among its evidence gaps. That is the STALE tier behaving correctly — an evidence gap, not a fault, and this ingest is its remedy. - do-not: do NOT re-dispatch `Release Integrate` to "refresh" this. The run is green and its artifact is live; a re-dispatch costs ~70 minutes of CI and proves nothing new. Only re-dispatch if the artifact has expired or main has moved. - repos-none-claimed: this entry claims NO repos — deliberately on one line, NOT as 2-space ` - Repo` bullets, because `worktree_check_conflict` treats any such bullet as a live claim. - -## release-drive-2026-08-03 -- issue: (no issue — a human-authorized manual release drive, not a dev task) -- session: claude --resume e0105850-b98b-47ff-9ada-cba04a455a65 -- status: SHIPPED 2026-08-07 — superseded by release-drive-2026-08-07 below, which carried this drive's payload to PyPI as 2026.8.7.1. This entry is now history; do NOT re-run it. (Was: Stage 2/3 CLEAR TO RE-RUN on request, 2026-08-04.) -- split-from: simulator-util-to-af-ex (#1444), closed out 2026-08-04 → complete/2026/08/simulator-util-to-af-ex.md. That task shipped; this release drive it opened did not, so it was split out rather than buried in a completion record. -- why a release is owed: workspace main now calls `af.ex.util` helpers that exist on PyAutoFit main but NOT on PyPI. Control test PROVED the breakage — AttributeError on the first simulator call against released autofit 2026.7.29.2 in a clean venv with PYTHONPATH unset — and HowToFit ships NO datasets (`dataset/` gitignored, 0 tracked files), so a new user gets no data at all. The `pending-release` gate was overridden 2026-08-03 on explicit human instruction ("all five once green") with that consequence stated, which makes publishing the remedy, not a preference. -- release-drive: human authorized driving a release 2026-08-03 (chose "Drive the release" over merging early or re-cutting workspace-only). Drive via `pyauto-brain release validate` — NOT the nightly driver; AUTONOMY.md forbids converting a manual release into the scheduled-nightly exception. -- release-progress: Stage 0/1 preflight PASS. Stage 2 rehearsal #1 (run 30841336540, dev69901) DISCARDED — PyAutoLens#686 merged 18:31:35Z DURING that build while the PyAutoLens job checked out at 18:27:31Z, so those wheels lacked #686; writing the post-build live-main sha would have attested to source the wheels never contained. Stage 2 rehearsal #2 = run 30841883371 SUCCESS → testpypi 2026.8.3.1.dev70001, verified every library main HEAD commit-time PREDATES its job checkout before writing commit_shas.json. Artifacts dir `~/.pyauto-heart/manual_validation_20260803_pm` (rehearsal.json + testpypi_version.txt + commit_shas.json). -- commit-shas (from rehearsal #2): PyAutoNerves e82c17fd / PyAutoFit 26033fb4 / PyAutoArray 54ba44e8 / PyAutoGalaxy 4249384b / PyAutoLens 4927738e. SUPERSEDED — do not use. The shipped SHAs are recorded under release-drive-2026-08-07 below; these were already stale on 2026-08-04 and the 2026-08-07 drive re-derived them from a fresh rehearsal. -- release-outcome (Stage 3 run 30842349506, COMPLETE): 30 jobs, exactly 2 failures, both diagnosed and both now FIXED AND MERGED. (1) autolens/point_source — #453 updated only ONE of the two model blocks in scripts/point_source/start_here.py, leaving prose saying PointSolved and code saying Point; with PyAutoLens#686 making solved all-to-all the default that raises PointProfileMismatchException. Regression proven: the same script passed 40.4s in run 30788224561 and failed 4.5s here. Fixed by autolens_workspace#461, MERGED 2026-08-03T21:41:40Z. (2) verify_install check D (`pip rc=0 import rc=1`) — fixed by PyAutoHeart#134 (`--pre` on the TestPyPI path), squash-MERGED 2026-08-04 as 46a331a, both pytest legs green on head fb0eff4b, one file `heart/checks/verify_install.sh` +24/-1. -- release-resume: re-rehearse (new wheels, so check D resolves the candidate family with `--pre`), re-dispatch integrate, then `gh run download -R PyAutoLabs/PyAutoHeart -n release-stage-report -D ` and `pyauto-brain release validate --ingest --commit-shas /commit_shas.json`. Never `--force` a RED/YELLOW without a fresh human ack. On GREEN the PUBLISH step is a SEPARATE human decision. -- release-scope-flag: RESOLVED 2026-08-04 — this release also carries PyAutoLens#686 (point-source defaults, 4927738e), which was not one of the fixes the drive set out to ship. Human answer: ship it. Verified before accepting: #686's exp-3 merge gate had landed two days before the merge, Tests+Docs green on 4927738e, it carries the `## API Changes` heading the breaking-change release notes require, and a workspace sweep found all 8 `al.AnalysisPoint` call sites relying on the new all-solved default compose `al.ps.PointSolved` — the only mismatch (point_source/start_here.py) was autolens_workspace#461, merged. -- do-not: do NOT pick up `ep.py` or any `ep*` script as a release blocker — human decision 2026-08-03 to park them for a while (PyAutoFit #1332 F10 tracks the underlying EP message-projection instability). They are parked NEEDS_FIX in autofit_workspace_test config/build/no_run.yaml via #82, which `run.py` loads unconditionally and profile-independently, so the parking holds for the RELEASE profile too. -- dep-floor-regression: CLEARED 2026-08-03 21:30Z — complete/2026/08/dep-floors-source-chain-ci.md (PyAutoNerves#146). Root cause was the `1.0.dev0` source stamp, not the floors; the floors stand unchanged. All five previously-blocked workspace PRs re-ran green and merged. -- correctives-worktree: REMOVED 2026-08-04 — both correctives merged, `~/Code/PyAutoLabs-wt/release-validate-correctives` gone (worktrees + local/remote `feature/release-validate-correctives` branches deleted in autolens_workspace and PyAutoHeart, `worktree prune` run in both; the dir held only tracked files, no output/ artifacts). -- heart-context (2026-08-04): verdict YELLOW, score 70, `red_reasons: []`. Reasons = workspace validation not passing / tenant-firewall manifest drift (2 mismatches vs repos.yaml) / release validation stale (source moved since rehearsal — expected, see commit-shas above). The workspace-validation reason has shrunk since the snapshot: autofit_workspace_test#84 fixed the jax_assertions entry it names. -- repos-none-claimed: this entry claims NO repos — deliberately listed on one line, NOT as ` - Repo` bullets, because `worktree_check_conflict` treats any 2-space ` - ` bullet as a live claim regardless of which field it sits under. - -## release-drive-2026-08-07 -- issue: (no issue — a human-authorized manual release drive, not a dev task) -- status: SHIPPED 2026-08-07. All five libraries published to PyPI at **2026.8.7.1**. Verified independently against pypi.org (not just the run conclusion): each package reports `latest = 2026.8.7.1`, and `autolens-2026.8.7.1-py3-none-any.whl` + `autogalaxy-2026.8.7.1-py3-none-any.whl` were downloaded from the live index as proof of installability. -- supersedes: release-drive-2026-08-03 (above). That drive's payload shipped here; its recorded commit_shas were stale and are NOT the shipped set. -- commit-shas (SHIPPED, verified still at origin/main immediately before ingest): PyAutoNerves 5a67f181 / PyAutoFit f02ea7ed / PyAutoArray 828d5c13 / PyAutoGalaxy 63d69b87 / PyAutoLens e4c7ba70. -- discharges: the mge-sigma-min-workspace-sweep RELEASE DEBT (see above) — `sigma_min` confirmed present in the released autogalaxy wheel. -- root-cause of the stalled nightly (2026-08-06): the GitHub Actions outage dropped push triggers, so several merges to main had NO Tests run and two runs failed on runner provisioning. Not a code fault. Remedy: `workflow_dispatch` added to PyAutoGalaxy + PyAutoLens `.github/workflows/main.yml` (PRs #560 / #693, both merged) so a main-HEAD run can be re-requested on demand without an empty commit. -- validation: Stage 0/1 preflight PASS. Stage 2 rehearsal run 31192317261 (PyAutoHands release.yml, rehearsal:true) → testpypi 2026.8.7.1.dev70601. Stage 3 integrate run 31193443960 (PyAutoHeart release-integrate.yml) → 51/51 jobs green, `status: pass`, 657p/0f/101s/0t, verify_install checks A–F all PASS. Artifacts `~/.pyauto-heart/manual_validation_20260807`. -- MGE-regression NOT reproduced: the 2026-08-06 integrate failed on `scripts/interferometer/features/multi_gaussian_expansion/likelihood_function.py` (numpy LinAlgError: Singular matrix). It passed here — and the script genuinely RAN rather than being silently dropped, proven by the count moving 324p+1f → 325p+0f with the total unchanged. -- ingest-trap (cost one cycle): the first readiness tick after a clean ingest came back RED with `stale_reasons: []` and only `PyAutoGalaxy/PyAutoLens: 4 commit(s) behind origin`. That is a LOCAL-clone signal, not a validation failure — the merged workflow_dispatch PRs had never been pulled. Fast-forwarding both local mains to the validated SHAs re-ticked GREEN score 100. Do not `--force` past this; sync the clone. -- pre_build-trap (caught before it fired): `pre_build` runs `black scripts/` then `git add scripts/`, and `git add` on a directory stages UNTRACKED files too — an uncommitted WIP script in `autolens_assistant/scripts/` would have been reformatted and pushed public inside the "pre build" commit. This is the same leak class the script's own comments describe fixing for `dataset/`/`config/` (#126); the `scripts/` path still has the hole. Mitigation used: move the file out of the repo before the run, restore after (verified byte-identical by md5). Worth a real fix so it is not left to operator vigilance. FIXED 2026-08-08 on branch `claude/automind-task-planning-163wk7` (PyAutoHands) — prompt `draft/bug/pyautohands/pre_build_stages_untracked_wip.md`. The hazard was first REPRODUCED against the pre-fix script on throwaway fixture repos (private file committed as "pre build" and pushed to the remote, exit 0, silently), then closed by two legs: a fail-fast preflight sweeping all 13 repos for untracked files under `notebooks/`/`scripts/`/`slam_pipeline/` before the first is touched (it must precede everything — run_workspace pushes each repo before moving to the next), and staging narrowed to `git add -u` plus explicit adds of run-created files, so the directory-wide form cannot return. Covered by `tests/test_pre_build_staging.py`, which runs the real script against fixture git repos with real bare remotes. No `--allow-dirty` override by design. Also answers the open atomicity question in PyAutoHands `docs/pre_build_failure_audit.md` §6. SHIPPED 2026-08-08: PyAutoHands#232 (issue, closed completed) → PyAutoHands#233 MERGED as a5bac76b, all three pytest legs green (3.12/3.13/3.14, 309 passed / 4 skipped — the count reconciles with local, confirming the new fixture tests actually ran rather than being collected-but-skipped). Mind bookkeeping PyAutoMind#152 MERGED as 62869a2e. -- release-run: 31200419263 (PyAutoHands release.yml, rehearsal:false → live). All five `release (...)` publish jobs SUCCESS; tags pushed and PyPI agree, so the line-428 hazard (upload timing out AFTER tagging) did not occur. -- post-publish failures (do NOT re-drive the release for these — the publish is complete and correct): (1) `wiki_currency_check` autolens — died on `No matching distribution found for autolens==2026.8.7.1` ~4 min after upload; a PyPI index-propagation race, proven by `wiki_currency_check_autofit` starting 2s earlier and PASSING, and by the wheel downloading fine minutes later. RESOLVED 2026-08-07 — autolens_assistant is CLEAN on all five legs, so the race was the whole story and no autolens follow-up is owed. A CI job re-run was impossible (`HTTP 403: The workflow run containing this job is already running`), so it was graded locally by the documented method instead: a fresh venv with `autolens==2026.8.7.1` from PyPI, `PYTHONPATH` cleared, and all four libraries verified to resolve to venv site-packages rather than this workspace's source checkouts (the `baseline-repin-TRAP` — grading against source installs would have been meaningless). Results: `--check-version` clean (baseline matches autolens 2026.8.7.1), `--scope all` 68 files / 143 symbols / **0 missing-broken**, `--lint-idioms` clean (214 files), `--check-citations` 105 files / 413 citations / **0 missing, 0 warnings**, `--check-provenance` **0 errors** (49 pages). Contrast autogalaxy's 5 provenance errors — the two failures shared a red badge but not a cause. (2) `wiki_drift_issue` — pure fallout, `Artifact not found: wiki-drift-report`, because (1) died before writing it. (3) `wiki_currency_check_autogalaxy` — REAL drift, see the provenance prompt filed under draft/maintenance/autogalaxy_assistant/. -- artifacts-are-laptop-only: Actions artifact downloads are blocked from cloud/mobile sessions (egress policy 403s `productionresultssa2.blob.core.windows.net` on CONNECT) — this is what stopped the cloud session finishing the ingest. Both wiki drift reports were captured to `~/.pyauto-heart/release_20260807_wiki_drift/` while on the laptop. -- do-not: do NOT use the nightly driver for a manual release — AUTONOMY.md forbids converting a manual release into the scheduled-nightly exception. -- repos-none-claimed: this entry claims NO repos — deliberately on one line, NOT as 2-space ` - Repo` bullets, because `worktree_check_conflict` treats any such bullet as a live claim. - ## version-stamp-sync-guards - issue: https://github.com/PyAutoLabs/PyAutoHands/issues/235 - prompt: active/version_stamp_sync_and_release_sed_guards.md diff --git a/complete/2026/08/release-drive-2026-08-03.md b/complete/2026/08/release-drive-2026-08-03.md new file mode 100644 index 00000000..f6743dbc --- /dev/null +++ b/complete/2026/08/release-drive-2026-08-03.md @@ -0,0 +1,17 @@ +## release-drive-2026-08-03 +- issue: (no issue — a human-authorized manual release drive, not a dev task) +- completed: 2026-08-07 +- session: claude --resume e0105850-b98b-47ff-9ada-cba04a455a65 +- status: SHIPPED 2026-08-07 — superseded by release-drive-2026-08-07 below, which carried this drive's payload to PyPI as 2026.8.7.1. This entry is now history; do NOT re-run it. (Was: Stage 2/3 CLEAR TO RE-RUN on request, 2026-08-04.) +- split-from: simulator-util-to-af-ex (#1444), closed out 2026-08-04 → complete/2026/08/simulator-util-to-af-ex.md. That task shipped; this release drive it opened did not, so it was split out rather than buried in a completion record. +- why a release is owed: workspace main now calls `af.ex.util` helpers that exist on PyAutoFit main but NOT on PyPI. Control test PROVED the breakage — AttributeError on the first simulator call against released autofit 2026.7.29.2 in a clean venv with PYTHONPATH unset — and HowToFit ships NO datasets (`dataset/` gitignored, 0 tracked files), so a new user gets no data at all. The `pending-release` gate was overridden 2026-08-03 on explicit human instruction ("all five once green") with that consequence stated, which makes publishing the remedy, not a preference. +- release-drive: human authorized driving a release 2026-08-03 (chose "Drive the release" over merging early or re-cutting workspace-only). Drive via `pyauto-brain release validate` — NOT the nightly driver; AUTONOMY.md forbids converting a manual release into the scheduled-nightly exception. +- release-progress: Stage 0/1 preflight PASS. Stage 2 rehearsal #1 (run 30841336540, dev69901) DISCARDED — PyAutoLens#686 merged 18:31:35Z DURING that build while the PyAutoLens job checked out at 18:27:31Z, so those wheels lacked #686; writing the post-build live-main sha would have attested to source the wheels never contained. Stage 2 rehearsal #2 = run 30841883371 SUCCESS → testpypi 2026.8.3.1.dev70001, verified every library main HEAD commit-time PREDATES its job checkout before writing commit_shas.json. Artifacts dir `~/.pyauto-heart/manual_validation_20260803_pm` (rehearsal.json + testpypi_version.txt + commit_shas.json). +- commit-shas (from rehearsal #2): PyAutoNerves e82c17fd / PyAutoFit 26033fb4 / PyAutoArray 54ba44e8 / PyAutoGalaxy 4249384b / PyAutoLens 4927738e. SUPERSEDED — do not use. The shipped SHAs are recorded under release-drive-2026-08-07 below; these were already stale on 2026-08-04 and the 2026-08-07 drive re-derived them from a fresh rehearsal. +- release-outcome (Stage 3 run 30842349506, COMPLETE): 30 jobs, exactly 2 failures, both diagnosed and both now FIXED AND MERGED. (1) autolens/point_source — #453 updated only ONE of the two model blocks in scripts/point_source/start_here.py, leaving prose saying PointSolved and code saying Point; with PyAutoLens#686 making solved all-to-all the default that raises PointProfileMismatchException. Regression proven: the same script passed 40.4s in run 30788224561 and failed 4.5s here. Fixed by autolens_workspace#461, MERGED 2026-08-03T21:41:40Z. (2) verify_install check D (`pip rc=0 import rc=1`) — fixed by PyAutoHeart#134 (`--pre` on the TestPyPI path), squash-MERGED 2026-08-04 as 46a331a, both pytest legs green on head fb0eff4b, one file `heart/checks/verify_install.sh` +24/-1. +- release-resume: re-rehearse (new wheels, so check D resolves the candidate family with `--pre`), re-dispatch integrate, then `gh run download -R PyAutoLabs/PyAutoHeart -n release-stage-report -D ` and `pyauto-brain release validate --ingest --commit-shas /commit_shas.json`. Never `--force` a RED/YELLOW without a fresh human ack. On GREEN the PUBLISH step is a SEPARATE human decision. +- release-scope-flag: RESOLVED 2026-08-04 — this release also carries PyAutoLens#686 (point-source defaults, 4927738e), which was not one of the fixes the drive set out to ship. Human answer: ship it. Verified before accepting: #686's exp-3 merge gate had landed two days before the merge, Tests+Docs green on 4927738e, it carries the `## API Changes` heading the breaking-change release notes require, and a workspace sweep found all 8 `al.AnalysisPoint` call sites relying on the new all-solved default compose `al.ps.PointSolved` — the only mismatch (point_source/start_here.py) was autolens_workspace#461, merged. +- do-not: do NOT pick up `ep.py` or any `ep*` script as a release blocker — human decision 2026-08-03 to park them for a while (PyAutoFit #1332 F10 tracks the underlying EP message-projection instability). They are parked NEEDS_FIX in autofit_workspace_test config/build/no_run.yaml via #82, which `run.py` loads unconditionally and profile-independently, so the parking holds for the RELEASE profile too. +- dep-floor-regression: CLEARED 2026-08-03 21:30Z — complete/2026/08/dep-floors-source-chain-ci.md (PyAutoNerves#146). Root cause was the `1.0.dev0` source stamp, not the floors; the floors stand unchanged. All five previously-blocked workspace PRs re-ran green and merged. +- correctives-worktree: REMOVED 2026-08-04 — both correctives merged, `~/Code/PyAutoLabs-wt/release-validate-correctives` gone (worktrees + local/remote `feature/release-validate-correctives` branches deleted in autolens_workspace and PyAutoHeart, `worktree prune` run in both; the dir held only tracked files, no output/ artifacts). +- heart-context (2026-08-04): verdict YELLOW, score 70, `red_reasons: []`. Reasons = workspace validation not passing / tenant-firewall manifest drift (2 mismatches vs repos.yaml) / release validation stale (source moved since rehearsal — expected, see commit-shas above). The workspace-validation reason has shrunk since the snapshot: autofit_workspace_test#84 fixed the jax_assertions entry it names. diff --git a/complete/2026/08/release-drive-2026-08-07.md b/complete/2026/08/release-drive-2026-08-07.md new file mode 100644 index 00000000..87fd6a48 --- /dev/null +++ b/complete/2026/08/release-drive-2026-08-07.md @@ -0,0 +1,16 @@ +## release-drive-2026-08-07 +- issue: (no issue — a human-authorized manual release drive, not a dev task) +- completed: 2026-08-07 +- status: SHIPPED 2026-08-07. All five libraries published to PyPI at **2026.8.7.1**. Verified independently against pypi.org (not just the run conclusion): each package reports `latest = 2026.8.7.1`, and `autolens-2026.8.7.1-py3-none-any.whl` + `autogalaxy-2026.8.7.1-py3-none-any.whl` were downloaded from the live index as proof of installability. +- supersedes: release-drive-2026-08-03 (above). That drive's payload shipped here; its recorded commit_shas were stale and are NOT the shipped set. +- commit-shas (SHIPPED, verified still at origin/main immediately before ingest): PyAutoNerves 5a67f181 / PyAutoFit f02ea7ed / PyAutoArray 828d5c13 / PyAutoGalaxy 63d69b87 / PyAutoLens e4c7ba70. +- discharges: the mge-sigma-min-workspace-sweep RELEASE DEBT (see above) — `sigma_min` confirmed present in the released autogalaxy wheel. +- root-cause of the stalled nightly (2026-08-06): the GitHub Actions outage dropped push triggers, so several merges to main had NO Tests run and two runs failed on runner provisioning. Not a code fault. Remedy: `workflow_dispatch` added to PyAutoGalaxy + PyAutoLens `.github/workflows/main.yml` (PRs #560 / #693, both merged) so a main-HEAD run can be re-requested on demand without an empty commit. +- validation: Stage 0/1 preflight PASS. Stage 2 rehearsal run 31192317261 (PyAutoHands release.yml, rehearsal:true) → testpypi 2026.8.7.1.dev70601. Stage 3 integrate run 31193443960 (PyAutoHeart release-integrate.yml) → 51/51 jobs green, `status: pass`, 657p/0f/101s/0t, verify_install checks A–F all PASS. Artifacts `~/.pyauto-heart/manual_validation_20260807`. +- MGE-regression NOT reproduced: the 2026-08-06 integrate failed on `scripts/interferometer/features/multi_gaussian_expansion/likelihood_function.py` (numpy LinAlgError: Singular matrix). It passed here — and the script genuinely RAN rather than being silently dropped, proven by the count moving 324p+1f → 325p+0f with the total unchanged. +- ingest-trap (cost one cycle): the first readiness tick after a clean ingest came back RED with `stale_reasons: []` and only `PyAutoGalaxy/PyAutoLens: 4 commit(s) behind origin`. That is a LOCAL-clone signal, not a validation failure — the merged workflow_dispatch PRs had never been pulled. Fast-forwarding both local mains to the validated SHAs re-ticked GREEN score 100. Do not `--force` past this; sync the clone. +- pre_build-trap (caught before it fired): `pre_build` runs `black scripts/` then `git add scripts/`, and `git add` on a directory stages UNTRACKED files too — an uncommitted WIP script in `autolens_assistant/scripts/` would have been reformatted and pushed public inside the "pre build" commit. This is the same leak class the script's own comments describe fixing for `dataset/`/`config/` (#126); the `scripts/` path still has the hole. Mitigation used: move the file out of the repo before the run, restore after (verified byte-identical by md5). Worth a real fix so it is not left to operator vigilance. FIXED 2026-08-08 on branch `claude/automind-task-planning-163wk7` (PyAutoHands) — prompt `draft/bug/pyautohands/pre_build_stages_untracked_wip.md`. The hazard was first REPRODUCED against the pre-fix script on throwaway fixture repos (private file committed as "pre build" and pushed to the remote, exit 0, silently), then closed by two legs: a fail-fast preflight sweeping all 13 repos for untracked files under `notebooks/`/`scripts/`/`slam_pipeline/` before the first is touched (it must precede everything — run_workspace pushes each repo before moving to the next), and staging narrowed to `git add -u` plus explicit adds of run-created files, so the directory-wide form cannot return. Covered by `tests/test_pre_build_staging.py`, which runs the real script against fixture git repos with real bare remotes. No `--allow-dirty` override by design. Also answers the open atomicity question in PyAutoHands `docs/pre_build_failure_audit.md` §6. SHIPPED 2026-08-08: PyAutoHands#232 (issue, closed completed) → PyAutoHands#233 MERGED as a5bac76b, all three pytest legs green (3.12/3.13/3.14, 309 passed / 4 skipped — the count reconciles with local, confirming the new fixture tests actually ran rather than being collected-but-skipped). Mind bookkeeping PyAutoMind#152 MERGED as 62869a2e. +- release-run: 31200419263 (PyAutoHands release.yml, rehearsal:false → live). All five `release (...)` publish jobs SUCCESS; tags pushed and PyPI agree, so the line-428 hazard (upload timing out AFTER tagging) did not occur. +- post-publish failures (do NOT re-drive the release for these — the publish is complete and correct): (1) `wiki_currency_check` autolens — died on `No matching distribution found for autolens==2026.8.7.1` ~4 min after upload; a PyPI index-propagation race, proven by `wiki_currency_check_autofit` starting 2s earlier and PASSING, and by the wheel downloading fine minutes later. RESOLVED 2026-08-07 — autolens_assistant is CLEAN on all five legs, so the race was the whole story and no autolens follow-up is owed. A CI job re-run was impossible (`HTTP 403: The workflow run containing this job is already running`), so it was graded locally by the documented method instead: a fresh venv with `autolens==2026.8.7.1` from PyPI, `PYTHONPATH` cleared, and all four libraries verified to resolve to venv site-packages rather than this workspace's source checkouts (the `baseline-repin-TRAP` — grading against source installs would have been meaningless). Results: `--check-version` clean (baseline matches autolens 2026.8.7.1), `--scope all` 68 files / 143 symbols / **0 missing-broken**, `--lint-idioms` clean (214 files), `--check-citations` 105 files / 413 citations / **0 missing, 0 warnings**, `--check-provenance` **0 errors** (49 pages). Contrast autogalaxy's 5 provenance errors — the two failures shared a red badge but not a cause. (2) `wiki_drift_issue` — pure fallout, `Artifact not found: wiki-drift-report`, because (1) died before writing it. (3) `wiki_currency_check_autogalaxy` — REAL drift, see the provenance prompt filed under draft/maintenance/autogalaxy_assistant/. +- artifacts-are-laptop-only: Actions artifact downloads are blocked from cloud/mobile sessions (egress policy 403s `productionresultssa2.blob.core.windows.net` on CONNECT) — this is what stopped the cloud session finishing the ingest. Both wiki drift reports were captured to `~/.pyauto-heart/release_20260807_wiki_drift/` while on the laptop. +- do-not: do NOT use the nightly driver for a manual release — AUTONOMY.md forbids converting a manual release into the scheduled-nightly exception. diff --git a/complete/index.md b/complete/index.md index 0f50c334..eb8ce119 100644 --- a/complete/index.md +++ b/complete/index.md @@ -6,7 +6,7 @@ Token-light navigation over the finished-work records (schema: only then grep a dated bucket. Curators: edit the band between the CURATED markers; everything below GENERATED is rebuilt. -1008 records across 7 buckets. +1010 records across 7 buckets. ## Highlights @@ -112,6 +112,8 @@ _(curate hard-won records here — survives regeneration.)_ - [reconcile-upstream-repo-mode](2026/08/reconcile-upstream-repo-mode.md) - [registry-integrity-check](2026/08/registry-integrity-check.md) - [regularization-jax-gradient-gaps](2026/08/regularization-jax-gradient-gaps.md) +- [release-drive-2026-08-03](2026/08/release-drive-2026-08-03.md) +- [release-drive-2026-08-07](2026/08/release-drive-2026-08-07.md) - [release-validation-tri-state](2026/08/release-validation-tri-state.md) - [resolve-border-relocator-hazard](2026/08/resolve-border-relocator-hazard.md) — Reconciled the likelihood hazard instrument after the border-relocator source fix. The resolved backend-diverg… - [resolve-curvature-floor-doc-drift](2026/08/resolve-curvature-floor-doc-drift.md) — Reconciled the curvature-floor documentation finding after PyAutoArray#444. The detector now requires both run… diff --git a/dashboard.md b/dashboard.md index 87e83866..cce502d3 100644 --- a/dashboard.md +++ b/dashboard.md @@ -11,7 +11,7 @@ Tasks only — the organism's health lives with the Heart (`/health`), not here. | [In flight](#in-flight) (`active/`) | 3 | | [Parked](#parked) (`parked.md`) | 1 | | [Planned](#planned) (`planned.md`) | 7 | -| [Backlog](#backlog) (`draft/`) | 137 | +| [Backlog](#backlog) (`draft/`) | 138 | Live on GitHub: [open issues](https://github.com/search?q=org%3APyAutoLabs+is%3Aissue+is%3Aopen&type=issues) · [open pull requests](https://github.com/search?q=org%3APyAutoLabs+is%3Apr+is%3Aopen&type=prs) @@ -34,6 +34,7 @@ Live on GitHub: [open issues](https://github.com/search?q=org%3APyAutoLabs+is%3A **Quick wins** (small enough, and safe enough to run unattended) +- [Has the falsified-by checkpoint stage gone rote after ten ships](draft/research/pyautobrain/has_the_falsified_by_checkpoint_stage_gone.md) — pyautobrain · small · safe · normal - [PyAutoFit CLI-noise batch: unclosed search.log handler + four small warning](draft/maintenance/pyautofit/cli_noise_pyautofit_batch.md) — pyautofit · small · safe · normal - [Tenant firewall: release_run.py carries an unlisted 'PyAutoLabs' instance fact](draft/bug/pyautoheart/tenant_firewall_release_run_instance_fact.md) — pyautoheart · small · safe · normal - [Silence the three autonerves-rooted CLI-noise sources (fits leak, pytest collection,](draft/maintenance/pyautonerves/cli_noise_autonerves_batch.md) — pyautonerves · small · safe · normal @@ -81,7 +82,7 @@ Scoped but not started; some are not yet prompt files. Full detail in [`planned. ## Backlog -**137** filed prompts, not started. Each section is sorted most-pickable first (priority, then size). +**138** filed prompts, not started. Each section is sorted most-pickable first (priority, then size).
bug — 38 @@ -162,7 +163,7 @@ Scoped but not started; some are not yet prompt files. Full detail in [`planned.
-research — 21 +research — 22 - [The `ell_comps` trapping was masked, not cleared — characterise it](draft/research/autolens_profiling/ell_comps_trapping_unmasked.md) — autolens_profiling · medium · supervised · high - [Optimize pixelized Prodigy settings on the laptop GPU](draft/research/autolens_workspace_developer/pixelized_prodigy_laptop_gpu_phase_2_settings.md) — autolens_workspace_developer · medium · human-required · high @@ -174,6 +175,7 @@ Scoped but not started; some are not yet prompt files. Full detail in [`planned. - [Delaunay-family JAX modules never hit the persistent compilation cache](draft/research/autoarray/delaunay_callback_persistent_cache_miss.md) — autoarray · medium · supervised · medium - [Quick-update plotting cost — minutes per update, and it is](draft/research/autolens/quick_update_plotting_cost.md) — autolens · medium · supervised · medium - [Use readthedocs or migrate to GitHub docs](draft/research/autobuild/git_docs.md) — autobuild · small · supervised · normal +- [Has the falsified-by checkpoint stage gone rote after ten ships](draft/research/pyautobrain/has_the_falsified_by_checkpoint_stage_gone.md) — pyautobrain · small · safe · normal - [Re-baseline the slacs0008 acceptance parity after the HAP-dedupe fix](draft/research/pyautoreduce/acceptance_noise_rebaseline.md) — pyautoreduce · small · supervised · normal - [Kernel-CDF bandwidth defaults — config-dependent quality, investigate adaptivity](draft/research/autoarray/rectangular_kernel_bandwidth_defaults.md) — autoarray · medium · supervised · normal - [We have lots of examples which profile how long JAX](draft/research/autolens_workspace_developer/jax_jit_profiling.md) — autolens_workspace_developer · medium · supervised · normal diff --git a/draft/research/pyautobrain/has_the_falsified_by_checkpoint_stage_gone.md b/draft/research/pyautobrain/has_the_falsified_by_checkpoint_stage_gone.md new file mode 100644 index 00000000..4cf989fb --- /dev/null +++ b/draft/research/pyautobrain/has_the_falsified_by_checkpoint_stage_gone.md @@ -0,0 +1,80 @@ +# Has the falsified-by checkpoint stage gone rote after ten ships + +Type: research +Target: PyAutoBrain +Repos: +- PyAutoBrain +Difficulty: small +Autonomy: safe +Priority: normal +Status: formalised + +## What this is + +The efficacy review that `docs/agent_failure_modes.md` §9 committed to when +mitigation 6 shipped: *"trial on the next ship series, review whether it went +rote after ~10 ships."* Spun out of PyAutoBrain#130 when that issue was closed +(2026-08-15) — it was the one §9 item still genuinely open, and nothing tracked +it. + +This is an investigation producing a written verdict from evidence. No code +change is committed up front; a fix may follow from the finding. + +## Background + +Mitigation 6 (PyAutoBrain#140, merged 2026-07-17, live) made the review faculty +lift **load-bearing empirical claims** out of a branch's commit messages into +the `ReviewSurface` as `claims to falsify` — the trigger vocabulary is `no-op`, +`byte-identical`, `does-not-affect`, `proven`, `behaviour-preserving`. `AGENTS.md` +step 2a then makes an unsupported one a FINDING of kind `unverified-claim`. + +Its design was deliberately reader-enforced rather than an author checklist, and +scoped to load-bearing phrasing only, precisely so it could not decay into the +"remember-to-run checklist" the campaign's own constraints ban. It targets the +A5/F3 failure class — confident-wrong effect-claims. + +## Why it needs reviewing + +The doc's constraint list bans checklists as a mechanism, and a routine +adversarial pass is the single mechanism most likely to decay into one. The +worry is explicit in the shipping comment: *"the one that needs care not to +become the banned checklist."* A stage that fires on every ship and is waved +through every time is worse than no stage, because it also carries false +assurance. + +## What to investigate + +Over the real ship history since 2026-07-17: + +1. **Firing rate** — on how many ships did `claims to falsify` populate at all, + versus come back empty? An always-empty surface means the vocabulary is too + narrow; an always-full one means it is too broad. +2. **Finding rate** — how many `unverified-claim` FINDINGS were actually raised, + and what happened to each? A stage that never produces a finding across ~10 + ships is either unnecessary or being rubber-stamped; distinguish those two. +3. **Were any load-bearing?** For each finding, did falsifying the claim change + the outcome — a correction, a held ship — or was it cosmetic? This is the + measure of whether the stage is earning its per-ship cost. +4. **Idle-phrasing exclusion** — is the load-bearing-only scoping holding, or has + the matcher started lifting incidental prose? Check for false positives of + the kind that trained bypass-by-default in the guard's first hour (the F5 + cost column). +5. **Rubber-stamping** — look for the signature: claims lifted, reviewed CLEAN, + no evidence cited in the review. That is the rote failure, and it looks + identical to a healthy pass unless you read what the reviewer actually did. + +## Deliverable + +A verdict with the numbers behind it, and one of: keep as-is, narrow/broaden the +trigger vocabulary, or retire the stage. If the finding is "it went rote", say +what would fire instead — per the campaign's own ranking, deleting the +possibility beats detecting it, and detecting beats reminding. + +## Method note + +Validate the instrument before trusting it: check that the review faculty's +claim-lifting still runs on a branch with known load-bearing claims before +concluding anything from a low firing rate. A null result that looks like a +finding (D1) is the exact failure this campaign catalogued. + +