Skip to content

Phase 6: benchmarks, hpc, output-surface skills, ledger retirement - #8

Merged
Jammy2211 merged 1 commit into
mainfrom
feature/phase-6-benchmarks-hpc
Aug 1, 2026
Merged

Phase 6: benchmarks, hpc, output-surface skills, ledger retirement#8
Jammy2211 merged 1 commit into
mainfrom
feature/phase-6-benchmarks-hpc

Conversation

@Jammy2211

Copy link
Copy Markdown
Contributor

Phase 6 — the close-out — of PyAutoLabs/PyAutoBrain#188.

  • Benchmarks: 4 cards total (medium MGE bulge-disk, hard multi-band pixelization, teacher workflow join the easy card); RESULTS.md regenerated from the empty runs/no score invented; calibration runs are the recorded post-birth follow-up.
  • hpc/ shipped galaxy-tuned (one galaxy per array task; parse_fit_args contract byte-identical to scripts/AGENTS.md), validated by execution on the bundled dataset (scaffold runs to the intended NotImplementedError); fixed a real trap where root .gitignore's output/ swallowed the batch keepers.
  • wiki/core complete: operations/hpc.md + hpc_infrastructure.md land (41 pages); the index now states what "complete" does and does not claim.
  • 27 skills: ag_to_notebook + ag_inspect_results_mcp (enum groups introspected from the wheel).
  • PENDING.md retired → ROADMAP.md (genuine backlog only: benchmark calibration, second dataset + mask_extra_galaxies.fits, two ungrounded skills, upstream fixes, leg-4 smoke, deferred paper//clone-profile); all 33 references updated and eight stale "not yet written" claims from earlier phases fixed in the sweep.

Newborn gate: legs 1–2 evidence in the commit (all wheel-venv checks clean, pytest 68, links 974/0); leg 3 = this PR's wiki-currency run; leg 4 (chat-surface smoke) recorded in ROADMAP.md as the human follow-up.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EVSbY4u1JLfqRMgXYtnN9v

Three new benchmark cards (medium MGE bulge-disk, hard multi-band
pixelization, teacher workflow; all rubrics parse and total 100);
RESULTS.md regenerated from the EMPTY runs dir — no run scored, no score
invented (calibration runs are the post-birth follow-up, autofit
precedent). hpc/ shipped galaxy-tuned with the parse_fit_args contract
preserved byte-for-byte and validated BY EXECUTION on the bundled
dataset (+ .gitignore negations so the batch output/error keepers
actually ship). operations/hpc.md + hpc_infrastructure.md complete the
core wiki (41 pages; index now says what 'complete' claims).
ag_to_notebook + ag_inspect_results_mcp -> 27 skills, 27/27 symlinks +
citation-map rows. PENDING.md retired -> ROADMAP.md (genuine backlog
only); all 33 references across 26 files updated, and the sweep fixed
eight stale not-yet-written claims left over from earlier phases.

Newborn gate legs 1-2 evidence: wheel-venv --scope all 46 files / 0
missing; --check-version clean; lint 122 files clean; citations 1059/0;
provenance 41 pages / 0 errors; literature validator clean; pytest 68
passed; link sweep 974/0; sparse-cone + banned-token sweeps clean.
Leg 3 = this PR's wiki-currency run; leg 4 recorded in ROADMAP.md.

Epic: PyAutoLabs/PyAutoBrain#188 (Phase 6)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EVSbY4u1JLfqRMgXYtnN9v
@Jammy2211
Jammy2211 merged commit a2af955 into main Aug 1, 2026
2 checks passed
@Jammy2211
Jammy2211 deleted the feature/phase-6-benchmarks-hpc branch August 1, 2026 18:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant