Skip to content

test(io): cover interrupted parser cache writes - #98

Merged
m1c0l merged 5 commits into
RosettaCommons:release/atomworks-3-0from
hwendler:fix/atomic-parse-cache-writes
Oct 1, 2026
Merged

m1c0l merged 5 commits into
RosettaCommons:release/atomworks-3-0from
hwendler:fix/atomic-parse-cache-writes

Conversation

@hwendler

@hwendler hwendler commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Retain @hwendler's interrupted-cache-write regression for AtomWorks 3.0.

The release parser already implements atomic cache publication: it serializes to a unique temporary file, explicitly selects compression, replaces the destination only after serialization succeeds, and cleans up on interruption. The original implementation change and unrelated documentation formatting are therefore removed from this PR's diff.

The remaining change is one test in tests/io/components/test_cache_atomic.py. It writes partial bytes and raises KeyboardInterrupt during serialization through the public parse(..., config=ParseConfig(...)) API, then checks that neither a cache entry nor any temporary file remains. It uses the checked-in 2HHB fixture and requires no external structure download. This complements the existing tests for failure during final file replacement.

Validation

  • Five focused cases passed: the new interruption regression, two existing buffer/cache round trips, cache-key disambiguation, and cache-directory defaults.
  • Ruff check and format check passed for the new test file.
  • Release-relative diff contains only the new test file.

Contributor commits are preserved; no force push or history rewrite. Target: release/atomworks-3-0.

Cache entries were written straight to their target path. A process that
died while serialising - a worker hitting a wall-clock limit or being
preempted, which is routine when the cache is populated from a batch
scheduler - left a truncated file behind that subsequent runs accepted as a
valid cache entry, so the failure surfaced later as an unrelated parse error.

Entries are now written to a temporary file that is moved into place once
complete. The temporary name carries host and process id so that several
workers sharing a cache directory, possibly on a network filesystem, cannot
overwrite each other's partial writes. An existing entry is not normally
rewritten, but two workers can pass that check at the same time and both
proceed, so the move has to tolerate an occupied destination: Path.replace
does, whereas Path.rename raises on Windows in that case.

Compression is now passed explicitly. pandas infers it from the file name,
and the temporary name does not carry the suffix that the destination has,
so without this the entries would silently be stored uncompressed. It is
derived from the destination, which keeps the stored format unchanged.

Tests cover an interrupted write leaving neither a cache entry nor a
temporary file, a destination created concurrently, and the stored format
still being gzip compressed.
Running `make format` reformats this file: the html_js_files entry uses five
spaces of indentation and single quotes, and the file lacks a trailing
newline. The change is whitespace and quoting only.

Kept as a separate commit so that the accompanying cache fix stays confined
to its own scope.
@nscorley

Copy link
Copy Markdown
Collaborator

Thanks for this fix! Just kicked-off the tests - once those pass likely good to merge

@nscorley nscorley left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you very much for catching this problem and putting up this PR! A few relatively minor comments, then would be excited to merge!

Comment thread src/atomworks/io/parser.py Outdated
Comment thread tests/io/components/test_caching.py Outdated
Comment thread src/atomworks/io/parser.py Outdated
Comment thread tests/io/components/test_caching.py Outdated
Comment thread tests/io/components/test_caching.py Outdated
hwendler and others added 2 commits August 18, 2026 21:30
Co-authored-by: Nathaniel Corley <nscorley@gmail.com>
Handle every suffix utils/compression recognises (.gz, .gzip, .zst) instead of
.gz alone, so an entry is never written uncompressed under a compressed name.
Note pandas infers .gz and .zst but not .gzip.

Drop two tests that add no coverage: the gzip-format test, since read_pickle
raises BadGzipFile on an uncompressed .pkl.gz and the end-to-end cache test
already round-trips; and the concurrent-destination test, which only
discriminates on Windows while CI runs ubuntu-latest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@hwendler
hwendler requested a review from nscorley August 18, 2026 19:58
@m1c0l

m1c0l commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

The interrupted-cache-write problem reported and fixed here by @hwendler is also addressed independently in the AtomWorks 3.0 parser on release/atomworks-3-0 (replacement writer).

I validated that replacement implementation against this PR's failure scenario:

  • Interrupting serialization after writing partial bytes publishes no cache entry and removes the temporary file.
  • An interrupted replacement preserves the previous complete, readable entry.
  • Actual SIGKILL during serialization never publishes a partial target, both with and without an existing entry. A subsequent write succeeds.
  • A real 2HHB parse/cache round trip succeeds, including coordinates, annotations and bonds; the cache hit is verified without invoking the CIF loader.

All three focused regression tests passed (the SIGKILL test covers both destination states). As expected, SIGKILL can leave an orphan temporary file, but the published cache remains valid or absent.

Credit to @hwendler for identifying the production failure and providing the interrupted-write regression. This validates the 3.0 replacement; it does not imply the fix is already on production or that this PR was merged.

@m1c0l m1c0l changed the title Fix/atomic parse cache writes test(io): cover interrupted parser cache writes Oct 1, 2026
@m1c0l
m1c0l changed the base branch from production to release/atomworks-3-0 October 1, 2026 19:15
@m1c0l
m1c0l merged commit 5e15154 into RosettaCommons:release/atomworks-3-0 Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants