Skip to content

feat(io): keep mmCIF label residue ids and entity sequence tables - #30

Merged
zmactep merged 5 commits into
zmactep:mainfrom
norsage:feat/io-label-fields
Oct 9, 2026
Merged

zmactep merged 5 commits into
zmactep:mainfrom
norsage:feat/io-label-fields

Conversation

@norsage

@norsage norsage commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Summary

mmCIF/bCIF readers now keep label residue ids and entity tables, which AF3-style models and _entity_poly_seq-based tools rely on:

  • AtomResidue gains label_asym_id, label_entity_id, label_seq_id (None for other formats).
  • _entity, _entity_poly, _entity_poly_seq are read into ObjectMolecule::entities. The polymer sequence includes unresolved residues and is indexed by num - 1, so polymer.monomer(label_seq_id) finds a residue's monomer. A skipped num is a parse error.
  • Both describe the file as loaded; edits do not update them.

PRS_FORMAT_VERSION goes 4 → 5 for the new ObjectMolecule field, as with assembly; version 4 sessions still load.

A separate perf(io) commit reuses the previous atom's residue when building molecules, so reading mmCIF is ~7% faster than main overall. The first commit fixes a clippy::collapsible_match error present on main.

Testing

  • README verification checks pass (fmt, clippy, workspace tests, python and wasm checks).
  • New unit tests for label/auth mismatch, missing label columns, older sessions, mutation, microheterogeneity, invalid num, typed bCIF columns.
  • Entities match gemmi on a test set of 487 structures

norsage and others added 5 commits October 8, 2026 11:35
Move the materialize precondition into a match guard. When the guard
fails, the action falls through to the existing no-op arm, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Atoms of one residue are contiguous in coordinate files, so an atom whose
residue fields match the previous atom's reuses its Arc<AtomResidue>
without building and hashing a cache key. The cache still shares
non-contiguous atoms of one residue.

On 487 PDB mmCIF files (5.2M atoms), instructions retired for reading
them drop from 86.5G to 75.9G.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Store label_asym_id, label_entity_id and label_seq_id on AtomResidue
when reading mmCIF and bCIF. AF3-style models number residues by the
label scheme, and label_seq_id is what links atoms to _entity_poly_seq.
Previously label_asym_id survived only in assembly membership and
label_entity_id was not read. bCIF encoders store entity ids either as
strings or as integers; both are read.

The fields are None for PDB and other formats, and default to None when
older sessions are deserialized. Topology grouping of models still uses
auth identity only, so grouped models keep the first model's labels.
Label ids describe the file as loaded: edits do not renumber them, and
build_mutant now clones the template residue so they survive mutation.

Label strings are shared with the previous atom when equal, so parsing
does not allocate them per atom.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Parse _entity, _entity_poly and _entity_poly_seq from mmCIF and bCIF into
ObjectMolecule::entities. The full polymer sequence includes unresolved
residues and is indexed by num - 1, so label_seq_id on a residue points
straight into it.

- Rows may come in any order, but num must run 1..N without gaps, as the
  PDBx/mmCIF dictionary requires; a skipped num is a parse error.
- num must be at most 1_000_000, which keeps a corrupt value from
  allocating a huge sequence.
- Microheterogeneity keeps the first monomer in the sequence and the rest
  in EntityPolymer::alternatives.
- Entity errors are reported only for blocks with atoms, as for assemblies.

Entity and polymer types map the PDBx/mmCIF enumerations and keep unknown
values verbatim. Other formats leave entities empty. Entities describe the
file as loaded; edits do not update them.

Sessions store the new field, so the PRS format version is 5. Version 4
files without entities still load.

With label ids and entity tables, reading 487 PDB mmCIF files costs 80.7G
instructions retired, against 86.5G before this series.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The APT cache hit restored only metadata and zero package archives, so
Fontconfig was missing when core tests compiled Slint. Install the same
Linux dependency list with apt-get and check Fontconfig before Cargo runs.
@zmactep
zmactep merged commit f253c36 into zmactep:main Oct 9, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants