Repository navigation
feat(io): keep mmCIF label residue ids and entity sequence tables - #30
Merged
Merged
Conversation
Move the materialize precondition into a match guard. When the guard fails, the action falls through to the existing no-op arm, as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Atoms of one residue are contiguous in coordinate files, so an atom whose residue fields match the previous atom's reuses its Arc<AtomResidue> without building and hashing a cache key. The cache still shares non-contiguous atoms of one residue. On 487 PDB mmCIF files (5.2M atoms), instructions retired for reading them drop from 86.5G to 75.9G. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Store label_asym_id, label_entity_id and label_seq_id on AtomResidue when reading mmCIF and bCIF. AF3-style models number residues by the label scheme, and label_seq_id is what links atoms to _entity_poly_seq. Previously label_asym_id survived only in assembly membership and label_entity_id was not read. bCIF encoders store entity ids either as strings or as integers; both are read. The fields are None for PDB and other formats, and default to None when older sessions are deserialized. Topology grouping of models still uses auth identity only, so grouped models keep the first model's labels. Label ids describe the file as loaded: edits do not renumber them, and build_mutant now clones the template residue so they survive mutation. Label strings are shared with the previous atom when equal, so parsing does not allocate them per atom. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Parse _entity, _entity_poly and _entity_poly_seq from mmCIF and bCIF into ObjectMolecule::entities. The full polymer sequence includes unresolved residues and is indexed by num - 1, so label_seq_id on a residue points straight into it. - Rows may come in any order, but num must run 1..N without gaps, as the PDBx/mmCIF dictionary requires; a skipped num is a parse error. - num must be at most 1_000_000, which keeps a corrupt value from allocating a huge sequence. - Microheterogeneity keeps the first monomer in the sequence and the rest in EntityPolymer::alternatives. - Entity errors are reported only for blocks with atoms, as for assemblies. Entity and polymer types map the PDBx/mmCIF enumerations and keep unknown values verbatim. Other formats leave entities empty. Entities describe the file as loaded; edits do not update them. Sessions store the new field, so the PRS format version is 5. Version 4 files without entities still load. With label ids and entity tables, reading 487 PDB mmCIF files costs 80.7G instructions retired, against 86.5G before this series. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The APT cache hit restored only metadata and zero package archives, so Fontconfig was missing when core tests compiled Slint. Install the same Linux dependency list with apt-get and check Fontconfig before Cargo runs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
mmCIF/bCIF readers now keep label residue ids and entity tables, which AF3-style models and
_entity_poly_seq-based tools rely on:AtomResiduegainslabel_asym_id,label_entity_id,label_seq_id(Nonefor other formats)._entity,_entity_poly,_entity_poly_seqare read intoObjectMolecule::entities. The polymer sequence includes unresolved residues and is indexed bynum - 1, sopolymer.monomer(label_seq_id)finds a residue's monomer. A skippednumis a parse error.PRS_FORMAT_VERSIONgoes 4 → 5 for the newObjectMoleculefield, as withassembly; version 4 sessions still load.A separate
perf(io)commit reuses the previous atom's residue when building molecules, so reading mmCIF is ~7% faster thanmainoverall. The first commit fixes aclippy::collapsible_matcherror present onmain.Testing
num, typed bCIF columns.