Skip to content
norsagePublic

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

affex

Structure-based prediction of protein–protein binding free energy (ΔG), using the PCANN model: an ESM2-conditioned residue-interface graph network. This repository ships the trained EXP-043 ensemble (25 cluster-aware cross-validation folds), the data needed to run and reproduce it, and the training/inference entry points.

Requirements

  • Python ≥ 3.12
  • uv (dependency manager)
  • ~7 GB free disk (≈1.5 GB downloaded structures + ≈0.1 GB checkpoints, plus ≈4.7 GB for the ESM2 embeddings you compute locally — see below; they are not shipped)

Install

Two mutually-exclusive install profiles:

make install            # uv sync --extra cpu --no-dev    — macOS, or Linux without a GPU (default)
make install-cu128      # uv sync --extra cu128 --no-dev  — Linux + NVIDIA (CUDA 12.8 wheels)

For CUDA, the cu128 wheels bundle the CUDA 12.8 runtime — the host needs only a recent NVIDIA driver, not a system CUDA toolkit. To target a different CUDA version, swap cu128 → cu126/cu129 in pyproject.toml. (For torch 2.8.0, PyG ships extension wheels only for cpu, cu126, cu128, cu129.)

Get data & checkpoints

The artifacts ship as two separate archives:

Component Contents Hosted
checkpoints the 25 EXP-043 models GitHub Release asset
pdb structures (large) out-of-band — download link distributed separately

Download the archives (see docs/data-prep.md for the exact links), drop them in the repo root, then:

make data        # unpacks any downloaded pcann_v2*.tar.gz in place

This populates logs/multiruns/EXP-043/ and data/raw/ppb-affinity/pdb/. You only need the components you'll use: checkpoints for inference, plus pdb for the test/train UIDs.

ESM2 embeddings are not shipped. The model reads per-residue esm2_t33_650M_UR50D embeddings from data/raw/ppb-affinity/esm/, and you must compute them before any inference or training — see ESM embeddings below. They are large (~4.7 GB) and fully reproducible from the shipped pdb archive.

ESM embeddings (required before any task)

The model needs one .pt embedding per structure at data/raw/ppb-affinity/esm/<uid>.pt. Generate them from the downloaded PDBs:

git clone https://github.com/facebookresearch/esm
export ESM_MODEL_DIR=$PWD/esm     # weights esm2_t33_650M_UR50D auto-download on first run
make esm                          # embeds every data/raw/ppb-affinity/pdb/*.pdb

make esm runs on the GPU automatically when one is visible, and falls back to CPU otherwise. To embed new PDBs later, re-run it (existing .pt files are skipped) — see docs/data-prep.md.

Run inference

make infer       # 25-fold ensemble on PPI-103 -> predictions_PPI-103.csv

or directly, for either test set — PPI-103 (data/test/PPI-103.csv) or FAB-75 (data/test/FAB-75.csv):

uv run python src/predict.py \
  --checkpoints-dir logs/multiruns/EXP-043 \
  --test-csv data/test/PPI-103.csv \
  --output predictions_PPI-103.csv          # add --folds 0,1 for a fast subset

Expected ensemble headline (CPU): PPI-103 (N=103) MAE ≈ 1.40, FAB-75 (N=75) MAE ≈ 1.40.

Reproduce EXP-043 training

make train       # 25-fold multirun, seed 42, CPU

EXP-043 ran on CPU (trainer.accelerator: cpu); a GPU only speeds it up. Exact numeric reproduction of the headline is a CPU statement.

Test

make test        # install + unpack + 2-fold inference + 1-epoch training smoke test

Docs

Citation / License

Released under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages