Structure-based prediction of protein–protein binding free energy (ΔG), using the PCANN model: an ESM2-conditioned residue-interface graph network. This repository ships the trained EXP-043 ensemble (25 cluster-aware cross-validation folds), the data needed to run and reproduce it, and the training/inference entry points.
- Python ≥ 3.12
uv(dependency manager)- ~7 GB free disk (≈1.5 GB downloaded structures + ≈0.1 GB checkpoints, plus ≈4.7 GB for the ESM2 embeddings you compute locally — see below; they are not shipped)
Two mutually-exclusive install profiles:
make install # uv sync --extra cpu --no-dev — macOS, or Linux without a GPU (default)
make install-cu128 # uv sync --extra cu128 --no-dev — Linux + NVIDIA (CUDA 12.8 wheels)For CUDA, the cu128 wheels bundle the CUDA 12.8 runtime — the host needs only a
recent NVIDIA driver, not a system CUDA toolkit. To target a different CUDA version,
swap cu128 → cu126/cu129 in pyproject.toml. (For torch 2.8.0, PyG ships
extension wheels only for cpu, cu126, cu128, cu129.)
The artifacts ship as two separate archives:
| Component | Contents | Hosted |
|---|---|---|
checkpoints |
the 25 EXP-043 models | GitHub Release asset |
pdb |
structures (large) | out-of-band — download link distributed separately |
Download the archives (see docs/data-prep.md for the exact links),
drop them in the repo root, then:
make data # unpacks any downloaded pcann_v2*.tar.gz in placeThis populates logs/multiruns/EXP-043/ and data/raw/ppb-affinity/pdb/. You only
need the components you'll use: checkpoints for inference, plus pdb for the
test/train UIDs.
ESM2 embeddings are not shipped. The model reads per-residue
esm2_t33_650M_UR50Dembeddings fromdata/raw/ppb-affinity/esm/, and you must compute them before any inference or training — see ESM embeddings below. They are large (~4.7 GB) and fully reproducible from the shippedpdbarchive.
The model needs one .pt embedding per structure at data/raw/ppb-affinity/esm/<uid>.pt.
Generate them from the downloaded PDBs:
git clone https://github.com/facebookresearch/esm
export ESM_MODEL_DIR=$PWD/esm # weights esm2_t33_650M_UR50D auto-download on first run
make esm # embeds every data/raw/ppb-affinity/pdb/*.pdbmake esm runs on the GPU automatically when one is visible, and falls back to CPU
otherwise. To embed new PDBs later, re-run it (existing .pt files are skipped) — see
docs/data-prep.md.
make infer # 25-fold ensemble on PPI-103 -> predictions_PPI-103.csvor directly, for either test set — PPI-103 (data/test/PPI-103.csv) or
FAB-75 (data/test/FAB-75.csv):
uv run python src/predict.py \
--checkpoints-dir logs/multiruns/EXP-043 \
--test-csv data/test/PPI-103.csv \
--output predictions_PPI-103.csv # add --folds 0,1 for a fast subsetExpected ensemble headline (CPU): PPI-103 (N=103) MAE ≈ 1.40, FAB-75 (N=75) MAE ≈ 1.40.
make train # 25-fold multirun, seed 42, CPUEXP-043 ran on CPU (trainer.accelerator: cpu); a GPU only speeds it up. Exact
numeric reproduction of the headline is a CPU statement.
make test # install + unpack + 2-fold inference + 1-epoch training smoke testdocs/data-prep.md— data layout, ESM extractiondocs/inference.md—predict.pycontractdocs/training.md— reproducing EXP-043
Released under the MIT License.