Skip to content

About

Near real-time data on the human neutralizing antibody landscape to influenza virus as of the summer of 2026 to inform vaccine-strain selection

Resources

Stars

2 stars

Watchers

1 watching

Forks

Repository files navigation

Near real-time data on the human neutralizing antibody landscape to influenza virus as of the summer of 2026 to inform vaccine-strain selection

This study led by Caroline Kikawa, Andrew Butler, John Huddleston, and Jesse Bloom uses sequencing-based neutralization assays to measure titers to influenza viruses with HAs from human seasonal H3N2 and H1N1 viruses representative of those circulating in mid-2026 against human sera collected in early to mid 2026.

The paper describing this study is Kikawa, Butler, et al (2026).

Here is an interactive summary of key results.

Quick links to key results

To jump directly to the key results:

Past similar studies

This study is the third in a series of studies by the Bloom lab using sequencing-based neutralization assays to provide near real-time data on the human neutralizing antibody landscape to influenza to inform vaccine strain selection. For past studies and their GitHub repositories, see:

Experimental assay

This repository analyzes data generated by sequencing-based neutralization assays. For details on these assays, see:

Structure of this repository

This repository contains a snakemake pipeline that uses seqneut-pipeline for the core analysis and contains additional rules for visualization, QC, and analysis.

Repository layout

flu-seqneut-2026/
├── Snakefile          # workflow: config, includes, `rule all`
├── config.yml         # all pipeline configuration
├── rules/             # rules for analyses outside seqneut-pipeline
├── scripts/           # scripts those rules call
├── envs/              # conda environments for individual rules
├── data/              # all input data
├── results/           # all output (key files tracked in git)
├── auspice/           # Nextstrain JSONs for the community build
└── non-pipeline_analyses/  # one-off analyses, not run by Snakefile

The git submodules are described under Submodules below. run_Hutch_cluster.bash submits the pipeline to the Fred Hutch SLURM cluster.

Input data

All input data for the pipeline are in ./data/. For the most part, these data should be self explanatory. Here are some key details:

Viral Library

Details on the viral library are in ./data/viral_libraries. There are CSVs defining the designed library and the actual library. The designed library is the set of strains initially designed for this project; the actual library is the set of strains used in the actual experiments (some strains were dropped due to low titers or other QC issues). The CSVs provide information about each barcoded HA used in the experiments. A few notes:

  • In some cases, the same HA may have multiple barcodes to provide internal replicates
  • For each strain, we take the HA ectodomain from a natural seasonal influenza strain and that ectodomain sequence is provided in the CSVs. The sequences in the CSVs do not include the signal peptide, endodomain, and cytoplasmic tail since those are constant across all strains as described in Loes et al (2024), Journal of Virology and Kikawa et al (2026), eLife. For H1 HAs, a "DTL" must be prepended to these sequences to get the full ectodomain since the barcoded constructs provide the first three ectodomain amino-acids from the lab-adapted A/WSN/1933 (H1N1) HA.
  • The columns each viral library CSV must have, and the values they may hold, are specified under viral_library_validations in config.yml.
  • derived_haplotype is the subclade plus the amino-acid mutations separating the strain from its subclade founder (for the ectodomain only), written as subclade:mut,mut (eg, K:F192V,HA2_R32K). HA1 mutations come first and are bare; HA2 mutations follow and are prefixed HA2_.
  • collection_date is a float (eg, 2025.5); in many cases it refers to the date that HA1 haplotype was last identified rather than the actual collection date of the particular named strain

Sera

Details on the sera used for the study are in data/sera_metadata/; note that titers may not have been measured for all sera listed here. There is a CSV for sera from each cohort, with the following required columns:

  • bloom_lab_id: unique identifier for each serum sample; must be non-null and unique across all files
  • cohort: cohort identifier; must be non-null
  • species: species of serum donor (e.g., "human"); must be non-null
  • age: age of donor; can be numeric (e.g., "45"), a range with or without 'y' suffix (e.g., "10-19y", "18-29"), or open-ended (e.g., "75+")
  • sex: sex of donor; accepts various formats ("M"/"F", "male"/"female", "Male"/"Female") which are normalized during aggregation
  • collection_date: date of serum collection in "YYYY-MM" format (month precision)

A CSV may carry further columns describing its cohort (eg, the vaccination status); multicohorts in config.yml can use these columns to further group sera into additional cohorts.

These per-cohort CSVs are aggregated and validated into results/sera_metadata/all_sera_metadata.csv, where bloom_lab_id is renamed to serum and sex and age are put in standard forms.

Plate configuration and FASTQ files

The configuration for the sequencing-based neutralization assay plates is specified in CSVs in ./data/plates and ./data/miscellaneous_plates. These CSVs point to the FASTQ files with the barcode sequencing results, which are held on the Fred Hutch computing cluster as they are too large for this repository. The barcode counts extracted from them are in ./results/ and are tracked here.

Configuration

All pipeline configuration parameters are specified in config.yml. This YAML should be self explanatory. Many of the configuration parameters relate to seqneut-pipeline; see seqneut-pipeline/README.md for detailed documentation on those parameters.

Results

All output generated by the pipeline goes in ./results/, although only some results are tracked in the GitHub repo (see .gitignore). The exception is that the Nextstrain tree JSONs are placed in ./auspice/, where they can be viewed as Nextstrain Community Builds. These results should be largely self-explanatory, here are the key ones to look at if you want to visualize the results:

Snakemake pipeline

The pipeline is specified in Snakefile, which includes:

Submodules

This repository uses the following git submodules:

Running the Pipeline

Clone the repository with its submodules:

git clone --recurse-submodules https://github.com/jbloomlab/flu-seqneut-2026.git

An existing clone needs git submodule update --init --recursive, as the pipeline will not run with empty submodule directories.

Then build the seqneut-pipeline conda environment in seqneut-pipeline/environment.yml:

# Create conda environment (only needs to be done once)
conda env create -f seqneut-pipeline/environment.yml

# Activate environment
conda activate seqneut-pipeline

Then to run the pipeline locally use:

snakemake -j <num_cores> --software-deployment-method conda

To run the pipeline on the Fred Hutch cluster using slurm, run:

sbatch run_Hutch_cluster.bash

Documentation on GitHub Pages

As described in seqneut-pipeline/README.md, the pipeline generates HTML documentation in ./results/docs (not tracked in git) that can be pushed to a gh-pages branch for publishing on GitHub Pages with:

./seqneut-pipeline/publish_docs_gh-pages.sh

This documentation is then displayed on GitHub Pages via the gh-pages branch at https://jbloomlab.github.io/flu-seqneut-2026.

The command to publish to the gh-pages branch must be run manually, and GitHub must be manually configured to display the Pages from this branch.

Non-pipeline analyses

The non-pipeline_analyses/ subdirectory contains one-off analyses that are not run by the Snakefile and are separate from the main pipeline. See the README in each of those subdirectories for details, and note some of these non-pipeline analyses may be obsolete.

License

MIT License

About

Near real-time data on the human neutralizing antibody landscape to influenza virus as of the summer of 2026 to inform vaccine-strain selection

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages