Near real-time data on the human neutralizing antibody landscape to influenza virus as of the summer of 2026 to inform vaccine-strain selection
This study led by Caroline Kikawa, Andrew Butler, John Huddleston, and Jesse Bloom uses sequencing-based neutralization assays to measure titers to influenza viruses with HAs from human seasonal H3N2 and H1N1 viruses representative of those circulating in mid-2026 against human sera collected in early to mid 2026.
The paper describing this study is Kikawa, Butler, et al (2026).
Here is an interactive summary of key results.
To jump directly to the key results:
-
results/final_titer_data has the final titer data and information on the viruses and sera for which these final data were obtained. Specifically:
- For human sera:
- results/final_titer_data/human_titers.csv: QC-ed set of titers for each virus/serum pair (keeping only viruses for which titers measured against most sera, and sera for which titers measured against most viruses).
- results/final_titer_data/human_sera.csv: detailed information about the sera for which these titers were measured.
- results/final_titer_data/human_sera_multicohort.csv: the same sera assigned to additional finer-grained cohorts (named by
multicohortsin config.yml), so each serum may appear in several rows. - results/final_titer_data/human_sera_summary.csv: table summarizing the sera in each sera set.
- results/final_titer_data/human_viruses.csv: detailed information about the viruses for which these titers were measured.
- results/final_titer_data/human_titers_summarized_by_virus.csv: summary statistics about the titers against each virus.
- For human sera:
-
The interactive documentation holds a detailed set of plots showing all of the fitted curves and per-plate and per-serum analysis and QC as well as plots summarizing the results. See the report summary for a narrated overview of the summarized results.
-
Interactive Nextstrain protein trees that can be colored by subclade, median titer, etc and also show individual titers in the Measurements panel can be viewed at:
This study is the third in a series of studies by the Bloom lab using sequencing-based neutralization assays to provide near real-time data on the human neutralizing antibody landscape to influenza to inform vaccine strain selection. For past studies and their GitHub repositories, see:
- flu-seqneut-2025to2026: Kikawa et al (2026) and https://github.com/jbloomlab/flu-seqneut-2025to2026
- flu-seqneut-2025: Kikawa et al (2025) and https://github.com/jbloomlab/flu-seqneut-2025
This repository analyzes data generated by sequencing-based neutralization assays. For details on these assays, see:
- Loes et al (2024), Journal of Virology
- Kikawa et al (2026), eLife
- Detailed experimental protocol on protocols.io
This repository contains a snakemake pipeline that uses seqneut-pipeline for the core analysis and contains additional rules for visualization, QC, and analysis.
flu-seqneut-2026/
├── Snakefile # workflow: config, includes, `rule all`
├── config.yml # all pipeline configuration
├── rules/ # rules for analyses outside seqneut-pipeline
├── scripts/ # scripts those rules call
├── envs/ # conda environments for individual rules
├── data/ # all input data
├── results/ # all output (key files tracked in git)
├── auspice/ # Nextstrain JSONs for the community build
└── non-pipeline_analyses/ # one-off analyses, not run by Snakefile
The git submodules are described under Submodules below.
run_Hutch_cluster.bash submits the pipeline to the Fred Hutch SLURM cluster.
All input data for the pipeline are in ./data/. For the most part, these data should be self explanatory. Here are some key details:
Details on the viral library are in ./data/viral_libraries. There are CSVs defining the designed library and the actual library. The designed library is the set of strains initially designed for this project; the actual library is the set of strains used in the actual experiments (some strains were dropped due to low titers or other QC issues). The CSVs provide information about each barcoded HA used in the experiments. A few notes:
- In some cases, the same HA may have multiple barcodes to provide internal replicates
- For each strain, we take the HA ectodomain from a natural seasonal influenza strain and that ectodomain sequence is provided in the CSVs. The sequences in the CSVs do not include the signal peptide, endodomain, and cytoplasmic tail since those are constant across all strains as described in Loes et al (2024), Journal of Virology and Kikawa et al (2026), eLife. For H1 HAs, a "DTL" must be prepended to these sequences to get the full ectodomain since the barcoded constructs provide the first three ectodomain amino-acids from the lab-adapted A/WSN/1933 (H1N1) HA.
- The columns each viral library CSV must have, and the values they may hold, are specified under
viral_library_validationsin config.yml. - derived_haplotype is the subclade plus the amino-acid mutations separating the strain from its subclade founder (for the ectodomain only), written as
subclade:mut,mut(eg,K:F192V,HA2_R32K). HA1 mutations come first and are bare; HA2 mutations follow and are prefixedHA2_. - collection_date is a float (eg, 2025.5); in many cases it refers to the date that HA1 haplotype was last identified rather than the actual collection date of the particular named strain
Details on the sera used for the study are in data/sera_metadata/; note that titers may not have been measured for all sera listed here. There is a CSV for sera from each cohort, with the following required columns:
- bloom_lab_id: unique identifier for each serum sample; must be non-null and unique across all files
- cohort: cohort identifier; must be non-null
- species: species of serum donor (e.g., "human"); must be non-null
- age: age of donor; can be numeric (e.g., "45"), a range with or without 'y' suffix (e.g., "10-19y", "18-29"), or open-ended (e.g., "75+")
- sex: sex of donor; accepts various formats ("M"/"F", "male"/"female", "Male"/"Female") which are normalized during aggregation
- collection_date: date of serum collection in "YYYY-MM" format (month precision)
A CSV may carry further columns describing its cohort (eg, the vaccination status); multicohorts in config.yml can use these columns to further group sera into additional cohorts.
These per-cohort CSVs are aggregated and validated into results/sera_metadata/all_sera_metadata.csv, where bloom_lab_id is renamed to serum and sex and age are put in standard forms.
The configuration for the sequencing-based neutralization assay plates is specified in CSVs in ./data/plates and ./data/miscellaneous_plates. These CSVs point to the FASTQ files with the barcode sequencing results, which are held on the Fred Hutch computing cluster as they are too large for this repository. The barcode counts extracted from them are in ./results/ and are tracked here.
All pipeline configuration parameters are specified in config.yml.
This YAML should be self explanatory.
Many of the configuration parameters relate to seqneut-pipeline; see seqneut-pipeline/README.md for detailed documentation on those parameters.
All output generated by the pipeline goes in ./results/, although only some results are tracked in the GitHub repo (see .gitignore). The exception is that the Nextstrain tree JSONs are placed in ./auspice/, where they can be viewed as Nextstrain Community Builds. These results should be largely self-explanatory, here are the key ones to look at if you want to visualize the results:
- Final processed and QC-ed titer data are ./results/final_titer_data as described in more detail above in the Quick links to key results section.
- Titers aggregated across all plates, before the subsetting to well-measured viruses and sera, are in ./results/aggregated_titers.
- Nextstrain trees are in ./auspice, they can be viewed as Nextstrain Community Builds at the links provided in the Quick links to key results section.
- Interactive views of protein structures are in ./results/prot-struct-viz; each page is a single self-contained HTML file, and only the validation report written beside it is tracked in the repo.
- Interactive summary of results and pipeline QC are written to
results/docsand are not tracked in the repo, but can be viewed interactively at https://jbloomlab.github.io/flu-seqneut-2026 if you follow the steps described in the Documentation on GitHub Pages section.
The pipeline is specified in Snakefile, which includes:
- seqneut-pipeline/seqneut-pipeline.smk: the core analysis workflow provided by the seqneut-pipeline submodule.
- rules/library_qc.smk: QC of the library, including sequencing of the library pool to balance the strain titers and single-well infections with each strain to confirm the absence of contaminanation.
- rules/analyze_titers.smk: analysis of the titers to generate plots and QC-ed aggregated results.
- rules/trees.smk: builds Nextstrain trees showing the library and results in a phylogenetic context.
- rules/prot-struct-viz.smk: renders interactive views of protein structures with prot-struct-viz, one page per directory of data/prot-struct-viz_config.
- rules/reports.smk: renders the hand-written narrative reports in data/reports to HTML pages of the documentation, and checks the links they make against the built site.
This repository uses the following git submodules:
- seqneut-pipeline for the core analysis workflow.
- nextstrain-prot-titers-tree for building interactive Nextstrain protein trees with titer data.
- bloomlab-coding-standards for the lab coding standards followed here.
Clone the repository with its submodules:
git clone --recurse-submodules https://github.com/jbloomlab/flu-seqneut-2026.gitAn existing clone needs git submodule update --init --recursive, as the pipeline will not run with empty submodule directories.
Then build the seqneut-pipeline conda environment in seqneut-pipeline/environment.yml:
# Create conda environment (only needs to be done once)
conda env create -f seqneut-pipeline/environment.yml
# Activate environment
conda activate seqneut-pipelineThen to run the pipeline locally use:
snakemake -j <num_cores> --software-deployment-method condaTo run the pipeline on the Fred Hutch cluster using slurm, run:
sbatch run_Hutch_cluster.bashAs described in seqneut-pipeline/README.md, the pipeline generates HTML documentation in ./results/docs (not tracked in git) that can be pushed to a gh-pages branch for publishing on GitHub Pages with:
./seqneut-pipeline/publish_docs_gh-pages.shThis documentation is then displayed on GitHub Pages via the gh-pages branch at https://jbloomlab.github.io/flu-seqneut-2026.
The command to publish to the gh-pages branch must be run manually, and GitHub must be manually configured to display the Pages from this branch.
The non-pipeline_analyses/ subdirectory contains one-off analyses that are not run by the Snakefile and are separate from the main pipeline. See the README in each of those subdirectories for details, and note some of these non-pipeline analyses may be obsolete.