Skip to content

Repository files navigation

Mytochondria

"It's the powerhouse of the cell." This one's mine.

Maintainers of research software are short on time, not on care. Mytochondria spends machine time on the part of their job that scales worst: reading the numerical core of a package line by line, reproducing every suspicion by execution, and turning what survives into a patch with a test that fails without it. What reaches a maintainer should be something they can merge or decline in a few minutes, not another report to triage.

Research software is unusually hard code: mathematically intricate, evolving alongside the science it serves, maintained for decades, and usually built by small teams under tight grant budgets. The packages read here are careful, well-engineered work. The premise is not that anyone was careless — it is that no team can read an entire codebase at the depth every line deserves, and that patient machine-assisted reading is a genuine complement to human testing. Nothing here carries any authority: every finding is an offer, the maintainers decide, and a "no" is a complete answer.

How it works

1. Survey. Harvest five years of research articles from six high-impact journals (Nature, Science, PNAS, NEJM, The Lancet, Cell) via the Europe PMC API, parse every openly readable full text, and extract which open-source packages each paper used, with a quotable evidence sentence and version where stated. Papers that could not be read are listed individually rather than silently dropped.

2. Target. For a package being read, re-mine the papers that used it to determine which parts they ran — commands, options, versions — so review effort lands on the code paths that carry published numbers.

3. Read and verify. Read those components adversarially, then verify: compile the routine under test verbatim against high-precision arithmetic, port it faithfully to Python and test against ground truth, or simulate the full pipeline. Only then does a finding get written up, with file:line evidence, a runnable harness, and a patch where the fix is crisp.

4. Ask whether it reaches real data. A defect verified on synthetic input has a precondition, and the next question is whether that precondition occurs in published work: take public data of the kind the cohort's papers describe, run the shipped code on it the way those papers do (default settings, or the options their methods name), and measure how often the precondition holds and whether the reported numbers cross the thresholds users apply. The answer goes into the issue text and decides the tier. The two checks that established this step went opposite ways: GSEA GS3 (an all-zero gene set under the default scheme) never occurred on airway bulk RNA-seq or PBMC 3k single-cell against every MSigDB collection, so its report now leads with the silent 0/0 path rather than with impact; BCFtools BC1 (32-bit overflow in the bias scores) fires across a tenth of the genome on two SARS-CoV-2 amplicon samples once the depth cap is raised as viral pipelines do, moving hundreds of filter decisions. Where no suitable public data is reachable, say so in the report rather than implying either answer.

5. File upstream, on the project's terms. Before anything is sent, read the project's CONTRIBUTING.md, any pinned policy issue, the NEWS/changelog, and the existing issue and PR history for the component. A finding that upstream has already fixed and announced is not a finding to file. Non-minor changes, and anything that changes numerical output, go to the project's preferred discussion venue first (for Bioconductor packages that is the support site, not a cold PR). Each filing kit records which documents were read and where the ask is being made. This step was added after the DESeq2 round, where the project filed four items without reading the project's pinned policy or CONTRIBUTING.md, claimed a NEWS entry was missing when it was not, and had all four closed by the maintainer within two hours.

6. File sparingly; hold the rest. Volume from one account is what reads as a campaign to maintainers, not any single report. File now only a finding that changes a number that ends up in a paper, under default or common settings, on the current release. Crashes, rare-option paths, API-only paths, documentation drift and design questions are held in the filing kit. At most two filings per repository until a maintainer responds; after a positive signal, the held items for that repository follow one at a time. A comment on an issue the maintainers already keep open is a separate, low-cost category and does not count against the cap. Every prepared finding carries its tier in audits/TRIAGE.md.

Some maintainers have said they do not want AI-generated contributions. That is their call and it is final for this project: the fork of their repository under github.com/cindykrafft gets the topic upstream-declines-ai-contributions, and from then on nothing is filed, commented, re-opened or pushed for that repository, held items included. Open filings are left to the maintainers to close as they see fit; the project stays published here, with the outcome recorded in its README and in the status ledger. Before preparing or opening anything, check the fork's topics; the ledger's site/audits.json and the filing console read the same topic. This step was added after the first week of filings put seven FieldTrip PRs up in two days and a reviewer's reply read as "fine, but minor".

7. Answer what the project already knows. These reading passes kept finding bugs that users had already reported and nobody had claimed (deepTools #1108 and #1118, BEDTools #1142, fastp #474 and #518 among them), and a patch on a thread the maintainers keep open is the lowest-friction contribution there is. So the project now alternates: for each repository already checked, read its open issues, pick one reproducible bug with a wrong-number or crash consequence and no existing PR, reproduce it on master by execution, fix it with a test that fails before and passes after, run the project's own tests and linter, and file the PR against the issue. Everything is recorded under audits/<package>/issue-fixes/, one directory per issue, with the reproduction, outputs, patch, PR body and the other candidates considered. The same limits apply: nothing for a repository that declines AI-generated contributions, and no new PR while the repository has an unanswered one from this project, unless it answers the maintainers' own issue. An issue with an assignee is someone's claimed work: no PR on it; a verified diagnosis goes on the thread as a comment with a link to the branch, and the assignee decides. The project method from steps 1 to 6 resumes once the open-issue backlog of the checked repositories is worked down.

What held up under the same scrutiny is recorded alongside what didn't, so findings are read in proportion.

Status

A live ledger of every filing and its state on GitHub, refreshed four times a day: https://cindykrafft.github.io/mytochondria/ (built from site/ by the Status page workflow).

Filing follows step 6 above; the per-finding tiers are in audits/TRIAGE.md. "Attributable" is the number of those papers whose use can be tied to the audited repository: every paper where the name is the program, and for GSEA and UMAP (method names that several programs implement) only the papers that name the audited program or a package that calls it; the rules, the per-package check of stated versions against the repository's release tags, and the full table are in survey/README.md and survey/scripts/attribute.py.

Package Papers exposed Attributable to the repository Findings Upstream
FreeSurfer 116 116 16, three reproduced numerically 5 fix PRs + 9 issues filed on GitHub (2026-08-30); the maintainers declined AI-generated analysis (2026-09-08): nothing further is filed, commented or pushed for FreeSurfer, the open items are theirs to close, and the remaining findings stay in this repository
FSL 114 114 12, six with patches reports and patches ready; filing via the FSL mailing list
SPM full-codebase audit (not survey-driven) 76 83 confirmed + ~38 plausible; three verified by executing the code (one unit-tested, two reproduced on SPM's tutorial data). Reanalysis of the merged artefact-window fix on ERP CORE (39 participants): on response-locked epochs the shipped code discarded 97 % of detected artefact windows, kept 98 % of trials where the fix keeps 49 %, and the group ERN changed from −10.4 to −8.3 µV (paired p = 0.003) 2 fix PRs merged upstream (M/EEG artefact-window baseline; downsample sampling rate); 3 issue+PR pairs open (χ² EC densities with new unit test; DCM free energy; parametric-modulation padding, filed 2026-09-02) ; issue #106 (assigned): diagnosis + branch offered as a comment 2026-09-03
AFNI 39 in the survey cohort; 57 adjudicated for the ReHo finding 39 36 (9 high-impact, 15 narrower, 12 likely), one reproduced numerically 12 PRs + 14 issues filed on GitHub: 6 PRs merged (ReHo tie handling, NIfTI slice timing, two test-suite repairs, 3dTshift -no_detrend, 3dROIstats -sigma) and the three R-script fixes of PR #960 (3dMEMA, 3dLMEr, 3dMVM) applied on master by the maintainer on 2026-09-22 with credit; 5 open. Reanalysis of the published ReHo design on open data (ds000030, 40 subjects): the pre-fix build returns NaN in 17 % of brain voxels and a third of the correct value elsewhere, and the SCZ-vs-control p < .001 maps from the two builds do not overlap ; issue-fix PR #974 (NIfTI time axis / 3dvolreg exit code, #73) filed 2026-09-04
DESeq2 886 886 3 confirmed (one live 2017–2025, fixed upstream in 1.49.4 with a NEWS entry), 3 verified-negligible 3 issues + 1 PR filed on GitHub 2026-09-01; all four closed by the maintainer the same day (the two NEWS requests were already met, the PR was declined under the project's results-stability policy). The filing skipped DESeq2's contribution guidelines and misread its NEWS; see the filing kit for the correction
MACS2 475 475 3 confirmed (one live in current MACS3, reproduced on shipped binary), 4 notes, 2 withdrawn by own review issue-fix PR #739 for #715 (bdgdiff scores truncated to integers) filed 2026-09-05 and merged 2026-09-18; the audit's own findings not yet filed. MACS has a CONTRIBUTING.md with a PR template and a Slack channel; read both before filing (step 4 above)
Kilosort 60 60 3 new verified on shipped code (one disables KS4's refractory split veto since v4.1.5), 5 by code reading, plus exposure map for the known 2024 "spike holes" bug; the refractory-veto fix (PR #1043) merged 2026-09-25 2 fix PRs + 3 issues filed on GitHub 2026-09-01 (#1043–#1047), all open; no maintainer response yet, but on 2026-09-24 an independent tester (kvnloo) verified PR #1043 in 12 cases and asked for a regression test, which is now on the branch
FieldTrip 42 42 4 verified by executing FieldTrip's own code under Octave (permutation p-values exclude ties, p = 0 with exhaustive permutations; correlationT df; PSI edge bins −Inf; a statfun branch that cannot run) + 1 by code reading 5 fix PRs + 2 issues filed on GitHub 2026-09-01/02 (#2607–#2613), all open; the two behavior-changing PRs carry new test_pullNNNN.m scripts as the project's bot requested; PR #2613 merged 2026-09-04; the maintainer pushed his own normalization fix and tests onto the #2610 branch and asked for test_pull2610 to be removed (done, 9af1fc4) and for a dpss issue (filed 2026-09-05 as #2614); replies to his questions on #2609/#2610 posted the same day; 2026-09-07: #2610 "almost ready to merge", #2614 judged minor (FAQ + Octave survey asked for instead); the FAQ page went to fieldtrip/website as PR #958 and was merged 2026-09-23; the Octave survey ran 2026-09-22 (471 tests, 191 pass, 286 with three fixes on a branch: a one-week-old external/stats path regression that also hits MATLAB, the startsWith/endsWith/contains shims, ft_fetch_data's addOptional); issue, PR and the #2614 comment ready in the kit
Suite2p 32 32 2 verified on the shipped 1.1.0 wheel (bidirectional-phase correction corrupts odd scan lines since v1.0; classifier bins values at the training minimum into the top bin) + 1 limit verified, 2 by code reading 3 fix PRs + 3 issues filed on GitHub 2026-09-02 (#1265–#1270), all open, no maintainer response yet
Seurat 767 767 review 1 (differential expression) done: 1 confirmed undocumented behaviour change (v5 fold-change formula, pseudocount 1/n per group — 79–88 % of lowest-expressed null genes in small clusters reported above +0.25; p-values unaffected), 5 notes, Wilcoxon/MAST paths held up. Profiling from the survey cache only (Europe PMC unreachable; full-text rerun pending) filing channel read; upstream issue #9346 already raised the asymmetry and was closed — its reply must be read before filing; nothing filed
Scanpy 200 200 2 confirmed on master by executing the shipped code (t-test silently run on exponentiated values when mean_in_log_space=False, on main and the 1.13 pre-releases, not in 1.12.4; score_genes top expression bin holds 1–12 genes so lists with the most-expressed genes get almost no controls or error, all versions), 3 notes; Wilcoxon/HVG/scale/normalize paths held up issue #4336 + PR #4337 (t-test, from the fork) filed 2026-09-02, open, release-note fragment renamed to the PR number; the score_genes PR waits on regenerating the paul15 reference pickle; filing kit
Scrublet 78 name Scrublet of 179 in the survey's doublet-tool group (lower bounds) – 4 confirmed by executing the shipped code — two in Scanpy's port (sc.pp.scrublet scores every cell over k_adj − 1 neighbours instead of k_adj, with the variant chosen by whether any two points coincide; log_transform=True transforms observed and simulated cells on different scales — both on main and 1.12.4) and two in the original (UMI-subsampled synthetic-doublet totals inflated by rate·(1−rate)·F; mean_center=True, normalize_variance=False crashes on scikit-learn ≥ 1.4 — master = PyPI 0.2.3, non-default options), 4 notes, 4 withdrawn; the default original pipeline equals an independent port bit for bit; the two packages' calls agree (Jaccard 0.986) while their scores do not (Spearman 0.43, 0.996 once given the same simulated doublets) not yet filed. Scrublet has no CONTRIBUTING/templates/tests and no maintainer activity since 2020 (plain issues + patches with new pytest files ready); Scanpy kit follows its bug-report template, Ruff and towncrier conventions, two patches with tests that fail on main; see the filing kit
CellPhoneDB 46 46 6 confirmed by executing the shipped code (permutation workers draw duplicate shuffles, so iterations is effectively divided by threads — 279 distinct permutations of 1,000 at the default 4; p-values count only shuffles strictly greater than the observed mean, so ties are dropped and p = 0 is reported for 42.6 % of tested entries on the project's own example data, verified against an exhaustive 1,680-permutation null and on every release v2.1.7–5.0.1; threshold inert in METHOD 1; deconvoluted.txt reports the complex minimum on every subunit row; ZeroDivisionError at threads=1, iterations<=50; scoring broken on pandas 3), 6 notes; cluster means, complex minima, threshold rule, p-value counting, significant_means, microenvs, DEG method and the v5 scoring protocol held up filing channel read (no CONTRIBUTING, no templates, no changelog convention); issue #231 (duplicate permutations across workers) filed 2026-09-03, open; the other five wait for a maintainer signal, and the replies on #179/#60 must be read before filing CPDB2; filing kit with 2 tested patches on the fork ; issue-fix PR #232 (create_db with an unused uniprot_N column, #224) filed 2026-09-04
umap-learn 1,111 name UMAP, which is a method as well as a program: 202 identify umap-learn (Scanpy, umap-learn, umap.UMAP, Python), 562 name a route that runs uwot instead (Seurat's RunUMAP, Monocle, ArchR, Signac), 347 name only the method 202 1 confirmed on master, 0.5.12 and 0.5.3 by executing the shipped code (smooth_knn_dist returns sigma = inf for every point with a neighbour pruned by disconnection_distance, so all its remaining edges get membership 1.0; fires on the exact, pynndescent and precomputed paths and, with the default settings, for metric="jaccard" on sparse binary data), 2 notes (duplicate rows collapse sigma to the floor; transform builds a different local graph than fit), 2 withdrawn (random_state vs n_jobs, issue #1080 is stale on master; sparse vs dense metrics); sigma/rho vs the paper's equation, symmetrisation, find_ab_params, spectral init, sparse vs dense metrics and random_state across n_jobs held up filing channel read (CONTRIBUTING only, no templates); issue #1286 + PR #1287 (from the fork) filed 2026-09-03; #1286 resolved and PR #1287 merged 2026-09-05 (no discussion); second issue-fix PR #1288 for #1194 (transform refusing precomputed distances above 4096 rows) filed 2026-09-08 and merged 2026-09-13; a third fix is being prepared; filing kit
Cutadapt 331 331 4 confirmed by executing the shipped code on main, 5.2 and older wheels (-e N loses one allowed error for adapter lengths 49, 98, … through floating-point rounding of N/len, all versions since 3.0; --max-ee/--max-aer ignore --quality-base, so Phred+64 data are never filtered; the demultiplexing index ranks barcodes by matching bases while --no-index ranks by the documented alignment score, since 5.0; the k-mer prefilter added in 4.3 silently drops ~40 % of reads carrying an anchored/non-internal adapter with one inserted base whenever exactly one error is allowed), 4 notes, 4 own suspicions withdrawn; quality trimming, expected errors, filters, pair-filter, interleaved/multi-core and the aligner invariants held up against independent ports on ~100k random cases filing channel read (CONTRIBUTING, bug-report template, CHANGES.rst convention); no prior issue for any of the four; issue #892 + PR #893 (-e N rounding, from the fork) filed 2026-09-03, open, no maintainer response yet; the other three patches wait for a maintainer signal; filing kit ; issue-fix PR #894 (--info-file offsets after 5' trimming, #518) filed 2026-09-04
deepTools 167 167 9 confirmed by executing the shipped code on master (= 3.5.6), 3.5.1 and 3.3.1 (bamCompare --skipZeroOverZero shifts every bin after a skipped one — open upstream as #1108/#1130 since 2021; plotPCA --log2 inert and --rowCenter inert at the default --ntop; --removeOutliers scales by the median instead of the MAD and removes nothing; --MNase counts four bases for odd fragments — open as #1118; BPM ≡ CPM against its documented definition; --smoothLength truncated at chunk edges; --ignoreDuplicates alone left out of the CPM/RPKM/RPGC and readCount denominators, 40 % low on a 40 %-duplicate sample; multiBigwigSummary reports zoom-level summaries, median 3.4 % and up to 245 % off at the default 10-kb bins; plotFingerprint over-counts a fragment's last bin when the sampling step equals the bin size), 11 notes, 2 withdrawn; bamCoverage/bamCompare arithmetic, computeMatrix binning, the profile/heatmap summaries, multiBamSummary, plotCorrelation, plotPCA default, bamPEFragmentSize, the fingerprint metrics and the coverage/filtering/enrichment counts held up against numpy/scipy and independent ports filing channel read (markdown issue/PR checklists, flake8 + pytest CI, no changelog rule); two findings already open upstream unanswered, seven new; filing kit with 9 issue/comment texts, MCVEs run on 3.5.6 and 3.3.1, and 7 git am-able patches whose new tests fail on unmodified master; nothing filed; 4.0.0 rewrite landed 2026-09-04; findings re-verified on it (5 survive, 2 fixed, 1 not applicable); two new pre-release defects filed 2026-09-05: #1457 + PR #1458 (bamCompare --operation), #1459 + PR #1460 (plotPCA loadings; closed 2026-09-18: bug confirmed by the maintainer and fixed his own way in #1472, merged the same day); PR #1451 for #1423 closed with the merge; the rest held
IQ-TREE 2 258 258 0 wrong numbers on master; 4 notes verified by executing the built binary (parametric aLRT prints the cube of 1 − p; UFBoot's 1-log-likelihood candidate cutoff gives 100 % where the standard bootstrap gives 61 % on a degenerate 4-taxon alignment, not reproduced on 6 taxa; SH-aLRT/UFBoot vary with thread count at fixed seed; rooted trees get NA/abort in gCF/sCF), 1 withdrawn; likelihoods, gamma rates, +I/+ASC, ModelFinder criteria and parameter counts, UFBoot/--bnni supports, SH-aLRT, aBayes, gCF/sCF held up on master and v2.4.0; the four notes reproduce on iqtree3 3.1.3 filing channel read (Issues for bugs, Discussions otherwise; no CONTRIBUTING or templates; development moved to iqtree/iqtree3); nothing rises to an Issue; kit holds two manual patches and a Discussion draft, none filed ; issue-fix PR #207 (outgroup absent from a partition/quartet, #203) filed 2026-09-04; issue-fix PR #207 (setRootNode assertion, #203/#89) merged 2026-09-08; second issue-fix PR #210 for #192 (+FU codon frequencies ignored by GY models since 3.x) filed 2026-09-08 and merged 2026-09-09 ; third issue-fix ready 2026-09-09 for #198 (SH-aLRT values inflated whenever -j/-J is given: the RELL replicates were jackknife samples, uncentred), PR #214 filed 2026-09-09
fastp 117 117 3 confirmed by executing the built binary (--cut_front/--cut_tail with --trim_front1/--trim_tail1 drop cut_window_size − 1 extra bases from every read — 52 nt of a 60 nt all-Q40 read instead of 55 — on every release built, 0.20.0 through 1.3.6 and master; an auto-detected built-in adapter longer than 60 nt is printed and then emptied by resize(0, 60), so 0 of 20,000 reads are trimmed and the JSON loses its adapter_cutting section, on 0.26.0 through master, hitting the 85 built-ins with no shorter prefix, i.e. the TruSeq Small RNA primers; the one-indel adapter search passes the read start instead of rdata + pos, so an adapter with an indel at the 3' end is never trimmed, 0 of 200 reads), 7 notes (overlap mismatch limits apply to the first 50 overlap bases only; insert-size histogram counted by thread 0 only; --overlap_len_require/--poly_g_min_len off by one; -m silently enables -c), 5 withdrawn; filters, filtering_result, Q20/Q30/GC/mean length, the k-mer table, poly-G/poly-X, base correction, UMI, --split, --reads_to_process and the duplication estimator held up against independent ports on 10^4–10^5 reads filing channel read (no CONTRIBUTING, no templates, no changelog; ./fastp test and scripts/*.sh are the test conventions); two of the three already have open upstream issues (#474, #518) so their texts are drafted as comments there, the third is a new issue; filing kit with three git am-able patches (fix + test, each new test failing on unmodified master); nothing filed, nothing pushed ; issue-fix PR #715 (--dedup drops counted as passed, #638) filed 2026-09-04
BEDTools 302 302 5 confirmed by executing the built binary on master (= 614e9a5), v2.31.1 and v2.30.0 (coverage -split counts overlapping blocks, not database records, and ignores -f/-F — wrong count since #673 in 2018; intersect -split tests -F/-r against the summed block length of all hits and clears the whole group, 530 false negatives + 48 false positives in 838 pairs — open upstream as #1142; closest -t first/last break a left/right tie by -D stream order not B-file order, 1027/2000 reverse-strand queries under -D a; reldist/subtract/flank/closest -d truncate 64-bit coordinates to 32-bit int on chromosomes > 2.15 Gb, dropping queries or emitting negative coordinates; slop -pct/flank -pct lose one base at 12 of 99 whole-percent values and absolute slop rounds above 2^24), 1 note (shuffle -incl lets features spill past the include interval — #1089) + 8 minor notes; the whole intersect/coverage/subtract/window/merge/cluster/map/groupby matrix, closest/genomecov (BED+BAM)/multicov, fisher (vs scipy.stats.fisher_exact), jaccard, reldist, nuc and shuffle/random uniformity held up against independent Python ports filing channel read (no CONTRIBUTING/templates/linter; CI is make test; docs/content/history.rst is the changelog); BT2 matches open #1142, BT1 relates to open #673, three findings new; filing kit with 3 git am-able patches (fix + tests failing on unmodified master, full suite passing) + 5 issue/comment texts + MCVEs; nothing filed ; issue-fix PR #1143 (flank -s drops strand-less records, #1123) filed 2026-09-04; BT2 filed 2026-09-05 as a comment on #1142 (cross-linked from #1141) with PR #1144; PR #1143 for #1123 open
HTSeq 161 161 2 confirmed by executing the shipped code on main, the 2.1.2 wheel and the 0.11.2, 0.12.4 and 0.13.5 sdists (htseq-count compares the MAPQ field of an unmapped mate record, 0 by aligner convention, with -a, so every pair whose mate is unmapped but present in the BAM goes to __too_low_aQual under the default -a 10 while the same pair counts when the record is absent, contrary to the FAQ — 113 of 113 such pairs lost on a 2,000-pair synthetic library, all versions; BAM_Reader[iv] calls pysam fetch with iv.start + 1 and misses alignments whose last base is iv.start, API only, all versions), 6 notes (--nonunique fraction divides by overlapping features not by NH, so with --secondary-alignments score the documented "sum equals the number of reads" fails; htseq-count-barcodes lets __no_feature/__alignment_not_unique/__too_low_aQual reads outvote a gene inside a UMI; docs say --secondary/--supplementary-alignments default to score but the code defaults to ignore since 0.10.0; a pair with a mate on a chromosome absent from the GTF is __no_feature as a whole; zero-length GenomicInterval.overlaps asymmetric; float32 count matrices), 1 withdrawn; the three overlap modes, all --nonunique modes, strandedness, -r name/-r pos pairing, NH/secondary/supplementary handling and the GenomicArrayOfSets/StepVector arithmetic equal an independent per-base port on all 324 single-end and paired-end configurations once HC2 is modelled, and the documentation figure's eight rows reproduce exactly filing channel read (markdown bug-report template, no CONTRIBUTING/PR template/linter; doc/history.rst is the changelog); no prior issue for either (nearest #99, #94, #96); filing kit with two git am-able patches (fix + test failing on main + history entry) ready; nothing filed; issue #109 + PR #110 filed 2026-09-03 and closed by the maintainer the same day; the project declines AI-generated contributions, so nothing further is filed (HC1 held); fork tagged upstream-declines-ai-contributions
featureCounts 331 (41 name 2.0.1) 331 2 confirmed by executing the built binary on master (= 2.0.6), the SourceForge 2.1.1, 2.0.3 and 2.0.1 tarballs (-p --countReadPairs --splitOnly counts every non-split fragment with only one mapped record — 1,000/1,000 singletons Assigned — on all four; the manual's read-type filter for single-end reads in stranded pair counting was removed with 2.0.2 while its Unassigned_Read_Type row stayed, so orphaned second reads of a dUTP library go to the antisense gene — S = 800, AS = 500 instead of 300/0 — on 2.0.3, 2.0.6 and 2.1.1, not 2.0.1), 8 notes (paired-end filter-order labels; .-strand features under -s 1/-s 2; --byReadGroup loses the -Q row; -p counts fragments on 2.0.1 and reads since 2.0.2; -C same-strand pairs; pre-2.0.6 --largestOverlap --fraction; -f row names; repeated GTF keys), 4 withdrawn; the documented assignment rules equal an independent implementation on 75 option sets × 3 seeds of random single-end and paired-end BAMs (0 count or status differences on master and 2.1.1), the pairer on 179k coordinate-sorted records, -J, threads, per-file -s, GTF/SAF parsing held up filing channel read (README only: no CONTRIBUTING, templates, changelog or linter; shell-script tests with .ora expectations); the GitHub tracker could not be read from this session (search returns nothing for the repository, list_issues not enabled, API and HTML blocked) and the repository is a bulk mirror of the SourceForge/Bioconductor releases; both findings held (rare option / mixed SE-PE input); filing kit with two git am-able patches whose new test/featureCounts cases fail on unmodified master, MCVEs run on four versions, FC2 drafted as a question-first issue; nothing filed
PLINK 184 (84 name a 1.9 build, 13 name 2.0) 184 1 confirmed by executing the built binaries on master, the v1.9.0-b.7.12 tag (2026-09-01), v1.90b6.21 and v1.90b4 (PLINK 1.9 --hwe/--hwe midp removes variants with 2 or 3 heterozygotes whose exact HWE p is above the threshold — SNPHWE_t never compares the last element of the observed tail, so the filter tests p − P(hets − 2) < t; every one of 9,798 two-het tables with n ≤ 200 flips below its p, by up to 5 % of p; --hardy prints the right p; plink2 unaffected), 5 notes (one open: plink2 --score center with imputed dosages differs from 1.9 and the reference by up to 6.5e-2, cause not traced), 3 own suspicions withdrawn; HWE and Fisher exact p on all tables with n ≤ 60 / N ≤ 40 to 1e-15, --hardy, --freq, --missing, --assoc/--model/--logistic/--linear/--glm, all --adjust columns, --r2/--ld, --genome, KING, --score held up against exact arithmetic, scipy/statsmodels and independent ports filing channel read (no CONTRIBUTING/templates/changelog; 1.9 has no unit-test runner); no prior issue (nearest #128); filing kit with issue text, MCVE run on four 1.9 builds, and a 28-line git am-able patch that brings the exhaustive count from 19,399 to 0; nothing filed; issue #380 + PR #381 filed 2026-09-03; the maintainer closed the PR and committed the same fix himself the same morning (1fe42e5), verified by rerunning the exhaustive harness at that commit: 0 tables wrongly removed
clusterProfiler + fgsea 244 (122 name fgsea, 720 GSEA) 244 8 confirmed by executing the shipped code — on the current release (4.20.0 + CRAN enrichit 0.1.4–0.2.1; fixed on devel, 0.2.2 not on CRAN): GSEA()/gseGO() tables filtered on raw p only (515 of 717 GO BP rows with p.adjust > 0.05; 292 "enriched" terms on pure noise); on every legacy version the cohort pins (≤ 4.18, DOSE engine; fixed on devel): the leading edge of a negative-ES set taken from the wrong peak, so core_enrichment lists extra genes in 6.1 % of suppressed pathways (median +12 %, up to +225 %), and query genes outside a user universe counted as draws (p ×1.21, 191 of 1,130 calls differ on GO BP); at devel: gsea(method = "sample"/"permute", adaptive) p ≈ half the GSEA convention against exact enumeration (non-default; patch + test), simplify() removes terms similar to no kept term (4–8 % of removals); three release-era engine defects fixed on CRAN (BH over zero-overlap sets = open #819, 513 vs 694 terms; NaN for small statistics; raw-size filter); 8 notes, 3 withdrawn; hypergeometric ORA vs scipy (1e-15), ES/leading edge/NES/sizes/gseaParam/scoreType vs an independent numpy GSEA and multilevel p-values vs exact enumeration held up for fgsea, enrichit and clusterProfiler filing channel read (issue template + CONTRIBUTING, maintainer's guide and book unreachable from the session — read first; enrichit and fgsea have no templates); nothing rises to file-now: CP1 fixed and announced upstream, CP2/CP3 legacy only, CP4 non-default; kit with one git am-able enrichit patch (test fails on devel, suite passes), a comment draft for #819, a CP1 heads-up and three MCVEs; nothing filed
lme4 312 312 2 confirmed by executing the shipped code on master, 2.0-6 and 1.1-35.1 (glmer's default finite-difference-Hessian standard errors ~100× too small in 9 of 150 ordinary Bernoulli random-intercept fits, z
edgeR 318 318 3 confirmed by executing six shipped builds (3.36.0, 4.0.16, 4.4.2, 4.10.1, the 4.10.5 release branch and devel 4.99.4) against ports of the published methods: filterByExpr drops a gene with exactly min.count reads in the median-size library for 12–17 % of library sizes because the C cpm() and the R cutoff round differently at the last bit (all versions, ~4 genes per 20,000 with an odd sample count); cpm()/rpkm() with an offset matrix counted library sizes twice in 4.10.0–4.10.2 and aveLogCPM(offset=) mis-fitted the first gene in 4.4.0–4.10.3 — both already fixed by the maintainers in August 2026 without a NEWS entry; 8 notes (a harsh-regime QL simulation among them), 4 withdrawn; TMM/RLE/UQ, log-CPM, the exact test, the Cox–Reid APL and WLEB dispersions, glmFit/glmLRT, both QL methods' arithmetic and BH held up to 2e-8 or better, and the default pipeline is calibrated in typical regimes filing channel read (Bioconductor support site, no tracker/PRs, edgeR-Tests.R convention); nothing filed; kit with the support-site post and git am-able patches for devel and RELEASE_3_23 (fix + test, R CMD Rdiff clean)
limma 234 234 2 confirmed by executing seven shipped builds (3.34.0, 3.42.2, 3.58.1, 3.62.2, 3.68.4, the 3.68.5 release branch and devel 3.99.0) against closed forms and ports of the published methods: arrayWeights(method="reml") with prior weights divides its convergence criterion by ngenes+prior.n twice, so it stops after 2 Fisher-scoring iterations instead of 5 and returns weights up to 15 % from the REML solution (every version since 2019; the path devel's voomLmFit(sample.weights=TRUE) now takes by default for 4.0.0; voomWithQualityWeights(method="reml") calls 137 instead of 183 genes in a simulation); and voom() (3.68.x) row-centres a DGEList's edgeR-style offset and adds it to log(lib.size), counting the library sizes twice (E shifted by −log2(lib.size/geomean) per column, every logFC by the groups' log library-size ratio, 5,374 vs 37 genes called; devel's voomLmFit reads the same slot as exp(offset); edgeR 4.10.5 voomLmFit affected too); 9 notes (the (2, 9998) df.prior bound of the unequal-df estimator, the documented contrasts.fit approximation at 4.6 % median under voom weights among them), 4 withdrawn; lmFit/GLS, contrasts.fit without weights, eBayes/treat/topTable/decideTests, voom, duplicateCorrelation vs lme4, camera, quantile normalisation and removeBatchEffect held up to 1e-8 or better, and the new C backends match the R code they replace filing channel read (Bioconductor support site, no tracker/PRs, changelog.txt + stopifnot test convention); nothing filed; kit with two support-site posts (LI1 first) and git am-able patches for devel (both) and RELEASE_3_23 (LI1), each with a test failing on the unmodified code
SAMtools 692 (the most-used unaudited tool; 196 state a version, 1.9 ×92, 1.10 ×36, 1.3.1 ×26) 692 2 confirmed by executing the built binaries on develop (ce612d2), 1.19.2, 1.10 and 1.9 (samtools stats computes its coverage distribution in a 1,500-cell ring buffer indexed modulo its size, so the far exon of a spliced RNA-seq alignment is added to positions next to the read start — on a simulated six-exon gene 2,591 covered positions for a true 2,197 and a -t/-g "percentage of target genome with coverage > 0" of 117.77 %; and the buffer reallocation for a longer read copies with byte lengths where element counts are needed, dropping pending counts on mixed-length and long-read data), 6 notes (stats -p misses overlaps of unequal-length mates; reads duplicated counts supplementary records while sequences does not; insert-size SD skips bin 0; depth -s broken by supplementary records; coverage --rf is any-of; stats -r mismatches per cycle ignore N), 3 withdrawn; flagstat, idxstats, view -c, depth (11 modes), coverage, the 31 stats SN numbers, markdup counts/optical/library size vs a Picard port, and mpileup depth with overlap removal, -Q, -d and BAQ all held up against independent Python truths filing channel read (CONTRIBUTING.md with a DCO sign-off and an AI policy: human-written commit messages, Assisted-by: line, no AI sign-off; markdown bug-report template; NEWS.md entries; make test via test/test.pl); no prior issue (nearest #1003, #640, #969); filing kit with one issue text, MCVE run on four builds, and a git am-able patch (fix + two regression tests failing on unmodified develop + NEWS entry); filed 2026-09-08: issue #2378, PR #2379 (branch fix/stats-coverage-ring-buffer on the fork, commit signed off by the submitter with the Assisted-by: trailer per the samtools AI policy); PR #2379 merged 2026-09-24 (daviesrob confirmed the results against the original code), #2378 closed; second kit ST3 prepared 2026-09-24: the duplicate-count note is the open maintainer question #696 (2017), so a comment there plus a PR closing it, with the insert-size SD fix as a second commit; filed 2026-09-24: comment on #696, PR #2388; 2026-09-28 a maintainer asked for isize 0 to be excluded rather than included (reworked commit prepared, reply drafted)
Trimmomatic 291 (116 name a version: 0.39 ×82, 0.36 ×53, 0.38 ×23) 291 2 confirmed by executing the built jar on main (= ef98d62, 0.41), the 0.41/0.40/0.39 release jars and the 0.38/0.36/0.33/0.32 tags built from source (ILLUMINACLIP palindrome mode charges int(Q/10) per mismatch instead of the README's and simple mode's Q/10 — a Q19 error costs 1, not 1.9 — so a 2x50 pair with a 42-nt insert and two Q19 errors scores 31.7 instead of 29.9 and is clipped at the default threshold 30, forward cut and reverse dropped: 600 of 7,200 grid pairs and 23 of 4,000 simulated 2x50 pairs flip, none with 75-nt or longer reads; MAXINFO's long score tables are scaled with Math.max of the two ratios so they saturate and wrap, trimming every read to 1 base for targetLength ≥ 248 at strictness 0.1, ≥ 495 at 0.2 or ≥ 711 anywhere — 41 of 150 grid cells, all versions), 8 notes (simple mode scores the best sub-range, not the sum; every-4th-position seeds for long adapters; N scored as Q0 everywhere; SLIDINGWINDOW drops reads shorter than the window; TRAILING never tests base 0; Phred auto-detection exits on all-Q27+ data and picks Phred+64 for Q47+; unequal PE inputs silently truncated; MAXINFO uses (Q+0.5)/10), 5 withdrawn; SLIDINGWINDOW, LEADING/TRAILING, the length/quality filters, TOPHRED, MAXINFO at typical settings, palindrome geometry (18 insert sweeps), simple-mode clip positions and fragment minima, PE routing, -summary, -trimlog and -threads invariance held up against independent Python references on 10^3–10^4 synthetic reads filing channel read (no CONTRIBUTING/templates/linter; versionHistory.txt last entry 0.40; CI mvn -B clean verify on JDK 25; JUnit 5 tests); tracker active, no prior issue for either (nearest #52, #16; #74, #22); filing kit with two issue texts + MCVEs run on main/0.41/0.39 and two git am-able patches (fix + JUnit test failing on unmodified main; suite 259 → 261 / 263 passing); TM1 ranked first; nothing TM1 filed 2026-09-24: issue #90, PR #91 (branch fix/palindrome-mismatch-penalty on the fork); PR #91 merged into V0.42 2026-09-28 (scheduled for 0.42), #90 closed
BCFtools 202 (56 state a version, 1.9 ×31, 1.10 ×12, 1.13 ×6) 202 1 confirmed by executing the built binaries on develop (7abcc0d6), 1.24 and 1.13 (bcftools mpileup computes the default INFO/MQBZ, BQBZ, RPBZ, SCBZ, MQSBZ and NMBZ Mann-Whitney Z-scores with 32-bit int accumulators, so the tie term wraps once one quality bin holds 1,291 reads and the scores shrink towards zero — 1,300 MAPQ-60 reference reads + 40 MAPQ-30 alternate reads give MQBZ −7.81 for a true −36.47, 44 single-sample BAMs under the default -d 250 give −6.97 for −36.73, and the score saturates near −11 at 12,000 reads for a true −109.72; 1.9 and 1.10.2 predate the annotations), 6 notes (merge takes unruled Number=A tags from the last file, not the first as documented; stats AF labels on a 100-grid for 99-grid bins; +fill-tags HWE at multiallelic sites; -d is a memory guard not a depth cap; flat PL → ./.; PSC nSingletons semantics), 3 withdrawn; mpileup DP/DP4/AD/SP/I16 under four filter settings, single-read PLs, +fill-tags counts and exact HWE/ExcHet, a port of call -m (QUAL/GT/GQ/AC/AN, -v, -P, --ploidy), stats SN/TSTV/SiS/AF/DP/PSC, norm left-alignment and split/join arithmetic, 42 filter expressions and merge all held up against independent Python truths, on develop and 1.24 filing channel read (CONTRIBUTING.md with a DCO sign-off and an AI policy: human-written commit messages, Assisted-by: line, no AI sign-off; no issue or PR template; NEWS entries; make test via test/test.pl); no prior issue (nearest #2003, #897, #1058); filing kit with one issue text, an awk MCVE run on four builds, and a git am-able patch (fix + regression test failing on unmodified develop + NEWS entry; make test 2482/0 with, 2480/2 without); not yet filed ; issue #2594 + PR #2595 filed 2026-09-23 from the fork, fork CI green; 2026-09-28 a maintainer posted a refactor that avoids the cube altogether, verified here; 2026-09-29 he opened #2598 as the replacement (our commit intact plus his refactor), now with pd3
STAR 489 (150 state a version: 2.7.10a ×26, 2.5.2b ×25, 2.7.9a ×24, 2.7.3a ×20) 489 2 confirmed by executing the built binaries on master (= 2.7.11b, b1edc12), 2.7.10a and 2.7.9a (--quantMode TranscriptomeSAM writes every alignment to a transcript whose GTF strand is . — StringTie single-exon transcripts, merged/de novo GTFs — with the strand flag inverted and the sequence reverse-complemented at the unchanged + position, 300/300 + 300/300 SE reads and 196/196 + 200/200 PE pairs on two . transcripts vs 0 on 14 stranded ones, so RSEM/salmon see reads mismatching at every base; STARsolo --soloStrand Forward counts 0 of 300 sense reads of a . gene in Gene/GeneFull/GeneFull_ExonOverIntron and counts them under Reverse, GeneFull_Ex50pAS 0 under both — same root cause, trStr==1 ? Str : 1-Str), 5 notes (XS from the annotation strand; spliced-vs-clipped ties reported as two loci; non-canonical junctions displaced to a nearby canonical motif; 2-pass junctions counted as annotated in Log.final.out; overhang ignores indels), 3 withdrawn; GeneCounts (19 rows × 3 columns × 4 configurations), TranscriptomeSAM for stranded transcripts, SJ.out.tab in 5 filter configurations, NH/HI/MAPQ/AS/nM/jM/jI/flags for 10,264 alignments, the mismatch/score/multimap filter boundaries, all 26 Log.final.out lines and 2-pass insertion held up against independent Python truths filing channel read (CONTRIBUTING.md: questions to the Google group, bug-report and PR guidelines, "default behaviour must not change"; no templates, CI, test runner or linter; CHANGES.md bullets); no prior issue (nearest #1922, #1880, #2679); filing kit with one issue text (STA1+STA2), an eight-read MCVE run on four builds, and a git am-able patch (five one-line fixes + regression script failing on unmodified master + CHANGES entry); nothing filed; filed 2026-09-24: issue #2707, PR #2708 (branch fix/undefined-strand-transcripts on the fork)
GSEA 720 name GSEA, which is a method as well as a program: 69 identify the desktop application (name, URL, or a stated desktop version), 233 name another implementation (fgsea, clusterProfiler, GSVA/ssGSEA, GSEApy, web tools), 418 name only the method 69 3 confirmed by executing the jar built from master (dc35c76 = v4.4.0) and the 4.4.0, 4.3.2, 4.1.0 and 4.0.3 tags (a gene set whose present members all have ranking score 0 gets a NaN hit weight under the default weighted scheme and is reported with the running sum before its first hit as ES plus an NES, p and FDR — ES −0.7056, NES −2.02, p = 0 for a 40-gene set on a list with 30 % zeros, no warning; weighted_p1.5 gives members with a negative score a hit weight of 1e-6 instead of r
AlphaFold 3 code review of the inference and data-pipeline code (2026-09-02), not survey-driven – 15 fix branches on the fork; the two next in line reproduced by execution on main (Input.to_json() reorders identical chains separated by another chain, so the documented two-stage _data.json workflow runs a different complex than the end-to-end run, since v3.0.1; summary_confidences.json chain_ids is per token, not per chain as documented, 150 entries against 2 on the shipped test output, v3.0.4 and main) PR #734 (a userCCD overlay mutated the shared cached CCD) filed 2026-09-03 and closed 2026-09-21 by the maintainer, who committed the same fix (3485995) with credit after the CLA check failed; kits for the next two ready 2026-09-22, CLA to sign first; filing kit
Picard 289 (143 name MarkDuplicates, 74 report a duplication rate; 66 state a version: 2.2x ×41, 2.0x ×28, 2.1x ×17, 3.x ×7) 289 3 confirmed by executing the built jar on master (c2a483d) and the 2.18.7, 2.27.4, 3.0.0 and 3.3.0 release jars (CollectRnaSeqMetrics never counts the last base of an alignment block in transcript coverage, so every internal exon end has ~0 coverage: CV 0.032 for a uniformly covered transcript, MEDIAN_CV_COVERAGE 0.430 → 0.436 and the 3′ end of the coverage curve −9 % on a realistic library; CollectHsMetrics ZERO_CVG_TARGETS_PCT divides unique uncovered targets by the raw interval count, 0.5 for a true 0.667, a regression since 2.19.0; CollectAlignmentSummaryMetrics in bisulfite mode compares against the reference base at the read's offset from the contig start, mismatch rate 0.24 for a true 0 and a crash on short contigs), 4 notes (GC-bias window off-by-one and contig-end binning, prior #1278; BAD_CYCLES block offset; het-sensitivity sampler; FOLD_80 undefined); MarkDuplicates (flags, seven metrics, optical clustering, library size), CollectInsertSizeMetrics, CollectWgsMetrics and the non-bisulfite alignment summary all held up against independent Python truths filing channel read (PR template with a checklist, issue template, no AI policy; nine prior threads read in full, none reports the three); filing kit with three issue texts, three single-commit branches (fix + TestNG regression test) and PR bodies; fork of broadinstitute/picard needed
Bowtie 2 477 (139 name paired-end use, 60 a MAPQ threshold, 27 an XS/"uniquely mapped" filter, 27 a -X value; 89 state a version: 2.4.2 ×24, 2.4.5 ×20, 2.3.5.1 ×19, 2.5.x ×23) 477 4 confirmed by executing the built master (58e34bf) and the 2.3.5.1, 2.4.2, 2.4.5, 2.5.1 and 2.5.4 release binaries on reads built from random references (a master-only regression from 2026-09-14 computes mate 2's MAPQ with its own length as the opposite mate's, so mates of unequal length get different MAPQs for the same concordant pair, 40/23 and 6/30; a mate reported as an unpaired alignment never carries XS:i even with a second-best alignment and MAPQ 0/1, every version; --no-mixed also suppresses discordant alignments while the manual says the opposite, 80 → 0 and the alignment rate 82 % → 58 %; fragments 18–33 bp longer than -X get one mate with a spurious 2–12-bp insertion at MAPQ 23–42 because the rescue alignment squeezed into the -X window is kept and the exact one is dropped as redundant, 0.4 % of mates on a N(380, 70) library at the default -X 500), 3 notes (-X bounds the mapped extent while TLEN includes outer soft clips; the -3/-5 manual claim; --score-min truncation); alignment scores and every tag, the default MAPQ for unpaired reads and equal-length pairs, pair classification under -I/-X/orientations/overlap options, the alignment summary in five modes and the --un-conc/--al-conc files all held up against independent Python truths filing channel read (no CONTRIBUTING.md, no templates, no AI policy; CI make simple-test); twenty prior threads read in full, none reports BW1, BW2 or BW4 (#430 is a different --no-mixed effect the maintainers said they were fixing in 2023); filing kit with four issue texts, two single-commit branches (one-token and six-line fixes, project tests run) and PR bodies; fork of BenLangmead/bowtie2 needed
FastQC 322 (85 state a version: 0.11.9 ×56, 0.11.8 ×30, 0.11.5 ×22, 0.12.1 ×13; MultiQC in 44; adapters named in 90, duplication in 68) 322 3 confirmed by running the compiled master (87fb336) and the 0.11.5, 0.11.9 and 0.12.1 archives on FASTQ files with known qualities, compositions and duplicate structure (a Phred+33 file with no quality below 31 is read as Illumina 1.5 and every quality is reported 31 too low, Q37 → 6, so the quality modules go red; each read's mean quality is truncated toward zero before binning, so the per-sequence quality histogram sits 0.5 too low on average; the GC module counts only the first 100 bases of 101–199-bp reads while its page says the whole read, 45 % vs 63 % on 150-bp reads with a poly-G tail), 4 notes (the documented 100,000-sequence tracking limit; integer %GC and mean length; unweighted group statistics; floor-rank percentiles); Basic Statistics, per-base quality statistics, content, N, length, the GC model, the duplication estimator (within 0.07 points), adapter content and overrepresented sequences all held up against exact truths and ports filing channel read (no CONTRIBUTING, templates, AI policy or test suite; twelve prior threads read in full, #147 is the encoding case, none reports FQ2 or FQ3); filing kit with three issue texts, one single-commit branch (one-line fix, verified on the patched build) and a PR body; fork of s-andrews/FastQC needed
VCFtools 112 (20 state a version: 0.1.16 ×24 mentions, 0.1.13 ×9, 0.1.14 ×6, 0.1.17 ×3; LD in 25, Fst in 25, --maf in 21, π in 20, --max-missing in 17, kinship in 13) 112 3 confirmed by running master (1f87a83) and the v0.1.16 and v0.1.13 tags built here plus Ubuntu's 0.1.16-3 package on generated VCFs against exact recomputations (every temporary-file output mode — --012, --plink, --hap-r2, --geno-r2 and nine more — aborts with "buffer overflow detected" on a source build with a current compiler's default fortification, a one-byte-short stack buffer at 19 sites; --relatedness2 scales the KING-robust kinship down by the pairwise call rate when genotypes are missing, duplicates 0.40 for 0.50 at 20 % missing; --max-missing-count counts missing alleles while the manual says genotypes), 3 notes (--missing-site and --max-missing disagree about depth-filtered genotypes; --TajimaD fixes n at 2 × individuals; --012 is biallelic-only with only a log warning); π, windowed π, Tajima's D on complete data, --het, the HWE exact test, Fst per site and windowed with and without missing data, LD r²/D/D', kinship on complete data and every site and genotype filter checked held up filing channel read (no CONTRIBUTING, templates, AI policy, tests or CI; twelve prior threads read in full, #28 is VT3 unanswered since 2016, none reports VT1 or VT2; no maintainer reply in any thread read); filing kit with three issue texts, two single-commit branches (both verified on patched builds) and PR bodies; fork of vcftools/vcftools needed
scDblFinder 179 (6 state a version: 1.16.0 ×2, 1.12 ×2, 1.13.13, 1.4.0; Seurat in 150, Scanpy in 65; Scrublet also used in 78, DoubletFinder in 73; per-sample runs stated in 14) 179 3 confirmed by running the Bioconductor 3.18 release 1.16.0 and 1.4.0, 1.8.0, 1.12.0 built from their release commits on simulated captures with known doublets and origins, devel (2d0e5e4) read (with clusters given, the size-adjusted quarter of the artificial doublets is returned out of order while the origin labels are attached by position, so about half are mislabelled and scDblFinder.mostLikelyOrigin is right for 24 % of heterotypic doublets, calls unchanged; a factor clusters argument is stored as integer codes so scDblFinder.stats shows 0 observed doublets and scDblFinder.cluster holds the codes; doubletThresholding(method="dbr") returns NaN on the documented minimal input), 2 notes (the dbr.sd default documented three ways; a failed xgboost training silently keeps the previous score); the expected-doublet arithmetic, cxds2 (double-loop port), the thresholding rules, createDoublets' sums and the end-to-end calls (AUC 1.000, recall 0.81, no false positive) held up; testthat 17/17 on the patched build filing channel read (bug-report template, no CONTRIBUTING or AI policy; twelve prior threads read in full, none reports these; responsive maintainer); filing kit with three issue texts, three single-commit branches and PR bodies; fork of plger/scDblFinder needed
NumPy 253 (26 state a version: 1.19 ×15 mentions, 1.24 ×11, 1.26 ×9, 1.21 ×9; SciPy in 147, Matplotlib in 135, pandas in 69, scikit-learn in 64; image / array processing in 86, sorting in 43, normalisation in 40, random numbers in 29) 253 0 confirmed by executing the 2.4.6 release and 1.23.5, 1.24.4 and 1.26.4 against exact recomputations (no wrong number for a documented definition), 5 notes (float32 mean/sum/var along the sample axis use a plain running sum rather than pairwise summation — 2.2e-6 relative on a 2,000-frame stack, 2.9e-2 on a 5,000,000 × 2 matrix, open upstream as numpy/numpy#22956 and #8869; percentile(method="closest_observation") tie rule fixed in 2.1.0; multivariate_normal on a non-PSD covariance samples from V|Λ|Vᵀ; float32 FFT in single precision since 2.0; np.round(2.675, 2) = 2.68); every quantile method, cov/corrcoef, histogram, digitize, polyfit, interp, trapezoid, gradient, lstsq, the decompositions, the FFT and the generators held up round 6: nothing to file. Full inspection (2026-09-25, nine more harness groups over the whole library, ledger): 14 bugs on 2.4.6 still on main. Five were already reported upstream (MT19937 jumped, quantile next to ±inf, datetime wraparound, timedelta64 // int, ma.corrcoef; three have fix PRs open). Two are to file now: np.unique(datetime, equal_nan=False) merges NaT on the 2.3+ hash path (fix verified on a source build of main), and np.select wraps an out-of-range Python int (300 → 44 in int8). The rest are queued. The kit is fact sheets, because the AI policy forbids AI-written issue and PR text; nothing filed
scikit-learn 322 (31 state a version: 1.0.2 ×8, 0.21.3 ×6, 1.2.2 ×5, 0.24.2 ×4, 1.3.2 ×3, 1.5.x ×4, 1.6.x ×2, 1.7.2; PCA in 55, random forests and logistic regression in 23 each, hierarchical clustering in 20, cross-validation in 16, k-means in 13, NMF in 11; SciPy in 111, UMAP in 92) 322 2 confirmed by executing the 1.9.1 wheel and 1.1.3, 1.2.2, 1.3.2, 1.5.2 against exact recomputations, main (857849927d) read (PCA's covariance_eigh solver, the svd_solver="auto" choice for dense data with ≤ 1000 features and ≥ 10× as many samples since 1.5, forms the covariance as X.T @ X − n·mean⊗mean and loses the variance of features with a common offset: float32 values around 100 give explained_variance_ 33 % off, around 1e3 meaningless, float64 values around 1e6 1.7 % off, 1e8 meaningless, where full — the auto choice up to 1.4 — is exact; NMF.reconstruction_err_ documented as the beta-divergence is sqrt(2·D_β) for the KL and IS losses, 40.1 for a divergence of 804.6), 1 note open upstream (#31210: euclidean_distances' dot-product expansion at large offsets, silhouette 0.717 for 0.674 at 1e8) + 4 notes (LogisticRegression default tol, KBinsDiscretizer quantile method since 1.9, roc_auc_score with one class returns nan in 1.9, KMeans centres on a tol stop); every ranking / classification / regression / clustering metric, PCA(full/arpack/randomized), the scalers, KMeans, GaussianMixture, DBSCAN, agglomerative clustering vs scipy, the splitters, GridSearchCV, Ridge, LogisticRegression, RandomForest, LDA, KDE, the discretisers and transformers held up filing channel read (Automated Contributions Policy: no AI-generated issue or PR text, AI use disclosed in the PR, code PRs only after "Needs Triage" is removed; seven prior threads read in full incl. the solver's PR #27491, none reports SK1 or SK2); filing kit with two fact sheets, two single-commit branches (PCA fix verified: new test fails on main, test_pca.py 388 passed on the patched install; NMF docstring) and PR bodies; fork of scikit-learn/scikit-learn needed. Full inspection (2026-09-25, nine more harness groups, ledger): 26 more findings on 1.9.1. Six were already reported upstream (#25380, #14054/#24732, #30332, #26658, #11395, #22283). Five have fixes on branches (GaussianNB partial_fit, DummyRegressor weighted median, GPC restarts, isotonic with one distinct X, PLS intercept_). The rest are queued behind SK1/SK2, led by f_regression p = 1 for a feature equal to the target, the cosine-kernel KDE normalisation and TSNE mutating a precomputed matrix

The survey itself covers 44,198 research articles, of which 20,501 were openly readable; 10,364 name at least one open-source package, across 270 distinct packages. Coverage is very uneven by journal (PNAS 78%, Science 5%) — see survey/README.md for the caveats that come with that, which matter for any count in this repository.

Layout

survey/
  README.md      the survey document: headline numbers, per-journal coverage,
                 most-used packages, and the caveats
  data/          papers.tsv, paper_software.tsv (36,424 usage records),
                 package_index.tsv, gap_list.tsv (23,697 unreadable papers),
                 pipelines.jsonl.gz
  packages/      one page per detected package (270), with evidence sentences
  scripts/       harvest + extraction pipeline (reproducible end to end)

audits/
  freesurfer/    component reviews, paper→feature exposure, numerical
                 reproductions, and the upstream filing kit
  fsl/           component reviews and paper→feature exposure
                 (reports, patches and harnesses: separate repo, below)
  spm/           eleven component reviews spanning the whole codebase,
                 verification harnesses incl. an Octave old-vs-new
                 regression, and the upstream filing kit
  afni/          nine subsystem reviews, paper→feature exposure, the ReHo
                 tie-handling reproduction, and the upstream filing queue

preprint/        LaTeX source, figures and compiled PDF of the AFNI paper
                 (method + the ReHo reanalysis on ds000030)
preprint-spm/    LaTeX source, figure and compiled PDF of the SPM companion
                 paper (the 83 findings, the five upstream fixes, the two
                 merged ones measured on SPM's MMN tutorial data)

Related repositories

  • fsl-bug-reports — the FSL findings as ready-to-send reports, git format-patch fixes, and verification harnesses. Kept separate because FSL takes reports on its mailing list, and those posts link to it directly.
  • cindykrafft/freesurfer — fork carrying the five FreeSurfer fix branches submitted upstream.
  • cindykrafft/spm — fork carrying the five SPM fix branches plus a new tests/test_spm_ECdensity.m (two merged, three open upstream).
  • cindykrafft/scanpy — fork carrying the two Scanpy fix branches (t-test input scale, opened as PR #4337; score_genes bins, PR pending the reference-pickle regeneration) and the two Scrublet-port branches, each with a test and a release-note fragment.
  • cindykrafft/cutadapt, cindykrafft/umap, cindykrafft/CellphoneDB — forks carrying the fix branches from those passes (one PR each opened 2026-09-03: cutadapt #893, umap #1287; CellPhoneDB has an issue, #231, and two branches held).

Reading these findings fairly

  • A finding is code-level unless it says otherwise. "CONFIRMED" means the defect was traced end to end in the source and, where stated, verified numerically — not that anyone has run the shipped binary on imaging data and watched it happen.
  • "Paper is exposed" means the paper used the affected feature or version. It does not mean its conclusions are wrong. Most of these biases are systematic and shared across subjects, which is exactly why they usually shift or attenuate measurements rather than reverse well-powered contrasts.
  • Reproduction sometimes shrinks a finding, and that is recorded too — one FreeSurfer finding turned out not to fire on the standard pipeline at all once measured, and a faithful port of AFNI's ReHo tie loop moved one published paper back out of the exposed set.
  • Upstream maintainers' reading wins. Where they have adjudicated, the outcome is recorded next to the finding: SPM and AFNI have merged fixes and left others open; DESeq2's maintainer closed all four filings, and on the central factual point (whether the weights fix had been announced) he was right and this project was wrong.

Reproducing the survey

survey/scripts/ runs the whole pipeline against the public Europe PMC API — metadata harvest by ISSN, full-text fetch, then extraction. It needs no credentials. Expect the fetch stage to take many hours and to hit articles that are marked open access but serve no XML; those are routed to the gap list rather than retried indefinitely.

License

Scripts are MIT (see LICENSE). The prose, tables and derived data are CC BY 4.0. Evidence sentences are short quotations from the source articles, each attributed by PMID and DOI; the underlying articles remain under their own licenses. No third-party source code is redistributed here — patches and harnesses only.

Releases

Packages

Contributors

Languages