Skip to content

PERF: ~1.5x faster denoise_image on MPS/CPU via separable box-sum filter - #70

Merged
ntustison merged 1 commit into
mainfrom
perf/denoise-box-filter-mps
Sep 29, 2026
Merged

ntustison merged 1 commit into
mainfrom
perf/denoise-box-filter-mps

Conversation

@stnava

@stnava stnava commented Sep 28, 2026

Copy link
Copy Markdown
Member

Summary

  • Replaces F.conv2d/F.conv3d-based box-sum filtering in antstorch.denoise_image's core SANLM loop with a new _box_filter_sum() helper that computes the same box-sum via direct sliding-window addition (sum of 2r+1 shifted views per axis).
  • The box-sum "convolution" (an all-ones kernel) is separable, so this is an exact reformulation, not an approximation. A cumsum/prefix-sum version was tried first and rejected — it subtracts two large accumulated sums and loses float32 precision; direct sliding-window summation avoids that cancellation.
  • Bundles the existing benchmark script enhancements in tools/benchmarks/compare_denoise_image.py (pipeline warm-up timing, per-run timing breakdown, HTML comparison report) that were used throughout to profile and validate this change against the ANTs reference.

Why

Profiling antstorch.denoise_image on a real 136x176x176 volume (r=2, p=1, Rician) showed the core SANLM loop's per-search-offset box-sum convolution (F.conv3d with a 3x3x3 ones kernel, called 124 times for the default search radius) was the dominant cost — PyTorch's MPS backend has very high per-call overhead for conv3d at this kernel/volume size (~40ms/call), independent of the actual FLOP count.

Results

Benchmarked with ants.denoise_image as reference (--ants-threads 1 for a deterministic single-threaded baseline):

Before After
ANTsTorch runtime (MPS) 10.78s 7.0s
Speed ratio (ANTs / ANTsTorch) 2.48x 3.77x
RMSE vs. ANTs 0.00836083 0.00836074
Correlation vs. ANTs 0.9999886049 0.9999886052
Run-to-run determinism exact exact

CPU also benefits from the same change. Accuracy is unchanged (differences are float32 noise-level, well within existing ANTs vs. ANTsTorch parity).

Test plan

  • pytest tests/test_denoise_image.py — 13/13 passed
  • tools/benchmarks/compare_denoise_image.py against ants.denoise_image on a real 3-D volume (MPS and CPU), confirming speedup and unchanged accuracy/determinism vs. the ANTs reference

…noise_image

antstorch.denoise_image's core SANLM loop used F.conv2d/conv3d with an
all-ones kernel to compute local box-sums (mean/variance, and the per
search-offset patch-similarity sum). PyTorch's MPS conv3d backend has
very high per-call overhead for this kernel/volume size (~40ms/call),
and this convolution is invoked once per search offset (124x for the
default r=2), making it the dominant cost of the whole filter.

Since the "convolution" is just a box-sum (all-ones kernel), it is
separable: replace it with a new _box_filter_sum() that sums 2r+1
shifted tensor views per axis directly, with no convolution op. A
cumsum/prefix-sum formulation was tried first but rejected: it
differences two large accumulated sums and loses float32 precision
(subtly larger RMSE vs. the ANTs reference); direct sliding-window
summation avoids that cancellation entirely.

Benchmarked against ants.denoise_image on a 136x176x176 volume
(r=2, p=1, Rician), --ants-threads 1 for a deterministic reference:
  - MPS:  10.78s -> 7.0s  (ANTs/ANTsTorch speed ratio 2.48x -> 3.77x)
  - CPU:  also improved
  - Accuracy vs. ANTs unchanged (RMSE 0.00836, correlation 0.99999),
    output remains bit-exact deterministic run-to-run.

Also includes the benchmark script's existing pipeline warm-up,
per-run timing breakdown, and HTML comparison report generation
(tools/benchmarks/compare_denoise_image.py), used throughout this
work to profile and validate the change against the ANTs reference.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@ntustison

Copy link
Copy Markdown
Member

Brilliant. Thanks @stnava

@ntustison
ntustison merged commit da6d7e4 into main Sep 29, 2026
3 checks passed
@ntustison
ntustison deleted the perf/denoise-box-filter-mps branch September 29, 2026 00:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants