Skip to content

Implement thread-parallel compression for vtk_grid - #168

Merged
jipolanco merged 10 commits into
JuliaVTK:masterfrom
efaulhaber:parallel-compression
Jun 22, 2026
Merged

jipolanco merged 10 commits into
JuliaVTK:masterfrom
efaulhaber:parallel-compression

Conversation

@efaulhaber

@efaulhaber efaulhaber commented Jun 16, 2026 •

Copy link
Copy Markdown
Contributor

When we run large simulations, we do so on machines with many CPU threads. We noticed that writing large VTK files takes a significant amount of time compared to the runtime of the actual simulation on these machines. We found that the culprit is not the actual write to disk but the compression that is applied by default.

The VTK format supports splitting appended data into multiple compressed blocks, but WriteVTK.jl currently uses a single block. This PR uses multiple blocks and compresses them in parallel, providing a significant speedup on large CPUs.

Benchmark

The speedups below are computed on the compression cost itself: I subtracted the corresponding compress=false runtime first and then compare serial vs parallel compression.

Intel Xeon w9-3475X, 36 threads, 256^3 Float32 data

Compression level Parallel time Serial time File size Compression speedup
none 38.8 ms 37.3 ms 64.0 MiB -
1 156 ms 1.74 s 57.6 MiB 14.5x
3 181 ms 1.93 s 57.5 MiB 13.3x
6 (default) 173 ms 2.08 s 57.3 MiB 15.3x
9 185 ms 2.08 s 57.3 MiB 13.9x

2 x AMD Epyc 9654, 192 threads, 512^3 Float32 data

Compression level Parallel time Serial time File size Compression speedup
none 151 ms 192 ms 512.0 MiB -
1 752 ms 16.3 s 460.5 MiB 26.9x
3 758 ms 17.7 s 460.2 MiB 28.8x
6 (default) 709 ms 19.3 s 458.1 MiB 34.2x
9 759 ms 19.3 s 458.1 MiB 31.4x
Benchmark code
using BenchmarkTools
using Printf
using WriteVTK

function write_file(path, data, compress, parallel_compression)
    vtk_grid(path, size(data)...; compress, parallel_compression) do vtk
        vtk["values"] = data
    end
end

compress_vec = (false, 1, 3, 6, 9)
dims = (512, 512, 512)

data = rand(Float32, dims)

println("mode,compression_level,seconds,bytes")

mktempdir() do dir
    for parallel_compression in (true, false)
        mode = parallel_compression ? "parallel" : "serial"
        for compress in compress_vec
            path = joinpath(dir, "compression_$(compress)_parallel_$(parallel_compression)")
            seconds = @belapsed write_file($path, $data, $compress, $parallel_compression)
            bytes = filesize(path * ".vti")
            @printf("%s,%d,%.6f,%d\n", mode, compress, seconds, bytes)
        end
    end
end

Real-life benchmark

The motivation for this change was a postprocessing workflow for a simulation that took 14 hours on a large data center GPU. The postprocessing step interpolates simulation data and writes 500 VTK files.

With serial compression, this was estimated to take about 23 hours on the same machine, with compression as the bottleneck because interpolation is already parallel. With parallel compression, the postprocessing time is reduced to about 1.5 hours.

The benchmark below writes one representative file on 72 threads. Interpolation before the VTK export takes 3.25 s on average across these runs, so the VTK export is clearly the bottleneck with serial compression. For the parallel runs, I also reported the runtime of just the blockwise parallel compression to demonstrate that the compression is not the dominant part anymore (when enough threads are used).

Compression level Serial VTK export Parallel VTK export Parallel compression File size VTK export speedup
none 4.04 s - - 5,042.7 MiB -
1 77.1 s 4.23 s 1.38 s 1,972.3 MiB 18.2x
3 90.3 s 4.47 s 1.50 s 1,971.9 MiB 20.2x
6 (default) 155 s 5.44 s 2.51 s 1,960.4 MiB 28.6x

@efaulhaber
efaulhaber marked this pull request as ready for review June 17, 2026 16:53
@efaulhaber

Copy link
Copy Markdown
Contributor Author

@jipolanco ready for review

@jipolanco

Copy link
Copy Markdown
Member

This seems really great, thank you! I'm unavailable this week for reviewing the code but I'll take a look at this early next week.

@efaulhaber

Copy link
Copy Markdown
Contributor Author

@jipolanco The issue in my first approach is that we get different block sizes on different hardware/number of threads, and therefore different files on different machines for the same data. I now choose the block size only based on field size. Trying to produce many blocks as soon as possible for small files, then scaling up block size to a reasonable 1MiB, then again increasing block size up to 1000 to get good scaling for huge CPUs. After reaching 1000 blocks, cap it at that and increase block size.

The result is this:

Input size Block size Number of blocks
2 MiB 128 KiB 16
10 MiB 128 KiB 80
25 MiB 256 KiB 100
100 MiB 1 MiB 100
500 MiB 1 MiB 500
1 GiB ~1 MiB ~1000
10 GiB ~10 MiB 1000
100 GiB ~102 MiB 1000

Does that make sense? I have no experience with Zlib compression and block sizes.

@codecov

codecov Bot commented Jun 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.49123% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 96.85%. Comparing base (384a53a) to head (7b8158c).

Files with missing lines Patch % Lines
src/write_data.jl 96.42% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master     #168      +/-   ##
==========================================
- Coverage   96.93%   96.85%   -0.08%     
==========================================
  Files          15       15              
  Lines         880      922      +42     
==========================================
+ Hits          853      893      +40     
- Misses         27       29       +2     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jipolanco

Copy link
Copy Markdown
Member

This looks seriously great, thanks.

The proposed heuristic to make things independent of the number of threads looks good for me. As a very minor suggestion, perhaps we could make the maximum number of blocks a power of 2 (e.g. 1024). This may (marginally) improve parallel performance (load balancing) when the number of threads is also a power of 2 (which can be quite common).

@efaulhaber

Copy link
Copy Markdown
Contributor Author

Done.

@jipolanco
jipolanco merged commit 457bbee into JuliaVTK:master Jun 22, 2026
5 of 7 checks passed
jipolanco referenced this pull request Jun 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants