Implement thread-parallel compression for vtk_grid - #168
Conversation
|
@jipolanco ready for review |
|
This seems really great, thank you! I'm unavailable this week for reviewing the code but I'll take a look at this early next week. |
|
@jipolanco The issue in my first approach is that we get different block sizes on different hardware/number of threads, and therefore different files on different machines for the same data. I now choose the block size only based on field size. Trying to produce many blocks as soon as possible for small files, then scaling up block size to a reasonable 1MiB, then again increasing block size up to 1000 to get good scaling for huge CPUs. After reaching 1000 blocks, cap it at that and increase block size. The result is this:
Does that make sense? I have no experience with Zlib compression and block sizes. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #168 +/- ##
==========================================
- Coverage 96.93% 96.85% -0.08%
==========================================
Files 15 15
Lines 880 922 +42
==========================================
+ Hits 853 893 +40
- Misses 27 29 +2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
This looks seriously great, thanks. The proposed heuristic to make things independent of the number of threads looks good for me. As a very minor suggestion, perhaps we could make the maximum number of blocks a power of 2 (e.g. 1024). This may (marginally) improve parallel performance (load balancing) when the number of threads is also a power of 2 (which can be quite common). |
|
Done. |
When we run large simulations, we do so on machines with many CPU threads. We noticed that writing large VTK files takes a significant amount of time compared to the runtime of the actual simulation on these machines. We found that the culprit is not the actual write to disk but the compression that is applied by default.
The VTK format supports splitting appended data into multiple compressed blocks, but WriteVTK.jl currently uses a single block. This PR uses multiple blocks and compresses them in parallel, providing a significant speedup on large CPUs.
Benchmark
The speedups below are computed on the compression cost itself: I subtracted the corresponding
compress=falseruntime first and then compare serial vs parallel compression.Intel Xeon w9-3475X, 36 threads,
256^3Float32 data2 x AMD Epyc 9654, 192 threads,
512^3Float32 dataBenchmark code
Real-life benchmark
The motivation for this change was a postprocessing workflow for a simulation that took 14 hours on a large data center GPU. The postprocessing step interpolates simulation data and writes 500 VTK files.
With serial compression, this was estimated to take about 23 hours on the same machine, with compression as the bottleneck because interpolation is already parallel. With parallel compression, the postprocessing time is reduced to about 1.5 hours.
The benchmark below writes one representative file on 72 threads. Interpolation before the VTK export takes 3.25 s on average across these runs, so the VTK export is clearly the bottleneck with serial compression. For the parallel runs, I also reported the runtime of just the blockwise parallel compression to demonstrate that the compression is not the dominant part anymore (when enough threads are used).