Skip to content

Commit 2076361

Browse files
GH-533: Add ALP (Adaptive Lossless floating-Point) encoding specification (#557)
* Add ALP (Adaptive Lossless floating-Point) encoding specification Add the encoding specification for ALP (encoding value 10) to Encodings.md. ALP compresses FLOAT and DOUBLE columns by converting values to integers via decimal scaling, then applying Frame of Reference encoding and bit-packing. Values that cannot be losslessly round-tripped are stored as exceptions. The spec covers: - Page layout: 7-byte header, offset array, compressed vectors - Vector format: AlpInfo, ForInfo, packed values, exception data - Encoding math: two-step multiplication for cross-language consistency - Parameter selection, exception detection, and decoding steps Based on the paper "ALP: Adaptive Lossless floating-Point Compression" (Afroozeh and Boncz, SIGMOD 2024). Wire format matches the C++ Arrow and Java parquet-java implementations. * Address review feedback on ALP encoding specification Incorporate review comments from emkornfield and alamb on PR #557: * Address remaining review feedback on ALP spec - Clarify no padding between vectors in offset array description - Use 'sizeof(encoded type) (float=4 and double=8)' per reviewer suggestion * Address alamb's second review: trim spec, fix wording, rework example - Remove Characteristics, Size Calculations, Constants Reference sections - Consolidate three examples into one worked example with f!=0 and exceptions - Remove incorrect sign-reversal claim for fast_round on negative values - Soften sampling recommendation from SHOULD to suggestion - Fix "individual vectors" → "individual values" for random access - Clarify power-of-10 interop as MUST requirement - Use consistent fast_round terminology throughout * Address review: clarify power-of-10 constants and add fast_round sign branching - Specify that power-of-10 constants must match IEEE 754 correctly-rounded decimal-to-binary conversion of literals (§5.12.2), not runtime pow() - Add negative-value branch to fast_round formula table to avoid landing in a binade where ULP < 1.0, which causes unnecessary exceptions * Address review: clarify parameter selection is encoder-only optimization Any valid (e,f) pair produces correct output — the decoder is agnostic. Reframe "minimize exceptions" as one heuristic; the actual target is smallest encoded size (bit-width + exception overhead). * Address review: fix out-of-range condition to reference correct integer types The example incorrectly referenced only INT32_MAX for all types. Reworded to specify int32 for FLOAT and int64 for DOUBLE, avoiding exact numeric limits that differ from INT_MAX due to float representability. * Address review: mark encoder pipeline as informative, clarify example and layout * Clarify fast_round arithmetic precision in ALP spec State explicitly that the magic-number add/subtract are floating-point operations (only the final cast is integer), and that the constants are written as integers merely because they are exact. Add a MUST requiring FLOAT fast_round to be evaluated in single precision and DOUBLE in double precision, since a promotion of the FLOAT path to double drops the value into a binade with ULP below 1.0, defeating the rounding entirely. * Use 2^n magic constant for ALP fast_round instead of 1.5*2^n The 1.5*2^n midpoint constant exists to make the magic-number rounding branchless, covering a symmetric range in a single formula. Since this spec branches on sign, the midpoint is the wrong choice: with a branch, 2^n doubles the usable domain (to +/-2^23 for FLOAT and +/-2^52 for DOUBLE) and makes the branch actually necessary, whereas with the midpoint the branch was redundant. Update the constants, the domain description, and the prose (the technique is no longer branchless). * Address review: add ALP=10 thrift enum and clarify spec Resolve open review comments on PR #557 (GH-533): - parquet.thrift: add missing Encoding.ALP = 10 (thread 50). The enum value was never added; prose and the summary table said "ALP = 10" but the machine-readable contract lacked it. - Fix paper authorship to three authors: Afroozeh, Kuffo, Boncz (62). - "For each data page..." intro, drop misleading "page-level header" (16/49). - num_elements is the non-null value count (56). - Correct vector offset base for compressed / rep-def pages: use alp_data_start (decoded page data) instead of page_data_start+7 (55). - Note Vector Format applies to compression_mode=0 / integer_encoding=0 (59). - Label vector header vs data section; forward-ref num_exceptions/bit_width (60/61). - Clarify deltas are non-negative/unsigned with no sign extension (64). - Exceptions stored as raw IEEE 754 bits, MUST NOT canonicalize NaN (63). - Header byte-diagram alignment and "parallel encoding/decoding" wording (49/52). Leaves the contested fast_round sign-branching (threads 31/47/48) untouched pending reviewer consensus. * Address review: pipeline wording, unsigned deltas, vector diagram Follow-up to review comments on PR #557 (GH-533): - Pipeline step 1 wording: "for this array" to match the diagram's "Input: float/double array" (thread 54). - Clarify FOR deltas are computed and stored as unsigned integers, explicitly to avoid signed overflow when max-min exceeds the encoded type's signed maximum; no sign extension on unpack (thread 64). - Annotate the Vector Format diagram with a bracket row marking the vector header vs the data section (thread 61). * Mark ALP fast_round rounding as informative * Address review: raw-bit exceptions, fast_round reference wording * Address review: trim informative Fast Rounding section --------- Co-authored-by: Prateek Gaur <prateek.gaur@snowflake.com>
1 parent c6a6967 commit 2076361

2 files changed

Lines changed: 463 additions & 1 deletion

File tree

0 commit comments

Comments
 (0)