Commit 2076361
* Add ALP (Adaptive Lossless floating-Point) encoding specification
Add the encoding specification for ALP (encoding value 10) to Encodings.md.
ALP compresses FLOAT and DOUBLE columns by converting values to integers via
decimal scaling, then applying Frame of Reference encoding and bit-packing.
Values that cannot be losslessly round-tripped are stored as exceptions.
The spec covers:
- Page layout: 7-byte header, offset array, compressed vectors
- Vector format: AlpInfo, ForInfo, packed values, exception data
- Encoding math: two-step multiplication for cross-language consistency
- Parameter selection, exception detection, and decoding steps
Based on the paper "ALP: Adaptive Lossless floating-Point Compression"
(Afroozeh and Boncz, SIGMOD 2024). Wire format matches the C++ Arrow
and Java parquet-java implementations.
* Address review feedback on ALP encoding specification
Incorporate review comments from emkornfield and alamb on PR #557:
* Address remaining review feedback on ALP spec
- Clarify no padding between vectors in offset array description
- Use 'sizeof(encoded type) (float=4 and double=8)' per reviewer suggestion
* Address alamb's second review: trim spec, fix wording, rework example
- Remove Characteristics, Size Calculations, Constants Reference sections
- Consolidate three examples into one worked example with f!=0 and exceptions
- Remove incorrect sign-reversal claim for fast_round on negative values
- Soften sampling recommendation from SHOULD to suggestion
- Fix "individual vectors" → "individual values" for random access
- Clarify power-of-10 interop as MUST requirement
- Use consistent fast_round terminology throughout
* Address review: clarify power-of-10 constants and add fast_round sign branching
- Specify that power-of-10 constants must match IEEE 754 correctly-rounded
decimal-to-binary conversion of literals (§5.12.2), not runtime pow()
- Add negative-value branch to fast_round formula table to avoid landing
in a binade where ULP < 1.0, which causes unnecessary exceptions
* Address review: clarify parameter selection is encoder-only optimization
Any valid (e,f) pair produces correct output — the decoder is agnostic.
Reframe "minimize exceptions" as one heuristic; the actual target is
smallest encoded size (bit-width + exception overhead).
* Address review: fix out-of-range condition to reference correct integer types
The example incorrectly referenced only INT32_MAX for all types. Reworded
to specify int32 for FLOAT and int64 for DOUBLE, avoiding exact numeric
limits that differ from INT_MAX due to float representability.
* Address review: mark encoder pipeline as informative, clarify example and layout
* Clarify fast_round arithmetic precision in ALP spec
State explicitly that the magic-number add/subtract are floating-point
operations (only the final cast is integer), and that the constants are
written as integers merely because they are exact. Add a MUST requiring
FLOAT fast_round to be evaluated in single precision and DOUBLE in double
precision, since a promotion of the FLOAT path to double drops the value
into a binade with ULP below 1.0, defeating the rounding entirely.
* Use 2^n magic constant for ALP fast_round instead of 1.5*2^n
The 1.5*2^n midpoint constant exists to make the magic-number rounding
branchless, covering a symmetric range in a single formula. Since this
spec branches on sign, the midpoint is the wrong choice: with a branch,
2^n doubles the usable domain (to +/-2^23 for FLOAT and +/-2^52 for
DOUBLE) and makes the branch actually necessary, whereas with the
midpoint the branch was redundant. Update the constants, the domain
description, and the prose (the technique is no longer branchless).
* Address review: add ALP=10 thrift enum and clarify spec
Resolve open review comments on PR #557 (GH-533):
- parquet.thrift: add missing Encoding.ALP = 10 (thread 50). The enum
value was never added; prose and the summary table said "ALP = 10"
but the machine-readable contract lacked it.
- Fix paper authorship to three authors: Afroozeh, Kuffo, Boncz (62).
- "For each data page..." intro, drop misleading "page-level header" (16/49).
- num_elements is the non-null value count (56).
- Correct vector offset base for compressed / rep-def pages: use
alp_data_start (decoded page data) instead of page_data_start+7 (55).
- Note Vector Format applies to compression_mode=0 / integer_encoding=0 (59).
- Label vector header vs data section; forward-ref num_exceptions/bit_width (60/61).
- Clarify deltas are non-negative/unsigned with no sign extension (64).
- Exceptions stored as raw IEEE 754 bits, MUST NOT canonicalize NaN (63).
- Header byte-diagram alignment and "parallel encoding/decoding" wording (49/52).
Leaves the contested fast_round sign-branching (threads 31/47/48)
untouched pending reviewer consensus.
* Address review: pipeline wording, unsigned deltas, vector diagram
Follow-up to review comments on PR #557 (GH-533):
- Pipeline step 1 wording: "for this array" to match the diagram's
"Input: float/double array" (thread 54).
- Clarify FOR deltas are computed and stored as unsigned integers,
explicitly to avoid signed overflow when max-min exceeds the encoded
type's signed maximum; no sign extension on unpack (thread 64).
- Annotate the Vector Format diagram with a bracket row marking the
vector header vs the data section (thread 61).
* Mark ALP fast_round rounding as informative
* Address review: raw-bit exceptions, fast_round reference wording
* Address review: trim informative Fast Rounding section
---------
Co-authored-by: Prateek Gaur <prateek.gaur@snowflake.com>
1 parent c6a6967 commit 2076361
2 files changed
Lines changed: 463 additions & 1 deletion
0 commit comments