Skip to content

[BUG] Parquet bloom filter pruning drops row groups that contain the value (int8/int16, TIME(MILLIS), -0.0) and fails for decimals #24319

Description

@kita-renji

Describe the bug

When a Parquet file has bloom filters, the libcudf reader uses them to prune row groups for col == literal predicates. It hashes the literal as the cudf type of the output column. The Parquet spec hashes the value's plain encoding, which follows the physical type (BloomFilter.md). Arrow's writer does that (cpp/src/parquet/xxhasher.cc), and so do files written by DuckDB. When the two encodings differ, the probe misses and row groups that contain the value get pruned, so the read silently returns fewer rows, usually none:

  • INT8, INT16, UINT8 and UINT16 columns are stored as INT32, but the literal is hashed as 1 or 2 bytes instead of 4.
  • TIME(MILLIS) is stored as INT32 and read as DURATION_MILLISECONDS, so it's hashed as 8 bytes instead of 4.
  • For FLOAT/DOUBLE, x == 0.0 doesn't find rows holding -0.0, and the other way round. The two compare equal but hash differently.

Related failures on the same path:

  • An equality on a BOOL column raises Bloom filters do not support boolean or compound types (bloom_filter_reader.cu:161) when another column in the predicate has a bloom filter. cudf-polars hits this too, e.g. with .filter((pl.col("id") == 5) & (pl.col("flag") == True)).
  • Through pylibcudf, every decimal equality raises Mismatched predicate column and literal types (bloom_filter_reader.cu:88), whether the decimal is stored as INT32, INT64 or FIXED_LEN_BYTE_ARRAY. INT96 timestamps (Spark's default) raise the same error. cudf-polars doesn't push these literals down, so these two only affect pylibcudf and C++ callers.

DuckDB writes bloom filters by default for dictionary-encoded columns, so a plain COPY ... TO 'file.parquet' is enough to hit this. With cudf-polars, pl.scan_parquet(path).filter(pl.col("status") == 3).collect(engine="gpu") returns 0 rows for a TINYINT column where CPU Polars returns 285,714.

Steps/Code to reproduce bug

import decimal

import polars as pl
import pyarrow as pa
import pyarrow.parquet as pq

import pylibcudf as plc
from pylibcudf.expressions import ASTOperator, ColumnNameReference, Literal, Operation

# int8 column with a bloom filter (pyarrow >= 24); one row group, value within min/max
pq.write_table(
    pa.table({"x": pa.array([i % 100 for i in range(10_000)], pa.int8())}),
    "int8.parquet",
    bloom_filter_options={"x": True},
)
q = pl.scan_parquet("int8.parquet").filter(pl.col("x") == 42)
print("int8  == 42  CPU:", q.collect().height, " GPU:", q.collect(engine="gpu").height)

# -0.0 in the data, filter on 0.0
pq.write_table(
    pa.table({"y": pa.array([-0.0, 1.5, 2.5] * 1000)}),
    "negzero.parquet",
    bloom_filter_options={"y": True},
)
q = pl.scan_parquet("negzero.parquet").filter(pl.col("y") == 0.0)
print("y == 0.0     CPU:", q.collect().height, " GPU:", q.collect(engine="gpu").height)

# decimal equality through pylibcudf
pq.write_table(
    pa.table({"d": pa.array([decimal.Decimal(i).scaleb(-2) for i in range(1000)], pa.decimal128(9, 2))}),
    "dec.parquet",
    bloom_filter_options={"d": True},
)
opts = plc.io.parquet.ParquetReaderOptions.builder(plc.io.SourceInfo(["dec.parquet"])).build()
expr = Operation(
    ASTOperator.EQUAL,
    ColumnNameReference("d"),
    Literal(plc.Scalar.from_arrow(pa.scalar(decimal.Decimal("0.42"), pa.decimal32(9, 2)))),
)
opts.set_filter(expr)
try:
    print("d == 0.42 rows:", plc.io.parquet.read_parquet(opts).tbl.num_rows())
except RuntimeError as e:
    print("d == 0.42 raised:", e)

# DuckDB with default settings
import duckdb

duckdb.sql(
    "COPY (SELECT (i % 7)::TINYINT AS status, (i % 50)::SMALLINT AS region "
    "FROM range(2000000) t(i)) TO 'duck.parquet' (FORMAT parquet)"
)
for col, v in (("status", 3), ("region", 17)):
    q = pl.scan_parquet("duck.parquet").filter(pl.col(col) == v)
    print(f"duckdb {col} == {v}  CPU:", q.collect().height, " GPU:", q.collect(engine="gpu").height)

Output on current main (nightly 26.12.00a199.post260929010848, commit 6b9c9bf):

int8  == 42  CPU: 100  GPU: 0
y == 0.0     CPU: 1000  GPU: 0
d == 0.42 raised: CUDF failure at: /__w/cudf/cudf/cpp/src/io/parquet/bloom_filter_reader.cu:88: Mismatched predicate column and literal types
duckdb status == 3  CPU: 285714  GPU: 0
duckdb region == 17  CPU: 40000  GPU: 0

Filters on int32, int64, uint32, uint64, non-zero floats, strings, dates and timestamps work. cudf.read_parquet(filters=...) isn't affected in practice, because it falls back to pyarrow filtering for Python scalars.

Expected behavior

The same rows as CPU Polars and pyarrow. A bloom filter should only prune a row group whose column chunk can't contain the value. When the reader can't probe a type (bool, decimals stored as BYTE_ARRAY), it should keep the row group.

Root cause

bloom_filter_caster in cpp/src/io/parquet/bloom_filter_reader.cu dispatches on the storage type of the output dtype and probes with XXHash_64<T> over sizeof(T) bytes. INT96 is the only physical type that gets converted first. The decimal and INT96 errors come from the type check, which compares the column with scalar_type_t of the storage or key type (an integer for decimals, string_view for INT96), so it never passes for them. The INT96 key itself would be wrong too: it's an int64 sign-extended to 12 bytes, while INT96 is nanoseconds since midnight followed by the Julian day. The existing tests don't catch this because libcudf can't write bloom filters (#23620) and cudf.read_parquet doesn't push these literals down.

Environment overview

  • Environment location: Cloud (RunPod), NVIDIA RTX PRO 4500 Blackwell, driver 580.167.08
  • Method of cuDF install: pip nightly wheels (cudf-cu12, pylibcudf-cu12, cudf-polars-cu12 26.12.0a199.post260929010848), pyarrow 25.0.1, polars 1.44.2, duckdb 1.5.6

Additional context

I have a fix that converts the literal to the physical type's encoding before probing. It widens or narrows integers, decimals, times and timestamps to INT32 or INT64, encodes FIXED_LEN_BYTE_ARRAY decimals as big-endian bytes and INT96 as nanoseconds plus Julian day, probes both +0.0 and -0.0 for a zero literal, and keeps all row groups for bool. It comes with pylibcudf tests that write the files with pyarrow. Happy to open a PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions