Describe the bug
When a Parquet file has bloom filters, the libcudf reader uses them to prune row groups for col == literal predicates. It hashes the literal as the cudf type of the output column. The Parquet spec hashes the value's plain encoding, which follows the physical type (BloomFilter.md). Arrow's writer does that (cpp/src/parquet/xxhasher.cc), and so do files written by DuckDB. When the two encodings differ, the probe misses and row groups that contain the value get pruned, so the read silently returns fewer rows, usually none:
- INT8, INT16, UINT8 and UINT16 columns are stored as INT32, but the literal is hashed as 1 or 2 bytes instead of 4.
- TIME(MILLIS) is stored as INT32 and read as DURATION_MILLISECONDS, so it's hashed as 8 bytes instead of 4.
- For FLOAT/DOUBLE,
x == 0.0 doesn't find rows holding -0.0, and the other way round. The two compare equal but hash differently.
Related failures on the same path:
- An equality on a BOOL column raises
Bloom filters do not support boolean or compound types (bloom_filter_reader.cu:161) when another column in the predicate has a bloom filter. cudf-polars hits this too, e.g. with .filter((pl.col("id") == 5) & (pl.col("flag") == True)).
- Through pylibcudf, every decimal equality raises
Mismatched predicate column and literal types (bloom_filter_reader.cu:88), whether the decimal is stored as INT32, INT64 or FIXED_LEN_BYTE_ARRAY. INT96 timestamps (Spark's default) raise the same error. cudf-polars doesn't push these literals down, so these two only affect pylibcudf and C++ callers.
DuckDB writes bloom filters by default for dictionary-encoded columns, so a plain COPY ... TO 'file.parquet' is enough to hit this. With cudf-polars, pl.scan_parquet(path).filter(pl.col("status") == 3).collect(engine="gpu") returns 0 rows for a TINYINT column where CPU Polars returns 285,714.
Steps/Code to reproduce bug
import decimal
import polars as pl
import pyarrow as pa
import pyarrow.parquet as pq
import pylibcudf as plc
from pylibcudf.expressions import ASTOperator, ColumnNameReference, Literal, Operation
# int8 column with a bloom filter (pyarrow >= 24); one row group, value within min/max
pq.write_table(
pa.table({"x": pa.array([i % 100 for i in range(10_000)], pa.int8())}),
"int8.parquet",
bloom_filter_options={"x": True},
)
q = pl.scan_parquet("int8.parquet").filter(pl.col("x") == 42)
print("int8 == 42 CPU:", q.collect().height, " GPU:", q.collect(engine="gpu").height)
# -0.0 in the data, filter on 0.0
pq.write_table(
pa.table({"y": pa.array([-0.0, 1.5, 2.5] * 1000)}),
"negzero.parquet",
bloom_filter_options={"y": True},
)
q = pl.scan_parquet("negzero.parquet").filter(pl.col("y") == 0.0)
print("y == 0.0 CPU:", q.collect().height, " GPU:", q.collect(engine="gpu").height)
# decimal equality through pylibcudf
pq.write_table(
pa.table({"d": pa.array([decimal.Decimal(i).scaleb(-2) for i in range(1000)], pa.decimal128(9, 2))}),
"dec.parquet",
bloom_filter_options={"d": True},
)
opts = plc.io.parquet.ParquetReaderOptions.builder(plc.io.SourceInfo(["dec.parquet"])).build()
expr = Operation(
ASTOperator.EQUAL,
ColumnNameReference("d"),
Literal(plc.Scalar.from_arrow(pa.scalar(decimal.Decimal("0.42"), pa.decimal32(9, 2)))),
)
opts.set_filter(expr)
try:
print("d == 0.42 rows:", plc.io.parquet.read_parquet(opts).tbl.num_rows())
except RuntimeError as e:
print("d == 0.42 raised:", e)
# DuckDB with default settings
import duckdb
duckdb.sql(
"COPY (SELECT (i % 7)::TINYINT AS status, (i % 50)::SMALLINT AS region "
"FROM range(2000000) t(i)) TO 'duck.parquet' (FORMAT parquet)"
)
for col, v in (("status", 3), ("region", 17)):
q = pl.scan_parquet("duck.parquet").filter(pl.col(col) == v)
print(f"duckdb {col} == {v} CPU:", q.collect().height, " GPU:", q.collect(engine="gpu").height)
Output on current main (nightly 26.12.00a199.post260929010848, commit 6b9c9bf):
int8 == 42 CPU: 100 GPU: 0
y == 0.0 CPU: 1000 GPU: 0
d == 0.42 raised: CUDF failure at: /__w/cudf/cudf/cpp/src/io/parquet/bloom_filter_reader.cu:88: Mismatched predicate column and literal types
duckdb status == 3 CPU: 285714 GPU: 0
duckdb region == 17 CPU: 40000 GPU: 0
Filters on int32, int64, uint32, uint64, non-zero floats, strings, dates and timestamps work. cudf.read_parquet(filters=...) isn't affected in practice, because it falls back to pyarrow filtering for Python scalars.
Expected behavior
The same rows as CPU Polars and pyarrow. A bloom filter should only prune a row group whose column chunk can't contain the value. When the reader can't probe a type (bool, decimals stored as BYTE_ARRAY), it should keep the row group.
Root cause
bloom_filter_caster in cpp/src/io/parquet/bloom_filter_reader.cu dispatches on the storage type of the output dtype and probes with XXHash_64<T> over sizeof(T) bytes. INT96 is the only physical type that gets converted first. The decimal and INT96 errors come from the type check, which compares the column with scalar_type_t of the storage or key type (an integer for decimals, string_view for INT96), so it never passes for them. The INT96 key itself would be wrong too: it's an int64 sign-extended to 12 bytes, while INT96 is nanoseconds since midnight followed by the Julian day. The existing tests don't catch this because libcudf can't write bloom filters (#23620) and cudf.read_parquet doesn't push these literals down.
Environment overview
- Environment location: Cloud (RunPod), NVIDIA RTX PRO 4500 Blackwell, driver 580.167.08
- Method of cuDF install: pip nightly wheels (
cudf-cu12, pylibcudf-cu12, cudf-polars-cu12 26.12.0a199.post260929010848), pyarrow 25.0.1, polars 1.44.2, duckdb 1.5.6
Additional context
I have a fix that converts the literal to the physical type's encoding before probing. It widens or narrows integers, decimals, times and timestamps to INT32 or INT64, encodes FIXED_LEN_BYTE_ARRAY decimals as big-endian bytes and INT96 as nanoseconds plus Julian day, probes both +0.0 and -0.0 for a zero literal, and keeps all row groups for bool. It comes with pylibcudf tests that write the files with pyarrow. Happy to open a PR.
Describe the bug
When a Parquet file has bloom filters, the libcudf reader uses them to prune row groups for
col == literalpredicates. It hashes the literal as the cudf type of the output column. The Parquet spec hashes the value's plain encoding, which follows the physical type (BloomFilter.md). Arrow's writer does that (cpp/src/parquet/xxhasher.cc), and so do files written by DuckDB. When the two encodings differ, the probe misses and row groups that contain the value get pruned, so the read silently returns fewer rows, usually none:x == 0.0doesn't find rows holding-0.0, and the other way round. The two compare equal but hash differently.Related failures on the same path:
Bloom filters do not support boolean or compound types(bloom_filter_reader.cu:161) when another column in the predicate has a bloom filter. cudf-polars hits this too, e.g. with.filter((pl.col("id") == 5) & (pl.col("flag") == True)).Mismatched predicate column and literal types(bloom_filter_reader.cu:88), whether the decimal is stored as INT32, INT64 or FIXED_LEN_BYTE_ARRAY. INT96 timestamps (Spark's default) raise the same error. cudf-polars doesn't push these literals down, so these two only affect pylibcudf and C++ callers.DuckDB writes bloom filters by default for dictionary-encoded columns, so a plain
COPY ... TO 'file.parquet'is enough to hit this. With cudf-polars,pl.scan_parquet(path).filter(pl.col("status") == 3).collect(engine="gpu")returns 0 rows for a TINYINT column where CPU Polars returns 285,714.Steps/Code to reproduce bug
Output on current main (nightly
26.12.00a199.post260929010848, commit 6b9c9bf):Filters on int32, int64, uint32, uint64, non-zero floats, strings, dates and timestamps work.
cudf.read_parquet(filters=...)isn't affected in practice, because it falls back to pyarrow filtering for Python scalars.Expected behavior
The same rows as CPU Polars and pyarrow. A bloom filter should only prune a row group whose column chunk can't contain the value. When the reader can't probe a type (bool, decimals stored as BYTE_ARRAY), it should keep the row group.
Root cause
bloom_filter_casterincpp/src/io/parquet/bloom_filter_reader.cudispatches on the storage type of the output dtype and probes withXXHash_64<T>oversizeof(T)bytes. INT96 is the only physical type that gets converted first. The decimal and INT96 errors come from the type check, which compares the column withscalar_type_tof the storage or key type (an integer for decimals,string_viewfor INT96), so it never passes for them. The INT96 key itself would be wrong too: it's an int64 sign-extended to 12 bytes, while INT96 is nanoseconds since midnight followed by the Julian day. The existing tests don't catch this because libcudf can't write bloom filters (#23620) andcudf.read_parquetdoesn't push these literals down.Environment overview
cudf-cu12,pylibcudf-cu12,cudf-polars-cu1226.12.0a199.post260929010848), pyarrow 25.0.1, polars 1.44.2, duckdb 1.5.6Additional context
I have a fix that converts the literal to the physical type's encoding before probing. It widens or narrows integers, decimals, times and timestamps to INT32 or INT64, encodes FIXED_LEN_BYTE_ARRAY decimals as big-endian bytes and INT96 as nanoseconds plus Julian day, probes both +0.0 and -0.0 for a zero literal, and keeps all row groups for bool. It comes with pylibcudf tests that write the files with pyarrow. Happy to open a PR.