Describe the bug
On a remote fsspec filesystem, read_parquet(..., prefetch_options={"method": "parquet"}) downloads the whole file instead of only the selected row groups and columns, and it rewrites the dict the caller passed in.
The cause is this line in python/cudf/cudf/io/parquet.py:
prefetch_options = prefetch_options.update({...})
dict.update returns None, so get_reader_filepath_or_buffer gets prefetch_options=None and _prefetch_remote_buffers falls back to method="all". The line came in with #16657, so the parquet prefetcher has never run from read_parquet.
Steps/Code to reproduce bug
import fsspec
import numpy as np
import pandas as pd
import cudf
from fsspec.implementations.memory import MemoryFileSystem
fs = MemoryFileSystem()
path = "memory://bucket/data.parquet"
pdf = pd.DataFrame({"a": np.arange(400_000), "b": np.arange(400_000) * 1.5})
with fs.open(path, "wb") as f:
pdf.to_parquet(f, row_group_size=100_000, index=False)
print("file size:", fs.size(path))
fetched = []
cat_ranges = MemoryFileSystem.cat_ranges
def counting_cat_ranges(self, *args, **kwargs):
out = cat_ranges(self, *args, **kwargs)
fetched.extend(len(b) for b in out)
return out
MemoryFileSystem.cat_ranges = counting_cat_ranges
opts = {"method": "parquet"}
df = cudf.read_parquet(path, row_groups=[1], prefetch_options=opts)
print("rows:", len(df), "bytes fetched:", sum(fetched))
print("prefetch_options after the call:", opts)
Output:
file size: 4857088
rows: 100000 bytes fetched: 4857088
prefetch_options after the call: {'method': 'parquet', 'columns': None, 'row_groups': [1]}
Expected behavior
The parquet prefetcher is used, so the whole file is not downloaded for one row group, and opts is left as {'method': 'parquet'}.
Environment overview (please complete the following information)
- Environment location: Bare-metal, Linux, RTX 5090 (driver 610.43.02)
- Method of cuDF install: pip, nightly
cudf-cu13 26.12.00a177.post260927053240 (commit 1c84b4e)
Environment details
- NVIDIA GeForce RTX 5090, driver 610.43.02, CUDA 13 (
cu13 wheels)
- Python 3.12.13, pandas 3.0.6, pyarrow 25.0.1, fsspec 2026.9.0
- The line is unchanged on
main at d75cb7f.
Additional context
I have a fix with tests ready and can open a PR. With it, the script above fetches 3,210,801 bytes and leaves opts unchanged. Replacing the update call with a new dict is not enough on its own: once the prefetcher runs, two more cases fail with Parquet header parsing failed on the file above, which is larger than fsspec's 1 MB footer sample:
columns=["b"], filters=[("a", ">", x)]: the filter columns are added to columns after the prefetch, so the bytes of a are never fetched. Moving that block before the prefetch fixes it.
case_sensitive_names=False with columns=["B"]: the prefetcher matches names exactly and fetches no column data. Passing columns=None to it in that case fixes it.
One question on the default. The same block is meant to pick "parquet" by default for a uniform row_groups selection, including one derived from filters, but because of this bug those reads have always used "all". Turning that default on is not a clear win: the parquet prefetcher reads a 1 MB footer sample and then fetches the footer range again with the data. On memory://, a 242,769-byte file read with row_groups=[0] goes from 242,769 bytes in one cat_ranges call with "all" to 485,538 bytes in two calls with "parquet". dask-cudf also passes per-fragment row_groups to cudf.read_parquet. Would you prefer to keep "all" as the default and only honour an explicit method="parquet", to enable the default as #16657 intended, or to remove the parquet prefetcher since the option is marked experimental?
Describe the bug
On a remote fsspec filesystem,
read_parquet(..., prefetch_options={"method": "parquet"})downloads the whole file instead of only the selected row groups and columns, and it rewrites the dict the caller passed in.The cause is this line in
python/cudf/cudf/io/parquet.py:dict.updatereturnsNone, soget_reader_filepath_or_buffergetsprefetch_options=Noneand_prefetch_remote_buffersfalls back tomethod="all". The line came in with #16657, so the parquet prefetcher has never run fromread_parquet.Steps/Code to reproduce bug
Output:
Expected behavior
The parquet prefetcher is used, so the whole file is not downloaded for one row group, and
optsis left as{'method': 'parquet'}.Environment overview (please complete the following information)
cudf-cu1326.12.00a177.post260927053240 (commit 1c84b4e)Environment details
cu13wheels)mainat d75cb7f.Additional context
I have a fix with tests ready and can open a PR. With it, the script above fetches 3,210,801 bytes and leaves
optsunchanged. Replacing theupdatecall with a new dict is not enough on its own: once the prefetcher runs, two more cases fail withParquet header parsing failedon the file above, which is larger than fsspec's 1 MB footer sample:columns=["b"], filters=[("a", ">", x)]: the filter columns are added tocolumnsafter the prefetch, so the bytes ofaare never fetched. Moving that block before the prefetch fixes it.case_sensitive_names=Falsewithcolumns=["B"]: the prefetcher matches names exactly and fetches no column data. Passingcolumns=Noneto it in that case fixes it.One question on the default. The same block is meant to pick
"parquet"by default for a uniformrow_groupsselection, including one derived fromfilters, but because of this bug those reads have always used"all". Turning that default on is not a clear win: the parquet prefetcher reads a 1 MB footer sample and then fetches the footer range again with the data. Onmemory://, a 242,769-byte file read withrow_groups=[0]goes from 242,769 bytes in onecat_rangescall with"all"to 485,538 bytes in two calls with"parquet". dask-cudf also passes per-fragmentrow_groupstocudf.read_parquet. Would you prefer to keep"all"as the default and only honour an explicitmethod="parquet", to enable the default as #16657 intended, or to remove the parquet prefetcher since the option is marked experimental?