Skip to content

[BUG] read_parquet ignores prefetch_options={"method": "parquet"} for remote files #24313

Description

@Arthur031221

Describe the bug
On a remote fsspec filesystem, read_parquet(..., prefetch_options={"method": "parquet"}) downloads the whole file instead of only the selected row groups and columns, and it rewrites the dict the caller passed in.

The cause is this line in python/cudf/cudf/io/parquet.py:

prefetch_options = prefetch_options.update({...})

dict.update returns None, so get_reader_filepath_or_buffer gets prefetch_options=None and _prefetch_remote_buffers falls back to method="all". The line came in with #16657, so the parquet prefetcher has never run from read_parquet.

Steps/Code to reproduce bug

import fsspec
import numpy as np
import pandas as pd
import cudf
from fsspec.implementations.memory import MemoryFileSystem

fs = MemoryFileSystem()
path = "memory://bucket/data.parquet"
pdf = pd.DataFrame({"a": np.arange(400_000), "b": np.arange(400_000) * 1.5})
with fs.open(path, "wb") as f:
    pdf.to_parquet(f, row_group_size=100_000, index=False)
print("file size:", fs.size(path))

fetched = []
cat_ranges = MemoryFileSystem.cat_ranges
def counting_cat_ranges(self, *args, **kwargs):
    out = cat_ranges(self, *args, **kwargs)
    fetched.extend(len(b) for b in out)
    return out
MemoryFileSystem.cat_ranges = counting_cat_ranges

opts = {"method": "parquet"}
df = cudf.read_parquet(path, row_groups=[1], prefetch_options=opts)
print("rows:", len(df), "bytes fetched:", sum(fetched))
print("prefetch_options after the call:", opts)

Output:

file size: 4857088
rows: 100000 bytes fetched: 4857088
prefetch_options after the call: {'method': 'parquet', 'columns': None, 'row_groups': [1]}

Expected behavior
The parquet prefetcher is used, so the whole file is not downloaded for one row group, and opts is left as {'method': 'parquet'}.

Environment overview (please complete the following information)

  • Environment location: Bare-metal, Linux, RTX 5090 (driver 610.43.02)
  • Method of cuDF install: pip, nightly cudf-cu13 26.12.00a177.post260927053240 (commit 1c84b4e)

Environment details

  • NVIDIA GeForce RTX 5090, driver 610.43.02, CUDA 13 (cu13 wheels)
  • Python 3.12.13, pandas 3.0.6, pyarrow 25.0.1, fsspec 2026.9.0
  • The line is unchanged on main at d75cb7f.

Additional context
I have a fix with tests ready and can open a PR. With it, the script above fetches 3,210,801 bytes and leaves opts unchanged. Replacing the update call with a new dict is not enough on its own: once the prefetcher runs, two more cases fail with Parquet header parsing failed on the file above, which is larger than fsspec's 1 MB footer sample:

  • columns=["b"], filters=[("a", ">", x)]: the filter columns are added to columns after the prefetch, so the bytes of a are never fetched. Moving that block before the prefetch fixes it.
  • case_sensitive_names=False with columns=["B"]: the prefetcher matches names exactly and fetches no column data. Passing columns=None to it in that case fixes it.

One question on the default. The same block is meant to pick "parquet" by default for a uniform row_groups selection, including one derived from filters, but because of this bug those reads have always used "all". Turning that default on is not a clear win: the parquet prefetcher reads a 1 MB footer sample and then fetches the footer range again with the data. On memory://, a 242,769-byte file read with row_groups=[0] goes from 242,769 bytes in one cat_ranges call with "all" to 485,538 bytes in two calls with "parquet". dask-cudf also passes per-fragment row_groups to cudf.read_parquet. Would you prefer to keep "all" as the default and only honour an explicit method="parquet", to enable the default as #16657 intended, or to remove the parquet prefetcher since the option is marked experimental?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions