Skip to content

[BUG] ORC reader applies the first stripe's writer timezone to every stripe #24315

Description

@vuule

Describe the bug
The ORC reader resolves the writer timezone from the first selected stripe only (reader_impl_chunking.cu, selected_stripes[0].stripe_footer->writerTimezone) and applies it to every stripe in the read. Both the transition table used to convert timestamps and the base epoch used for the negative timestamp borrow (#24259, #24265) come from that one stripe. The writerTimezone of the other stripes is parsed but never checked, so a mismatch is not detected or reported.

A single file normally has one writer timezone, but a multi-source read aggregates stripes from several files, which can be written in different timezones. In that case:

  • Timezone honored: every stripe is converted with the first file's offsets, so timestamps from the other files are shifted by the difference between the zones.
  • Timezone ignored (ignore_timezone_in_stripe_footer): values stay on each file's wall clock, but the borrow is decided in the first file's frame. Timestamps within the timezone offset of 1970-01-01 that have a fractional part can come back one second off. This includes a UTC file read after a non-UTC one, which decoded correctly before Decide the ORC timestamp borrow in the writer's timezone #24265.

Found by code inspection while addressing review on #24265; not yet reproduced.

Steps/Code to reproduce bug

  1. Write two ORC files with timestamp columns, one with writer timezone Asia/Shanghai and one with UTC (or any two different zones).
  2. Read both in a single cudf::io::read_orc call, listing the Shanghai file first, with ignore_timezone_in_stripe_footer both unset and set.
  3. Compare against reading each file on its own.

Expected behavior
Open question, to confirm with the Spark team before deciding:

  • Resolve each distinct writer timezone once and have each stripe use its own transition table and base epoch, so a multi-source read matches reading each file separately.
  • Or reject mixed writer timezones in one read, or warn about them, if resolving per stripe isn't worth supporting.

Whichever is chosen, the transition table and the base epoch should come from the same place so the two read modes stay consistent.

Environment overview (please complete the following information)

  • Environment location: any
  • Method of cuDF install: from source (main)

Additional context
Raised in review of #24265: #24265 (comment). The limitation predates that PR for the honored path, since the transition table was already built from the first stripe only.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ? - Needs TriagebugSomething isn't workinglibcudfAffects libcudf (C++/CUDA) code.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions