You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Describe the bug
The ORC reader resolves the writer timezone from the first selected stripe only (reader_impl_chunking.cu, selected_stripes[0].stripe_footer->writerTimezone) and applies it to every stripe in the read. Both the transition table used to convert timestamps and the base epoch used for the negative timestamp borrow (#24259, #24265) come from that one stripe. The writerTimezone of the other stripes is parsed but never checked, so a mismatch is not detected or reported.
A single file normally has one writer timezone, but a multi-source read aggregates stripes from several files, which can be written in different timezones. In that case:
Timezone honored: every stripe is converted with the first file's offsets, so timestamps from the other files are shifted by the difference between the zones.
Timezone ignored (ignore_timezone_in_stripe_footer): values stay on each file's wall clock, but the borrow is decided in the first file's frame. Timestamps within the timezone offset of 1970-01-01 that have a fractional part can come back one second off. This includes a UTC file read after a non-UTC one, which decoded correctly before Decide the ORC timestamp borrow in the writer's timezone #24265.
Found by code inspection while addressing review on #24265; not yet reproduced.
Steps/Code to reproduce bug
Write two ORC files with timestamp columns, one with writer timezone Asia/Shanghai and one with UTC (or any two different zones).
Read both in a single cudf::io::read_orc call, listing the Shanghai file first, with ignore_timezone_in_stripe_footer both unset and set.
Compare against reading each file on its own.
Expected behavior
Open question, to confirm with the Spark team before deciding:
Resolve each distinct writer timezone once and have each stripe use its own transition table and base epoch, so a multi-source read matches reading each file separately.
Or reject mixed writer timezones in one read, or warn about them, if resolving per stripe isn't worth supporting.
Whichever is chosen, the transition table and the base epoch should come from the same place so the two read modes stay consistent.
Environment overview (please complete the following information)
Environment location: any
Method of cuDF install: from source (main)
Additional context
Raised in review of #24265: #24265 (comment). The limitation predates that PR for the honored path, since the transition table was already built from the first stripe only.
Describe the bug
The ORC reader resolves the writer timezone from the first selected stripe only (
reader_impl_chunking.cu,selected_stripes[0].stripe_footer->writerTimezone) and applies it to every stripe in the read. Both the transition table used to convert timestamps and the base epoch used for the negative timestamp borrow (#24259, #24265) come from that one stripe. ThewriterTimezoneof the other stripes is parsed but never checked, so a mismatch is not detected or reported.A single file normally has one writer timezone, but a multi-source read aggregates stripes from several files, which can be written in different timezones. In that case:
ignore_timezone_in_stripe_footer): values stay on each file's wall clock, but the borrow is decided in the first file's frame. Timestamps within the timezone offset of 1970-01-01 that have a fractional part can come back one second off. This includes a UTC file read after a non-UTC one, which decoded correctly before Decide the ORC timestamp borrow in the writer's timezone #24265.Found by code inspection while addressing review on #24265; not yet reproduced.
Steps/Code to reproduce bug
Asia/Shanghaiand one withUTC(or any two different zones).cudf::io::read_orccall, listing the Shanghai file first, withignore_timezone_in_stripe_footerboth unset and set.Expected behavior
Open question, to confirm with the Spark team before deciding:
Whichever is chosen, the transition table and the base epoch should come from the same place so the two read modes stay consistent.
Environment overview (please complete the following information)
main)Additional context
Raised in review of #24265: #24265 (comment). The limitation predates that PR for the honored path, since the transition table was already built from the first stripe only.