Describe the bug
astype(str) on a timedelta column formats negative values differently from pandas, and pandas reads the cudf strings back as different values. Parsing has the mirror problem: cudf reads pandas-formatted negative timedelta strings that have a time part as different values. The same happens under cudf.pandas, where pd.Series(...).astype(str) returns '-0 days 00:00:01' for -1 second.
Steps/Code to reproduce bug
import cudf
import pandas as pd
s = pd.Series(pd.to_timedelta([-1, 5, -90061], unit="s"))
s.astype(str).tolist()
# ['-1 days +23:59:59', '0 days 00:00:05', '-2 days +22:58:59']
cudf.Series(s).astype(str).to_pandas().tolist()
# ['-0 days 00:00:01', '0 days 00:00:05', '-1 days 01:01:01']
pd.to_timedelta(cudf.Series(s).astype(str).to_pandas()).tolist()
# [Timedelta('0 days 00:00:01'), Timedelta('0 days 00:00:05'), Timedelta('-1 days +01:01:01')]
cudf.Series(["-1 days +23:59:59"]).astype("timedelta64[s]").to_pandas().tolist()
# [Timedelta('-2 days +00:00:01')]
Expected behavior
The same strings as pandas, and "-1 days +23:59:59" parsed as -1 second. cudf's own CSV writer and reader already do this: to_csv writes -1 days +23:59:59.000000000 for -1 second, and read_csv with a timedelta64[s] dtype reads -1 days +23:59:59 as -1 second.
Environment overview (please complete the following information)
- Environment location: Bare-metal, Linux, RTX 5090 (driver 610.43.02)
- Method of cuDF install: pip, nightly
cudf-cu13 26.12.00a177 (commit 1c84b4e), pandas 3.0.6, Python 3.12
Environment details
- Ubuntu 24.04, Linux 7.0
- NVIDIA GeForce RTX 5090, driver 610.43.02, CUDA 13 (
cu13 wheels)
- Python 3.12.13,
cudf-cu13 26.12.0a177 (commit 1c84b4e), pandas 3.0.6, pyarrow 25.0.1
Additional context
TimeDeltaColumn._as_string_pandas_compat passes the column to from_durations with "%D days %H:%M:%S", and StringColumn.as_timedelta_column uses to_durations with the same format. libcudf puts a single sign in front of the whole duration, as convert_durations.hpp documents ("If duration value is negative, only one negative sign is written to output string."), while pandas writes floored days followed by a non-negative time of day and in the N days HH:MM:SS form applies a leading minus to the days only.
I have a Python-side change that fixes both directions:
- when the output has a time part, format the time of day (
self % 1 day) with "%H:%M:%S" and prepend the floored days from .days, with " days +" for negative values;
- after
to_durations, rows starting with - come back as x = -(d + t), and are corrected to -d + t = x + 2 * ((-x) % 1 day).
With it, astype(str) matches pandas for s, ms, us and ns in my tests, and the existing python/cudf/cudf/tests and the related pandas tests under cudf.pandas show no new failures. Two things for whoever decides:
- It changes the string output for negative values, and
'-0 days 00:00:01' written by current cudf would then read back as +1 second, as pandas reads it. So it is probably a breaking change.
- For
ms, us and ns, negative values between pd.Timedelta.min (-106752 days +00:12:43.145224193) and -106751 days would format as -106752 days +.... These round-trip through cudf today, but cudf cannot parse the new strings back, because to_durations already returns wrong values for a day count of 106752 in these units ("106752 days 00:00:00" cast to timedelta64[ns] gives -106752 days +00:25:26.290448384 today). pandas raises OutOfBoundsDatetime for these strings.
If fixing this in the Python layer is the direction you want, I can open a PR with the change and tests.
Describe the bug
astype(str)on a timedelta column formats negative values differently from pandas, and pandas reads the cudf strings back as different values. Parsing has the mirror problem: cudf reads pandas-formatted negative timedelta strings that have a time part as different values. The same happens undercudf.pandas, wherepd.Series(...).astype(str)returns'-0 days 00:00:01'for -1 second.Steps/Code to reproduce bug
Expected behavior
The same strings as pandas, and
"-1 days +23:59:59"parsed as -1 second. cudf's own CSV writer and reader already do this:to_csvwrites-1 days +23:59:59.000000000for -1 second, andread_csvwith atimedelta64[s]dtype reads-1 days +23:59:59as -1 second.Environment overview (please complete the following information)
cudf-cu1326.12.00a177 (commit 1c84b4e), pandas 3.0.6, Python 3.12Environment details
cu13wheels)cudf-cu1326.12.0a177 (commit 1c84b4e), pandas 3.0.6, pyarrow 25.0.1Additional context
TimeDeltaColumn._as_string_pandas_compatpasses the column tofrom_durationswith"%D days %H:%M:%S", andStringColumn.as_timedelta_columnusesto_durationswith the same format. libcudf puts a single sign in front of the whole duration, asconvert_durations.hppdocuments ("If duration value is negative, only one negative sign is written to output string."), while pandas writes floored days followed by a non-negative time of day and in theN days HH:MM:SSform applies a leading minus to the days only.I have a Python-side change that fixes both directions:
self % 1 day) with"%H:%M:%S"and prepend the floored days from.days, with" days +"for negative values;to_durations, rows starting with-come back asx = -(d + t), and are corrected to-d + t = x + 2 * ((-x) % 1 day).With it,
astype(str)matches pandas for s, ms, us and ns in my tests, and the existingpython/cudf/cudf/testsand the related pandas tests undercudf.pandasshow no new failures. Two things for whoever decides:'-0 days 00:00:01'written by current cudf would then read back as +1 second, as pandas reads it. So it is probably a breaking change.ms,usandns, negative values betweenpd.Timedelta.min(-106752 days +00:12:43.145224193) and -106751 days would format as-106752 days +.... These round-trip through cudf today, but cudf cannot parse the new strings back, becauseto_durationsalready returns wrong values for a day count of 106752 in these units ("106752 days 00:00:00"cast totimedelta64[ns]gives-106752 days +00:25:26.290448384today). pandas raisesOutOfBoundsDatetimefor these strings.If fixing this in the Python layer is the direction you want, I can open a PR with the change and tests.