Describe the bug
Two builtin character classes are compiled wrong when a regex program is created with regex_flags::ASCII. Both are in cpp/src/strings/regex/regcomp.cpp.
\S outside brackets matches only 0-9. In regex_parser::lex(), the ASCII branch of case 'S' builds the class but does not return, so it falls through into case 'd', which replaces the class with the ASCII digit class and returns CCLASS. The non-ASCII branch has its return NCCLASS; inside the else. The \W and \D cases return NCCLASS after their if/else.
\W inside brackets treats _ as a non-word character. add_ascii_word_class(ranges, true) adds the range {'Z' + 1, 'a' - 1}, which contains _ (0x5F). The comment on that line, // {'_'-1, '_' + 1}, looks like the split that was intended. Outside brackets \W is built from the non-negated word class, so \W and [\W] disagree on _, and [^\W] does not match _ while \w does.
Both come from #11404. v26.08.01 and main (1c84b4e) have the same code.
Steps/Code to reproduce bug
cudf.Series.str.contains rejects re.ASCII (NotImplementedError: unsupported value for `flags` parameter), so this goes through pylibcudf with the raw value of regex_flags::ASCII. The same flag is RegexFlag.ASCII in the Java bindings.
import pyarrow as pa
import pylibcudf as plc
ASCII = 256 # cudf::strings::regex_flags::ASCII
col = plc.Column.from_arrow(pa.array(["a_b", "12", " \t", "_", "x-y"]))
for pat in [r"\S", r"[^\s]", r"\d", r"[\W]", r"[^\w]", r"\W"]:
for flags in (0, ASCII):
prog = plc.strings.regex_program.RegexProgram.create(pat, flags)
out = plc.strings.contains.count_re(col, prog).to_arrow().to_pylist()
print(f"count_re {pat!r:8} flags={flags:<3} -> {out}")
Output:
count_re '\\S' flags=0 -> [3, 2, 0, 1, 3]
count_re '\\S' flags=256 -> [0, 2, 0, 0, 0]
count_re '[^\\s]' flags=0 -> [3, 2, 0, 1, 3]
count_re '[^\\s]' flags=256 -> [3, 2, 0, 1, 3]
count_re '\\d' flags=0 -> [0, 2, 0, 0, 0]
count_re '\\d' flags=256 -> [0, 2, 0, 0, 0]
count_re '[\\W]' flags=0 -> [0, 0, 2, 0, 1]
count_re '[\\W]' flags=256 -> [1, 0, 2, 1, 1]
count_re '[^\\w]' flags=0 -> [0, 0, 2, 0, 1]
count_re '[^\\w]' flags=256 -> [0, 0, 2, 0, 1]
count_re '\\W' flags=0 -> [0, 0, 2, 0, 1]
count_re '\\W' flags=256 -> [0, 0, 2, 0, 1]
Expected behavior
The input is all ASCII, so every pattern should give the same counts with and without the flag. Python's re.findall with re.ASCII gives [3, 2, 0, 1, 3] for \S and [0, 0, 2, 0, 1] for [\W]. With the flag, \S gives the counts of \d, and [\W] counts the _ in "a_b" and "_".
A possible fix:
--- a/cpp/src/strings/regex/regcomp.cpp
+++ b/cpp/src/strings/regex/regcomp.cpp
@@ -274,7 +274,8 @@ class regex_parser {
ranges.push_back({'_', '_'});
} else {
ranges.back().last = 'A' - 1;
- ranges.push_back({'Z' + 1, 'a' - 1}); // {'_'-1, '_' + 1}
+ ranges.push_back({'Z' + 1, '_' - 1});
+ ranges.push_back({'_' + 1, 'a' - 1});
ranges.push_back({'z' + 1, MAX_REGEX_CHAR});
}
}
@@ -492,8 +493,8 @@ class regex_parser {
} else {
if (_id_cclass_s < 0) { _id_cclass_s = _prog.add_class(cclass_s); }
_cclass_id = _id_cclass_s;
- return NCCLASS;
}
+ return NCCLASS;
}
case 'd': {
if (is_ascii(_flags)) {
I compiled regcomp.cpp on the host, with and without this change, and printed the class each pattern compiles to under ASCII. With the change, \S becomes NCCLASS over [9-32], the same class [^\s] compiles to on main, and [\W] no longer contains _. \s, \d, \D, \w, \W and [\S] compile to the same classes as on main. I have not built libcudf with the change or run the gtests.
StringsContainsTests.ASCII uses \S only as [^\S], and none of its input strings contain _.
Environment overview (please complete the following information)
- Environment location: Bare-metal, Linux, RTX 5090 (driver 610.43.02)
- Method of cuDF install: pip, nightly
pylibcudf-cu13 26.12.0a177 (commit 1c84b4e), Python 3.12
Environment details
- Ubuntu 24.04, Linux 7.0
- NVIDIA GeForce RTX 5090, driver 610.43.02, CUDA 13 (
cu13 wheels)
Describe the bug
Two builtin character classes are compiled wrong when a regex program is created with
regex_flags::ASCII. Both are incpp/src/strings/regex/regcomp.cpp.\Soutside brackets matches only0-9. Inregex_parser::lex(), the ASCII branch ofcase 'S'builds the class but does not return, so it falls through intocase 'd', which replaces the class with the ASCII digit class and returnsCCLASS. The non-ASCII branch has itsreturn NCCLASS;inside theelse. The\Wand\Dcases returnNCCLASSafter theirif/else.\Winside brackets treats_as a non-word character.add_ascii_word_class(ranges, true)adds the range{'Z' + 1, 'a' - 1}, which contains_(0x5F). The comment on that line,// {'_'-1, '_' + 1}, looks like the split that was intended. Outside brackets\Wis built from the non-negated word class, so\Wand[\W]disagree on_, and[^\W]does not match_while\wdoes.Both come from #11404. v26.08.01 and main (1c84b4e) have the same code.
Steps/Code to reproduce bug
cudf.Series.str.containsrejectsre.ASCII(NotImplementedError: unsupported value for `flags` parameter), so this goes through pylibcudf with the raw value ofregex_flags::ASCII. The same flag isRegexFlag.ASCIIin the Java bindings.Output:
Expected behavior
The input is all ASCII, so every pattern should give the same counts with and without the flag. Python's
re.findallwithre.ASCIIgives[3, 2, 0, 1, 3]for\Sand[0, 0, 2, 0, 1]for[\W]. With the flag,\Sgives the counts of\d, and[\W]counts the_in"a_b"and"_".A possible fix:
I compiled
regcomp.cppon the host, with and without this change, and printed the class each pattern compiles to underASCII. With the change,\SbecomesNCCLASSover[9-32], the same class[^\s]compiles to on main, and[\W]no longer contains_.\s,\d,\D,\w,\Wand[\S]compile to the same classes as on main. I have not built libcudf with the change or run the gtests.StringsContainsTests.ASCIIuses\Sonly as[^\S], and none of its input strings contain_.Environment overview (please complete the following information)
pylibcudf-cu1326.12.0a177 (commit 1c84b4e), Python 3.12Environment details
cu13wheels)