Bug description:
\p{Lu} matches uppercase Unicode letters. It breaks when mixed with a standard letter with re.IGNORECASE:
import re
property_only = re.compile(r"[\p{Lu}]", re.IGNORECASE)
property_plus_one = re.compile(r"[\p{Lu}1]", re.IGNORECASE)
property_plus_a = re.compile(r"[\p{Lu}a]", re.IGNORECASE)
assert property_only.fullmatch("B")
assert property_plus_one.fullmatch("B")
assert property_plus_a.fullmatch("B"), "adding 'a' broke the match for 'B'"
Expected: adding "a" to the character class shouldn't break the existing match
Actual: AssertionError: adding 'a' broke the match for 'B'
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
Bug description:
\p{Lu}matches uppercase Unicode letters. It breaks when mixed with a standard letter withre.IGNORECASE:Expected: adding
"a"to the character class shouldn't break the existing matchActual:
AssertionError: adding 'a' broke the match for 'B'CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
re.IGNORECASEfor Unicode property categories #155299