Repository navigation
regexr 0.2.2
Fixed
- A malformed escape at the start of a pattern is rejected instead of compiling to the empty pattern, which matched every input.
\e,\a,\q,\x,\u{,\p{and a lone\all reported success and matched everywhere; only the first token was affected, so the same escape one character later already errored. - Inline flag negation works:
(?-i),(?im-sx)and(?-x:…)were rejected as invalid groups because-never reached the flag parser. - A mid-pattern
(?flags)applies to what follows it rather than to the whole pattern. The change previously reached only the AST's single global flag set, so(?i)ab(?-i)CDmatched case-sensitively throughout. - An escape that is invalid inside a character class names itself:
[\b]reportedinvalid escape sequence '\?'.
Added
- Extended mode (
x) is implemented. Unescaped ASCII whitespace and#comments outside a character class are no longer part of the pattern; previously the flag was accepted and ignored. \Xmatches one extended grapheme cluster, following UAX #29 in full — Hangul syllables, emoji ZWJ sequences, regional-indicator flag pairs and Indic conjuncts included. Checked against the complete Unicode grapheme-break test suite.\Q…\Equotes everything between it literally, including metacharacters, extended-mode whitespace and character-class syntax. An unterminated\Qruns to the end of the pattern and a stray\Eis ignored, as in PCRE.\a(alert, U+0007) and\e(escape, U+001B).\followed by any non-alphanumeric ASCII character is that character, matching PCRE, Perl, Java and theregexcrate.\,\",\@and\#previously failed to compile.
Changed
escapealso escapes#and ASCII whitespace, so its output is safe to splice into an extended-mode pattern.Parser::newreturnsResult<Self>, andLexer::read_identis replaced by the infallibleLexer::read_ident_rest. Both are internal parsing types exposed throughregexr::parser;RegexandRegexBuilderare unaffected.
Full diff: v0.2.1...v0.2.2