# Regular Expression

A regular expression (or RE) specifies a set of strings that matches it; the functions in this module let you check if a particular string matches a given regular expression (or if a given regular expression matches a particular string, which comes down to the same thing).

An "r" or "R" prefix is present, a character following a backslash is included in the string without change, and all backslashes are left in the string. For example, the string literal r"\n" consists of two characters: a backslash and a lowercase "n". String quotes can be escaped with a backslash, but the backslash remains in the string; for example, r"\"" is a valid string literal consisting of two characters: a backslash and a double quote; r"\" is not a valid string literal (even a raw string cannot end in an odd number of backslashes). Specifically, a raw string cannot end in a single backslash (since the backslash would escape the following quote character). Note also that a single backslash followed by a newline is interpreted as those two characters as part of the string, not as a line continuation.

Regular expressions can contain both special and ordinary characters. Most ordinary characters, like 'A', 'a', or '0', are the simplest regular expressions; they simply match themselves. You can concatenate ordinary characters, so last matches the string 'last'. (In the rest of this section, we’ll write RE’s in this special style, usually without quotes, and strings to be matched 'in single quotes'.)

## Functions

'`re.compile(pattern, flags=0)`

compile a regular expression pattern into a `regular expression object`, which can be used to matching using its `match()`,`search()` and other methods.

import re

prog=re.compile(pattern)

result=program.match(string)

In [6]:
import re

pattern=r'[0-9]'
string1="123abc"
string2="456def"

prog=re.compile(pattern)
result1=prog.match(string1)
result2=prog.match(string2)

In [7]:
print(result1)
print(result2)

<re.Match object; span=(0, 1), match='1'>
<re.Match object; span=(0, 1), match='4'>


but using re.compile() and saving the resulting regular expression object for reuse is more efficient when the expression will be used several times in a single program.

**Note:** 

The compiled versions of the most recent patterns passed to re.compile() and the module-level matching functions are cached, so programs that use only a few regular expressions at a time needn’t worry about compiling regular expressions.

## re.search(pattern, string, flags=0)

Scan through string looking for the first location where the regular expression pattern produces a match, and return a corresponding `Match`. Return None if no position in the string matches the pattern; note that this is different from finding a zero-length match at some point in the string.

In [10]:
import re

text = "The quick brown fox jumps over the lazy dog."

# search for the word 'fox' in the text
match = re.search(r'fox',text)

if match:
    print(f'Match found {match.group()} at position {match.start()}-{match.end()}')
else:
    print("No match found")

Match found fox at position 16-19


## re.match(pattern, string, flags=0)

If zero or more characters at the beginning of string match the regular expression pattern, return a corresponding `Match`. Return `None` if the string does not match the pattern; note that this is different from a zero-length match.

**Note** that even in `MULTILINE` mode, `re.match()` will only match at the `beginning` of the string and not at the beginning of each line.

In [12]:
import re

text = "Hello World"

match = re.match(r'Hello',text)

print(match)

<re.Match object; span=(0, 5), match='Hello'>


## re.fullmatch(pattern, string, flags=0)

If the whole string matches the regular expression pattern, return a corresponding Match. Return None if the string does not match the pattern; note that this is different from a zero-length match.

In [14]:
import re

text="Hello World"

match=re.fullmatch(r'Hello World',text)
print(match)

<re.Match object; span=(0, 11), match='Hello World'>


## re.split(pattern, string, maxsplit=0, flags=0)

Split string by the occurrences of pattern. If capturing parentheses are used in pattern, then the text of all groups in the pattern are also returned as part of the resulting list. If maxsplit is nonzero, at most maxsplit splits occur, and the remainder of the string is returned as the final element of the list.

In [15]:
text="python is a powerful language"
result=text.split()
print(result)

['python', 'is', 'a', 'powerful', 'language']


In [17]:
text = "one two three four five"
result=text.split(' ',3)
print(result)

['one', 'two', 'three', 'four five']


## re.findall(pattern, string, flags=0)

Return all non-overlapping matches of pattern in string, as a list of strings or tuples. The string is scanned left-to-right, and matches are returned in the order found. Empty matches are included in the result.

In [18]:
import re

text = "An apple a day keeps all the ants away."

match=re.findall(r'day',text)
print(match)

['day']


Regular expressions can contain both special and oridinary characters. Most ordinary characters, like `'A','a', or '0'`, are the simplest regular expressions they simply match themselves. You can concatenate ordinary characters. so `last` matches the string `last`.

Some characters, like `'|' or '('`, are special. Special characters either stand for classes of ordinary characters, or affect how the regular expressions around them are interpreted.

Repetition operators or quantifiers `(*, +, ?, {m,n}, etc)` cannot be directly nested. This avoids ambiguity with the non-greedy modifier suffix ?, and with other modifiers in other implementations. To apply a second repetition to an inner repetition, parentheses may be used. For example, the expression (?:a{6})* matches any multiple of six `'a'` characters.

### The special characters are:

**(Dot.)**

(Dot.) in the default mode, this matches any character except a newline. If the `DOTALL` flag has been specified this matches any character including a newline.

In [4]:
import re

text = """Hello
World"""

# Using dot (.) to match any character except newline
result = re.findall(r".", text)

print(result)

['H', 'e', 'l', 'l', 'o', 'W', 'o', 'r', 'l', 'd']


## IGNORECASE

The `re.IGNORECASE` (or re.I) flag is used in Python's re module to perform case-insensitive matching. This means that when a pattern is matched against a string, it will ignore whether letters are uppercase or lowercase.

In [7]:
import re

text = "Hello World"

pattern=r'hello'

match=re.findall(pattern,text,re.IGNORECASE)

print(match)

['Hello']


**^(Caret.)**

`^(Caret.) Matches the **start of the string**, and in MULTILINE mode also matches immediately after each newline.


In [8]:
import re

text = "Python is a programming language"

match=re.search(r'^Python',text) # with ^ caret.
print(match)

match=re.search(r'^is',text) # with ^ caret.
print(match)

match=re.search(r'is',text) # without ^ caret.
print(match)

<re.Match object; span=(0, 6), match='Python'>
None
<re.Match object; span=(7, 9), match='is'>


**$**

Matches the end of the string or just before the newline at the `end of the string`, and in MULTILINE mode also matches before a newline.

In [14]:
import re

text = 'Python is a programming language'

match=re.search(r'language$',text)
if match:
    print(f'match found: {match.group()}')
else:
    print('match not found')

match found: language


**asterisk (*)**

Causes the resulting RE to match 0 or more repetitions of the preceding RE, as many repetitions as are possible. ab* will match ‘a’, ‘ab’, or ‘a’ followed by any number of ‘b’s.

In [23]:
import re

text = ['uday','kiran','jyostna','manoj','karan']
for name in text:
    match=re.fullmatch(r'u.*y|j.*a',name)
    if match:
        print(name)

uday
jyostna


In [28]:
import re

text = "a ab aab aaab abab"

match=re.findall(r'a*b',text)
print(match)

['ab', 'aab', 'aaab', 'ab', 'ab']


### +

Causes the resulting RE to match 1 or more repetitions of th 
preceding RE. ab+ will match ‘a’ followed by any non-zero number of ‘b,

it t
will not match just ‘a’.

In [34]:
import re

text="a ab abc aab abab"

match=re.search(r'a+b',text) # search return the first occurence in the text
print(match)
match=re.findall(r'ab+',text) # findall returns the all occurence in the text
print(match)
match=re.match(r'ab+',text) # match returns the begining of the text if it found else it returns None
print(match)

<re.Match object; span=(2, 4), match='ab'>
['ab', 'ab', 'ab', 'ab', 'ab']
None


### ?

Causes the resulting RE to match 0 or 1 repetitions of the preceding RE. ab? will match either ‘a’ or ‘ab’.


In [64]:
import re

text = "a aabc ababc abab aabb c"

match = re.findall(r'ab?c+',text) # ab? match 0 or 1 repetitions of the preceding re and + 'a' followed by 'b' repetitions 1 or more repetitions

print(match)

['abc', 'abc']


**{m}**

Specifies that exactly m copies of the previous RE should be matched; fewer matches cause the entire RE not to match. For example, a{6} will match exactly six 'a' characters, but not five.



In [49]:
import re

text = "aaabbb aabbbb aabbbb abababab"

match= re.findall(r'ab{4}',text) # Pattern to match 'a' followed by exactly 4 'b's

print(match)

['abbbb', 'abbbb']


## {m,n}

Causes the resulting RE to match from m to n repetitions of the preceding RE, attempting to match as many repetitions as possible. For example, a{3,5} will match from 3 to 5 'a' characters. Omitting m specifies a lower bound of zero, and omitting n specifies an infinite upper bound. As an example, a{4,}b will match 'aaaab' or a thousand 'a' characters followed by a 'b', but not 'aaab'. The comma may not be omitted or the modifier would be confused with the previously described form.



In [51]:
import re

text ="aaabbb aabbbb aabbbb abababab aaabbbbbb"

match= re.findall(r'ab{4,6}',text) # Pattern to match 'a' followed by exactly minimum 4 and maximum 6 'b's
print(match)

['abbbb', 'abbbb', 'abbbbbb']


### {m,n}?

Causes the resulting RE to match from m to n repetitions of the preceding RE, attempting to match as few repetitions as possible. This is the non-greedy version of the previous quantifier. For example, on the 6-character string 'aaaaaa', a{3,5} will match 5 'a' characters, while a{3,5}? will only match 3 characters.

In [55]:
import re

text = "aaaaaa aaaa a aa aaaaab"

# Pattern to match 'a' occurring between 2 and 5 times, non-greedy
pattern = r'a{2,5}?'

matches = re.findall(pattern, text)
print(matches)


['aa', 'aa', 'aa', 'aa', 'aa', 'aa', 'aa', 'aa']


### {m,n}+

The {m,n}+ quantifier in a regular expression matches a sequence of characters that appears at least m times but no more than n times, and it will try to match as many characters as possible (greedy).

In [67]:
import re

text = "aaaaaa a ab aaab"

pattern = r'a{2,4}+'

match = re.findall(pattern,text)

print(match)

['aaaa', 'aa', 'aaa']


### []

Ranges of characters can be indicated by giving two characters and separating them by a '-', for example [a-z] will match any lowercase ASCII letter, [0-5][0-9] will match all the two-digits numbers from 00 to 59, and [0-9A-Fa-f] will match any hexadecimal digit

In [73]:
# extract date and time from string
import re

text = "Date of joining 20-09-2024"

match=re.search(r'([0-9]{2})-([0-9]{2})-([0-9]{4})',text)

print(match)
print(match.group(0))
print(match.group(1))
print(match.group(2))
print(match.group(3))

<re.Match object; span=(16, 26), match='20-09-2024'>
20-09-2024
20
09
2024


### /

Either escapes special characters (permitting you to match characters like '*', '?', and so forth), or signals a special sequence; special sequences are discussed below.

If you’re not using a raw string to express the pattern, remember that Python also uses the backslash as an escape sequence in string literals; if the escape sequence isn’t recognized by Python’s parser, the backslash and subsequent character are included in the resulting string. However, if Python would recognize the resulting sequence, the backslash should be repeated twice. This is complicated and hard to understand, so it’s highly recommended that you use raw strings for all but the simplest expressions.

In [74]:
import re

text = 'This is sentence with period.'

match=re.search(r'\.',text)

print(match)

<re.Match object; span=(28, 29), match='.'>


### |

A|B, where A and B can be arbitrary REs, creates a regular expression that will match either A or B. An arbitrary number of REs can be separated by the 'A|B' in this way.

In [87]:
str1="joining date is 12-06-2024 and time 12:45:54"

str2="joining time 12:45:54 and date is 12-06-2024"

match=re.search(r'([0-9]{2})-([0-9]{2})-([0-9]{4})|([0-9]{2}):([0-9]{2}):([0-9]{2})',str1)

print(match)

match=re.search(r'([0-9]{2})-([0-9]{2})-([0-9]{4})|([0-9]{2}):([0-9]{2}):([0-9]{2})',str2)

print(match)

<re.Match object; span=(16, 26), match='12-06-2024'>
<re.Match object; span=(13, 21), match='12:45:54'>


### \A

Matches only at the start of the string.

In [4]:
import re

text = """python is a programming language
python is a object oriented language"""

match= re.findall(r'\Apython',text,re.MULTILINE) # \A matches only the begining of the string, doesn't consider breaks of line
print(match)

['python']


In [6]:
import re

text = """python is a programming language
python is a object oriented language"""

match= re.findall(r'^python',text,re.MULTILINE) # ^caret it consider the multiline breaks and beginning of the line
print(match)

['python', 'python']


### \b

Matches the empty string, but only at the beginning or end of a word. A word is defined as a sequence of word characters. Note that formally, \b is defined as the boundary between a \w and a \W character (or vice versa), or between \w and the beginning or end of the string. This means that r'\bat\b' matches 'at', 'at.', '(at)', and 'as at ay' but not 'attempt' or 'atlas'.


In [4]:
import re

text = "Python is an object oriented programming language"

match = re.search(r'\bprogramming\b', text)

print(match)

<re.Match object; span=(29, 40), match='programming'>


### \B

Matches the empty string, but only when it is not at the beginning or end of a word. This means that r'at\B' matches 'athens', 'atom', 'attorney', but not 'at', 'at.', or 'at!'. \B is the opposite of \b, so word characters in Unicode (str) patterns are Unicode alphanumerics or the underscore, although this can be changed by using the ASCII flag. Word boundaries are determined by the current locale if the LOCALE flag is used.


In [11]:
import re

text = "Python is an object oriented programming language."

match = re.search(r'\Bgram\B',text)

print(match)

<re.Match object; span=(32, 36), match='gram'>


### \d

For Unicode (str) patterns:
Matches any Unicode decimal digit (that is, any character in Unicode character category [Nd]). This includes [0-9], and also many other digit characters.

Matches [0-9] if the ASCII flag is used.

For 8-bit (bytes) patterns:
Matches any decimal digit in the ASCII character set; this is equivalent to [0-9].

In [24]:
import re

text = "i have 12 apples and 4 oranges."

match = re.findall(r'\d',text) # \d will match each digit individually

print(match)

['1', '2', '4']


In [25]:
import re

text = "i have 12 apples and 24 oranges."

match = re.findall(r'\d+',text) # \d+ will match sequences of digits

print(match)

['12', '24']


### \D

Matches any character which is not a decimal digit. This is the opposite of \d.

Matches [^0-9] if the ASCII flag is used.

In [26]:
import re

text = "i have 12 apples and 24 oranges."

match = re.findall(r'\D',text) # \D will match sequences of digits

print(match)

['i', ' ', 'h', 'a', 'v', 'e', ' ', ' ', 'a', 'p', 'p', 'l', 'e', 's', ' ', 'a', 'n', 'd', ' ', ' ', 'o', 'r', 'a', 'n', 'g', 'e', 's', '.']


### \s

For Unicode (str) patterns:
Matches Unicode whitespace characters (as defined by str.isspace()). This includes [ \t\n\r\f\v], and also many other characters, for example the non-breaking spaces mandated by typography rules in many languages.

Matches [ \t\n\r\f\v] if the ASCII flag is used.

For 8-bit (bytes) patterns:
Matches characters considered whitespace in the ASCII character set; this is equivalent to [ \t\n\r\f\v].

In [29]:
import re

text = "Hello World!\nThis\tis text string."

match = re.findall(r'\s',text)

print(match)

[' ', '\n', '\t', ' ', ' ']


### \S

Matches any character which is not a whitespace character. This is the opposite of \s.

Matches [^ \t\n\r\f\v] if the ASCII flag is used.

In [31]:
import re

text = "Hello World!\nThis\tis text string."

match = re.findall(r'\S',text) # Use \S to find haracter which is not a whitespace character. This is the opposite of \s.

print(match)

['H', 'e', 'l', 'l', 'o', 'W', 'o', 'r', 'l', 'd', '!', 'T', 'h', 'i', 's', 'i', 's', 't', 'e', 'x', 't', 's', 't', 'r', 'i', 'n', 'g', '.']


### \w

For Unicode (str) patterns:
Matches Unicode word characters; this includes all Unicode alphanumeric characters (as defined by str.isalnum()), as well as the underscore (_).

Matches [a-zA-Z0-9_] if the ASCII flag is used.

In [35]:
import re

text = "Python is popular a programming language.It was released in 1991."

match = re.findall(r'\w+',text)

print(match)

['Python', 'is', 'popular', 'a', 'programming', 'language', 'It', 'was', 'released', 'in', '1991']


### \W

Matches any character which is not a word character. This is the opposite of \w. By default, matches non-underscore (_) characters for which str.isalnum() returns False.

Matches [^a-zA-Z0-9_] if the ASCII flag is used.

In [38]:
import re

text = "i have 2 mangoes and 5 apples the cost of 100/-."

match = re.findall(r'\W',text) # return not worf characters

print(match)

[' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', ' ', '/', '-', '.']


### \Z

Matches only at the end of the string.

In [43]:
import re

text = "Hello World!"

match = re.search(r'World!\Z',text)

if match:
    print("match found")

else:
    print("match not found")

match found


`(?P<name>....)`

similar to regular parentheses. But the substring matched by the group is accessible via the symbolic group name.

In [53]:
import re

str1 = "SBI IFSC CODE SBIN0070256"

m = re.search(r'(?P<ifsc>SBIN[0-9]{7})',str1)

print(m)

print(m.group(0))

print(m.group('ifsc'))

<re.Match object; span=(14, 25), match='SBIN0070256'>
SBIN0070256
SBIN0070256


In [59]:
import re

str1 = "SBI IFSC CODE SBIN0070256"

m = re.search(r'(?P<ifsc>SBIN\d{7})',str1)

print(m)

print(m.group(0))

print(m.group('ifsc'))

<re.Match object; span=(14, 25), match='SBIN0070256'>
SBIN0070256
SBIN0070256


`(?P=name)`

A backreference to a named group it matches whatever text was matched by the earlier group name nam.

In [64]:
import re

str1 = "aa-aa"

m = re.search(r'(?P<x>aa)',str1)

print(m)

str2 = "aa-aa"

m = re.search(r'(?P<x>aa)-(?P=x)',str2)

print(m)

str3 = "aa-ab"

m = re.search(r'(?P<x>aa)-(?P=x)',str2)

print(m)   

<re.Match object; span=(0, 2), match='aa'>
<re.Match object; span=(0, 5), match='aa-aa'>
<re.Match object; span=(0, 5), match='aa-aa'>


### ?=

Matches if... matches next, but doesn't consume any of the string. This is called a lookahead assertion. For example, Issac (?=Asimov) will match 'Isaac' only if it's followed by 'Asimov'.

In [67]:
import re

text = "I love python programming"

match = re.search(r'python(?= programming)',text)

if match:
    print(f'match found: {match.group(0)}')

else:
    print("match not found")

match found: python


### ?!...

Matches if... doesn't match next. This is a negative lookahead assertion. For example, Issac(?! Asimov) will match 'Isaac' onlu if it's not followed by 'Asimov'.

In [70]:
import re

text = "I love python language"

match = re.search(r'python(?! programming)',text)

if match:
    print(f'match found: {match.group(0)}')

else:
    print("match not found")

match found: python


### (?<=....)

Matches if the current position in the string is preceded by a match for... that ends at the current position. This is called a positive lookbehind assertion. (?<=abc)def will find a match in 'abcdef'. Since the lookbehind will back up 3 characters and check if the contained pattern matches. The contained pattern must only match strings of some fixed length, meaning that abc or a|b are allowed, but a* and a{3,4} are not.


In [72]:
import re

str1 = 'abcdef'

match = re.search('(?<=abc)def',str1)

print(match)

<re.Match object; span=(3, 6), match='def'>


### (?(id/name)yes-pattern|no-pattern)

Will try to match with yes-pattern if the group with given id or name exists, and with no-pattern if it doesn’t. no-pattern is optional and can be omitted. For example, (<)?(\w+@\w+(?:\.\w+)+)(?(1)>|$) is a poor email matching pattern, which will match with '<user@host.com>' as well as 'user@host.com', but not with '<user@host.com' nor 'user@host.com>'.

In [73]:
import re

pattern = r'(<)?(\w+@\w+(?:\.\w+)+)(?(1)>|$)'

text = "<user@host.com>"
print(re.match(pattern, text))  # Match

text = "user@host.com"
print(re.match(pattern, text))  # Match

text = "<user@host.com"
print(re.match(pattern, text))  # No match

text = "user@host.com>"
print(re.match(pattern, text))  # No match


<re.Match object; span=(0, 15), match='<user@host.com>'>
<re.Match object; span=(0, 13), match='user@host.com'>
None
None


### ?:....

When you don't need to capture the content, non-capturing groups make the regex more efficient. It shows that you're grouping for pattern logic, not for capturing content.

In [76]:
import re

text = "I have red car and blue car"

match = re.findall(r'(red|blue) (?:car)',text)

print(match)

['red', 'blue']


In [77]:
import re

text = "The events are on 12-06-2024, 12/06/2024, and 12.06.2024."

# Pattern to match dates with any separator (-, /, .)
pattern = r'(\d{2})(?:[-/.])(\d{2})(?:[-/.])(\d{4})'

matches = re.findall(pattern, text)
print(matches)


[('12', '06', '2024'), ('12', '06', '2024'), ('12', '06', '2024')]
