<B>3.1.2 Python For Text</B>

Python can manipulate text (represented by type str, so-called “strings”) as well as numbers. This includes characters “!”, words “rabbit”, names “Paris”, sentences “Got your back.”, etc. “Yay! :)”. They can be enclosed in single quotes ('...') or double quotes ("...") with the same result.

In [1]:
# Single quote text
'spam eggs'

'spam eggs'

In [2]:
# Double Quote text
"spam eggs"

'spam eggs'

In [3]:
# Digits in quotes are also strings
'1975'

'1975'

To quote a quote, we need to “escape” it, by preceding it with \. Alternatively, we can use the other type of quotation marks:

In [4]:
# Use \' to escape the single quote
'doesn\'t'

"doesn't"

In [5]:
# Using double quotes insterad
'"Yes," they said.'

'"Yes," they said.'

In [6]:
# Or like this
"\"Yes,\" they said."

'"Yes," they said.'

In the Python shell, the string definition and output string can look different. The print() function produces a more readable output, by omitting the enclosing quotes and by printing escaped and special characters:

In [7]:
# \n means newline
s = 'First line.\nSecond line.'

In [8]:
# What not to do
s

'First line.\nSecond line.'

In [9]:
# What to do
print(s)

First line.
Second line.


If you don’t want characters prefaced by \ to be interpreted as special characters, you can use raw strings by adding an r before the first quote:

In [10]:
# Wrong way
print('C:\some\name')

C:\some
ame


In [11]:
# The correct way
print(r'C:\some\name')

C:\some\name


There is one subtle aspect to raw strings: a raw string may not end in an odd number of \ characters; see the FAQ entry for more information and workarounds.

String literals can span multiple lines. One way is using triple-quotes: """...""" or '''...'''. End-of-line characters are automatically included in the string, but it’s possible to prevent this by adding a \ at the end of the line. In the following example, the initial newline is not included:

In [12]:
print("""\
Usage: thingy [OPTIONS]
     -h                        Display this usage message
     -H hostname               Hostname to connect to
""")

Usage: thingy [OPTIONS]
     -h                        Display this usage message
     -H hostname               Hostname to connect to



Strings can be concatenated (glued together) with the + operator, and repeated with *:

In [13]:
# 3 times 'un', followed by 'ium'
3 * 'un' + 'ium'

'unununium'

Two or more string literals (i.e. the ones enclosed between quotes) next to each other are automatically concatenated.

In [14]:
"Py" "thon"

'Python'

In [15]:
# Me Playing around
"Un" "ob" "tan" "ium"

'Unobtanium'

This feature is particularly useful when you want to break long strings:

In [18]:
text = ('Put several strings within parentheses ' 
        'to have them joined together.')

In [19]:
text

'Put several strings within parentheses to have them joined together.'

This only works with two literals though, not with variables or expressions:

In [23]:
prefix = 'Py'

In [24]:
prefix "thon"

SyntaxError: invalid syntax (1520653236.py, line 1)

If you want to concatenate variables or a variable and a literal, use +:

In [25]:
prefix + 'thon'

'Python'

Strings can be indexed (subscripted), with the first character having index 0. There is no separate character type; a character is simply a string of size one:

In [26]:
word = "Python"

In [30]:
# Letter spot 1
word[0]

'P'

In [31]:
# Letter spot 6
word[5]

'n'

In [34]:
# Since there are no more letters, this should fail
word[6]

IndexError: string index out of range

Indices may also be negative numbers, to start counting from the right:

In [39]:
# This is the last character
word[-1]

'n'

In [40]:
# This is the second to last character
word[-2]

'o'

In [38]:
word[-6]

'P'

Note that since -0 is the same as 0, negative indices start from -1.

In addition to indexing, slicing is also supported. While indexing is used to obtain individual characters, slicing allows you to obtain a substring:

In [41]:
# character from the beginning to position 2 (excluded)
word[:2]

'Py'

In [42]:
# characters from position 4 (included) to the end
word[4:]

'on'

In [43]:
# characters from the second-last (included) to the end
word[-2:]

'on'

Note how the start is always included, and the end always excluded. This makes sure that s[:i] + s[i:] is always equal to s:

In [45]:
word[:2] + word[2:]

'Python'

In [46]:
word[:4] + word[4:]

'Python'

One way to remember how slices work is to think of the indices as pointing between characters, with the left edge of the first character numbered 0. Then the right edge of the last character of a string of n characters has index n, for example:

 +---+---+---+---+---+---+<br>
 |  P |  y |  t |  h |  o |  n |<br>
 +---+---+---+---+---+---+<br>
 0   1   2   3   4   5   6<br>
-6  -5  -4  -3  -2  -1<br>

The first row of numbers gives the position of the indices 0…6 in the string; the second row gives the corresponding negative indices. The slice from i to j consists of all characters between the edges labeled i and j, respectively.

For non-negative indices, the length of a slice is the difference of the indices, if both are within bounds. For example, the length of word[1:3] is 2.

Attempting to use an index that is too large will result in an error:

In [47]:
word[42]

IndexError: string index out of range

However, out of range slice indexes are handled gracefully when used for slicing:

In [48]:
word[4:42]

'on'

In [49]:
word[42:]

''

Python strings cannot be changed — they are immutable. Therefore, assigning to an indexed position in the string results in an error:

In [50]:
word[0] = 'J'

TypeError: 'str' object does not support item assignment

In [51]:
word[2:] = 'py'

TypeError: 'str' object does not support item assignment

If you need a different string, you should create a new one:

In [52]:
'J' + word[1:]

'Jython'

In [54]:
word[:2] + 'py'

'Pypy'

The built-in function len() returns the length of a string:

In [55]:
x = 'supercalifragilisticexpialidocious'
len(x)

34

See also
Text Sequence Type — str
Strings are examples of sequence types, and support the common operations supported by such types.

String Methods
Strings support a large number of methods for basic transformations and searching.

f-strings
String literals that have embedded expressions.

Format String Syntax
Information about string formatting with str.format().

printf-style String Formatting
The old formatting operations invoked when strings are the left operand of the % operator are described in more detail here.