Strings learning resource: https://docs.python.org/3/tutorial/introduction.html#strings

The following learnings are using section 3.1.2.

Besides numbers, Python can also manipulate strings, which can be expressed in several ways. They can be enclosed in single quotes ('...') or double quotes ("...") with the same result 2. \ can be used to escape quotes:

In [1]:
'span eggs' #single quotes

'span eggs'

In [2]:
'doesn\'t' #use \ to escape the single quote

"doesn't"

In [3]:
"doesn't" #or use double quotes instead

"doesn't"

In [4]:
'"Yes," they said.'

'"Yes," they said.'

In [5]:
'\'Yes,\' they said.'

"'Yes,' they said."

In [6]:
'\'Isn\'t,\' they said.'

"'Isn't,' they said."

In the interactive interpreter, the output string is enclosed in quotes and special characters are escaped with backslashes. While this might sometimes look different from the input (the enclosing quotes could change), the two strings are equivalent. The string is enclosed in double quotes if the string contains a single quote and no double quotes, otherwise it is enclosed in single quotes. The print() function produces a more readable output, by omitting the enclosing quotes and by printing escaped and special characters:

In [7]:
'"Isn\'t it," they said.'

'"Isn\'t it," they said.'

In [8]:
print('"Isn\'t it," they said.')

"Isn't it," they said.


In [12]:
s = 'First line.\nSecond line.' #\n means new line
s #without print(), \n is included in the output

'First line.\nSecond line.'

In [14]:
print(s) #using print() to produce a more readable output

First line.
Second line.


If you don’t want characters prefaced by \ to be interpreted as special characters, you can use raw strings by adding an r before the first quote:

In [15]:
print('C:\some\name') #here \n means new line

C:\some
ame


In [16]:
print(r'C:\some\name')

C:\some\name


There is one subtle aspect to raw strings: a raw string may not end in an odd number of \ characters; see the FAQ entry for more information and workarounds.

String literals can span multiple lines. One way is using triple-quotes: """...""" or '''...'''. End of lines are automatically included in the string, but it’s possible to prevent this by adding a \ at the end of the line. The following example:

In [17]:
print("""
Usage: thingy[OPTIONS]
    -h                      Display this usage message
    -H hostname             Histname to connect to
""")


Usage: thingy[OPTIONS]
    -h                      Display this usage message
    -H hostname             Histname to connect to



Strings can be concatenated (glued together) with the + operator, and repeated with *

In [18]:
# 3 times 'un', followed by 'ium'
3 * 'un' + 'ium'

'unununium'

Two or more string literals (i.e. the ones enclosed between quotes) next to each other are automatically concatenated.

In [19]:
'Py' 'thon'

'Python'

In [20]:
print('Py' 'thon')

Python


This feature is particularly useful when you want to break long strings:

In [22]:
text = ('Put several strings within parentheses'
       'to have them joined together.')
text

'Put several strings within parenthesesto have them joined together.'

This only works with two literals though, not with variables or expressions:

In [23]:
prefix = 'Py'
prefix 'thon' #can't concat a variable and a string literal

SyntaxError: invalid syntax (3207490062.py, line 2)

If you want to concatenate variables or a variable and a literal, use +:

In [25]:
prefix = 'Py'
prefix + 'thon'

'Python'

Strings can be indexed (subscripted), with the first character having index 0. There is no separate character type; a character is simply a string of size one:

In [26]:
word = 'Python'
word[0] #index first character in the string

'P'

In [27]:
word[5] #indexing the 6th character in the string

'n'

Indices may also be negative numbers, to start counting from the right:

In [29]:
word[-1] #indexing the last character in the string, same as the last character

'n'

In [30]:
word[-2] #indexing the 2nd to last character

'o'

Note that since -0 is the same as 0, negative indices start from -1.

In addition to indexing, slicing is also supported. While indexing is used to obtain individual characters, slicing allows you to obtain substring:

In [31]:
word[2:5]

'tho'

In [32]:
word[0:2]

'Py'

Slice indices have useful defaults; an omitted first index defaults to zero, an omitted second index defaults to the size of the string being sliced.

In [33]:
word[:2] #excluding 2

'Py'

In [35]:
word[3:] #including 3, aka h, and showing characters on and after h

'hon'

In [36]:
word[4:]

'on'

In [40]:
word[:-3] # characters from the first character to the third-last (excluded)

'Pyt'

In [39]:
word[-2:] # characters from the second-last (included) to the end

'on'

Note how the start is always included, and the end always excluded. This makes sure that s[:i] + s[i:] is always equal to s:

In [41]:
word[:2] + word[2:]

'Python'

One way to remember how slices work is to think of the indices as pointing between characters, with the left edge of the first character numbered 0. Then the right edge of the last character of a string of n characters has index n, for example:

In [50]:
 print('''
  +---+---+---+---+---+---+
  | P | y | t | h | o | n |
  +---+---+---+---+---+---+
  0   1   2   3   4   5   6
-6  -5  -4  -3  -2  -1
'''
 )


 +---+---+---+---+---+---+
 | P | y | t | h | o | n |
 +---+---+---+---+---+---+
 0   1   2   3   4   5   6
-6  -5  -4  -3  -2  -1



The first row of numbers gives the position of the indices 0…6 in the string; the second row gives the corresponding negative indices. The slice from i to j consists of all characters between the edges labeled i and j, respectively.

For non-negative indices, the length of a slice is the difference of the indices, if both are within bounds. For example, the length of word[1:3] is 2.

Attempting to use an index that is too large will result in an error:

In [51]:
word[42] #the word Python does not have the 43rd character

IndexError: string index out of range

However, out of range slice indexes are handled gracefully when used for slicing:

In [52]:
word[0:42]

'Python'

Indexing means picking a particular element out of a string (or list)

Slicing means picking a substring out of a string (or list)

Python strings cannot be changed — they are immutable. Therefore, assigning to an indexed position in the string results in an error:

In [54]:
word[0] = 'p'

TypeError: 'str' object does not support item assignment

If you need a different string, you should create a new one:

In [55]:
'p' + word

'pPython'

In [56]:
word[:2] + 'Py'

'PyPy'

The built-in function len() returns the length of a string:

In [57]:
s = 'supercalifragilisticexpialidocious'
len(s)

34