Skip to content

Pandas cheat sheet

Zakir Syed edited this page Feb 13, 2019 · 20 revisions

Data Frame - cleaning column names

df.columns = df.columns.str.replace(r'\s+', '_') #This replaces ' ' with '_' in columns  

df.columns is of type index hence can be converted to string and cleaned up in a chained fashion

df.columns = df.columns.str.replace('.', '_').str.replace('(','').replace(')','')

Data Frame - Efficiently iterating over all rows

for row in df.itertuples():
    print(f"{row.Col_1} : {row.Col_2}")

Data Frame - Filtration

This Stack Overflow Link has in-depth analysis with performance plotted in a great amount of detail. Here are a few convenient ways:-

df[df['col'] == val] # Slows down with df length
df[df['col'].values == val] # np version scales up very well
df.query('col == val') # Scales up well

Data Frame - Filtration

This Stack Overflow Link has in-depth analysis with performance plotted in a great amount of detail. Here are a few convenient ways:-

df[df['col'] == val] # Slows down with df length
df[df['col'].values == val] # np version scales up very well
df.query('col == val') # Scales up well

Data Frame - Slicing

df.loc['row_1':'row_2', 'col_1':'col_2'] #Label based
df.iloc[10:20, 2:6] #Index based    

Data Frame - Delete/Insert column

df.drop(['col_name'], axis=1, inplace=True)
df.insert(loc=3, column='col_name', value=new_col)

Data Frame - Transform categorical features to numeric

embarked_dummies = pd.get_dummies(df['Embarked'], prefix='Embarked@')
embarked_dummies.head()

data frame before

df = pd.merge(df, embarked_dummies, left_index=True, right_index=True).drop(['Embarked'], axis=1)
df.head()

data frame after


Clone this wiki locally