Ref: notebook från Adit Deshpande (https://github.com/adeshpande3)

# Introduction

In [None]:
import pandas as pd
import numpy as np

# Loading in Data

The first step in any ML problem is identifying what format your data is in, and then loading it into whatever framework you're using. 

We're going to be looking at a sports dataset that shows the results from NCAA basketball games from 1985 to 2016. This dataset is in a CSV file, and the function we're going to use to read in the file is called **pd.read_csv()**. This function returns a **dataframe** variable. The dataframe is the golden jewel data structure for Pandas. It is defined as "a two-dimensional size-mutable, potentially heterogeneous tabular data structure with labeled axes (rows and columns)".

Just think of it as a table for now. 

Pandas import can handle most things:
Comma-separated values (CSV),
XLSX,
XML,
DOCX,
ZIP-files,
Text format (.txt),
JSON,
PDF,
HTML,
Images,
MP3,
MP4,
SQL

In [None]:
df = pd.read_csv('data/RegularSeasonCompactResults.csv')

# The Basics

Now that we have our dataframe in our variable df, let's look at what it contains. We can use the function **head()** to see the first couple rows of the dataframe (or the function **tail()** to see the last few rows).

In [None]:
df.head()

In [None]:
df.tail()

We can see the dimensions of the dataframe using the the **shape** attribute

In [None]:
df.shape

We can also extract all the column names as a list, by using the **columns** attribute and can extract the rows with the **index** attribute

In order to get a better idea of the type of data that we are dealing with, we can call the **describe()** function to see statistics like mean, min, etc about each column of the dataset. 

In [None]:
df.describe()

The following is a summary of the 3 data structures used in Pandas

![](DataStructures.png)

# Extracting Rows and Columns

The bracket indexing operator is one way to extract certain columns from a dataframe.

In [None]:
df.columns

In [None]:
df['Season'].head()

In [None]:
df.Season.head()

In [None]:
df[['Wscore', 'Lscore']].head()

Notice that you can acheive the same result by using the loc function. Loc is a veryyyy versatile function that can help you in a lot of accessing and extracting tasks. 

In [None]:
df.loc[:, ['Wscore', 'Lscore']].head()

Note the difference is the return types when you use brackets and when you use double brackets. 

In [None]:
type(df['Wscore'])

In [None]:
type(df[['Wscore']])

You've seen before that you can access columns through df['col name']. You can access rows by using slicing operations. 

In [None]:
df[0:3]

Here's an equivalent using iloc

In [None]:
df.iloc[0:3,:]

# Acessing Values

Then, in order to get attributes about the game, we need to use the **iloc[]** function. Iloc is definitely one of the more important functions. The main idea is that you want to use it whenever you have the integer index of a certain row that you want to access. As per Pandas documentation, iloc is an "integer-location based indexing for selection by position."

In [None]:
df['Wscore'].idxmax()

In [None]:
df.iloc[[df['Wscore'].idxmax()]]

Let's take this a step further. Let's say you want to know the game with the highest scoring winning team (this is what we just calculated), but you then want to know how many points the losing team scored. 

In [None]:
df.iloc[[df['Wscore'].idxmax()]]['Lscore']

When you see data displayed in the above format, you're dealing with a Pandas **Series** object, not a dataframe object.

In [None]:
type(df.iloc[[df['Wscore'].idxmax()]]['Lscore'])

In [None]:
type(df.iloc[[df['Wscore'].idxmax()]])


The other really important function in Pandas is the **loc** function. Contrary to iloc, which is an integer based indexing, loc is a "Purely label-location based indexer for selection by label". Since all the games are ordered from 0 to 145288, iloc and loc are going to be pretty interchangable in this type of dataset

https://stackoverflow.com/questions/31593201/how-are-iloc-ix-and-loc-different

loc gets rows (or columns) with particular labels from the index.
iloc gets rows (or columns) at particular positions in the index (so it only takes integers).


In [None]:
df[:3]

In [None]:
df.iloc[:3]

In [None]:
df.loc[:3]

Notice the slight difference in that iloc is exclusive of the second number, while loc is inclusive. 

If you'd like to see more discussion on how loc and iloc are different, check out this great Stack Overflow post: http://stackoverflow.com/questions/31593201/pandas-iloc-vs-ix-vs-loc-explanation. Just remember that **iloc looks at position** and **loc looks at labels**. Loc becomes very important when your row labels aren't integers. 

In [None]:
s = pd.Series(np.nan, index=[49,48,47,46,45, 1, 2, 3, 4, 5]);s

In [None]:
s.iloc[:3] # slice the first three rows

In [None]:
s.loc[:3] # slice up to and including label 3

In [None]:
s.iloc[:6]

In [None]:
s.loc[:6]

# Useful functions

Okay, so now let's looking at information that we want to extract from the dataframe. Let's say I wanted to know the max value of a certain column. The function **max()** will show you the maximum values of all columns

In [None]:
df.max()

In [None]:
df.dtypes

Then, if you'd like to specifically get the max value for a particular column, you pass in the name of the column using the bracket indexing operator

In [None]:
df['Wscore'].max()

If you'd like to find the mean of the Losing teams' score. 

In [None]:
df['Lscore'].mean()

But what if that's not enough? Let's say we want to actually see the game(row) where this max score happened. We can call the **argmax()** function to identify the row index

In [None]:
df['Wscore'].argmax()

One of the most useful functions that you can call on certain columns in a dataframe is the **value_counts()** function. It shows how many times each item appears in the column. This particular command shows the number of games in each season

In [None]:
df['Season'].value_counts()

In [None]:
df_season = df['Season'].value_counts()
df_season.plot(kind='bar', figsize=(14,7))

# Sorting

Let's say that we want to sort the dataframe in increasing order for the scores of the losing team

In [None]:
df.sort_values('Lscore').head()

In [None]:
#df.groupby('Lscore')

# Filtering Rows Conditionally

Now, let's say we want to find all of the rows that satisy a particular condition. For example, I want to find all of the games where the winning team scored more than 150 points. The idea behind this command is you want to access the column 'Wscore' of the dataframe df (df['Wscore']), find which entries are above 150 (df['Wscore'] > 150), and then returns only those specific rows in a dataframe format (df[df['Wscore'] > 150]).

In [None]:
df[df['Wscore'] > 150]

This also works if you have multiple conditions. Let's say we want to find out when the winning team scores more than 150 points and when the losing team scores below 100. 

In [None]:
df[(df['Wscore'] > 150) & (df['Lscore'] < 100)]

# Grouping

Another important function in Pandas is **groupby()**. This is a function that allows you to group entries by certain attributes (e.g Grouping entries by Wteam number) and then perform operations on them. The following function groups all the entries (games) with the same Wteam number and finds the mean for each group. 

In [None]:
df.groupby('Wteam')['Wscore'].mean().head()

This next command groups all the games with the same Wteam number and finds where how many times that specific team won at home, on the road, or at a neutral site

In [None]:
df.groupby('Wteam')['Wloc'].value_counts().head(9)

Each dataframe has a **values** attribute which is useful because it basically displays your dataframe in a numpy array style format

In [None]:
df.values

Now, you can simply just access elements like you would in an array. 

In [None]:
df.values[0][0]

# Dataframe Iteration

In order to iterate through dataframes, we can use the **iterrows()** function. Below is an example of what the first two rows look like. Each row in iterrows is a Series object

In [None]:
for index, row in df.iterrows():
    print(row)
    if index == 1:
        break

# Data Cleaning

One of the big jobs of doing well in Kaggle competitions is that of data cleaning. A lot of times, the CSV file you're given (especially like in the Titanic dataset), you'll have a lot of missing values in the dataset, which you have to identify. The following **isnull** function will figure out if there are any missing values in the dataframe, and will then sum up the total for each column. In this case, we have a pretty clean dataset.

In [None]:
df.isnull().sum()

If you do end up having missing values in your datasets, be sure to get familiar with these two functions. 
* **dropna()** - This function allows you to drop all(or some) of the rows that have missing values. 
* **fillna()** - This function allows you replace the rows that have missing values with the value that you pass in.

In [None]:
df_travel = pd.read_csv('data/travel-times.csv')

In [None]:
df_travel.head()

In [None]:
df_travel['FuelEconomy'] = pd.to_numeric(df_travel['FuelEconomy'])

In [None]:
df_travel['FuelEconomy'] = df_travel['FuelEconomy'].replace(to_replace='-', value=np.nan)

In [None]:
df_travel['missing FuelEconomy value'] = df_travel['FuelEconomy'].isnull()

In [None]:
df_travel.head()

In [None]:
df_travel["FuelEconomy"].fillna(df_travel['FuelEconomy'].mean(), inplace=True)

In [None]:
df_travel.head(10)

In [None]:
#df_travel['FuelEconomy'] = df_travel['FuelEconomy'].fillna(df_travel['FuelEconomy'].mean()); df_travel.head()

# Visualizing Data

An interesting way of displaying Dataframes is through matplotlib. 

In [None]:
import matplotlib.pyplot as plt
%matplotlib inline

In [None]:
ax = df['Wscore'].plot.hist(bins=20)
ax.set_xlabel('Points for Winning Team')

# Other Useful Functions

* **drop()** - This function removes the column or row that you pass in (You also have the specify the axis). 
* **agg()** - The aggregate function lets you compute summary statistics about each group
* **apply()** - Lets you apply a specific function to any/all elements in a Dataframe or Series
* **get_dummies()** - Helpful for turning categorical data into one hot vectors.
* **drop_duplicates()** - Lets you remove identical rows

# Lots of Other Great Resources

Pandas has been around for a while and there are a lot of other good resources if you're still interested on getting the most out of this library. 
* http://pandas.pydata.org/pandas-docs/stable/10min.html
* https://www.datacamp.com/community/tutorials/pandas-tutorial-dataframe-python
* http://www.gregreda.com/2013/10/26/intro-to-pandas-data-structures/
* https://www.dataquest.io/blog/pandas-python-tutorial/
* https://drive.google.com/file/d/0ByIrJAE4KMTtTUtiVExiUGVkRkE/view
* https://www.youtube.com/playlist?list=PL5-da3qGB5ICCsgW1MxlZ0Hq8LL5U3u9y