# Notebook 7 - pandas
[pandas](http://pandas.pydata.org) provides high-level data structures and functions designed to make working with structured or tabular data fast, easy and expressive. The primary objects in pandas that we will be using are the `DataFrame`, a tabular, column-oriented data structure with both row and column labels, and the `Series`, a one-dimensional labeled array object.

pandas blends the high-performance, array-computing ideas of NumPy with the flexible data manipulation capabilities of spreadsheets and relational databases. It provides sophisticated indexing functionality to make it easy to reshape, slice and perform aggregations.

While pandas adopts many coding idioms from NumPy, the most significant difference is that pandas is designed for working with tabular or heterogeneous data. NumPy, by contrast, is best suited for working with homogeneous numerical array data.
<br>

## Table of Contents:
- [Data Structures](#structures)
    - [Series](#series)
    - [DataFrame](#dataframe)
- [Essential Functionality](#ess_func)
    - [Removing Entries](#removing)
    - [Indexing](#indexing)
    - [Arithmetic Operations](#arithmetic)
- [Summarizing and Computing Descriptive Statistics](#sums)
- [Loading and storing data](#loading)
- [Data Cleaning and preperation](#cleaning)
    - [Handling missing data](#missing)
    - [Data transformation](#transformation)

The common pandas import statment is shown below:

In [1]:
# Common pandas import statement
import numpy as np
import pandas as pd

# 1. Data Structures <a name="structures"></a>
## 1.1 Series <a name="series"></a>
A Series is a one-dimensional array-like object containing a sequence of values and an associated array of data labels called its index.

The easiest way to make a Series is from a list:

In [2]:
data = pd.Series([4, 7, -5, 3])

Now try printing out data

In [3]:
data

0    4
1    7
2   -5
3    3
dtype: int64

The string representation of a Series displayed interactively shows the index on the left and the values on the right. Because we didn't specify an index, the default one is simply integers 0 through N-1.

It is possible to get only the index or only the values of a Serioes with:

In [4]:
data.values

array([ 4,  7, -5,  3])

In [5]:
data.index

RangeIndex(start=0, stop=4, step=1)

You can specify custom indices when intialising the Series

In [6]:
data2 = pd.Series([4, 7, -5, 3], index=["a", "b", "c", "d"])

In [7]:
data2

a    4
b    7
c   -5
d    3
dtype: int64

Another way to think about Series is as a fixed-length ordered dictionary. Furthermore, you can actually define a Series in a similar manner to a dictionary

In [10]:
cities = {"Glasgow" : 599650, "Edinburgh" : True, "Aberdeen" : "196670", "Dundee" : np.array([0,1])}
data3 = pd.Series(cities)

In [11]:
data3

Glasgow      599650
Edinburgh      True
Aberdeen     196670
Dundee       [0, 1]
dtype: object

## 1.2 DataFrame <a name="dataframe"></a>
A DataFrame represents a rectangular table of data and contains an ordered collection of columns, each of which can be a different value type. The DataFrame has both row and column index and can be thought of as a dict of Series all sharing the same index.

The most common way to create a DataFrame is with dicts:

In [12]:
data = {"cities" : ["Glasgow", "Edinburgh", "Aberdeen", "Dundee"],
        "population" : [599650, 464990, 196670, 147710],
        "year" : [2011, 2013, 2013, 2013]}
frame = pd.DataFrame(data)

Try printing it out

In [13]:
frame

Unnamed: 0,cities,population,year
0,Glasgow,599650,2011
1,Edinburgh,464990,2013
2,Aberdeen,196670,2013
3,Dundee,147710,2013


Jupyter Notebooks prints it out in a nice table but the basic version of this is also just as readable!

Additionally you can also specify the order of columns, or the row index,  during initialisation

In [14]:
frame2 = pd.DataFrame(data, columns=["year", "cities", "population"], index=["a", "b", "c", "d"])
frame2

Unnamed: 0,year,cities,population
a,2011,Glasgow,599650
b,2013,Edinburgh,464990
c,2013,Aberdeen,196670
d,2013,Dundee,147710


You can retrieve a particular column from a DataFrame with
```python
frame["cities"]
```
The result is going to be a Series.

In [15]:
frame2["cities"]

a      Glasgow
b    Edinburgh
c     Aberdeen
d       Dundee
Name: cities, dtype: object

It is also possible to add and modify the columns of a DataFrame

In [16]:
frame2["area"] = 100
frame2

Unnamed: 0,year,cities,population,area
a,2011,Glasgow,599650,100
b,2013,Edinburgh,464990,100
c,2013,Aberdeen,196670,100
d,2013,Dundee,147710,100


In [17]:
frame2["area"] = [175, 264, 65.1, 60]  # in km^2
frame2

Unnamed: 0,year,cities,population,area
a,2011,Glasgow,599650,175.0
b,2013,Edinburgh,464990,264.0
c,2013,Aberdeen,196670,65.1
d,2013,Dundee,147710,60.0


Here is a table of different ways of initialising a DataFrame for your reference

| Type | Notes |
| --- | --- |
| 2D ndarray | A matrix of data; passing optional row and column labels |
| dict of arrays, lists, or tuples | Each sequence becomes a column in the DataFrame; all sequences must be the same length |
| dict of Series | Each value becomes a column; indexes from each Series are unioned together to<br>form the result's row index if not explicit index is passed |
| dict of dicts | Each inner dict becomes a column; keys are unioned to form the row<br>index as in the "dict of Series" case |
| List of dicts or Series | Each item becomes a row in the DataFrame; union of dict keys or<br>Series indices becomes the DataFrame's column labels |
| List of lists or tuples | Treated as the "2D ndarray" case |

# 2. Essential Functionality <a name="ess_func"></a>
In this section, we will go through the fundamental mechanics of interacting with the data contained in a Series or DaraFrame.

## 2.1Removing columns/indices <a name="removing"></a>
Similar to above, it is easy to remove entries. This is done with the `drop()` method and can be applied to both columns and indices:

In [18]:
# define new DataFrame
data = np.reshape(np.arange(9), (3,3))
df = pd.DataFrame(data, index=["a", "b", "c"],
                  columns=["Edinburgh", "Glasgow", "Aberdeen"])

df

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5
c,6,7,8


In [19]:
df.drop("b")  # remove row (index)

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
c,6,7,8


In [20]:
# You can also drop from a column
df.drop(["Aberdeen", "Edinburgh"], axis="columns")

Unnamed: 0,Glasgow
a,1
b,4
c,7


Note that the original data frame is unchanged: `df.drop()` gives us a new data frame with the desired data dropped, and leaves the original data intact. We can ask `.drop()` to operate directly on the original data frame by setting the argument `inplace=True`.

In [21]:
df

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5
c,6,7,8


## 2.2 Indexing <a name="indexing"></a>

Indexing in pandas works simiarly to numpy, but we have to use `.iloc` instead of just normal indexing

In [22]:
data = np.reshape(np.arange(9), (3,3))
df = pd.DataFrame(data, index=["a", "b", "c"],
                  columns=["Edinburgh", "Glasgow", "Aberdeen"])

df

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5
c,6,7,8


If you want to get the data of column. You can just use the column name.

In [23]:
df["Edinburgh"]  # Use column names to get the column

a    0
b    3
c    6
Name: Edinburgh, dtype: int64

But what if we want to get the data of rows?

In [24]:
df[2]   # this won't work

KeyError: 2

In [25]:
df.iloc[2]  # this works and return a series

Edinburgh    6
Glasgow      7
Aberdeen     8
Name: c, dtype: int64

You can also slice a dataframe as usual:

In [26]:
df.iloc[:2]

Unnamed: 0,Edinburgh,Glasgow,Aberdeen
a,0,1,2
b,3,4,5


You can index into many dimensions as seen in NumPy:

In [27]:
df.iloc[2,0]

6

Finally, when the indices/column names of your data frame are not integers, you can use `.loc` to index:

In [28]:
frame2

Unnamed: 0,year,cities,population,area
a,2011,Glasgow,599650,175.0
b,2013,Edinburgh,464990,264.0
c,2013,Aberdeen,196670,65.1
d,2013,Dundee,147710,60.0


In [29]:
frame2.loc["c", "area"]

65.1

In [30]:
frame2.iloc[2,3]

65.1

## 2.3 Arithmetic <a name="arithmetic"></a>
When you are performing arithmetic operations between two objects, if any index/column pairs are not the same, the respective indexes/columns in the result will be the union of the index pair. Let's have a look

In [31]:
df1 = pd.DataFrame(np.arange(12).reshape((3,4)),
                  columns=list("abcd"))
df1

Unnamed: 0,a,b,c,d
0,0,1,2,3
1,4,5,6,7
2,8,9,10,11


In [32]:
df2 = pd.DataFrame(np.arange(12).reshape((3,4)),
                  columns=list("cdef"))
df2

Unnamed: 0,c,d,e,f
0,0,1,2,3
1,4,5,6,7
2,8,9,10,11


In [33]:
# adding the two
df1+df2

Unnamed: 0,a,b,c,d,e,f
0,,,2,4,,
1,,,10,12,,
2,,,18,20,,


Notice how where we don't have matching values from `df1` and `df2` the output of the addition operation is `NaN` since there are no two numbers to add.

Well, we can "fix" that by filling in the `NaN` values. This effectively tells pandas where there are no two values to add, assume that the missing value is just zero.

In [34]:
df1.add(df2, fill_value=0)

Unnamed: 0,a,b,c,d,e,f
0,0.0,1.0,2,4,2.0,3.0
1,4.0,5.0,10,12,6.0,7.0
2,8.0,9.0,18,20,10.0,11.0


Here's a list of all arithmetic operations within panda. They are similar to NumPy!

| Operator | Method | Description |
| -- | -- | -- |
| + | add | Addition |
| - | sub | Subtraction |
| / | div | Division |
| // | floordiv | Floor division |
| * | mul | Multiplication |
| ** | pow | Exponentiation |


## 2.4 Summarizing and computing descriptive stats <a name="sums"></a>
`pandas` is equipped with common mathematical and statistical methods. Most of which fall into the category of reductions or summary statistics. These are methods that extract a single value from a list of values. For example, you can extract the sum of a `Series` object like this:

In [35]:
df = pd.DataFrame(np.arange(20).reshape(5,4),
                 columns=["a", "b", "c", "d"])
df

Unnamed: 0,a,b,c,d
0,0,1,2,3
1,4,5,6,7
2,8,9,10,11
3,12,13,14,15
4,16,17,18,19


In [36]:
df.sum()

a    40
b    45
c    50
d    55
dtype: int64

Notice how that created the sum of each column?

Well you can actually make that the other way around by adding an extra option to `sum()`

In [37]:
df.sum(axis="columns")

0     6
1    22
2    38
3    54
4    70
dtype: int64

A similar method also exists for obtaining the mean of data:

In [39]:
df.mean(axis="columns")

0     1.5
1     5.5
2     9.5
3    13.5
4    17.5
dtype: float64

Finally, the mother of the methods we discussed here is `describe()` 

In [40]:
df.describe()

Unnamed: 0,a,b,c,d
count,5.0,5.0,5.0,5.0
mean,8.0,9.0,10.0,11.0
std,6.324555,6.324555,6.324555,6.324555
min,0.0,1.0,2.0,3.0
25%,4.0,5.0,6.0,7.0
50%,8.0,9.0,10.0,11.0
75%,12.0,13.0,14.0,15.0
max,16.0,17.0,18.0,19.0


Here are some of the summary methods:

| Method | Description |
| -- | -- |
| count          | Return number of non-NA values |
| describe       | Set of summary statistics |
| min, max       | Minimum, maximum values |
| argmin, argmax | Index locations at which the minimum or maximum value is obtained | 
| sum            | Sum of values |
| mean           | Mean of values |
| median         | Arithmetic median of values |
| std            | Sample standard deviation of values
| value_counts() | Counts the number of occurrences of each unique element in a column |

### Exercise 1
A dataset of random numbers is created below. Obtain all rows starting from row 85 to 97.

*Note: Remember that Python uses 0-based indexing*

In [41]:
df = pd.DataFrame(np.reshape(np.arange(10000), (100,100)))
df.iloc[84:97]

Unnamed: 0,0,1,2,3,4,5,6,7,8,9,...,90,91,92,93,94,95,96,97,98,99
84,8400,8401,8402,8403,8404,8405,8406,8407,8408,8409,...,8490,8491,8492,8493,8494,8495,8496,8497,8498,8499
85,8500,8501,8502,8503,8504,8505,8506,8507,8508,8509,...,8590,8591,8592,8593,8594,8595,8596,8597,8598,8599
86,8600,8601,8602,8603,8604,8605,8606,8607,8608,8609,...,8690,8691,8692,8693,8694,8695,8696,8697,8698,8699
87,8700,8701,8702,8703,8704,8705,8706,8707,8708,8709,...,8790,8791,8792,8793,8794,8795,8796,8797,8798,8799
88,8800,8801,8802,8803,8804,8805,8806,8807,8808,8809,...,8890,8891,8892,8893,8894,8895,8896,8897,8898,8899
89,8900,8901,8902,8903,8904,8905,8906,8907,8908,8909,...,8990,8991,8992,8993,8994,8995,8996,8997,8998,8999
90,9000,9001,9002,9003,9004,9005,9006,9007,9008,9009,...,9090,9091,9092,9093,9094,9095,9096,9097,9098,9099
91,9100,9101,9102,9103,9104,9105,9106,9107,9108,9109,...,9190,9191,9192,9193,9194,9195,9196,9197,9198,9199
92,9200,9201,9202,9203,9204,9205,9206,9207,9208,9209,...,9290,9291,9292,9293,9294,9295,9296,9297,9298,9299
93,9300,9301,9302,9303,9304,9305,9306,9307,9308,9309,...,9390,9391,9392,9393,9394,9395,9396,9397,9398,9399


### Exercise 2
Create a (3,3) DataFrame and square all elements in it.

In [42]:
df = pd.DataFrame(np.reshape(np.arange(9), (3,3)))
df**2

Unnamed: 0,0,1,2
0,0,1,4
1,9,16,25
2,36,49,64


### Exercise 3

A random DataFrame is created below. Find the median value of each column.

In [43]:
df = pd.DataFrame(np.random.uniform(0, 10, (100, 10)))
df.median()

0    4.616608
1    5.000067
2    5.852447
3    4.721273
4    5.153451
5    4.985477
6    5.205821
7    5.052857
8    4.413456
9    5.626979
dtype: float64

# 3. Data Loading and Storing <a name="loading"></a>
Accessing data is a necessary first step for data science.

A very common format is the `.csv`. This is an easy to read file format which is usually visualised like a spreadsheet. The data itself is usually separated with a `,` which is called the **delimiter**.

Here is an example of a `.csv` file:
```
name,sex
Joseph,M	
Andrew,M	
Joshua,M	
Abigail,F	
Benjamin,M	
Anthony,M	
James,M	
Emily,F	
Elijah,M	
Matthew,M	
Daniel,M	
Aiden,M	
Alexander,M	
Ava,F	
Jayden,M	
Michael,M	
Liam,M	
William,M	
Olivia,F	
Noah,M	
```

It details some names and sexes of newborn babies. The first line is called the header, and you can imagine that it is the name of the columns of a spreadsheet.

Let's now see how we can load this data and analyse it. The file is located in the current folder and is called `yob2012.csv`. We can read it like this:

In [44]:
babes = pd.read_csv("./yob2012.csv")
babes

Unnamed: 0,name,sex
0,Sophia,F
1,Emma,F
2,Isabella,F
3,Olivia,F
4,Ava,F
...,...,...
33679,Zylin,M
33680,Zymari,M
33681,Zyrin,M
33682,Zyrus,M


Easy right?

### Exercise 4
Load the file `./homes.csv` and find the mean selling price of these houses.

In [46]:
prices = pd.read_csv('./homes.csv')
print(prices)
prices.loc[:,'Sell'].mean()

    Sell  List  Living  Rooms  Beds  Baths  Age  Acres  Taxes
0    142   160      28     10     5      3   60      0   3167
1    175   180      18      8     4      1   12      0   4033
2    129   132      13      6     3      1   41      0   1471
3    138   140      17      7     3      1   22      0   3204
4    232   240      25      8     4      3    5      2   3613
5    135   140      18      7     4      3    9      1   3028
6    150   160      20      8     4      3   18      4   3131
7    207   225      22      8     4      2   16      2   5158
8    271   285      30     10     5      2   30      1   5702
9     89    90      10      5     3      1   43      0   2054
10   153   157      22      8     3      3   18      0   4127
11    87    90      16      7     3      1   50      1   1445
12   234   238      25      8     4      2    2      2   2087
13   106   116      20      8     4      1   13      0   2818
14   175   180      22      8     4      2   15      2   3917
15   165

170.36

###### The `read_csv` function has a lot of optional arguments (more than 50). It's impossible to memorise all of them - it's usually best just to look up the particular functionality when you need it. 

You can search `pandas read_csv` online and find all of the documentation.

There are also many other functions that can read textual data. Here are some of them:

| Function | Description
| -- | -- |
| read_csv       | Load delimited data from a file, URL, or file-like object. The default delimiter is a comma `,` |
| read_table     | Load delimited data from a file, URL, or file-like object. The default delimiter is tab `\t` |
| read_clipboard | Reads the last object you have copied (Ctrl-C) |
| read_excel     | Read tabular data from Excel XLS or XLSX file |
| read_hdf       | Read HDF5 file written by pandas |
| read_html      | Read all tables found in the given HTML document |
| read_json      | Read data from a JSON string representation |
| read_sql       | Read the results of a SQL query |

*Note: there are also other loading functions which are not touched upon here*

# 4. Data Cleaning <a name="cleaning"></a>
While doing data analysis and modeling, a significant amount of time is spent on data preparation: loading, cleaning, transforming and rearranging. Such tasks are often reported to take **up to 80%** or more of a data analyst's time. Often the way the data is stored in files isn't in the correct format and needs to be modified. Researchers usually do this using programming languages like Python.

In this chapter, we will discuss tools for handling missing data, duplicate data, string manipulation, and some other analytical data transformations.

## 4.1 Handling missing data <a name="missing"></a>
Missing data occurs commonly in many data analysis applications. One of the goals of pandas is to make working with missing data as painless as possible.

In pandas, missing numeric data is represented by `NaN` (Not a Number) and can easily be handled:

In [47]:
string_data = pd.Series(['orange', 'tomato', np.nan, 'avocado'])
string_data

0     orange
1     tomato
2        NaN
3    avocado
dtype: object

In [48]:
string_data.isnull()

0    False
1    False
2     True
3    False
dtype: bool

The pandas `NaN` is functionally equlevant to the standard Python type `NoneType` which can be defined with `x = None`.

In [49]:
string_data.iloc[3] = None
string_data.isnull()

0    False
1    False
2     True
3     True
dtype: bool

Here are some other methods which you can find useful:
    
| Method | Description |
| -- | -- |
| dropna | Filter axis labels based on whether the values of each label have missing data|
| fillna | Fill in missing data with some value |
| isnull | Return boolean values indicating which values are missing |
| notnull | Negation of isnull |

Just like `.drop()`, these methods all return a new object, leaving the original unchanged (this behaviour can be overridden using the argument `inplace=True`).

### Exercise 5
Remove the missing data below using the appropriate method

In [50]:
data = pd.Series([1, None, 3, 4, None, 6])
data

0    1.0
1    NaN
2    3.0
3    4.0
4    NaN
5    6.0
dtype: float64

In [51]:
data.dropna()

0    1.0
2    3.0
3    4.0
5    6.0
dtype: float64

`dropna()` by default removes any row/column that has a missing value. What if we want to remove only rows in which all of the data is missing though?

In [52]:
data = pd.DataFrame([[1., 6.5, 3.], [1., None, None],
                    [None, None, None], [None, 6.5, 3.]])
data

Unnamed: 0,0,1,2
0,1.0,6.5,3.0
1,1.0,,
2,,,
3,,6.5,3.0


In [53]:
data.dropna()

Unnamed: 0,0,1,2
0,1.0,6.5,3.0


In [54]:
data.dropna(how="all")

Unnamed: 0,0,1,2
0,1.0,6.5,3.0
1,1.0,,
3,,6.5,3.0


### Exercise 6
That's fine if we want to remove missing data, what if we want to fill in missing data? Do you know of a way? Try to fill in all of the missing values from the data below with **0s**

In [55]:
data = pd.DataFrame([[1., 6.5, 3.], [2., None, None],
                    [None, None, None], [None, 1.5, 9.]])
data

Unnamed: 0,0,1,2
0,1.0,6.5,3.0
1,2.0,,
2,,,
3,,1.5,9.0


In [56]:
data.fillna(0)

Unnamed: 0,0,1,2
0,1.0,6.5,3.0
1,2.0,0.0,0.0
2,0.0,0.0,0.0
3,0.0,1.5,9.0


## 4.2 Data Transformation  <a name="transformation"></a>
### Removing duplicates
Duplicate data can be a serious issue, luckily pandas offers a simple way to remove duplicates

In [57]:
data = pd.DataFrame([1, 2, 3, 4, 3, 2, 1])
data

Unnamed: 0,0
0,1
1,2
2,3
3,4
4,3
5,2
6,1


In [58]:
data.drop_duplicates()

Unnamed: 0,0
0,1
1,2
2,3
3,4


You can also select which rows to keep

In [59]:
data.drop_duplicates(keep="last")

Unnamed: 0,0
3,4
4,3
5,2
6,1


### Replacing data
You've already seen how you can fill in missing data with `fillna()`. That is actually a special case of more general value replacement. That is done via the `replace()` method.

Let's consider an example where the dataset given to us had `-999` as a mark for missing data instead of `NaN`.

In [60]:
data = pd.DataFrame([1., -999., 2., -999., 3., 4., -999, -999, 7.])
data

Unnamed: 0,0
0,1.0
1,-999.0
2,2.0
3,-999.0
4,3.0
5,4.0
6,-999.0
7,-999.0
8,7.0


In [61]:
data.replace(-999, np.nan)

Unnamed: 0,0
0,1.0
1,
2,2.0
3,
4,3.0
5,4.0
6,
7,
8,7.0


### Detecting and Filtering Outliers
Filtering or transforming outliers is largely a matter of applying array operations. Consider a DataFrame with some normally distributed data:

In [62]:
data = pd.DataFrame(np.random.randn(50, 4))
data.describe()

Unnamed: 0,0,1,2,3
count,50.0,50.0,50.0,50.0
mean,-0.05423,-0.036326,0.100026,-0.240913
std,1.019708,0.931018,1.065112,0.968294
min,-2.489258,-1.735833,-2.118575,-2.168071
25%,-0.779301,-0.689771,-0.414737,-0.856904
50%,-0.050181,-0.116291,-0.04997,-0.36493
75%,0.539847,0.442804,0.632083,0.315485
max,2.054866,2.76971,2.719965,2.724311


Suppose you now want to lower all values exceeding 2 from one of the columns.

In [63]:
col = data[3]
col[col > 2]

39    2.724311
Name: 3, dtype: float64

We could set a ceiling for these values at 2.

In [64]:
data[data > 2] = 2
data.describe()

Unnamed: 0,0,1,2,3
count,50.0,50.0,50.0,50.0
mean,-0.055327,-0.051721,0.071125,-0.2554
std,1.017419,0.889093,1.003598,0.92759
min,-2.489258,-1.735833,-2.118575,-2.168071
25%,-0.779301,-0.689771,-0.414737,-0.856904
50%,-0.050181,-0.116291,-0.04997,-0.36493
75%,0.539847,0.442804,0.632083,0.315485
max,2.0,2.0,2.0,2.0


What we just did is called **boolean indexing**.

In [66]:
data = pd.DataFrame(np.random.randn(50, 4), columns=["F1", "F2", "F3", "Label"])
print(data)

data["Label"][data["Label"] > 2] = np.nan
print(data)

new_data = data.dropna()
print("The length of data is : ", len(data))
print("The length of new_data is : ", len(new_data))

          F1        F2        F3     Label
0  -1.283309 -0.059069 -0.179424 -0.076210
1   1.361824 -1.030212  1.247154 -0.227082
2  -0.166962 -1.050645  0.700060  0.305703
3  -0.095045 -0.147598 -1.361449 -0.002426
4  -0.856997 -0.728431 -1.214735 -0.662769
5  -2.625800  0.058670  0.164477 -0.539076
6   0.589927 -0.950813 -0.561226 -1.796971
7  -0.146477 -0.666354 -0.239791 -0.689825
8   0.797708  0.207686  0.207607  1.206474
9   0.641967 -0.027258  0.164671 -0.960228
10  1.436676  0.608191 -0.923032  1.578068
11  0.304771 -1.533107  0.988270 -0.425986
12  0.378705  1.599103  0.592270  1.176312
13  0.544053  0.241947 -1.837720  0.904436
14  2.052230 -1.353009  0.772985  0.658189
15 -0.810750 -0.412684  1.274516 -0.736468
16  1.236898  1.954967 -0.734401 -0.215669
17 -0.233633  0.733871  0.335801 -0.177351
18  1.165159  0.918069 -0.031935  0.006165
19 -1.079139 -1.394755 -0.179080  0.698132
20  0.270777 -0.655113 -1.403339 -0.715509
21  0.172754  1.272820  1.669479  0.741634
22 -0.88108

### Exercise 7
Let's load again our file with home prices and filter out homes based on our preference:
1. Load up the file `./homes.csv`
2. The data contains some duplicates. Filter them out.
3. Let's say that the most we can spend on a house is £150. Keep only houses that have a **sell**ing price less than £150 and remove the rest
4. Select only houses that have 4 or more bedrooms
5. Select only houses that have 3 or more baths

You should end up with only 2 houses

In [67]:
homes = pd.read_csv('./homes.csv')
homes = homes.drop_duplicates()
homes

Unnamed: 0,Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
0,142,160,28,10,5,3,60,0,3167
1,175,180,18,8,4,1,12,0,4033
2,129,132,13,6,3,1,41,0,1471
3,138,140,17,7,3,1,22,0,3204
4,232,240,25,8,4,3,5,2,3613
5,135,140,18,7,4,3,9,1,3028
6,150,160,20,8,4,3,18,4,3131
7,207,225,22,8,4,2,16,2,5158
8,271,285,30,10,5,2,30,1,5702
9,89,90,10,5,3,1,43,0,2054


In [68]:
homes1 = homes[homes['Sell']<150]
homes1

Unnamed: 0,Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
0,142,160,28,10,5,3,60,0,3167
2,129,132,13,6,3,1,41,0,1471
3,138,140,17,7,3,1,22,0,3204
5,135,140,18,7,4,3,9,1,3028
9,89,90,10,5,3,1,43,0,2054
11,87,90,16,7,3,1,50,1,1445
13,106,116,20,8,4,1,13,0,2818
17,136,140,19,7,3,1,22,1,3607
18,148,160,17,7,3,2,13,0,3648
26,110,115,16,8,4,1,26,0,3103


In [69]:
homes2 = homes1[homes1['Beds']>=4]
homes2

Unnamed: 0,Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
0,142,160,28,10,5,3,60,0,3167
5,135,140,18,7,4,3,9,1,3028
13,106,116,20,8,4,1,13,0,2818
26,110,115,16,8,4,1,26,0,3103
27,135,145,18,7,4,1,35,0,3363
37,127,130,20,8,4,1,65,0,3334
40,128,135,17,9,4,1,25,1,3374
41,110,120,15,8,4,2,11,1,3119
42,123,130,18,8,4,2,43,0,3268
44,145,145,18,8,4,2,44,0,2783


In [70]:
homes3 = homes2[homes2['Baths']>=3]
homes3

Unnamed: 0,Sell,List,Living,Rooms,Beds,Baths,Age,Acres,Taxes
0,142,160,28,10,5,3,60,0,3167
5,135,140,18,7,4,3,9,1,3028
