# Data Wrangling: Join, Combine, and Reshape
- __in this chapter focuses on tools to help combine, join, and rearrange data__

## Hierarchical Indexing

- __It is one of the most important features of pandas, it enable us to multiple (two or more) index level on axis__
- __It provide a way to work with high dimensional data in lower dimensional form__
- __let's start__

In [4]:
import pandas as pd
import numpy as np

data = pd.Series(np.random.randn(9), 
                 index=[['a', 'a', 'a', 'b', 'b', 'c', 'c', 'd', 'd'],[1, 2, 3, 1, 3, 1, 2, 2, 3]])
data

a  1    0.647778
   2   -0.062371
   3    0.259765
b  1   -0.793568
   3    0.979211
c  1   -0.591073
   2    0.868192
d  2    1.250819
   3   -0.656085
dtype: float64

In [5]:
data.index

MultiIndex([('a', 1),
            ('a', 2),
            ('a', 3),
            ('b', 1),
            ('b', 3),
            ('c', 1),
            ('c', 2),
            ('d', 2),
            ('d', 3)],
           )

In [6]:
data['b']

1   -0.793568
3    0.979211
dtype: float64

In [7]:
data['b':'c']

b  1   -0.793568
   3    0.979211
c  1   -0.591073
   2    0.868192
dtype: float64

In [9]:
data[['b','d']]

b  1   -0.793568
   3    0.979211
d  2    1.250819
   3   -0.656085
dtype: float64

- __Selection is even possible__

In [11]:
data.loc[:,2]

a   -0.062371
c    0.868192
d    1.250819
dtype: float64

- __Hierarchical indexing play a important role in reshaping and group-based operations like forming a pivot table__
- __for example we could rearrange the data into a DataFrame using unstack method__


In [13]:
data.unstack()

Unnamed: 0,1,2,3
a,0.647778,-0.062371,0.259765
b,-0.793568,,0.979211
c,-0.591073,0.868192,
d,,1.250819,-0.656085


- __The inverse operation of unstack is stack:__

In [14]:
data.unstack().stack()

a  1    0.647778
   2   -0.062371
   3    0.259765
b  1   -0.793568
   3    0.979211
c  1   -0.591073
   2    0.868192
d  2    1.250819
   3   -0.656085
dtype: float64

- __Another example__

In [15]:
frame = pd.DataFrame(np.arange(12).reshape((4, 3)),
                     index=[['a', 'a', 'b', 'b'], [1, 2, 1, 2]],
                     columns=[['Ohio', 'Ohio', 'Colorado'],
                              ['Green', 'Red', 'Green']]
                    )
frame

Unnamed: 0_level_0,Unnamed: 1_level_0,Ohio,Ohio,Colorado
Unnamed: 0_level_1,Unnamed: 1_level_1,Green,Red,Green
a,1,0,1,2
a,2,3,4,5
b,1,6,7,8
b,2,9,10,11


In [18]:
#let's see the data base on state

frame['Ohio']

Unnamed: 0,Unnamed: 1,Green,Red
a,1,0,1
a,2,3,4
b,1,6,7
b,2,9,10


In [20]:
frame['Colorado']

Unnamed: 0,Unnamed: 1,Green
a,1,2
a,2,5
b,1,8
b,2,11


## Reordering and Sorting Levels

- __At times you need to rearange the order of the levels on an axis or sort the data__
- __A swaplevel takes two level number of names and returns a new object with levels interchanges__

In [21]:
frame_two = pd.DataFrame(np.arange(12).reshape((4, 3)),
                     index=[['a', 'a', 'b', 'b'], [1, 2, 1, 2]],
                     columns=[['Ohio', 'Ohio', 'Colorado'],
                              ['Green', 'Red', 'Green']]
                    )
frame_two

Unnamed: 0_level_0,Unnamed: 1_level_0,Ohio,Ohio,Colorado
Unnamed: 0_level_1,Unnamed: 1_level_1,Green,Red,Green
a,1,0,1,2
a,2,3,4,5
b,1,6,7,8
b,2,9,10,11


In [29]:
frame_two.index.names = ['key1','key2']
frame_two.columns.names = ['State','Color']

In [30]:
frame_two

Unnamed: 0_level_0,State,Ohio,Ohio,Colorado
Unnamed: 0_level_1,Color,Green,Red,Green
key1,key2,Unnamed: 2_level_2,Unnamed: 3_level_2,Unnamed: 4_level_2
a,1,0,1,2
a,2,3,4,5
b,1,6,7,8
b,2,9,10,11


In [33]:
frame_two.swaplevel('key1','key2')

Unnamed: 0_level_0,State,Ohio,Ohio,Colorado
Unnamed: 0_level_1,Color,Green,Red,Green
key2,key1,Unnamed: 2_level_2,Unnamed: 3_level_2,Unnamed: 4_level_2
1,a,0,1,2
2,a,3,4,5
1,b,6,7,8
2,b,9,10,11


- __'sort_index' funtion sorts the data using the only the values in single level__


In [38]:
frame_two.sort_index(level=1)

Unnamed: 0_level_0,State,Ohio,Ohio,Colorado
Unnamed: 0_level_1,Color,Green,Red,Green
key1,key2,Unnamed: 2_level_2,Unnamed: 3_level_2,Unnamed: 4_level_2
a,1,0,1,2
b,1,6,7,8
a,2,3,4,5
b,2,9,10,11


In [41]:
frame_two.swaplevel(0,1).sort_index(level=1)

Unnamed: 0_level_0,State,Ohio,Ohio,Colorado
Unnamed: 0_level_1,Color,Green,Red,Green
key2,key1,Unnamed: 2_level_2,Unnamed: 3_level_2,Unnamed: 4_level_2
1,a,0,1,2
2,a,3,4,5
1,b,6,7,8
2,b,9,10,11


## Summary Statistics by Level

- __Many descriptive and statistice on DataFrmae and Series have level option in which you can specify the level you want to aggregate by on particular axis__
- __Let's aggregate above DataFrame by level like on rows and columns__

In [43]:
frame_two.sum(level='key2')

State,Ohio,Ohio,Colorado
Color,Green,Red,Green
key2,Unnamed: 1_level_2,Unnamed: 2_level_2,Unnamed: 3_level_2
1,6,8,10
2,12,14,16


In [44]:
frame_two.sum(level='key1')

State,Ohio,Ohio,Colorado
Color,Green,Red,Green
key1,Unnamed: 1_level_2,Unnamed: 2_level_2,Unnamed: 3_level_2
a,3,5,7
b,15,17,19


In [52]:
frame_two.sum(level='Color', axis=1)

Unnamed: 0_level_0,Color,Green,Red
key1,key2,Unnamed: 2_level_1,Unnamed: 3_level_1
a,1,2,1
a,2,8,4
b,1,14,7
b,2,20,10


In [51]:
frame_two.sum(level='State', axis=1)

Unnamed: 0_level_0,State,Ohio,Colorado
key1,key2,Unnamed: 2_level_1,Unnamed: 3_level_1
a,1,1,2
a,2,7,5
b,1,13,8
b,2,19,11


## Indexing with a DataFrame’s columns

- __It’s not unusual to want to use one or more columns from a DataFrame as the row index;__
- __If you want to move the row index into the DataFrame's columns__

In [54]:
frame_three = pd.DataFrame({'a': range(7), 
                      'b': range(7, 0, -1),
                      'c': ['one', 'one', 'one', 'two', 'two','two', 'two'],
                      'd': [0, 1, 2, 0, 1, 2, 3]}
                    )
frame_three

Unnamed: 0,a,b,c,d
0,0,7,one,0
1,1,6,one,1
2,2,5,one,2
3,3,4,two,0
4,4,3,two,1
5,5,2,two,2
6,6,1,two,3


- __DataFrame's 'set_index' function create a new DataFrame using one or more of it's columns as the index__

In [56]:
frame_three.set_index(['c','d'])

Unnamed: 0_level_0,Unnamed: 1_level_0,a,b
c,d,Unnamed: 2_level_1,Unnamed: 3_level_1
one,0,0,7
one,1,1,6
one,2,2,5
two,0,3,4
two,1,4,3
two,2,5,2
two,3,6,1


- __By default the coulmns are removed from the DataFrame__
- __If you don't want to remove then use drop = false__
- __If we 'set_reindex' then it do the opposite of 'set_index'__