****
## JSON exercise | Data Provided by World Bank

Using data in file 'data/world_bank_projects.json' and the techniques demonstrated above,
1. Find the 10 countries with most projects
2. Find the top 10 major project themes (using column 'mjtheme_namecode')
3. In 2. above you will notice that some entries have only the code and the name is missing. Create a dataframe with the missing names filled in.

In [1]:
#Import libraries
import pandas as pd
import json
from pandas.io.json import json_normalize
import numpy as np

In [2]:
#Read/Save DataFrame
df = pd.read_json('data/world_bank_projects.json')
#Inspect Data
df.info()

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 500 entries, 0 to 499
Data columns (total 50 columns):
_id                         500 non-null object
approvalfy                  500 non-null int64
board_approval_month        500 non-null object
boardapprovaldate           500 non-null object
borrower                    485 non-null object
closingdate                 370 non-null object
country_namecode            500 non-null object
countrycode                 500 non-null object
countryname                 500 non-null object
countryshortname            500 non-null object
docty                       446 non-null object
envassesmentcategorycode    430 non-null object
grantamt                    500 non-null int64
ibrdcommamt                 500 non-null int64
id                          500 non-null object
idacommamt                  500 non-null int64
impagency                   472 non-null object
lendinginstr                495 non-null object
lendinginstrtype            495 non

## 1. Ten Countries with the Most Projects 

In [3]:
#Ten Countries with the Most Projects
df.countryname.value_counts()

Republic of Indonesia                       19
People's Republic of China                  19
Socialist Republic of Vietnam               17
Republic of India                           16
Republic of Yemen                           13
People's Republic of Bangladesh             12
Kingdom of Morocco                          12
Nepal                                       12
Republic of Mozambique                      11
Africa                                      11
Federative Republic of Brazil                9
Islamic Republic of Pakistan                 9
Burkina Faso                                 9
Republic of Tajikistan                       8
United Republic of Tanzania                  8
Republic of Armenia                          8
Kyrgyz Republic                              7
Hashemite Kingdom of Jordan                  7
Lao People's Democratic Republic             7
Federal Republic of Nigeria                  7
Republic of Peru                             6
Islamic State

### It appears that China and Indonesia have the Most Projects through the World Bank

## 2. Find the top 10 major project themes (using column 'mjtheme_namecode').

In [4]:
#Must load json file as string since df cannot handle nested lists within dataframes
df = json.load((open('data/world_bank_projects.json')))
df = json_normalize(df, 'mjtheme_namecode')
df.head()

Unnamed: 0,code,name
0,8,Human development
1,11,
2,1,Economic management
3,6,Social protection and risk management
4,5,Trade and integration


In [5]:
df =pd.DataFrame(df)
df.name.value_counts()

Environment and natural resources management    223
Rural development                               202
Human development                               197
Public sector governance                        184
Social protection and risk management           158
Financial and private sector development        130
                                                122
Social dev/gender/inclusion                     119
Trade and integration                            72
Urban development                                47
Economic management                              33
Rule of law                                      12
Name: name, dtype: int64

Must fill in the missing values for the complete answer

## 3. In 2. above you will notice that some entries have only the code and the name is missing. Create a dataframe with the missing names filled in.

We need to sort values by code so that we can properly align all that is missing with backfill

In [6]:
df = df.sort_values(['code', 'name'])
df.name[df['name'] == ''] = np.nan
df.head(100)

Unnamed: 0,code,name
212,1,
363,1,
1024,1,
1114,1,
1437,1,
2,1,Economic management
88,1,Economic management
175,1,Economic management
204,1,Economic management
205,1,Economic management


In [7]:
df.name.value_counts().head(10)

Environment and natural resources management    223
Rural development                               202
Human development                               197
Public sector governance                        184
Social protection and risk management           158
Financial and private sector development        130
Social dev/gender/inclusion                     119
Trade and integration                            72
Urban development                                47
Economic management                              33
Name: name, dtype: int64

### The top ten project themes are above!