
# World Bank Projects (JSON Exercise)
****
**Goal::**
Using the world bank projects dataset (in file 'dataset/world_bank_projects.json'), provide a solution for below queries.

1. Find the 10 countries with most projects
2. Find the top 10 major project themes (using column 'mjtheme_namecode')
3. In point 2 above, there are some entries that have only the code and the name is missing. Create a dataframe with the missing names filled in.

### Let's start with loading the dataset and taking a look at the data


In [227]:
#Importing the pandas and json packages
import pandas as pd
import json
from pandas.io.json import json_normalize
from pandas.io.json import to_json

In [228]:
# Read Json data into a DataFrame
df = pd.read_json('dataset/world_bank_projects.json')

In [229]:
df.shape

(500, 50)

In [230]:
df.head()

Unnamed: 0,_id,approvalfy,board_approval_month,boardapprovaldate,borrower,closingdate,country_namecode,countrycode,countryname,countryshortname,...,sectorcode,source,status,supplementprojectflg,theme1,theme_namecode,themecode,totalamt,totalcommamt,url
0,{'$oid': '52b213b38594d8a2be17c780'},1999,November,2013-11-12T00:00:00Z,FEDERAL DEMOCRATIC REPUBLIC OF ETHIOPIA,2018-07-07T00:00:00Z,Federal Democratic Republic of Ethiopia!$!ET,ET,Federal Democratic Republic of Ethiopia,Ethiopia,...,"ET,BS,ES,EP",IBRD,Active,N,"{'Percent': 100, 'Name': 'Education for all'}","[{'code': '65', 'name': 'Education for all'}]",65,130000000,130000000,http://www.worldbank.org/projects/P129828/ethi...
1,{'$oid': '52b213b38594d8a2be17c781'},2015,November,2013-11-04T00:00:00Z,GOVERNMENT OF TUNISIA,,Republic of Tunisia!$!TN,TN,Republic of Tunisia,Tunisia,...,"BZ,BS",IBRD,Active,N,"{'Percent': 30, 'Name': 'Other economic manage...","[{'code': '24', 'name': 'Other economic manage...",5424,0,4700000,http://www.worldbank.org/projects/P144674?lang=en
2,{'$oid': '52b213b38594d8a2be17c782'},2014,November,2013-11-01T00:00:00Z,MINISTRY OF FINANCE AND ECONOMIC DEVEL,,Tuvalu!$!TV,TV,Tuvalu,Tuvalu,...,TI,IBRD,Active,Y,"{'Percent': 46, 'Name': 'Regional integration'}","[{'code': '47', 'name': 'Regional integration'...",52812547,6060000,6060000,http://www.worldbank.org/projects/P145310?lang=en
3,{'$oid': '52b213b38594d8a2be17c783'},2014,October,2013-10-31T00:00:00Z,MIN. OF PLANNING AND INT'L COOPERATION,,Republic of Yemen!$!RY,RY,Republic of Yemen,"Yemen, Republic of",...,JB,IBRD,Active,N,"{'Percent': 50, 'Name': 'Participation and civ...","[{'code': '57', 'name': 'Participation and civ...",5957,0,1500000,http://www.worldbank.org/projects/P144665?lang=en
4,{'$oid': '52b213b38594d8a2be17c784'},2014,October,2013-10-31T00:00:00Z,MINISTRY OF FINANCE,2019-04-30T00:00:00Z,Kingdom of Lesotho!$!LS,LS,Kingdom of Lesotho,Lesotho,...,"FH,YW,YZ",IBRD,Active,N,"{'Percent': 30, 'Name': 'Export development an...","[{'code': '45', 'name': 'Export development an...",4145,13100000,13100000,http://www.worldbank.org/projects/P144933/seco...


<br>
We can see that the dataset has 500 observations and 50 attributes. Let’s move ahead with our ** Query # 1 : Finding 10 countries with the most number of projects.**

We will analyze the features first and decide which features could be useful for our solution. We also need to see if there is any missing value in chosen attributes and how we can take care of that.

In [231]:
# check the DataFrame details
df.info()

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 500 entries, 0 to 499
Data columns (total 50 columns):
_id                         500 non-null object
approvalfy                  500 non-null int64
board_approval_month        500 non-null object
boardapprovaldate           500 non-null object
borrower                    485 non-null object
closingdate                 370 non-null object
country_namecode            500 non-null object
countrycode                 500 non-null object
countryname                 500 non-null object
countryshortname            500 non-null object
docty                       446 non-null object
envassesmentcategorycode    430 non-null object
grantamt                    500 non-null int64
ibrdcommamt                 500 non-null int64
id                          500 non-null object
idacommamt                  500 non-null int64
impagency                   472 non-null object
lendinginstr                495 non-null object
lendinginstrtype            495 non

In [232]:
# Find total number of countries and projects
df[['countryname','project_name']].count()

countryname     500
project_name    500
dtype: int64

In [233]:
# Find unique number of countries and projects
df[['countryname', 'project_name']].nunique()

countryname     118
project_name    500
dtype: int64

It looks like columns **countryname** and **project_name** can be used here. These columns have no missing values either. No data imputing is needed and let’s  straight away go to finding  the solution. 

One interesting part to note here is that each observation represents a project. There are 500 projects which involves 118 countries. We need to find out top 10 countries with most projects. We can do this by grouping the data by 'countryname' and aggregate it on count of rows. Sorting it in descending order would give us the top countries.

Another way to get the same result would be to use value_counts() on 'countryname'.

Lets see both the approaches now.
<br>

In [234]:
#Using groupby
top_countries = df.groupby('countryname')['project_name'].count().sort_values(ascending = False)
top_countries.head(10)

countryname
People's Republic of China         19
Republic of Indonesia              19
Socialist Republic of Vietnam      17
Republic of India                  16
Republic of Yemen                  13
Nepal                              12
People's Republic of Bangladesh    12
Kingdom of Morocco                 12
Africa                             11
Republic of Mozambique             11
Name: project_name, dtype: int64

In [235]:
#Using value_counts
df['countryname'].value_counts().sort_values(ascending = False).head(10)

Republic of Indonesia              19
People's Republic of China         19
Socialist Republic of Vietnam      17
Republic of India                  16
Republic of Yemen                  13
People's Republic of Bangladesh    12
Nepal                              12
Kingdom of Morocco                 12
Africa                             11
Republic of Mozambique             11
Name: countryname, dtype: int64

### Query # 2 : Find the top 10 major project themes (using column 'mjtheme_namecode') 

As we can see that column **mjtheme_namecode** is a list of json objects itself, we would need to flatten it into a dataframe in order to work on it.

In [237]:
# function to flatten the theme json data
def flatten_theme():
    themeJsonList = df['mjtheme_namecode'].tolist()
    dfs=[json_normalize(themeJson)for themeJson in themeJsonList]
    return pd.concat(dfs, ignore_index=True)

In [238]:
#flatten the theme and create a dataframe
theme_df=flatten_theme_test()
theme_df.head()

Unnamed: 0,code,name
0,8,Human development
1,11,
2,1,Economic management
3,6,Social protection and risk management
4,5,Trade and integration


In [239]:
# Check for the NaN values
theme_df.info()

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 1499 entries, 0 to 1498
Data columns (total 2 columns):
code    1499 non-null object
name    1499 non-null object
dtypes: object(2)
memory usage: 23.5+ KB


In [240]:
#Check for the blank code and name
print('blank code:')
print(theme_df[theme_df.code==''].count())
print()
print('blank name:')
print(theme_df[theme_df.name==''].count())

blank code:
code    0
name    0
dtype: int64

blank name:
code    122
name    122
dtype: int64



We can see there are 122 missing entries for name. For now, we will ignore the blank names while generating our report. In query # 3, we will see how to handle these missing names.
</br>

In [241]:
# Get the top 10 theme names
topThemes = theme_df[theme_df.name != ''].name.value_counts().sort_values(ascending = False)
topThemes.head(10)

Environment and natural resources management    223
Rural development                               202
Human development                               197
Public sector governance                        184
Social protection and risk management           158
Financial and private sector development        130
Social dev/gender/inclusion                     119
Trade and integration                            72
Urban development                                47
Economic management                              33
Name: name, dtype: int64

### Query # 3 : In query 2 above, there are some theme observations where the name is missing. Create a dataframe with the missing names filled in.

My approach would be finding out relation between code and name. If each code has a unique non-blank project name, we can easily fill the missing project names based on the code.

In [242]:
# find count of unique code 
theme_df['code'].nunique()

11

We have 11 codes for the project themes. Let's find out the corresponding project names.

In [243]:
# Find the unique combination of code and name 
unique_theme_df = theme_df.drop_duplicates()
unique_theme_df = unique_theme_df[unique_theme_df.name != '']

unique_theme_df

Unnamed: 0,code,name
0,8,Human development
2,1,Economic management
3,6,Social protection and risk management
4,5,Trade and integration
5,2,Public sector governance
6,11,Environment and natural resources management
8,7,Social dev/gender/inclusion
11,4,Financial and private sector development
18,10,Rural development
53,9,Urban development


Now we have a new DataFrame with unique code & name combination. Let's go ahead and merge the two dataframes to fill in all the blank project names.

In [244]:
filled_theme_df = pd.merge(theme_df, unique_theme_df, on='code', suffixes=('_x','')).drop('name_x' , axis=1)
filled_theme_df.head()

Unnamed: 0,code,name
0,8,Human development
1,8,Human development
2,8,Human development
3,8,Human development
4,8,Human development


In [245]:
filled_theme_df['name'].value_counts()

Environment and natural resources management    250
Rural development                               216
Human development                               210
Public sector governance                        199
Social protection and risk management           168
Financial and private sector development        146
Social dev/gender/inclusion                     130
Trade and integration                            77
Urban development                                50
Economic management                              38
Rule of law                                      15
Name: name, dtype: int64

These are the top themes after we handled the missing theme names

In [246]:
# Verifying that there are no more blank values in theme name
filled_theme_df[filled_theme_df.name == ''].count()

code    0
name    0
dtype: int64

With this, we have solution to all our 3 queries on the World Bank Projects dataset.