# JSON examples and exercise
****
+ get familiar with packages for dealing with JSON
+ study examples with JSON strings and files 
+ work on exercise to be completed and submitted 
****
+ reference: http://pandas.pydata.org/pandas-docs/stable/io.html#io-json-reader
+ data source: http://jsonstudio.com/resources/
****

## JSON exercise

Using data in file 'data/world_bank_projects.json' and the techniques demonstrated above,
1. Find the 10 countries with most projects
2. Find the top 10 major project themes (using column 'mjtheme_namecode')
3. In 2. above you will notice that some entries have only the code and the name is missing. Create a dataframe with the missing names filled in.

In [1]:
import numpy as np
import pandas as pd
import json
from pandas.io.json import json_normalize

In [2]:
json_data = pd.read_json('data/world_bank_projects.json')
json_data.head()

Unnamed: 0,_id,approvalfy,board_approval_month,boardapprovaldate,borrower,closingdate,country_namecode,countrycode,countryname,countryshortname,...,sectorcode,source,status,supplementprojectflg,theme1,theme_namecode,themecode,totalamt,totalcommamt,url
0,{'$oid': '52b213b38594d8a2be17c780'},1999,November,2013-11-12T00:00:00Z,FEDERAL DEMOCRATIC REPUBLIC OF ETHIOPIA,2018-07-07T00:00:00Z,Federal Democratic Republic of Ethiopia!$!ET,ET,Federal Democratic Republic of Ethiopia,Ethiopia,...,"ET,BS,ES,EP",IBRD,Active,N,"{'Percent': 100, 'Name': 'Education for all'}","[{'code': '65', 'name': 'Education for all'}]",65,130000000,130000000,http://www.worldbank.org/projects/P129828/ethi...
1,{'$oid': '52b213b38594d8a2be17c781'},2015,November,2013-11-04T00:00:00Z,GOVERNMENT OF TUNISIA,,Republic of Tunisia!$!TN,TN,Republic of Tunisia,Tunisia,...,"BZ,BS",IBRD,Active,N,"{'Percent': 30, 'Name': 'Other economic manage...","[{'code': '24', 'name': 'Other economic manage...",5424,0,4700000,http://www.worldbank.org/projects/P144674?lang=en
2,{'$oid': '52b213b38594d8a2be17c782'},2014,November,2013-11-01T00:00:00Z,MINISTRY OF FINANCE AND ECONOMIC DEVEL,,Tuvalu!$!TV,TV,Tuvalu,Tuvalu,...,TI,IBRD,Active,Y,"{'Percent': 46, 'Name': 'Regional integration'}","[{'code': '47', 'name': 'Regional integration'...",52812547,6060000,6060000,http://www.worldbank.org/projects/P145310?lang=en
3,{'$oid': '52b213b38594d8a2be17c783'},2014,October,2013-10-31T00:00:00Z,MIN. OF PLANNING AND INT'L COOPERATION,,Republic of Yemen!$!RY,RY,Republic of Yemen,"Yemen, Republic of",...,JB,IBRD,Active,N,"{'Percent': 50, 'Name': 'Participation and civ...","[{'code': '57', 'name': 'Participation and civ...",5957,0,1500000,http://www.worldbank.org/projects/P144665?lang=en
4,{'$oid': '52b213b38594d8a2be17c784'},2014,October,2013-10-31T00:00:00Z,MINISTRY OF FINANCE,2019-04-30T00:00:00Z,Kingdom of Lesotho!$!LS,LS,Kingdom of Lesotho,Lesotho,...,"FH,YW,YZ",IBRD,Active,N,"{'Percent': 30, 'Name': 'Export development an...","[{'code': '45', 'name': 'Export development an...",4145,13100000,13100000,http://www.worldbank.org/projects/P144933/seco...


# Part 1: Top 10 Countries with the Most Projects

In [3]:
country = json_data.groupby('countryname')['id'].agg('count').sort_values(ascending=False)
print(country.head(10))

countryname
People's Republic of China         19
Republic of Indonesia              19
Socialist Republic of Vietnam      17
Republic of India                  16
Republic of Yemen                  13
Nepal                              12
People's Republic of Bangladesh    12
Kingdom of Morocco                 12
Africa                             11
Republic of Mozambique             11
Name: id, dtype: int64


The 9th country on this list is Africa, which is a continent and not a country. I want to take a look at the borrower column for the countries that have a country of Africa to see if there are values that can be updated.

In [4]:
africa = json_data[json_data['countryname'] == 'Africa']
print(africa[['countrycode', 'borrower', 'countryshortname']])

    countrycode                       borrower countryshortname
45           3A                         ECOWAS           Africa
46           3A                 UGANDA-COMOROS           Africa
51           3A               OSS, IUCN, CILSS           Africa
58           3A                  BANK EXECUTED           Africa
65           3A        BURUNDI,RWANDA,TANZANIA           Africa
99           3A             GOVERNMENT OF MALI           Africa
167          3A  MINISTRIES OF TELECOM AND ICT           Africa
184          3A                       TANZANIA           Africa
287          3A     MALAWI, MOZAMBIQUE, ZAMBIA           Africa
353          3A          NILE BASIN INITIATIVE           Africa
449          3A            GOVERNMENT OF NIGER           Africa


The table above shows the borrower column for the rows that had Africa as the countryname. It looks like there are 3 rows in this list that apply to countries (Government of Mali, Tanzania, and Government of Niger).

In [6]:
json_data.set_value(449, 'countrycode', 'NE')
json_data.set_value(449, 'countryname', 'Republic of Niger')
json_data.set_value(449, 'countryshortname', 'Republic of Niger');
json_data.iloc[449]

_id                                      {'$oid': '52b213b38594d8a2be17c941'}
approvalfy                                                               2013
board_approval_month                                                  October
boardapprovaldate                                        2012-10-02T00:00:00Z
borrower                                                  GOVERNMENT OF NIGER
closingdate                                              2020-04-01T00:00:00Z
country_namecode                                                  Africa!$!3A
countrycode                                                                NE
countryname                                                 Republic of Niger
countryshortname                                            Republic of Niger
docty                       Implementation Status and Results Report,Summa...
envassesmentcategorycode                                                    A
grantamt                                                        

The code, name, and country short name were replaced with the values for Niger in the code above.

In [8]:
json_data.set_value(184, 'countryname', 'United Republic of Tanzania')
json_data.set_value(184, 'countryshortname', 'United Republic of Tanzania')
json_data.set_value(184, 'countrycode', 'TZ');
print(json_data.iloc[184])

_id                                      {'$oid': '52b213b38594d8a2be17c838'}
approvalfy                                                               2013
board_approval_month                                                      May
boardapprovaldate                                        2013-05-21T00:00:00Z
borrower                                                             TANZANIA
closingdate                                              2018-12-31T00:00:00Z
country_namecode                                                  Africa!$!3A
countrycode                                                                TZ
countryname                                       United Republic of Tanzania
countryshortname                                  United Republic of Tanzania
docty                       Implementation Status and Results Report,Proje...
envassesmentcategorycode                                                    B
grantamt                                                        

The name, code, and short name were replaced with values for Tanzania. The same is done below for Mali.

In [9]:
json_data.set_value(99, 'countryname', 'Republic of Mali')
json_data.set_value(99, 'countryshortname', 'Republic of Mali')
json_data.set_value(99, 'countrycode', 'ML');

The remaining projects for the country of 'Africa' have multiple countries associated to them so those values cannot be updated.

In [10]:
country_updated = json_data.groupby('countryname')['id'].agg('count').sort_values(ascending=False)
print(country_updated.head(10))

countryname
People's Republic of China         19
Republic of Indonesia              19
Socialist Republic of Vietnam      17
Republic of India                  16
Republic of Yemen                  13
Nepal                              12
People's Republic of Bangladesh    12
Kingdom of Morocco                 12
Republic of Mozambique             11
United Republic of Tanzania         9
Name: id, dtype: int64


The table above shows the top 10 countries with the most projects.

# Part 2: Top 10 Major Project Themes

In [12]:
json_string = json.load((open('data/world_bank_projects.json')))
json_normal = json_normalize(json_string, 'mjtheme_namecode')
print(json_normal.shape)
print(json_normal.head(20))

(1499, 2)
   code                                          name
0     8                             Human development
1    11                                              
2     1                           Economic management
3     6         Social protection and risk management
4     5                         Trade and integration
5     2                      Public sector governance
6    11  Environment and natural resources management
7     6         Social protection and risk management
8     7                   Social dev/gender/inclusion
9     7                   Social dev/gender/inclusion
10    5                         Trade and integration
11    4      Financial and private sector development
12    6         Social protection and risk management
13    6                                              
14    2                      Public sector governance
15    4      Financial and private sector development
16   11  Environment and natural resources management
17    8           

As you can see, there are missing values for the name field. We'll have to remove these in order to see the top ten project themes.

In [13]:
json_no_nulls = json_normal[json_normal['name'] != '']
print(json_no_nulls.shape)
json_no_nulls.groupby('name')['code'].agg('count').sort_values(ascending=False).head(10)

(1377, 2)


name
Environment and natural resources management    223
Rural development                               202
Human development                               197
Public sector governance                        184
Social protection and risk management           158
Financial and private sector development        130
Social dev/gender/inclusion                     119
Trade and integration                            72
Urban development                                47
Economic management                              33
Name: code, dtype: int64

After removing the blanks for the name field, we can see the top 10 project themes.

# Part 3: Filling in the Missing Values

There are two ways I saw that this could be done.

1) Replace the name value of a row using the apply() method on a dataframe and 
2) Sort the dataframe and backfill the data

## Replace using apply()

In [14]:
def fill_blanks(row):
    value = ''
    if row['code'] == '1':
        value = 'Economic management'
    if row['code'] == '2':
        value = 'Public sector governance'
    if row['code'] == '3':
        value = 'Rule of law'
    if row['code'] == '4':
        value = 'Financial and private sector development'
    if row['code'] == '5':
        value = 'Trade and integration'
    if row['code'] == '6':
        value = 'Social protection and risk management'
    if row['code'] == '7':
        value = 'Social dev/gender/inclusion'
    if row['code'] == '8':
        value = 'Human development'
    if row['code'] == '9':
        value = 'Urban development'
    if row['code'] == '10':
        value = 'Rural development'
    if row['code'] == '11':
        value = 'Environment and natural resources management'
    row['name'] = value
    return row

The function above takes a row and replaces the name value based on the value in the code column.

In [17]:
json_new = json_normal.copy()
json_new = json_new.apply(fill_blanks, axis=1)
print(json_new.head(20))
print(json_new.groupby('name')['code'].agg('count').sort_values(ascending=False).head(10))

   code                                          name
0     8                             Human development
1    11  Environment and natural resources management
2     1                           Economic management
3     6         Social protection and risk management
4     5                         Trade and integration
5     2                      Public sector governance
6    11  Environment and natural resources management
7     6         Social protection and risk management
8     7                   Social dev/gender/inclusion
9     7                   Social dev/gender/inclusion
10    5                         Trade and integration
11    4      Financial and private sector development
12    6         Social protection and risk management
13    6         Social protection and risk management
14    2                      Public sector governance
15    4      Financial and private sector development
16   11  Environment and natural resources management
17    8                     

The function is then applied to the dataframe of codes and names using the apply() method. This apply() method is performed over the rows of the dataframe.

The table at the bottom shows the top 10 project themes after filling in the data.

## Sort and Backfill

In [18]:
json_sorted = json_normal.sort_values(['code', 'name'], ascending=True)
json_sorted.head(10)

Unnamed: 0,code,name
212,1,
363,1,
1024,1,
1114,1,
1437,1,
2,1,Economic management
88,1,Economic management
175,1,Economic management
204,1,Economic management
205,1,Economic management


The values of the blank names are ''. In order to use backfilling, we'll have to convert these values to NaN using np.nan.

In [19]:
json_sorted['name'][json_sorted['name'] == ''] = np.nan
json_sorted.head(10)

Unnamed: 0,code,name
212,1,
363,1,
1024,1,
1114,1,
1437,1,
2,1,Economic management
88,1,Economic management
175,1,Economic management
204,1,Economic management
205,1,Economic management


In [20]:
json_filled = json_sorted.fillna(method='bfill')
json_filled.sort_index().head(20)

Unnamed: 0,code,name
0,8,Human development
1,11,Environment and natural resources management
2,1,Economic management
3,6,Social protection and risk management
4,5,Trade and integration
5,2,Public sector governance
6,11,Environment and natural resources management
7,6,Social protection and risk management
8,7,Social dev/gender/inclusion
9,7,Social dev/gender/inclusion


After backfilling the data and re-sorting the index, the values are correctly filled in for the codes that had missing values.

In [21]:
json_filled.groupby('name')['code'].agg('count').sort_values(ascending=False).head(10)

name
Environment and natural resources management    250
Rural development                               216
Human development                               210
Public sector governance                        199
Social protection and risk management           168
Financial and private sector development        146
Social dev/gender/inclusion                     130
Trade and integration                            77
Urban development                                50
Economic management                              38
Name: code, dtype: int64

The table above shows the top 10 project themes with filled in values for the blank rows. It turns out to the same as the table from the apply() method.