# JSON exercise

Download the data from [**here**](https://drive.google.com/file/d/1DGaX5AVfYhmWeb15lI-MzUbSKTYSz9fQ/view?usp=sharing) and answer following questions:
1. Find the 10 countries with most projects
2. What are top 10 sectors with projects?
3. Find the top 10 major project themes (using column 'mjtheme_namecode')
4. In 3. above you will notice that some entries have only the code and the name is missing. Create a dataframe with the missing names filled in.

In [1]:
import json
import pandas as pd 
from pandas.io.json import json_normalize
from pandas.io.json import to_json

In [2]:
# Read Json data into a dataframe 
df = pd.read_json("/Users/Oswal/Documents/GitHub/Other_data_types_exercise/world_bank_projects.json")

In [3]:
df.shape

(500, 50)

In [4]:
df.head()

Unnamed: 0,sector,supplementprojectflg,projectfinancialtype,prodline,mjtheme,idacommamt,impagency,project_name,mjthemecode,closingdate,...,majorsector_percent,board_approval_month,theme_namecode,countryname,url,source,projectstatusdisplay,ibrdcommamt,sector_namecode,_id
0,"[{'Name': 'Primary education'}, {'Name': 'Seco...",N,IDA,PE,[Human development],130000000,MINISTRY OF EDUCATION,Ethiopia General Education Quality Improvement...,811,2018-07-07T00:00:00Z,...,"[{'Percent': 46, 'Name': 'Education'}, {'Perce...",November,"[{'code': '65', 'name': 'Education for all'}]",Federal Democratic Republic of Ethiopia,http://www.worldbank.org/projects/P129828/ethi...,IBRD,Active,0,"[{'code': 'EP', 'name': 'Primary education'}, ...",{'$oid': '52b213b38594d8a2be17c780'}
1,[{'Name': 'Public administration- Other social...,N,OTHER,RE,"[Economic management, Social protection and ri...",0,MINISTRY OF FINANCE,TN: DTF Social Protection Reforms Support,16,,...,"[{'Percent': 70, 'Name': 'Public Administratio...",November,"[{'code': '24', 'name': 'Other economic manage...",Republic of Tunisia,http://www.worldbank.org/projects/P144674?lang=en,IBRD,Active,0,"[{'code': 'BS', 'name': 'Public administration...",{'$oid': '52b213b38594d8a2be17c781'}
2,[{'Name': 'Rural and Inter-Urban Roads and Hig...,Y,IDA,PE,"[Trade and integration, Public sector governan...",6060000,MINISTRY OF TRANSPORT AND COMMUNICATIONS,Tuvalu Aviation Investment Project - Additiona...,52116,,...,"[{'Percent': 100, 'Name': 'Transportation'}]",November,"[{'code': '47', 'name': 'Regional integration'...",Tuvalu,http://www.worldbank.org/projects/P145310?lang=en,IBRD,Active,0,"[{'code': 'TI', 'name': 'Rural and Inter-Urban...",{'$oid': '52b213b38594d8a2be17c782'}
3,[{'Name': 'Other social services'}],N,OTHER,RE,"[Social dev/gender/inclusion, Social dev/gende...",0,LABOR INTENSIVE PUBLIC WORKS PROJECT PMU,Gov't and Civil Society Organization Partnership,77,,...,"[{'Percent': 100, 'Name': 'Health and other so...",October,"[{'code': '57', 'name': 'Participation and civ...",Republic of Yemen,http://www.worldbank.org/projects/P144665?lang=en,IBRD,Active,0,"[{'code': 'JB', 'name': 'Other social services'}]",{'$oid': '52b213b38594d8a2be17c783'}
4,[{'Name': 'General industry and trade sector'}...,N,IDA,PE,"[Trade and integration, Financial and private ...",13100000,MINISTRY OF TRADE AND INDUSTRY,Second Private Sector Competitiveness and Econ...,54,2019-04-30T00:00:00Z,...,"[{'Percent': 50, 'Name': 'Industry and trade'}...",October,"[{'code': '45', 'name': 'Export development an...",Kingdom of Lesotho,http://www.worldbank.org/projects/P144933/seco...,IBRD,Active,0,"[{'code': 'YZ', 'name': 'General industry and ...",{'$oid': '52b213b38594d8a2be17c784'}


### We can see that the dataset has 500 observations and 50 attributes. Let’s move ahead with our 
 
### ** Query # 1 : Finding 10 countries with the most number of projects.**

### We will analyze the features first and decide which features could be useful for our solution.
### We also need to see if there is any missing value in chosen attributes and how we can take care of that.

In [None]:
## check the dataframe details
df.info()

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 500 entries, 0 to 499
Data columns (total 50 columns):
 #   Column                    Non-Null Count  Dtype 
---  ------                    --------------  ----- 
 0   sector                    500 non-null    object
 1   supplementprojectflg      498 non-null    object
 2   projectfinancialtype      500 non-null    object
 3   prodline                  500 non-null    object
 4   mjtheme                   491 non-null    object
 5   idacommamt                500 non-null    int64 
 6   impagency                 472 non-null    object
 7   project_name              500 non-null    object
 8   mjthemecode               500 non-null    object
 9   closingdate               370 non-null    object
 10  totalcommamt              500 non-null    int64 
 11  id                        500 non-null    object
 12  mjsector_namecode         500 non-null    object
 13  docty                     446 non-null    object
 14  sector1                   

In [10]:
## Find the total number of countries and projects 

df[['countryname', 'project_name']].count()

countryname     500
project_name    500
dtype: int64

In [11]:
df[['countryname', 'project_name']].nunique()

countryname     118
project_name    500
dtype: int64

### It looks like columns countryname and project_name can be used here. These columns have no missing values either. 
### No data imputing is needed and let’s straight away go to finding the solution.

### One interesting part to note here is that each observation represents a project. 
### There are 500 projects which involves 118 countries. We need to find out top 10 countries with most projects. 
### We can do this by grouping the data by 'countryname' and aggregate it on count of rows. Sorting it in descending order would give us the top countries.

### Another way to get the same result would be to use value_counts() on 'countryname'.

### Lets see both the approaches now.

In [12]:
# using groupby

top_countries = df.groupby('countryname')['project_name'].count().sort_values(ascending = False)


In [13]:
top_countries.head(10)

countryname
People's Republic of China         19
Republic of Indonesia              19
Socialist Republic of Vietnam      17
Republic of India                  16
Republic of Yemen                  13
People's Republic of Bangladesh    12
Nepal                              12
Kingdom of Morocco                 12
Republic of Mozambique             11
Africa                             11
Name: project_name, dtype: int64

In [14]:
# Using value counts 
df['countryname'].value_counts().sort_values(ascending = False).head(10)

Republic of Indonesia              19
People's Republic of China         19
Socialist Republic of Vietnam      17
Republic of India                  16
Republic of Yemen                  13
People's Republic of Bangladesh    12
Nepal                              12
Kingdom of Morocco                 12
Africa                             11
Republic of Mozambique             11
Name: countryname, dtype: int64

# Query # 2 : What are the top ten (10) sectors with projects?

In [17]:
# check dataframe details 
df.info()

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 500 entries, 0 to 499
Data columns (total 50 columns):
 #   Column                    Non-Null Count  Dtype 
---  ------                    --------------  ----- 
 0   sector                    500 non-null    object
 1   supplementprojectflg      498 non-null    object
 2   projectfinancialtype      500 non-null    object
 3   prodline                  500 non-null    object
 4   mjtheme                   491 non-null    object
 5   idacommamt                500 non-null    int64 
 6   impagency                 472 non-null    object
 7   project_name              500 non-null    object
 8   mjthemecode               500 non-null    object
 9   closingdate               370 non-null    object
 10  totalcommamt              500 non-null    int64 
 11  id                        500 non-null    object
 12  mjsector_namecode         500 non-null    object
 13  docty                     446 non-null    object
 14  sector1                   

In [20]:
# Using value counts 
df['mjsector_namecode'].value_counts().sort_values(ascending = False).head(10)

[{'code': 'BX', 'name': 'Public Administration, Law, and Justice'}]                                                                       33
[{'code': 'LX', 'name': 'Energy and mining'}]                                                                                             23
[{'code': 'JX', 'name': 'Health and other social services'}]                                                                              18
[{'code': 'JX', 'name': 'Health and other social services'}, {'code': 'BX', 'name': 'Public Administration, Law, and Justice'}]           18
[{'code': 'EX', 'name': 'Education'}]                                                                                                     13
[{'code': 'TX', 'name': 'Transportation'}, {'code': 'BX', 'name': 'Public Administration, Law, and Justice'}]                             11
[{'code': 'TX', 'name': 'Transportation'}]                                                                                                10
[{'code': 'BX

In [21]:
# Using value counts 
df['majorsector_percent'].value_counts().sort_values(ascending = False).head(10)

[{'Percent': 100, 'Name': 'Public Administration, Law, and Justice'}]                                                                33
[{'Percent': 100, 'Name': 'Energy and mining'}]                                                                                      23
[{'Percent': 100, 'Name': 'Health and other social services'}]                                                                       18
[{'Percent': 100, 'Name': 'Education'}]                                                                                              13
[{'Percent': 100, 'Name': 'Transportation'}]                                                                                         10
[{'Percent': 100, 'Name': 'Agriculture, fishing, and forestry'}]                                                                      7
[{'Percent': 100, 'Name': 'Finance'}]                                                                                                 5
[{'Percent': 100, 'Name': 'Water, sanitation and

# Query # 3 : Find the top 10 major project themes (using column 'mjtheme_namecode')




### As we can see that column mjtheme_namecode is a list of json objects itself, we would need to flatten it into a dataframe in order to work on it.

In [24]:
# function to flatten the theme json data
def flatten_theme(): 
    themeJsonList = df['mjtheme_namecode'].tolist()
    dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
    return pd.concat(dfs, ignore_index = True)


In [25]:
# flatten the theme and create a dataframe 
theme_df = flatten_theme()
theme_df.head()

  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonList]
  dfs = [json_normalize(themeJson) for themeJson in themeJsonL

Unnamed: 0,code,name
0,8,Human development
1,11,
2,1,Economic management
3,6,Social protection and risk management
4,5,Trade and integration


In [26]:
# check for the NaN values ]
theme_df.info()

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 1499 entries, 0 to 1498
Data columns (total 2 columns):
 #   Column  Non-Null Count  Dtype 
---  ------  --------------  ----- 
 0   code    1499 non-null   object
 1   name    1499 non-null   object
dtypes: object(2)
memory usage: 23.5+ KB


In [29]:
# check for the blank code and name 
print('blank code:')
print(theme_df[theme_df.code == ''].count())
print()
print('blank name:')
print(theme_df[theme_df.name == ''].count())

blank code:
code    0
name    0
dtype: int64

blank name:
code    122
name    122
dtype: int64


### We can see there are 122 missing entries for name. 

### For now, we will ignore the blank names while generating our report. In query # 3, we will see how to handle these missing names.

In [31]:
# Get the top 10 theme names
topThemes = theme_df[theme_df.name != ''].name.value_counts().sort_values(ascending = False)
topThemes.head(10)

Environment and natural resources management    223
Rural development                               202
Human development                               197
Public sector governance                        184
Social protection and risk management           158
Financial and private sector development        130
Social dev/gender/inclusion                     119
Trade and integration                            72
Urban development                                47
Economic management                              33
Name: name, dtype: int64

# Query # 3 : In query 2 above, there are some theme observations where the name is missing. Create a dataframe with the missing names filled in.

###  Find out relation between code and name. If each code has a unique non-blank project name, we can easily fill the missing project names based on the code.

In [34]:
# find count of unique code 
theme_df['code'].nunique()

11

#### We have 11 codes for the project themes. Let's find out the corresponding project names 

In [35]:
# Find the unique combination of code and name 
unique_theme_df = theme_df.drop_duplicates()
unique_theme_df = unique_theme_df[unique_theme_df.name != '']

unique_theme_df = unique_theme_df.reset_index().drop('index', axis = 1)
unique_theme_df

Unnamed: 0,code,name
0,8,Human development
1,1,Economic management
2,6,Social protection and risk management
3,5,Trade and integration
4,2,Public sector governance
5,11,Environment and natural resources management
6,7,Social dev/gender/inclusion
7,4,Financial and private sector development
8,10,Rural development
9,9,Urban development


#### Now we have a new dataframe with unique code & name combination. Let's go ahead and merge the two dataframes to fill in all the blank project names.

In [38]:
filled_theme_df = pd.merge(theme_df, unique_theme_df, on = 'code', suffixes = ('_x','')).drop('name_x', axis = 1)
filled_theme_df.head()

Unnamed: 0,code,name
0,8,Human development
1,8,Human development
2,8,Human development
3,8,Human development
4,8,Human development


In [39]:
filled_theme_df['name'].value_counts()

Environment and natural resources management    250
Rural development                               216
Human development                               210
Public sector governance                        199
Social protection and risk management           168
Financial and private sector development        146
Social dev/gender/inclusion                     130
Trade and integration                            77
Urban development                                50
Economic management                              38
Rule of law                                      15
Name: name, dtype: int64

#### These are the top ten themes after we handled the missing theme names. 

In [41]:
# Verifying that there are no more blank values in theme name
filled_theme_df[filled_theme_df.name == ''].count()

code    0
name    0
dtype: int64