# Week 03 Assignment Covid

New types of data and new data science technologies enable new research. These new technologies are technologies such as the ability to combine existing data or the ability to generate synthetic data from existing knowledge. This week casus is based on such research. Data is generated by Synthea's COVID-19 module. The data was constructed using three peer-reviewed publications published in the early stages of the global pandemic, when less was known, along with emerging resources, data, publications, and clinical knowledge. The simulation outputs synthetic Electronic Health Records (EHR), including the daily consumption of Personal Protective Equipment (PPE) and other medical devices and supplies. The Data is stored in separate tables to avoid redundancy, with as a concequence that tables need to be combined and reorganized in dataframes for analysing purpose.

Keywords: merge data, subset data, clean data, generate data

You will learn about combining data with pandas and numpy and you will learn to visualize with bokeh. Concretely, you will preprocess the partly Synthetic Covid data in an appropiate format in order to conduct statistical and visual analysis. Learning objectives

- Combine multiple data sources for analysis
- Read, inspect, clean, reshape data
- Visualize data using bokeh
- Maintain development environment 
- Apply coding standards and FAIR principles
- Reshape the dataset into a format suitable for visual and statistical analysis
- Use widgets to make the plot interactive 
- Use GIS libraries to plot geographical data

Tutorials about combining data: https://github.com/fenna/BFVM23PROG2/blob/main/tutorials/tutorial_combine_data.ipynb

study case combining data:https://github.com/fenna/BFVM23PROG2/blob/main/study_cases/adults_who_binge_drank_in_hot_towns.ipynb


Please add the topics you want to learn about here: <padlet>


Your job is to **visualize the lab values taken for COVID-19 patients of survived versus not survived patients**. 

The assignment consists of 6 parts:

- [part 1: load the data](#0)
     - [Exercise 1.1](#ex-11)
- [part 2: data wrangling](#1)
     - [Exercise 2.1](#ex-21)
- [part 3: more wrangling](#2)
     - [Exercise 3.1](#ex-31)
- [part 4: plot the data](#3)
     - [Exercise 4.1](#ex-41)
- [part 5: plot patient location](#5)
     - [Exercise 5.1](#ex-51)


Part 1 and 4 are mandatory, part 5 is optional (bonus)
Mind you that you cannot copy code without referencing the code. If you copy code you need to be able to explain your code verbally and you will not get the full score. 


## About the data

The data is generated by Synthea's COVID-19 module. The data was constructed using three peer-reviewed publications published in the early stages of the global pandemic, when less was known, along with emerging resources, data, publications, and clinical knowledge. The simulation outputs synthetic Electronic Health Records (EHR), including the daily consumption of Personal Protective Equipment (PPE) and other medical devices and supplies. For this assignment the `conditions`, `patients`, `observations`, `careplans` and `encounters` table will be used. The Data is stored in separate tables to avoid redundancy, with as a concequence that tables need to be combined and reorganized in dataframes for analysing purpose.

Source: Walonoski J, Klaus S, Granger E, Hall D, Gregorowicz A, Neyarapally G, Watson A, Eastman J. Synthea™ Novel coronavirus (COVID-19) model and synthetic data set. Intelligence-Based Medicine. 2020 Nov;1:100007. https://doi.org/10.1016/j.ibmed.2020.100007

Please <a href = "https://synthetichealth.github.io/synthea-sample-data/downloads/10k_synthea_covid19_csv.zip">download</a> the data

#### Covid Patients
Patients are considered Covid patients if they are identified with `CODE` `840539006`


#### Survivors
Patients that had covid and where tested negative after isolation have tested code `94531-1`,  SARS-CoV-2 RNA Pnl Resp NAA+probe (covid-sars test) + a value of `Not detected (qualifier value)`. These patients are considered to be survived covid patients. 

#### Non-Survivors
Patients that did not survived Covid have a `DEATHDATE` which is not null. 


#### Lab values  COVID-19 patients

Patients are monitored for blood and heart conditions once they are admitted in Hospital or under treatment. The lab values of interest are as follow: 

- `48065-7`  Fibrin D-dimer FEU [Mass/volume] in Platelet poor plasma
- `26881-3`   Interleukin 6 [Mass/volume] in Serum or Plasma
- `2276-4` Ferritin [Mass/volume] in Serum or Plasma
- `89579-7` Troponin I.cardiac [Mass/volume] in Serum or Plasma by High sensitivity method
- `731-0` Lymphocytes [#/volume] in Blood by Automated count
- `14804-9` Lactate dehydrogenase [Enzymatic activity/volume] in Serum or Plasma by Lactate to pyruvate reaction


---

<a name='0'></a>
## Part 1: Load the data (20 pt)

Instructions: Load the data of the following files. 
Preferably we read the data not with a hard coded data path but using a config file. See https://fennaf.gitbook.io/bfvm23prog2/data-processing/configuration-files/yaml

- conditions.csv
- patients.csv
- observations.csv
- careplans.csv
- encounters.csv

Get yourself familiar with the data. Create some meaningful overviews. Answer the following questions

1. How many patients are there
2. How many covid-patients are there
3. How many patients do have a 'Hospital admission for isolation' encounter
    
<details>    
<summary>
    <font size="3" color="darkgreen"><b>Hints</b></font>
</summary>
    <ul><li>use a unique dataframe for each file, use a meaningful name</li>
    <li>pandas.read_csv() method can be used to read a csv file</li>
    <li>pandas.DataFrame.head() method is often used to inspect the dataframe</li>
    <li>.unique() returns a list of unique values of a column</li>
</ul>
</details>

<a name='ex-11'></a>
### 1.1 Code your solution

In [2]:
import pandas as pd
import yaml

def get_config():
    with open("config.yaml", 'r') as stream:
        config = yaml.safe_load(stream)
    return config

config = get_config()

conditions_df = pd.read_csv(config['conditions'])
patients_df = pd.read_csv(config['patients'])
observations_df = pd.read_csv(config['observations'])
careplans_df = pd.read_csv(config['careplans'])
encounters_df = pd.read_csv(config['encounters'])

In [3]:
num_pat = patients_df['Id'].nunique()
num_cov = conditions_df.groupby('DESCRIPTION').size()['COVID-19']
num_admitted = encounters_df.groupby('DESCRIPTION').size()['Hospital admission for isolation (procedure)']
num_died = patients_df['DEATHDATE'].count()

### 1.2 Test your solution
The following function needs to be called. You can use this as a test. There are however more meaningful overviews 
you can create. 

In [4]:
def part1(num_pat, num_cov, num_admitted, num_died):
    print(f'There are {num_pat} patients in total')
    print(f'There are {num_cov} covid patients')
    print(f'There are {num_admitted} admitted patients')
    print(f'{num_died} patients died')
part1(num_pat, num_cov, num_admitted, num_died)


There are 12352 patients in total
There are 8820 covid patients
There are 1867 admitted patients
2352 patients died


### Expected outcome

---

<a name='1'></a>
## Part 2: Data Wrangling: set up the dataframe (30 pt)

In this part we are going to combine data to create a dataframe with values of interest for the lab values analysis. 

We would like a dataframe containing the following information per record (only Covid patients!!!)

- `PATIENT` - the ID of the covid patient
- `days` - the number of days the patient is under observation
- `CODE-Y` - the code of the observation  
- `VALUE` - the lab value of the observation

where only the following observation codes needs to be selected:

- `48065-7`  Fibrin D-dimer FEU [Mass/volume] in Platelet poor plasma
- `26881-3`   Interleukin 6 [Mass/volume] in Serum or Plasma
- `2276-4` Ferritin [Mass/volume] in Serum or Plasma
- `89579-7` Troponin I.cardiac [Mass/volume] in Serum or Plasma by High sensitivity method
- `731-0` Lymphocytes [#/volume] in Blood by Automated count
- `14804-9` Lactate dehydrogenase [Enzymatic activity/volume] in Serum or Plasma by Lactate to pyruvate reaction

The days information is not primarely available and needs to be calculated by substracting observation DATE - START. 

An example of such a dataframe is given below:

<details>    
<summary>
    <font size="3" color="darkgreen"><b>Hints</b></font>
</summary>
    <ul><li>you can use pandas.DataFrame.merge() to merge dataframes</li>
    <li>df = df[(df.CODE == condition1 | df.CODE == condition1 )] selects rows with CODE of 2 conditional values</li>
    <li>df.DATE - df.START return days if DATE and START are datetime format</li>
    <li>pd.to_datetime() can be used to typecast to datetime</li>
</ul>
</details>

<a name='ex-21'></a>
### 2.1 Code your solution

In [5]:
#Possible approach:
codes = ['48065-7', '26881-3', '2276-4', '89579-7', '731-0', '14804-9']
#Select all the patients with covid from the conditions table
covid_df = conditions_df[conditions_df['DESCRIPTION'] == 'COVID-19']
#Combine conditions table (only covid patients) with the patient table into a covid_patient table
patients_df = patients_df.rename(columns={'Id':'PATIENT'})
covid_patient = covid_df.merge(patients_df, on='PATIENT')
#select the only the relevant lab observations from the observations table into a lab_obs table
lab_obs = observations_df.drop(columns=[ 'ENCOUNTER','DESCRIPTION','TYPE'])
#merge the covid_patient table with the lab_obs table into a covid_patients_obs table
covid_patients_obs = covid_patient.merge(lab_obs, on='PATIENT')
covid_patients_obs = covid_patients_obs[['PATIENT','DATE','START','CODE_y','VALUE','UNITS']]
#clean the covid_patients_obs table (rename columns, select only relevant columns, sort, typecast, add days column)
covid_patients_obs['START'] = pd.to_datetime(covid_patients_obs['START'])
covid_patients_obs['DATE'] = pd.to_datetime(covid_patients_obs['DATE'])
covid_patients_obs['days'] = (covid_patients_obs['DATE'] - covid_patients_obs['START']).dt.days
covid_patients_obs = covid_patients_obs.drop(columns=['START','DATE'])
# select the rows with correct observation codes
covid_patients_obs = covid_patients_obs[covid_patients_obs['CODE_y'].isin(codes)]

# cast the values of the values column into floats for further analysis
covid_patients_obs['VALUE'] = covid_patients_obs['VALUE'].astype(float)

---

<a name='2'></a>
## Part 3: Data Wrangling, split into survived and not survived (10 pt)

Now we have the required data we would like to split the data into survived and not survived. First we fetch all the ids of the survived and deceased patients. We can use these ids to select the records of the survived patients and the patients that did not survived

Your job is to split the data into survived and not survived records. There are multiple ways to do this. One way is the  `.isin()` method

In [6]:
import numpy as np
#the following code is given, RUN THIS CELL
#get survived and deceased ids
completed_isolation_patients = careplans_df[(careplans_df.CODE == 736376001) & (careplans_df.STOP.notna()) \
                                          & (careplans_df.REASONCODE == 840539006)].PATIENT
negative_covid_patient_ids = observations_df[(observations_df.CODE == '94531-1') \
                                          & (observations_df.VALUE == 'Not detected (qualifier value)')].PATIENT.unique()
survivor_ids = np.union1d(completed_isolation_patients, negative_covid_patient_ids)
deceased_ids = patients_df[patients_df.DEATHDATE.notna()].PATIENT

<a name='ex-31'></a>
### 3.1 Code your solution

In [7]:
survived_patients = covid_patients_obs[covid_patients_obs['PATIENT'].isin(survivor_ids)]
deceased_patients = covid_patients_obs[covid_patients_obs['PATIENT'].isin(deceased_ids)]
survived = survived_patients['PATIENT'].count()
deceased = deceased_patients['PATIENT'].count()

### 3.2 Test your solution

In [8]:
def test3(survived, died):
    print(f'patients records survived: {survived}, patients records deceased {died}')
#call the test3

test3(survived, deceased)

patients records survived: 57303, patients records deceased 16793


#### Expected outcome

---

<a name='3'></a>
## Part 4: Plot the data (20 pt)

Create plots with the lab data, for each code one plot. Separate the survivors and the deceased by color. An example of such a plot is given below. You can create 6 plots in one grid (for each code one plot) or use a widget (for instance a drop down menu widget) to select a lab CODE. Plot on the x-axis the days, on the y-axis the VALUE. Use proper labels, titles and legends.

<img src="../images/week3_plot.png" width="500" height="500"/>

<a name='ex-41'></a>
### 4.1 Code your solution

In [9]:
from bokeh.plotting import figure, output_file
# define a function for making a scatterplot that takes in the arguments survived and deceased
def scatterplot(survived, deceased):
    '''This function is able to create new scatterplots that can distinguish two
    groups from a certain dataframe
    input: survived  and deceased dataframe with as the name of the columns, UNITS as name of 
    the column that contains the units, VALUE for the column that contain value of the data and 
    a list of codes that are used to select certain patients with certain observation codes
    output: plots that visualize the difference between two groups in the form of a scatterplot
    '''
    # define p as a figure with a certain title
    p = figure(title = f"Survived vs Deceased {code}")
    # define the figure label of the x-axis 
    p.xaxis.axis_label = 'days'
    # define the figure label of the y axis with an f string to make it more re-usable
    p.yaxis.axis_label = f"{survived['UNITS'].values[0]}"
    # create figures with the specified dataframes and columns
    p.circle(survived['days'].values, survived['VALUE'].values, fill_alpha=0.4, size=12, fill_color="magenta", legend_label = "Survived")
    p.circle(deceased['days'].values, deceased['VALUE'].values, fill_alpha=0.4, size=12, fill_color="lime", legend_label = "Deceased")
    # output the file as an .html files with an f string and title 
    output_file(f"COVID-19{codes}.html", title="Survived vs Deceased COVID-19 patients")
    return p

In [10]:
from bokeh.io import show
from bokeh.models import TabPanel, Tabs
# define an empty list to store the tabs(plots) for the TabPanel in
tabs = []
# loop through the list of codes and select only survived and deceased patients from the dataframe and store them in the tabs list
for code in codes:
    tabs.append(TabPanel(child = scatterplot(survived_patients[survived_patients['CODE_y'] == code], deceased_patients[deceased_patients['CODE_y'] == code]), title=f'{code}'))

# show the defined tabs
show(Tabs(tabs=tabs))

<a name='4'></a>
## Part 5: Plot the location of the patients (10 pt)

This is a bonus part. Can you plot the patients location on a map? See also 
https://docs.bokeh.org/en/latest/docs/user_guide/geo.html 

You can use either package folium or geopandas. You need the Latitude and Longitude information from the patient tabel


<a name='ex-51'></a>
### 5.1 Code your solution

In [32]:
# Importing Libraries
import pandas as pd
import folium

# merge the patient dataframe (because of the presence of LON and LAT) with the survival and deceased patients
# select only the necessary columns of the dataframes 
survived_geo = survived_patients[['PATIENT']].merge(patients_df[['PATIENT', 'LON', 'LAT']], on='PATIENT')
deceased_geo = deceased_patients[['PATIENT']].merge(patients_df[['PATIENT','LON', 'LAT']], on='PATIENT')
# put the lattitude and longitude in a list and zip them to loop over them further on
survived_latlon = list(zip(survived_geo['LAT'].unique(), survived_geo['LON'].unique()))
deceased_latlon = list(zip(deceased_geo['LAT'].unique(), deceased_geo['LON'].unique()))
# initialize a map with the median of the lattitude of the patients df and the longitutde of the
# patients df
map = folium.Map((patients_df.LAT.median(), patients_df.LON.median()), zoom_start=7, min_zoom=10, max_zoom=10)
# add the survived to the map and mark them red
survived = folium.FeatureGroup("survived").add_to(map)
# loop through the zipped values and add them to the map
for lat, lon in survived_latlon:
    folium.CircleMarker((lat, lon), color="green", fill=True, radius = 5).add_child(folium.Popup("Survived")).add_to(survived)

deceased = folium.FeatureGroup("deceased").add_to(map)
for lat, lon in deceased_latlon:
    folium.CircleMarker((lat, lon), color="blue", fill=True, radius = 5).add_child(folium.Popup("Deceased")).add_to(deceased)

# add a layer to select from deceased and survived patients
folium.LayerControl().add_to(map)
# save the map as an html file
map.save('Location_map.html')
