Generate text based on the provided prompt, which is related to identifying causality between medical conditions. The function replaces a placeholder in the prompt with the actual test case, and the generated text is printed and returned. Finally, the function is called once with a specific test case ("A: heart rate, B: sepsis").

In [6]:
import requests
import re

# For local streaming, the websockets are hosted without ssl - http://
HOST = 'localhost:5000'
URI = f'http://{HOST}/api/v1/generate'

# For reverse-proxied streaming, the remote will likely host with ssl - https://
# URI = 'https://your-uri-here.trycloudflare.com/api/v1/generate'


def run(prompt,test):
    request = {
        'prompt': prompt.replace("{test}",test),
        'max_new_tokens': 250,
        'temperature': 0.0,
        'top_p': 0.1,
        'length_penalty': 5,
        'early_stopping': True,
        'seed': 0,
    }

    response = requests.post(URI, json=request)
    print(response.json()['results'][0]['text'])

    return response.json()['results'][0]['text']


prompt = '''
[INST]
You are a medical expert. You are tasked to identify causality between two conditions. 
Return yes if A causes B. IF B causes A or if you are not sure, return no. 
Keep your responses short and succinct. 
Input
{test}
[/INST]
'''
run(prompt,"A: heart rate, B: sepsis")

Yes, A (heart rate) can cause B (sepsis). Elevated heart rate can lead to increased cardiac output, which can increase the delivery of bacteria-carrying blood to the site of infection, leading to sepsis.


'Yes, A (heart rate) can cause B (sepsis). Elevated heart rate can lead to increased cardiac output, which can increase the delivery of bacteria-carrying blood to the site of infection, leading to sepsis.'

In [60]:
prompt2 = '''
[INST]
You are a medical researcher with expertise in diagnosing diseases. 
Your task is to identify what biological markers, demographic data, or symptoms are most crucial for diagnosing a specific disease.

Input Disease: [Disease_Name]

Please list the top biological markers, demographic data, and symptoms that are essential for diagnosing [Disease_Name]. 
Keep your response short and to the point.
Input
{test}
[/INST]
'''

In [61]:
run(prompt2,"Input Disease: uti")

For diagnosing Urinary Tract Infections (UTIs), the following biological markers, demographic data, and symptoms are essential:

Biological Markers:

1. Leukocyte esterase test: This is a rapid dipstick test that detects the presence of white blood cells in urine, which is a common finding in UTIs.
2. Nitrite test: This test detects the presence of nitrite in urine, which is produced by certain bacteria that cause UTIs.
3. Urinalysis: This includes examining urine for red blood cells, white blood cells, and bacteria.

Demographic Data:

1. Age: UTIs are more common in women and children.
2. Sex: Women are more likely to develop UTIs than men.
3. History of previous UTIs: Patients with a history of recurrent UTIs may be more likely to have a current UTI.

Symptoms:

1. Dysuria (painful urination)
2. Frequent urination
3. Cloudy or strong


'For diagnosing Urinary Tract Infections (UTIs), the following biological markers, demographic data, and symptoms are essential:\n\nBiological Markers:\n\n1. Leukocyte esterase test: This is a rapid dipstick test that detects the presence of white blood cells in urine, which is a common finding in UTIs.\n2. Nitrite test: This test detects the presence of nitrite in urine, which is produced by certain bacteria that cause UTIs.\n3. Urinalysis: This includes examining urine for red blood cells, white blood cells, and bacteria.\n\nDemographic Data:\n\n1. Age: UTIs are more common in women and children.\n2. Sex: Women are more likely to develop UTIs than men.\n3. History of previous UTIs: Patients with a history of recurrent UTIs may be more likely to have a current UTI.\n\nSymptoms:\n\n1. Dysuria (painful urination)\n2. Frequent urination\n3. Cloudy or strong'

In [7]:
import pandas as pd
import numpy as np

Use the local to focus on cause and effect. Like what symptoms are causing sepsis.

In [8]:
symptom = ['uti', 'low oxygen', 'low neutrophils', 'low BP', 'high BP', 'age>50', 'pregnancy', 'age>70']
df = pd.DataFrame(symptom, columns=['symptom'])
df

Unnamed: 0,symptom
0,uti
1,low oxygen
2,low neutrophils
3,low BP
4,high BP
5,age>50


In [11]:
# Apply 'run' function to each row of DataFrame
df['relation'] = df.apply(lambda x: run(prompt, f"A: {x['symptom']}, B: sepsis"), axis=1)

Yes, A (uti) causes B (sepsis).
Yes, A (low oxygen) causes B (sepsis). Low oxygen levels can lead to sepsis, a life-threatening condition that can cause organ failure and death.
Yes, A (low neutrophils) causes B (sepsis). Low neutrophil counts can impair the body's ability to fight off infections, which can lead to the development of sepsis.
No. Sepsis can cause low blood pressure (BP), but low BP alone does not cause sepsis.
Yes, A (high BP) can cause B (sepsis). High blood pressure can lead to damage to blood vessels, which can increase the risk of bacterial infection and sepsis.
Yes, A (age > 50) causes B (sepsis). As people age, their immune systems weaken, making them more susceptible to infections like sepsis.


In [15]:
df['symptom_short'] = df['relation'].apply(lambda x: 'yes' if 'yes' in x.lower() else 'no')

In [59]:
df

Unnamed: 0,symptom,relation,symptom_short
0,uti,"Yes, A (uti) causes B (sepsis).",yes
1,low oxygen,"Yes, A (low oxygen) causes B (sepsis). Low oxy...",yes
2,low neutrophils,"Yes, A (low neutrophils) causes B (sepsis). Lo...",yes
3,low BP,"No. Sepsis can cause low blood pressure (BP), ...",no
4,high BP,"Yes, A (high BP) can cause B (sepsis). High bl...",yes
5,age>50,"Yes, A (age > 50) causes B (sepsis). As people...",yes


# mean time calculations

Calculating the mean time (of each test and final test)

In [None]:
lab_path = 'CITS5553/Project/data-science-capstone-project/data/LABEVENTS.csv'
df_lab = pd.read_csv(lab_path)

In [41]:
# calculate mean time - of each test and final test
from datetime import datetime

In [52]:
def computeTdiff(x):
    x = x['CHARTTIME']
    x = x.dropna()
    x = sorted(x)
    
    if len(x)>1:
        #return datetime.strptime(x[-1]['CHARTTIME'], '%Y-%m-%d %H:%M%:%S').timestamp() - datetime.strptime(x[0]['CHARTTIME'], '%Y-%m-%d %H:%M%:%S').timestamp()
        return datetime.strptime(x[-1], '%Y-%m-%d %H:%M:%S').timestamp() - datetime.strptime(x[0], '%Y-%m-%d %H:%M:%S').timestamp()
    else:
        return float("nan")

In [37]:
datetime.strptime('2101-10-12 16:06:00', '%Y-%m-%d %H:%M:%S').timestamp()


4158547560.0

In [58]:
#np.nanmean(df_lab.iloc[:10000].groupby('HADM_ID').apply(lambda x:computeTdiff(x)).values())/(24*3600)
df_lab.iloc[:10000].groupby('HADM_ID').apply(lambda x:computeTdiff(x.dropna()))/(24*3600)

HADM_ID
103251.0     0.000000
107064.0    15.914583
109235.0     2.923611
109451.0    12.885417
112213.0    12.672917
118037.0     0.000000
143045.0     6.463194
145834.0    10.510417
150750.0     4.939583
157681.0     4.754861
159514.0     3.875000
161087.0     3.961806
163353.0     0.015278
178980.0     0.000000
184167.0     6.697917
185777.0     6.616667
188822.0     1.839583
194023.0     3.597222
194540.0    25.347222
dtype: float64

In [56]:
# checking the HADM_ID shows that only one 
df_lab[df_lab.HADM_ID == 178980]

Unnamed: 0,ROW_ID,SUBJECT_ID,HADM_ID,ITEMID,CHARTTIME,VALUE,VALUENUM,VALUEUOM,FLAG
2853,3421,5,178980.0,51143,2103-02-02 06:10:00,0.0,0.0,%,
2854,3422,5,178980.0,51144,2103-02-02 06:10:00,0.0,0.0,%,
2855,3423,5,178980.0,51146,2103-02-02 06:10:00,0.0,0.0,%,
2856,3424,5,178980.0,51200,2103-02-02 06:10:00,0.0,0.0,%,
2857,3425,5,178980.0,51221,2103-02-02 06:10:00,43.0,43.0,%,
2858,3426,5,178980.0,51222,2103-02-02 06:10:00,14.9,14.9,g/dL,
2859,3427,5,178980.0,51244,2103-02-02 06:10:00,16.0,16.0,%,abnormal
2860,3428,5,178980.0,51248,2103-02-02 06:10:00,37.6,37.6,pg,abnormal
2861,3429,5,178980.0,51249,2103-02-02 06:10:00,34.7,34.7,%,
2862,3430,5,178980.0,51250,2103-02-02 06:10:00,109.0,109.0,fL,abnormal


In [None]:
def drop_nans()

In [None]:
binaryEncode
# lazy function 

# give it fields, give it OHE
# testtrainsplit
# drop nans
# drop columns - drop target
# runmodels
# metrics


# topic modelling
TfidfVectorizer()
vectorizer.fit_transform()
num_topics = 15
lda_model = LatentDiric

# use topics instead