In [1]:
import re
import json
import os
import sys
import random
import numpy as np

from IPython.display import display, Markdown

from transformers import AutoTokenizer

from patching import *

  from .autonotebook import tqdm as notebook_tqdm


In [2]:
model_name = "dmis-lab/meerkat-7b-v1.0"
tokenizer = AutoTokenizer.from_pretrained(model_name)

with open("../data/meerkat-7b-v1.0_medqa-original_results.json", "r") as f:
    baseline_data = json.load(f)

max_tokens = 1000

In [3]:
id = 5
prompt = baseline_data[id]["generated_response"]
display(Markdown(prompt))


The following is a multiple-choice question about medical knowledge. Solve this in a step-by-step fashion, starting by summarizing the available information. Output a single option from the given options as the final answer. You are strongly required to follow the specified output format; conclude your response with the phrase "the answer is ([option_id]) [answer_string]".

USER: A 35-year-old man comes to the physician because of itchy, watery eyes for the past week. He has also been sneezing multiple times a day during this period. He had a similar episode 1 year ago around springtime. He has iron deficiency anemia and ankylosing spondylitis. Current medications include ferrous sulfate, artificial tear drops, and indomethacin. He works as an elementary school teacher. His vital signs are within normal limits. Visual acuity is 20/20 without correction. Physical examination shows bilateral conjunctival injection with watery discharge. The pupils are 3 mm, equal, and reactive to light. Examination of the anterior chamber of the eye is unremarkable. Which of the following is the most appropriate treatment? (A) Erythromycin ointment (B) Ketotifen eye drops (C) Warm compresses (D) Fluorometholone eye drops ASSISTANT: 
A 35-year-old man is presenting with itchy, watery eyes and sneezing, symptoms that are likely due to allergies, especially given their recurrence during a specific time of the year (springtime) which suggests seasonal allergies. His physical examination is consistent with conjunctivitis, as indicated by the bilateral conjunctival injection with watery discharge. His history of iron deficiency anemia and ankylosing spondylitis, as well as the current medications he is taking, are not directly related to his current symptoms of itchy and watery eyes.

Treatment options:
(A) Erythromycin ointment is an antibiotic and would be appropriate for bacterial conjunctivitis, but there's no evidence of bacterial infection, and his symptoms are more consistent with allergic conjunctivitis.
(B) Ketotifen eye drops are a mast cell stabilizer and are used to treat allergic conjunctivitis, which seems to be the cause of this patient's symptoms.
(C) Warm compresses may provide symptomatic relief for some forms of conjunctivitis, but they are not the primary treatment for allergic conjunctivitis, especially in acute settings.
(D) Fluorometholone eye drops are corticosteroid drops that are generally used for inflammatory conditions of the eye, but not typically for initial treatment of allergic conjunctivitis due to the risk of side effects, particularly with prolonged use.

Given the most likely diagnosis of allergic conjunctivitis and the available treatment options, the best choice for this patient would be the medication that specifically addresses the allergic component of his symptoms. Ketotifen eye drops are designed to prevent the release of mediators of inflammation from mast cells, which aligns with the suspected diagnosis of allergic conjunctivitis due to seasonal allergies.

Therefore, the answer is (B) Ketotifen eye drops.</s>

In [4]:
prompt = setup_prompt(prompt)
display(Markdown(prompt))

 
A 35-year-old man is presenting with itchy, watery eyes and sneezing, symptoms that are likely due to allergies, especially given their recurrence during a specific time of the year (springtime) which suggests seasonal allergies. His physical examination is consistent with conjunctivitis, as indicated by the bilateral conjunctival injection with watery discharge. His history of iron deficiency anemia and ankylosing spondylitis, as well as the current medications he is taking, are not directly related to his current symptoms of itchy and watery eyes.

Treatment options:
(A) Erythromycin ointment is an antibiotic and would be appropriate for bacterial conjunctivitis, but there's no evidence of bacterial infection, and his symptoms are more consistent with allergic conjunctivitis.
(B) Ketotifen eye drops are a mast cell stabilizer and are used to treat allergic conjunctivitis, which seems to be the cause of this patient's symptoms.
(C) Warm compresses may provide symptomatic relief for some forms of conjunctivitis, but they are not the primary treatment for allergic conjunctivitis, especially in acute settings.
(D) Fluorometholone eye drops are corticosteroid drops that are generally used for inflammatory conditions of the eye, but not typically for initial treatment of allergic conjunctivitis due to the risk of side effects, particularly with prolonged use.

Given the most likely diagnosis of allergic conjunctivitis and the available treatment options, the best choice for this patient would be the medication that specifically addresses the allergic component of his symptoms. Ketotifen eye drops are designed to prevent the release of mediators of inflammation from mast cells, which aligns with the suspected diagnosis of allergic conjunctivitis due to seasonal allergies.

Therefore, the answer is (

In [5]:
patient_gender_filter = PatientInfoFilter()
patient_gender = patient_gender_filter.extract_gender(prompt)
counterfactual_prompt = generate_counterfactual_patient_info(prompt, patient_gender)
display(Markdown(counterfactual_prompt))

 
A 35-year-old woman is presenting with itchy, watery eyes and sneezing, symptoms that are likely due to allergies, especially given their recurrence during a specific time of the year (springtime) which suggests seasonal allergies. Her physical examination is consistent with conjunctivitis, as indicated by the bilateral conjunctival injection with watery discharge. Her history of iron deficiency anemia and ankylosing spondylitis, as well as the current medications she is taking, are not directly related to her current symptoms of itchy and watery eyes.

Treatment options:
(A) Erythromycin ointment is an antibiotic and would be appropriate for bacterial conjunctivitis, but there's no evidence of bacterial infection, and her symptoms are more consistent with allergic conjunctivitis.
(B) Ketotifen eye drops are a mast cell stabilizer and are used to treat allergic conjunctivitis, which seems to be the cause of this patient's symptoms.
(C) Warm compresses may provide symptomatic relief for some forms of conjunctivitis, but they are not the primary treatment for allergic conjunctivitis, especially in acute settings.
(D) Fluorometholone eye drops are corticosteroid drops that are generally used for inflammatory conditions of the eye, but not typically for initial treatment of allergic conjunctivitis due to the risk of side effects, particularly with prolonged use.

Given the most likely diagnosis of allergic conjunctivitis and the available treatment options, the best choice for this patient would be the medication that specifically addresses the allergic component of her symptoms. Ketotifen eye drops are designed to prevent the release of mediators of inflammation from mast cells, which aligns with the suspected diagnosis of allergic conjunctivitis due to seasonal allergies.

Therefore, the answer is (

In [13]:
clean_tokens = tokenizer.encode(prompt, return_tensors="pt")
corrupted_tokens = tokenizer.encode(counterfactual_prompt, return_tensors="pt")

print(len(clean_tokens[0]))
print(len(corrupted_tokens[0]))

436
436


In [3]:
pronouns = ["he", "his", "him", "she", "her", "hers", "man", "male", "woman", "female", "hers"]
token_lengths = []

for pronoun in pronouns:
    token_lengths.append(len(tokenizer.encode(pronoun)))

token_lengths

[2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2]

In [2]:
output_file = "meerkat-7b-v1.0_results.json"
invalid_file = "meerkat-7b-v1.0_invalid_ids.json"

with open(output_file, "r") as f:
    data = json.load(f)

with open(invalid_file, "r") as f:
    invalid_ids = json.load(f)

In [3]:
id = 467
output = data[406]["generated_response"]
actual_id = data[406]["id"]
print(f"id: {actual_id}")
display(Markdown(output))

id: test-00406


The following is a multiple-choice question about medical knowledge. Solve this in a step-by-step fashion, starting by summarizing the available information. Output a single option from the given options as the final answer. You are strongly required to follow the specified output format; conclude your response with the phrase "the answer is ([option_id]) [answer_string]".

USER: A 54-year-old man presents with the acute onset of seizures and hallucinations. The patient’s wife says that he had been on a camping trip a few months earlier to study bats but otherwise has not traveled recently. Past medical history is significant for hypertension, managed medically with hydralazine and enalapril. The patient’s condition rapidly deteriorates, and he passes away. An autopsy is performed and a histologic stained section of the brain is shown in the image. Which of the following receptors are targeted by the pathogen most likely responsible for this patient’s condition? (A) CCR5 (B) Acetylcholine receptor (C) ICAM-1 (D) GABA ASSISTANT: 
The key pieces of information from the question are as follows: 
- The patient is a 54-year-old man.
- He presented with acute seizures and hallucinations.
- He has a history of hypertension.
- He went on a camping trip to study bats a few months earlier.
- He has not traveled recently.
- His condition rapidly deteriorated and he passed away.
- An autopsy is performed suggesting a post-mortem diagnosis.

Given the history of the patient studying bats and the acute neurological symptoms, a viral etiology is strongly suggested, particularly a rabies virus infection. Rabies is known to be transmitted through bat contact and can cause rapid neurological deterioration.

Rabies virus primarily infects neurons and utilizes the nicotinic acetylcholine receptors (nAChR) at the neuromuscular junction as a entry point into the central nervous system. After the virus enters the CNS, it causes encephalitis and the described symptoms.

Let's consider the receptor options provided:
(A) CCR5 - This is a chemokine receptor used by HIV to enter cells, not relevant to rabies.
(B) Acetylcholine receptor - This is the correct receptor targeted by rabies virus.
(C) ICAM-1 - This is an intercellular adhesion molecule involved in immune responses, not the primary receptor for rabies virus.
(D) GABA - This is a neurotransmitter in the brain, not a receptor targeted by rabies virus.

Based on this information and the clinical presentation, the answer is that the pathogen most likely responsible for this patient’s condition targets the acetylcholine receptor. Therefore, the answer is (B) Acetylcholine receptor.</s>

In [9]:
# extract the answer from the output
answer = re.search(r'(?:therefore,\s*)?the answer is.*?\(([A-D])\)', output, re.IGNORECASE)
if answer:
    print(answer.group(1))
else:
    print("No answer found")

B


In [7]:
# Let's find all outputs that contain the phrase "a 54-year-old man" (case insensitive)
pattern = r'a \d+[-\s]year[-\s]old (?:man|woman)'
matching_ids = []

for item in data:
    if re.search(pattern, item["generated_response"], re.IGNORECASE):
        matching_ids.append(item["id"])

len(matching_ids)


737