In [1]:
import re
import json
import os
import sys
import random
import numpy as np

from IPython.display import display, Markdown

from transformers import AutoTokenizer

from patching import *

  from .autonotebook import tqdm as notebook_tqdm


In [2]:
model_name = "dmis-lab/meerkat-7b-v1.0"
tokenizer = AutoTokenizer.from_pretrained(model_name)

with open("../data/meerkat-7b-v1.0_medqa-original_results.json", "r") as f:
    baseline_data = json.load(f)

max_tokens = 1000

In [3]:
id = 3
prompt = baseline_data[id]["generated_response"]
display(Markdown(prompt))

The following is a multiple-choice question about medical knowledge. Solve this in a step-by-step fashion, starting by summarizing the available information. Output a single option from the given options as the final answer. You are strongly required to follow the specified output format; conclude your response with the phrase "the answer is ([option_id]) [answer_string]".

USER: A 39-year-old man presents to the emergency department because of progressively worsening chest pain and nausea that started at a local bar 30 minutes prior. The pain radiates to the epigastric area. He has a 5-year history of untreated hypertension. He has smoked 1 pack of cigarettes daily for the past 5 years and started abusing cocaine 2 weeks before his emergency room visit. The patient is diaphoretic and in marked distress. What should be the first step in management? (A) Diltiazem (B) Labetalol (C) Propranolol (D) Reassurance and continuous monitoring ASSISTANT: 

The presented case describes a 39-year-old man who is experiencing progressively worsening chest pain, nausea, has a history of untreated hypertension, smokes, and has recently started abusing cocaine. The description of his pain radiating to the epigastric area and the association with nausea suggests the possibility of an acute coronary syndrome, a condition that can be precipitated by cocaine abuse due to its effects on the cardiovascular system, including vasospasm and hypertension.

Given the acute presentation and the high likelihood that the patient is experiencing a coronary event, the initial management should focus on stabilization, which includes continuous monitoring, assessment for revascularization procedures, and medical management to address the acute coronary syndrome and hypertension.

Options A (Diltiazem), B (Labetalol), and C (Propranolol) are beta-blockers and calcium channel blockers which can be used in hypertensive emergencies. However, in the context of acute coronary syndrome, non-selective beta-blockers such as propranolol are not recommended as first-line agents due to potential adverse effects on myocardial contractility and coronary perfusion. Labetalol has both alpha and beta-blocking effects and could be used cautiously. Diltiazem is a calcium channel blocker that can be used to control blood pressure in acute coronary syndromes.

Reassurance and continuous monitoring (option D) would be an appropriate first step in the management of this patient to stabilize his condition while preparing for further diagnostic and therapeutic interventions. It is essential to monitor the patient for any changes in his condition that might necessitate immediate action.

Therefore, the answer is (D) Reassurance and continuous monitoring.</s>

In [4]:
prompt = setup_prompt(prompt)
display(Markdown(prompt))

 

The presented case describes a 39-year-old man who is experiencing progressively worsening chest pain, nausea, has a history of untreated hypertension, smokes, and has recently started abusing cocaine. The description of his pain radiating to the epigastric area and the association with nausea suggests the possibility of an acute coronary syndrome, a condition that can be precipitated by cocaine abuse due to its effects on the cardiovascular system, including vasospasm and hypertension.

Given the acute presentation and the high likelihood that the patient is experiencing a coronary event, the initial management should focus on stabilization, which includes continuous monitoring, assessment for revascularization procedures, and medical management to address the acute coronary syndrome and hypertension.

Options A (Diltiazem), B (Labetalol), and C (Propranolol) are beta-blockers and calcium channel blockers which can be used in hypertensive emergencies. However, in the context of acute coronary syndrome, non-selective beta-blockers such as propranolol are not recommended as first-line agents due to potential adverse effects on myocardial contractility and coronary perfusion. Labetalol has both alpha and beta-blocking effects and could be used cautiously. Diltiazem is a calcium channel blocker that can be used to control blood pressure in acute coronary syndromes.

Reassurance and continuous monitoring (option D) would be an appropriate first step in the management of this patient to stabilize his condition while preparing for further diagnostic and therapeutic interventions. It is essential to monitor the patient for any changes in his condition that might necessitate immediate action.

Therefore, the answer is (

In [6]:
patient_gender_filter = PatientInfoFilter()
patient_gender = patient_gender_filter.extract_gender(prompt)
counterfactual_prompt = generate_counterfactual_patient_info(prompt, patient_gender)
display(Markdown(counterfactual_prompt))

 

The presented case describes a 39-year-old woman who is experiencing progressively worsening chest pain, nausea, has a history of untreated hypertension, smokes, and has recently started abusing cocaine. The description of her pain radiating to the epigastric area and the association with nausea suggests the possibility of an acute coronary syndrome, a condition that can be precipitated by cocaine abuse due to its effects on the cardiovascular system, including vasospasm and hypertension.

Given the acute presentation and the high likelihood that the patient is experiencing a coronary event, the initial management should focus on stabilization, which includes continuous monitoring, assessment for revascularization procedures, and medical management to address the acute coronary syndrome and hypertension.

Options A (Diltiazem), B (Labetalol), and C (Propranolol) are beta-blockers and calcium channel blockers which can be used in hypertensive emergencies. However, in the context of acute coronary syndrome, non-selective beta-blockers such as propranolol are not recommended as first-line agents due to potential adverse effects on myocardial contractility and coronary perfusion. Labetalol has both alpha and beta-blocking effects and could be used cautiously. Diltiazem is a calcium channel blocker that can be used to control blood pressure in acute coronary syndromes.

Reassurance and continuous monitoring (option D) would be an appropriate first step in the management of this patient to stabilize her condition while preparing for further diagnostic and therapeutic interventions. It is essential to monitor the patient for any changes in her condition that might necessitate immediate action.

Therefore, the answer is (

In [7]:
clean_tokens = tokenizer.encode(prompt, return_tensors="pt")
corrupted_tokens = tokenizer.encode(counterfactual_prompt, return_tensors="pt")

print(len(clean_tokens[0]))
print(len(corrupted_tokens[0]))

397
397


In [3]:
pronouns = ["he", "his", "him", "she", "her", "hers", "man", "male", "woman", "female", "hers"]
token_lengths = []

for pronoun in pronouns:
    token_lengths.append(len(tokenizer.encode(pronoun)))

token_lengths

[2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2]

In [2]:
output_file = "meerkat-7b-v1.0_results.json"
invalid_file = "meerkat-7b-v1.0_invalid_ids.json"

with open(output_file, "r") as f:
    data = json.load(f)

with open(invalid_file, "r") as f:
    invalid_ids = json.load(f)

In [3]:
id = 467
output = data[406]["generated_response"]
actual_id = data[406]["id"]
print(f"id: {actual_id}")
display(Markdown(output))

id: test-00406


The following is a multiple-choice question about medical knowledge. Solve this in a step-by-step fashion, starting by summarizing the available information. Output a single option from the given options as the final answer. You are strongly required to follow the specified output format; conclude your response with the phrase "the answer is ([option_id]) [answer_string]".

USER: A 54-year-old man presents with the acute onset of seizures and hallucinations. The patient’s wife says that he had been on a camping trip a few months earlier to study bats but otherwise has not traveled recently. Past medical history is significant for hypertension, managed medically with hydralazine and enalapril. The patient’s condition rapidly deteriorates, and he passes away. An autopsy is performed and a histologic stained section of the brain is shown in the image. Which of the following receptors are targeted by the pathogen most likely responsible for this patient’s condition? (A) CCR5 (B) Acetylcholine receptor (C) ICAM-1 (D) GABA ASSISTANT: 
The key pieces of information from the question are as follows: 
- The patient is a 54-year-old man.
- He presented with acute seizures and hallucinations.
- He has a history of hypertension.
- He went on a camping trip to study bats a few months earlier.
- He has not traveled recently.
- His condition rapidly deteriorated and he passed away.
- An autopsy is performed suggesting a post-mortem diagnosis.

Given the history of the patient studying bats and the acute neurological symptoms, a viral etiology is strongly suggested, particularly a rabies virus infection. Rabies is known to be transmitted through bat contact and can cause rapid neurological deterioration.

Rabies virus primarily infects neurons and utilizes the nicotinic acetylcholine receptors (nAChR) at the neuromuscular junction as a entry point into the central nervous system. After the virus enters the CNS, it causes encephalitis and the described symptoms.

Let's consider the receptor options provided:
(A) CCR5 - This is a chemokine receptor used by HIV to enter cells, not relevant to rabies.
(B) Acetylcholine receptor - This is the correct receptor targeted by rabies virus.
(C) ICAM-1 - This is an intercellular adhesion molecule involved in immune responses, not the primary receptor for rabies virus.
(D) GABA - This is a neurotransmitter in the brain, not a receptor targeted by rabies virus.

Based on this information and the clinical presentation, the answer is that the pathogen most likely responsible for this patient’s condition targets the acetylcholine receptor. Therefore, the answer is (B) Acetylcholine receptor.</s>

In [9]:
# extract the answer from the output
answer = re.search(r'(?:therefore,\s*)?the answer is.*?\(([A-D])\)', output, re.IGNORECASE)
if answer:
    print(answer.group(1))
else:
    print("No answer found")

B


In [7]:
# Let's find all outputs that contain the phrase "a 54-year-old man" (case insensitive)
pattern = r'a \d+[-\s]year[-\s]old (?:man|woman)'
matching_ids = []

for item in data:
    if re.search(pattern, item["generated_response"], re.IGNORECASE):
        matching_ids.append(item["id"])

len(matching_ids)


737