### Explore Biomedical QA in OpenBioMed

BioMedGPT-R1-17B is composed of a reasoning language model (BioMedGPT-R1-LM-14B), a molecule encoder (GraphMVP), a protein encoder (ESM2-3B) and two modality adaptors. Approximately 70GB GPU Memory is required to load BioMedGPT-R1-17B (40GB GPU with fp16).

Feel free to [download](https://huggingface.co/PharMolix/BioMedGPT-R1) our trained models, put them under checkpoints, and explore the biomedical QA!

In [None]:
# Change working directory
import os
import sys
parent = os.path.dirname(os.path.abspath(''))
print(parent)
sys.path.append(parent)
os.chdir(parent)

import logging
logging.basicConfig(level=logging.ERROR)

In OpenBioMed, we provide an unified interface for multimodal biomedical chatbot. To build an agent, you just need to configure the path to the model checkpoint and which device to deploy the model.

In [2]:
from open_biomed.models.foundation_models.biomedgpt import BioMedGPTR14Chat
from open_biomed.data import Molecule, Protein, Text

# Initialize chatbot
agent = BioMedGPTR14Chat.from_pretrained("./checkpoints/biomedgpt-r1", "cuda:7")

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.


Here we provide examples on molecule QA and protein QA. 



In [3]:
# Molecule QA
molecule = "CN1CC[C@]23[C@@H]4[C@H]5[C@@H](CC2=O)C(=CCO[C@H]5CC(=O)N4C6=C3C=C(C=C6)O)C1"
question = "Please describe this molecule."

# Initialize a molecule via SMILES string and append it to context
# Then ask a question about it
agent.reset()
molecule = Molecule.from_smiles(molecule)
agent.append_molecule(molecule)
answer = agent.chat(Text.from_str(question))

print("Human: ", question)
print ("Assistant: ", answer)

Human:  Please describe this molecule.
Assistant:  The molecule is a monoterpenoid indole alkaloid with formula C21H24N2O4, originally isolated from Tabernaemontana corymbosa. It has a role as a plant metabolite. It is a delta-lactam, a monoterpenoid indole alkaloid and an organic heterohexacyclic compound.


In [4]:
# Protein QA
protein = "MDCVLRSYLLLAFGFLICLFLFCLVVFIWFVYKQILFRTTAQSNEARHNHSTVV"
question = "There is a mutation from G to A at position 14. What is the mutational effects?"

# Initialize a protein with an amino acid sequence and append it to context
# Then ask a question about it
agent.reset()
protein = Protein.from_fasta(protein)
agent.append_protein(protein)
answer = agent.chat(Text.from_str(question))

print("Human: ", question)
print ("Assistant: ", answer)

Human:  There is a mutation from G to A at position 14. What is the mutational effects?
Assistant:  Okay, I'm trying to figure out the mutational effects. Hmm, that looks a bit complicated, but I can try to break it down.

First, I'm thinking about the function of the provided protein. Transports viral genome to neighboring plant cells via plasmodesmata, resulting in systemic infection. The movement protein oligomerizes plasmodesmata and enlarges their pore size (Probable). 

Then there's a mutation from G to A at position 14. What is the mutational effects? Wait. I know it.
</think>

Complete loss of movement.


Of course, you can also have some regular text-only chat.

In [8]:
# Text-only Chat
question = "Give me some brief advice for future R&D on EGFR-targeted drugs."

agent.reset()
answer = agent.chat(Text.from_str(question))

print("Human: ", question)
print ("Assistant: ", answer)

Human:  Give me some brief advice for future R&D on EGFR-targeted drugs.
Assistant:  Okay, so I need to come up with some brief advice for future R&D on EGFR-targeted drugs. Hmm, where do I start? I remember that EGFR stands for epidermal growth factor receptor, and it's involved in cell proliferation and survival. Mutations in the EGFR gene can lead to cancer, so targeting EGFR could be a way to treat various cancers.

The user mentioned resistance mechanisms like T790M mutation and C797T mutation. I think the T790M mutation occurs in the exon 20 region of the EGFR gene and causes resistance to first-generation EGFR tyrosine kinase inhibitors (TKIs) like gefitinib and erlotinib. The C797T mutation occurs in the exon 19 region of the EGFR gene and causes resistance to third-generation EGFR TKIs like osimertinib. 

So, what can be done about these resistance mechanisms? One idea is to combine EGFR TKIs with other targeted therapies to overcome resistance. For example, combining EGFR TKI

We also provide an intermediate version of BioMedGPT-R1, [BioMedGPT-R1-DS-Qwen](https://huggingface.co/PharMolix/BioMedGPT-R1-DS-Qwen), which is only finetuned for cross-modal alignment. Here is an example on multimodal reasoning.

In [2]:
from open_biomed.models.foundation_models.biomedgpt import BioMedGPTR14Chat
from open_biomed.data import Molecule, Protein, Text

# Initialize chatbot
agent = BioMedGPTR14Chat.from_pretrained("./checkpoints/biomedgpt-r1-ds-qwen", "cuda:7")

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.


In [8]:
# Multimodal Reasoning
molecule = "CN1C(=O)CN=C(C2=C1C=CC(=C2)Cl)C3=CC=CC=C3"
question = "Please describe this molecule and figure out potential applications in pharmaceuticals."

# Initialize a molecule via SMILES string and append it to context
# Then ask a question about it
agent.reset()
molecule = Molecule.from_smiles(molecule)
agent.append_molecule(molecule)
answer = agent.chat(Text.from_str(question))

print("Human: ", question)
print ("Assistant: ", answer)

Human:  Please describe this molecule and figure out potential applications in pharmaceuticals.
Assistant:  Okay, so I'm trying to figure out what this molecule is and how it might be useful in pharmaceuticals. Let me start by looking at the SMILES notation provided: CN1C(=O)CN=C(C2=C1C=CC(=C2)Cl)C3=CC=CC=C3. Hmm, that looks a bit complicated, but I think I can break it down.

First, I notice that there's a ring structure involved because of the numbers in the SMILES string. The "C1" and "C2" suggest that there are two rings connected somehow. The part "CN1C(=O)CN=C(C2=C1C=CC(=C2)Cl)" seems to describe a bicyclic structure with a ketone group (C(=O)) and a chlorine substituent (Cl). The other part, "C3=CC=CC=C3", looks like a benzene ring attached to the rest of the molecule.

So, putting it all together, this molecule has a bicyclic core with a ketone and a chlorine atom, and it's connected to a benzene ring. I wonder what kind of properties this structure might have. Maybe it's a typ