In [2]:
from datasets import load_dataset, Dataset
from tqdm import tqdm
import openai
from dotenv import load_dotenv
import os
from openai import OpenAI

load_dotenv()

  from .autonotebook import tqdm as notebook_tqdm


True

In [3]:
client = OpenAI(api_key=os.environ['OPENAI_API_KEY'])

In [4]:
dataset = Dataset.from_json('/home/tkdrnjs0621/work/kmel-reasoning3/dataset/original/hiv_train_2k.jsonl')

In [5]:
dataset[0]

{'SELFIES': '[C][C][S][C][Branch1][C][N][=C][Branch1][S][S][=Branch1][C][=O][=Branch1][C][=O][C][=C][C][=C][C][=C][Ring1][=Branch1][C][=Ring1][#C][C]',
 'result': 'Yes.',
 'id': '28-6970c00840fe46729c1aef69c35b2819'}

In [7]:
completion = client.chat.completions.create(
    model="gpt-4.1",
    messages= [
    {"role": "system", "content": """You are a chemical domain expert specializing in molecular property prediction.
You will be provided with a SELFIES representation of a molecule.
Your task is to determine whether the given molecule functions as an inhibitor of the human immunodeficiency virus (HIV).

Analyze the molecule and provide a detailed explanation of your reasoning, citing relevant chemical properties, structural features, or known biological activity.

After your reasoning, insert a new line.

On the following line, present your conclusion in the exact format of "ANSWER: YES", if the molecule can inhibit HIV replication, and "ANSWER: NO" if it cannot.
""".strip()},
    {"role": "user", "content": dataset[0]['SELFIES']},
],
    temperature=1,
    top_p=1,
)

print(completion.choices[0].message.content)

Let's analyze the SELFIES string:  
[C][C][O][P][=Branch1][C][=O][Branch1][=N][N][C][=C][C][=C][C][Branch1][C][Cl][=C][Ring1][#Branch1][O][C][C]

**Stepwise translation and interpretation:**
- [C][C][O] : Suggests an ethoxy group.
- [P][=Branch1][C][=O] : Phosphoryl group with adjacent carbonyl.
- [Branch1][=N][N][C][=C][C][=C][C] : Likely part of a heterocycle, maybe a pyrazine or pyrimidine system.
- [Branch1][C][Cl][=C][Ring1] : Cyclized ring possibly with a chloro substituent.
- [#Branch1][O][C][C] : Terminal group, possibly an alcohol or ether.

**Identifying structural motifs:**  
- Contains a phosphorus center with C, O, and (possibly) N linkages, reminiscent of phosphonomethyl moieties.
- Heterocyclic content with nitrogen involvement (N–N–C connectivity), possibly indicating a substituted pyrimidine/purine or related heterocycle.
- There is a chloro substituent and what appears to be an aromatic (or heteroaromatic) fused ring.

**Comparison to known HIV inhibitors:**  
- Many 

In [7]:
txt = "SELFIES: "+dataset[0]['SELFIES']+'\nGT Answer: '+dataset[0]['result'][:-1]
completion = client.chat.completions.create(
    model="gpt-4.1",
    messages= [
    {"role": "system", "content": """You are a chemical domain expert specializing in molecular property prediction.
You will be provided with a SELFIES representation of a molecule and a ground truth answer.
Your task is to give a detailed reasoning and determine if the given molecule functions as an inhibitor of the human immunodeficiency virus (HIV).

You must give an detailed explanation of the reasoning process. Use numbering and bold for each relevant point categories.
For every relevant point, provide a clear, concise paragraph explaining its role and significance in the context of potential HIV inhibition. 

The ground truth answer is given. The final conclusion must be identical to the ground truth, but you must not use or reference the ground truth explicitly while reasoning.
Do not write phrases such as “since the answer is…” or any variant that indicates reliance on the label.

After your reasoning, insert a new line.

On the following line, present your conclusion in the exact format of "ANSWER: YES", if the molecule can inhibit HIV replication, and "ANSWER: NO" if it cannot.
""".strip()},
    {"role": "user", "content": txt},
],
    temperature=1,
    top_p=1,
)

print(completion.choices[0].message.content)

**1. Presence of Functional Groups Associated with HIV Inhibition**  
A key consideration for HIV inhibition is the presence of chemical functionalities known to interfere with viral enzymes, particularly reverse transcriptase or protease. The SELFIES string indicates the existence of carbonyl groups (C=O), thiol (–S–), and nitrile (–C#N) functionalities. Carbonyl groups often participate in hydrogen bonding with enzyme active sites, while sulfur-containing groups are capable of forming disulfide bridges or binding to metal cofactors present in viral enzymes.

**2. Molecular Framework and Conjugation**  
The molecule exhibits extended conjugation, as evidenced by multiple alternating double bonds ([=C][C][=C][C][=C][C][=C]). Such conjugated systems can facilitate strong π-π stacking interactions, which improve binding affinity to nucleic acid-binding sites of viral proteins. Aromatic or pseudo-aromatic frameworks are common in known HIV inhibitors, enhancing their ability to intercalat

In [6]:
print(""""Yes.","id":"25-d8151750ccf44ee4954a838e11c9ae3c","reasoning":"Let's analyze the SELFIES string:\n\n1. Structure Overview:\n- The representation suggests a tricyclic or bicyclic system, with branches, sulfur (S), multiple double bonds (=C), oxygen atoms ([O]), and a terminal amine ([N]).\n- There's an ether linkage ([O][C]), a sulfur heterocycle ([S] within rings), and other functionalities, as well as several branches suggesting substituents.\n\n2. Blood-Brain Barrier (BBB) Penetration Determinants:\n- Key properties influencing BBBP:\n  - Molecular weight (MW) ideally < 450 Da.\n  - High lipophilicity (logP between 1\u20134).\n  - Low number of hydrogen bond donors (\u22642) and acceptors (\u22645).\n  - Low polar surface area (PSA, < 90 \u00c5\u00b2 favored).\n  - Few or no ionizable groups at physiological pH.\n\n3. Functional Group Analysis:\n- The molecule has two oxygens likely as ether or alcohol (potential hydrogen bond acceptors).\n- It has a terminal primary\/secondary amine ([N]) and a possible thioether or thiophene ([S]).\n- Extensive hydrocarbon skeleton and aromatic or unsaturated rings, likely increasing lipophilicity.\n- Branching and conjugation probably further increase hydrophobicity.\n- Two hydrogen bond donors (OHs or NHs) and possibly 3\u20134 acceptors (O, N, S).\n\n4. Size Estimate:\n- Given the multiple rings, several branches, and length, the molecular weight is likely above 400 Da but possibly below 550 Da, depending on exact structure.\n\n5. PSA and logP:\n- PSA is likely moderate (expected 40\u201380 \u00c5\u00b2).\n- The hydrocarbon content and heterocyclic framework suggest a logP in the range consistent with CNS drugs (likely 2\u20134).\n\n6. Structural Red Flags:\n- No obvious strongly ionizable groups at physiological pH.\n- No quaternary ammonium, carboxylate, or sulfonate groups.\n- Aromatic and sulfur-containing heterocycles are seen in several CNS-active drugs.\n\n7. Comparisons to Known BBB+ Structures:\n- Structural features (heterocycles, low-to-moderate H-bonding, moderate size, aromaticity) are common among molecules that do cross the BBB (e.g., tricyclic antidepressants like imipramine, antipsychotics).\n\n8. Possible Limitation:\n- The presence of [S] can indicate higher PSA, but if embedded in an aromatic ring (thiophene, thiazole), its impact is reduced.\n\n**Summary:**\nThis molecule's estimated size, hydrogen bonding profile, and lipophilicity align with trends seen in BBB-permeable drugs. Its features do not strongly preclude BBB penetration.\n\nANSWER: YES""")

"Yes.","id":"25-d8151750ccf44ee4954a838e11c9ae3c","reasoning":"Let's analyze the SELFIES string:

1. Structure Overview:
- The representation suggests a tricyclic or bicyclic system, with branches, sulfur (S), multiple double bonds (=C), oxygen atoms ([O]), and a terminal amine ([N]).
- There's an ether linkage ([O][C]), a sulfur heterocycle ([S] within rings), and other functionalities, as well as several branches suggesting substituents.

2. Blood-Brain Barrier (BBB) Penetration Determinants:
- Key properties influencing BBBP:
  - Molecular weight (MW) ideally < 450 Da.
  - High lipophilicity (logP between 1–4).
  - Low number of hydrogen bond donors (≤2) and acceptors (≤5).
  - Low polar surface area (PSA, < 90 Å² favored).
  - Few or no ionizable groups at physiological pH.

3. Functional Group Analysis:
- The molecule has two oxygens likely as ether or alcohol (potential hydrogen bond acceptors).
- It has a terminal primary\/secondary amine ([N]) and a possible thioether or thioph