In [3]:
# 在整个过程中使用到的package
import json
from openai import OpenAI
import os
import re

分布遍历循环的整个过程：
1，将所需要的数据集导入：形态学矩阵信息（以JSON format导入）以及调用API_key
2，第一次调用API为整个数据集中所有的species选择一个合适的分类character，将调用的结果（通过initial character根据不同状态将species分为几个组）来作为初始的分类结果
3，将这个初始的分类结果保存下来，通过解析这个分类结果，将这个分类结果存储为字典形式
4，通过分别依次遍历每个组中的内容，为每个组中的species选择分类character，通过记录下不同组中的分类结果，将其存储在环境中
5，通过将多次分布遍历的结果，整合，然后对格式的要求在完善一些，从而来得到一个完整的分类检索表，最终在整合结果的过程中选择合适的方法和格式进行优化

1，导入形态学矩阵数据集信息（JSON format）and 调用API_key

In [4]:
# 调用环境中的API_KEY
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
# 首先导入需要的数据集
with open("D:/桌面/TEST-KG/nexus fix/matrix_knowledge_graph_22.json", "r", encoding="utf-8") as file:
    matrix_data = json.load(file)
# 由于输入的内容为JSON格式，因此在调用给API时需要将这个JSON文件
matrix_data_str = json.dumps(matrix_data, ensure_ascii=False)

2，调用API选择初始分类character
通过基于information gain方法来为整个数据集中的所有species选择initial character，根据initial character的不同特征状态来获得不同的分组情况

In [8]:
# 这是输入给API的相关prompt，通过使用client.chat.completions.create接口来进行多轮次对话。通过赋予对话信息中不同的角色身份，包括user，assistant，system，来实现完整的将所有的输入信息传递给调用的API模型
messages = [
            # 对模型设置系统身份，使其能够专项于分类学任务，同时在系统构建中强调其对形态学矩阵，information gain以及分类检索表的构建
            {"role":"system",
             "content":
                 "You are a helpful taxonomist assistant."+ \
                 "You are skilled at calculating the correct information gain to choose the character that best divides species into even groups based on their states." +\
                 "Based on the selected character, classify the species into different groups according to their states." + \
                 "For each group with more than two species, continue selecting characters to further classify this group until each group only has one species." + \
                 "After multiple classifications, determine the final classification levels and record each classifying character and its state. " + \
                 "Finally, generate a taxonomic key."},
            # 通过输入主要的任务请求，构建检索表以及强调在进行分类时所需要采取的细节：主要目的，information gain进行计算，忽略无效状态干扰，明确多特征状态，information gain选择的意义指引其正确的选择，规范输出结果的最终格式，以便于在后续对该API结果的提取
            {"role":"system",
            "content": """
                Generate the taxonomic key based on the provided morphological matrix. The matrix includes all species and their different states for each character.\n
                The process involves selecting a character to classify the species into groups. Repeat this classification within each subgroup until each group contains only one species.\n
                Information gain measures how much the uncertainty in the dataset is reduced after using a character for classification. It helps in selecting characters that minimize the entropy of the subset after classification, leading to better classification results.\n
                Please select the initial classification character for all species based on the morphological matrix and information gain methods.\n
                In the morphological matrix, 'Missing' and 'Not applicable' are invalid states. If a character has invalid states for the group being classified, it should be ignored.\n
                States are represented by numbers, and like the '1 and 2' means multiple states should be treated as a single state type. The initial character should have no more than three state types.\n
                you need to calculate the information gain for each character, and choose the highest information gain result, The higher the Information Gain result, the greater the contribution of the feature to the classification.\n
                Now I will show you the morphological matrix. Please provide only the initial classification character and the categorization of species based on its state, and you should singly show this called as # initial character classify result # 
            """},
            # 通过使用assistant来对输入给API的prompt内容进行总结和细化，通过assistant中的对话在深化给API对于这个内容的理解，来一定程度上的控制API的响应结果的控制，规范API响应结果的输出格式
            {"role": "assistant",
            "content": """
                Understood. I will generate the taxonomic key based on the provided morphological matrix. Here is a summary of the steps I will follow:\n
                1. The matrix includes all species and their different states for each character.\n
                2. I will select a character to classify the species into groups and repeat this classification within each subgroup until each group contains only one species.\n
                3. I will use information gain to measure how much the uncertainty in the dataset is reduced after using a feature for classification. This helps in selecting features that minimize the entropy of the subset after classification, leading to better classification results.\n
                4. I will select the initial classification character for all species based on the morphological matrix and information gain methods.\n
                5. In the morphological matrix, 'Missing' and 'Not applicable' are considered invalid states. If a character has invalid states for the group being classified, it will be ignored.\n
                6. States are represented by numbers. For example, '1 and 2' means multiple states should be treated as a single state type. The initial character should have no more than three state types.\n
                7. I will use information gain to calculate all character and choose the highest information gain result, The higher the Information Gain result, the greater the contribution of the feature to the classification. \n
                8. The final result will provide only the initial classification character and the categorization of species based on its state. \n
                9. Don't need to show how i calculate, only need to show the final result, and i will store result in #initail character classify result# block
                Please provide the morphological matrix data so that I can proceed with the initial classification.
            """
            },
            # 输入对应的形态学矩阵信息，让其进行分析
            {"role":"user",
             "content":
                 f"Here is morphological matrix:{matrix_data_str}"}
            ]
# 设置在响应API时需要调控的各个参数，我感觉温度最好是调成0，然后max_token数要么不参与设置，要不然就不能设置太多，一是为了节约成本，二是为了生成对应的结果并不需要太长，如果token数量太大但是又没有设置终止符就会导致最终的结果出现复杂的冗余以及错误
initial_character_info =  client.chat.completions.create(
        model="gpt-4o",
        messages=messages,
        stop=None,
        max_tokens=512,
        temperature=0,
        n = 1)
# 存储API调用响应的结果，存储为文件（当讲后续分布API调用循环所连接起来时，需要考虑存储在环境变量中，以便于之后连续的调用修饰）
initial_response = initial_character_info.choices[0].message.content
with open("D:/桌面/TEST-KG/nexus fix/initial_character_info", 'w') as file:
    file.write(initial_response)
print(initial_response)

# Initial Character Classify Result #

Based on the provided morphological matrix, I have calculated the information gain for each character and selected the character with the highest information gain for the initial classification. The character with the highest information gain is **Character1**.

Here is the categorization of species based on the states of **Character1**:

- **State: 1**
  - Equisetum_litorale
  - Equisetum_palustre

- **State: 2**
  - Equisetum_pratense
  - Equisetum_hyemale
  - Equisetum_moorei
  - Equisetum_ramosissimum
  - Equisetum_trachyodon
  - Equisetum_variegatum

- **State: 1 and 2**
  - Equisetum_arvense
  - Equisetum_telmateia
  - Equisetum_sylvaticum
  - Equisetum_fluviatile

This is the initial classification based on **Character1**. Further classification within each group will be performed in subsequent steps.


3，保存初始调用API的结果，进行解析，存储为字典模式作为一个新的变量（**这个地方存在一个可能的挑战就是这个解析的过程，为了后续的分布循环进行，这一步的解析结果十分关键，但是解析的function需要对于API响应的结果进行调用，尽管在API的响应的prompt中明确规范了最终输出的内容，但是由于LLM的随机性等原因，我也无法保证每次解析都能够设置准确的结果（设置种子或者调整解析格式））

In [17]:
# 读取初始字符信息
with open("D:/桌面/TEST-KG/nexus fix/initial_character_info", "r", encoding="utf-8") as file:
    initial_character = file.read()
    # 可以通过使用 print(initial_character)
# 定义用于解析API响应结果的函数
def parse_classification_result(result_text):
    classification = {}
    # 使用正则表达式匹配每个状态块
    state_sections = re.findall(r'- \*\*State: ([\d\s\w&]+)\*\*\n((?:\s*-\s[A-Za-z_]+\n)+)', result_text)
    for state, species_block in state_sections:
        # 匹配物种列表
        species_list = re.findall(r'- ([A-Za-z_]+)', species_block)
        classification[state.strip()] = species_list
    return classification

# 解析初始分类结果
parsed_initial_classification = parse_classification_result(initial_character)
print(parsed_initial_classification)

# 示例输出解析结果
# for state, species in parsed_initial_classification.items():
    # print(f"State {state}: {', '.join(species)}")


{'1': ['Equisetum_litorale', 'Equisetum_palustre'], '2': ['Equisetum_pratense', 'Equisetum_hyemale', 'Equisetum_moorei', 'Equisetum_ramosissimum', 'Equisetum_trachyodon', 'Equisetum_variegatum'], '1 and 2': ['Equisetum_arvense', 'Equisetum_telmateia', 'Equisetum_sylvaticum', 'Equisetum_fluviatile']}
State 1: Equisetum_litorale, Equisetum_palustre
State 2: Equisetum_pratense, Equisetum_hyemale, Equisetum_moorei, Equisetum_ramosissimum, Equisetum_trachyodon, Equisetum_variegatum
State 1 and 2: Equisetum_arvense, Equisetum_telmateia, Equisetum_sylvaticum, Equisetum_fluviatile


此处的结果显示为（通过解析存储为一种字典格式来分别调用哪一行）：

In [65]:
# 此处为通过使用解析function对于initial character中所调用响应的结果

initial_classification = {
    "1": ["Equisetum_litorale", "Equisetum_palustre"],
    "2": ["Equisetum_pratense", "Equisetum_hyemale", "Equisetum_moorei", "Equisetum_ramosissimum", "Equisetum_trachyodon", "Equisetum_variegatum"],
    "1 and 2": ["Equisetum_arvense", "Equisetum_telmateia", "Equisetum_sylvaticum", "Equisetum_fluviatile"]
}

# 用于存储所有值的列表
groups = []
# 遍历字典中的每一个键值对，并将值存储在列表中，这是为后续设置分布循环设置基础，对于一个字典内容中的各个键值调用存储到一个变量中然后可以多次循环这几个变量最终存储结果
for state, species_list in initial_classification.items():
    groups.append(species_list)
# 打印存储在列表中的结果
for i, group in enumerate(groups, 1):
    print(f"Group {i} species:", group)
group_1_species = groups[0]
group_2_species = groups[1]
group_3_species = groups[2]
print(f"State 1: {group_1_species}",f"State 2: {group_2_species}",f"State 3: {group_3_species}")

['Equisetum_litorale', 'Equisetum_palustre']
['Equisetum_pratense', 'Equisetum_hyemale', 'Equisetum_moorei', 'Equisetum_ramosissimum', 'Equisetum_trachyodon', 'Equisetum_variegatum']
['Equisetum_arvense', 'Equisetum_telmateia', 'Equisetum_sylvaticum', 'Equisetum_fluviatile']
State 1: ['Equisetum_litorale', 'Equisetum_palustre'] State 2: ['Equisetum_pratense', 'Equisetum_hyemale', 'Equisetum_moorei', 'Equisetum_ramosissimum', 'Equisetum_trachyodon', 'Equisetum_variegatum'] State 3: ['Equisetum_arvense', 'Equisetum_telmateia', 'Equisetum_sylvaticum', 'Equisetum_fluviatile']


4，通过分别依次遍历每个组中的内容，为每个组中的species选择分类character（即我想要实现的分布循环调用API）（**但是目前我还没有很好的构建这个循环的结果，因此我可以目前先使用一种比较简单的方式，来测试对于不同组别中的调用的信息进行API调用）

In [48]:
# 通过对initial character的不同特征状态得到的species分类结果：对于key 1 即state 1进行进一步的分类
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
# 在每一次需要进行分类分析时，可以直接加载完整的形态学矩阵信息重新写入给API调用，这样可以避免API在进行reload时出现计算冗余，因此产生额外的计算负担
# 如果需要构建一个循环的话，或许可以使用这个方法将这个result设置为一个存储的空列表
# 设置一个空list用于存储results = {}
# 将所有的result[0],result[1],result[2]有一个有效的方法是通过将这个序号之前字典中对键值的存储序号所统一，因此统一这个[i]
results = {}

with open("D:/桌面/TEST-KG/nexus fix/matrix_knowledge_graph_22.json", "r") as file:
    data = json.load(file)

    messages2 = [
            {"role": "system",
             "content":"""
                 You are a helpful taxonomist assistant.\n
                 You are skilled at calculating the correct information gain to choose the character that best divides species into even groups based on their states.\n
                 Based on the selected character, classify the species into different groups according to their states.\n
                 For each group with more than two species, continue selecting characters to further classify this group until each group only has one species.\n
                 After multiple classifications, determine the final classification levels and record each classifying character and its state.\n
                 Finally, generate a taxonomic key.
                 """},
            {"role": "system",
             "content":"""
                Generate the taxonomic key based on the provided morphological matrix. \n
                The process involves selecting a character to classify the species into groups. Repeat this classification within each subgroup until each group contains only one species.
                Information gain measures how much the uncertainty in the dataset is reduced after using a character for classification. It helps in selecting characters that minimize the entropy of the subset after classification, leading to better classification results.
                Please select the classification character for these group's species based on the morphological matrix and information gain methods.
                In the morphological matrix, 'Missing' and 'Not applicable' are invalid states. If a character has invalid states for the group being classified, it should be ignored.
                States are represented by numbers. For example, '1 and 2' means multiple states should be treated as a single state type and this multi-state characterization should not be confused with the single states within it.The initial character should have no more than three state types.
                You need to calculate the information gain for each character and choose the highest information gain result. The higher the information gain result, the greater the contribution of the feature to the classification.
                After selecting the initial classification character and categorizing the species based on its state, repeat the process within each subgroup. For each subgroup, select the character with the highest information gain to further classify the species. Continue this process recursively until each group contains only one species.
                Now I will show you the morphological matrix. Please provide the classification character and the categorization of species based on its state. Then, continue to classify each subgroup recursively, showing the chosen character and categorization for each subgroup. Please present the result in a structured format, with each step clearly labeled.
                please don't show how you analysis and calculate, please show me the final result           
            """},
            {"role": "assistant",
             "content": """
                Understood. I will generate the taxonomic key based on the provided morphological matrix. Here is a summary of the steps I will follow:\n
                1. The matrix includes all species and their different states for each character.\n
                2. I will select a character to classify the species into groups and repeat this classification within each subgroup until each group contains only one species.\n
                3. I will use information gain to measure how much the uncertainty in the dataset is reduced after using a feature for classification. This helps in selecting features that minimize the entropy of the subset after classification, leading to better classification results.\n
                4. I will select the classification character for the group's species based on the morphological matrix and information gain methods.\n
                5. In the morphological matrix, 'Missing' and 'Not applicable' are considered invalid states. If a character has invalid states for the group being classified, it will be ignored.\n
                6. States are represented by numbers. For example, '1 and 2' means multiple states should be treated as a single state type, and This multi-state characterization should not be confused with the individual states within it. The classification character should have no more than three state types.\n
                7. I will use information gain to calculate all character and choose the highest inforamtion gain result, The higher the Information Gain result, the greater the contribution of the feature to the classification. \n
                8. The final result will provide only the initial classification character and the categorization of species based on its state. \n
                9. Don't need to show how the process about choose, only need to show the final result, and i will store result in #character classify result# block
                Please provide the group morphological matrix data so that I can proceed with the classification.
             """},
            # 在此处导入数据类型的时候，考虑直接通过之前在字典中存储的内容，来选择调用字典中不同的键值，以及再结合最开始的
            {"role": "user", "content": f"Here is the group information need to be classify: {groups[0]}, and here is the morphological matrix{json.dumps(data)}"}
        ]
    response =  client.chat.completions.create(
            model="gpt-4o",
            messages=messages2,
            stop=None,
            temperature=0,
            max_tokens=1000,
            n = 1
    )
    results[0] = response.choices[0].message.content
    result1 = results[0]
print(results[0])

# Character Classification Result

## Step 1: Initial Classification

### Selected Character: Character1

- **State 1**: Equisetum_litorale, Equisetum_palustre

## Step 2: Further Classification within State 1

### Selected Character: Character4

- **State 1**: Equisetum_litorale
- **State 1 and 2**: Equisetum_palustre

# Taxonomic Key

1. **Character1**
   - **State 1**
     1. **Character4**
        - **State 1**: Equisetum_litorale
        - **State 1 and 2**: Equisetum_palustre


通过再一次调用API来通过调用存储在字典中不同的键值，来实现对整个过程的分布循环，例如在这个循环的结果中需要调用的是字典中的键2的结果，则实际上是initial character通过state 2得到的为species的最终分类的结果，因此此处的调用是通过group_2_species，实际上来说可以通过调用groups[2]来调用这个相关的键值，来作为input，如果在后续使用groups[i]，可以通过循环的调用initial_classification字典中存储的结果，通过调用不同的键值来实现调用不同初始特征状态的分类结果，可以使code更加简洁明了，通过使用循环的方式，来减少代码的冗余

In [ ]:
# 通过对initial character的不同特征状态得到的species分类结果：对于key 2 即state 2进行进一步的分类
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
# 在每一次需要进行分类分析时，可以直接加载完整的形态学矩阵信息重新写入给API调用，这样可以避免API在进行reload时出现计算冗余，因此产生额外的计算负担
with open("D:/桌面/TEST-KG/nexus fix/matrix_knowledge_graph_22.json", "r") as file:
    data = json.load(file)

    messages3 = [
            {"role": "system",
             "content":"""
                 You are a helpful taxonomist assistant.\n
                 You are skilled at calculating the correct information gain to choose the character that best divides species into even groups based on their states.\n
                 Based on the selected character, classify the species into different groups according to their states.\n
                 For each group with more than two species, continue selecting characters to further classify this group until each group only has one species.\n
                 After multiple classifications, determine the final classification levels and record each classifying character and its state.\n
                 Finally, generate a taxonomic key.
                 """},
            {"role": "system","content":"""
                Generate the nested taxonomic key based on the provided morphological matrix. \n
                The process involves selecting a character to classify the species into groups. Repeat this classification within each subgroup until each group contains only one species.
                Information gain measures how much the uncertainty in the dataset is reduced after using a character for classification. It helps in selecting characters that minimize the entropy of the subset after classification, leading to better classification results.
                Please select the classification character for these group's species based on the morphological matrix and information gain methods.
                In the morphological matrix, 'Missing' and 'Not applicable' are invalid states. If a character has invalid states for the group being classified, it should be ignored.
                States are represented by numbers. For example, '1 and 2' means multiple states should be treated as a single state type and this multi-state characterization should not be confused with the single states within it (the state of '3' and '2 and 3' is different state, when you choose the character to based on the state to distinguish need to careful handle).The initial character should have no more than three state types.
                You need to calculate the information gain for each character and choose the highest information gain result. The higher the information gain result, the greater the contribution of the feature to the classification.
                After selecting the initial classification character and categorizing the species based on its state, repeat the process within each subgroup. For each subgroup, select the character with the highest information gain to further classify the species. Continue this process recursively until each group contains only one species.
                Now I will show you the morphological matrix. Please provide the classification character and the categorization of species based on its state. Then, continue to classify each subgroup recursively, showing the chosen character and categorization for each subgroup. Please present the result in a structured format, with each step clearly labeled.
                please don't show how you analysis and calculate, please show me the final result           
            """},
            {"role": "assistant",
             "content": """
                Understood. I will generate the nested taxonomic key based on the provided morphological matrix. Here is a summary of the steps I will follow:\n
                1. The matrix includes all species and their different states for each character.\n
                2. I will select a character to classify the species into groups and repeat this classification within each subgroup until each group contains only one species.\n
                3. I will use information gain to measure how much the uncertainty in the dataset is reduced after using a feature for classification. This helps in selecting features that minimize the entropy of the subset after classification, leading to better classification results.\n
                4. I will select the classification character for the group's species based on the morphological matrix and information gain methods.\n
                5. In the morphological matrix, 'Missing' and 'Not applicable' are considered invalid states. If a character has invalid states for the group being classified, it will be ignored.\n
                6. States are represented by numbers. For example, '2 and 3' means multiple states should be treated as a single state type, and This multi-state characterization should not be confused with the individual states(like '2', '3') within it (such as '3' and '2 and 3'  is the different state, these are two separate states, when i choose character to based on different state to distinguish the species). The classification character should have no more than three state types.\n
                7. I will use information gain to calculate all character and choose the highest information gain result, The higher the Information Gain result, the greater the contribution of the feature to the classification. \n
                8. The final result will provide only the initial classification character and the categorization of species based on its state. \n
                9. Don't need to show how the process about choose, only need to show the final result as nested structure, and i will store result in #character classify result# block
                Please provide the group morphological matrix data so that I can proceed with the classification.
             """},
            # 在此处导入数据类型的时候，考虑直接通过之前在字典中存储的内容，来选择调用字典中不同的键值，以及再结合最开始的
            {"role": "user", "content": f"Here is the group information need to be classify: {groups[1]}, and here is the morphological matrix{json.dumps(data)}"}
        ]
    response =  client.chat.completions.create(
            model="gpt-4o",
            messages=messages3,
            stop=None,
            temperature=0,
            max_tokens=1000,
            n = 1
    )
    result2 = response.choices[0].message.content
print(result2)

重复上面的过程，通过对initial character中的不同特征状态得到的结果，继续重复循环字典initial_classification中的存储的键值，来实现完整的分类

In [ ]:
# 通过对initial character的不同特征状态得到的species分类结果：对于key 2 即state 2进行进一步的分类
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
# 在每一次需要进行分类分析时，可以直接加载完整的形态学矩阵信息重新写入给API调用，这样可以避免API在进行reload时出现计算冗余，因此产生额外的计算负担
with open("D:/桌面/TEST-KG/nexus fix/matrix_knowledge_graph_22.json", "r") as file:
    data = json.load(file)

    messages3 = [
            {"role": "system",
             "content":"""
                 You are a helpful taxonomist assistant.\n
                 You are skilled at calculating the correct information gain to choose the character that best divides species into even groups based on their states.\n
                 Based on the selected character, classify the species into different groups according to their states.\n
                 For each group with more than two species, continue selecting characters to further classify this group until each group only has one species.\n
                 After multiple classifications, determine the final classification levels and record each classifying character and its state.\n
                 Finally, generate a taxonomic key.
                 """},
            {"role": "system","content":"""
                Generate the taxonomic key based on the provided morphological matrix. \n
                The process involves selecting a character to classify the species into groups. Repeat this classification within each subgroup until each group contains only one species.
                Information gain measures how much the uncertainty in the dataset is reduced after using a character for classification. It helps in selecting characters that minimize the entropy of the subset after classification, leading to better classification results.
                Please select the classification character for these group's species based on the morphological matrix and information gain methods.
                In the morphological matrix, 'Missing' and 'Not applicable' are invalid states. If a character has invalid states for the group being classified, it should be ignored.
                States are represented by numbers. For example, '1 and 2' means multiple states should be treated as a single state type and this multi-state characterization should not be confused with the single states within it.The initial character should have no more than three state types.
                You need to calculate the information gain for each character and choose the highest information gain result. The higher the information gain result, the greater the contribution of the feature to the classification.
                After selecting the initial classification character and categorizing the species based on its state, repeat the process within each subgroup. For each subgroup, select the character with the highest information gain to further classify the species. Continue this process recursively until each group contains only one species.
                Now I will show you the morphological matrix. Please provide the classification character and the categorization of species based on its state. Then, continue to classify each subgroup recursively, showing the chosen character and categorization for each subgroup. Please present the result in a structured format, with each step clearly labeled.
                please don't show how you analysis and calculate, please show me the final result           
            """},
            {"role": "assistant",
             "content": """
                Understood. I will generate the taxonomic key based on the provided morphological matrix. Here is a summary of the steps I will follow:\n
                1. The matrix includes all species and their different states for each character.\n
                2. I will select a character to classify the species into groups and repeat this classification within each subgroup until each group contains only one species.\n
                3. I will use information gain to measure how much the uncertainty in the dataset is reduced after using a feature for classification. This helps in selecting features that minimize the entropy of the subset after classification, leading to better classification results.\n
                4. I will select the classification character for the group's species based on the morphological matrix and information gain methods.\n
                5. In the morphological matrix, 'Missing' and 'Not applicable' are considered invalid states. If a character has invalid states for the group being classified, it will be ignored.\n
                6. States are represented by numbers. For example, '1 and 2' means multiple states should be treated as a single state type, and This multi-state characterization should not be confused with the individual states within it (should distinguish '3' and '2 and 3'，it is different state). The classification character should have no more than three state types.\n
                7. I will use information gain to calculate all character and choose the highest information gain result, The higher the Information Gain result, the greater the contribution of the feature to the classification. \n
                8. The final result will provide only the initial classification character and the categorization of species based on its state. \n
                9. Don't need to show how the process about choose, only need to show the final result, and i will store result in #character classify result# block
                Please provide the group morphological matrix data so that I can proceed with the classification.
             """},
            # 在此处导入数据类型的时候，考虑直接通过之前在字典中存储的内容，来选择调用字典中不同的键值，以及再结合最开始的
            {"role": "user", "content": f"Here is the group information need to be classify: {groups[2]}, and here is the morphological matrix{json.dumps(data)}"}
        ]
    response =  client.chat.completions.create(
            model="gpt-4o",
            messages=messages3,
            stop=None,
            temperature=0,
            max_tokens=1000,
            n = 1
    )
    result3 = response.choices[0].message.content
print(result3)

5，通过将多次分布遍历的结果（目前存在三次遍历结果的重复循环，因为基于初始的character的状态特征分为了三个组），整合（然后通过分别记录每个loop的结果，将其存储提取在一个list中，例如我可以规定在prompt中需要对每个分类的最终结果都以#taxonomic_information#存储，因此为了避免出现其他分析的冗余结果，我们可以将这个block（#taxonomic_information#）单独提取出来信息，来进行整合处理），然后对格式的要求在完善一些，从而来得到一个完整的分类检索表，最终在整合结果的过程中选择合适的方法和格式进行优化

In [1]:
#通过对之间进行分布遍历循环获得的每一步的结果，即对initial character中的所有不同特征状态的分类结果，再次调用API时进行分类的结果
# print(result1)
# print(result2)
# print(result3)
# 需要提取出这三个部分中最关键的部分，即在每个subgroup中存在的species的分类路径
# 在每个分布循环的结果中，我们需要规定对于最终产生的中应该将对于每个subgroup中的species的分类结果存储在block（Final Taxonomic Key），然后通过构建提取的函数来摘取出其中最终得到的分类信息

# 定义一个函数来提取每个结果中的所需部分
def extract_final_taxonomic_key(result):
    match = re.search(r'Final Taxonomic Key(.*)', result, re.DOTALL)
    if match:
        return match.group(1).strip()
    else:
        return None

# 提取每个结果中的所需部分
final_taxonomic_key1 = extract_final_taxonomic_key(result1)
final_taxonomic_key2 = extract_final_taxonomic_key(result2)
final_taxonomic_key3 = extract_final_taxonomic_key(result3)

# 打印提取的结果
print("Final Taxonomic Key 1:\n", final_taxonomic_key1)
print("Final Taxonomic Key 2:\n", final_taxonomic_key2)
print("Final Taxonomic Key 3:\n", final_taxonomic_key3)

NameError: name 'result1' is not defined