## File search assistant
The McKessonMDTranscripts demonstrates the creation and execution of an assistant using the file_search feature with a knowledge base based on earnings call transcripts downloaded from the internet.

## First steps
- First, download AntRunLib from Nuget
- At least once, set up the environment using **[0-AI-settings](0-AI-settings.ipynb)**

In [1]:
#r "nuget: AntRunnerLib, 0.9.5"

using AntRunnerLib;
using AntRunnerLib.Identity;
using static AntRunnerLib.ClientUtility;
using System.IO;

#!import config/Settings.cs 

var envVariables = Settings.GetEnvironmentVariables();
foreach (var kvp in envVariables)
{
    Environment.SetEnvironmentVariable(kvp.Key, kvp.Value);
}

var config = AzureOpenAiConfigFactory.Get();
var client = GetOpenAiClient(config);


## Ensure the State_DOE_AI_Guides assistant exists
The definition of this assistant is located in the ".\AssistantDefinitions\McKessonMDTranscripts" folder

".\AssistantDefinitions\" is a default path. You can override it by setting the **ASSISTANTS_BASE_FOLDER_PATH** environment variable.

### Explanation
`AssistantUtility.Create` will create the vector store if necessary, upload the files, and then update the manifest with the vector store ID before creating the assistant.

In [2]:
var assistantId = await AssistantUtility.Create("State_DOE_AI_Guides", config);
Console.WriteLine(assistantId)

asst_ZZEIKcs9t1C6CybVLwfmyE72


## Run the Assistant

`output.Dialog` shows the conversation.

In [23]:
var assistantRunOptions = new AssistantRunOptions() {
    AssistantName = "State_DOE_AI_Guides",
    Instructions = "What does the California policy say about plagiarism detectors with regards to detecting AI created content. Do not use any policy except for California's when answering"
};
var output = await AntRunnerLib.AssistantRunner.RunThread(assistantRunOptions, config);
output.Dialog.DisplayAs("text/markdown");


User: What does the California policy say about plagiarism detectors with regards to detecting AI created content. Do not use any policy except for California's when answering

Assistant: I searched my knowledge base for the answer.

The California Department of Education (CDE) has specific guidelines regarding the use of plagiarism detection tools in relation to AI-created content. The CDE advises caution in using AI detectors due to their current limitations and potential for inaccuracy. Specifically, the guidance suggests that AI detectors, such as Turnitin's AI detection tool and GPT Zero, often incorrectly detect the use of AI, particularly when students use formal academic language. Therefore, while these tools can be used to inform feedback provided to students, the CDE discourages relying solely on these systems to determine instances of cheating or plagiarism【8:12†source】【8:13†source】【8:16†source】.


In [25]:
var assistantRunOptions = new AssistantRunOptions() {
    AssistantName = "State_DOE_AI_Guides",
    Instructions = "What specific AI detectors and vendors are mentioned in the policies and what do they say about each? A valid answer includes the names of specific **AI Detection** products, not general guidelines or chatbots like ChatGPT or Copilot"
};
var output = await AntRunnerLib.AssistantRunner.RunThread(assistantRunOptions, config);
output.Dialog.DisplayAs("text/markdown");


User: What specific AI detectors and vendors are mentioned in the policies and what do they say about each? A valid answer includes the names of specific **AI Detection** products, not general guidelines or chatbots like ChatGPT or Copilot

Assistant: I searched my knowledge base for the answer.

The state Department of Education guidelines mention several specific AI detection tools and vendors. Here are the products and what the guidelines say about each:

1. **AI Content Detector**
   - Described as an AI classifier that evaluates text to determine if it is AI-generated. It is noted that these tools do not accurately detect AI-generated material, especially material containing only a few hundred words. Additionally, they can falsely identify the writings of English Language students as generated material simply because their writing skills are still developing【4:0†source】.

2. **GPTZero**
   - Similar to AI Content Detector, GPTZero is mentioned as an AI classifier used to evaluate whether text is AI-generated. The guidelines caution that these tools are not very accurate in detecting AI-generated content and have similar issues with false positives【4:0†source】.

3. **Turnitin**
   - Known primarily as a plagiarism checker, Turnitin now features AI classifier detection. It scans for exact or close matches of keywords, phrases, and sentences to identify potential plagiarism, and it has been updated to include AI detection capabilities【4:0†source】【11:12†source】.

4. **Google’s Originality Reports**
   - Another plagiarism checker mentioned for its utility in reviewing text for plagiarism by scanning for exact or close matches of keywords, phrases, and sentences. It is useful for detecting unoriginal content, but there is no specific mention of its AI detection capabilities【4:0†source】.

These tools are recommended with caution, as the guidelines emphasize that AI detectors are not completely reliable and should not be used as the sole factor in determining if a student has used AI-generated content【4:1†source】【9:9†source】【11:11†source】.


In [4]:
output.LastMessage.DisplayAs("text/markdown");

The policies on the use of AI detectors by teachers and staff generally emphasize caution and highlight their unreliability. Here are the key points from the documents:

1. **Dependability Issues**: AI detectors are not considered dependable and should not be used as the sole factor in determining if a student has cheated. They have a high frequency of false positives, especially for non-native English speakers and creative writers, and false negatives for students skilled at using AI【4:0†source】【4:1†source】【4:2†source】.

2. **Educational Approach**: Suspicion of AI overuse should be seen as a teachable moment rather than a punitive one. The focus should be on educating students about the appropriate use of AI tools, similar to working with a tutor or in a collaborative group【4:0†source】【4:1†source】.

3. **Communication**: Educators should clearly communicate the appropriate uses of AI on each assignment. This may include referencing an AI Acceptable Use Scale to clarify the level of generative AI usage permitted【4:0†source】【4:1†source】.

4. **Critical Thinking and Bias**: Both teachers and students should evaluate AI outputs for accuracy and potential bias, and develop the critical thinking skills needed to navigate AI-generated content successfully【4:2†source】.

5. **Human Oversight**: AI technologies should not replace human decision-making in education. Instead, they should enhance learning experiences with proper human oversight and evaluation of AI-generated content【4:4†source】【4:14†source】.

These guidelines stress the importance of human judgment and ethical considerations in the use of AI tools in educational settings.

## Annotations

In [5]:
Console.WriteLine(JsonSerializer.Serialize(output.Annotations, new JsonSerializerOptions() { WriteIndented = true }));

[
  {
    "type": "file_citation",
    "text": "\u30104:0\u2020source\u3011",
    "start_index": 489,
    "end_index": 501,
    "file_citation": {
      "file_id": "assistant-F4vYCqSsKEjIycGkV5R7ae3n",
      "file_name": "",
      "quote": null
    }
  },
  {
    "type": "file_citation",
    "text": "\u30104:0\u2020source\u3011",
    "start_index": 489,
    "end_index": 501,
    "file_citation": {
      "file_id": "assistant-ApyhbGNmrcnrD5IpIxyAMnOr",
      "file_name": "",
      "quote": null
    }
  },
  {
    "type": "file_citation",
    "text": "\u30104:1\u2020source\u3011",
    "start_index": 501,
    "end_index": 513,
    "file_citation": {
      "file_id": "assistant-F4vYCqSsKEjIycGkV5R7ae3n",
      "file_name": "",
      "quote": null
    }
  },
  {
    "type": "file_citation",
    "text": "\u30104:1\u2020source\u3011",
    "start_index": 501,
    "end_index": 513,
    "file_citation": {
      "file_id": "assistant-ApyhbGNmrcnrD5IpIxyAMnO

### Formatting Output
The current version of the API leaves some things to be desired. First, because of chunking it can cite the same source twice. Second, 'quote' is supposed to contain a snippet, but currently, it is always null

In [26]:
var files = new Dictionary<string, string>();
var annotatedMessage = output.LastMessage;
foreach(var annotation in output.Annotations)
{
    // Only get the file name once per id
    if(!files.ContainsKey(annotation.FileCitation.FileId))
    {
        files[annotation.FileCitation.FileId] = await annotation.GetFileName(config);
    }
    // But the annotation itself may be a duplicate citation to the same file, so one must process them all
    // You can format the replacement however you like. In this case it just makes an md link
    annotatedMessage = annotatedMessage.Replace(annotation.Text, $" [{files[annotation.FileCitation.FileId]}]({files[annotation.FileCitation.FileId]})");
}

foreach(var file in files)
{
    Console.WriteLine($"{file.Key} = {file.Value}");
}
Console.WriteLine("-------------------");
annotatedMessage.DisplayAs("text/markdown");

assistant-mXASPcBcveVF7Zx7CtP2qyEP = State_DOE_AI_Guides_Mississippi_AI_guidance.pdf
assistant-ZY6Uf5tkIBXrf3ilRZ0Q5kHg = State_DOE_AI_Guides_Kentucky_AI_guidance.pdf
assistant-rMFVLDOfcoPyG4X3uKck4wFJ = State_DOE_AI_Guides_Louisiana_AI_guidance.pdf
assistant-F4vYCqSsKEjIycGkV5R7ae3n = State_DOE_AI_Guides_North_Carolina_AI_guidance.pdf
assistant-psRibNDvO3uqOTRYgqv9IW6v = State_DOE_AI_Guides_Washington_AI_guidance.pdf
-------------------


The state Department of Education guidelines mention several specific AI detection tools and vendors. Here are the products and what the guidelines say about each:

1. **AI Content Detector**
   - Described as an AI classifier that evaluates text to determine if it is AI-generated. It is noted that these tools do not accurately detect AI-generated material, especially material containing only a few hundred words. Additionally, they can falsely identify the writings of English Language students as generated material simply because their writing skills are still developing [State_DOE_AI_Guides_Mississippi_AI_guidance.pdf](State_DOE_AI_Guides_Mississippi_AI_guidance.pdf).

2. **GPTZero**
   - Similar to AI Content Detector, GPTZero is mentioned as an AI classifier used to evaluate whether text is AI-generated. The guidelines caution that these tools are not very accurate in detecting AI-generated content and have similar issues with false positives [State_DOE_AI_Guides_Mississippi_AI_guidance.pdf](State_DOE_AI_Guides_Mississippi_AI_guidance.pdf).

3. **Turnitin**
   - Known primarily as a plagiarism checker, Turnitin now features AI classifier detection. It scans for exact or close matches of keywords, phrases, and sentences to identify potential plagiarism, and it has been updated to include AI detection capabilities [State_DOE_AI_Guides_Mississippi_AI_guidance.pdf](State_DOE_AI_Guides_Mississippi_AI_guidance.pdf) [State_DOE_AI_Guides_Louisiana_AI_guidance.pdf](State_DOE_AI_Guides_Louisiana_AI_guidance.pdf).

4. **Google’s Originality Reports**
   - Another plagiarism checker mentioned for its utility in reviewing text for plagiarism by scanning for exact or close matches of keywords, phrases, and sentences. It is useful for detecting unoriginal content, but there is no specific mention of its AI detection capabilities [State_DOE_AI_Guides_Mississippi_AI_guidance.pdf](State_DOE_AI_Guides_Mississippi_AI_guidance.pdf).

These tools are recommended with caution, as the guidelines emphasize that AI detectors are not completely reliable and should not be used as the sole factor in determining if a student has used AI-generated content [State_DOE_AI_Guides_North_Carolina_AI_guidance.pdf](State_DOE_AI_Guides_North_Carolina_AI_guidance.pdf) [State_DOE_AI_Guides_Washington_AI_guidance.pdf](State_DOE_AI_Guides_Washington_AI_guidance.pdf) [State_DOE_AI_Guides_Louisiana_AI_guidance.pdf](State_DOE_AI_Guides_Louisiana_AI_guidance.pdf).

## Clean Up

In [7]:
var assistant = (await client.AssistantList())?.Data?.FirstOrDefault(o => o.Name == "McKessonMDTranscripts");
if(assistant != null) {
    await client.AssistantDelete(assistant.Id);
    Console.WriteLine("Deleted assistant");
}
else
{
    Console.WriteLine("Didn't find MsGraphUserProfile");
}

Deleted assistant
