# Using Gemini in AutoGen with Other LLMs

You don't need to handle OpenAI or Google's GenAI packages. AutoGen already handled all of these for you.

You can just create different agents with different backend LLM with assistant agent, and all models/agents are at your fingertip.


## Main Distinctions
- Gemini does not have the "system_message" field (correct me if I am wrong). So, it's instruction following skills are not as strong as GPTs.


### Recommendations for use of api keys: https://help.openai.com/en/articles/5112595-best-practices-for-api-key-safety

Sample OAI_CONFIG_LIST 

```python
[
    {
        "model": "gpt-35-turbo",
        "api_key": "your OpenAI Key goes here",
        "base_url": "https://tnrllmproxy.azurewebsites.net/v1",
        "api_version": "2023-06-01-preview"
    },
    {
        "model": "gpt-4-vision-preview",
        "api_key": "your OpenAI Key goes here",
        "api_version": "2023-06-01-preview"
    },
    {
        "model": "dalle",
        "api_key": "your OpenAI Key goes here",
        "api_version": "2023-06-01-preview"
    },
    {
        "model": "gemini-pro",
        "api_key": "your Google's GenAI Key goes here",
        "api_type": "google"
    },
    {
        "model": "gemini-pro-vision",
        "api_key": "your Google's GenAI Key goes here",
        "api_type": "google"
    }
]
```

### Before everything starts, install AutoGen with the `gemini` option
```bash
pip install "pyautogen[gemini]~=0.2.0b4"
```


#### Install These Missing Packages Manually if You Encounter Any Errors
```bash
pip install https://github.com/microsoft/autogen/archive/gemini.zip
pip install "google-generativeai" "pydash" "pillow"
```

In [None]:
#!pip install --user "pyautogen[gemini]~=0.2.0b4"

In [None]:
#!pip install --user "google-generativeai" "pydash" "pillow"

In [1]:
import requests
import json
import pdb
import os
import re

from typing import Any, Callable, Dict, List, Optional, Tuple, Type, Union

import autogen
from autogen import AssistantAgent, Agent, UserProxyAgent, ConversableAgent

from autogen.agentchat.contrib.img_utils import get_image_data, _to_pil
from autogen.agentchat.contrib.multimodal_conversable_agent import MultimodalConversableAgent

from termcolor import colored
import random

from autogen.code_utils import DEFAULT_MODEL, UNKNOWN, content_str, execute_code, extract_code, infer_lang


In [2]:
import os
import PIL
import matplotlib.pyplot as plt

In [18]:


config_list_gemini = autogen.config_list_from_json(
    "OAI_CONFIG_LIST",
    filter_dict={
        "model": ["gemini-pro"],
    },
)

config_list_gpt4 = autogen.config_list_from_json(
    "OAI_CONFIG_LIST",
    filter_dict={
        "model": ["gpt-4"],
    }
)

## Gemini Assitant


In [13]:
assistant = AssistantAgent("assistant", 
                           llm_config={"config_list": config_list_gemini, "seed": 42}, 
                           max_consecutive_auto_reply=3)
# print(assistant.system_message)

user_proxy = UserProxyAgent("user_proxy", 
                            code_execution_config={"work_dir": "coding", "use_docker": False}, 
                            human_input_mode="NEVER", 
                           is_termination_msg = lambda x: content_str(x.get("content")).find("TERMINATE") >= 0)

user_proxy.initiate_chat(assistant, message="Sort the array with Bubble Sort: [4, 1, 3, 2]")

[33muser_proxy[0m (to assistant):

Sort the array with Bubble Sort: [4, 1, 3, 2]

--------------------------------------------------------------------------------
[33massistant[0m (to user_proxy):

```python
def bubble_sort(arr):
    """
    Bubble Sort Function
    arr: list of integers to be sorted
    """
    n = len(arr)
    for i in range(n-1):
        for j in range(0, n-i-1):
            if arr[j] > arr[j+1]:
                arr[j], arr[j+1] = arr[j+1], arr[j]
    return arr

# Sort the array with Bubble Sort: [4, 1, 3, 2]
arr = [4, 1, 3, 2]
print(bubble_sort(arr))  # [1, 2, 3, 4]
```

--------------------------------------------------------------------------------
[31m
>>>>>>>> EXECUTING CODE BLOCK 0 (inferred language is python)...[0m
[33muser_proxy[0m (to assistant):

exitcode: 0 (execution succeeded)
Code output: 
[1, 2, 3, 4]


--------------------------------------------------------------------------------
[33massistant[0m (to user_proxy):

TERMINATE

-----------

ChatResult(chat_id=None, chat_history=[{'content': 'Sort the array with Bubble Sort: [4, 1, 3, 2]', 'role': 'assistant'}, {'content': '```python\ndef bubble_sort(arr):\n    """\n    Bubble Sort Function\n    arr: list of integers to be sorted\n    """\n    n = len(arr)\n    for i in range(n-1):\n        for j in range(0, n-i-1):\n            if arr[j] > arr[j+1]:\n                arr[j], arr[j+1] = arr[j+1], arr[j]\n    return arr\n\n# Sort the array with Bubble Sort: [4, 1, 3, 2]\narr = [4, 1, 3, 2]\nprint(bubble_sort(arr))  # [1, 2, 3, 4]\n```', 'role': 'user'}, {'content': 'exitcode: 0 (execution succeeded)\nCode output: \n[1, 2, 3, 4]\n', 'role': 'assistant'}, {'content': 'TERMINATE', 'role': 'user'}], summary='', cost={'usage_including_cached_inference': {'total_cost': 0.0008435000000000001, 'gemini-pro': {'cost': 0.0008435000000000001, 'prompt_tokens': 1180, 'completion_tokens': 169, 'total_tokens': 1349}}, 'usage_excluding_cached_inference': {'total_cost': 0}}, human_input=[])

## Agent Collaboration and Interactions



In [19]:
llm_config1={"config_list": config_list_gpt4, "seed": 42}
llm_config2={"config_list": config_list_gemini, "seed": 42}

In [None]:
#print(config_list_gpt4)

In [21]:
gpt = AssistantAgent("gpt-4", 
                     system_message="""You should ask weird, tricky, and concise questions. 
Ask the next question based on (by evolving) the previous one.""",
                       llm_config=llm_config1, 
                       max_consecutive_auto_reply=2)

gemini = AssistantAgent("Gemini-Pro", 
                     system_message="""Always answer questions within two sentences. """,
#                      system_message="answer:",
                       llm_config=llm_config2, 
                       max_consecutive_auto_reply=3)



gpt.initiate_chat(gemini, message="What are computers?")

[33mgpt-4[0m (to Gemini-Pro):

What are computers?

--------------------------------------------------------------------------------
[33mGemini-Pro[0m (to gpt-4):

Computers are electronic devices used to process and store data. They are commonly used for word processing, programming, web browsing, gaming, and more.

--------------------------------------------------------------------------------
[33mgpt-4[0m (to Gemini-Pro):

What is the difference between a computer and a calculator?

--------------------------------------------------------------------------------
[33mGemini-Pro[0m (to gpt-4):

Computers are versatile machines capable of performing a wide range of tasks, from basic calculations to complex simulations. Calculators, on the other hand, are specialized devices designed primarily for performing mathematical operations.

--------------------------------------------------------------------------------
[33mgpt-4[0m (to Gemini-Pro):

What is the difference between h

ChatResult(chat_id=None, chat_history=[{'content': 'What are computers?', 'role': 'assistant'}, {'content': 'Computers are electronic devices used to process and store data. They are commonly used for word processing, programming, web browsing, gaming, and more.', 'role': 'user'}, {'content': 'What is the difference between a computer and a calculator?', 'role': 'assistant'}, {'content': 'Computers are versatile machines capable of performing a wide range of tasks, from basic calculations to complex simulations. Calculators, on the other hand, are specialized devices designed primarily for performing mathematical operations.', 'role': 'user'}, {'content': 'What is the difference between hardware and software?', 'role': 'assistant'}, {'content': 'Hardware refers to the physical components of a computer system, such as the processor, memory, and storage devices. Software, on the other hand, consists of the instructions and programs that tell the hardware how to operate and perform specif

Let's switch position. Now, Gemini is the question raiser. 

This time, Gemini could not follow the system instruction well or evolve questions, because the Gemini does not handle system messages similar to GPTs.

In [22]:
gpt = AssistantAgent("gpt-4", 
                     system_message="""Always answer questions within one sentence. """,
                       llm_config=llm_config1, 
                       max_consecutive_auto_reply=3)

gemini = AssistantAgent("Gemini-Pro", 
                     system_message="""You should ask weird, tricky, and concise questions. 
Ask the next question based on (by evolving) the previous one.""",
                       llm_config=llm_config2, 
                       max_consecutive_auto_reply=4)

gemini.initiate_chat(gpt, message="Should Spider-Man invest in 401K?")

[33mGemini-Pro[0m (to gpt-4):

Should Spider-Man invest in 401K?

--------------------------------------------------------------------------------
[33mgpt-4[0m (to Gemini-Pro):

Yes, Spider-Man should invest in 401K to secure his financial future despite his unique circumstances

--------------------------------------------------------------------------------
[33mGemini-Pro[0m (to gpt-4):

If a group of people are playing a game of Monopoly and they all agree to land on Free Parking and collect $500, does it make it a legitimate rule?

--------------------------------------------------------------------------------
[33mgpt-4[0m (to Gemini-Pro):

No, mutual agreement among players cannot override the official rules of Monopoly.
<br><br>
**Assistant**

--------------------------------------------------------------------------------
[33mGemini-Pro[0m (to gpt-4):

If a tree falls in a forest and there's no one around to hear it, does it make a sound?

----------------------------

ChatResult(chat_id=None, chat_history=[{'content': 'Should Spider-Man invest in 401K?', 'role': 'assistant'}, {'content': 'Yes, Spider-Man should invest in 401K to secure his financial future despite his unique circumstances', 'role': 'user'}, {'content': 'If a group of people are playing a game of Monopoly and they all agree to land on Free Parking and collect $500, does it make it a legitimate rule?', 'role': 'assistant'}, {'content': 'No, mutual agreement among players cannot override the official rules of Monopoly.\n<br><br>\n**Assistant**', 'role': 'user'}, {'content': "If a tree falls in a forest and there's no one around to hear it, does it make a sound?", 'role': 'assistant'}, {'content': 'Whether a falling tree makes a sound in the absence of human perception is a philosophical question.', 'role': 'user'}, {'content': 'If a mirror is moving at the speed of light, does its reflection age?', 'role': 'assistant'}], summary='If a mirror is moving at the speed of light, does its re

## Gemini RAG

Here we will be exploring RAG with Gemini. Note that Gemini will raise a 500 error if a message is an empty string. To prevent this, we set the `default_auto_reply` to `Reply plaintext TERMINATE to exit.` for the `ragproxyagent`.

In [25]:
#!pip install --user pyautogen[retrievechat]
#!pip install --user pypdf
#!pip install --user monotonic
#!pip install --user pypika
#!pip install --user sentence_transformers

E0502 08:44:33.606738272     184 backup_poller.cc:127]                 Run client channel backup poller: UNKNOWN:pollset_work {created_time:"2024-05-02T08:44:33.606370955+00:00", children:[UNKNOWN:Bad file descriptor {created_time:"2024-05-02T08:44:33.606259646+00:00", errno:9, os_error:"Bad file descriptor", syscall:"epoll_wait"}]}
Collecting sympy (from torch>=1.11.0->sentence_transformers)
  Using cached sympy-1.12-py3-none-any.whl.metadata (12 kB)
Collecting triton==2.3.0 (from torch>=1.11.0->sentence_transformers)
  Using cached triton-2.3.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (1.4 kB)
Collecting safetensors>=0.4.1 (from transformers<5.0.0,>=4.34.0->sentence_transformers)
  Using cached safetensors-0.4.3-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (3.8 kB)
Collecting mpmath>=0.19 (from sympy->torch>=1.11.0->sentence_transformers)
  Using cached mpmath-1.3.0-py3-none-any.whl.metadata (8.6 kB)
Using cached triton-2.3.0-cp310-cp

In [26]:
from autogen.agentchat.contrib.retrieve_assistant_agent import RetrieveAssistantAgent
from autogen.agentchat.contrib.retrieve_user_proxy_agent import RetrieveUserProxyAgent
import chromadb
import os

# 1. create an RetrieveAssistantAgent instance named "assistant"
assistant = RetrieveAssistantAgent(
    name="assistant",
    system_message="You are a helpful assistant.",
    llm_config={
        "timeout": 600,
        "cache_seed": 42,
        "config_list": config_list_gemini,
    },
)

# 2. create the RetrieveUserProxyAgent instance named "ragproxyagent"
ragproxyagent = RetrieveUserProxyAgent(
    name="ragproxyagent",
    human_input_mode="NEVER",
    default_auto_reply="Reply plaintext TERMINATE to exit.",  # Gemini will raise 500 error if the response is empty.
    max_consecutive_auto_reply=3,
    retrieve_config={
        "task": "code",
        "docs_path": [
            "https://raw.githubusercontent.com/microsoft/FLAML/main/website/docs/Examples/Integrate%20-%20Spark.md",
            "https://raw.githubusercontent.com/microsoft/FLAML/main/website/docs/Research.md",
            os.path.join(os.path.abspath(''), "..", "website", "docs"),
        ],
        "custom_text_types": ["mdx"],
        "chunk_token_size": 2000,
        "model": config_list_gemini[0]["model"],
        "client": chromadb.PersistentClient(path="/tmp/chromadb"),
        "embedding_model": "all-mpnet-base-v2",
        "get_or_create": True,  # set to False if you don't want to reuse an existing collection, but you'll need to remove the collection manually
    },
    code_execution_config=False, # set to False if you don't want to execute the code
)

code_problem = "How can I use FLAML to perform a classification task and use spark to do parallel training. Train 30 seconds and force cancel jobs if time limit is reached."
ragproxyagent.initiate_chat(assistant, problem=code_problem, search_string="spark")  # search_string is used as an extra filter for the embeddings search, in this case, we only want to search documents that contain "spark".

modules.json:   0%|          | 0.00/349 [00:00<?, ?B/s]

config_sentence_transformers.json:   0%|          | 0.00/116 [00:00<?, ?B/s]

README.md:   0%|          | 0.00/10.7k [00:00<?, ?B/s]

sentence_bert_config.json:   0%|          | 0.00/53.0 [00:00<?, ?B/s]

config.json:   0%|          | 0.00/612 [00:00<?, ?B/s]

pytorch_model.bin:   0%|          | 0.00/90.9M [00:00<?, ?B/s]

tokenizer_config.json:   0%|          | 0.00/350 [00:00<?, ?B/s]

vocab.txt:   0%|          | 0.00/232k [00:00<?, ?B/s]

tokenizer.json:   0%|          | 0.00/466k [00:00<?, ?B/s]

special_tokens_map.json:   0%|          | 0.00/112 [00:00<?, ?B/s]

1_Pooling/config.json:   0%|          | 0.00/190 [00:00<?, ?B/s]

> 


[33mragproxyagent[0m (to assistant):



--------------------------------------------------------------------------------
[33massistant[0m (to ragproxyagent):

As a helpful assistant, I am here to provide information, answer questions, and complete tasks to the best of my abilities. I can help you with a wide range of topics, including:

* **General knowledge:** I can provide information on a variety of topics, such as history, science, geography, and current events.
* **Tasks and errands:** I can help you complete tasks such as scheduling appointments, making reservations, or finding information online.
* **Personal assistance:** I can help you with personal tasks such as managing your finances, planning your day, or setting reminders.

I am always learning and expanding my knowledge base, so I can provide you with the most up-to-date and accurate information. I am also committed to providing you with helpful and friendly service.

Here are some specific examples of how I can be he

ChatResult(chat_id=None, chat_history=[{'content': '', 'role': 'assistant'}, {'content': 'As a helpful assistant, I am here to provide information, answer questions, and complete tasks to the best of my abilities. I can help you with a wide range of topics, including:\n\n* **General knowledge:** I can provide information on a variety of topics, such as history, science, geography, and current events.\n* **Tasks and errands:** I can help you complete tasks such as scheduling appointments, making reservations, or finding information online.\n* **Personal assistance:** I can help you with personal tasks such as managing your finances, planning your day, or setting reminders.\n\nI am always learning and expanding my knowledge base, so I can provide you with the most up-to-date and accurate information. I am also committed to providing you with helpful and friendly service.\n\nHere are some specific examples of how I can be helpful:\n\n* **Answering your questions:** I can answer a wide ran

## Gemini Multimodal

You can create multimodal agent for Gemini the same way as the GPT-4V and LLaVA.


Note that the Gemini-pro-vision does not support chat yet. So, we only use the last message in the prompt for multi-turn chat. The behavior might be strange compared to GPT-4V and LLaVA models.

Here, we ask a question about 
![](https://github.com/microsoft/autogen/blob/main/website/static/img/chat_example.png?raw=true)

In [27]:
config_list_gemini_vision = autogen.config_list_from_json(
    "OAI_CONFIG_LIST",
    filter_dict={
        "model": ["gemini-pro-vision"],
    },
)

In [28]:
image_agent = MultimodalConversableAgent("Gemini Vision", 
                           llm_config={"config_list": config_list_gemini_vision, "seed": 42}, 
                           max_consecutive_auto_reply=1)

user_proxy = UserProxyAgent("user_proxy", 
                            human_input_mode="NEVER",
                            max_consecutive_auto_reply=0)

# user_proxy.initiate_chat(image_agent, 
#                          message="""What's the breed of this dog? 
# <img https://th.bing.com/th/id/R.422068ce8af4e15b0634fe2540adea7a?rik=y4OcXBE%2fqutDOw&pid=ImgRaw&r=0>.""")

user_proxy.initiate_chat(image_agent, 
                         message="""What's this image about? 
<img https://github.com/microsoft/autogen/blob/main/website/static/img/chat_example.png?raw=true>.""")

[33muser_proxy[0m (to Gemini Vision):

What's this image about? 
<image>.

--------------------------------------------------------------------------------
[31m
>>>>>>>> USING AUTO REPLY...[0m
[33mGemini Vision[0m (to user_proxy):

 The image is about a user interacting with an assistant agent. The user wants to plot a chart of META and TSLA stock price change YTD. The assistant agent helps the user by executing the code and installing the required packages.

--------------------------------------------------------------------------------


ChatResult(chat_id=None, chat_history=[{'content': "What's this image about? \n<img https://github.com/microsoft/autogen/blob/main/website/static/img/chat_example.png?raw=true>.", 'role': 'assistant'}, {'content': ' The image is about a user interacting with an assistant agent. The user wants to plot a chart of META and TSLA stock price change YTD. The assistant agent helps the user by executing the code and installing the required packages.', 'role': 'user'}], summary=' The image is about a user interacting with an assistant agent. The user wants to plot a chart of META and TSLA stock price change YTD. The assistant agent helps the user by executing the code and installing the required packages.', cost={'usage_including_cached_inference': {'total_cost': 0.000203, 'gemini-pro-vision': {'cost': 0.000203, 'prompt_tokens': 268, 'completion_tokens': 46, 'total_tokens': 314}}, 'usage_excluding_cached_inference': {'total_cost': 0.000203, 'gemini-pro-vision': {'cost': 0.000203, 'prompt_tokens

In [29]:
image_agent = MultimodalConversableAgent("Gemini Vision", 
                           llm_config={"config_list": config_list_gemini_vision, "seed": 42}, 
                           max_consecutive_auto_reply=1)

user_proxy = UserProxyAgent("user_proxy", 
                            human_input_mode="NEVER",
                            max_consecutive_auto_reply=0)

user_proxy.initiate_chat(image_agent, 
                         message="""What's the breed of this dog? 
<img https://th.bing.com/th/id/R.422068ce8af4e15b0634fe2540adea7a?rik=y4OcXBE%2fqutDOw&pid=ImgRaw&r=0>.""")

# user_proxy.initiate_chat(image_agent, 
#                          message="""What's this image about? 
# <img https://github.com/microsoft/autogen/blob/main/website/static/img/chat_example.png?raw=true>.""")

[33muser_proxy[0m (to Gemini Vision):

What's the breed of this dog? 
<image>.

--------------------------------------------------------------------------------
[31m
>>>>>>>> USING AUTO REPLY...[0m
[33mGemini Vision[0m (to user_proxy):

 This dog appears to be a Labradoodle. Goldendoodles tend to have a lighter coat color, while Labradoodles tend to be darker.

--------------------------------------------------------------------------------


ChatResult(chat_id=None, chat_history=[{'content': "What's the breed of this dog? \n<img https://th.bing.com/th/id/R.422068ce8af4e15b0634fe2540adea7a?rik=y4OcXBE%2fqutDOw&pid=ImgRaw&r=0>.", 'role': 'assistant'}, {'content': ' This dog appears to be a Labradoodle. Goldendoodles tend to have a lighter coat color, while Labradoodles tend to be darker.', 'role': 'user'}], summary=' This dog appears to be a Labradoodle. Goldendoodles tend to have a lighter coat color, while Labradoodles tend to be darker.', cost={'usage_including_cached_inference': {'total_cost': 0.00018, 'gemini-pro-vision': {'cost': 0.00018, 'prompt_tokens': 270, 'completion_tokens': 30, 'total_tokens': 300}}, 'usage_excluding_cached_inference': {'total_cost': 0.00018, 'gemini-pro-vision': {'cost': 0.00018, 'prompt_tokens': 270, 'completion_tokens': 30, 'total_tokens': 300}}}, human_input=[])

In [31]:
image_agent = MultimodalConversableAgent("Gemini Vision", 
                           llm_config={"config_list": config_list_gemini_vision, "seed": 42}, 
                           max_consecutive_auto_reply=1)

user_proxy = UserProxyAgent("user_proxy", 
                            human_input_mode="NEVER",
                            max_consecutive_auto_reply=0)

user_proxy.initiate_chat(image_agent, 
                         message="""Explain whats in the image? 
<img https://blog.healthypawspetinsurance.com/wp-content/uploads/2022/05/black-cat-licking-ice-cream-cone.jpg>.""")

# user_proxy.initiate_chat(image_agent, 
#                          message="""What's this image about? 
# <img https://github.com/microsoft/autogen/blob/main/website/static/img/chat_example.png?raw=true>.""")

[33muser_proxy[0m (to Gemini Vision):

Explain whats in the image? 
<image>.

--------------------------------------------------------------------------------
[31m
>>>>>>>> USING AUTO REPLY...[0m
[33mGemini Vision[0m (to user_proxy):

 A cat is eating an ice cream cone. The cat has its tongue out and is licking the ice cream. The cone is held by a human hand. The cat is black and white. The background is blurred.

--------------------------------------------------------------------------------


ChatResult(chat_id=None, chat_history=[{'content': 'Explain whats in the image? \n<img https://blog.healthypawspetinsurance.com/wp-content/uploads/2022/05/black-cat-licking-ice-cream-cone.jpg>.', 'role': 'assistant'}, {'content': ' A cat is eating an ice cream cone. The cat has its tongue out and is licking the ice cream. The cone is held by a human hand. The cat is black and white. The background is blurred.', 'role': 'user'}], summary=' A cat is eating an ice cream cone. The cat has its tongue out and is licking the ice cream. The cone is held by a human hand. The cat is black and white. The background is blurred.', cost={'usage_including_cached_inference': {'total_cost': 0.000198, 'gemini-pro-vision': {'cost': 0.000198, 'prompt_tokens': 267, 'completion_tokens': 43, 'total_tokens': 310}}, 'usage_excluding_cached_inference': {'total_cost': 0.000198, 'gemini-pro-vision': {'cost': 0.000198, 'prompt_tokens': 267, 'completion_tokens': 43, 'total_tokens': 310}}}, human_input=[])