# Adding Different LLM Modalities to Exisiting Agents

### This notebook showcases how to add image generation modality as a conversable agent capability 

First lets import all the required modules to run this example

In [15]:
import os
import re
from typing import Dict, Optional

import matplotlib.pyplot as plt
from PIL.Image import Image

import autogen
from autogen.agentchat import groupchat
from autogen.agentchat.contrib import img_utils
from autogen.agentchat.contrib.capabilities import generate_images
from autogen.oai import openai_utils

Lets define our config list with the models we want to experiment with. For this example, we wil be using Dalle image generator to interact with the GPT4V agent. 

In [11]:
OAI_CONFIG_LIST = [
    {"model": "dall-e-3", "api_key": os.environ.get("OAI_API_KEY")},
    {"model": "gpt-4-vision-preview", "api_key": os.environ.get("OAI_API_KEY")},
    {"model": "gpt-3.5-turbo", "api_key": os.environ.get("OAI_API_KEY")},
]

Lets define our LLM configs (you can experiment with different params)

In [12]:
def gpt_config() -> Dict:
    filtered_configs = openai_utils.filter_config(OAI_CONFIG_LIST, filter_dict={"model": ["gpt-3.5-turbo"]})

    return {"config_list": filtered_configs, "timeout": 120, "temperature": 0.7}


def gpt_v_config() -> Dict:
    filtered_configs = openai_utils.filter_config(OAI_CONFIG_LIST, filter_dict={"model": ["gpt-4-vision-preview"]})

    return {"config_list": filtered_configs, "timeout": 120, "temperature": 0.7, "max_tokens": 1000}


def dalle_config() -> Dict:
    filtered_configs = openai_utils.filter_config(OAI_CONFIG_LIST, filter_dict={"model": ["dall-e-3"]})

    return {"config_list": filtered_configs, "timeout": 120, "temperature": 0.7}

Our system will consist of 3 main agents:
1. User proxy
2. Image generator agent
3. Critic

We will use the user proxy to define the initial prompt. The image generator agent will have to generate the images based on the prompts given to it. The critic is there to improve the output of the image generator.

The image generator agent and the critic will communicate with each other through a groupchat, acting like an art studio.

In [17]:
CRITIC_SYSTEM_MESSAGE = """You need to improve the prompt of the figures you saw.
How to create a figure that is better in terms of color, shape, text (clarity), and other things.
Reply with the following format:

CRITICS: the image needs to improve...
PROMPT: here is the updated prompt!
"""

In [22]:
def image_generator_agent() -> autogen.ConversableAgent:
    agent = autogen.ConversableAgent(name="dalle", llm_config=gpt_config())
    dalle_gen = generate_images.DalleImageGenerator(llm_config=dalle_config(), cache_settings={"directory": ".cache/"})
    image_gen_capability = generate_images.ImageGeneration(image_generator=dalle_gen)

    image_gen_capability.add_to_agent(agent)
    return agent

def critic_agent() -> autogen.ConversableAgent:
    return autogen.ConversableAgent(name="critic", llm_config=gpt_v_config(), system_message=CRITIC_SYSTEM_MESSAGE)

def art_studio() -> groupchat.GroupChatManager:
    gc = groupchat.GroupChat(
        admin_name=None,
        agents=[image_generator_agent(), critic_agent()],
        messages=[],
        speaker_selection_method="round_robin",
        max_round=4,
    )
    return groupchat.GroupChatManager(groupchat=gc, llm_config=False)

We'll define `extract_img` to help us extract the image generated by the image generator agent.

In [23]:
def extract_img(sender: autogen.Agent, recipient: autogen.Agent) -> Optional[Image]:
    # From notebook/agentchat_dalle_and_gpt4v.ipynb
    """
    Extracts an image from the last message of an agent and converts it to a PIL image.

    This function searches the last message sent by the given agent for an image tag,
    extracts the image data, and then converts this data into a PIL (Python Imaging Library) image object.

    Parameters:
        agent (Agent): An instance of an agent from which the last message will be retrieved.

    Returns:
        PIL.Image: A PIL image object created from the extracted image data.

    Note:
    - The function assumes that the last message contains an <img> tag with image data.
    - The image data is extracted using a regular expression that searches for <img> tags.
    - It's important that the agent's last message contains properly formatted image data for successful extraction.
    - The `_to_pil` function is used to convert the extracted image data into a PIL image.
    - If no <img> tag is found, or if the image data is not correctly formatted, the function may raise an error.
    """
    last_message = recipient.last_message(sender)["content"]
    img_data = None

    if isinstance(last_message, str):
        img_data = re.findall("<img (.*)>", last_message)
        if img_data:
            img_data = img_data[0]
    elif isinstance(last_message, list):
        # The GPT-4V format, where the content is an array of data
        assert isinstance(last_message[0], dict)
        img_data = last_message[0].get("image_url", {}).get("url")

    if img_data:
        return img_utils._to_pil(img_data)
    else:
        return None

In [24]:
def main():
    user = autogen.UserProxyAgent(name="user", human_input_mode="NEVER", max_consecutive_auto_reply=0)
    manager = art_studio()

    user.initiate_chat(
        manager,
        message="""Create an image with black background, a happy robot is showing a sign with "I Love AutoGen".""",
    )

    dalle = manager.groupchat.agent_by_name("dalle")

    assert dalle is not None
    image = extract_img(dalle, manager)
    if image is not None:
        plt.imshow(image)
        plt.axis("off")  # Turn off axis numbers
        plt.show()

In [21]:
main()

[33muser[0m (to chat_manager):

Create an image with black background, a happy robot is showing a sign with "I Love AutoGen".

--------------------------------------------------------------------------------
[31m
>>>>>>>> USING AUTO REPLY...[0m


  super().__init__(


[33mdalle[0m (to chat_manager):

Generated an image with the prompt: To create the image described in the text, you would prompt the image generator with the following specifications:
- Set the background color to black.
- Include a happy robot character.
- Instruct the robot to hold a sign with the text "I Love AutoGen" displayed on it.<image>

--------------------------------------------------------------------------------
[31m
>>>>>>>> USING AUTO REPLY...[0m
[33mcritic[0m (to chat_manager):

CRITICS: the image needs to improve on the clarity of the text on the sign, the robot's expression to clearly convey happiness, and the color contrast to make the image more visually appealing.

PROMPT: here is the updated prompt!
- Choose a vibrant color for the robot to stand out against the black background.
- Ensure the robot has a distinctly joyful expression, with features like a smile or 'sparkling' eyes.
- The sign should have a contrasting color (e.g., white or bright color) to ma

TypeError: Object of type PngImageFile is not JSON serializable