Repository navigation
Support multimodal image content from defineTool() results #565
Minoo7
started this conversation in
Feature Request
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
Allow a
defineTool()tool to return model-facing image content to the parent agent, alongside text and structured JSON output.Flue already accepts images on user messages and on
session.prompt()/session.task()/session.skill(). Pi also supports image blocks in tool results. The missing path is an image produced or loaded by a tool during an agent turn: current custom-tool output is JSON-serialized into one text block, so the parent model never receives the pixels.Background & Motivation
I'm building an agent that works with PDFs, plans, photographs, screenshots, and other files in its sandbox. I do not want to attach every image up front. The useful pattern is on-demand inspection:
Today step 4 is not expressible through Flue's public tool API.
ToolRunEnvelope.outputis JSON-serializable, andcreateCustomTools()turns it into a Pi text block withJSON.stringify(output). Returning an object shaped like{ type: "image", data, mimeType }therefore gives the model JSON text (and a very large base64 string), not an image.MCP tools show the same underlying gap from another direction:
formatMcpResult()currently renders image items as placeholders such as[Image: image/png, … base64 chars]. This was raised for MCP in #291, but I think the useful primitive is broader than an MCP-specific fix: any Flue tool should be able to return image content to its caller.Why the existing options are not equivalent
A tool can make a nested vision call:
That is useful when the tool itself owns a fixed analysis task, but it is not equivalent to returning the image:
Putting base64 in
outputis also not a workaround: it is serialized as text, wastes context and persistence budget, and still does not become a vision input.Use cases this would unblock
This matters because images are often discovered during a task. User-message attachments solve initial input, but not agent-driven retrieval after the conversation has started.
Goals
defineTool()result include model-facing text and image blocks that are delivered to the parent model as one tool result.outputavailable separately for application/UI data.Example
The exact API is a design choice, but a separate model-facing
contentfield would preserve the current meaning ofoutput:Conceptually,
contentwould map to Pi's(TextContent | ImageContent)[]and Flue's canonical model-facing tool-result content, whileoutputwould retain its existing structured-data role.Non-goals
harness.prompt()call merely to expose an image to the parent agent.Related work
prompt()/skill()/task()inputs.Flue already appears to have most of the internal representation needed: Pi tool results can contain image blocks, and Flue persists tool-result images as attachment references. The request is to expose that capability through the public tool authoring path without weakening the durability and budget constraints around it.
All reactions