Feature idea: for text-only models, route images to a vision model automatically and feed the description to the base model #3512
Replies: 4 comments
|
+1 on auto vision routing for text-only chat models. A lot of users hit “current model does not support images” and stop, when the desired UX is: keep the conversation model, run a small vision pass, inject a short caption/OCR into the next turn. Community plugins already explore pieces of this (vision companions / Pi bridges); a host-owned route would be more reliable than every plugin reinventing it. If it ships, please make the vision model + prompt template configurable per profile so cost/latency stay under user control. |
|
@ylwl1997 mentioned "Pi bridges" — that's the thing I maintain, so let me put the prior art on the table for whoever ends up designing the native version. Not a pitch: you already built your own plugin, and if this lands upstream both of ours retire. pi2dsh has shipped this shape since 0.9.0, and the design detail worth stealing is how the route gets created, not the vision call itself. Rather than intercepting at the request layer, the engine walks the model directory and registers, for every text-only route, a companion route named
It is automatic with zero configuration; Boundaries, stated plainly:
Disclosure: pi2dsh is mine. |
|
@weijiafu14 Thank you so much for sharing the prior art and walking through the companion-route design — especially the hot-follow mechanism on the model directory. That's really valuable context for whoever ends up building the native version. One thing worth noting: DeepSeek now supports multimodal input as well, and I believe multimodal base models are the inevitable trend. That's great news for all of us users — it removes a whole category of workarounds and lets the conversation model see images directly. Hope we can keep exchanging ideas in the future. Thanks again! 🙌 |
|
I turned this proposal into a source-backed English operator guide for DeepSeek Harness: https://github.com/sandbaseai/deepseek-harness-handbook/blob/main/docs/en/integrations/text-model-vision-fallback.md — it keeps the selected text Agent, treats the vision result as untrusted derived evidence, and covers privacy, caching, cost, and acceptance tests. Feedback welcome. |
Uh oh!
There was an error while loading. Please reload this page.
Small feature suggestion: right now, if the chat model doesn't accept images (e.g. DeepSeek), sending an image is rejected outright with "the current model does not support images". But a lot of the time we just want to use DeepSeek for the task and have it look at an image as part of it.
What I'd like: without switching the base model, when the user attaches an image, the harness automatically sends it to a configurable vision model, feeds the resulting text description back to the current model, and lets the current model answer. The UI still shows the original image — seamless, same feel as native multimodal.
The building blocks all seem to be there: attachment storage (content-addressed sha256), model modality detection (
inputModalities), thellm/streamwaterfall interception point, and the server-sidehost.*native RPCs. The general shape is: admit images at the entry layer + rewrite image blocks into description text at the request layer.I wrote an unofficial plugin to validate this approach (dsh-vision-fallback): attach image → auto-call a vision model (default: Doubao) → description fed to DeepSeek → normal answer. Feel free to use it as a reference if you plan to build this natively; the plugin retires itself once it lands upstream.
All reactions