[Plugin] dsh-vision-primitives — precise interactive visual reasoning, native, no external MCP #1322
Replies: 3 comments
|
v1.1.0 — WebUI configuration card (https://github.com/zouyuanqing/dsh-vision-primitives/releases/tag/v1.1.0) The plugin now registers the |
|
v1.3.0 — Visual Evidence Protocol (https://github.com/zouyuanqing/dsh-vision-primitives/releases/tag/v1.3.0): pasted images (and the new |
|
v1.6.0 — send-time image bridge (https://github.com/zouyuanqing/dsh-vision-primitives/releases/tag/v1.6.0): text-only models can now paste/drag images into the composer (native thumbnails, no "model does not support images"); at send time the image is cached into the workspace and handed to the model as |
Uh oh!
There was an error while loading. Please reload this page.
dsh-vision-primitives — a native interactive visual-reasoning plugin for DeepSeek Harness, with zero external MCP servers.
Repo: https://github.com/zouyuanqing/dsh-vision-primitives · Install:
dsh plugin --profile <name> add github:zouyuanqing/dsh-vision-primitivesWhat it does
Gives text-only agents "precise eyes": a Set-of-Mark numbered grid overlay turns fuzzy visual perception into exact pixel coordinates (grid →
vision_resolve(cell)→ cell-center coordinates →vision_zoomlossless upscaling with a coordinate-mapping chain back to the original frame), then deterministic verification via annotate / measure / diff / color segmentation / native OCR.Highlights
vision_*tools: capture (multi-monitor screen or PNG), grid, resolve, zoom, annotate, measure, diff, find_color, ocr, describe, locate, state, resetvision_describe/vision_locate, plus a registeredmimoLlmAdapter (streaming / tool calls / image input)subprocessservice; no desktop controlDesign is a native re-implementation of my earlier vision-primitives-mcp.
Feedback, issues, and PRs welcome!
All reactions