dsh-mindseye
Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, GUI automation, through layered evidence memory and cache.
为 DSH 上的纯文本模型提供视觉插件:通过分层证据记忆与缓存,实现图片理解与生成的原生交互及 GUI 自动化。
How to install
dsh plugin add dsh-mindseye About
MindsEye 让 DeepSeek 原生看图 —— model-driven vision tools for DeepSeek Harness MindsEye 是一个 DeepSeek Harness(dsh)vision 插件。粘贴图片后,图片原样显示在会话里,DeepSeek 继续负责思考,视觉模型负责看图。插件暴露一组按任务拆分的视觉工具,由模型根据用户意图选择工具,每个工具固定映射到对应的意图和模型路由,返回结构化 JSON,并通过缓存与证据复用减少重复开销。 核心体验 **粘贴即看图**:接管 deepseek-official 路由,图片原生进入会话;接管不可用时自动降级为路径粘贴,新图始终能发出去 **模型选工具,插件管模型**:mindseye_read_image、mindseye_ocr、mindseye_ground、mindseye_colors 各自固定意图,模型按用户问题选工具,插件按工具映射到对应的模型链 **图片轮自动挂载**:检测到图片消息时自动注册视觉工具;纯文本轮默认只保留一个激活入口,避免常驻占用模型上下文 **多图一次读**:批量读取多张图片,批量遇 4xx 按指数拆分降级,失败只影响单张 **旧会话不毒化**:历史带图会话在回退模式下也能正常对话,图片块自动替换为附件标记 **每次调用透明**:返回 provider、model、…
Recommendation signals
Meta
- License
- MIT
- Language
- TypeScript
- GitHub stars
- 1
- mo. downloads
- –
- Last push
- 2026-08-29
- Created
- 2026-08-16
Basic safety check
- Findings
- None
- Sources
- curated:awesome-dsh-plugin.com, curated:awesome-dsh-plugin/awesome-dsh-plugin
- Topics
- agent, deepseek-harness, dsh-plugin, gui-automation, memory, multimodal, text-only-llm, vision
Related plugins
modlens
liustack/modlens
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
dsh-web
zhu1090093659/dsh-web-ui/tree/main/packages/dsh-tool-describe-image
A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.
dsh-vision-router
ysr666/dsh-vision-router
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
dsh-vision-toolkit
Anionex/dsh-vision-toolkit
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.