dsh-plugin-multimodal
Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.
在纯文本 DeepSeek 线路上开放贴图准入,用视觉 sidecar 把附件转成文字,原生视觉模型不改写。
How to install
dsh plugin add github:shinjiyu/dsh-plugin-multimodal About
dsh-plugin-multimodal DeepSeek Harness 官方线路是**纯文本**。Web 里贴图会被拒:当前模型不支持图片。 这个插件补的是**贴图准入**,不是视觉工具箱。 1. 主模型本身收图(Claude / GPT / 自建视觉网关)→ 原样把图交给模型,不转文字 2. 主模型是纯文本(官方 DeepSeek)→ GUI 先收下图,sidecar 转成文字再发给主模型 3. see_image 给磁盘上的截图用 官方 PI adapter 已经会做第 1 步,不会做第 2 步:不支持就 UNSUPPORTED_CONTENT。Anionex 的 dsh-vision-toolkit 是另一条路:给 agent 一堆 vision_* 工具,要模型自己去调。本插件让**粘贴不被拒**。 对应需求:Discussions #588 · 介绍帖:Show and tell #1709 安装 powershell dsh plugin --profile web add github:shinjiyu/dsh-plugin-multimodal 本地路径: powershell dsh plugin --profile web add D:\tempWorkspace\dsh-plugin-multimodal 然后**重启** dsh web。旧会话的工具表…
Recommendation signals
Meta
- License
- MIT
- Language
- TypeScript
- GitHub stars
- 3
- mo. downloads
- –
- Last push
- 2026-08-16
- Created
- 2026-08-15
Links
Basic safety check
- Findings
- curated 收录但无 npm 包/安装命令
- Sources
- curated:awesome-dsh-plugin.com, curated:awesome-dsh-plugin/awesome-dsh-plugin
- Topics
- deepseek-harness, dsh-plugin, vision
Related plugins
modlens
liustack/modlens
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
dsh-web
zhu1090093659/dsh-web-ui/tree/main/packages/dsh-tool-describe-image
A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.
dsh-vision-router
ysr666/dsh-vision-router
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
dsh-vision-toolkit
Anionex/dsh-vision-toolkit
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.