dsh-vision-proxy
DeepSeek brain + automatic image transcription: attach images in the GUI and each one is transcribed via the official deepseek-v4-flash-vision-exp by default (a pure-text V4-Pro brain can see images), with any OpenAI-compatible VLM or local Ollama as alternatives.
DeepSeek 大脑 + 自动识图:GUI 附加的每张图片自动经官方 deepseek-v4-flash-vision-exp 原生识图转译,再交给纯文本的 DeepSeek 作答(纯文本 V4-Pro 也能看图);支持任意 OpenAI 兼容 VLM 与本地 Ollama 作为备选。
How to install
dsh plugin add dsh-vision-proxy About
dsh-vision-proxy English | 简体中文 **保持 DeepSeek 作为对话大脑,图片照样直接发。** 为 DeepSeek Harness 打造:GUI 附加图片自动转译,纯文本 DeepSeek 也能识图。 为什么需要它 DeepSeek Harness 原生按模型声明的 inputModalities 决定是否放行图片附件。DeepSeek 的 chat-completions 线路是纯文本的,所以选中 DeepSeek 时附加图片会被原生拒绝。已有的视觉插件提供 view_image 等*工具*(适用于文件路径),但 **GUI 图片附件对纯文本模型依然失败**。 本插件补上这个缺口:注册一条新提供商路由(deepseek-vision),包装真正的 DeepSeek 适配器——对外声明支持图片输入(附件预检放行),并在请求流里**把每张附加图片转译成文字**后再委托给 DeepSeek。对话仍然由 DeepSeek 作答,识图只是附加能力。 用户附加图片 ──▶ deepseek-vision 路由 ──▶ 经 VLM 转译(OCR+版式+细节) │ │ ▼ ▼ DeepSeek 作答 ◀── 纯文本对话(图片已替换为 [图片转译] 文字) 特性 **绝不卡死**。匿名端点强制 20 秒超时上限(免费档挂起也拖不住整轮对话);匿名端点遇到 HTTP…
Recommendation signals
Meta
- License
- MIT
- Language
- JavaScript
- GitHub stars
- 14
- mo. downloads
- 5.8K
- Last push
- 2026-08-26
- Created
- 2026-08-13
Basic safety check
- Findings
- package.json 含 postinstall 脚本
- Sources
- curated:awesome-dsh-plugin.com, curated:awesome-dsh-plugin/awesome-dsh-plugin, curated:0xsline/awesome-deepseek-harness
- Topics
- dashscope, deepseek-harness, dsh-plugin, image-understanding, multimodal, ocr, qwen, vision, vlm
Related plugins
modlens
liustack/modlens
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
dsh-web
zhu1090093659/dsh-web-ui/tree/main/packages/dsh-tool-describe-image
A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.
dsh-vision-router
ysr666/dsh-vision-router
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
dsh-vision-toolkit
Anionex/dsh-vision-toolkit
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.