← Back to list

modlens

DeepSeek HarnessClaude CodeCodex Other Low risk

Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).

为纯文本模型架起视觉桥梁:粘贴图片,输出结构化 JSON 证据(OCR、版面、语义)。

How to install

DeepSeek Harness dsh plugin add @liustack/modlens
Claude Code git clone https://github.com/liustack/modlens ~/.claude/skills/modlens
Codex npx -y @liustack/modlens

About

The flagship DeepSeek and GLM chat models are text-only and cannot read images. ModLens is a plug-in vision engine that gives a text-only model sight. **ModLens reads images pasted straight into the chat**, no saving to a file and passing a path first. Talk to us Issues are welcome any time: open one. And come find me on X: **@liustack**. What you built with it, which harness you are on, what should come next. New releases land there first, and a proper community space is on the way. Highlights **🥇 The most capable vision plugin for DeepSeek Harness (dsh):** one command, npx -y @deepseek-ai/d…

Recommendation signals

92 Tool quality · Based on stars, downloads, maintenance, security and docs
– User interest · Adjusted by in-site views, install copies and download clicks
92 Overall
0views
0unique visitors
0install copies
0download clicks
0outbound clicks

Meta

License
MIT
Language
TypeScript
GitHub stars
4.0K
mo. downloads
106.1K
Last push
2026-09-24
Created
2026-02-22

Links

Basic safety check

Findings
None
Sources
curated:awesome-dsh-plugin.com, curated:awesome-dsh-plugin/awesome-dsh-plugin, curated:0xsline/awesome-deepseek-harness
Topics
agent-skills, claude-code, claude-skills, codex, cordis, deepseek, dsh, dsh-plugin, glm, harness, harness-engineering, hermes-agent, image-to-text, multimodal, ocr, openclaw, pi-agent, text-only-llm, vision, vision-transformer

Related plugins

DeepSeek Harness Featured
Score72

dsh-web

zhu1090093659/dsh-web-ui/tree/main/packages/dsh-tool-describe-image

A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.

☆ 8.0K ↓ – Other ↗
DeepSeek Harness
Score70

dsh-vision-router

ysr666/dsh-vision-router

Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.

☆ 1.1K ↓ 66.0K Other ↗
DeepSeek Harness
Score67

dsh-vision-toolkit

Anionex/dsh-vision-toolkit

Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.

☆ 883 ↓ 42.3K Other ↗
DeepSeek Harness
Score67

dsh-design-qa

sunxin-ai/dsh-design-qa

Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.

☆ 44 ↓ 560 Other ↗