40 #multimodal DeepSeek Harness (DSH) Plugins
Every indexed repository carrying the GitHub topic multimodal, sorted by stars. The tag is the author's own word for it — the install verdict on each card is ours.
Repositories tagged #multimodal, most-starred first
Modlens
liustack
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网第一个 DeepSeek Harness 视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
Agent Vision Toolkit
Anionex
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
Dsh Vision Router
ysr666
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
Dsh Vision
oil-oil
Near-native image understanding for DeepSeek Harness
Dsh Vision Complete
Yts1919
给 DeepSeek 补上「眼睛和耳朵」的多模态视觉插件:看图 / OCR / 物体检测 / 视频理解 / 语音转写 / 截图直读,一键安装(DSH 插件)。
Dsh Vision Proxy
Flyvhidbwo
DeepSeek Harness 插件:DeepSeek 大脑 + 自动识图。GUI 附加图片自动经 OpenAI 兼容 VLM 转译成文字后交给 DeepSeek 作答;支持百炼/智谱/OpenRouter 等任意 OpenAI 兼容端点(默认 qwen3.7-flash),无 key 自动探测本地 Ollama(图片不出本机);安装时有一问式确认
Dsh Visual Plugin
jyh20030112
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
Dsh Vision Provider
libinyam
Config-only DeepSeek Harness bundle for OpenAI-compatible vision models.
Dsh Plugin Deepeye
Favio8
DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.
Dsh Xiapan Media
dongsheng123132
Native vision, gpt-image-2 and Seedance plugins for DeepSeek Harness via Xiapan Cloud
Dsh Multimodal
MC5lan
给 DeepSeek 安装一双眼睛和一支画笔:会话里直接贴截图/图片,GLM 视觉模型先精确转写图片内容(报错信息、代码、界面逐字保留),然后 DeepSeek 继续处理你的问题——同一轮完成,全程无感;需要配图时,DeepSeek 自动调用文生图后端出图并显示在会话中。
Dsh Vision
reimu-create
DSH plugin: text-only models (e.g. DeepSeek-V4) automatically see images via a vision model. Official surface-replace, cache-friendly, human transcript untouched. 纯文本模型自动识图桥
GrassVison
moduqishi
给纯文本大模型装上原生视觉:流式真实思考链 · 跨轮次无感重看 · 像素级证据与 SVG 图元 · OpenAI/Anthropic/Responses 三协议兼容 | Native vision for text-only LLMs: streaming real thinking chain, cross-turn re-view, pixel-level evidence & SVG primitives, OpenAI/Anthropic/Responses compatible.
Deepseek Omnimodal
good-boy4069
Open-source multimodal MCP plugin for text-only AI agents: recognize and generate images, video, and audio through Qwen/DashScope. Supports Codex, Claude Code, and DeepSeek Harness ecosystem.
Shadow Vision
WardLu
Open-source MCP vision server that gives text-only LLMs and AI agents image understanding, OCR, visual analysis, UI inspection, and multimodal capabilities.
Deepseek Hsrness Devkit
2472786266-spec
DSH DevKit: multimodal gallery + multi-agent supervision console (DeepSeek Harness dynamic Cordis plugin)
Dsh Deepseek Vision Router
mochgolf
Transparent image preprocessing route for DeepSeek Harness
Dsh Qwen Mm
RRRosmontis
Qwen-MM-Plugins integration bundle for DeepSeek Harness (dsh) — multimodal MCP tools (vision, OCR, ASR, search, video, Blender, FreeCAD) + image attachment bridge. 让 DeepSeek Harness 原生支持多模态。
Dsh Qwen Mm Plugins
ShuiHan268
DeepSeek Harness 的 Qwen-MM-Plugins 集成插件:12 个多模态 MCP 工具(视觉/OCR/定位/ASR/音视频)、Web 设置页(粘贴 Qwen API Key 即用)、内置技能与一键安装器
Dsh See Image
tiefeiyu
A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).
Ds Vision Plugin
Sorwcyra
Paste images into DeepSeek Harness with a four-model vision race, OCR, and an automatic text bridge.
Dsh Vision Bridge
x-Xin23
给 DeepSeek Harness 纯文本模型装上原生视觉(Windows):粘贴即看图——预注入描述,模型首轮就看见,不用选模型、不用调工具;see_image 精查;自定义视觉后端(任意 OpenAI 兼容模型)+ 四后端容灾;换主模型视觉自动跟随。| Give text-only DeepSeek Harness models native-feeling vision on Windows: paste and the model just sees it — pre-injected descriptions, see_image tool, custom backends, 4-backend failover.
Dsh Plugin Vision Toolkit
YYTbit
Vision toolkit for DeepSeek Harness -- give text-only agents eyes
Deepsee
chang416
Vision + smart model routing for DeepSeek Harness. Gemini sees. DeepSeek codes.
Dsh Open Eyes
Hyp6666
A lightweight DeepSeek Harness vision delegation tool for text-only routes, with native OpenAI Responses, Chat Completions, and Anthropic Messages adapters.
Sidesight
ZhuXinAI
CLI-first vision sidecar for text-only coding agents. Analyze screenshots, diagrams, charts, UI diffs, and videos with OpenAI-compatible multimodal models.
Dsh Voice
zhuiyueya
Voice for DeepSeek Harness(dsh) — speech-to-text input + read-aloud TTS for text-only DeepSeek, zero API key.
Dsh Multimodal Skill
v587d
给纯文本 LLM 一双慧眼。 一个 DeepSeek Harness(DSH)原生 skill + 零依赖 Python CLI, 为 DeepSeek 等纯文本模型补上图像理解与文档解析(OCR、表格、公式、PDF → Markdown), 使用免费额度优先的三方多模态 API,国内网络直连、无需代理。
Dsh Mmx Multimodal
welsione
MiniMax multimodal capability hub for DeepSeek Harness (DSH): image understanding (VLM), text/image-to-video, speech, music, audio cover, web search, quota — one mmx_multimodal model tool wrapping the mmx-cli.
Dsh Guide Dog
AtropinolTT
Guide Dog for DSH — MiniMax multimodal plugin: image/video/music/speech generation & vision tools, hard-metric voice mode (host event-driven auto TTS, session-safe playback), voice input (mic → faster-whisper → composer), recorder page, two-layer settings.
Dsh Vision
yyking12
dsh-vision · 开眼 — 给 DeepSeek Harness 的 Agent 装上眼睛:非多模态模型也能看图(analyze_image + 发图/粘贴/拖放标记管线,任意 OpenAI 兼容供应商)
Dsh Qwen Vision
Asanagl
DeepSeek Harness 插件:Qwen-MM-Plugins 视觉 MCP + 聊天框粘贴图片 + 看图技能
Analyze Image Tool
CaseyTso
给纯文本 DeepSeek Harness 模型加上识图能力:analyze_image 把图片转发到任意 OpenAI 兼容视觉端点 | Vision bridge for text-only DSH models
Mimo Vision
wulusai2333
DeepSeek Harness (DSH) native plugin — describe_image tool: a vision bridge (image → mimo-v2.5 → text description) over the ctx.fs / ctx.credentials seams
Dsh Qwen Multimodal
wuwangmao
DSH bundle: Qwen multimodal bridge — vision (qwen3-vl), speech-to-text (qwen3-asr), text-to-image (qwen-image), for DeepSeek Harness
Vision Kit
Seom-ingit
Make your AI agent a math tutor. Structured extraction of vectors, matrices & geometry from math figures, with dimension-consistency + geometric self-check. Vision plugins for DeepSeek Harness, opencode (MCP) & CLI. Verify, don't believe.
Deepseek Harness Extension
hypergraphdev
DeepSeek Harness browser-extension edition: side panel with page-context awareness, vision-bridge multimodal image reading, and voice input
Dsh Visionary
zhuiyueya
Give text-only DeepSeek models eyes — a DeepSeek Harness plugin that transparently converts chat images into OCR text + vision-model descriptions before they reach the LLM. Configure vision backends (GLM-4V, Qwen-VL, Gemini, Ollama…) right in the Models settings page; multi-backend fallback chain, double-layer caching, no config files.
Dsh Tool Vision
pzqian123
DSH plugin: read images through a user-configured vision-capable model when the main model route cannot accept image input
Dsh Vision Relay
junhongchashui
零修改、零切换的 DeepSeek Harness 视觉能力插件:纯文本模型粘贴即读图片,云端 + 本地 Ollama 双后端自动切换,ModLens v2 风格结构化证据输出。