Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.
dsh plugin add kaixinbaba/dsh-vision-recognizer
Reviewed DeepSeek Harness extensions — filter by category, search by capability, and jump straight to each plugin's public source.
Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.
dsh plugin add kaixinbaba/dsh-vision-recognizer
Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, GUI automation, through layered evidence memory and cache.
dsh plugin add kanchengw/dsh-mindseye
Turn every idea into an image or video with Kling AI. In DeepSeek Harness, use natural language for text-to-image, image-to-image, text-to-video, image-to-video, reference-image creation, task tracking, and result previews.
dsh plugin add klingai-dev/deepseek-plugin
Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).
dsh plugin add ld-1101/dsh-vision-plugin
GPT Image 2 `image_gen` with Codex subscription OAuth by default or explicit API-key mode: developing card, up to three live API partials, durable attachment replay/lightbox/download, text-only model output, and bounded credential-safe requests.
dsh plugin add LeemanCheung/dsh-image-gen
On-demand vision for text-only DeepSeek models: upload images, and the model calls a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).
dsh plugin add Leeminjing/dsh-eyes
Structured vision evidence (OCR/layout/semantics) plus a USB camera capture tool for a "shoot-look-adjust" debug loop; backends: Ollama, DeepSeek, Xiaomi.
dsh plugin add lhbsaa/dsh-visibridge
External vision plugin for DeepSeek Harness: whale-button config panel, image recognition with auto-reply, and agent screenshot/recognize tools.
dsh plugin add linenxi-ctrl/dsh-vision
Local image, voice, music and SFX generation plus transcription, with pinned identity: characters, animals, objects and actor voices are defined once and reused on every later call, degenerate output (a near-flat image, silent audio) is rejected instead of returned as success, and a generated line can be read back as text so a clone that swallowed its ending becomes visible. The engines unload when idle, and image generation and transcription can each be pointed at an OpenAI-shaped API instead of the local Vulkan backend.
dsh plugin add linxuhao/Deepseek-Continuity
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
dsh plugin add liustack/modlens
Adds a take screenshot tool plus two optional automatic capture points, putting the screen into the conversation as an image block; requires an image-capable model.
dsh plugin add lkh081231/screenshot-feedback-hook-mcp#dsh-screenshot-feedback-hook-mcp
Gives dsh the ability to generate images and videos through the grok2api API.
dsh plugin add lsjspl/dsh-plugin-grok2api-media-tool
Dual-engine screen capture for DSH: inside the SSiD desktop shell, a global hotkey or tray opens a fullscreen box-select overlay (all monitors, per-display pixel-perfect frames) with in-place red-box annotation and a WeChat-style toolbar; in plain DSH (browser), the composer camera button captures the display via getDisplayMedia and reuses the same single-phase box-select + annotation overlay in-page. The cropped image lands in the current conversation composer through the official attachment intake.
dsh plugin add Max-Null/dsh-capture
Local OCR for attached images via Tesseract: only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
dsh plugin add maxwell-feng/dsh-tesseract-ocr
Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
dsh plugin add maxwell-feng/dsh-windows-ocr
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.
dsh plugin add MicroHEROX/dsh-koboldcpp-hands
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.
dsh plugin add MicroHEROX/dsh-unsloth-hands
Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.
dsh plugin add MJorgin/dsh-media-skills
Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.
dsh plugin add moon09300731/dsh-vision-tools
Vision-augmented DeepSeek adapter: a vision-capable model describes image input, then a text-only DeepSeek model reasons over the description.
dsh plugin add NagasakiSoyo-ui/dsh-llm-deepseek-vision
Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.
dsh plugin add niuniuaba/dsh-subagent-vision
Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.
dsh plugin add niyongsheng/free-vision-skill
Auto-discovery vision bridge for text-only DeepSeek Harness agents: automatically finds an image-capable model from your configured providers and returns picture descriptions as plain text via a vision tool.
dsh plugin add NormanFxxkingRockwell/dsh-auto-vision
Control a HarmonyOS phone from the DSH web UI with AI: live H.264 screen mirroring, mouse touch and system keys, hilog streaming, and agentic tools that let the model read the screen, locate UI controls, then tap, long-press, press keys, or type.
dsh plugin add ns-zzj/dsh-hos-scrcpy
Page 3 of 5
Filter by category or keyword — every listing is a public GitHub project.
Commands follow the upstream list; check each repo's README for exact package names.
Plugins run third-party code with your permissions — read the source first.
This directory is a snapshot of the community-maintained awesome-dsh-plugin list (updated 2026-09-07). Listings link to the authors' repositories; inclusion is not an endorsement or a security review. Missing a plugin? Contribute to the upstream list.