Tools & Capabilities
Give text-only coding agents eyes: image Q&A, long-screenshot OCR, frontend UI restoration and GUI automation, as a vision toolkit plus a skill, with optional drop-in support for Codex, Claude Code and more.
Eyes for text-only DeepSeek Harness agents: a built-in free vision chain (no API key) plus pixel-level vision tools — Q&A, grounding, crop, pixel diff, OCR and more.
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
Capability-aware Auto routing. Keep configured order for deterministic control, or explicitly enable Auto to prioritize already-configured models using measured capability evidence. Auto never infers capability from model names, and enabling Auto alone does not start benchmarks. - Verifiable model profiling. Exact Test Vision sends one request to one exact model; Quick and Full benchmark OCR / general / structured / document / grounding capabilities. Background profiling is separately authorized and yields to real foreground vision work. - Original pixels, real answers. The vision chain reads the image at original resolution (auto-downscaled only to protect latency/quota); the agent's question travels with the image, so answers are about your question, not a generic description. - Automatic failover with classified errors. Region blocks, ToS filtering, 402 quota, 429 rate limits, context overflow, network failures — the chain walks providers one by one and only reports after all of them failed, with actionable advice. A 429 immediately advances to the next backend and opens a Retry-After-aware cooldown instead of sleeping inside the request. - Image memory. Vision answers are cached by attachment content hash; later text turns substitute the recorded description (marked as untrusted evidence), so DeepSeek genuinely remembers earlier images without re-spending vision calls.…
dsh-vision-router is a DeepSeek Harness ecosystem resource maintained by ysr666. Eyes for text-only DeepSeek Harness agents: a built-in free vision chain (no API key) plus pixel-level vision tools — Q&A, grounding, crop, pixel diff, OCR and more.
Source code and usage instructions are available at https://github.com/ysr666/dsh-vision-router. Follow the repository README for the correct setup steps.
dsh-vision-router is a community open-source project released under the MIT license. Review the repository license and documentation before use.
Tools & Capabilities
Give text-only coding agents eyes: image Q&A, long-screenshot OCR, frontend UI restoration and GUI automation, as a vision toolkit plus a skill, with optional drop-in support for Codex, Claude Code and more.
Tools & Capabilities
An open-source Smartisan-style notes app: self-hostable with one click via Docker, supports skill invocation and dsh plugin, even generates WeChat-official-account format output.
Tools & Capabilities
Detect and download video and audio from pages inside DeepSeek Harness.