Anionex/dsh-vision-toolkit

[dsh]为纯文本模型设计更强大的视觉工具箱:安装免费使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

Bundle 已验证 MIT TypeScript v0.1.31
Bundle 已验证

已收录

659

Vision

Bundle 已验证

版本v0.1.31
语言TypeScript
许可证MIT
在 GitHub 查看

预览

第 1 个预览,共 2 个:Anionex/dsh-vision-toolkit
第 2 个预览,共 2 个:Anionex/dsh-vision-toolkit

功能介绍

让纯文本模型更好地做视觉任务:带意图的图片问答、长截图 OCR、UI 还原等。

适合

  • 需要围绕一张或多张图片进行任务导向问答的纯文本 DSH 模型。
  • 涉及长截图 OCR、视觉定位、裁剪、描摹或查找 UI 元素的流程。
  • 需要 UI 还原或像素对比的前端复刻与视觉验证流程。

不适合

  • 不需要图像理解或视觉验证的纯文本流程。
  • 持续用量超过内置免费服务所注明的每台机器每天 100 张图片额度,且没有其他合适服务配置的场景。
  • 明确要求基础文本模型自身具备原生多模态能力,而不是借助外部视觉工具的流程。

README

DSH Vision Toolkit helps text-only DeepSeek Harness agents understand images and complete visual tasks

DSH Vision Toolkit

Anionex%2Fdsh-vision-toolkit | Trendshift

Recommended by dshfind dshfind score: 94 — highest-rated plugin agentic leaderboard

npm MIT DSH

A more powerful vision toolkit—give text-only models in DeepSeek Harness eyes: image Q&A, long-screenshot OCR, UI restoration, and GUI visual tasks in one toolkit and Skill.

🚀 Paste an image and ask directly Install with one command Built-in free vision Broad use cases
Highlights Quick start Toolbox Configuration and limits Troubleshooting Community
🌐 English 中文

🏆 This project is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem: it was initiated before internal beta and built during the beta with reference to agent-vision-toolkit.

Original work: The system and division of responsibilities behind these visual tools, together with the vision-skills Skill, were personally created and continuously refined by the author through long-term real-world use and repeated iteration.

If this project helps you or gives you some inspiration, feel free to star 🌟 & fork.

Highlights

  • Paste an image and ask directly. In DSH Web, pasting an image switches the text-only model to its (Vision Toolkit) variant automatically — no manual path copying or model changes. Native thumbnails, session history, and workspace paths stay intact; Web can preview artifacts.
  • One command to install. The built-in free Gemini 3.7 Flash vision service is ready after installation, with no API key required.
  • Built-in free vision quota. The shared service works immediately after installation with a quota of 100 images per machine per day.
  • Not just a caption — the content that matters. The model does not produce a generic description; it extracts evidence around the current task, such as “Where is the error?” or “Where is the button?”.
  • A battle-tested visual-task methodology. The bundled Skill tells the agent what to look at for different visual tasks, which tool to choose, how to proceed, and how to verify the result.

agent-vision-toolkit gives an agent more than image captions: it can read, locate, crop, trace, rebuild, and verify visual work. DSH Vision Toolkit is its native DeepSeek Harness integration, bringing that workflow into Web and Headless Profiles.

This project has two layers:

  1. Visual tools and a Skill: the agent learns when to inspect, ground, OCR, crop, trace, or compare pixels.
  2. Native DSH integration: those capabilities live inside Profiles, sessions, Settings, Artifacts, and the Web UI, with a free Gemini 3.7 Flash vision service ready after installation.

Install and use it immediately. The default setup includes a free Gemini 3.7 Flash vision service and requires no API key.

dsh plugin --profile web add @anionex/dsh-vision-toolkit

Upstream toolkit: Anionex/agent-vision-toolkit · Project website: agent-vision.anionex.me

❤️ Sponsor

Want to sponsor this project? See FUNDING.md or email davidyang042@gmail.com.

常见问题常见问题

在启用了 DSH 的终端中执行已验证命令 dsh plugin --profile default add github:Anionex/dsh-vision-toolkit。命令会解析公开 package 元数据,并保持插件与本页展示的目录身份一致。