kanchengw/dsh-mindseye
Plug-in vision for text-only models on DSH, with native multimodal interaction, layered evidence memory and cache, and intent router.
已收录
1
Vision
Bundle 已验证
预览
功能介绍
为纯文本 DeepSeek Harness 模型提供视觉能力:原生图片粘贴、分层证据记忆与缓存、意图驱动的工具选择。
适合
- 需要视觉问答、OCR、布局、图表、颜色、像素差异或坐标定位等意图专用工具的 DSH 用户。
- 可受益于结构化证据、精确缓存和持久科研记忆的重复图片分析工作流。
- 同时需要图片理解,以及可配置图片生成或编辑路由的用户。
不适合
- 未配置通用视觉模型端点、密钥和模型 ID 的环境。
- 仅需一次性图片转译,分层记忆、路由和多工具机制价值有限的场景。
- 期望生成图片自动保存到项目或自动验证的工作流;该插件两者都不执行。
README
MindsEye

Let DeepSeek see images natively — model-driven vision tools for DeepSeek Harness
| English | 中文 |
Current version: 0.2.4
MindsEye is a vision plugin for DeepSeek Harness (dsh). Pasted images stay visible in the conversation while DeepSeek keeps reasoning and the vision model does the seeing. The plugin exposes task-specific vision tools that the model selects by intent; each tool maps to a fixed intent and model route, returns structured JSON, and reduces repeated calls through caching and evidence reuse.
Core Experience
-
Paste and see: automatically attempts the
deepseek-officialbridge so images enter the conversation natively; when unavailable it falls back to path-based paste, so new images always get through -
Model picks the intent, plugin routes the model:
mindseye_read_imagetakes anintent(visual-qa / ocr / layout / chart / color / pixel-diff / general) plus optionalextractfor combined structured evidence in one call;mindseye_groundstays separate for coordinates -
Generated images appear in the conversation:
mindseye_generate_imagedelegates to a dedicated image generation model and returns the result as a dsh attachment, without auto-saving to the project or running automatic verification - Automatic mounting on image turns: vision tools are registered when an image message arrives; text-only turns keep only one activation entry so tools do not occupy model context permanently
- Batch reads in one call: multiple images are read together, with exponential split fallback on batch 4xx so a failure affects only the failed image
- Old sessions stay clean: image-bearing history remains usable in fallback mode, with image blocks rewritten into attachment markers
- Every call is transparent: provider, model, latency, token usage, and fallback markers are returned for auditability

Implemented Features
Image Input
- Native paste/drag (automatic bridge, no duplicate model selector entry)
-
paste-to-pathfallback: pasted images are converted to path text in text-only model scenarios -
mindseye_read_imagegeneral vision, supporting local paths, single attachment ids, and batch attachment ids
Tools and Routing
| Tool | Intent | Route | Batch |
|---|---|---|---|
mindseye_read_image |
General vision QA + intent tasks (ocr / layout / chart / color / pixel-diff / general), optional extract for combined evidence |
understand / extract per intent | Yes |
mindseye_ground |
Target pixel coordinate location | locate | No |
-
understand / extract / locatemodel routes are independently configurable and fall back to the general understanding model when unset - Vision tools auto-mount on image turns; text-only turns keep only
mindseye_vision_activateso tools do not permanently consume model context - Structured JSON:
images/evidence/answer/meta;metaincludes real token usage, call attempts, and fallback markers - Exact cache: image sha256 + normalized query + region + baseUrl + model + prompt version; a hit skips the vision model call
Image Generation
-
mindseye_generate_image(intentId): text-to-image throughimage.generate -
mindseye_edit_image(intentId, attachmentId): image-to-image throughimage.edit, sending the reference image to the provider - Generated results are displayed as dsh attachments with a
(token_usage=..., widthxheight, size)audit line
Providers
- OpenAI-compatible Chat Completions and Responses protocols
-
vision.fallbacksis the recognition-chain fallback list; image generation and editing use the orderedimage.generateandimage.editchains directly - Image routes expose configurable
endpoint,bodyMode(jsonormultipart), andimageFieldfor provider compatibility - Multi-image batch calls with exponential fallback (batch 4xx retries by halving;
locatedoes not support batch)
Memory
- Image-level hard facts are persisted by sha256, with evidence evicted by capacity LRU (default 1000 entries)
- Soft memory uses BM25 retrieval of historical Q&A injected as context, evicted by capacity (default 1000 entries)
-
mindseye_memory_put / get / search / diffare exposed as dsh tools; calls are visible in the session and audited
Data Handling and Security
- Native attachments first: image-capable models keep native dsh attachments; MindsEye associates images by attachment id and does not ask the user to choose local files manually
-
Automatic temporary path fallback: when the current model is confirmed text-only and
paste-to-pathis enabled, freshly pasted PNG, JPEG, WebP, or GIF files (up to 25 MiB each) are validated, stored in an isolated system temp directory, and returned as a path. Temp files use0600mode - External vision calls: image bytes and question text are sent only to the vision provider’s Base URL when a MindsEye tool executes; configure only services you trust
- Credentials and cache: API keys resolve from environment variables, dsh Credentials, or plugin settings and are sent as Bearer auth to the matching provider only. The exact cache lives only in the current dsh process memory (max 500 entries), is never persisted, and clears on process exit
- Execution boundary: the plugin never starts shells, child processes, or executes downloaded code. The normal web paste fallback only reads the temporary image just created by the plugin; tools also accept dsh attachment ids
dsh Web Settings Card
-
understand / extract / locateroutes can be added as needed; unset routes fall back to the default model - Base URL, API key (masked with eye toggle), model id, protocol (explicit), and common Max Tokens values
- Model takeover is always attempted at startup; a failed attempt restores the official adapter and continues in path-paste mode
Installation
npx @deepseek-ai/dsh plugin --profile web add dsh-mindseye
Restart dsh web to paste images natively. Takeover is automatic and is not exposed as a user setting.
On first use, configure one general vision model (Base URL, API key, model id) in the MindsEye settings card; unconfigured OCR / locate routes fall back to the general model automatically.
Development
pnpm install
pnpm test
pnpm typecheck
pnpm build
常见问题常见问题
在启用了 DSH 的终端中执行已验证命令 dsh plugin --profile default add github:kanchengw/dsh-mindseye。命令会解析公开 package 元数据,并保持插件与本页展示的目录身份一致。
兼容性以页面上展示的 bundle 与 profile 状态为准。如果某个 profile 尚未检测到,请先保持禁用,并在生产启用前阅读仓库文档。
GitHub 链接和 activity 元数据是 release 与维护状态的来源。新版本发布后重新查看本页,确认目录已经观察到最新版本。