s3yf1337/dsh-easyvision

DeepSeek Harness plugin: describe images through a dedicated vision model from the dsh model list, called over the harness's own LLM runtime

Bundle verified MIT JavaScript Unknown
Bundle verified

Listed

0

Vision

Bundle verified

VersionUnknown
LanguageJavaScript
LicenseMIT
View on GitHub

Preview

Preview 1 of 2: s3yf1337/dsh-easyvision
Preview 2 of 2: s3yf1337/dsh-easyvision

What it does

Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.

Best for

  • Users who prefer a text-only conversation model but occasionally need to inspect attached images or image files.
  • DSH setups that already include a vision-capable model and want to reuse the harness LLM runtime and credentials.
  • Workflows that need image descriptions returned as plain text to the main model.

Not ideal for

  • Setups with no configured model that declares image-input support.
  • Users whose main model already handles images directly and do not need delegated descriptions.
  • Installations where the required host patch cannot be applied or maintained after dsh upgrades.

README

dsh-easyvision

version license dsh

Give your text-only agent eyes — with one command and zero extra APIs.

A DeepSeek Harness (dsh) plugin that lets a text-only conversation model “see” images by delegating them to a vision model from your own dsh model list, called over the harness’s own LLM runtime.

Why

Your main model (e.g. deepseek-v4-flash) is text-only, so dsh’s built-in read_image tool refuses to send image blocks to it. dsh-easyvision fixes that in two complementary ways:

  • Attached images in the web chat just work. When you drop an image into the composer and send it, the message is admitted and the image is described through the vision model — no more “The current model does not support images; switch to a model that does” refusal. This happens only while the plugin is active, configured, and resolves a vision-capable model; if anything is wrong with the plugin you get an actionable “configure EasyVision” error instead.
  • describe_image tool — the model can also inspect image files on its own by calling the tool, which hands the picture to the vision-capable model and returns the description as plain text.

No external API keys. No extra plumbing. Just a model that can see, picked from the models you already have.

Features

  • One command install — idempotent, safe to re-run
  • Zero external APIs — the vision call goes through ctx.llm, the exact same runtime the agent loop uses: your keys, your retry policy, your middleware
  • Any vision model — pick anything from your dsh model list in Settings → EasyVision; no vendor lock-in
  • Live configuration — model changes apply immediately, no restart
  • Multiple images per call — validated PNG/JPEG/WebP/GIF, same attachment pipeline as read_image
  • Composer image drops — images attached to a chat message are described automatically when the conversation model is text-only

Screenshots

Configure the vision model in the dsh Settings UI — no file editing:

Settings → EasyVision

Quick start

curl -fsSL https://raw.githubusercontent.com/s3yf1337/dsh-easyvision/main/install.sh | bash

That’s it. Then open Settings → EasyVision and pick a vision-capable model from your list (the default is qwen3.7-plus on opencode-go).

Only models that declare image input work — a text-only pick is refused by the tool with a clear message.

Demo

$ dsh "what's in testpics/1.jpg?"

  ✦ describe_image(file_paths=["testpics/1.jpg"])
  ✓ qwen3.7-plus (opencode-go) · 1024×1024

  A futuristic cityscape at night — glowing cyan and blue towers
  under three moons, rendered in a digital painting style.

How it works

Two entry points, one pipeline:

composer: drop an image into the chat ──▶ host session.prompt admission
                                              │  model text-only?
                                              ▼
                          easyvision bridge (ctx service, health check)
                                              │  admitted AS-IS: the message
                                              ▼  keeps its real image blocks
                        the chat shows the picture; the model request is
                        transformed at dispatch time (llm prepareCall/stream):
                                              │
                         ctx.llm.stream(provider, model, messages=[image blocks + prompt])
                                              │
                                   vision model (e.g. qwen3.7-plus)
                                              │
                    description text replaces the image blocks in the request
                                              │
                               text-only model reads the description

model turn: describe_image(paths, prompt?) ──▶ same vision pipeline, called as a tool

The install.sh / scripts/patch-dsh-host.mjs host patch changes the session.prompt admission: when the conversation model is text-only and the message carries images, the host asks the plugin’s easyvision service to confirm it can take them instead of refusing outright. The prompt is admitted only while the service is present (plugin active) and the configured model resolves and declares image input (plugin configured and healthy) — and the message keeps its real image blocks, so the web chat shows the picture. The describing itself happens at dispatch time: the plugin wraps the shared LLM runtime’s prepareCall/stream, and any request whose model is text-only has its user-message image blocks replaced by the EasyVision description before the adapter sees them. Otherwise the client gets an actionable error naming the fix (see Troubleshooting).

The plugin validates the configured model against your dsh model list and refuses to run when it is missing or does not declare image input — before any image bytes are accepted.

Configuration

Everything is optional — the defaults work out of the box.

Control Meaning
Vision model (picker) Any model from your dsh model list, grouped by provider — the same catalog the composer’s picker uses.
Advanced → Max tokens Optional output cap for the vision call.
Advanced → System prompt System prompt sent to the vision model before every call.
Advanced → Default prompt Question used when describe_image is called without a prompt.

Profile-level defaults live in the plugin’s entry config (cordis.patch.yml) and act as the base layer: Settings overrides inherit from it, and “Reset to default” restores it.

Tool

describe_image(file_paths: string[], prompt?: string) — sends one or more image files to the vision model and returns its description, plus the resolved provider/model, per-image dimensions, and whether the response hit the token limit.

Troubleshooting

Frequently Asked QuestionsFAQ

Use the verified command dsh plugin --profile default add github:s3yf1337/dsh-easyvision in a DSH-enabled shell. The command resolves the public package metadata and keeps the plugin attached to the catalog identity shown on this page.