3274375092/dsh-voice
Voice input plugin for DeepSeek Harness: mic → local/browser speech recognition → text submitted as a normal chat message. Input-only and preset-agnostic.
已收录
4
Voice
Bundle 已验证
预览
功能介绍
语音输入插件:对着麦克风说话,识别后的文字作为普通聊天消息发送(本地模型或浏览器语音识别)。
适合
- 希望用语音代替键盘输入聊天消息的 DSH Web 用户。
- 需要只改变输入方式、不修改 preset、persona 或系统提示的用户。
- 可以安装可选 sherpa-onnx 本地识别运行时与模型的隐私敏感工作流。
- 浏览器支持 Web Speech、希望零额外配置使用语音输入的用户。
不适合
- 需要语音输出、文字转语音或语音 Agent 对话的用户;该插件只负责输入。
- 没有浏览器麦克风采集和输入框界面的 headless 或非 Web 工作流。
- 不愿安装可选本地运行时及约 100 MB 模型的离线识别环境。
- Web Speech 不可用或被禁用,且未配置本地模型的纯浏览器环境。
README
dsh-voice 🎤
| English | 中文 |
A voice input plugin for DeepSeek Harness: click 🎤 in the web UI (or press a hotkey), speak, and the recognized text is submitted as a normal chat message. Input only — it never touches the agent preset/persona, so it behaves like “another input method” in every mode.
Features
- 🎤 Voice input: microphone button (platform design-system UI) + configurable global hotkey (Ctrl+Space by default)
- ⚡ Live recognition: streaming partial transcripts are echoed while you speak; VAD finalization commits on stop (0.6s tail padding keeps sentence endings)
- 🧠 Adaptive dual engine: host-native ASR (sherpa-onnx-node zipformer2 + silero VAD, offline/private) with automatic fallback to browser Web Speech (zero extra dependencies)
- 🔌 Preset-agnostic: does not touch the persona/system prompt; works with code/standard/minimal/custom presets
- 📦 Optional models: zero-config out of the box; run
dsh-voice-modelswhen you want native offline recognition
Installation
# Plugin
dsh plugin --profile web add @nn12138/dsh-voice
# Optional: offline native recognition (the plugin does not auto-install this runtime)
dsh plugin --profile web add sherpa-onnx-node
dsh-voice-models # one-shot model download (~100MB) → ./dsh-voice-models
# Optional: configuration (edit ~/.dsh/profiles/web/cordis.patch.yml)
- id: voice
config:
modelDir: './dsh-voice-models' # native ASR model directory
hotkey: 'ctrl+space' # global hotkey
vadThreshold: 0.3 # lower = less clipping at sentence boundaries
tailPadSeconds: 0.6 # tail-padding duration
engine: auto # auto (default) | native | browser
The row-level config is received by the host half. engine and hotkey are synced to the browser half over the /voice.config loopback RPC, so there is no separate client config to write. auto probes host native capability: with a model it uses native; without one it falls back to Web Speech, so zero-config users keep working. Restart dsh web after changing the config.
dsh web # 🎤 button appears on the left of the composer, or press Ctrl+Space
See USAGE.md and INSTALL.md (Chinese) for details.
How it works
Browser captures mic audio (auto-resampled to 16 kHz)
→ PCM base64 chunks (256 ms) → /voice RPC channel (loopback)
→ host: silero VAD + zipformer2 streaming decode
→ partials returned per chunk (live echo) / finals committed
(VAD segmentation + 0.6s tail padding)
→ conversation service submits the text (same path as typing)
Engine selection: the host resolves the effective engine (config + model-load result) and the client consumes it via /voice.ping — native unavailable falls back to browser Web Speech. /voice.config carries the row-level engine/hotkey from host to client.
Development
pnpm install --ignore-workspace # standalone deps (no DSH monorepo needed); prepare auto-builds
pnpm --ignore-workspace test # unit tests (including real-model smoke tests)
pnpm --ignore-workspace typecheck # type check
pnpm --ignore-workspace build # build (tsc host half + tsdown client half)
Real-model smoke tests look for the local voxelf assets and skip when absent; override with:
DSH_VOICE_MODEL_DIR (model directory) / DSH_VOICE_TEST_WAV (test wav) / DSH_VOICE_DOWNLOADED_MODELS (downloaded model directory).
Layout: src/index.ts (host half) / src/client/ (browser half) / src/core/ (recognition core) / tools/ (wire-protocol smoke tools + model downloader).
License
MIT
常见问题常见问题
在启用了 DSH 的终端中执行已验证命令 dsh plugin --profile default add github:3274375092/dsh-voice。命令会解析公开 package 元数据,并保持插件与本页展示的目录身份一致。
兼容性以页面上展示的 bundle 与 profile 状态为准。如果某个 profile 尚未检测到,请先保持禁用,并在生产启用前阅读仓库文档。
GitHub 链接和 activity 元数据是 release 与维护状态的来源。新版本发布后重新查看本页,确认目录已经观察到最新版本。