Letter2025/dsh-model-failover
Two-level model circuit breaker with failover for DeepSeek Harness: trip a model or a whole provider after repeated request failures and route the next request to a configured fallback
已收录
3
Workflow
Bundle 已验证
功能介绍
两级模型熔断与回退:模型或平台连续失败后自动熔断,并把下一个请求路由到配置好的备用模型。
适合
- 需要在模型或平台请求连续失败后自动恢复的 DSH 部署。
- 已明确配置兼容备用路由的多平台环境。
- 已采用有界重试,并希望对重试后仍然存在的故障进行熔断的工作流。
不适合
- 要求多个实例共享熔断状态的部署,因为状态仅存在于进程内,并会在插件重载时重置。
- 会话标题、压缩或自定义 `ctx.llm.stream` 等辅助模型调用;插件只路由 agent loop 请求。
- 使用无限重试的 `retryPolicy.mode: 'always'` 配置,因为失败不会传递给熔断器。
- 未选择兼容路由时,使用上下文窗口更小或不兼容的备用模型的场景。
README
dsh-model-failover
| English | 中文 |
Two-level model circuit breaker with failover for DeepSeek Harness. When a model (or a whole provider) starts failing repeatedly, the plugin opens a circuit and routes the next model request to a configured fallback — no core changes, installable with dsh plugin.
What it does
-
Model circuit — a
provider/modelroute opens aftermodelCircuitThresholdfailures insideburstWindowMs. -
Platform circuit — a provider opens when
platformCircuitThresholddistinct models under it are open at once (a platform outage usually takes down every model on it). -
Failover — the next request goes to the first healthy fallback in
fallbacks; the switch is recorded by the loop itself (request/headerchange) and optionally announced as a user-visible message. -
Recovery probes — an open model circuit is probed after
modelCooldownMswith a tiny real call; a successful probe closes the circuit, a failed one extends the cooldown. -
Composes with
llm-retry— per-request backoff retries stay owned by the bundledllm-retrypolicy; the breaker observes the failures that escape it, so transient blips that recover after a retry never trip a circuit.
Install
dsh plugin --profile web add dsh-model-failover
Then configure fallbacks (the only field you must set) and, if needed, the thresholds in your profile cordis.patch.yml — see the plugin row in cordis.patch.yml for the full default config.
A companion skill, configure-model-failover, walks the agent through setting the fallback models (AI probes the current model config, writes the fallback override, then asks you to confirm). It is installed automatically with the plugin: the package ships skills/configure-model-failover/SKILL.md and the plugin registers it as a bundled skill whenever the skills service is present — no copy step needed. For a standalone install (profile without the plugin), copy it into ~/.dsh/skills/configure-model-failover/ (user-level skills are picked up live).
How it works
The plugin decorates two agent-loop waterfalls (both official extension points, no core changes):
| Waterfall | Role |
|---|---|
agent/request-error |
Records failures whose code is in tripCodes into the circuit breaker, then delegates through next() so llm-retry still owns retries. |
agent/request |
await next() for the resolved config, then returns the healthy primary route, or the first healthy fallback when the primary’s circuit is open. |
request ──> agent/request ──> primary (mock/m1) ──> fail ×2 ──> circuit open
│
next request ──> agent/request ──> primary open ──> fallback (mock2/m2) ✔
│
probe after cooldown ──> success ──> circuit closed
Configuration
| Field | Default | Meaning |
|---|---|---|
enabled |
true |
Master switch. |
fallbacks |
[] |
Ordered fallback routes {provider, model}; must point at providers with a registered adapter. |
tripCodes |
RATE_LIMIT, SERVER, TIMEOUT, TRANSPORT, QUOTA, EMPTY_RESPONSE |
Failure codes that count toward a circuit; e.g. AUTH/INVALID_CREDENTIAL stay terminal. |
modelCircuitThreshold |
2 |
Failures inside the burst window that open a model circuit. |
modelCooldownMs |
60000 |
Cooldown before an open model circuit is probed. |
platformCircuitThreshold |
2 |
Distinct open models that open the whole provider. |
platformCooldownMs |
120000 |
Provider-wide cooldown. |
burstWindowMs |
300000 |
Failures older than this start a fresh burst. |
enableProbe |
true |
Probe open models after cooldown to recover circuits. |
probeMaxTokens |
8 |
Output cap for probe calls. |
stripReasoningEffort |
true |
Drop the primary’s reasoning effort when failing over (fallbacks may not support it). |
notifyUser |
true |
Append a user-visible message when a route switches. |
Events
Plugin-defined (emit) events, typed via the @deepseek-ai/cordis augmentation in src/types.ts:
-
model-failover/circuit-opened—{provider, model, level: 'model' | 'platform'} -
model-failover/circuit-closed—{provider, model, level: 'model'} -
model-failover/failover—{from, to, agentId} -
model-failover/probe—{provider, model, ok, message?}
Known Limitations and Deferred Work
- Process-local state — circuit state lives in memory and resets on plugin reload (like every harness registry). Cross-instance sharing is deferred.
-
Agent-loop calls only —
agent/requestcovers the main conversation loop. Auxiliary calls (session-title,compaction, hand-builtctx.llm.stream) are not routed. -
retryPolicy.mode: 'always'— the bundledllm-retrynever delegates a failure to this breaker in that mode, so failover stays idle by design (the operator chose unbounded retries). -
No context-window adaptation — a fallback with a smaller context window may hit
CONTEXT_WINDOW_EXCEEDED; setstripReasoningEffortand pick compatible fallbacks. - No platform probe — the platform circuit recovers by cooldown expiry; only model circuits are probed.
License
MIT
常见问题常见问题
在启用了 DSH 的终端中执行已验证命令 dsh plugin --profile default add github:Letter2025/dsh-model-failover。命令会解析公开 package 元数据,并保持插件与本页展示的目录身份一致。
兼容性以页面上展示的 bundle 与 profile 状态为准。如果某个 profile 尚未检测到,请先保持禁用,并在生产启用前阅读仓库文档。
GitHub 链接和 activity 元数据是 release 与维护状态的来源。新版本发布后重新查看本页,确认目录已经观察到最新版本。