Relistencode/dsh-recall

Conversation history recall for DeepSeek Harness (DSH) — literal/fuzzy/semantic retrieval of every past conversation, fully local & offline. AI never forgets what you told it.

Bundle 已验证 MIT JavaScript v0.2.1
Bundle 已验证

已收录

1

Memory

Bundle 已验证

版本v0.2.1
语言JavaScript
许可证MIT
在 GitHub 查看

功能介绍

对话历史回忆:字面/模糊/语义三层检索全部历史会话的原始文本,完全本地离线——AI 再也不会忘记你说的话。一条命令安装(自带 `dsh.bundle.patch`),语义推理跑在 worker 线程。

适合

  • 对话跨越大量会话或长到难以翻阅的重度用户。
  • 需要逐字找回旧设定、情节细节或人物关系的写作及角色扮演工作流。
  • 需要追溯以往决策依据的代码和文档维护者。
  • 需要在对话压缩后仍通过本地离线方式检索原文的用户。

不适合

  • 历史会话很短、容易直接翻阅的用户,回忆检索价值有限。
  • 无法接受文档所述语义模型及运行时体积的安装环境;完整安装约 37 MB。
  • 无法接受首次搜索建立索引及数分钟后台语义预热的用户。
  • 明确省略或关闭语义模型、却仍要求语义检索的工作流;此时只剩字面和模糊检索。

README

dsh-recall

🌏 中文 · English

Conversation history recall plugin — a “memory maze” for your DeepSeek Harness agent

Version License Node Platform [Offline

AI never forgets what you told it.

A native DeepSeek Harness (DSH) plugin that gives the agent a memory maze — corridors and rooms built from every conversation you have had together. Every decision, setting, discussion, or casually mentioned requirement is remembered. Ask “where were we?” and it walks the maze, brings back the conversation verbatim, and answers as naturally as if it had never forgotten — you won’t even notice it thought for a moment.

Conversation history recall · Three-layer retrieval (literal / fuzzy / semantic) · Fully local & offline · Compaction-proof

  • While searching, a quiet sweeping light appears in the corner:

    Recalling

  • When done, no trace:

    Recall complete

Who is it for

  • Heavy users of long sessions — conversations spanning days and hundreds of turns, too long to scroll back through
  • Writers / RP / tavern players — settings, foreshadowing, and character relationships scattered across months of chat
  • Code & doc maintainers — the reasoning behind past decisions and pitfalls, now reduced to a one-line summary
  • Anyone who has said “didn’t we discuss this before?” — it brings back the original words instead of making you retell them

Conversely, if your sessions are short and easy to scroll, you probably don’t need it — it is built for “history too long, memory compacted” scenarios.

Quick start

dsh plugin --profile web add dsh-recall@0.2.2

One command: the package ships its own composition patch (bundle layer), so the plugin and the search index it needs are wired up automatically. Restart dsh web. Nothing else to do — the model ships with the package (~37MB full install), the index builds on first search, and semantic warm-up finishes quietly in the background (a few minutes, imperceptible to you).

You can also install / disable / uninstall dsh-recall from the Add-ons block of the Plugin Management tab in dsh-extension-hub.

Install from source (git clone):

dsh plugin --profile web add git+https://github.com/Relistencode/dsh-recall.git

The repository tracks the model (models/model_merged.onnx) and the vendored runtime, so a git install is fully functional offline with no build step and no allowBuilds entry. The optional dsh-recall-models dependency is still attempted from npm; if it fails to resolve, the in-repo model is used instead — either way the semantic layer works. The harness resolves all paths through $DSH_HOME (default ~/.dsh), so this works identically regardless of where your harness home lives.

Optional configuration
- id: recall
  name: dsh-recall
  config:
    semantic: false   # disable the semantic layer (literal + fuzzy only, smaller package)
    warmup: gentle    # slower warm-up, lower background CPU (only during warm-up; zero afterwards)

What it is not

  • ❌ Not context engineering — it does not cram history into the model window
  • ❌ Not prompt engineering — it does not rely on prompts to make the model “pretend to remember”
  • ❌ Not a memory-document system — no MEMORY.md or manual notes to maintain
  • ✅ It is actual recall: on-demand retrieval of the original records — including history already compacted away (compaction only summarizes; the original text stays searchable forever)

Capability overview

Capability Implementation
Three-layer hybrid retrieval Literal / fuzzy / semantic merged automatically with a coverage gate (≥90%) and a silent degradation chain
Progressive disclosure Light coarse recall by default (titles + snippets + events, ~100–800 tokens); detail drills into the original text — hit list / exact window / paged browsing
Event aggregation Repeated mentions of one topic merge into events ([startSeq..endSeq], ≤5 text blocks apart) — one complete episode instead of scattered fragments; the full event text is one detail browse away
Proactive recall The agent recalls on its own when needed (after compaction, when details are missing); explicit user requests also work
Compaction anchor On compaction/summary, one lightweight anchor is injected automatically (summary + key original fragments, expires after 3 turns)
Scope control Current session only by default; workspace / all only on explicit user request
Compaction-proof Index covers the full history, including shadowed (compacted) events
Incremental indexing Live sessions via ctx.sessions, persisted via sessionPersistence, append-only deltas
Background warm-up Worker-thread embedding (~10 texts/sec), host event loop never blocked
Invisible UI “Recalling…” sweep → one quiet “Recall complete” line; results never enter the UI, the agent presents them naturally
Fully local & offline Zero npm runtime dependencies; no external model APIs; works with no network at all

Architecture

dsh-recall architecture: turn lifecycle on top, capability layers below

  • Turn lifecycle (top): one recall is a straight line — the user asks, the agent calls the recall tool, the three layers are searched, hits are grouped into events per session, and the agent receives either a light coarse recall or a drill-down window, depending on what it needs.
  • Retrieval layer: three independent retrieval channels (literal / fuzzy / semantic) that merge under a coverage gate (see Three-layer hybrid retrieval).
  • Index & data: everything is read through official services (ctx.sessions / ctx.sessionPersistence / ctx.sessionQuery) — no .zstd parsing, no private formats. The plugin’s own recall-index.db (SQLite) holds the fuzzy index, the vectors and the trigram FTS.
  • Governance & scope: the scope red line (session by default), the coverage gate, the degradation chain, and the token budget live here.
  • Automatic layer: a compaction/summary listener that turns every compaction into one lightweight anchor, so the agent keeps its bearings after history is folded away.

Core mechanisms

Three-layer hybrid retrieval
Layer Technique Covers
Literal Official FTS5 full-text index Exact keyword matches
Fuzzy Self-built trigram + char-bigram index (zero dependencies) Rough wording, remembered fragments, typos / missing chars
Semantic Local bge-small-zh model (int8, 24MB, bundled) Paraphrase, word substitution, “roughly what it was about”

Three-layer retrieval: query fans out to literal/fuzzy/semantic, merges through the coverage gate

  • The fuzzy layer is the primary path (it already covers the literal layer’s ground with far more tolerance); the official FTS5 layer is the fallback; the semantic layer joins the mix only when it covers ≥90% of the literal/fuzzy hits — otherwise it stays silent rather than dragging the ranking down.
  • Any layer failure degrades silently to the layer below — semantic → fuzzy → literal, never an error. The recall tool always answers.
  • Inference runs in a worker thread (WASM on the main thread would block the host event loop; measured ~9.6 texts/sec with zero main-thread impact).
  • Everything runs locally and offline — no external model APIs, no network.
Progressive disclosure

Recall happens in two stages, and the second stage only fires when the agent actually needs it:

Stage What the agent gets Cost
1 — coarse recall (default) Session titles + snippets + same-topic events, grouped, ranked ~100–800 tokens for up to 10 sessions
2 — detail drill-down A session’s hit list / the exact original-text window (readEvent) / paged browsing ~300 tokens per session (e.g. a ±3 event window)

Measured live on a real instance: coarse recall saved ~80% of tokens versus the old full-context windows (2500–3000 → ~600 on a 10-session hit, pre-aggregation). Event aggregation keeps the same discipline — snippets only, full event text one drill-down away — so a coarse call stays under ~800 tokens. Irrelevant content never enters the context — and when it matters, the original text is always one drill-down away.

Compaction anchors

Compaction is where memories get lost — the harness summarizes, the original text is shadowed. dsh-recall listens for compaction/summary and immediately injects one lightweight anchor into the compacted session:

  • Content: the LLM summary + up to 3 key original fragments (user messages first, then longest text blocks).
  • Expiry: after 3 assemblies, the anchor disappears — it is a bearing, not a crutch.
  • Escape hatch: the exact original text stays one detail drill-down away, always.
  • Verified end-to-end on a live instance: a real /compact produced the anchor in the very next assembly, with the correct content, expiring automatically after 3 turns.
Scope & privacy
  • Default scope is the current session only — cross-session (workspace) and cross-project (all) searches happen only on the user’s explicit request.
  • The reply layer is invisible: a quiet “Recalling…” sweep, one “Recall complete” line, nothing else. Results never enter the UI — the agent presents them naturally.
  • Data stays on this machine: no external APIs, no telemetry, no network.

Measured

Token benefit (live, v0.2.1)
Measurement Result
Coarse recall cost (default) ~100–800 tokens per call
Old full-context windows (10 sessions) ~2500–3000 tokens — 3–4× more
detail ±3 window ~300 tokens per session
Compaction anchor Verified live: real /compact → anchor injected next assembly, correct content, auto-expires after 3 turns
Semantic warm-up ~10 texts/sec in a worker thread, host event loop zero-blocked
Retrieval quality (golden set)

Synthetic 4-session corpus (32 docs) with 23 hand-annotated queries (exact / fuzzy-typo / paraphrase / cross-session), run in-memory with the real model — repro: node eval/run-golden.mjs:

Variant recall@5 MRR nDCG@10
Literal only (simulated official FTS5) 0.196 0.217 0.201
Fuzzy only 0.587 0.652 0.579
Semantic only 0.533 0.609 0.529
Hybrid (production path) 0.696 0.761 0.687
  • The hybrid merge beats every single layer (+19% recall@5 over the best solo layer) — all three layers contribute, none is decoration.
  • The literal layer alone is the weakest (exact match only; FTS5 unicode61 is word-splitting-blind for Chinese) — confirming its fallback role.
  • The fuzzy layer is the primary path (beats semantic solo); the semantic layer adds recall on paraphrase and word-swap queries.
  • Coverage gate verified: at half warm-up the gate correctly falls back to fuzzy-only (0.587 = fuzzy-only); running the half-warmed semantic layer anyway yields a small gain on this small corpus (0.674) — the 0.90 gate is a conservative safety default for real long sessions, not tuned to this set.
  • Known misses (documented boundaries): zero literal-overlap paraphrases below the semantic min-score (e.g. “打码” for “脱敏”) and abstract-concept queries (e.g. “方案”).

Recent updates

常见问题常见问题

在启用了 DSH 的终端中执行已验证命令 dsh plugin --profile default add github:Relistencode/dsh-recall。命令会解析公开 package 元数据,并保持插件与本页展示的目录身份一致。