JunNanLYS/dsh-layered-memory
让 DeepSeek Harness 拥有长期记忆:对话自动蒸馏为事实/场景/画像三层记忆,每步自动召回注入——让 AI 基于证据说话,零操作无感使用 | Long-term memory for DeepSeek Harness: conversations auto-distilled into facts, scenes & persona, recalled before every step — so the AI speaks from evidence, not guesses. Zero effort.
Listed
1
Memory
Bundle verified
What it does
Layered memory pipeline for DSH: auto-distills conversations into atomic facts, scene summaries and a persona profile (L0–L3), with hybrid BM25 + vector retrieval, chat/work family separation, and context injection before each model step.
Best for
- DSH users who want long-term memory distilled automatically into facts, scene summaries, and a persona profile.
- Mixed personal-chat and work workflows that need separate memory families and per-session memory modes.
- Agents that benefit from automatic hybrid retrieval and relevant context injection before each model step.
Not ideal for
- Environments below Node.js 22.16.
- Privacy-sensitive workflows that should not persist conversations, facts, scenes, profiles, and a local memory database.
- Users who do not want background LLM calls for extraction, consolidation, and persona distillation.
- Workflows requiring branch-aware memory or direct Claude Code/Codex memory import, which are documented as roadmap items rather than current features.
README

dsh-layered-memory
DeepSeek Harness 的分层蒸馏记忆插件:对话在后台自动完成 L0 捕获 → L1 原子记忆 → L2 场景整合 → L3 画像蒸馏,模型每一步前自动把相关记忆注入上下文。
快速开始
需要 Node ≥ 22.16。两种调用方式任选(npx 前缀可替换下面任何 dsh 命令):
# 方式一:npx 直接跑官方 CLI(无需预装 dsh;可 pin 版本,如 dsh-layered-memory@0.8.0)
npx -y @deepseek-ai/dsh plugin --profile web add dsh-layered-memory
# 方式二:已装 dsh CLI(dsh 是 pnpm 转发器,未装 pnpm 时先 npm i -g pnpm)
dsh plugin --profile web add dsh-layered-memory
# 包源备选:GitHub 仓库 / 本地路径(开发调试,link: 指向仓库,npm run build + 重启 dsh 即生效)
dsh plugin --profile web add https://github.com/JunNanLYS/dsh-layered-memory
dsh plugin --profile web add /path/to/dsh-layered-memory
让 Agent 安装(推荐)
如果当前 Agent 可以执行终端命令,把下面这段话完整发送给它:
请为 DeepSeek Harness 的 web Profile 安装 dsh-layered-memory 插件。
只执行下面两条命令,不要修改其他 Profile:
dsh plugin --profile web add dsh-layered-memory
dsh --profile web --dump-config
确认输出中出现 dsh-layered-memory 后告诉我安装结果。
不要替我关闭或重启正在运行的 DSH;安装完成后提醒我手动重启 DSH Web Host。
Agent 应当返回安装结果,并明确告诉你配置中是否已经出现 dsh-layered-memory。
本包声明了 dsh.bundle 组合包层(cordis.patch.yml),安装后会自动挂载插件行——
不需要再手改 $DSH_HOME/profiles/web/cordis.patch.yml。然后重启 DeepSeek Harness,
验证:~/.dsh/memory/ 下出现 conversations/ records/ scenes/ 目录和 memory.db
即插件 apply 成功;设置页出现”记忆”页面、输入栏出现档位 pill 即 client 半边就绪。
卸载:dsh plugin --profile web remove dsh-layered-memory + 重启。数据保留在
~/.dsh/memory/,不需要时手动删除整个目录即可。
从源码开发
git clone https://github.com/JunNanLYS/dsh-layered-memory
cd dsh-layered-memory
npm install && npm run build
dsh plugin --profile web add . # link: 安装,改代码后 npm run build + 重启 dsh 即生效
npm run smoke # 冒烟测试(先重编:见下方命令)
npx tsc src/smoke.ts --outDir dist-smoke --module nodenext --moduleResolution nodenext --target es2022 --strict --skipLibCheck --esModuleInterop
运行时数据流
插件挂在 dsh 原生事件上(session/event 捕获、agent/pre-step 注入),蒸馏调用复用宿主 ctx.llm。召回以消息侧注入呈现:相关记忆作为一条合成消息排在用户新消息之前,会话流里显示为“上下文注入 · memory”行(点开看命中内容)——用户能直接看到”记忆生效了”;注入内容有长度预算与时间预算,超限截断/超时跳过,绝不拖慢对话。
记忆工具(3):
- memory_search
- conversation_search
- memory_read_scene
真机实录:召回注入与工具调用在对话里的样子——”上下文注入 · memory”行先带出相关记忆,模型再按需调 memory_read_scene 读取场景块,凭记忆直接作答:
<img src=”./assets/img/MemoryTools.png” width=”60%” alt=”对话界面实录(浅色主题):用户消息”我们最近要干什么?”上方可见”上下文注入 · memory”行;助手回答前列出 4 次 memory_read_scene 工具调用(参数为 scenes 场景块的 .md 文件名),随后凭记忆梳理近期目标与推进路线”>
在只开放代码执行入口的受限会话中,模型经由 run_code 间接调用记忆工具(轨迹视图中的 SUBTOOL 嵌套):
分层记忆(L0–L3)
会话级记忆档位
-
控件:输入栏内、模式选择器右侧的 pill(
记忆·自动),点击在上方浮出档位滑块深浅主题自适应; - 每会话的选择按 sessionId 持久化到
session-modes.json,重启/恢复会话不丢; 与全局开关叠加(全局是总闸);L2/L3 完全分类,分类内容不渗透。
界面预览
存储布局
向量能力默认关闭(纯 FTS)。DSH 的 ctx.llm 无 embeddings 端点,语义检索由
三态嵌入源提供(关闭 / 远程 / 本地),设置页可运行时切换——见下节。
语义检索(嵌入源)
设置页(记忆 → 概览 → 语义检索)选择嵌入源,即时生效、无需改配置重启:
三种嵌入源:关闭(默认,纯 BM25 关键词检索)、远程(自备任意 OpenAI 兼容
/embeddings 服务,embedding.* 四件套配齐才可选)、本地(内置模型目录选一款,
ONNX 量化 CPU 推理——无需 API Key,数据不出本机)。本地模型目录是插件内置
白名单(每款锁定 revision + 每文件 sha256,不可下载任意仓库)。
-
下载:模型卡一键下载(默认镜像
hf-mirror.com,断点续传 + sha256 完整性 校验;直连不可达时可走代理——默认自动探测HTTPS_PROXY/ALL_PROXY等环境 变量,见embedding.proxy)。单文件失败自动重试且换缓存键(?dshmem-retry=N, 绕开镜像 CDN 偶发的坏缓存对象),校验失配从零重下、网络错误保留断点续传; 落盘数据目录models/<id>/,不用了随时在设置页删除; -
按需运行时:首次切换本地档才安装推理运行时(transformers.js,约 100~200MB,
装进数据目录
runtime/——不进插件依赖树,不碰插件安装目录); - 活切换:一键换源——自动后台全量重嵌(进度可见、可取消,期间检索自动降级 关键词,不影响对话;维度变化时向量表按新维度重建);切换失败保持旧源,重启仍按原源运行;
-
生效规则 = 部署上限 AND 运行时选择:
embedding.allowLocalModels=false可整体 禁用本地档、未配embedding.*四件套则远程档不可选(企业部署可收口),状态持久 化在embedding-source.json。
配置
覆盖配置写在 profile 自己的 cordis.patch.yml,用顶层裸 patch 条目(直接 id:,
不要包在 insert: 里——insert 与 bundle 层同 id 追加会导致 duplicate loader entry id
启动失败):
- id: dsh-memory
name: dsh-layered-memory
config: # 键按行整体替换(不深合并),按需写全要保留的键
family: auto # 新会话默认档:auto | chat | work
llm: # 蒸馏模型静态路由(双字段齐 = 部署 pin,优先于设置页选择;
provider: '' # 留空则跟随设置页"蒸馏模型"选择器或当前默认模型)
model: ''
| 字段 | 默认 | 说明 |
|---|---|---|
family |
auto |
新会话默认记忆档位:auto(双族自动)| chat(个人)| work(工作);会话内可用输入栏控件临时切换 |
dataDir |
$DSH_HOME/memory |
数据目录 |
capture.enabled |
true |
L0 捕获 |
capture.stripCodeBlocks |
true |
助手消息剥离代码块 |
capture.maxMessageChars |
4000 |
单条消息最大字符数 |
extract.enabled |
true |
L1 抽取 |
extract.minMessages |
6 |
稳态触发阈值:单会话攒够 N 条新消息跑一次 L1 抽取。起步阶段生效阈值从 1 翻倍爬坡到此值(首轮即出记忆,随后自动攒批省调用) |
extract.idleSeconds |
300 |
闲置兜底:会话静默 N 秒后把未蒸馏切片落袋(接住”没攒够阈值用户就离开”);0 关闭 |
extract.backgroundMessages |
10 |
抽取时附带的背景消息条数(按会话从 L0 现查,会话间互不污染) |
extract.candidatePool |
5 |
去重候选池大小 |
l2.enabled |
true |
L2 场景整合 |
l2.minNewMemories |
5 |
距上次 L2 整合的新记忆阈值 |
l2.maxScenes |
12 |
场景块数量上限 |
l2.sceneContextLimit |
3 |
L2 prompt 附带的相似场景全文上限 |
l3.enabled |
true |
L3 画像蒸馏 |
l3.interval |
20 |
L3 蒸馏间隔(新记忆条数) |
recall.enabled |
true |
自动召回 |
recall.maxResults |
5 |
每条新用户消息前注入的 L1 条数上限 |
recall.maxCharsPerMemory |
500 |
单条注入记忆的字符上限(超限截断并提示用记忆工具查全文);0 不限 |
recall.maxTotalRecallChars |
2000 |
整轮注入总字符上限(超限按相关性丢尾部);0 不限 |
recall.timeoutMs |
5000 |
召回总预算(ms):超时跳过本轮注入、不阻塞对话;0 不限时 |
recall.includePersona |
true |
系统提示注入画像上下文(<user-persona>,稳定区) |
recall.includeSceneNav |
true |
系统提示注入场景导航(<scene-navigation>,稳定区) |
recall.strategy |
hybrid |
检索策略:keyword / embedding / hybrid
|
recall.scoreThreshold |
0.3 |
召回分数阈值(低于不注入;仅 keyword/embedding 策略生效,hybrid 融合前不过滤;工具路径不过滤) |
embedding.enabled |
false |
向量检索开关;关闭即纯 FTS 运行 |
embedding.baseUrl |
空 | OpenAI 兼容 /embeddings 地址(如 https://api.siliconflow.cn/v1) |
embedding.apiKey |
空 | API Key |
embedding.model |
空 | embedding 模型名 |
embedding.dimensions |
0 |
向量维度(启用时必填,须与模型输出一致) |
embedding.maxInputChars |
5000 |
单条文本最大字符数(超长截断) |
embedding.timeoutMs |
10000 |
单次 embedding 调用超时(ms) |
embedding.allowLocalModels |
true |
允许本地嵌入档(部署上限:关闭后设置页不能下载模型、不能切本地档) |
embedding.mirror |
https://hf-mirror.com |
本地模型下载镜像根地址(可改回官方 https://huggingface.co) |
embedding.proxy |
'' |
模型下载代理三态:''(默认)= 自动探测代理环境变量(HTTPS_PROXY/ALL_PROXY 等,尊重 NO_PROXY);none = 禁用强制直连;其他值 = 代理 URL(如 http://127.0.0.1:7890)。镜像直连在国内网络间歇不可达(直连超时与污染字节交替出现过),开代理的机器建议保持默认自动探测 |
llm.provider/model |
空 | 蒸馏模型静态路由(部署 pin):provider 与 model 双字段齐时锁定蒸馏路由,优先于设置页的运行时选择与默认模型(部署可强制蒸馏走指定路由);留空则跟随”设置页选择 → 默认模型”。运行时可在设置页 → 记忆 → 概览的”蒸馏模型”选择器从已配置的供应商(含 dsh 设置 → 模型里添加的自定义供应商)中切换,即时生效无需重启 |
llm.maxTokens |
65536 |
未分层调用的兜底输出总闸。各蒸馏层有独立预算(抽取 16k / 去重 8k / L2 32k / L3 16k;思考档 high/max 时自动 ×4,防 reasoning 吃光预算),分层预算可在设置页 → 记忆 → 概览 → 蒸馏参数运行时调整(留空/0 = 跟随内置默认) |
llm.reasoningEffort |
off |
蒸馏思考档位(部署默认):off / high / max,空串不传(跟随模型默认)。蒸馏是结构化抽取任务,默认关思考——推理模型(如 v4-flash)默认 high 档的思考可把任意输出预算全部吃光导致正文 0 字符;非推理模型不认识 effort 时需设为空串。运行时可在设置页 → 记忆 → 概览临时切换(选”跟随配置”即回退本值) |
llm.temperature |
0.3 |
蒸馏温度 |
llm.maxInputChars |
700000 |
单次蒸馏输入字符预算(超限的 L1 输入自动分块抽取);运行时可在设置页 → 蒸馏参数 → 输入预算调整(留空/0 = 跟随本值) |
llm.timeoutMs |
120000 |
单次蒸馏调用超时(ms) |
tools |
true |
是否注册模型可调用的记忆工具 |
与 MemoryCore 的差异
- 内嵌完整管线(不依赖外部 Gateway),蒸馏复用 DSH 自己的 LLM;
- L2/L3 由”LLM 操作文件工具”改为”LLM 输出操作 JSON / 完整文档,工程侧执行”;
- 召回注入点在
agent/pre-step(消息侧合成消息,官方 pre-step 替换语义)+ agent 作用域systemPrompt.context(画像/导航稳定区,DSH 原生事件/服务); - 存储/检索即官方 sqlite 后端的单机裁剪版(裁掉多租户隔离列、TCVDB 云后端、审计表; 分词与官方一致用 jieba——@node-rs/jieba 预编译二进制 + CJK 二元组并集, 词元供 BM25 精确整词命中、二元组保子词召回;加载失败自动回退纯二元组, FTS 索引按分词器版本戳自动重建)。
路线图
以下为规划中的功能,欢迎在 Issues 反馈需求与优先级:
- Git 分支感知:记忆与当前 git 分支关联,召回可按分支过滤/加权(与现有记忆档位正交)
-
Claude Code / Codex 记忆导入:一键迁移既有记忆资产(
CLAUDE.md、Claude Code 记忆文件、CodexAGENTS.md等),导入后进入分层蒸馏管线
致谢
记忆核心能力(分层蒸馏管线、Prompt 设计、双写存储架构)参考自 TencentCloud/TencentDB-Agent-Memory 项目中的 MemoryCore,感谢原项目开放的设计与实现。
License
Frequently Asked QuestionsFAQ
Use the verified command dsh plugin --profile default add github:JunNanLYS/dsh-layered-memory in a DSH-enabled shell. The command resolves the public package metadata and keeps the plugin attached to the catalog identity shown on this page.
Compatibility follows the bundle and profile status shown above. If a profile is not detected, keep the plugin disabled there and check the repository documentation before enabling it in production.
The GitHub link and activity metadata are the source of truth for releases and maintenance. Revisit this page after a new release to confirm the catalog has observed the latest version.