zhtx2024/dsh-pdf
PDF parsing toolkit for DeepSeek Harness: pdf_info, pdf_extract_text, pdf_render_page with hybrid pdfjs/built-in engines and system-font rendering for non-embedded CJK PDFs.
已收录
1
Docs
Bundle 已验证
功能介绍
DSH 的 PDF 解析工具:pdf_info、pdf_extract_text、pdf_render_page 三件套,pdfjs 与自研渲染器双引擎,未内嵌 CJK 字体的 PDF 也能用系统字体渲染出图。
适合
- 需要 PDF 元数据、字体信息、逐页文本提取或 PNG 页面渲染的 DSH 工作流。
- 处理 EasyEDA 或嘉立创 EDA 原理图导出文件,且其中未内嵌 CJK 字体难以由 pdfjs 单独处理的工程师。
- 需要可选行级文本坐标和字号信息的文档分析场景。
不适合
- 对扫描版或纯图片 PDF 做文本提取的工作流;这类文件没有文本层。
- 使用特殊 CID 字体、且不符合系统字体渲染器 Unicode 码点假设的文档。
- 在内存受限环境中处理大型 PDF;文件会被完整解析到内存。
- 需要 PDF 编辑、表单填写或 OCR,而非解析和渲染的用户。
README
@zhtx2026/dsh-pdf
PDF parsing toolkit for DeepSeek Harness — three agent-facing tools (
pdf_info/pdf_extract_text/pdf_render_page) with hybrid engines, including system-font rendering for PDFs with non-embedded CJK fonts (e.g. EasyEDA / JLC EDA schematic exports).
| English | 中文 |
Tags: dsh-plugin · pdf · pdfjs · canvas · schematic · tools
Install
dsh plugin --profile web add @zhtx2026/dsh-pdf
Or from a local checkout during development:
dsh plugin --profile web add link:<path-to-repo>
Then restart DSH (or open a new session) — the three tools and an agent guidance block become available.
Tools
| Tool | What it does |
|---|---|
pdf_info |
Page count, page sizes (pt), document metadata, and the font list with per-font embedded flags. |
pdf_extract_text |
Extracts text per page; optional page and withPosition (per-line x/y coordinates + font size). |
pdf_render_page |
Renders one page to a PNG (local file) with configurable scale/outDir; engine auto-selection. |
Why a custom renderer?
pdfjs-dist cannot render text in PDFs whose fonts are not embedded and lack a
ToUnicode map — the classic case is JLC EDA / EasyEDA schematic exports
(SimSun/SimHei Type0 fonts with UniGB-UCS2-H encoding): pdfjs emits
translateFont failed and drops the glyphs.
dsh-pdf solves this with a two-engine design:
- pdfjs engine — used when all fonts are embedded (standard PDFs).
-
sysfont engine — a built-in content-stream renderer that decodes the
PDF text operators (
BT/ET,Tf,Tm,Td,Tj,TJ) itself, mapsUniGB-UCS2-H/UTF-16BE strings to Unicode, and draws them with system fonts (SimSun,SimHei,Microsoft YaHei, …) via@napi-rs/canvas.
Text extraction uses the same auto-selection: embedded fonts → pdfjs text layer; non-embedded fonts → the built-in decoder (which correctly recovers CJK text pdfjs loses).
Example
pdf_info("D:/Downloads/SCH_SA-V11A.pdf")
→ 3 pages, A3 landscape, jsPDF producer, 6 fonts (all non-embedded)
pdf_extract_text("D:/Downloads/SCH_SA-V11A.pdf", { page: 1 })
→ net names, pin numbers, and Chinese labels ("主控板", "嘉立创", …)
pdf_render_page("D:/Downloads/SCH_SA-V11A.pdf", { page: 1, scale: 2 })
→ D:/Downloads/SCH_SA-V11A-pages/page-1.png (2396×1698, engine: sysfont)
Architecture
-
lib/index.js— cordis plugin entry (name,inject,apply); registers the three tools as native tool objects (plain JSON-Schemaparametersandoutput— no@deepseek-ai/dsh-toolsdependency needed) and injects the agent guidance block (systemPrompt.section, order 160). -
lib/pdf-core.js— self-contained engine: PDF object parser, font scanner, pdfjs loading with fs-basedCMapReaderFactory/StandardFontDataFactory(Node 24’s globalfetchdoesn’t speakfile://), the built-in text extractor, and the sysfont renderer. -
cordis.patch.yml— thedsh.bundle.patchinsert row.
Limitations
- Scanned/image-only PDFs have no text layer — extraction returns nothing (rendering still works, as a plain image).
- The sysfont renderer assumes
Type0CID codes equal Unicode code points (true forUniGB-UCS2-Hexports); exotic CID fonts fall back to pdfjs. - Large PDFs are parsed fully into memory.
License
常见问题常见问题
在启用了 DSH 的终端中执行已验证命令 dsh plugin --profile default add github:zhtx2024/dsh-pdf。命令会解析公开 package 元数据,并保持插件与本页展示的目录身份一致。
兼容性以页面上展示的 bundle 与 profile 状态为准。如果某个 profile 尚未检测到,请先保持禁用,并在生产启用前阅读仓库文档。
GitHub 链接和 activity 元数据是 release 与维护状态的来源。新版本发布后重新查看本页,确认目录已经观察到最新版本。