dsh-files
已验证 · 实测可装 taxueseek
功能简介
DeepSeek Harness双面插件:会话隔离文件上传,彩色卡片+文档读取工具,含嗅探与LRU缓存
可用 — 实测通过,社区增长中
DeepSeek Harness双面插件:会话隔离文件上传,彩色卡片+文档读取工具,含嗅探与LRU缓存 实测能干净安装、正常启动。社区在增长,是个稳妥选择。
「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。
README
dsh-files
A DeepSeek Harness plugin that fills one official gap per stage of a file's session lifecycle:
- Ingest: the folder button next to the native paperclip (plus a same-named entry in the official "+" command menu) — the browser flattens the directory (Office lock files,
.DS_Store,.envand other system/hidden files are filtered), and every file enters the official native attachment pipeline - Read: the
read_documenttool — structured text extraction for binary documents (PDF / DOC / DOCX / XLSX) plus enhanced text reading (encoding fallback, paging, sheet-level access) - Manage: the
attachment_list/export_attachmenttools — make the attachment store visible to the model (name/size/sha) and copy a file into the workspace for read/edit/bash to work on - Fetch back: the attachment dock (one official pill below the composer card) plus download/export routes and an
@attachment source — the store becomes visible to users and files can be pulled into the browser (the right path for remote/LAN deployments); the@menu inserts the official handle line, identical to what the model saw at upload
Upload, images and
@reference were removed in 0.5.0 — harness 0.1.3 ships them natively (universal file upload, the image vision pipeline, unified@file/@sessionreference), and does it better. This plugin is part of the taxueseek plugin matrix; the flagship is argo.
Why it exists
Native upload in harness 0.1.3 stores files as byte objects and hands the model one handle line (name, size, digest, read-only path) to read with file tools — but the built-in read tool rejects binary content with FS_NOT_TEXT. Structured text extraction for PDF / DOC / DOCX / XLSX, plus attachment-store listing and export (the official GC is on the roadmap and the store is invisible to the model today), are the gaps the official stack leaves open; this plugin fills them.
Capabilities
- Content sniffing: PDF header / OLE Compound File (Word 97-2003) / ZIP central-directory members / UTF-8 (fatal) / UTF-16 BOM / GB18030 — decided from bytes, never from extensions; disguised files (an exe renamed .pdf) are rejected. The format hint is only a last resort when bytes are fully unknown
- Legacy .doc: macOS uses the system
textutil(most complete body and date lines in the gold-standard comparison), other platforms fall back to pure-JSword-extractor - Encoding chain: UTF-16 BOM → UTF-8 (fatal, NUL rejected) → GB18030 (fatal) → UTF-16 without BOM (high-confidence guard); GBK Chinese and BOM-less UTF-16 both read
- Paged reads: line numbers + offset/limit; the per-call character budget differs by format (text full, xlsx 3/4, pdf/doc/docx 1/2), overflow truncates with an explicit remaining-lines marker
- Line-number policy: text (code/config) carries line numbers for precise edits; PDF/DOC/DOCX/XLSX are paragraph flows without line numbers (saves tokens)
- XLSX sheet-level reads:
list_sheetsnames the sheets, thesheetparameter reads one sheet in full (no row cap), out-of-range errors list the available sheets
dsh-files
A DeepSeek Harness plugin that fills one official gap per stage of a file's session lifecycle:
- Ingest: the folder button next to the native paperclip (plus a same-named entry in the official "+" command menu) — the browser flattens the directory (Office lock files,
.DS_Store,.envand other system/hidden files are filtered), and every file enters the official native attachment pipeline - Read: the
read_documenttool — structured text extraction for binary documents (PDF / DOC / DOCX / XLSX) plus enhanced text reading (encoding fallback, paging, sheet-level access) - Manage: the
attachment_list/export_attachmenttools — make the attachment store visible to the model (name/size/sha) and copy a file into the workspace for read/edit/bash to work on - Fetch back: the attachment dock (one official pill below the composer card) plus download/export routes and an
@attachment source — the store becomes visible to users and files can be pulled into the browser (the right path for remote/LAN deployments); the@menu inserts the official handle line, identical to what the model saw at upload
Upload, images and
@reference were removed in 0.5.0 — harness 0.1.3 ships them natively (universal file upload, the image vision pipeline, unified@file/@sessionreference), and does it better. This plugin is part of the taxueseek plugin matrix; the flagship is argo.
Why it exists
Native upload in harness 0.1.3 stores files as byte objects and hands the model one handle line (name, size, digest, read-only path) to read with file tools — but the built-in read tool rejects binary content with FS_NOT_TEXT. Structured text extraction for PDF / DOC / DOCX / XLSX, plus attachment-store listing and export (the official GC is on the roadmap and the store is invisible to the model today), are the gaps the official stack leaves open; this plugin fills them.
Capabilities
- Content sniffing: PDF header / OLE Compound File (Word 97-2003) / ZIP central-directory members / UTF-8 (fatal) / UTF-16 BOM / GB18030 — decided from bytes, never from extensions; disguised files (an exe renamed .pdf) are rejected. The format hint is only a last resort when bytes are fully unknown
- Legacy .doc: macOS uses the system
textutil(most complete body and date lines in the gold-standard comparison), other platforms fall back to pure-JSword-extractor - Encoding chain: UTF-16 BOM → UTF-8 (fatal, NUL rejected) → GB18030 (fatal) → UTF-16 without BOM (high-confidence guard); GBK Chinese and BOM-less UTF-16 both read
- Paged reads: line numbers + offset/limit; the per-call character budget differs by format (text full, xlsx 3/4, pdf/doc/docx 1/2), overflow truncates with an explicit remaining-lines marker
- Line-number policy: text (code/config) carries line numbers for precise edits; PDF/DOC/DOCX/XLSX are paragraph flows without line numbers (saves tokens)
- XLSX sheet-level reads:
list_sheetsnames the sheets, thesheetparameter reads one sheet in full (no row cap), out-of-range errors list the available sheets
Install
Requires harness ≥ 0.1.3-alpha.1.
curl -fsSL https://raw.githubusercontent.com/taxueseek/dsh-files/main/install.sh | sh
# restart dsh web
Manual equivalent:
dsh plugin --profile web add git+https://github.com/taxueseek/dsh-files.git
# restart dsh web
The npm package named
dsh-filesis an unrelated third-party placeholder — install only via the script or the git command above.
Compatibility
| dsh-files | Harness | Notes |
|---|---|---|
| 0.5.3 | 0.1.7-alpha.1 (verified) | Current. Pins @deepseek-ai/dsh-fs/dsh-tools/dsh-client-ui-primitives at 0.1.7-alpha.1; client icons follow the host's Regular/Medium naming. |
| 0.5.x | ≥ 0.1.3-alpha.1 | Older SDK pins (0.1.0-rc.x); attachment dock and @ source predate the host's conversation.composer.dock slot. |
| 0.6.x | — | Never released (folded into 0.5.2/0.5.3); do not use. |
The plugin targets the alpha line the maintainer runs locally (0.1.7-alpha.1); npm latest (0.1.5-rc.3 at the time of writing) is older, so prefer the git install above over any registry version.
Configuration
- id: files-toolkit
name: 'dsh-files'
config:
maxFileBytes: 25165824 # byte cap for one document read
readLimit: 2000 # lines returned per call (paging is cheap)
sheetRowLimit: 200 # rows kept per worksheet
maxSheets: 5 # sheets read per workbook
maxOutputChars: 24000 # per-call window character budget (truncated with a marker)
readTimeoutMs: 120000 # per-call timeout (raise for huge PDFs)
# attachmentsDir: /path/to/attachments/v1 # attachment store root; empty = DSH_HOME / ~/.dsh autodetect
attachmentsEnabled: true # master switch for the attachment loop (dock/download/export/@ source)
maxDownloadBytes: 209715200 # per-download/export byte cap (answers 413)
trustedHosts: [] # non-loopback host[:port] authorities; required for LAN/domain (same semantics as --trusted-host)
Security
- Parsing dependencies are read-only and maintained:
pdfjs-dist(Mozilla),mammoth,read-excel-file,word-extractor(.doc fallback) - ZIP central-directory probing never expands members; malicious archives are rejected safely
- Reads and export destinations go through
ctx.fs, inheriting the session sandbox, same rights as the built-in read tool; the attachment-store scan is a host-side read-only walk with internally-constructed paths - Attachment download/export flows through the official
AttachmentStore.readFileStream(integrity check, no absolute paths), on top of the Host trust fence +sha256:reference whitelist + size cap; there is no delete route (content-addressed objects may be referenced by historical messages; deletion stays with the official future retention) - The dock and
@source are UI-layer data: no systemPrompt injection, no model tools, zero tokens
Development
pnpm install
pnpm test
pnpm build
npx tsc --noEmit
License
MIT
安装
装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:
dsh plugin add dshbase-catalog 然后对 agent 说「帮我装 dsh-files」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。
该插件是 GitHub 源码(未发 npm)——直接从仓库装:
Web profile:
dsh plugin --profile web add github:taxueseek/dsh-files Headless(CLI)profile:
dsh plugin --profile headless add github:taxueseek/dsh-files 实测报告
验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。
使用场景
给 agent 持久存储——数据库、文件存储或持久层——让状态跨会话留存。
适合谁
任务需要读写结构化数据并跨运行保留的人。
二次开发建议
存储后端和数据模型是缝——插新数据库、加 schema 或暴露查询工具。