dsh-tool-ocr
已验证 · 实测可装 ferstar
功能简介
本地 OCR 插件:让纯文本生成 LLM 也能读懂图片 | Local OCR plugin: give text-only generative LLMs the ability to read images
可用 — 实测通过,早期项目
本地 OCR 插件:让纯文本生成 LLM 也能读懂图片 | Local OCR plugin: give text-only generative LLMs the ability to read images 实测能干净安装、正常启动。早期项目,但功能可用。
「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。
README
dsh-tool-ocr
English | 中文
Local image text recognition for DeepSeek Harness models without vision input, backed by the standalone newbee-ocr (nbocr) engine over PP-OCRv6 models. Fully out-of-tree: depends only on published dsh base packages.
Features
| Item | Description |
|---|---|
ocr tool |
recognize / status / install / check actions |
| Image inputs | Local path, or attachment_id (an image already in the conversation) |
| Output | Reading-ordered text, per-block bounding boxes, review flags (low confidence / amounts / numbers / dates / quantities), heuristic Markdown tables |
| Robustness | Survives MNN diagnostics in engine stdout; honors cancellation and timeout |
| Settings card | On dsh ≥ 0.1.0-rc.7 the plugin registers the tool-ocr namespace: the web Settings → Plugins page renders a card that edits the engine command, model tier, and run bounds live — no composition edit needed |
Quick start
1. Install the engine — grab nbocr from the newbee-ocr-cli releases (one-line installer or prebuilt archive):
curl -LsSf https://github.com/zibo-chen/newbee-ocr-cli/releases/latest/download/newbee_ocr_cli-installer.sh | sh
# Windows PowerShell: irm https://github.com/zibo-chen/newbee-ocr-cli/releases/latest/download/newbee_ocr_cli-installer.ps1 | iex
2. Install the plugin
dsh plugin --profile web add dsh-tool-ocr
# or: cd ~/.dsh/profiles/web && pnpm add dsh-tool-ocr
3. Mount it in your profile's cordis.patch.yml:
- insert:
- id: ocr
name: 'dsh-tool-ocr'
inject: [tools, subprocess, systemPrompt]
config:
command: 'C:/path/to/nbocr.exe' # required — the nbocr executable
That's it — restart dsh, then point the model at an image: ocr { path: "C:/screenshot.png" }. To tune recognition (language, model tier, limits), open Settings → Plugins → Plugin config and edit the OCR card; changes apply on save, no restart needed. command is the only field that must come from the composition.
Image inputs
pathworks on every dsh build: the model reads the file from disk.attachment_idneeds a build that admits images for text-only models (e.g. the dsh fork withapi-gateway.allowImagePlaceholder: true): images enter the session, the llm layer substitutes[image attachment <id>]text, and the model hands the id toocr.
Configuration
| Field | Default | Meaning |
|---|---|---|
command |
— (required) | The nbocr executable: absolute path or PATH-resolved name. |
args |
[] |
Extra arguments before the nbocr subcommand (no shell). |
env |
{} |
Extra environment entries. |
language |
chinese |
Recognition model/language alias. |
detModel |
v6-tiny |
Detection model tier. |
modelsDir |
'' |
Model directory for non-embedded models; empty uses embedded models. |
maxImageBytes |
26214400 |
Largest accepted image in bytes. |
maxOutputBytes |
2000000 |
Largest collected engine stdout in bytes. |
maxTextChars |
12000 |
Largest recognized text returned in text. |
timeoutMs |
600000 |
Tool-call timeout budget in ms. |
Tool usage
ocr { path | attachment_id, action?, include_boxes?, table?, max_text_chars? }
status probes readiness without engine work; check runs a real end-to-end 1x1 probe.
Development
pnpm typecheck # tsc --noEmit
pnpm test # vitest
pnpm lint # oxlint
Real-engine smoke (skipped unless both env vars are set):
NBOCR_BIN=/path/to/nbocr OCR_E2E_IMAGE=/path/to/image.png pnpm test
License
MIT
安装
装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:
dsh plugin add dshbase-catalog 然后对 agent 说「帮我装 dsh-tool-ocr」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。
该插件是 GitHub 源码(未发 npm)——直接从仓库装:
Web profile:
dsh plugin --profile web add github:ferstar/dsh-tool-ocr Headless(CLI)profile:
dsh plugin --profile headless add github:ferstar/dsh-tool-ocr 实测报告
验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。
使用场景
给模型装上眼睛——图像理解、OCR 或屏幕定位——让它读视觉而非靠猜。
适合谁
会把截图、图表或照片交给模型、想被原生理解的人。
二次开发建议
视觉后端和预处理是缝——加 OCR、区域裁剪,或调分辨率和模型路由。