dshbase

插件目录 / Developer / dsh-deepseek-vision

dsh-deepseek-vision

已验证 · 实测可装 Argonaut790

✓ 持续维护 基于 17 个官方 DSH 包 纯 TypeScript

查看 GitHub ↗ ← 返回插件目录

3Stars
0Forks
1未关闭 issue
TypeScript语言
2026-08-14最近推送
跨平台平台

功能简介

图像理解、OCR、持久视觉证据

✅
我们的评价
可用 — 实测通过,早期项目

图像理解、OCR、持久视觉证据 实测能干净安装、正常启动。早期项目,但功能可用。

「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

README

DSH DeepSeek Vision

CI
License: MIT

DSH DeepSeek Vision is an open-source DeepSeek Harness (DSH) vision plugin
that adds image understanding, full-screen OCR, and persistent visual evidence
to text-only DeepSeek models without replacing the parent model.

Unlike provider-pool or CLI interception tools, this plugin keeps DeepSeek
Harness in charge of models, attachments, sessions, and UI. It adds:

  • see_image with latest, all, and explicit image selection
  • one conversation-scoped vision analyst with follow-up memory
  • structured summaries, question answers, exhaustive OCR, and uncertainties
  • a read-only Evidence tab and per-call evidence cards
  • a global Vision: … provider/model picker beside Choose Model
  • live route changes; changing the route starts a new analyst

Screenshots

Vision-enabled DeepSeek Harness composer

DeepSeek Harness composer using Grok Latest as the vision model and DeepSeek V4 Pro High as the parent model

The parent DeepSeek model stays in control while the separate Vision route
handles image understanding and OCR.

Compact vision model selector

Compact DSH Vision model selector beside the DeepSeek parent model selector

The global selector makes the active image-capable model visible and lets users
change the visual-analysis route without changing the conversation model.

Evidence card in a conversation

Vision analysis reply with an evidence card in a DeepSeek Harness conversation

Each see_image call renders an evidence card with the structured summary,
question answers, and any uncertainties, so the analysis stays reviewable in
the conversation.

GitHub project overview

The open-source DSH DeepSeek Vision repository on GitHub

Requirements

  • Node.js ^22.19.0 or >=24
  • DeepSeek Harness 0.1.0-rc.6
  • an image-capable model registered in the Harness catalog
  • the DSH spawn subagent provider

The Harness must provide delegated-image prompt admission, model input
modalities, the see-image-model settings namespace, and the Web conversation
slots. This plugin cannot retrofit those contracts into an older release.

Do not mount this package while equivalent in-tree see-image-model,
tool-subagent-image, or vision-picker rows are enabled. Duplicate services
and tools will conflict.

Install from GitHub

This project is not published to npm. Build a checkout and add that local
package to the Web profile:

git clone https://github.com/Argonaut790/dsh-deepseek-vision.git
cd dsh-deepseek-vision
corepack yarn install --frozen-lockfile
corepack yarn build
dsh plugin --profile web add .

The included cordis.patch.yml mounts the global route service and
see_image; its package metadata exposes the Web picker and Evidence UI.

Configure

Open a conversation and select an image-capable route from the Vision: …
chip. Models are listed only when the Harness catalog explicitly declares
image input.

For a headless profile, configure the same global route in
$DSH_HOME/settings.yaml:

see-image-model:
  provider: openrouter
  model: '~x-ai/grok-latest'
  maxTokens: 8192

The provider and model names are examples. They must match routes registered
in your Harness. The supported output-token range is 1–32768.

An optional static fallback may be set on the tool row:

- id: deepseek-vision-tool
  name: dsh-deepseek-vision/tool
  config:
    provider: spawn
    agentOptions:
      provider: openrouter
      model: '~x-ai/grok-latest'
      maxTokens: 8192

The global picker takes precedence when it contains a complete route.

How it works

  1. Harness retains pasted images as durable delegated-image attachments.
  2. The text-only parent calls see_image with questions and an image
    selection.
  3. The plugin reuses the newest matching vision analyst for that conversation,
    forwarding only images the analyst has not already received.
  4. The analyst receives no tools, uses a fixed anti-prompt-injection persona,
    and must return strict JSON.
  5. The parent receives concise model-facing text while the complete structured
    record is retained for evidence cards and the Evidence tab.
  6. If durable continuation is unavailable, the plugin performs an isolated
    one-shot structured readback.

Calls are serialized per conversation by the Harness tool runtime. A route
change creates a new analyst rather than mutating the model behind an existing
child.

Privacy, trust, and cost

  • Selected images are sent to the configured vision provider. Review that
    provider's retention, region, and privacy terms before use.
  • Each analyst turn consumes the selected model's tokens and may incur
    provider charges. Follow-ups can reuse visual context but are still model
    calls.
  • OCR and visual conclusions are model-generated evidence, not guaranteed
    facts. Verify high-impact decisions independently.
  • Text found inside images is treated as untrusted data, never as
    instructions. The analyst has no tools or external-action authority.
  • Evidence records keep attachment identifiers and derived text in the
    conversation history; they do not embed image bytes.

Image selection

see_image supports:

  • latest (default): images from the newest conversation event containing
    delegated images
  • all: the de-duplicated conversation image catalog
  • ids: exact attachment IDs already present in that catalog

A call accepts up to 12 questions, 2,000 characters per question, and 8,000
characters in total.

Development

Use Corepack-managed Yarn:

corepack yarn install --frozen-lockfile
corepack yarn typecheck
corepack yarn build
corepack yarn test

The build emits Host entries at lib/index.js and lib/tool.js, declarations
under lib/types, and a browser __ModuleLoader__ bundle at lib/client.js.
See CONTRIBUTING.md, SECURITY.md, and
CHANGELOG.md.

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 dsh-deepseek-vision」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。

该插件是 GitHub 源码(未发 npm)——直接从仓库装:

Web profile:

dsh plugin --profile web add github:Argonaut790/dsh-deepseek-vision

Headless(CLI)profile:

dsh plugin --profile headless add github:Argonaut790/dsh-deepseek-vision

实测报告

验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。

使用场景

扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。

适合谁

想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。

二次开发建议

工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。

安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7797 个插件 →