dshbase

插件目录 / Developer / dsh-plugin-describe-image

dsh-plugin-describe-image

已验证 · 实测可装 whitelonng

✓ 持续维护 基于 8 个官方 DSH 包 纯 TypeScript

查看 GitHub ↗ ← 返回插件目录

6Stars
2Forks
1未关闭 issue
TypeScript语言
2026-08-15最近推送
跨平台平台

功能简介

describe_image插件:通过OpenAI兼容VLM端点赋予纯文本模型视觉

✅
我们的评价
可用 — 实测通过,早期项目

describe_image插件:通过OpenAI兼容VLM端点赋予纯文本模型视觉 实测能干净安装、正常启动。早期项目,但功能可用。

「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

README

dsh-plugin-describe-image

English | 中文

DeepSeek Harness 图片理解插件 — a vision-language describe_image tool that gives a text-only model (DeepSeek V4 and friends) the ability to understand images.

A DeepSeek Harness plugin: the model-facing describe_image tool. It loads one image — a local file path, an http(s) URL, or a durable attachment reference — and asks a vision-language model (VLM) at an OpenAI-compatible endpoint (Qwen-VL, GLM-4V, GPT-4o, or a local Ollama endpoint) to describe it. Only the returned text crosses into the conversation; the image itself never enters the session log. Keywords: DeepSeek Harness plugin, describe_image tool, image understanding, image description, multimodal, vision-language model, VLM, text-only model, Qwen-VL, GLM-4V, GPT-4o, Ollama.

Install

dsh plugin --profile web add github:whitelonng/dsh-plugin-describe-image

The desktop app's plugin list accepts the same spec in its install box (github:whitelonng/dsh-plugin-describe-image); the plugin loads after an application restart.

Features

  • Three input forms: local path, http(s) URL, or the JSON of an [image attachment …] note (resolved through the harness attachment service — copy the note verbatim into image).
  • Live configuration card: the Web GUI's Settings → Plugins → "Image understanding" card edits baseURL, model, and the API key (via the credential seam) with immediate effect — no restart.
  • Per-call API key resolution: inline apiKey → credential seam (apiKeyEnv, default VISION_API_KEY) → launch environment.
  • Security and bounds: redirects refused on every request, maxBytes / maxOutputTokens / timeoutMs bounds, magic-byte media-type gate, bounded error excerpts, secrets never logged.
  • Companion harness changes (shipped in the harness repo, not this subtree): the DeepSeek text-only route flattens image blocks into the copyable [image attachment …] notes, and the host accepts image prompts on text-only routes — together they close the "send an image to a text-only model" loop.

Quick start (in a DeepSeek Harness checkout)

# cordis.yml
- id: describe-image
  name: '@deepseek-ai/dsh-tool-describe-image'
  config:
    baseURL: https://dashscope.aliyuncs.com/compatible-mode/v1
    model: qwen-vl-max
    apiKey: !!js process.env.VISION_API_KEY

FAQ

What does this plugin do?
It adds the describe_image tool to DeepSeek Harness: the agent (or user) hands the tool an image, the tool asks a configured vision-language model to describe it, and only the description text goes back into the conversation.

Which vision models work?
Any OpenAI-compatible vision endpoint: Qwen-VL (https://dashscope.aliyuncs.com/compatible-mode/v1), GLM-4V, GPT-4o, or a local Ollama endpoint. Set baseURL and model in Settings → Plugins → "Image understanding".

Does the image itself enter the conversation or the session log?
No. The image is loaded, checked, and sent only to the vision endpoint; the session log and the model see only the returned description text.

How do I install it?
Run dsh plugin --profile web add github:whitelonng/dsh-plugin-describe-image, or paste the same spec into the desktop app's plugin install box. Restart the application afterwards.

How is the API key configured?
Three layers, in order: an inline apiKey in config, the credential seam (apiKeyEnv, default VISION_API_KEY), then the launch environment. The key is never written into logs.

Is it safe against malicious input?
Redirects are refused, media type is gated by magic bytes, sizes and output tokens are bounded, and error excerpts are truncated — a hostile image or endpoint cannot exfiltrate secrets.

Repository layout

packages/vision/
├── README.md                  # vision capability family
└── tool-describe-image/       # the plugin package (source + tests + docs)

This repository holds the plugin subtree as it lives inside deepseek-harness: package dependencies stay workspace:^, and building, type-checking, and testing happen inside a harness checkout (see INTEGRATION.md). The harness tree is the build environment, not this repo. Keep the two in sync with:

git subtree push --prefix packages/vision dsh-describe-image main   # from the harness checkout

Acknowledgments

  • LINUX DO — This project is continuously shared and discussed in the LINUX DO community.

License

MIT

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 dsh-plugin-describe-image」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。

该插件是 GitHub 源码(未发 npm)——直接从仓库装:

Web profile:

dsh plugin --profile web add github:whitelonng/dsh-plugin-describe-image

Headless(CLI)profile:

dsh plugin --profile headless add github:whitelonng/dsh-plugin-describe-image

实测报告

验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。

使用场景

扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。

适合谁

想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。

二次开发建议

工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。

安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7797 个插件 →