插件目录 / Developer / dsh-qwen-multimodal
dsh-qwen-multimodal
已验证 · 实测可装 wuwangmao
✓ 持续维护 2 位贡献者
0Stars
0Forks
0未关闭 issue
JavaScript语言
2026-08-15最近推送
跨平台平台
功能简介
DSH捆绑:Qwen多模态桥接,视觉/语音/文生图。
我们的评价
可用 — 实测通过,早期项目
可用 — 实测通过,早期项目
DSH捆绑:Qwen多模态桥接,视觉/语音/文生图。 实测能干净安装、正常启动。早期项目,但功能可用。
「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。
README
dsh-qwen-multimodal
A DSH bundle that gives text-only main models (e.g. DeepSeek) **three multimodal skills in one plugin** through Qwen APIs: vision, speech-to-text, and text-to-image — with a built-in generate-then-verify quality loop. | Tool | Capability | Backend | |---|---|---| |describe_image | Image / screenshot / OCR / chart understanding (multiple images at once) | Qwen VL (default qwen3-vl-flash) |
| transcribe_audio | Speech / recording transcription (wav/mp3/m4a/aac/flac/ogg/amr) | Qwen3-ASR (qwen3-asr-flash) |
| generate_image | Generate images from text and save them locally | Qwen-Image (qwen-image-plus) |
Media never enters the main model context: visual/audio content is converted to text, and generated images are saved to local files with paths returned by the tool.
How it works
All API calls reuse the original Python scripts inskills/deepseek-vision/scripts/*.py and the .env configuration (vision/audio use the Alibaba Cloud Bailian OpenAI-compatible endpoint; image generation uses the native multimodal-generation endpoint). The plugin itself is a pure-JS Cordis bundle depending only on the host's mounted subprocess / tools services — **no build step required for git installs**.
Install
From GitHub
``sh
dsh plugin --profile demo add github:wuwangmao/dsh-qwen-multimodal
`
Local checkout / tarball
`sh
dsh plugin --profile demo add ./dsh-qwen-multimodal
or
pnpm pack # then
dsh plugin --profile demo add ./dsh-qwen-multimodal-0.1.0.tgz
`
Before first use, configure your API key: copy skills/deepseek-vision/.env.example to
skills/deepseek-vision/.env and fill in VISION_API_KEY (create one in the Alibaba Cloud Bailian console; new users get free quota, college students get a ¥300 annual voucher). Vision/audio/image reuse the same key by default, or configure them separately (see .env.example).
Python
**Python 3.10+ is required** (the image-generation script uses int | None type-annotation syntax).
python is resolved from the system PATH by default. If it cannot be resolved, restate the plugin row in your profile's cordis.patch.yml and set config.pythonPath.
Configuration overrides
The plugin uses the bundled skill directory by default. To point it at an external directory (e.g. to reuse an existing .env and scripts, or to keep your key outside node_modules), restate the row:
`yaml
- insert:
- id: qwen-multimodal
name: dsh-qwen-multimodal
config:
skillDir: 'D:/qwen-vision'
pythonPath: 'C:/path/to/python.exe'
`
Usage
Once loaded, the model can call the three tools directly:
- describe_image({ images: ['screenshot.png'] }) — verbatim extraction of text/code/errors in images
- describe_image({ images: ['chart.png'], prompt: '逐字提取图中所有文字,保留原样' }) — custom prompt
- transcribe_audio({ audios: ['recording.m4a'], language: 'zh' }) — specify language for accuracy
- generate_image({ prompt: 'a cute orange cat on a windowsill watching the sunset', out_dir: './out' }) — generate and save locally
- generate_image({ prompt: '...', out_dir: './out', verify: true }) — generate, then automatically re-check the result with Qwen VL against the prompt (quality loop)
Layout
`
dsh-qwen-multimodal/
├── package.json # dsh.bundle manifest
├── cordis.patch.yml # bundle layer: inserts the plugin row
├── src/index.js # plugin: registers the three model tools (pure JS)
├── scripts/selfcheck.mjs # self-check: node scripts/selfcheck.mjs
└── skills/deepseek-vision/ # skill assets: SKILL.md + Python scripts + .env.example
``
License
MIT安装
🧩 让 Agent 自动装(推荐)
装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:
dsh plugin add dshbase-catalog 然后对 agent 说「帮我装 dsh-qwen-multimodal」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。
该插件是 GitHub 源码(未发 npm)——直接从仓库装:
Web profile:
dsh plugin --profile web add github:wuwangmao/dsh-qwen-multimodal Headless(CLI)profile:
dsh plugin --profile headless add github:wuwangmao/dsh-qwen-multimodal 实测报告
验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。
使用场景
扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。
适合谁
想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。
二次开发建议
工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。