Plugin directory / Developer / dsh-image-reader
dsh-image-reader
Verified · install-tested on dsh zcXie777
What it does
Give DeepSeek Harness agents native image reading: a read_image tool backed by any OpenAI-compatible vision endpoint.
Works — verified, early-stage project
Give DeepSeek Harness agents native image reading: a read_image tool backed by any OpenAI-compatible vision endpoint. It installs cleanly and boots without issues in our testing. It's early-stage but functional.
“Verified” means our automated CI actually ran dsh plugin add in a clean profile and it booted — nothing more. Feature descriptions and version compatibility are the author’s claims. This is not a security audit and not an endorsement of third-party code.
README
dsh-image-reader
Give a text-only DeepSeek Harness agent the ability to read images directly: one model-facing read_image tool that asks any OpenAI-compatible vision endpoint about an image by its workspace path.
Why
DeepSeek Harness is "everything is a plugin". This bundle mounts a single tool so the model can look at a screenshot, diagram, or photograph and answer questions about it, instead of only ever reasoning over text.
Verification status
- Verified locally:
npm run typecheck,npm run build, andnpm test(16 tests) all pass. - Not yet verified: a real end-to-end read against a live vision endpoint. The request/response logic is covered by a mocked-fetch unit test, but the plugin has not been smoke-tested inside a running dsh profile against a real multimodal model. Do that once with a real
VISION_API_KEYbefore relying on it.
Install
git clone https://github.com/zcXie777/dsh-image-reader.git
cd dsh-image-reader
npm install
npm run build # lib/ is not committed; build once after cloning
cd ..
dsh plugin --profile web add "$PWD/dsh-image-reader"
dsh plugin --profile headless add "$PWD/dsh-image-reader"
dsh --profile web --dump-config | grep image-reader
Restart a running Web profile after installing.
Configure
provider.baseUrl and provider.model are required; the plugin never assumes a vendor. Override them in the profile patch row with the same id:
- id: image-reader
config:
provider:
baseUrl: https://api.openai.com/v1
model: gpt-4o-mini
apiKeyEnv: VISION_API_KEY
lang: zh
timeoutMs: 60000
maxImageBytes: 10485760
allowedDirs: []
Set the key in the environment before starting the profile:
export VISION_API_KEY=sk-...
Use
In a conversation, point the model at an image path and ask:
read_image image="screenshot.png" query="What error is shown in this dialog?"
read_image image="diagram.png"
Configuration fields
| Field | Default | Contract |
|---|---|---|
provider.baseUrl |
— (required) | OpenAI-compatible chat/completions base URL |
provider.model |
— (required) | Multimodal model name |
provider.apiKeyEnv |
VISION_API_KEY |
Environment variable holding the API key |
lang |
zh |
Answer language: zh or en |
timeoutMs |
60000 |
Whole-request deadline, 1000–600000 ms |
maxImageBytes |
10485760 |
Encoded-byte limit per image |
allowedDirs |
[] |
Extra realpath-resolved input roots; the workspace is always allowed |
Security
- Inputs resolve against the workspace and
allowedDirsthroughrealpath, so a symlink cannot escape the fence. - Images are size-limited and extension-checked before upload.
- The key is read from the environment per call, never stored in config.
Development
npm install
npm run typecheck
npm run build
Publish
Tag the repo with the dsh-plugin topic so it is discoverable, and publish to npm when ready.
License
MIT
Install
Install the catalog once, then DeepSeek Harness can find and install any plugin from this site automatically:
dsh plugin add dshbase-catalog Then say "install dsh-image-reader for me" — your agent finds it in the directory and installs it. Docs: dshbase-catalog · verified packs.
This plugin is GitHub source (not published to npm) — install it straight from the repo:
Web profile:
dsh plugin --profile web add github:zcXie777/dsh-image-reader Headless (CLI) profile:
dsh plugin --profile headless add github:zcXie777/dsh-image-reader Test report
Verified: L1 install + L2 load + L3 runtime from GitHub source on dsh 0.1.0-rc.6.
When to use it
Extend the agent's coding surface — give it a new tool, workflow, or integration so it handles a dev task it couldn't before.
Who it's for
Developers who want dsh to behave like a teammate on real codebases — editing, running, and verifying changes rather than just answering.
For developers — extending it
The tool/command surface is the seam: expose more of the SDK, add smarter context wiring, or tighten the loop between code changes and verification.