Plugin directory / Developer / dsh-plugin-vision-toolkit
dsh-plugin-vision-toolkit
Verified · install-tested on dsh YYTbit
What it does
Vision toolkit for DeepSeek Harness -- give text-only agents eyes
Works — verified, early-stage project
Vision toolkit for DeepSeek Harness -- give text-only agents eyes It installs cleanly and boots without issues in our testing. It's early-stage but functional.
“Verified” means our automated CI actually ran dsh plugin add in a clean profile and it booted — nothing more. Feature descriptions and version compatibility are the author’s claims. This is not a security audit and not an endorsement of third-party code.
README
dsh-plugin-vision-toolkit
Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.
What it does
Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.
Tools
glance-- describe, ask about, or OCR an imageground-- locate a specific element (returns bounding box)detect-- find all instances of an element kindcrop-- cut a region from an image
Install
dsh plugin --profile your-profile add dsh-plugin-vision-toolkit
Configuration
Set environment variables:
export VISION_API_KEY=sk-xxx # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://... # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2 # Vision model name
Usage examples
# Describe an image
glance screenshot.png
# Ask a question
glance screenshot.png -q "What error is shown?"
# OCR
glance screenshot.png --ocr
# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870
# Find all buttons
detect screenshot.png "buttons"
# Crop a region
crop screenshot.png 450,820,620,870 button.png
How it works
The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:
- Reads the image file
- Encodes it as base64
- Sends it to the vision API with a prompt
- Returns the text response
The agent never sees raw pixels -- it gets text descriptions it can reason about.
Supported vision providers
- DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
- OpenAI GPT-4V / GPT-4o
- Any OpenAI-compatible multimodal endpoint
License
MIT -- YYTbit
Install
Install the catalog once, then DeepSeek Harness can find and install any plugin from this site automatically:
dsh plugin add dshbase-catalog Then say "install dsh-plugin-vision-toolkit for me" — your agent finds it in the directory and installs it. Docs: dshbase-catalog · verified packs.
Web profile:
dsh plugin --profile web add dsh-plugin-vision-toolkit Headless (CLI) profile:
dsh plugin --profile headless add dsh-plugin-vision-toolkit Package
npm: dsh-plugin-vision-toolkit · version 0.1.0 · tested on dsh 0.1.0-rc.6
Test report
Verified end-to-end: L1 install + L2 load + L3 runtime Q&A on dsh 0.1.0-rc.6.
When to use it
Extend the agent's coding surface — give it a new tool, workflow, or integration so it handles a dev task it couldn't before.
Who it's for
Developers who want dsh to behave like a teammate on real codebases — editing, running, and verifying changes rather than just answering.
For developers — extending it
The tool/command surface is the seam: expose more of the SDK, add smarter context wiring, or tighten the loop between code changes and verification.