Plugin directory / Developer / dsh-eval
dsh-eval
Verified · install-tested on dsh hccccc01333
What it does
(no description)
Works — verified, early-stage project
(no description) It installs cleanly and boots without issues in our testing. It's early-stage but functional.
“Verified” means our automated CI actually ran dsh plugin add in a clean profile and it booted — nothing more. Feature descriptions and version compatibility are the author’s claims. This is not a security audit and not an endorsement of third-party code.
README
dsh-eval
Agent Evaluation Platform for deepseek-harness.
Run benchmarks against headless dsh profiles, harvest persisted session logs as traces, fold automatic metrics, grade task success and tool selection, and report or compare runs 锟斤拷 one benchmark.yaml in, one JSON run + Markdown report out.
The dsh ecosystem already has observability and debugging tools (
dsh-trace,dsh-tps,dsh-context-doctor). dsh-eval fills the missing slot: an evaluation platform.
Highlights
dsh eval run benchmark.yaml锟斤拷 orchestrate one headlessdshsubprocess per case 锟斤拷 trial- Trace harvesting from persisted session logs (everything a model sees is reconstructable from the log)
- Automatic metrics: task success, tool success, tool-selection accuracy, steps, tokens, latency, cost, retry, invalid tool calls, context usage
- Scripted grading:
expected.tool(tool-selection accuracy) andexpected.check(task success) - LLM judge: final-answer score and hallucination flags from a judge model
- Subagent trace merging: child session logs fold into the trial metrics
- Paired A/B: same-case win/lose/tie statistics across two runs
- Keyless replay: record once with a key, replay in CI from recorded logs
- Cross-harness import:
dsh eval import codex|claude-code <log> --out run.json dsh eval report run.json锟斤拷 Markdown report with per-trial scores and pooled ratesdsh eval compare v1.json v2.json锟斤拷 signedB - Acomparison table
Status
npm 0.3.0 锟斤拷 113 tests 锟斤拷 100% branch/line coverage on src 锟斤拷 typecheck clean.
Quick start
Install the package directly:
pnpm add dsh-eval
dsh plugin --profile eval add dsh-eval
The npm package targets the official @deepseek-ai/* releases (0.1.0-rc.6 peers). For the source flow, clone this repo and link it to a deepseek-harness checkout:
git clone https://github.com/hccccc01333/dsh-eval.git
cd dsh-eval
Windows (junction):
New-Item -ItemType Junction -Path harness -Target D:\path\to\deepseek-harness
macOS / Linux (symlink):
ln -s /path/to/deepseek-harness harness
Then:
pnpm install
pnpm --filter dsh-eval build
pnpm --filter dsh-eval test
With a dsh launcher from the harness checkout:
dsh plugin --profile eval add dsh-eval
dsh eval run benchmark.yaml --out eval-run.json
dsh eval report eval-run.json
dsh eval compare eval-v1.json eval-v2.json
dsh eval import codex ~/.codex/sessions/.../session.jsonl --out codex-run.json
Benchmark document
name: skill-regression
model: deepseek-v4
profile: headless
command: [dsh]
trials: 3
timeoutMs: 600000
seed: 42
cases:
- id: fix-tests-001
prompt: Fix the failing tests in this workspace.
workspace: ./fixtures/fix-tests
expected:
tool: bash
check: ./check.sh
pricing:
deepseek-v4:
inputUsdPerMTokens: 0.27
cacheReadUsdPerMTokens: 0.07
cacheWriteUsdPerMTokens: 0.27
outputUsdPerMTokens: 1.10
Each trial runs in a private temp workspace with an isolated DSH_HOME and non-interactive permissions. The primary session log becomes the trial's trace; scripted grading pools taskSuccess and toolSelectionAccuracy into run-level rates. See packages/eval/README.md for the full field reference.
Metrics
| Metric | Source |
|---|---|
| Task success | expected.check exit 0 |
| Tool success / rate | tool results in the session log |
| Tool-selection accuracy | expected.tool substring match |
| Steps / turns | session log turns and tool events |
| Tokens / context usage | disjoint token buckets + billed context |
| Latency | llmMs / toolMs / ttftMs / latencyMs |
| Cost | per-model pricing table (pricing) |
| Retry | llm/retry events |
| Invalid tool call | tool results carrying an internal failure identity |
| Final answer score / hallucination | LLM judge verdict from judge config |
CLI
| Command | What it does |
|---|---|
dsh eval run benchmark.yaml --out run.json |
Execute the benchmark and write the JSON run |
dsh eval report run.json |
Render a run as Markdown |
dsh eval compare base.json candidate.json |
Compare two runs with signed B - A deltas |
dsh eval import codex|claude-code log.jsonl --out run.json |
Import an external session log as a one-trial run |
Roadmap
- Per-arm leaderboards and significance testing over paired trials
- Parallel trial execution across cases
- Web UI dashboard for run reports and comparisons
Repository layout
packages/eval/ plugin bundle + CLI app + tests
harness/ local deepseek-harness checkout (gitignored junction/symlink)
License
MIT
Install
Install the catalog once, then DeepSeek Harness can find and install any plugin from this site automatically:
dsh plugin add dshbase-catalog Then say "install dsh-eval for me" — your agent finds it in the directory and installs it. Docs: dshbase-catalog · verified packs.
Web profile:
dsh plugin --profile web add dsh-eval Headless (CLI) profile:
dsh plugin --profile headless add dsh-eval Package
npm: dsh-eval · version 0.3.0 · tested on dsh 0.1.0-rc.6
Test report
Verified end-to-end: L1 install + L2 load + L3 runtime Q&A on dsh 0.1.0-rc.6.
When to use it
Extend the agent's coding surface — give it a new tool, workflow, or integration so it handles a dev task it couldn't before.
Who it's for
Developers who want dsh to behave like a teammate on real codebases — editing, running, and verifying changes rather than just answering.
For developers — extending it
The tool/command surface is the seam: expose more of the SDK, add smarter context wiring, or tighten the loop between code changes and verification.