dshbase

Plugin directory / Developer / dsh-eval

dsh-eval

Verified · install-tested on dsh hccccc01333

✓ Actively maintained 2 contributors Builds on 6 official DSH packages Pure TypeScript

View on GitHub ↗ ← Back to plugin directory

0Stars
0Forks
0Open issues
TypeScriptLanguage
2026-08-14Last push
Cross-platformPlatform

What it does

(no description)

✅
Our take
Works — verified, early-stage project

(no description) It installs cleanly and boots without issues in our testing. It's early-stage but functional.

“Verified” means our automated CI actually ran dsh plugin add in a clean profile and it booted — nothing more. Feature descriptions and version compatibility are the author’s claims. This is not a security audit and not an endorsement of third-party code.

README

dsh-eval

Agent Evaluation Platform for deepseek-harness.

npm version
license

Run benchmarks against headless dsh profiles, harvest persisted session logs as traces, fold automatic metrics, grade task success and tool selection, and report or compare runs 锟斤拷 one benchmark.yaml in, one JSON run + Markdown report out.

The dsh ecosystem already has observability and debugging tools (dsh-trace, dsh-tps, dsh-context-doctor). dsh-eval fills the missing slot: an evaluation platform.

Highlights

  • dsh eval run benchmark.yaml 锟斤拷 orchestrate one headless dsh subprocess per case 锟斤拷 trial
  • Trace harvesting from persisted session logs (everything a model sees is reconstructable from the log)
  • Automatic metrics: task success, tool success, tool-selection accuracy, steps, tokens, latency, cost, retry, invalid tool calls, context usage
  • Scripted grading: expected.tool (tool-selection accuracy) and expected.check (task success)
  • LLM judge: final-answer score and hallucination flags from a judge model
  • Subagent trace merging: child session logs fold into the trial metrics
  • Paired A/B: same-case win/lose/tie statistics across two runs
  • Keyless replay: record once with a key, replay in CI from recorded logs
  • Cross-harness import: dsh eval import codex|claude-code <log> --out run.json
  • dsh eval report run.json 锟斤拷 Markdown report with per-trial scores and pooled rates
  • dsh eval compare v1.json v2.json 锟斤拷 signed B - A comparison table

Status

npm 0.3.0 锟斤拷 113 tests 锟斤拷 100% branch/line coverage on src 锟斤拷 typecheck clean.

Quick start

Install the package directly:

pnpm add dsh-eval
dsh plugin --profile eval add dsh-eval

The npm package targets the official @deepseek-ai/* releases (0.1.0-rc.6 peers). For the source flow, clone this repo and link it to a deepseek-harness checkout:

git clone https://github.com/hccccc01333/dsh-eval.git
cd dsh-eval

Windows (junction):

New-Item -ItemType Junction -Path harness -Target D:\path\to\deepseek-harness

macOS / Linux (symlink):

ln -s /path/to/deepseek-harness harness

Then:

pnpm install
pnpm --filter dsh-eval build
pnpm --filter dsh-eval test

With a dsh launcher from the harness checkout:

dsh plugin --profile eval add dsh-eval
dsh eval run benchmark.yaml --out eval-run.json
dsh eval report eval-run.json
dsh eval compare eval-v1.json eval-v2.json
dsh eval import codex ~/.codex/sessions/.../session.jsonl --out codex-run.json

Benchmark document

name: skill-regression
model: deepseek-v4
profile: headless
command: [dsh]
trials: 3
timeoutMs: 600000
seed: 42
cases:
  - id: fix-tests-001
    prompt: Fix the failing tests in this workspace.
    workspace: ./fixtures/fix-tests
    expected:
      tool: bash
      check: ./check.sh
pricing:
  deepseek-v4:
    inputUsdPerMTokens: 0.27
    cacheReadUsdPerMTokens: 0.07
    cacheWriteUsdPerMTokens: 0.27
    outputUsdPerMTokens: 1.10

Each trial runs in a private temp workspace with an isolated DSH_HOME and non-interactive permissions. The primary session log becomes the trial's trace; scripted grading pools taskSuccess and toolSelectionAccuracy into run-level rates. See packages/eval/README.md for the full field reference.

Metrics

Metric Source
Task success expected.check exit 0
Tool success / rate tool results in the session log
Tool-selection accuracy expected.tool substring match
Steps / turns session log turns and tool events
Tokens / context usage disjoint token buckets + billed context
Latency llmMs / toolMs / ttftMs / latencyMs
Cost per-model pricing table (pricing)
Retry llm/retry events
Invalid tool call tool results carrying an internal failure identity
Final answer score / hallucination LLM judge verdict from judge config

CLI

Command What it does
dsh eval run benchmark.yaml --out run.json Execute the benchmark and write the JSON run
dsh eval report run.json Render a run as Markdown
dsh eval compare base.json candidate.json Compare two runs with signed B - A deltas
dsh eval import codex|claude-code log.jsonl --out run.json Import an external session log as a one-trial run

Roadmap

  • Per-arm leaderboards and significance testing over paired trials
  • Parallel trial execution across cases
  • Web UI dashboard for run reports and comparisons

Repository layout

packages/eval/    plugin bundle + CLI app + tests
harness/          local deepseek-harness checkout (gitignored junction/symlink)

License

MIT

Install

🧩 Let your agent install it (recommended)

Install the catalog once, then DeepSeek Harness can find and install any plugin from this site automatically:

dsh plugin add dshbase-catalog

Then say "install dsh-eval for me" — your agent finds it in the directory and installs it. Docs: dshbase-catalog · verified packs.

Web profile:

dsh plugin --profile web add dsh-eval

Headless (CLI) profile:

dsh plugin --profile headless add dsh-eval

Package

npm: dsh-eval · version 0.3.0 · tested on dsh 0.1.0-rc.6

Test report

Verified end-to-end: L1 install + L2 load + L3 runtime Q&A on dsh 0.1.0-rc.6.

When to use it

Extend the agent's coding surface — give it a new tool, workflow, or integration so it handles a dev task it couldn't before.

Who it's for

Developers who want dsh to behave like a teammate on real codebases — editing, running, and verifying changes rather than just answering.

For developers — extending it

The tool/command surface is the seam: expose more of the SDK, add smarter context wiring, or tighten the loop between code changes and verification.

Security: not yet scanned — our daily static scan will cover it shortly.

Share this badge

More in Developer

Browse all 7797 plugins →