dsh-eval
已验证 · 实测可装 hccccc01333
功能简介
(无描述)
可用 — 实测通过,早期项目
(无描述) 实测能干净安装、正常启动。早期项目,但功能可用。
「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。
README
dsh-eval
Agent Evaluation Platform for deepseek-harness.
Run benchmarks against headless dsh profiles, harvest persisted session logs as traces, fold automatic metrics, grade task success and tool selection, and report or compare runs 锟斤拷 one benchmark.yaml in, one JSON run + Markdown report out.
The dsh ecosystem already has observability and debugging tools (
dsh-trace,dsh-tps,dsh-context-doctor). dsh-eval fills the missing slot: an evaluation platform.
Highlights
dsh eval run benchmark.yaml锟斤拷 orchestrate one headlessdshsubprocess per case 锟斤拷 trial- Trace harvesting from persisted session logs (everything a model sees is reconstructable from the log)
- Automatic metrics: task success, tool success, tool-selection accuracy, steps, tokens, latency, cost, retry, invalid tool calls, context usage
- Scripted grading:
expected.tool(tool-selection accuracy) andexpected.check(task success) - LLM judge: final-answer score and hallucination flags from a judge model
- Subagent trace merging: child session logs fold into the trial metrics
- Paired A/B: same-case win/lose/tie statistics across two runs
- Keyless replay: record once with a key, replay in CI from recorded logs
- Cross-harness import:
dsh eval import codex|claude-code <log> --out run.json dsh eval report run.json锟斤拷 Markdown report with per-trial scores and pooled ratesdsh eval compare v1.json v2.json锟斤拷 signedB - Acomparison table
Status
npm 0.3.0 锟斤拷 113 tests 锟斤拷 100% branch/line coverage on src 锟斤拷 typecheck clean.
Quick start
Install the package directly:
pnpm add dsh-eval
dsh plugin --profile eval add dsh-eval
The npm package targets the official @deepseek-ai/* releases (0.1.0-rc.6 peers). For the source flow, clone this repo and link it to a deepseek-harness checkout:
git clone https://github.com/hccccc01333/dsh-eval.git
cd dsh-eval
Windows (junction):
New-Item -ItemType Junction -Path harness -Target D:\path\to\deepseek-harness
macOS / Linux (symlink):
ln -s /path/to/deepseek-harness harness
Then:
pnpm install
pnpm --filter dsh-eval build
pnpm --filter dsh-eval test
With a dsh launcher from the harness checkout:
dsh plugin --profile eval add dsh-eval
dsh eval run benchmark.yaml --out eval-run.json
dsh eval report eval-run.json
dsh eval compare eval-v1.json eval-v2.json
dsh eval import codex ~/.codex/sessions/.../session.jsonl --out codex-run.json
Benchmark document
name: skill-regression
model: deepseek-v4
profile: headless
command: [dsh]
trials: 3
timeoutMs: 600000
seed: 42
cases:
- id: fix-tests-001
prompt: Fix the failing tests in this workspace.
workspace: ./fixtures/fix-tests
expected:
tool: bash
check: ./check.sh
pricing:
deepseek-v4:
inputUsdPerMTokens: 0.27
cacheReadUsdPerMTokens: 0.07
cacheWriteUsdPerMTokens: 0.27
outputUsdPerMTokens: 1.10
Each trial runs in a private temp workspace with an isolated DSH_HOME and non-interactive permissions. The primary session log becomes the trial's trace; scripted grading pools taskSuccess and toolSelectionAccuracy into run-level rates. See packages/eval/README.md for the full field reference.
Metrics
| Metric | Source |
|---|---|
| Task success | expected.check exit 0 |
| Tool success / rate | tool results in the session log |
| Tool-selection accuracy | expected.tool substring match |
| Steps / turns | session log turns and tool events |
| Tokens / context usage | disjoint token buckets + billed context |
| Latency | llmMs / toolMs / ttftMs / latencyMs |
| Cost | per-model pricing table (pricing) |
| Retry | llm/retry events |
| Invalid tool call | tool results carrying an internal failure identity |
| Final answer score / hallucination | LLM judge verdict from judge config |
CLI
| Command | What it does |
|---|---|
dsh eval run benchmark.yaml --out run.json |
Execute the benchmark and write the JSON run |
dsh eval report run.json |
Render a run as Markdown |
dsh eval compare base.json candidate.json |
Compare two runs with signed B - A deltas |
dsh eval import codex|claude-code log.jsonl --out run.json |
Import an external session log as a one-trial run |
Roadmap
- Per-arm leaderboards and significance testing over paired trials
- Parallel trial execution across cases
- Web UI dashboard for run reports and comparisons
Repository layout
packages/eval/ plugin bundle + CLI app + tests
harness/ local deepseek-harness checkout (gitignored junction/symlink)
License
MIT
安装
装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:
dsh plugin add dshbase-catalog 然后对 agent 说「帮我装 dsh-eval」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。
Web profile:
dsh plugin --profile web add dsh-eval Headless(CLI)profile:
dsh plugin --profile headless add dsh-eval 包信息
npm:dsh-eval · 版本 0.3.0 · 实测环境 dsh 0.1.0-rc.6
实测报告
端到端验证通过:dsh 0.1.0-rc.6 上 L1 安装 + L2 加载 + L3 运行问答。
使用场景
扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。
适合谁
想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。
二次开发建议
工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。