dshbase

插件目录 / Developer / dsh-eval

dsh-eval

已验证 · 实测可装 hccccc01333

✓ 持续维护 2 位贡献者 基于 6 个官方 DSH 包 纯 TypeScript

查看 GitHub ↗ ← 返回插件目录

0Stars
0Forks
0未关闭 issue
TypeScript语言
2026-08-14最近推送
跨平台平台

功能简介

(无描述)

✅
我们的评价
可用 — 实测通过,早期项目

(无描述) 实测能干净安装、正常启动。早期项目,但功能可用。

「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

README

dsh-eval

Agent Evaluation Platform for deepseek-harness.

npm version
license

Run benchmarks against headless dsh profiles, harvest persisted session logs as traces, fold automatic metrics, grade task success and tool selection, and report or compare runs 锟斤拷 one benchmark.yaml in, one JSON run + Markdown report out.

The dsh ecosystem already has observability and debugging tools (dsh-trace, dsh-tps, dsh-context-doctor). dsh-eval fills the missing slot: an evaluation platform.

Highlights

  • dsh eval run benchmark.yaml 锟斤拷 orchestrate one headless dsh subprocess per case 锟斤拷 trial
  • Trace harvesting from persisted session logs (everything a model sees is reconstructable from the log)
  • Automatic metrics: task success, tool success, tool-selection accuracy, steps, tokens, latency, cost, retry, invalid tool calls, context usage
  • Scripted grading: expected.tool (tool-selection accuracy) and expected.check (task success)
  • LLM judge: final-answer score and hallucination flags from a judge model
  • Subagent trace merging: child session logs fold into the trial metrics
  • Paired A/B: same-case win/lose/tie statistics across two runs
  • Keyless replay: record once with a key, replay in CI from recorded logs
  • Cross-harness import: dsh eval import codex|claude-code <log> --out run.json
  • dsh eval report run.json 锟斤拷 Markdown report with per-trial scores and pooled rates
  • dsh eval compare v1.json v2.json 锟斤拷 signed B - A comparison table

Status

npm 0.3.0 锟斤拷 113 tests 锟斤拷 100% branch/line coverage on src 锟斤拷 typecheck clean.

Quick start

Install the package directly:

pnpm add dsh-eval
dsh plugin --profile eval add dsh-eval

The npm package targets the official @deepseek-ai/* releases (0.1.0-rc.6 peers). For the source flow, clone this repo and link it to a deepseek-harness checkout:

git clone https://github.com/hccccc01333/dsh-eval.git
cd dsh-eval

Windows (junction):

New-Item -ItemType Junction -Path harness -Target D:\path\to\deepseek-harness

macOS / Linux (symlink):

ln -s /path/to/deepseek-harness harness

Then:

pnpm install
pnpm --filter dsh-eval build
pnpm --filter dsh-eval test

With a dsh launcher from the harness checkout:

dsh plugin --profile eval add dsh-eval
dsh eval run benchmark.yaml --out eval-run.json
dsh eval report eval-run.json
dsh eval compare eval-v1.json eval-v2.json
dsh eval import codex ~/.codex/sessions/.../session.jsonl --out codex-run.json

Benchmark document

name: skill-regression
model: deepseek-v4
profile: headless
command: [dsh]
trials: 3
timeoutMs: 600000
seed: 42
cases:
  - id: fix-tests-001
    prompt: Fix the failing tests in this workspace.
    workspace: ./fixtures/fix-tests
    expected:
      tool: bash
      check: ./check.sh
pricing:
  deepseek-v4:
    inputUsdPerMTokens: 0.27
    cacheReadUsdPerMTokens: 0.07
    cacheWriteUsdPerMTokens: 0.27
    outputUsdPerMTokens: 1.10

Each trial runs in a private temp workspace with an isolated DSH_HOME and non-interactive permissions. The primary session log becomes the trial's trace; scripted grading pools taskSuccess and toolSelectionAccuracy into run-level rates. See packages/eval/README.md for the full field reference.

Metrics

Metric Source
Task success expected.check exit 0
Tool success / rate tool results in the session log
Tool-selection accuracy expected.tool substring match
Steps / turns session log turns and tool events
Tokens / context usage disjoint token buckets + billed context
Latency llmMs / toolMs / ttftMs / latencyMs
Cost per-model pricing table (pricing)
Retry llm/retry events
Invalid tool call tool results carrying an internal failure identity
Final answer score / hallucination LLM judge verdict from judge config

CLI

Command What it does
dsh eval run benchmark.yaml --out run.json Execute the benchmark and write the JSON run
dsh eval report run.json Render a run as Markdown
dsh eval compare base.json candidate.json Compare two runs with signed B - A deltas
dsh eval import codex|claude-code log.jsonl --out run.json Import an external session log as a one-trial run

Roadmap

  • Per-arm leaderboards and significance testing over paired trials
  • Parallel trial execution across cases
  • Web UI dashboard for run reports and comparisons

Repository layout

packages/eval/    plugin bundle + CLI app + tests
harness/          local deepseek-harness checkout (gitignored junction/symlink)

License

MIT

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 dsh-eval」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。

Web profile:

dsh plugin --profile web add dsh-eval

Headless(CLI)profile:

dsh plugin --profile headless add dsh-eval

包信息

npm:dsh-eval · 版本 0.3.0 · 实测环境 dsh 0.1.0-rc.6

实测报告

端到端验证通过:dsh 0.1.0-rc.6 上 L1 安装 + L2 加载 + L3 运行问答。

使用场景

扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。

适合谁

想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。

二次开发建议

工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。

安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7797 个插件 →