dshbase

插件目录 / Developer / dsh-continual-evolve

dsh-continual-evolve

已验证 · 实测可装 ZK-Andy

✓ 持续维护 基于 4 个官方 DSH 包 纯 TypeScript

查看 GitHub ↗ ← 返回插件目录

18Stars
0Forks
0未关闭 issue
TypeScript语言
2026-09-04最近推送
跨平台平台

功能简介

DeepSeek Harness 的持续自我进化插件:从会话轨迹提炼版本化、可审计、可回滚的 harness 状态,并带基准驱动的验证循环。

✅
我们的评价
可用 — 实测通过,早期项目

DeepSeek Harness 的持续自我进化插件:从会话轨迹提炼版本化、可审计、可回滚的 harness 状态,并带基准驱动的验证循环。 实测能干净安装、正常启动。早期项目,但功能可用。

「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

README

dsh-continual-evolve

中文 | English

awesome · DSH plugin
npm
CI
License: MIT
Node
Tests

Continual self-evolution for DeepSeek Harness: a versioned, auditable, rollback-safe harness state layer — prompt notes, memories, skills, subagent specs — refined from session trajectories.

The model proposes, the code guarantees. Every mechanical safety property — schema validation, atomic writes, snapshots, versioning, audit trail, acceptance decisions — is enforced in code, never by prompt discipline.

Why

Agents accumulate reusable experience (repeated failures, durable facts, reusable procedures) and forget it next session. This plugin turns that experience into first-class state:

  • Local scope per session; global scope across sessions with merge semantics — plus mechanical promotion guards so only portable, substantial, non-duplicate knowledge reaches global
  • Deterministic rollback: inverse edits generated from applied results — no LLM re-guessing
  • Benchmark loop: candidate refinements are evaluated against frozen cases by a separate scorer before acceptance (rubric encrypted at rest)
  • Store hygiene: /evolve consolidate turns write-time conflict hints and zero-use staleness into one approved, fully reversible batch of archives — with merge, near-duplicate content folds into the surviving original

How it works

  1. Sediment — the model creates entries via evolve_add, or the automatic review gate proposes them from the session trajectory (turn-interval + compaction checkpoints).
  2. Guard — code-enforced validation: edit schema, blast-radius/scope coherence, and the promotion policy (project-scoped markers, thin content, near-duplicate detection, credential screening keep the global store clean — secrets are rejected at every write sink, including mount materialization). Global creates that near-duplicate an existing entry are rejected at write time (≥0.8 similarity); moderate overlaps carry a conflictHint for later consolidation.
  3. Approve — global writes require explicit human approval; local-fate proposals are consulted before they land.
  4. Apply & inject — atomic apply with snapshot + audit event. Prompt notes and delegation specs inject into the system prompt (capped, relevance-ranked, contradicted entries demoted, zero tokens when empty); memories/skills appear as a capped directory index.
  5. Validate & roll back — benchmarks score candidates against frozen cases; rejected candidates roll back deterministically and are captured as draft regression cases (auto_regression benchmark).

Install

# from npm (installs and activates — ships its own bundle patch)
dsh plugin add dsh-continual-evolve

# or from source (first GitHub installs require approving the allowBuilds step)
dsh plugin add ZK-Andy/dsh-continual-evolve

Restart dsh web after installing or updating.

Usage

Commands (in-session):

Command Effect
/evolve help + current local store
/evolve list · history · rollback <id> inspect and revert (add global for the cross-session store)
/evolve plan [msg] run the LLM planner against the store
/evolve wrapup assess this session's local entries: promote / archive / keep
/evolve archive · unarchive · demote <id> hide from injection (data kept, restorable) — demote targets global noise
/evolve consolidate [apply] [merge] report (or apply) one batch archive of conflict-hinted + stale zero-use global entries; merge folds near-duplicate content into the survivors
/evolve failures aggregated failure classes (gate + benchmark)
/evolve log [tail N] [session <id>] plugin log
/evolve export · import <path> backup / restore a store
/evolve mount · unmount <skillId> hot-mount an executable skill as a live plugin
/evolve goal [objective · done · block] round-driven auto-review goal
/evolve benchmark … case lifecycle, runs, acceptance

Model tools: evolve_list / add / update / delete / rollback.

For third-party consumers: every applied evolution (gate or manual) appends a structured evolve_complete event to reviews.jsonl (src/evolve-event.ts defines the shape) alongside the human-readable audit records.

Injection shape: prompt notes and delegation specs inject with content (≤6/kind × 180 chars, relevance-ranked). Memories and skills appear as a directory index ([kind:id] title, capped at 15 lines with a fold counter) — full text via evolve_list. Empty store = zero injected tokens.

Configuration

Key Default Meaning
baseDir resolved DSH home root for the evolve/ stores
autoReview false enable the automatic review gate
reviewIntervalTurns 6 gate cadence on the turn-interval path
maxReviewInputChars 40000 trajectory slice handed to the gate
reviewBudgetTokens 4096 output budget for the gate call
notifyOnAutoReview true visible follow-up notice after an applied gate run
requireGlobalApproval true global edits ask for explicit approval
localFate true gate audits local entries and proposes promote/archive (consulted, never silent)
fateIntervalTurns follows reviewIntervalTurns minimum turns between fate assessments
goalBlockedWrapupTurns 3 consecutive blocked-goal gate runs trigger one fate assessment (0 disables)
promotionBlockPatterns POSIX paths, session ids, ~/.dsh content matching these is project-scoped and never promoted to global
promotionMinChars 100 whole promotions below this length stay local
injectionDirectoryLines 15 entry-directory lines per build before folding into a counter
sectionOrder 118 system-prompt section order
skillsDir <dshHome>/skills where skill entries materialize as SKILL.md bundles
rubricKey auto-generated key file AES-256-GCM passphrase for benchmark rubrics (DSH_EVOLVE_RUBRIC_KEY overrides)
logToFile / logLevel / logMaxBytes true / 1 / 5 MiB plugin-owned JSONL file log with rotation
autoRollbackOnReject true deterministic rollback after a benchmark rejection
autoCase true failed evolution attempts are captured as draft regression cases (auto_regression benchmark)
reviewModel agent's own optional cheaper model for the gate ("provider/model")

Example profile patch:

- id: continual-evolve
  config:
    autoReview: true
    reviewIntervalTurns: 6

Development

pnpm install && pnpm build   # deps + tsc -> lib/
pnpm test                    # vitest (573 tests)
pnpm test:coverage           # v8 coverage, thresholds enforced in CI
pnpm lint                    # oxlint src test

Project layout:

├── src/                   # engine, tools, commands, gate, fate, benchmark, usage…
├── test/                  # vitest suites (36 files)
├── lib/                   # build output (tsc)
├── docs/
│   ├── design.md          # full design doc (hardening matrix)
│   ├── FAQ.md             # real failure/fix records
│   ├── gap-analysis.md    # vs prime-agent /refine + penguin-harness
│   ├── research/pi-dsh-competitor-gap-analysis.md  # pi/dsh ecosystem competitors
│   ├── experiment-bootstrap.md
│   ├── archive/           # closed point-in-time reports
│   └── research/          # penguin report + prime-agent annotated source
├── examples/README.md     # seed benchmark cases
└── .agents/               # AI collaboration layer (AGENTS.md, skills, ADR notes)

Docs & provenance

License

MIT

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 dsh-continual-evolve」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。

该插件是 GitHub 源码(未发 npm)——直接从仓库装:

Web profile:

dsh plugin --profile web add github:ZK-Andy/dsh-continual-evolve

Headless(CLI)profile:

dsh plugin --profile headless add github:ZK-Andy/dsh-continual-evolve

实测报告

验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。

使用场景

扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。

适合谁

想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。

二次开发建议

工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。

安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7797 个插件 →