dshbase

Blog · Concept

Self-evolving software: miracle or chaos?

August 14, 2026 · dshbase

DeepSeek Harness has been called both the prototype of self-evolving software and a recipe for even more chaos. The split comes down to one feature: an agent that can modify its own running capabilities. Here's what that actually means in practice — including where it goes wrong — and why the cautionary tales matter.

What "self-evolving" actually means in DSH

It's not science fiction. In Create mode, the agent can inspect its own running Cordis environment, experiment with plugins in memory, and — critically — create a plugin and hot-attach it to the running flow. The "make your own wrench" metaphor: the agent notices it has no wrench, forges one, attaches it to its hand, and keeps working.

This works because Cordis guarantees two properties: temporal composability (unloading a plugin fully undoes its side effects) and spatial composability (dependencies re-resolve as siblings change). Without those, hot-swapping your own code mid-run would corrupt state. We've watched the happy path on dsh 0.1.0-rc.6: dsh plugin add dsh-memory installs and, on next launch, the plugin is already in the config tree — no rebuild, no restart dance. The kernel discovers and loads it.

But there's a contract underneath. A self-created plugin only activates if it ships a dsh.bundle manifest. Without it, the package installs silently as a plain dependency — the agent "forged a wrench" that never connects to its hand. When we audited 101 community plugins, silent non-activation and outright install failure were the most common ways "self-evolution" fell short of the promise.

The safety rails are real — and load-bearing

The difference between "the agent extended itself" and "the agent ran arbitrary code with your permissions" is a handful of prompts and guarantees:

  • allowBuilds confirmation — when a GitHub-sourced plugin ships a prepare build script, pnpm blocks it and asks you to confirm. For a self-modifying agent this is a genuine permission boundary, not friction.
  • Clean unload — temporal composability means a bad experiment can be removed and the state rolls back to before it was added.
  • The trajectory log — every tool call and permission change is an event, so you can review exactly what the agent did to itself.

The cautionary tale: OpenClaw

The skepticism isn't hypothetical. OpenClaw, an earlier agent that promised a similar "everything is modular" vision, ran into exactly the chaos critics predict — a reminder that self-modification without safety rails produces unmaintainable state, not evolution. The difference DSH argues for is Cordis's composability guarantees plus the rails above, but that's still a claim that needs proving at scale. The 101-plugin audit cuts both ways here: a fast-growing ecosystem proves the mechanics work, but its long quality tail (81 GitHub-only, several failing outright) shows the rails aren't automatic.

A practical guide to self-evolution

If you want to use this carefully, treat it as a controlled experiment, not free rein:

  • Scope it down — start with read-only presets (e.g. "read code, never modify files") before anything that writes.
  • Use the trajectory log — every tool call and permission change is an event; review what the agent actually did to itself.
  • One change at a time — have the agent add a single plugin, verify, then continue, rather than letting it rewrite everything at once.
  • Verify the manifest — after the agent "creates" a plugin, confirm it has a dsh.bundle; otherwise it installs silently and never activates.
  • Keep a restore path — profiles and presets mean you can always revert to a known-good config.

When a self-created plugin misbehaves, the failure is usually one of the documented modes — a ERR_REQUIRE_ESM from a CJS plugin depending on ESM-only packages, or a missing manifest. The troubleshooting guide covers each with a one-line fix, and the plugin directory shows which community plugins actually activate.

The honest take

Self-evolving software is neither a miracle nor automatically chaos — it's a capability with a failure mode. The miracle framing oversells it; the chaos framing underestimates Cordis's safeguards. The truth will show up in whether real teams can use it without losing control. For now, the practical path is scoped, logged, reversible self-modification — and an honest read of the manifest contract before you trust anything the agent "evolved" for you.

FAQ

Can the agent really change its own code? Yes — in Create mode it can create a plugin and hot-attach it to the running flow. But "create" is bounded by the dsh.bundle manifest contract: no manifest, no activation.

How do I keep self-evolution from wrecking my setup? Scope it to read-only first, use the trajectory log to review every change, add one plugin at a time, and keep a known-good profile you can restore to.

Why did OpenClaw fail if the idea is sound? It had the modular vision without Cordis's composability guarantees and rails. Modularity alone produces unmaintainable state; you need clean unload and permission boundaries on top.

Related: session log observability · the Create mode.

All articles →