dshbase

Blog · Analysis

Why harness choice moves the bill 7x

August 14, 2026 · dshbase

DeepSeek's models are already cheap per token — but the harness you run them in can multiply the bill anyway. A coding agent re-sends the system prompt, tool definitions, history, and code on every step. Get the request shape wrong and you pay for the same context over and over; get it right and the provider hands most of it back to you for free.

Prefix caching

If the beginning of a request matches a previous one, the provider reuses it. The cache hit rate is the share of input tokens that reused prior computation — higher = cheaper. A changed timestamp or reordered tool invalidates the cache, which is why the cheapest harnesses are obsessive about keeping the request prefix byte-for-byte stable.

Measured: 99.93% hit rate

The open-source agent Pi (on GitHub, ~86k stars) running DeepSeek reported a ~99.93% cache hit rate — a miss rate of 0.07%. Pi gives the model just four tools by default and keeps sessions append-only. Reported impact: ~¥19 per 1B tokens with caching vs ¥900+ without.

Read that again: ~¥19 versus ¥900+. That's not a rounding error — it's a ~47x swing on the input bill driven entirely by how the harness assembles the request. The model didn't change; the discipline around context did.

The 7x gap

A third-party benchmark (Composio) ran DeepSeek V4 Flash across 8 harnesses on real tasks: Pi ~$0.028 per successful task (cheapest), Claude Code ~$0.195 (nearly 7x Pi). Same model, different harness, very different bill.

What actually drives the number

LeverCache-friendlyCache-hostile
Tool setFew tools, fixed order (Pi: 4)Many tools, reordered per turn
SessionAppend-only, never rewrites historyRewrites / summarizes earlier turns
System promptStable, no timestampInjected date, changing metadata
Intermediate dataKept out of context (DSH PTC)Full tool output re-sent every step

This is why DSH ships a Minimal mode (two tools, for benchmarking) and a PTC mode (the model chains tool operations inside one run_code call, keeping intermediate results out of context). They're the same levers Pi pulls, exposed as first-class modes. More on the mode mechanics in our four modes guide.

Make your own preset cache-friendly

The levers aren't hidden — they're just easy to skip when you're assembling an agent preset. Four checks that move the number:

  • Cap the tool set — start from a minimal preset and add tools one at a time; every schema you add lives in the request prefix.
  • Keep the system prompt static — no injected date or changing metadata that would invalidate the prefix.
  • Prefer append-only sessions — don't rewrite or re-summarize earlier turns mid-task.
  • Route long chains through PTC — batch read/search operations into a single run_code call to cut round-trips.

The cost you don't measure: your time

Token cost isn't the only bill. A young plugin ecosystem has a hidden engineering-time cost, and we measured it directly. While testing 101 community plugins on dsh 0.1.0-rc.6 for the plugin directory, the install path itself ate real hours:

  • ERR_PNPM_FETCH_404 — npm mirror lag on freshly published packages; a retry loop disguised as a bug.
  • ERR_REQUIRE_ESM — CommonJS plugins depending on ESM-only packages; the author's problem, your time.
  • Silently inert plugins — no dsh.bundle manifest, so the plugin installs as a plain dependency and never loads.
  • allowBuilds prompts — git-hosted plugins with a prepare script blocked by pnpm by default.

Of those 101 plugins, 11 npm packages worked out of the box, 4 needed a Web/TUI context, 3 failed to install, and 81 were GitHub-source. The one-line fixes are in our troubleshooting guide, but the larger point stands: when people say DSH is cheap, they mean tokens — not the hours you'll spend being your own quality gate in week one.

FAQ

What's a good cache hit rate for DeepSeek?

Pi's ~99.93% is the reference point. Anything below ~90% means you're re-billing a meaningful chunk of context every turn and should look at your tool set and session structure.

Why does adding tools increase cost?

Every tool's schema sits in the system context, and every new tool grows the request prefix. More tools also mean the model can chain more round-trips, each of which re-sends context. Fewer, stable tools = more cache hits.

Does PTC mode actually save money?

Yes, in the right tasks. By batching multiple tool operations into one run_code, PTC removes whole round-trips and keeps intermediate data out of context — directly improving the cache hit rate.

Is DeepSeek Harness cheaper than Pi?

Not by default. Pi is single-mindedly minimal. DSH's Standard mode carries a larger surface; its Minimal and PTC modes are how you get Pi-like economics while keeping the framework.

Takeaway

  • Minimal tool sets and append-only sessions are the biggest levers for cache hits.
  • DSH aligns: small active surface, everything-as-plugin, PTC keeps intermediate data out of context.
  • Remember the second bill: a young ecosystem costs engineering time, not just tokens.

Community-reported measurements, not official DeepSeek benchmarks — verify current pricing. Plugin counts reflect our directory as of August 2026.

All articles →