Blog · Analysis
How a minimal harness hits a 99.93% cache hit rate
August 14, 2026 · dshbase
The cheapest way to run DeepSeek isn't just a cheap model — it's a harness that reuses the work the model already did. Community measurements of the open-source agent Pi showed a 99.93% prefix-cache hit rate. That number isn't a model property — it's a harness property, and the levers behind it are concrete and transferable.
The cost driver
A coding agent re-sends the system prompt, tool definitions, history, and code on every step. Automatic prefix caching reuses the computed prefix if the start of the request is unchanged. The cache hit rate is the share of tokens reused — higher is cheaper. When you pay per token, every re-sent block you can avoid is money back, and the biggest re-sent blocks are exactly the ones a harness controls: the system prompt and the tool schema.
Why caching is fragile
Add a timestamp, reorder a tool, or rewrite earlier content and the cache misses. Cache-friendly harnesses keep the request prefix stable. This is where most harnesses quietly lose the game: they inject a fresh timestamp or reshuffle tool definitions on every call, and the hit rate collapses even though nothing about the task changed.
Pi's recipe
Pi gives the model just four tools by default (read, write, edit, run) with everything else opt-in, and sessions append rather than rewrite. Two consequences follow:
- Fewer tools = smaller, more stable prefix. The tool schema is a fixed block at the front of every request. Four tools is a tiny block that almost never changes; forty tools is a large block that changes the moment you reorder or toggle one.
- Append-only = nothing earlier ever changes. When history rewrites, everything after the edit point loses cache. When it only appends, the prefix stays byte-identical.
Reported result: ~99.93% hit rate, roughly ¥19 per 1B tokens with caching vs ¥900+ without. The gap between those two numbers is the entire point of harness design.
The 7x gap
A Composio benchmark ran DeepSeek V4 Flash across 8 harnesses: Pi ~$0.028 per successful task (cheapest), Claude Code ~$0.195 (nearly 7x). Same model, different harness, very different bill. That's the strongest argument that caching is a harness property, not a model one.
What this means for DSH
DSH sits on the cache-friendly end of the same spectrum, and its design maps directly onto Pi's two levers:
- Minimal mode gives the model just two tools — a persistent bash shell and a file editor — the smallest possible stable prefix for benchmarking (see the modes guide).
- PTC mode keeps intermediate data out of the context by collapsing many tool round-trips into a single program execution.
- Append-only sessions — the trajectory log is an append-only event log, which is also a stable request prefix.
The cautionary flip side is tool sprawl. Standard mode ships a large tool surface, and the broader plugin ecosystem adds more — our audit of 101 community plugins shows how fast the surface can grow. Every tool you add is a block at the front of every request, and a block that can change. That's why the cache play is a discipline, not a default: minimal active surface, everything else opt-in.
Takeaway
- Minimal tool sets + append-only sessions = the biggest levers for cache hits.
- The 7x gap between harnesses on the same model is the real cost story — model price is one input, harness cache behavior is the other.
- DSH aligns: small active surface, everything-as-plugin, PTC keeps intermediate data out of context.
How to measure your own hit rate
You don't have to take Pi's number on faith — you can read your own off any single run. The API's usage block separates cached from uncached input tokens, so your per-call hit rate is visible in the response. When it drops, the usual culprits are a reordered tool schema or a timestamp that slipped into the system prompt — both prefix problems, not model problems. That's the practical payoff of treating the prefix as something you design rather than inherit.
FAQ
Can I hit 99.93% with any harness? Only if the harness keeps a small, stable prefix and appends rather than rewrites. A harness that reshuffles tool definitions or injects timestamps won't get close, no matter how cheap the model.
Why does tool count affect the cache? The tool schema is a fixed block at the front of every request. More tools means a bigger block, and any reorder or toggle invalidates it. Fewer tools means a smaller, more stable prefix.
Is the ¥19 vs ¥900+ gap realistic? It reflects cached vs uncached pricing at DeepSeek's listed rates — the spread is the reason a harness's cache behavior can matter more than the model's base price.
Does this only work for DeepSeek? No — automatic prefix caching is a provider-level feature, but the levers are harness-level. A minimal tool set and append-only sessions improve cache behavior on any provider that offers prefix caching.