dshbase

Blog · Analysis

Why X keeps calling DSH token efficient

September 28, 2026 · dshbase · X pulse

“DSH is token efficient” is having a week on X. The phrase sounds like a property of the product, the way a car has a mileage rating. It is not. Under the posts we actually read, three different claims are sharing one slogan: the prefix cache is being kept warm, someone switched to Minimal, or someone pointed the harness at a Flash-class model and compared it with a heavier one. Those levers are real. They are not automatic, they are not the same lever, and a fourth idea — that PTC mode is a discount code — keeps sneaking into the replies. It is not.

This note separates the claims, points at the on-site explainers that already measured the shape of the problem, and ends with the control that the slogan never includes: a cost cap. We are not publishing a new price table. If a number is not already in a post we link, it does not belong in a rumor roundup.

Three sentences hiding inside one compliment

@TheInsiderrz and @frede_rico sit in the cluster that talks about DSH as if the savings were baked in. Read them as mood, then ask which lever they touched. @BranchGhost is more specific and more useful: a preference for keeping the cache, which is an operational habit rather than a model review. On the other side, @ihero1020 is the counterexample the slogan needs. A lot of tokens got spent. Same family of tools, opposite anecdote. If both posts can be true, “efficient” is not a label you can stamp on the binary.

@bitflipgremlin makes the comparison that actually teaches something: the same model, with Minimal and a different chain-of-thought setting, moves the outcome a lot. That is one person’s run, not a benchmark we are adopting. The direction matches what the harness is designed to do. Minimal is a smaller tool surface and a frozen prompt. Change the surface and you change both the behavior and the bill. The model card did not change between those runs. The request prefix did.

The harness changes the tool surface, and the cache notices

DeepSeek-style prefix caching reuses the front of a request when it matches a previous one byte for byte. Tool schemas live in that front. So does the system prompt. So does a plugin that injects a date, a random ordering, or a “helpful” summary of yesterday. The cost and caching guide walks the levers: few tools in a fixed order, append-only sessions, a static prompt, intermediate data kept out of the transcript. The Pi note is the extreme version of the same idea — a tiny tool list and a session that only grows — and it is why people learned to talk about hit rate at all.

DSH is efficient when you let it be boring. It is expensive when every turn rearranges the prefix. Adding a plugin is not free just because the plugin is small: its schema joins the prefix for every later step. Reordering tools between turns is a cache miss you paid for on purpose. Rewriting history to “save context” can destroy the cache that was saving you more than the rewrite removed. The append-only session log is an audit feature first, and it is also the cache-friendly shape. Fight it and you will feel clever for a turn and poorer for the session.

@BranchGhost’s preference is the practical form. Decide the tool list before the task, leave it alone while the task runs, and treat a mid-task plugin install as a new session with a cold cache. dsh --profile web --dump-config is the list to diff when the bill surprises you.

Minimal versus Standard

The four presets are in the modes guide. Standard is the full coding agent: files, shell, search, skills, planning, sub-agents. That surface is why Standard feels capable, and why its prefix is harder to keep stable. Minimal cuts this to a persistent bash and a file editor, freezes the system prompt, and turns off most of the rest. It exists so you can compare models without the harness doing the work. It is also, incidentally, the cheapest shape, because the prefix is small and still.

@bitflipgremlin’s observation belongs here. Minimal plus a tighter chain of thought is not a skin. It removes tools the model would otherwise have called, and it removes the tokens those calls would have pasted back into the transcript. People experience that as “DSH got smarter and cheaper” when what happened is “the run was no longer allowed to wander.” Use Minimal to measure, or for a narrow repair where bash and an editor are honestly enough. If you compare two settings, keep the model id identical and write the mode next to it. Otherwise you will publish a model review that was actually a prompt review.

Flash versus Pro, without a fake price list

The other half of the slogan is model choice. Flash-class and Pro-class models do not cost the same and do not fail the same way, and X collapses that into the harness because the dropdown lives there. Our three-model fix test already showed the spread on one bug: a fast path, a Flash path, a Pro path that stalled. The V4.1 Flash beta note was one checkpoint with an expiry in the id, not a permanent coupon. Do not cite it as the model you are running unless your config still says so. If the invoice dropped after a model change, thank the price, then check whether the tools changed too. If the id stayed and the invoice dropped, thank the cache and the mode. This week’s V4.1 talk is tangled with the 0.2 rumor, including a hope that a V4.1 Pro “for DSH” arrives the same day. That pairing is not a pricing plan. Pin the id you have, and compare a new one on a trajectory you already ran — the method in the Space Bunny note.

PTC does not automatically save tokens

Programmatic tool calling lets the model write a small program that chains tool operations inside one run_code, so intermediate files and search hits stay out of the transcript. When the task is a repetitive pipeline, that can cut round-trips and protect the cache. When the task is a judgment call every step, PTC adds a planning tax: the model writes code, the code fails, the model explains the failure, and you have paid for a meta-loop. The orange-book field notes say this plainly, and the modes guide says it again: PTC is not the cheap preset. Minimal is. PTC is the preset you turn on after you have watched Standard thrash on the same tool pattern.

A reply that says “just use PTC, it’s efficient” is spending your money. Measure a single repeated task both ways. Count input, cache hits if your provider reports them, and output. If PTC’s fixed overhead is larger than the round-trips it removed, switch back. The slogan does not get a vote.

Set a cost cap before you believe the slogan

The first-month report already flagged the hole: the loop can run, and the product still leaves the maximum bill as an exercise for the reader. One discussion described a run that continued until the account was empty. That is not a cache miss. That is an unbounded agent. Efficiency talk without a stop condition is how the counterexample in @ihero1020’s lane happens to people who thought they were in @BranchGhost’s lane.

Practical caps, none of which require a secret setting:

  • A provider-side budget or a key that is not your production key. When it stops, the experiment stops.
  • A session you are willing to abandon. Do not “continue” a runaway thread to save the cache; the cache is not worth the tail.
  • A tool list you wrote down. If the agent starts asking for plugins mid-task, that is a new decision, not an automatic yes.
  • A hop limit when more than one agent is in the room. The shared-board note is about coordination spend, which is token spend with extra steps.
  • A human check before Create mode or a team is allowed to edit the same files the budget is funding.

Write the cap next to the mode and the model id. “We use DSH because it is efficient” is not a control. “This profile is Minimal, this id, this monthly key, stop at N” is a control. The X posts are a reason to check which of those you have. They are not a measurement.

X sources

Community signal only. Not a pricing sheet and not an official efficiency claim.

  • @TheInsiderrz — part of the “DSH saves tokens” cluster.
  • @frede_rico — same cluster; read as mood until a mode and a model id are named.
  • @bitflipgremlin — same model, Minimal and chain-of-thought changed the run a lot.
  • @BranchGhost — preference for protecting the cache.
  • @ihero1020 — counterexample: a run that spent a lot of tokens anyway.

All articles →