Docs / How it works

How savings are measured

Updated

TL;DR — Every compaction writes a row to `~/.tokenade/gain.jsonl` with tokens before and after. Savings are capped at what the agent would really have seen, negligible rows are dropped, and estimates (code intelligence, re-reads avoided) are kept in a separate ledger and never mixed into `saved:`.

Tokenade measures each saving where it happens: the size of the output before it acted, the size after, and the difference. It then corrects that figure for what your agent would really have received, and keeps anything it cannot measure directly in a separate, clearly labelled estimate. This page explains each step, so you can judge the numbers tokenade gain and the dashboard show.

The ledger

Every time Tokenade changes what reaches the model, it appends one row to ~/.tokenade/gain.jsonl. All projects write to the same file. Each row records when it happened, the channel (hook, wrap, proxy…), what acted, the agent and model when known, the working directory, the program name, and the tokens before, after and saved.

The file rotates at 10 MiB and keeps the last 30 days of rows. Your totals over a longer period are kept in a rollup.

Counting tokens

Tokens are counted locally, accurate to within a few percent. The same counter is used on both sides of every row, so before and after are always comparable. No text is sent anywhere to be counted.

Three corrections applied to every row

1. Only real savings are booked

  • If the output after Tokenade is larger than before, the saving is zero (and the raw output was returned anyway).
  • Negligible savings are not recorded: that is measurement noise, not a saving.
  • Implausible rows are zeroed and logged.

2. The agent's own limits are taken into account

Agents already truncate or refuse very large tool results. Counting the full raw size as "saved" would overstate the benefit, so Tokenade caps each row at what the agent would actually have delivered to the model.

When the raw output would have been truncated, only the part the agent would have seen counts, plus a limited credit for the extra reading a truncated or refused result usually triggers. That credit is flagged as approximate: tokenade gain shows it on its own line, and a receipt reports it separately.

3. What Tokenade adds is charged

Tokenade adds a little text of its own: the style note, the rules it writes into your agent's instructions file. That cost is recorded as overhead and subtracted, on every request that carries it, on every agent.

Measured versus estimated

Some benefits can't be measured after the fact. They go to a separate ledger, ~/.tokenade/gain_estimated.jsonl, with every row marked "estimated": true, and they are never added to saved:.

EstimateWhy it can't be measured exactly
Code intelligence (skeleton, query, map, impact, semantic)The input is measured, but the alternative (reading the whole file) is assumed
Re-reads avoidedA folded result keeps saving on every later request that would have re-sent it; counted conservatively
Batching and recoveryFewer round-trips and fewer retries are inferred, not observed

tokenade gain shows these on est.: lines below the measured total. The dashboard shows them as "estimated", next to the measured figure, never inside it.

The LLM proxy

For agents served by the LLM proxy, only turns the provider answered successfully are booked. Each folded block is counted once, as measured, the first time it is sent. When later requests re-send it folded, those are booked as estimated re-reads, conservatively. The style note the proxy adds is charged on every request that carries it.

Money figures

The dashboard prices each row at the rate of the model that actually ran it, using the model recorded on the row or, failing that, the session's model. Input, cache-read and output rates are applied to the matching tokens, and overhead is subtracted.

Local figures and your account

tokenade gain and the dashboard read your local ledger. Your tokenade.net account counts savings from the usage events Tokenade reports (numbers and program names only; see what Tokenade sends). That server count, shown as metered: in tokenade gain, deduplicates more coarsely, so it can be slightly lower than the local saved:. Estimated rows are never counted toward your plan's quota. Plan limits are on the pricing page.

Check the numbers yourself

tokenade gain --history # one line per ledger row
tokenade gain --json # totals, by_op[], by_source[]
tokenade receipt --json # signed summary, verifiable with --verify