KV Cache
What is a KV cache?
A KV cache (key-value cache) is the buffer a transformer keeps, during a single request, of the attention "keys" and "values" it has already computed for every token in the context so far. When the model generates the next token, it doesn't recompute attention over the whole prompt from scratch — it reuses those stored keys and values and only computes the new ones. That's the whole trick that makes autoregressive generation tractable instead of quadratically miserable.
I find it's the piece of LLM internals that people billing for tokens most often confuse with prompt caching. They're related but not the same thing, and the distinction matters once you're paying real money.
Why the KV cache matters in 2026
It matters because the KV cache is where your context window physically lives at inference time, and it scales linearly with the number of tokens in flight. Every token you keep in context is a row in that cache occupying GPU memory for the duration of the turn. Take a long agentic session with a fat system prompt, a pile of tool definitions, and three files the agent re-read for no reason. That is not just expensive on the invoice. It is a chunk of HBM the provider has to hold open for you, and that cost gets passed through.
This is also the mechanism that makes prompt caching possible at the API level. When a provider lets you reuse a stable prefix, what they're really doing is persisting the KV cache for that prefix across requests so they don't have to "prefill" it again. That's why cache reads are billed at a fraction of the input rate rather than free: the keys and values still have to be loaded and attended to, they just don't have to be recomputed. As of October 2026, Anthropic's pricing page lists cache hits at 0.1× input on most Claude models, 0.05× on Opus 5.5 and 0.025× on Fable 5.1. On Opus 5.5 that turns a $4/MTok input prefix into $0.20/MTok on a hit — worth structuring your context to earn. (Writing the cache isn't free either: 1.25× input for a 5-minute entry, 2× for a 1-hour one.)
The newer twist is that the discount itself deepened: earlier Claude models (Opus 4.5 through Opus 5) all billed hits at 0.1×, and the newest top models halved or quartered that. It means the gap between a cache-friendly agent (stable prefix, append-only history) and one that reshuffles its context every turn is now wider in dollars than it was a year ago.
The practical takeaway is the same one I keep coming back to: the cheapest token is the one you never put in the cache. Pruning what the agent reads — via semantic code search instead of whole-file dumps, and output filtering on noisy command logs — shrinks the KV cache, the bill, and the latency at the same time. That's the lever Tokenade pulls automatically, and you can watch the saved tokens add up on the dashboard.
When the KV cache is the wrong thing to optimise
- When the real problem is too much context, not caching. If your agent is slow and pricey because it loads everything every turn, no caching strategy saves you — you need to put less in the window. Fix the input first; see context engineering.
- For one-shot calls with no shared prefix. A single isolated request builds its KV cache, uses it once, and throws it away. There's nothing to reuse and nothing to tune; the cache is just an implementation detail you'll never feel.
- When you'd be hand-managing it yourself. Unless you run your own inference stack, you don't control the KV cache directly — you influence it through what you send. Reach for prompt caching and a smaller context, not for knobs the API doesn't expose.