Claude Code /compact: What It Costs and What Survives

A warm /compact costs cents, a cold one costs dollars on the same history. What compaction keeps, what it drops, and when /clear or /rewind is the cheaper move.

Profile photo of Paul Irolla

By Paul Irolla

Founder · AI & developer tools · Tokenade

Ph.D. in AI · builds token-optimization tooling for AI coding agents

View author page
Updated 14 min read
Cite this page

The same /compact can cost you a few cents or a couple of dollars, and the context size does not decide which. The cache does. Anthropic's own prompt caching documentation for Claude Code says it plainly: to write the summary, Claude Code sends a separate request carrying your whole history, and "while the cache is warm, that request reads your prefix from the cache, so a mid-session /compact costs a fraction of what the context size suggests". After a break longer than the cache lifetime, the same request "reprocesses the full history as uncached input". That is why /compact costs the most right after you resume an old session.

Most guides on Claude Code compaction stop at "it summarizes your conversation" and a rule of thumb for when to run it. This one puts the cost mechanics beside the survival rules, and names the alternatives, so you can tell which move is cheap at the moment you are about to make it.

TL;DR

  • /compact sends one extra request containing your full conversation plus a summarization instruction. While the prompt cache is warm it reads that history at cache price; once the cache has expired it pays full input price.
  • The turn after compaction is cheap: it only rebuilds the cache for the short summary.
  • CLAUDE.md, auto memory, the plan and a fresh git status come back from disk. Path-scoped rules, nested CLAUDE.md files and anything a hook injected earlier are summarized away.
  • Auto-compact fires near the model's limit: around 967K tokens on 1M-window models, at 200K on 200K-window setups. /autocompact moves that point.
  • /clear costs nothing, /rewind reuses a cache that already exists, and /recap summarizes without touching the cache. Pick /compact when you need continuity and the cache is still warm.

What /compact actually does

Compaction replaces your message history with a summary written by the model. Anthropic's session management post (Thariq Shihipar, April 15, 2026) describes it as "lossy, but you didn't have to write anything yourself", and you can steer it with instructions such as /compact focus on the auth refactor, drop the test debugging.

Mechanically, the prompt caching page lists four steps:

  1. Claude Code sends a separate request with the same system prompt, tools and history as your conversation, with the summarization instruction appended as a final user message.
  2. The model writes the summary. As of v2.1.198 that request inherits your session's extended thinking setting, according to the context window page.
  3. Claude Code swaps the history for the summary and reloads project context (CLAUDE.md, memory) from disk.
  4. Your next turn builds a new cache entry for the shorter history.

Two consequences follow. First, compaction invalidates the conversation layer of the cache by design, because the new history shares no prefix with the old one. The system prompt layer survives. Second, because CLAUDE.md is reloaded from disk, any edit you made to it mid-session takes effect at that point: the docs note that CLAUDE.md is "read once at session start" and that a new version "loads on the next /clear, /compact, or restart". If you edited CLAUDE.md since the session started, that reload is a cache miss on the project-context layer too.

In a fresh session, /compact does nothing useful. The costs page says it prints Not enough messages to compact. when there is no history to summarize.

If you want the broader picture of why every turn re-sends the whole conversation, our explainer on how the context window drives your token bill covers the base mechanism this section builds on.

What compaction costs: warm cache vs cold cache

The number that matters is not your context size. It is whether the cached prefix is still alive when the summarization request goes out.

Cache lifetime depends on how you pay. The TTL table in the Claude Code docs splits requests into two buckets:

Request bucketClaude subscription, within plan usageUsage credits, API key, or cloud provider
Main conversationOne hourFive minutes
Everything else (subagents, forks, compaction, titles)Five minutes, except some server-controlled helpersFive minutes

Compaction sits in the second bucket, but the docs specify that the summarization call uses the same prefix-sharing approach as subagents, so it reads the cache your main conversation built. What decides whether that cache exists is how long you have been idle: under an hour on a subscription, under five minutes on an API key.

Here is what that means in dollars, using Anthropic's published prompt caching prices for Claude Opus 5.5 as of October 1, 2026: $4 per million input tokens, $0.20 per million for cache hits, $5 and $8 per million for 5-minute and 1-hour cache writes. This is arithmetic on list prices, not a measured bill, and it leaves out the output tokens of the summary itself.

  • 400K tokens of history, warm cache: 400,000 × $0.20 / 1M = $0.08 of input for the summarization request.
  • 400K tokens of history, cold cache: 400,000 × $4 / 1M = $1.60 of input, twenty times more for the same command.

On Sonnet 5.5 ($2 input, $0.20 cache hit per million on the same page), the same cold compaction costs $0.80 of input and the warm one stays at $0.08. The gap between warm and cold is a ratio of input price to cache-hit price, so it widens on the more expensive models.

On a Pro or Max plan you do not see dollars, but the same tokens count against your limits. The costs page lists compaction among the reasons usage climbs in a long session: "/compact reads the conversation it summarizes, so compacting a large context is itself a large request." For the plan side of this, see how Claude usage limits work.

The practical rule that falls out: compact while you are still working, not after lunch. If you come back to a large session after the cache has expired, Claude Code on Pro or Max offers to resume from a summary, which is the same trade made at the cheapest available moment.

What survives compaction, and what quietly disappears

The context window documentation gives a per-mechanism table. Condensed:

Kept or restoredLost into the summary
System prompt and output stylePath-scoped rules (paths: frontmatter) until a matching file is read again
Project-root CLAUDE.md and unscoped rules, re-read from diskNested CLAUDE.md files in subdirectories, same condition
Auto memory, re-read from diskContext that hooks injected earlier
The plan written in plan mode, re-read from diskEverything else in the conversation not captured by the summary
A fresh git status snapshot
Up to five recently modified files Claude read or edited
Invoked skill bodies, capped at 5,000 tokens each and 25,000 in total

Three details in that table cause most of the "Claude forgot X after compaction" reports.

Files over 5,000 tokens come back as a path only. Claude Code re-reads up to five files, most recently modified first, but a large file returns as Referenced file without its content. If the model needs that file, it reads it again, and you pay for the read again. Our guide to skeleton compression for file reads covers ways to make those re-reads smaller.

Skills get truncated from the bottom. Truncation keeps the start of SKILL.md, and the oldest invoked skills drop first once the 25,000-token budget is exceeded. The docs advise putting the most important instructions near the top of the file. More on what skills cost in context in Claude Code skills: what they are, what they cost.

Rules scoped to paths vanish until triggered again. A rule with paths: frontmatter loads into message history when its trigger file is read, so compaction summarizes it away. If a rule must hold for the whole session, the docs say to drop the paths: frontmatter or move it to the project-root CLAUDE.md. That moves its tokens into every turn, so it is a trade worth making only for rules that matter. How a bloated CLAUDE.md inflates your bill shows what that trade costs.

When auto-compact fires

Auto-compact is the same pass as /compact, triggered for you. The model configuration page lists the defaults:

  • Models running with a native 1M window compact "at about 967K tokens by default". On the Anthropic API that covers Sonnet 5, the Fable models and Opus 4.7 and later.
  • Sonnet 4.6 and Opus 4.6 without extended context compact at the 200K boundary, as do Opus 4.8 and later when they run with a 200K window on Bedrock, Google Cloud or Foundry.
  • Setting CLAUDE_CODE_DISABLE_1M_CONTEXT=1 holds 1M models to 200K, and they compact there.
  • An unrecognized model ID, such as a gateway alias, compacts at the window Claude Code assumes for it.

You can move the trigger with /autocompact 500k (saved to your user settings as autoCompactWindow), with the --autocompact flag for a single launch, or with the CLAUDE_CODE_AUTO_COMPACT_WINDOW environment variable, which overrides both. The command and flag accept 100K to 1M tokens.

Waiting for the default has a quality cost on top of the token cost. The Anthropic post on session management puts it this way: "due to context rot, the model is at its least intelligent point when compacting." It also explains bad compactions: they happen "when the model can't predict the direction your work is going", for example when a long debugging session gets summarized and your next message asks about a warning the summary dropped. Our glossary entry on context rot covers the underlying effect.

There is also a timing cost. Auto-compact fires mid-task, wherever the threshold falls. The prompt caching page recommends running /compact "at a natural break in your work, such as between tasks, instead of waiting for auto-compaction to trigger mid-task", so you choose when the overhead lands.

/compact, /clear, /rewind, /recap: which one, when

Four commands shrink or reshape what the model sees. They differ in cost and in what they keep.

/clear starts a new conversation. The costs page is direct: "When you want a fresh start instead of continuity, /clear costs nothing." The price is effort: you write the brief yourself ("we're refactoring the auth middleware, the constraint is X, the files that matter are A and B"), and the model starts from what you decided was relevant. Anthropic's rule of thumb is that a new task deserves a new session.

/rewind (double Esc) truncates back to an earlier turn. Because "the remaining history is the same content the cache was built from at that point", the next request hits the earlier cache entry. Use it when Claude went down a wrong path: keep the useful file reads, drop the failed attempt, re-prompt with what you learned. /rewind also offers "Summarize from here" and "Summarize up to here" to compact part of the conversation.

/recap generates a summary for display in your terminal. Unlike /compact, it appends the summary as command output rather than replacing your history, so the cached prefix stays intact. It reorients you; it does not free any context.

/compact keeps continuity without you writing anything. It is the right call mid-task when the session is full of stale exploration you will not need again, and the cache is still warm.

A fifth option prevents the problem: send work that produces a lot of throwaway output to a subagent, so the file reads stay in its context window and only the result comes back. Anthropic's test is "will I need this tool output again, or just the conclusion?" Subagents have their own startup cost; measuring AI agent token usage explains how to check whether delegation paid off in your session.

Steering what the summary keeps

You have three levers on the summary's content, all documented.

Instructions on the command line. /compact Focus on code samples and API usage tells Claude what to preserve. The costs page uses this exact example.

A standing section in CLAUDE.md. The same page shows a # Compact instructions heading in the project-root CLAUDE.md, for example "When you are using compact, please focus on test output and code changes". Since project-root CLAUDE.md is re-injected after every compaction, these instructions stay available for every summary.

A SessionStart hook on the compact source. The hooks guide shows a SessionStart hook with "matcher": "compact" whose stdout is added to the compacted context:

{
"hooks": {
"SessionStart": [
{
"matcher": "compact",
"hooks": [
{
"type": "command",
"command": "git log --oneline -5"
}
]
}
]
}
}

The guide suggests git log --oneline -5 as a dynamic replacement for a static echo. Whatever this hook prints, you pay for it on every turn until the next compaction, so keep it short. PreCompact and PostCompact hooks also exist, with manual and auto matchers, if you want to log or react to each compaction. Hooks are one of several levers listed in our guide to reducing Claude Code token usage.

Measuring what compaction did to your session

Claude Code reports cache behaviour directly. Since v2.1.251, /usage adds a Prompt cache (main) line with the request count, the share of input served from cache, misses, and "expected rebuilds", which the costs page defines as misses caused by Claude Code itself rewriting the conversation "by compaction or by clearing old tool results". The example line in the docs reads 1 expected rebuild (compaction or tool-result clearing).

On a Pro, Max, Team or Enterprise plan, the /usage breakdown also flags behaviours such as long context or cache misses when one accounts for 10% or more of recent usage. If compaction keeps showing up there, the fix is usually timing (compact earlier, while warm) or scope (clear between tasks).

/context gives a live breakdown of what fills the window by category, including which CLAUDE.md and memory files loaded, per the context window page. Run it before and after a compaction to see what came back. For a tool-independent way to count tokens across sessions, see how to measure AI agent token usage and our Claude Code token usage statistics.

Common mistakes

  • Compacting right after resuming an old session. The cache has expired, so the summarization request pays full input on the whole history. Use the "resume from a summary" offer on Pro or Max, or /clear with a short brief.
  • Using /compact to switch tasks. You pay for a summary of work you no longer need, and the summary carries it into the new task. /clear is free and cleaner.
  • Correcting a failed attempt instead of rewinding. "That didn't work, try X" keeps the failed attempt in context. /rewind drops it and hits a cache that already exists.
  • Relying on path-scoped rules across a long session. They disappear at compaction until a matching file is read again. Move critical ones to the project-root CLAUDE.md.
  • Burying key instructions at the bottom of a skill. Re-injected skills are truncated from the end once they pass 5,000 tokens.
  • Letting auto-compact pick the moment. It fires mid-task, when context rot is at its worst, and summarizes without knowing where you are going next.

What to do this week

  1. Add a # Compact instructions section to your project-root CLAUDE.md. Two lines naming what the summary must keep in this repo (test output, the files under active change). Expected result: fewer "it forgot X" moments after auto-compact, at a few dozen tokens per turn.
  2. Lower your auto-compact window if you work on a 1M model. /autocompact 500k is the example the docs use. Expected result: compaction runs earlier, on a smaller and less rotted context, so the summarization request is cheaper and better.
  3. Run /usage at the end of a long session. Look at the Prompt cache (main) line: misses versus expected rebuilds. If misses dominate, the cache is expiring between your turns; compact before breaks, not after.
  4. Make /clear your default between tasks. Keep /compact for mid-task cleanup while the cache is warm, and /rewind for dead ends.

FAQ

What does /compact do in Claude Code? It replaces your conversation history with a summary the model writes. Claude Code sends a separate request with your full history and a summarization instruction, swaps the history for the result, then reloads CLAUDE.md, auto memory, the plan and a fresh git status from disk. Up to five recently modified files are re-read, and invoked skills are re-injected within a 5,000-token-per-skill and 25,000-token total budget, per the context window docs.

Does /compact cost tokens? Yes. The summarization request carries your whole conversation. While the cache is warm it reads that history at cache-hit price, which on Opus 5.5 is $0.20 per million tokens against $4 for uncached input (pricing as of October 1, 2026). After the cache expires, it pays full input price on the whole history. The summary's output tokens come on top.

Is /clear cheaper than /compact? Yes. The costs page states that /clear costs nothing. The trade-off is continuity: you write the brief for the next session yourself, while /compact writes it for you. For unrelated work, /clear is the better choice on both cost and quality.

When does auto-compact trigger? On models with a native 1M window, at about 967K tokens by default. On 200K-window setups, at the 200K boundary. An unrecognized model ID compacts at the window Claude Code assumes for it. You can change the point with /autocompact, the --autocompact flag or CLAUDE_CODE_AUTO_COMPACT_WINDOW, between 100K and 1M tokens (model configuration docs).

Why did Claude forget something after compaction? Either the summary dropped it, or it came from a mechanism that does not survive: path-scoped rules, nested CLAUDE.md files, context injected by hooks, or the content of a file over 5,000 tokens. Anthropic notes that bad compactions happen when the model cannot predict where the work is going, so passing instructions to /compact helps.

Does editing CLAUDE.md mid-session take effect after /compact? Yes. Project-root and user-level CLAUDE.md files are read once at session start and held in memory; a new version loads on the next /clear, /compact or restart, according to the prompt caching docs.

Can I keep important context across compactions? Put it in the project-root CLAUDE.md (re-read from disk every time), in a # Compact instructions section, or in a SessionStart hook with the compact matcher, whose output is added after each compaction. Each of these adds tokens to every turn, so keep them short.

What is the difference between /compact and /recap? /compact replaces your history with a summary and frees context. /recap prints a summary in your terminal and appends it as command output, leaving the history and the cached prefix untouched. Use /recap to reorient, /compact to make room.

FAQ

Cut your AI coding agent's token bill.

Ranked #1 on the Token-Harness Optimizer Leaderboard. Zero config.

Tokenade cuts the token bill of AI coding agents: one install, zero config, and it trims what your agent sends to the model.

$ npm install -g @tokenade/cli
$ tokenade install
$ tokenade login