The Claude Code hooks reference now lists 33 events, from SessionStart to ElicitationResult (Anthropic hooks reference, read 2026-09-24). Most people meet hooks through a list of ten copy-paste examples: a linter after every edit, a desktop ping, a guard on rm -rf. Then they read that a hook cut someone's token usage by 42% or 90% and wonder whether their own hooks help or hurt. Both happen. A hook that prints to stdout on UserPromptSubmit adds its text to the conversation on every prompt you send. A hook that rewrites a tool call before it runs can stop a 3,000-line log from ever reaching the model. The difference sits in two details the example lists skip: which event the hook listens to, and which JSON field it returns. This article maps every token-relevant event to what it adds or removes, using only the official reference, and lists the traps that leave a hook silently doing nothing.
TL;DR
- Four events add plain stdout to Claude's context:
UserPromptSubmit,UserPromptExpansion,SessionStartandPostModelSwitch. On every other event, stdout from a successful hook goes to the debug log and costs nothing. additionalContextis a token cost by design: Claude Code wraps it in a system reminder and inserts it where the hook fired, and it stays in the conversation after that.- Two fields remove tokens:
updatedInputonPreToolUse(rewrite the call before it runs) andupdatedToolOutputonPostToolUse(replace the result before Claude reads it). - Every injected string is capped at 10,000 characters; past that, Claude gets a file path and a 2,000-character preview.
- Exit 1 does not block, a timed-out
PreToolUsehook lets the call through, and a mistyped script path disables a gate without stopping anything.
What a hook is, in cost terms
A hook is a handler that Claude Code runs at a fixed point of a session: a shell command, an HTTP endpoint, a tool on a connected MCP server, an LLM prompt or a subagent (reference). The reference groups the events by cadence. SessionStart and SessionEnd fire once per session. UserPromptSubmit, Stop and StopFailure fire once per turn. PreToolUse and PostToolUse fire on every tool call inside the agentic loop, so their count grows with every file read, search and command the agent runs.
That cadence is the first half of the cost question. A handler that fires once per session and injects 400 characters costs you those characters once, plus their re-reading on each later request. The same 400 characters on PreToolUse land next to every tool result in the session. Anything that enters the conversation is re-sent with the rest of the history on each following model request, which is why prompt caching matters so much for agents: cached re-reads are cheaper than fresh input, but they are never free.
The second half is what the handler returns. A hook can do four things with tokens:
- Nothing. It runs, exits 0, and its stdout goes to the debug log. Most formatting, notification and logging hooks live here, and they have zero token cost.
- Add context: plain stdout on the four events that accept it, or
additionalContextin JSON. - Remove context before it exists: rewrite a tool's arguments with
updatedInput, or replace its result withupdatedToolOutput. - Call a model: prompt-based and agent-based hooks send the hook input to Claude, which is a separate billed request.
The rest of this article sorts the events into those four buckets. If you have never measured where your session's tokens go, do that first; measuring agent token usage shows how to read the usage fields in the session transcript.
The hooks that add tokens
Claude Code's rule for plain stdout is short. On exit 0, "for most events, Claude Code writes stdout to the debug log and doesn't show it in the transcript. The exceptions are UserPromptSubmit, UserPromptExpansion, SessionStart, and PostModelSwitch, where Claude Code adds plain-text stdout as context that Claude can see and act on" (reference, exit code 0).
So an echo in a SessionStart hook is a small, one-time addition. The same echo in a UserPromptSubmit hook is an addition on every prompt. The reference also notes that neither plain stdout nor additionalContext on UserPromptSubmit produces a visible transcript entry: each is injected as a system reminder that starts with the hook's name. You pay for text you never see on screen, which is the most common way a hook setup grows a bill without anyone noticing.
additionalContext is the structured version, and it works on more events. For PreToolUse and PostToolUse it is a "string added to Claude's context alongside the tool result." For SessionStart it is added "at the start of the conversation, before the first prompt." For SubagentStart and PostModelSwitch it is the only thing the hook can do: context, with no blocking (reference, decision control). When several hooks return additionalContext for the same event, Claude receives all of the values.
A third path is less obvious. On PostToolUse and PostToolUseFailure, exiting 2 "shows stderr to Claude; the tool already ran." The reference recommends it as the way to surface a warning from those events. Every warning is text in the conversation, so a linter hook that exits 2 with 80 lines of findings on each edit adds 80 lines per edit.
Stop hooks cost differently. Exit 2 on Stop "prevents Claude from stopping, continues the conversation." That is a whole new model turn, with the full history re-sent. A check that fails on every turn keeps the session going until you intervene. The prompt-based variant has a guard for that: when the model returns impossible with ok: false, Claude Code lets the turn end instead of feeding the reason back. Command hooks have no such guard; the loop is yours to bound.
The pattern to keep: context injection is fine when the text is short, changes the model's next move, and fires at the lowest cadence that still does the job. A branch name and the open issue at SessionStart pass that test. A 2,000-character style guide on every UserPromptSubmit belongs in CLAUDE.md, where it is at least part of the cached prefix.
The hooks that cut tokens
Two fields let a hook shrink what the model reads, and they act at opposite ends of a tool call.
updatedInput on PreToolUse "modifies the tool's input parameters before execution. Replaces the entire input object, so include unchanged fields alongside modified ones" (reference, PreToolUse decision control). This is where most token-saving hooks work. A Bash call to npm test can be rewritten to pipe through a filter that keeps failures only; a Read of a 5,000-line file can be narrowed with offset and limit; a Grep without a path can be scoped to src/. The model only ever reads the trimmed result, and nothing extra enters the context. Two details from the reference matter here. Permission rules are evaluated "against the input your hook returns, not the input Claude sent," so a rewrite cannot smuggle a command past a deny rule. And the replacement is whole-object: forget a field and the tool runs without it.
updatedToolOutput on PostToolUse "replaces the tool's output with the provided value before it is sent to Claude. The value must match the tool's output shape." The reference now recommends it over the older updatedMCPToolOutput, which works on MCP tools only. This is the field to use when you cannot know in advance how large a result will be: let the call run, then keep the error lines of a build log, drop the progress bars, or cut a JSON response down to the keys the model needs. The reference example replaces a Bash result's stdout with a redacted string, which shows the shape.
Note what decision: "block" on PostToolUse does and does not do: it "adds the reason next to the tool result. Claude still sees the original output; to replace it, use updatedToolOutput." A blocking PostToolUse hook meant to hide a huge output adds text and removes nothing.
The published results for this approach are single-team benchmarks, so read them as orders of magnitude. SitePoint reported on 2026-08-20 that Graft, a hook-based framework, cut mean token usage from 8,070 to 4,650 per task on its authors' internal SWE-bench Verified runs, a 42% drop (SitePoint). Spotify Engineering published an engineer's account in September 2026 titled "Portal by Spotify cut my Claude Code token usage by 90%," which argues that most of what a coding agent does is I/O, such as reading five files to answer a question about one method (Spotify Engineering). Neither is a controlled test on your repository. The mechanism behind both is the one described above, and the general technique has its own page: output filtering for command logs.
The 10,000-character ceiling
Every string a hook injects has a hard limit. The reference: "A hook's additionalContext, systemMessage, and initialUserMessage strings, and its plain stdout, are capped at 10,000 characters." Each JSON field is measured separately, plain stdout is measured whole, and several hooks on one event are each measured on their own (reference, JSON output).
Over the limit, Claude Code does not truncate. It "saves the output to a file in the session directory and replaces it with the file path and a preview of up to the first 2,000 characters." Claude can read that file, but Claude Code doesn't ask it to. Large valid Bash results are handled the same way. Unlike the Bash ceiling, the hook cap "has no setting or environment variable to raise it."
Two consequences for cost:
- The cap is a ceiling per string, and nothing stops you from sending 9,999 characters on every prompt. It protects against runaway hooks, and a hook that stays under it can still be the largest single item in your context.
- Crossing the cap can make things worse. If the model decides it needs the full text, it reads the file with a tool call, and you pay for the preview, the tool call and the full content.
If a hook regularly produces more than a few hundred characters, the better design is usually to write the details to a file yourself and inject one line pointing to it.
Hooks that call a model
Prompt-based hooks (type: "prompt") "send the hook input and your prompt to a Claude model, Haiku by default"; the model field overrides it (reference, prompt-based hooks). Agent-based hooks go further and run a subagent that can use tools before deciding. Both are billed requests on top of your main session.
That changes the arithmetic for high-cadence events. A prompt hook on Stop runs once per turn, which is usually tolerable. A prompt hook on PreToolUse runs a model request for every tool call, and each one carries the hook input, including the full tool arguments. For a check that a regular expression or jq can do, a command hook costs nothing in tokens and answers in milliseconds.
Use prompt and agent hooks where judgment is the point: deciding whether a task is really finished before letting Claude stop, or whether a diff touches something a rule forbids. Keep them off the per-tool-call path unless you have measured the extra requests. Agentic coding costs follow the same logic: every fresh model context starts with its own overhead.
PreModelSwitch: the first native cost signal in a hook
For a long time, the complaint in GitHub issue #11008, opened in November 2025, held: hooks received no token or cost data, so a budget hook had to parse the transcript itself. PreModelSwitch changes that for one expensive moment. Its input includes context_tokens (the tokens the next request re-sends as its prompt), prompt_cache_warm, cache_ttl, and estimated_cache_write_usd, the estimated cost of writing that context to the new model's cache (reference, PreModelSwitch input).
Why it matters: a model switch forfeits a warm prompt cache. The reference's own example shows a /model opus request from a Sonnet 5 session with 182,340 context tokens and an estimated cache write of $1.14 at list price. A hook can return permissionDecision: "ask" with that figure in the reason, so the switch happens with the price in front of you. The default timeout for this event is 30 seconds, and unlike PreToolUse, a PreModelSwitch hook that times out blocks the switch.
For every other event, the transcript remains the source. Each hook receives transcript_path, and the reference warns that the file "is written asynchronously and may lag the in-memory conversation." A cost hook reading it should expect the latest turn to be missing. For per-session numbers without writing a hook, see how to reduce Claude Code token usage, which covers the existing counters.
Common mistakes
- Using exit 1 to block. "Without valid JSON on stdout, Claude Code treats exit code 1 as a non-blocking error and proceeds with the action." Policy hooks must exit 2, or return a JSON decision.
- Trusting a slow gate. A
command,httpormcp_toolhook onPreToolUsethat reaches its timeout "doesn't block the tool call." If your filter shells out to something slow, the unfiltered call goes through. - A typo in the script path. The shell exits 127, Claude Code shows a non-blocking notice, and "a mistyped path in
settings.jsonleaves the gate silently disabled." Watch the first run. - Matching MCP tools with a bare prefix.
mcp__memoryis compared as an exact string and matches no tool;mcp__memory__.*matches all of that server's tools. - Printing a banner from the shell profile. Stdout must contain only the JSON object. A profile that echoes on startup breaks parsing, and the JSON fields are ignored.
- Expecting
suppressOutputto save anything. The reference says it "has no effect": Claude Code accepts the field and does not act on it. Successful hook stdout already stays out of the transcript on most events.
What to do this week
- List your hooks by event. Open
/hooksand write down every handler onUserPromptSubmit,SessionStart,PreToolUseandPostToolUse. For each, note whether it prints plain stdout, returnsadditionalContext, or exits 2 with stderr. Those are your token-adding hooks. - Move per-prompt text to a lower cadence. Anything on
UserPromptSubmitthat does not change between prompts goes toSessionStartor to CLAUDE.md. Expected outcome: the same guidance, paid once instead of per prompt. - Add one output-shrinking hook to your noisiest tool. Check the transcript for the tool whose results are largest (usually
Bashrunning tests or builds) and add aPostToolUsehook that returnsupdatedToolOutputwith errors and the summary line only. Compare total input tokens over two similar sessions before and after. - Make every policy hook exit 2 and test it. Run one command it should block and confirm the block appears. A gate you have not seen fire is not a gate.
- Add a PreModelSwitch hook with
"ask". Quotecontext_tokensandestimated_cache_write_usdin the reason, so a mid-session model switch shows its price before it happens.
FAQ
What are Claude Code hooks?
Hooks are handlers Claude Code runs automatically at fixed points of a session: before and after tool calls, when you submit a prompt, when a session starts or ends, before compaction, and about thirty other events. A handler can be a shell command, an HTTP endpoint, an MCP tool, an LLM prompt or a subagent. It receives the event as JSON (on stdin for commands) and can return a decision, extra context, or a rewritten tool input or output. You configure them in settings files or through the /hooks menu.
Do Claude Code hooks use tokens?
Only when their output enters the conversation or when they call a model. On most events, stdout from a successful hook goes to the debug log and costs nothing. Plain stdout on UserPromptSubmit, UserPromptExpansion, SessionStart and PostModelSwitch is added to context, additionalContext is added wherever it is supported, and stderr from an exit-2 PostToolUse hook is shown to Claude. Prompt and agent hooks are separate model requests, Haiku by default for prompt hooks.
Can a hook reduce Claude Code token usage?
Yes, through two fields. updatedInput on PreToolUse rewrites a tool call before it runs, for example narrowing a file read or piping a test command through a filter. updatedToolOutput on PostToolUse replaces a tool's result before Claude reads it. Published results come from single-team benchmarks: SitePoint reported a 42% drop per task for Graft on its authors' own SWE-bench runs in August 2026. Measure on your own sessions before trusting a figure.
What is the difference between PreToolUse and PostToolUse?
PreToolUse fires before a tool call executes and can block it, ask for confirmation, or rewrite its arguments with updatedInput. PostToolUse fires after a call succeeds; it cannot undo the call, but it can add context, show a warning via exit 2, or replace the result with updatedToolOutput. For token savings, PreToolUse avoids producing a large output at all, and PostToolUse handles output whose size you cannot predict.
Why does my hook not block anything?
The usual causes are exit code 1 instead of 2, a script path that does not exist, a matcher that matches no tool (an MCP prefix without .*), text printed before the JSON on stdout, or a PreToolUse hook that timed out. The reference is explicit that a timed-out command hook on PreToolUse lets the tool call continue. Enable debug logging to see the parse or validation message, and test each gate with a call it should stop.
Is there a size limit on hook output?
Yes. additionalContext, systemMessage, initialUserMessage and plain stdout are each capped at 10,000 characters. Above that, Claude Code writes the text to a file in the session directory and gives Claude the path plus a preview of up to 2,000 characters. No setting or environment variable raises this cap. The notes a hook attaches for the auto mode classifier have a separate 2,000-character cap per tool call.
Can a hook see how many tokens the session has used?
Not directly on most events: GitHub issue #11008 asked for token and cost data in hook inputs. Every hook receives transcript_path, and the transcript's usage fields can be summed, keeping in mind that the file is written asynchronously and may lag the latest turn. PreModelSwitch is the exception: its input carries context_tokens and estimated_cache_write_usd for the switch it is about to allow or block.
Should I use prompt-based hooks?
Use them where a judgment call is the point, such as checking before Stop whether the task is really done. Avoid them on PreToolUse or PostToolUse unless you have measured the cost, because they add one model request per tool call. Anything a regular expression, jq or a short script can decide belongs in a command hook, which costs no tokens.
Related reads
- Output filtering for command logs: the technique most token-saving
PostToolUsehooks implement. - How to reduce Claude Code token usage: the wider list of levers, hooks included.
- CLAUDE.md token bloat: where static guidance belongs instead of a per-prompt hook.
- Prompt caching and your input bill: why injected context keeps costing after it lands.
- How to measure agent token usage: read the transcript before and after adding a hook.
- Claude Code skills and their cost: the other way to load instructions only when needed.
- Lazy MCP loading: the same idle-cost problem, on the tool-definition side.
- Best Claude Code token optimizers: tools that package these hooks for you.
FAQ
Cut your AI coding agent's token bill.
Ranked #1 on the Token-Harness Optimizer Leaderboard. Zero config.
Tokenade cuts the token bill of AI coding agents: one install, zero config, and it trims what your agent sends to the model.
$ npm install -g @tokenade/cli$ tokenade install$ tokenade login