Docs / Commands
The LLM proxy
Updated
tokenade llm-proxy is a local HTTP proxy between your agent and its model provider. On the way out it folds tool results that are already in the conversation history, removes exact duplicates and scrubs credentials; on the way back it reads the provider's own token usage. It is the fallback for agents, or parts of agents, that hooks can't reach.
Where it fits
Tokenade reaches an agent through the best channel available, in this order:
- hooks (or a plugin bridge), which act before a tool result is ever sent;
- the shell wrapper (
tokenade wrap); - the MCP proxy (
mcp-wrap); - the LLM proxy.
The LLM proxy is last because it can't change the turn on which a raw tool result first arrives: it removes that result's cost from every later turn. When an agent's hooks do run, the proxy only covers what they don't, and nothing is done twice.
Usage
[--remote] [--no-compaction]
tokenade llm-proxy install --agent <name> [--listen <addr>] [--autostart]
tokenade llm-proxy uninstall --agent <name>
tokenade llm-proxy status
Alias: llm_proxy.
| Flag | Default | Effect |
|---|---|---|
--upstream <url> | Provider base URL to forward to. Without it, llm-proxy serves every agent recorded by install | |
--listen <addr> | 127.0.0.1:8787 | Address to bind. install gives each agent its own port (8787, 8788, …) |
--agent <name> | Agent this port serves; savings are credited to it | |
--remote | off | Required to bind anything other than loopback. The process holds your API key, so a non-loopback bind is reachable from other machines |
--no-compaction | off | Relay without folding |
--autostart | off | With install: write the platform's service definition and print the commands that enable it |
--enable | off | With install --autostart: also run those commands |
Point an agent at it
install first says what it will change and where traffic will go, then rewrites the agent's configuration to use the local proxy:
| Agent | What is changed |
|---|---|
| Codex (CLI and app) | ~/.codex/config.toml: model_provider = "tokenade" and a [model_providers.tokenade] base_url |
| Claude Code | ANTHROPIC_BASE_URL in the env block of settings.json. Used for exact cost figures only; payloads are relayed untouched, because hooks already do the folding |
| Qwen Code, OpenClaw | OPENAI_BASE_URL in the agent's settings |
| Copilot CLI | COPILOT_PROVIDER_BASE_URL |
| Cursor | CURSOR_API_ENDPOINT in your shell profile |
| Others | Aider, OpenCode, Kilo Code, Grok, Cline, Hermes, Droid, Pi, Copilot in VS Code, Devin Desktop: their own config file |
Claude Desktop, Claude Cowork and the Codex app as a separate target are refused. The Codex app shares ~/.codex/config.toml with the CLI, so --agent codex covers both.
uninstall --agent <name> restores the agent's original endpoint.
aider ~/.aider.conf.yml on 127.0.0.1:8787 → https://api.openai.com/v1
grok ~/.grok/config.toml on 127.0.0.1:8788 → https://api.x.ai/v1
qwen-code ~/.qwen/.env on 127.0.0.1:8789 → https://openrouter.ai/api/v1
aider
(— = not supported · partial = partly · yes = supported)
Command compaction partial → yes
Web search partial → yes
…
status lists each installed agent, its port and upstream, and which features the proxy adds on top of what the agent already gets.
Automatic setup
tokenade install, tokenade upgrade and the daily background check set up the proxy on their own, but only for agents where it adds something the agent can't get otherwise (for example Codex while its hooks are not approved). Three guards apply:
- an agent already covered by its hooks is left alone;
- the proxy is started and asked whether it answers; if it doesn't, the agent's configuration is put back exactly as it was;
tokenade llm-proxy uninstall --agent <name>is remembered (~/.tokenade/llm-proxy-declined), so the automatic setup never undoes your choice. Runningllm-proxy installfor that agent again clears the refusal.
Start at boot
A base URL pointing at a stopped proxy makes every turn fail, so the proxy needs to run whenever the agent does. With --autostart, each agent's proxy gets its own service:
| System | Service |
|---|---|
| Linux | systemd user unit ~/.config/systemd/user/tokenade-llm-proxy*.service. Run loginctl enable-linger so it survives a reboot without a login |
| macOS | LaunchAgent plist in ~/Library/LaunchAgents |
| Windows | Scheduled Task |
Tokenade prints the commands that enable the service; pass --enable to have it run them. It never runs sudo itself: any command that needs administrator rights is printed for you to run. Adding --autostart to an agent that is already installed keeps the existing setup and adds the startup entry.
Codex is set up automatically
Codex only runs hooks you have approved in Codex. tokenade install marks Tokenade's Codex hooks as trusted (tokenade codex-trust re-applies that), and also sets up the LLM proxy for Codex, CLI and desktop app, as described above. Until the hooks actually run, the proxy carries compaction, credential scrubbing, the style note and savings figures; once they run, the proxy only measures. Codex keeps using your ChatGPT sign-in through the proxy.
What the proxy does to a request
- Understands the Anthropic Messages, OpenAI Chat and OpenAI Responses formats. Anything else is relayed byte for byte.
- Folds tool results in the outgoing history with the compactor for their command, and replaces an exact repeat of an earlier result with a short pointer.
- Scrubs credentials from every relayed tool result, folded or not.
- Adds the style note only for agents that get it no other way.
- On an internal error, relays your original request.
- Without a valid license, or past your plan's quota, it only scrubs credentials.
Savings are booked once per folded block, the first time it is sent, and conservatively as re-reads saved on later requests. See how savings are measured.
Gotchas
- If the proxy is down, the agent's requests fail. Check
tokenade llm-proxy status, start the service, or runtokenade llm-proxy uninstall --agent <name>to restore the direct endpoint. - Keep it on loopback.
--remoteexposes a process that holds your provider key. - Exit code
2when there is nothing to install for that agent, or the configuration can't be planned.