What Claude Code consumes, and how to cut it without losing quality: cache, model, effort, context, tracking usage. The other chapters point here for the details.
Why care, even on a subscription
A token is the unit the model reads and writes: roughly a chunk of a word. Everything is counted in tokens, input (what Claude reads) as well as output (what it writes).
On a subscription, you don't pay per token. But you have two limits: a session window that resets every 5 hours, and a weekly limit. Both are shared between claude.ai, Claude Code, the desktop app and Claude Design. Anthropic doesn't publish any numeric quota per plan. When you hit the limit, you either wait or switch to usage credits, billed at API rates, with a monthly cap you set yourself (on Team and Enterprise, an admin turns them on).
Since October 7, 2026, Max and Team include monthly API credits. They cover the API, Managed Agents and the Agent SDK, but not Claude Code.
The official ballpark figures, on the enterprise side: about $13 per developer per active day on average ($150 to $250 per month), and under $30 per active day for 90% of users.
My numbers, measured on October 1, 2026:
1,255 sessions = $27,964 at API prices, for less than $2,000 in subscription fees. And an empty session, before I even send my first message, weighs 63,514 tokens, 98.6% of them read from cache.
The subscription absorbs a lot, but the limits hit fast if you waste: two hours in Claude Design ate half of my Max plan (chapter 15).
How Claude Code consumes
The model has no memory between two messages. With every message, Claude Code sends the whole package again from the start. That package is the context: everything the model has in front of it to answer. It's stacked in this order:
- The system prompt and tool definitions. They rarely change.
- The project context: CLAUDE.md, auto memory, rules.
- The conversation: your messages, the replies, every file read, every command output.
Whatever enters the conversation stays there until the next /clear or /compact. A 3,000-line file read at the third message is sent again with every message after that.
On the output side, everything Claude writes is billed at the output rate, the most expensive one. Internal reasoning (thinking) counts as output too. On the 5.5 models and Fable, you can't turn it off. Your lever is effort: how much thinking you ask the model to do before answering (low, medium, high, xhigh, max). Less effort means less thinking, which means less output.
The cache, in detail
How it works
The cache is Anthropic's server-side memory of the beginning of your package. If the beginning is identical, character for character, to the previous request, it's reread at a discount instead of being processed again. The smallest change upstream sends everything after it back to full price.
An empty session is 98.6% read from cache because the system prompt and the tools don't change from one message to the next. For the same reason, plan mode and skills add their instructions at the bottom, in the conversation: the prefix stays put.
Prices
Price per million tokens, API, October 2026:
| Model | Input | 5-min cache write | 1-h cache write | Cache read | Output |
|---|---|---|---|---|---|
| Fable 5.1 | $10 | $12.50 | $20 | $0.25 | $50 |
| Opus 5.5 | $4 | $5 | $8 | $0.20 | $20 |
| Sonnet 5.5 | $2 | $2.50 | $4 | $0.10 | $10 |
| Haiku 5.5 (prompt ≤ 100K) | $0.10 | $0.125 | $0.20 | $0.01 | $0.50 |
As multipliers of the input price: a cache read costs 0.05x on Opus 5.5 and Sonnet 5.5, 0.025x on Fable 5.1, 0.1x on the others. A write costs 1.25x with a 5-minute cache, 2x with a 1-hour cache.
How long it lasts (the TTL)
How long the cache survives without activity. Every message that reads it restarts the countdown, for free. Past that delay, the next message pays for the whole package again. The durations are in the table below.
| How you pay | Main conversation | Subagents, workflows, forks, compaction |
|---|---|---|
| Subscription, within your quota | 1 hour | 5 minutes |
| Usage credits, API key, cloud | 5 minutes | 5 minutes |
What breaks the cache, what keeps it
The table repeats the figure, with a few extra cases.
| Breaks the cache (the next turn starts at full price) | Keeps the cache |
|---|---|
Switching models with /model (each model has its own cache) | Editing repo files |
opusplan, which switches models entering and leaving plan mode | Editing CLAUDE.md mid-session (the change applies at the next /clear, /compact or restart) |
A skill with a different model in its frontmatter | Changing permission mode or output style |
| Changing effort, except on Opus 5.5, Sonnet 5.5, Haiku 5.5 and Fable 5.1 | Changing effort on Opus 5.5, Sonnet 5.5, Haiku 5.5 and Fable 5.1 |
Turning on /fast (once per conversation) | Invoking a skill, entering plan mode |
| Connecting or removing an MCP server whose tools load upfront | Connecting an MCP server with tool search (the default behavior) |
/compact | /recap, /rewind, /cd, /fork |
| Lots of images in the conversation | Launching a subagent |
| Updating Claude Code (the first turn after the update starts from scratch) |
The effort exception applies with an API key or a subscription. On Bedrock or Google Cloud, changing effort still breaks the cache.
The Claude Code team explained on Anthropic's blog on April 30, 2026 ("Lessons from building Claude Code: prompt caching is everything") that they declare an incident when the cache hit rate drops. Switching models to "save money" often backfires. On a 100,000-token conversation, going from Opus 5.5 to Sonnet 5.5 forces you to rewrite the whole cache on Sonnet: $0.25 ($2.50/M), versus $0.02 to reread it on Opus ($0.20/M). The savings only show up after several turns. For a simple subtask, launch a subagent on Haiku instead.
The cold restart, in numbers
The first message sent after a pause longer than the TTL. The cache has expired: Claude Code rereads the whole conversation with no discount and writes it back into the cache, at the write price. The message itself can be three words long. What you pay for again is the conversation behind it.
Example: a 300,000-token conversation on Opus 5.5.
- Warm, rereading it costs 300,000 × $0.20 / M = $0.06.
- Cold, everything goes back through a cache write: 300,000 × $8 / M = $2.40 with the 1-hour cache, 300,000 × $5 / M = $1.50 with the 5-minute cache.
Between 25 and 40 times more expensive, for the same message. The 1-hour cache costs more to write, but it saves you a cold restart when you go to lunch for 40 minutes. On credits or an API key, with a 5-minute TTL, a coffee break is enough to cool everything down.
Over 30 days, I counted 147 cold restarts and 11 model switches mid-session. Each one made me pay for an entire conversation again.
Two ways around it:
- Resume from a summary. On Pro and Max, when you resume a big session after a long break, Claude Code offers to restart from a summary instead of the full history.
- Start clean. If the task is done,
/clearand a new session cost less than a cold restart on 300,000 tokens.
Check your TTL
claude -p "just reply ok" --output-format json
In the JSON, look at usage.cache_creation. 1-hour cache writes show up under ephemeral_1h_input_tokens, 5-minute ones under ephemeral_5m_input_tokens.
Set the TTL yourself
# 1 h for the main conversation (useful on an API key or usage credits)
export CLAUDE_CODE_PROMPT_CACHE_TTL=1h
# 1 h everywhere, conversation and subagents
export ENABLE_PROMPT_CACHING_1H=1
# 5 min everywhere (to debug or compare)
export FORCE_PROMPT_CACHING_5M=1
The settings.json equivalent: promptCacheTtl and subagentPromptCacheTtl, set to "5m" or "1h". You need Claude Code v2.1.242 or later. The 1-hour cache pays off if you take breaks longer than 5 minutes. On short bursts, you pay double for the write and get nothing back.
Choosing the model and effort
In Claude Code, the default model is Opus 5.5, on every plan, with medium effort. The four current models (Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 5.5) all have a 1-million-token context window.
| Task | Recommended model | Effort |
|---|---|---|
| Most coding: features, bugs, tests | Sonnet 5.5 | medium |
| Architecture, multi-step reasoning, large refactors | Opus 5.5 | medium, then high or xhigh |
| Simple subagents: search, reading logs, inventory | Haiku 5.5 | low or medium |
| A very hard problem where only quality matters | Fable 5.1 | Use sparingly: the priciest ($10 / $50) |
The model choice follows the Claude Code costs documentation. For effort, start from the default and only go up if the answer falls short. Personally I run high on Opus 5.5, knowing what it costs in output.
The commands:
/model sonnet
/effort medium
claude --model sonnet --effort medium
export CLAUDE_CODE_SUBAGENT_MODEL=haiku
CLAUDE_CODE_SUBAGENT_MODEL sets the model for subagents, teammates and workflow agents that don't have one assigned. For a specific subagent, put model: haiku in its frontmatter. The max level only applies to the current session; settings don't accept it.
Decide at the start. The model is part of the cache key. Switching models midway makes you pay for the whole conversation again. Effort, on the other hand, can be changed without breaking the cache on the 5.5 models and Fable 5.1.
Three settings switch models without you noticing:
/fastkeeps the same model and answers about 2.5 times faster, but at the fast rate: $8 / $40 per million on Opus 5.5, double the normal price. And turning it on recomputes the whole context once, without cache. If you want it, turn it on at the start of the session.opusplanuses Opus in plan mode and Sonnet for execution. Every switch is a model change.- A skill with
model:in its frontmatter switches models for the length of its turn.
Keeping the context light
/clear, /compact, /rewind: what each one costs
| Command | What it does | Cost |
|---|---|---|
/clear | Starts over from an empty conversation | Free. Also resets the /usage counter. Run /rename first so you can find the session again with /resume. |
/compact | Rereads the whole conversation and replaces it with a summary | Warm cache: a fraction of the price. Cold cache: the most expensive operation of your day. |
/rewind | Goes back to an earlier turn (code and conversation) | Cheaper than compacting: you land on a prefix that's already cached. |
/recap | Shows a summary of the session without touching the history | Keeps the cache. |
The rule: compact when you're continuing the same task and running out of room, clear when you change subjects. To steer the summary, pass instructions (/compact keep the decisions about the database schema) or add a "Compact instructions" section to your CLAUDE.md. In /rewind, the "Summarize up to here" option compacts the conversation up to the point you pick.
Auto-compact
On 1-million-token models, Claude Code compacts on its own at around 967,000 tokens by default. That's late, and every message at 800,000 tokens costs you. You can bring the threshold forward:
/autocompact 500k
/autocompact auto
The equivalent setting is autoCompactWindow in settings.json, between 100,000 and 1 million.
What bloats the context, and what to do about it
- CLAUDE.md under 200 lines. It's sent with every message. Specialized workflows go into skills, which only load when they're needed.
- Skills.
/skill-doctorshows what each skill costs in context and how often it's used. Cut the ones you don't use. - MCP servers. MCP tools are deferred by default (tool search): they announce themselves by name and weigh almost nothing at rest. The risk comes from their responses, capped at 25,000 tokens by default. When a CLI exists (
gh,aws), it's often leaner than an MCP server. Details in chapter 17. - Command output. A hook can filter before Claude reads. The example from the docs: a
PreToolUsehook that rewrites the test command to keep only the failures. A 10,000-line log turns into a few hundred tokens. - Screenshots. Paste the error text instead of an image. Lots of images also break the cache.
- Verbose output. Hand it to a subagent: running the test suite, reading logs, digging through docs. Its output stays in its own context, and only a summary comes back. Careful: a subagent doesn't read your cache and only gets a 5-minute TTL, even on a subscription. A fork, on the other hand, starts from the conversation's cache: the forked subagent from
/subtask, like the session copy from/fork. - Vague prompts. "Look at the project and improve performance" makes Claude read half the repo. Tell it where to look. And if Claude heads the wrong way, hit
Escright away: every useless turn costs you.
Third-party tools
| Tool | What it does | What's been measured |
|---|---|---|
| Caveman | Makes Claude answer in telegraphic style | The repo claims -65% output. Not independently measured. Output is only part of the bill. |
| RTK | Filters tool output before Claude reads it | It filters 74% of the text in tool output (self-reported figure). An A/B test by JetBrains shows no net gain. |
| ccusage | Reads Claude Code's local JSONL logs and prices your usage | Open source, handy (ccusage blocks --live). It's an estimate at API prices, not your subscription bill. |
Before installing a tool, write down your usage in /usage, then compare a week later.
Tracking your usage
| Command or tool | What it shows |
|---|---|
/usage | Estimated cost, plan limits, and on a subscription the breakdown by skill, subagent, plugin and MCP server. Aliases: /cost, /stats. |
/context | What's filling the window right now, as a grid, with suggestions. /context all for the details. |
/insights | An HTML report on how you work (friction points, suggestions). It consumes tokens itself. |
| The status line | Your numbers, always visible under the input field. |
| OpenTelemetry | For a team: tokens, costs and activity per person, exported to your own observability stack. |
In /usage, the Prompt cache (main) line gives your hit rate, tells you whether the cache is warm or cold, and names the likely cause of the last miss (for example, tool definitions that changed). Check it after every session that felt expensive.
A status line that shows the cost, the 5-hour limit and the cache state:
#!/bin/bash
input=$(cat)
COST=$(echo "$input" | jq -r '.cost.total_cost_usd // 0')
FIVE_H=$(echo "$input" | jq -r '.rate_limits.five_hour.used_percentage // "?"')
HIT=$(echo "$input" | jq -r '.prompt_cache.hit_ratio // 0')
WARM=$(echo "$input" | jq -r 'if .prompt_cache.warm then "warm" else "cold" end')
echo "\$$COST | 5h: ${FIVE_H}% | cache ${WARM} (hit ${HIT})"
rate_limits only shows up on Pro and Max, after the first response. cost.total_cost_usd is an estimate at list price, not your bill. How to wire up the status line is shown in my setup.
Since October 1, 2026, a mod (code that runs inside Claude Code) can also show you these numbers. Anything it does without calling a model, like a command it serves or a line under the prompt, starts no Claude turn and costs no tokens. Chapter 20 builds one in 39 lines: it shows how full the context is, the share of the last turn read from cache and the time left before the cache goes cold, so you see the cold restart coming. What does cost money is a mod that calls a model ($.model.complete) or submits a prompt: read the calls: line of claude plugin validate before you install one.
When you hit a limit, /usage-credits opens the usage credits settings. /rate-limit-options also offers to wait for the window to reset and resume the task automatically.
The action plan: 10 habits
Ranked by likely savings, from the biggest to the most marginal.
- Pick the model at the start and stick with it. Sonnet 5.5 for everyday code, Opus 5.5 when the work calls for reasoning.
- Avoid cold restarts. After a long break, resume from a summary or
/clear. /clearevery time you change subjects, after a/rename.- Lower the effort on simple tasks. On the 5.5 models, it doesn't break the cache.
- Hand verbose work to a subagent on Haiku 5.5: tests, logs, exploration.
- Tell Claude where to look instead of letting it dig through the repo.
- Keep your CLAUDE.md under 200 lines and move the rest into skills.
- Filter output with a hook (keep only the failing tests).
- Paste text, not screenshots.
- Check
/usageonce a week, especially the cache line and the breakdown by skill and by MCP server.
Measuring cost and CO2
/usage gives you your usage for the session and the week. For a monthly cost, or a carbon equivalent, you need another tool.
I wanted to see my sessions in euros and CO2, not just as a percentage of a limit. So I wrote claude-carbon, the plugin covered in chapter 16, then TokenClimate to run the same calculation on a team's AI usage. For a ballpark figure without installing anything, there's the calculator, and the course goes over the levers from this chapter from the billing side.
Question
Which of these moves makes the next turn start without cache?
Pick an answer to see the explanation.
Question
Max subscription, Opus 5.5, a 300,000-token conversation. You go from /effort high to /effort low. What happens to the cache?
Pick an answer to see the explanation.
Key takeaways
- With every message, Claude Code sends everything again. The cache makes that bearable, as long as the beginning of the package doesn't change.
- Switching models or resuming a cooled-down session makes you pay for the whole conversation again: up to 40 times the price of a warm message on 300,000 tokens.
- Sonnet 5.5 for most code, Opus 5.5 for reasoning, Haiku 5.5 for subagents. Effort can be lowered without breaking the cache on the 5.5 models.
/clearis free./compacton a cold cache is the most expensive operation.- Measure before you optimize:
/usage, thePrompt cache (main)line, then claude-carbon or TokenClimate for estimated cost and CO2.
Sources
- Track and reduce costs (Claude Code docs)
- How Claude Code uses prompt caching (Claude Code docs) and Lessons from building Claude Code: prompt caching is everything (blog, April 30, 2026)
- Pricing and Prompt caching (Claude Platform docs)
- Model configuration (Claude Code docs)
- Understanding usage and length limits (help center)
My tools and measurements: claude-carbon, TokenClimate, what Claude Code costs, the cost of a long context (free course), and the "Claude Code without the waste" sheet (two months, 230 measured sessions).