So you stop hitting "limit reached" in the middle of your work: what eats your limits, how caching works, and ten habits to last longer on the same subscription, whether you code or not.
The unit: the token
The chunk of a word that Claude reads and writes. A common word often fits in one token, a long or rare word takes several. A page of text is roughly a few hundred tokens.
Everything is counted in tokens: what you send (input), what Claude writes (output) and what it rereads. Because with every message, Claude rereads the whole conversation from the start, plus your instructions, your files and the description of its tools. On my setup, an empty Claude Code session, before I've typed anything, already weighs 63,514 tokens.
Two details that matter. Languages other than English usually take more tokens than English to say the same thing (details in this TokenClimate course). And since Opus 4.7, a new way of splitting text produces about 30% more tokens for the same text.
Two ways to pay
The subscription (Pro, Max, Team, Enterprise). A fixed monthly price, with two limits: a session limit that resets every 5 hours, and a weekly limit across all models, reset at a fixed time specific to your account.
These limits are shared between the Claude app, Claude Code, Claude Desktop and Claude Design. An afternoon in Design also eats into your Claude Code quota. Anthropic doesn't publish any quota figures, for any plan.
I ran the numbers on 1,255 of my sessions: $27,964 at API prices, for less than $2,000 of subscription. For heavy use, the subscription costs far less than the API, as long as you stay within the limits.
Pay as you go. Two options: usage credits when you go past your subscription, and the API, which is paid entirely per token (see chapter 10).
Pay-as-you-go billing that takes over when you hit your subscription limit. You keep going at API rates, with a monthly cap you set and optional auto-reload. On Pro and Max, you turn them on yourself at claude.ai/settings/usage; on Team and Enterprise, the admin does. Nothing to do with the API credits included in Max and Team, which are only for the API and the Agent SDK.
Plans and prices
| Plan | Price | Good to know |
|---|---|---|
| Free | $0 | Sonnet and Haiku. Docs, Slides and Design included since October 8, 2026. |
| Pro | $20/month ($17/month billed annually) | Opus, Sonnet, Haiku. Claude Code and Claude Design included. |
| Max 5x | $100/month | All models. $100 of API credits per month. |
| Max 20x | $200/month | All models. $200 of API credits per month. |
| Team | Standard seat $20 (annual) or $25 (monthly), Premium $100 or $125 | 2 to 150 seats. Pooled API credits, capped at $500. |
| Enterprise | $20 per seat per month self-serve, plus usage at API rates | No API credits included. |
Access to Fable 5.1, the most expensive model, depends on the plan and has changed several times since June 2026. Check claude.com/pricing before you count on it.
What uses the most
From the biggest to the most discreet:
- The model. At API prices, Fable 5.1 costs 5 times Sonnet 5.5 and 100 times Haiku 5.5 for input. On a subscription, a bigger model burns through your limits faster.
- Effort. That's how much Claude thinks before answering (from low to max). This reasoning is billed as output, the most expensive part: 5 times the input price.
- Conversation length. At message 50, Claude rereads the 49 before it. A conversation that drags on costs more with every exchange.
- Long tasks and agents. Since September 16, 2026, Claude decides on its own how much work a request deserves. A long task means many turns, file reads and tool calls.
- Claude Design. I burned 51% of my Max plan in 2 hours redoing my website with it.
- Connectors and tools. Each connected tool adds its description to the context. And so does what it brings back (a thread of 40 emails, a whole Notion page).
- Research. A Deep Research goes through many web pages. On the API, web search costs $10 per 1,000 searches, on top of tokens.
Caching, explained simply
When you send a message, Claude rereads everything before it. If that opening hasn't changed since your last message, Anthropic keeps it aside: that's the cache. Rereading it from the cache costs 10 to 40 times less than reading it fresh (10 times on most models, 20 times on Opus 5.5 and Sonnet 5.5, 40 times on Fable 5.1).
In my empty Claude Code session, 98.6% of the input is read from cache, so at the reduced rate.
But it expires.
| Where | Cache duration |
|---|---|
| API | 5 minutes by default, 1 hour as an option |
| Claude Code on a subscription | 1 hour for the main conversation, 5 minutes for subagents |
| Claude Code on usage credits, an API key or a cloud session | 5 minutes |
| Claude app | Anthropic doesn't publish the duration |
Every read resets the timer. As long as you keep sending messages, the cache stays warm. Two situations break it:
- Coming back after a break. You return to a long conversation after lunch: the cache has expired, and everything is reread at full price. A 300,000-token conversation restarted cold on Opus 5.5 costs about $2.40 at API prices with a 1-hour cache ($1.50 with a 5-minute cache). Warm, the same reread costs around $0.06. Over 30 days, I counted 147 cold restarts.
- Switching models midway. Each model has its own cache. Going from Opus to Sonnet in the middle of a conversation means everything gets reread from scratch. In my measurements: 11 model switches in 30 days.
Ten habits, from the most to the least worthwhile
- Start a new conversation for each topic. It's Anthropic's official recommendation, and the strongest lever: you start again from a light context.
- Pick the model for the task. Haiku or Sonnet to rephrase, sort or summarize. Opus for long reasoning. Fable when nothing else works.
- Lower the effort when the task is simple. An email to rephrase doesn't need max. Only go up if the answer disappoints.
- Keep the same model from start to finish in a conversation. If you need to switch, open a new one.
- After a long break, restart from a summary. Ask Claude to summarize the conversation, then paste the summary into a new one.
- Shorten your instructions. Your personal instructions and your project instructions are reread with every message. Ten useful lines beat three pages.
- Disconnect what you don't use. Connectors, web search, tools: turn off whatever the conversation doesn't need.
- Paste text instead of a screenshot. A copied error message costs less than an image.
- Ask for short answers, and group your questions. Output costs 5 times input. Three questions in one message means one reread of the conversation instead of three.
- For volume, use Haiku and batch. On the API, Haiku 5.5 costs $0.10 per million input tokens, and batch cuts the bill in half again.
Question
You're at message 80 of a conversation on Opus. To save money, you switch to Sonnet without changing conversations. What happens on the next message?
Pick an answer to see the explanation.
Tracking your usage
- In the Claude app: Settings > Usage (claude.ai/settings/usage), to see where you stand on your limits and turn on usage credits.
- In Claude Code:
/usage(aliases/costand/stats). You get an estimate in dollars, a breakdown by skill, subagent and MCP server, and a line just for the cache with the read rate and the likely cause of a cache miss.
/usage
/context
/context shows what fills your context window, and suggests ways to slim it down.
When you hit the limit
Three options. First, wait for the end of the 5-hour window. Since September 22, 2026, subscribers also get a "limit reset": a reset they can keep in reserve and use when they get stuck. The last option is to turn on usage credits, with a monthly cap (in Claude Code: /usage-credits). Careful: in Claude Code, switching to usage credits drops the main conversation's cache to 5 minutes.
The API credits included with Max and Team since October 7, 2026 don't help here: they cover the API, the Agent SDK and Managed Agents, not the app or Claude Code (details in chapter 10).
Measuring for real
The app's gauge shows where you stand on your limits. For the actual cost in money, the model that weighs most in your team or the carbon footprint, you need something else.
I had no way to see what a team's AI cost in euros and in CO2, so I built TokenClimate, which does the math from tokens. Its calculator gives you a first estimate, and the free course helps you choose between plans and the API.
For devs, the costs chapter of the Claude Code guide goes much further: cache settings, /compact, subagents, MCP.
Question
Weekly limit reached in the middle of your work, on Max 5x, with $100 of included API credits this month. What can take over?
Pick an answer to see the explanation.
Key takeaways
- The habit that pays off most: one conversation per topic.
- The app, Claude Code and Design share your 5-hour and weekly limits.
- Caching makes rereading 10 to 40 times cheaper. After a break or a model switch, everything is reread at full price.
- For a simple task, Haiku and low effort are enough.
- The API credits in Max and Team cover neither the app nor Claude Code.