Claude plans meter usage by tokens processed, not by the number of messages you send or hours you are logged in. A quick question costs almost nothing. A long coding session where you paste several screenshots per review costs vastly more, because those images and the growing conversation around them are re-sent, in full, on every single turn.

Updated August 2026. None of this is a glitch, and support cannot fix it for you. It is how large language models process a conversation, and it applies to every plan and every provider. Once the mechanic is visible the fixes are mostly free and take minutes. This is part of the wider picture in Claude Code token costs.

  • 2,691 tokens — one 1080p screenshot on Claude 4.7 and later
  • 10,764 tokens — four screenshots covering one page
  • 80,730 tokens — 30 seconds of screen recording frame-dumped at 1 fps

What actually counts against your limit?#

Every token processed: the text and code in your message, any images or files attached, Claude's reply, and in Claude Code the output of every tool call and file read. Message count is not the meter. A one-line question and a message carrying a screenshot are billed very differently even though the interface shows both as one message.

Anthropic does not publish a message or token quota for Pro or Max. Claude's pricing page says only that usage limits apply and describes Max as 5x or 20x more usage than Pro; it gives no per-window number. This article does not state a quota, a reset hour, or a dollar cap for those plans, because none of those figures are published in a form worth standing behind.

What Anthropic does document is the shape of the limits. Claude Code's error reference names three: a session limit, a weekly limit, and a separate Opus limit, each shown with its own reset time. The session and weekly limits are shared across all models, so switching models with /model does not restore access — the Opus limit is the one that does, because it applies only to Opus requests. On Teams and Enterprise plans specifically, Claude Code's cost documentation states that each member's allowance resets on a rolling five-hour window and a weekly window.

Why does context re-send on every turn?#

Because the API is stateless. To keep a thread coherent, every message carries the entire conversation so far back to the model: your first message, every reply, every screenshot, every file. A screenshot pasted at turn two is still part of the payload at turn twenty unless something removes it.

That is why a long visual-review session compounds. If turn two costs 2,691 tokens for one screenshot on Claude 4.7 and later and nothing is trimmed, the same 2,691 tokens is reprocessed at turn three, turn four, and every turn after, stacked on whatever you are adding. The per-turn version of this arithmetic is worked through in why UI work burns more Claude Code tokens than backend work.

Prompt caching changes the price of that repetition, not its existence. Claude Code's cost documentation is explicit: with prompt caching it re-reads that history at the cached token rate, so a one-line question in a session that has been open all day still draws usage for the whole conversation. Anthropic's prompt-caching documentation puts the cache read at 0.1x the base input rate, and the cache lifetime is an hour on a subscription, dropping to five minutes once you are drawing on usage credits — so the first message after a long break reprocesses your full context at full price.

What does a long session do to consumption?#

A long session compounds two ways at once: each turn adds fresh tokens, and each turn re-sends everything before it. Twenty separate 30-second visual reviews crammed into one sprawling thread can push a single conversation's payload into the hundreds of thousands of tokens.

Same 30-second visual review, sent as Tokens Type
One screenshot 2,691, Claude 4.7 and later calculated
Four screenshots covering one page, no narration 10,764 calculated
Full recording, frame-dumped at 1 fps 80,730 calculated
Voice and pointer bundle ~3,500 measured, typical

Take twenty reviews in a day as a working assumption, not a measurement. Frame-dumped, that is 20 × 80,730 = 1,614,600 tokens (calculated), mostly pixels the model is re-deriving from a frame it saw a second earlier. As voice-and-pointer bundles the same twenty reviews come to 20 × 3,500 = 70,000 tokens: a 23x reduction (80,730 ÷ 3,500). Priced like API access at $10 per million input tokens, that is $0.81 for one frame-dumped review against $0.04 for one bundled review, both calculated. Max is a flat subscription rather than metered per token, but the ratio holds however the bill is structured. The measurement method behind the bundle figure is published in the visual-context token cost research.

Even without full video, one 30-second frame dump alone consumes roughly 40% of a 200,000-token context window before you have made a single fix — and that is the same context then re-sent, in full, on every turn after it.

What else drains a plan while you are not typing?#

Several documented behaviours, none of them obvious from the interface. Claude Code's cost documentation lists the ones that make usage climb in a session that has been open for hours, and most have nothing to do with how much you typed.

  1. Long context. The full conversation goes with every request, and each tool use sends another request carrying that batch of results.
  2. Cache misses. The first message after a break longer than the cache lifetime reprocesses your full context at full price.
  3. Scheduled tasks. A scheduled task fires on its interval even while the session is idle, sending your full context each time.
  4. Cross-session messages. A message delivered from another of your sessions arrives as a new turn, sending your full context again.
  5. Agent teammates. Each active teammate keeps consuming tokens until it exits, and agent teams use roughly 7x the tokens of a standard session when teammates run in plan mode.
  6. Compaction itself. /compact reads the conversation it summarises, so compacting a large context is itself a large request. /clear costs nothing.

On a Pro, Max, Team or Enterprise plan, the /usage breakdown flags any of these that account for 10% or more of your recent usage, which is the fastest way to find out which one is yours.

Are images the hidden cost?#

Yes. Anthropic prices an image at one visual token per 28x28-pixel patch, capped per image, so a full 1080p screenshot lands at exactly 2,691 tokens on Claude 4.7 and later. That number is invisible in the interface, which shows "image attached", not a cost. Four screenshots documenting one page run 10,764 tokens before a word of feedback is written.

The formula is published in Anthropic's vision documentation: tokens = min(4784, ceil(width / 28) × ceil(height / 28)) on the high-resolution tier used by Claude 4.7 and later. For 1920x1080 that is 69 × 39 = 2,691 exactly, comfortably under the cap. Standard-tier models downscale to 1456x819 first and charge 1,560. A smaller real test screenshot at 1288x811 measured 1,334 tokens: still real cost, still invisible. See how many tokens a screenshot actually costs for the breakdown across common resolutions.

Most people do not paste one screenshot per review. They paste several, because one angle rarely tells the whole story. That is the honest problem: nobody pasting four screenshots is doing anything wrong, and nobody is shown the running total either.

What changes if you compress feedback?#

Pixels are expensive and low in meaning per token; words are cheap and can be precise. A spoken description runs roughly 100–200 tokens (measured) and can name the location, the expectation and the gap, while a screenshot at more than ten times the cost still cannot say what you care about.

That is the lever, and it costs nothing: describe more, screenshot less. When a picture is genuinely necessary — a layout bug, a colour mismatch — crop it to the relevant region, since a smaller image is fewer patches under the same formula. A precise sentence such as "the save button sits 8px right of the card edge and should align with the title above it" often replaces a screenshot entirely.

Some bugs really are easier to point at than to describe, which is why tools exist that turn a screen-and-voice walkthrough into a compressed, text-first bundle before it reaches Claude. Walkie is one option built for exactly that. It is a purchase decision, not a requirement, and everything above works with nothing but habits you already have.

What can you control today?#

Five changes cost nothing and take effect immediately. None require buying a tool or changing how you code — only how you hand Claude your feedback.

  1. Crop before you paste. A smaller image is fewer patches: cut the screenshot to the region that matters instead of the full window.
  2. Describe first, screenshot second. Name the location and the expected-against-actual behaviour in a sentence before reaching for the capture tool.
  3. Start new threads for new tasks. A fresh conversation does not carry forward every prior screenshot and file read. An old one does, on every turn.
  4. Do not re-paste what Claude already has. Referencing a file or screenshot by name costs far less than pasting it again.
  5. Watch /usage for a week before changing anything else. The attribution breakdown will usually name one or two habits doing most of the damage.

None of this changes what you are building. It changes how much of your limit gets spent showing Claude what you mean instead of building it. For the fuller ranked list, see 11 ways to cut Claude Code token costs, and for the single-session view of the same arithmetic, what one Claude Code session actually costs.