Claude Code tokens go to four places: the conversation history re-sent on every turn, tool calls and their results, files and images pasted into context, and the model's own output. Images are the outsized cost in that list — a single 1080p screenshot costs 2,691 tokens on Claude 4.7 and later, where any one image is capped at 4,784 — and a raw video frame-dumped into context multiplies that by dozens.

Updated August 2026. The clearest way to see where the money goes is one worked example: the same 30-second visual bug review, captured three ways. Frame-dumped at 1 frame per second it costs about 80,730 tokens. Described out loud and paired with the frames the narration points at, the same review runs about 3,500 tokens — 23x less for the same 30 seconds. Every figure below is labelled measured or calculated, and the formula is shown, not just the result.

What actually consumes tokens in a Claude Code session?#

Four things: the full conversation history, which Claude Code re-sends on every turn; tool calls and their results, such as file contents and command output; anything you paste or attach, including images; and the model's own response, including extended thinking. History dominates long sessions because it never shrinks on its own.

Anthropic's context-window documentation is explicit about the scope: everything in the request counts — the system prompt, every message in the conversation including tool results, images and documents, and your tool definitions — plus the output the model generates for that turn. The model has no memory between calls, so the transcript is the memory, and the transcript rides along every time.

Claude Code's own cost documentation names the same mechanic when it explains why usage climbs in a long session: the full conversation goes with every request, and each tool use sends another request carrying that batch of results. It also lists the quieter drains: cache misses after a break, scheduled tasks firing while the session is idle, and agent teammates still running.

Tool definitions have a footprint of their own, separate from what a tool returns. Every connected MCP server and every skill loaded at session start sits in context before you have asked for anything. Output matters too: the model's reply, including any extended thinking, is billed as output at a higher per-token rate than input, so a long deliberative answer costs more than a short direct one before a single tool call happens.

For a line-item look at where one two-hour session's tokens landed, see what one Claude Code session actually costs.

How much does an image cost an agent to read?#

Claude prices images by patch, not by file size: one visual token per 28x28-pixel block, capped at 4,784 tokens on the high-resolution tier used by Claude 4.7 and later. A full 1920x1080 screenshot works out to 2,691 tokens — ceil(1920/28) × ceil(1080/28) = 69 × 39. A smaller 1288x811 screenshot measured 1,334 tokens.

Anthropic's vision documentation publishes both the patch rule and the tier limits: a 2576 px long edge and 4,784 visual tokens on high-resolution models, a 1568 px long edge and 1,568 visual tokens on standard-tier models, which downscale a 1080p image to 1456x819 and charge 1,560 for it. The formula is worked through step by step in how Claude calculates image tokens, or you can run your own dimensions through the Walkie token calculator.

What you send Tokens Type
One 1080p screenshot, 1920x1080, high-resolution tier 2,691 calculated, exact
The same image on a standard-tier model 1,560 calculated
One real test screenshot, 1288x811 1,334 measured
Four screenshots to explain one page, no narration 10,764 calculated
A 15-second spoken description, transcribed 65 measured
A typical spoken walkthrough, transcribed ~150 measured, range 100–200

Four screenshots strung together to explain one page — the empty state, the error state, the hover state, the mobile width — cost four times a single frame, and none of the four says which one matters or why. Words are the cheapest input in this table by a wide margin.

The cap has a practical edge: past the tier's limits, a bigger and sharper screenshot costs nothing extra, because the patch count maxes out at 4,784. Above that threshold resolution becomes a legibility question, not a cost question. The full breakdown across real screen sizes is in how many tokens a screenshot costs an AI agent.

Why does the same task cost more on the second try?#

Because nothing about a retry repeats cheaply. Claude Code re-sends the entire session history with every attempt, so the second try inherits the tokens spent on the first plus its own. Long gaps also trigger cache misses, and the whole context gets reprocessed at full price instead of the discounted cache-read rate.

Anthropic's prompt-caching documentation puts numbers on that gap: a cache read costs 0.1x the base input rate, a five-minute cache write costs 1.25x, and a one-hour write costs 2x. Claude Code's cost page adds the practical detail that the cache lifetime is an hour on a subscription and drops to five minutes once you are drawing on usage credits. Miss the window and you pay full input price for context you already sent.

This compounds hardest around visual work. A fix that does not land usually gets another screenshot pasted in to show what is still wrong, and that screenshot does not replace the previous one — it stacks on it, along with everything the session has accumulated since. A bug reported three times with a fresh screenshot each time is not three separate 2,691-token charges on Claude 4.7 and later; it is three charges plus the growing weight of the conversation each one attaches to. That per-turn multiplication is the subject of why UI work burns more Claude Code tokens than backend work.

The cheapest retry is the one that never happens. A bug report specific enough to act on the first time skips the second round of screenshots entirely.

What does a context window have to do with cost?#

A context window is the fixed token budget one request can hold — system prompt, files, tool calls, images and history all draw from the same pool. Every token in that budget is a token billed for that turn, so a window filled with frame-dumped video leaves little room for the fix itself.

On a 200,000-token window, one 30-second review frame-dumped at 1 fps consumes roughly 40% of the budget before the fix begins (80,730 ÷ 200,000, calculated). Anthropic now ships 1M-token context windows by default on Claude Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 and Sonnet 4.6, while other models including Sonnet 4.5 remain at 200,000 tokens. A larger window raises the ceiling; it does not lower the price per token once you are in it.

There is a second cost that is not measured in dollars. Anthropic's own documentation names it: as token count grows, accuracy and recall degrade, a phenomenon it calls context rot. A session bloated with old screenshots and abandoned attempts is not only more expensive to keep running, it is more likely to produce a worse answer, because the signal you need is diluted across everything else still in the window. For what a context window is and what fills it, see what is a context window.

Which parts of your workflow are the expensive ones?#

Visual review is the most expensive category. Showing the agent your screen costs far more per turn than describing a bug in text, and raw video is the worst format for it: a 30-second recording frame-dumped at 1 fps costs about 80,730 tokens, while the same review as a narrated bundle costs roughly 3,500 — about 23x less, a 96% reduction.

Format Tokens, 30-second review Cost at $10 per million input tokens
Frame-dumped video, 1 fps 80,730 $0.81
Four screenshots, no narration 10,764 $0.11
Narrated bundle ~3,500 $0.04

Quote the rate, never the model: input prices differ per model and move over time, so "at $10 per million input tokens" stays checkable while "on model X" goes stale. The measurement method behind the bundle figure is published in the visual-context token cost research.

Retries and multi-step debugging are the second-largest category, for the reason above — every retry drags the full history behind it. Reading and editing code is comparatively cheap per turn. It is the visual proof that something works, not the code itself, that blows up a session's token count.

  • Re-sent conversation history, not new typing, is why long sessions cost more than they feel like they should.
  • Images are billed by patch and capped at 4,784 tokens each, regardless of file size.
  • Frame-dumped video is the single most expensive way to show an agent your screen: about 23x more than a narrated bundle covering the same 30 seconds.

How do you cut cost without cutting quality?#

Three habits carry most of the savings: clear context between unrelated tasks instead of letting history grow unbounded, match the model to the task instead of defaulting to the most expensive one, and stop sending raw video or screenshot dumps when a smaller, targeted visual reference says the same thing.

  1. Clear context between unrelated tasks. History never shrinks on its own; /clear is the biggest lever you control directly and it costs nothing.
  2. Match the model to the job. Reserve the most capable model for genuinely hard reasoning, not routine edits.
  3. Send one representative screenshot, not four. One well-chosen 1080p screenshot costs 2,691 tokens on Claude 4.7 and later; four covering the same page cost 10,764, for less clarity rather than more.
  4. Never frame-dump a screen recording. At 1 fps a 30-second clip runs about 80,730 tokens, and Claude Code cannot watch the video anyway — those tokens buy still frames, not understanding.
  5. Say what is wrong instead of only showing it. A spoken or written bug description runs 65–150 tokens, measured: the cheapest input on this list and the one thing a screenshot cannot do alone.
  6. If you are recording anyway, bundle it. A narrated screen-and-voice bundle — the format Walkie produces — measures about 3,500 tokens typical, close to the cost of a couple of screenshots and far closer in meaning to sitting next to someone.

The full set of eleven tactics, ranked by tokens saved per minute of effort, is in 11 ways to cut Claude Code token costs.

What should you measure?#

Run /usage inside Claude Code to see the current session's token breakdown by model, including cache reads and writes, or configure the status line to show context usage continuously. On a Pro, Max, Team or Enterprise plan, /usage also attributes recent consumption to skills, subagents, plugins and individual MCP servers, and flags behaviours such as long context or cache misses when one accounts for 10% or more of recent usage.

Two caveats worth knowing, both from Claude Code's cost documentation. The dollar figure in the Session block is computed locally from token counts at standard list rates, so it does not reflect promotional pricing or contracted discounts and can differ from an actual bill; for authoritative billing, the Claude Console usage page is the source. And the figures are computed from local session history on that machine, so usage from other devices is not included.

Watch /usage for a week before changing anything. Which turns spike — a pasted screenshot, a long retry, a session left open since morning — usually names one or two habits doing most of the damage. The point is not to spend nothing; it is to stop paying for tokens that never made the fix land any faster.