Updated August 2026. Every screenshot you paste into Claude Code, Cursor, or any other AI coding agent is priced before the agent reads a word of your bug report. This article prices it exactly, using Anthropic's published patch formula plus figures measured on a real test recording, its transcript, and the review bundle built from it.

The short version: a 1920x1080 screenshot costs 2,691 tokens on Claude 4.7 and later, where a single image is capped at 4,784 visual tokens (calculated). A real 1288x811 test screenshot measured 1,334 tokens. The same review delivered as narration plus the frames it points at measured about 3,500 tokens in total. Every figure below is labelled measured or calculated, because in a topic this checkable an unlabelled number is worth nothing.

Five ways of showing an agent the same broken screen, side by side:

What you send Token cost Type
One 1080p screenshot, 1920x1080, high-resolution tier 2,691 tokens calculated
The same image on a standard-tier model, downscaled to 1456x819 1,560 tokens calculated
A real measured screenshot, 1288x811 1,334 tokens measured
A spoken review transcript, typical ~150 tokens, range 100–200 measured
A complete narrated review bundle: transcript, pointed frames, summary ~3,500 tokens measured
A 30-second screen recording frame-dumped at 1 fps 80,730 tokens calculated

Two things jump out. A spoken transcript alone is roughly 18x cheaper than one 1080p screenshot on Claude 4.7 and later (2,691 ÷ 150 = 17.9) while still carrying an explanation a screenshot cannot. And the naive move — recording the screen and handing over every frame — is by far the most expensive option on the table, worse than pasting a stack of screenshots. The rest of this article works through why, one question at a time. You can run any of these numbers yourself in the Walkie token calculator.

How many tokens does a 1080p screenshot cost an AI coding agent?#

A single 1920x1080 screenshot costs 2,691 tokens on Claude 4.7 and later, the high-resolution tier, where any one image is capped at 4,784 visual tokens (calculated). That is the price of the agent looking at the pixels once, before any explanation of what is wrong in the image and before it does anything about it.

The number is not an estimate. Anthropic's vision documentation publishes a table of image sizes against token counts, and the 1920x1080 row reads 2,691 tokens on the high-resolution tier, or 1,560 on the standard tier after the image is downscaled to 1456x819. A smaller real-world capture costs less: a 1288x811 test screenshot on an M-series Mac measured 1,334 tokens. Image cost tracks pixel geometry, not file size or format.

The practical consequence is that resolution is a cost lever most developers never pull. A screenshot cropped to the broken component rather than the whole display costs proportionally less — but even a tight crop still costs hundreds of tokens before a single word of explanation is attached to it. For context, 2,691 tokens — the high-resolution-tier price of one 1080p capture, against a 4,784-token per-image cap — is roughly the weight of a few hundred lines of source. Pasting three or four screenshots into one message can cost more than the bug report, the relevant function, and the fix combined, and none of those tokens contain a line of code the agent can act on.

How does Anthropic actually calculate image tokens?#

Claude views images in patches, not pixels. Each 28x28-pixel block is one visual token, so an image costs ceil(width / 28) × ceil(height / 28) visual tokens, capped at 4,784 on the high-resolution tier used by Claude 4.7 and later. Written as one expression: tokens = min(4784, ceil(width / 28) × ceil(height / 28)).

Anthropic's vision documentation states the rule directly and publishes both tiers: high-resolution models accept a long edge up to 2576 px and a maximum of 4,784 visual tokens, while standard-tier models downscale to a 1568 px long edge and a 1,568-token maximum. Anything larger than a tier's limits is downscaled before it is priced.

Working it by hand takes seconds. For 1920x1080: ceil(1920/28) = 69, ceil(1080/28) = 39, and 69 × 39 = 2,691 tokens, under the 4,784 cap, so it is exact rather than approximate. For the measured 1288x811 test file: ceil(1288/28) = 46, ceil(811/28) = 29, and 46 × 29 = 1,334 — the same number the API actually billed. Formula and measurement agreeing to the token is the reason this article treats the formula as usable rather than theoretical.

One consequence catches people out: a screenshot of a blank settings page costs exactly the same as a screenshot of a dense code editor at the same resolution. The model is not charging for what is interesting in the image. It is charging for the image's geometry. The full worked derivation, including phone and cropped sizes, is in how Claude calculates image tokens.

Does a higher-resolution screenshot cost more?#

Yes, until it hits the cap. Cost scales with pixel area, so a 1920x1080 capture at 2,691 tokens costs roughly double a 1288x811 capture at 1,334 — almost exactly the ratio of their pixel counts on Claude 4.7 and later, where the ceiling is 4,784 visual tokens per image.

Above the tier's limits, the scaling stops. Anthropic's published table prices a 3840x2160 capture at 4,784 tokens on the high-resolution tier, because the image is first downsized to 2576x1449 and then hits the cap — not at four times the 1080p figure, which is what naive area scaling would predict. The same 4K image on a standard-tier model is downsized to 1456x819 and costs 1,560 tokens (Anthropic vision documentation).

So capturing a Retina display at full resolution is the most expensive way to hand an agent a bug report, but it is expensive by a factor of about 1.8, not 4. The bigger lever is the crop. Cutting a full 1080p capture down to an 800x600 region takes cost from 2,691 tokens, the high-resolution-tier price under the 4,784 cap, to ceil(800/28) × ceil(600/28) = 29 × 22 = 638 tokens, a 76% reduction, and it is free.

How much does a screenshot cost in dollars?#

At $10 per million input tokens, one 1080p screenshot on Claude 4.7 and later costs about $0.027 (2,691 ÷ 1,000,000 × $10). Volume is where it bites: a 30-second review frame-dumped at 1 frame per second runs 80,730 tokens, about $0.81 at that rate, against roughly $0.04 for a complete narrated bundle covering the same 30 seconds.

Four full-page screenshots covering one page, with no narration explaining what is wrong in any of them, run 10,764 tokens — about $0.11 at $10 per million input tokens, before the agent has been told what to fix. Quote the rate, not a model: Anthropic's pricing page lists input rates from $1 per million tokens on Claude Haiku 4.5 to $10 per million on Claude Fable 5, and those rates change while the arithmetic does not.

For most individual developers the dollar figure is a translation device rather than the real cost. On a Pro or Max subscription the currency is the plan's rolling allowance, and Claude Code's own cost documentation notes that the dollar figure it computes locally is priced at standard list rates and is not what a subscriber is billed. An $0.81 frame dump, at $10 per million input tokens, never appears on an invoice for a Max user. It appears as the agent running out of room to think three exchanges earlier than it otherwise would have.

Is a screen recording cheaper than a screenshot for an agent to read?#

No, because coding agents cannot watch video at all. A raw recording is cheap to store — 6.7 MB for a 15-second capture, measured — but it has to be converted into individual frames before any agent can read it, and a frame-dumped recording costs far more than a handful of well-chosen screenshots.

Recording feels like the more complete way to show a bug, because it captures everything rather than one static moment. "Everything" is the problem. A 30-second recording converted at a modest 1 frame per second produces 30 separate images, each priced like a full screenshot: 80,730 tokens (30 × 2,691 on Claude 4.7 and later, under the 4,784-token cap). That is 23x more than the ~3,500-token narrated bundle covering the same 30 seconds, a 96% reduction.

Turning the frame rate down does not rescue it. A fixed-interval sampler oversamples a static settings screen and undersamples a fast scroll, so a lower rate buys a cheaper bill and a higher chance of missing the second that mattered. There is no frame rate that makes blind sampling cheap; the fix is selecting frames by where a person actually pointed or spoke. The full mechanics are in why your coding agent cannot watch a screen recording.

How many screenshots can I paste before I fill the context window?#

Four full-page screenshots already cost 10,764 tokens with no narration attached. A 30-second recording frame-dumped at 1 fps costs 80,730 tokens — close to 40% of a 200,000-token context window gone (calculated) before the agent has read a line of your explanation.

Window sizes have moved, and it does not change the per-image price. Anthropic's context-window documentation states that Claude Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, and Sonnet 4.6 carry a 1M-token context window by default, while other models including Sonnet 4.5 stay at 200,000 tokens. The same page is explicit that everything in the request counts toward the window — system prompt, every message including images and tool results, and your tool definitions. A larger window raises the ceiling. It does not make a screenshot cheaper, and it does not stop images being re-sent on later turns, which is the compounding effect covered in why UI work burns more Claude Code tokens than backend work.

Four habits keep the number down:

  1. Crop before you paste. Cost scales with pixel area, so a tighter crop is a direct saving: 800x600 is 638 tokens against 2,691 for full 1080p on Claude 4.7 and later.
  2. Send fewer, better screenshots. Four screenshots at 10,764 tokens with no explanation communicate less than one screenshot plus a sentence of context.
  3. Never frame-dump a recording. Check how many frames a tool will produce before you send it; 30 seconds becomes 30 full-price images.
  4. Prefer narration to repetition. A spoken description costs roughly 150 tokens and replaces several near-identical screenshots of the same broken state.

Is a spoken description cheaper than a screenshot?#

Yes, by roughly 18x. A typical spoken review transcript measures about 150 tokens (measured, range 100–200) against 2,691 tokens for one 1080p screenshot on Claude 4.7 and later, capped at 4,784. Unlike a screenshot, narration tells the agent why something is wrong, not only what it looked like at one instant.

A real test makes it concrete: a 15-second spoken walkthrough of a screen transcribed to just 65 tokens (measured), while carrying intent, urgency, and which part of the screen actually matters. That is the asymmetry most developers have never priced: screenshots are expensive and mute, a few spoken sentences are cheap and explicit.

Transcripts alone are not the whole answer. A spoken review still has to point at something, or "it is broken over here" carries no more information in text than it did out loud. The 65-token and ~150-token figures are narration alone. Pair the transcript with one or two frames tied to what was said and the total stays close to the transcript's cost while the information content jumps past screenshot level, because the frame is no longer doing all the explaining by itself.

What should you send instead of raw screenshots?#

Send a bundle that pairs short narration with only the frames it points to. Measured, that combination — transcript, pointed frames, and a written summary — runs about 3,500 tokens for a complete review, against 80,730 tokens for the frame-dumped equivalent of the same 30 seconds: 23x fewer, a 96% reduction.

Same 30-second review, sent as Tokens Cost at $10 per million input tokens Type
Frame dump, 1 fps, 1080p 80,730 $0.81 calculated
Four screenshots, no narration 10,764 $0.11 calculated
Narrated bundle: transcript, pointed frames, summary ~3,500 $0.04 measured
Transcript only, no frames ~150 under $0.01 measured

This is the mechanic behind visual feedback for AI coding agents: instead of showing an agent everything and letting it guess what matters, you narrate what is wrong while pointing at it, and only the frames tied to your voice get included. Several tools now convert screen recordings into agent-readable bundles this way, and they differ mainly in how frames get selected — see how Walkie compares to Clipy and Loom for the current field, including which are subscriptions and which are bought once. The measurement method behind the 3,500-token figure, and the full baseline comparison, is published in the visual-context token cost research.

How do you check a screenshot's cost before you paste it?#

Multiply the patch counts. Divide the width by 28 and round up, do the same for the height, multiply the two, and cap the result at 4,784 for Claude 4.7 and later. That is the whole calculation, and it takes about ten seconds:

  1. Read the pixel dimensions. Most screenshot tools show them in the file's properties or export dialog, for example 1920x1080 or 1288x811.
  2. Apply the patch formula. ceil(width / 28) × ceil(height / 28), then cap at 4,784 on the high-resolution tier, or 1,568 on standard-tier models.
  3. Multiply by the number of images you are about to send in one turn — an agent reads everything in the turn together.
  4. Convert to your budget. Tokens ÷ 1,000,000 × your input rate for dollars, or tokens ÷ window size for the share of context it consumes.
  5. Crop and recount if the number looks high, since cropping reduces both terms of the product.

If you would rather not do the arithmetic, Anthropic's token-counting endpoint accepts the same message payload as a real request, including images, and returns the exact input-token count before you send anything. It is free to call. Inside Claude Code, /context shows what is currently occupying the window and /usage shows the session's token breakdown by model.

The bottom line#

Screenshots are not useless — one well-cropped image is often exactly the right amount of information. The narrow case against them is habit: pasting four by reflex, or frame-dumping a recording, spends 10,764 or 80,730 tokens on pixels when a sentence of narration plus one pointed frame would have communicated more for about 3,500.

Walkie is one entry in that category: a Mac app that records your screen and voice while you point at what is wrong, then hands your coding agent a bundle built for reading rather than watching. It is not the only tool doing this, and it does not replace a well-cropped screenshot when that is genuinely all a bug needs. It is built for the moment a screenshot cannot carry enough meaning and a full recording carries far too much. For the session-level view of where the rest of the tokens go, see what one Claude Code session actually costs.