Updated August 2026. Claude prices an image the same way it prices text, in tokens, but the rule behind that price is rarely spelled out end to end. This article does exactly that: the published formula, both resolution tiers, a worked example on a standard screen, and a second example measured against a real API bill, so you can price any image yourself before you paste it into a coding agent.
It is a companion to how many tokens a screenshot costs your agent, which covers why the cost matters in a coding workflow. This one stays on the arithmetic.
What is the image token formula Claude uses?#
Claude views images in patches rather than pixels. Each 28x28-pixel block is one visual token, so an image costs ceil(width / 28) × ceil(height / 28) visual tokens, capped at 4,784 on the high-resolution tier used by Claude 4.7 and later. As a single expression: tokens = min(4784, ceil(width / 28) × ceil(height / 28)).
Anthropic's vision documentation states the patch rule directly and publishes the tier limits it caps against. Worked on a screen everyone has:
ceil(1920 / 28) = 69 patches wide
ceil(1080 / 28) = 39 patches tall
69 × 39 = 2,691 visual tokens
2,691 is under the 4,784-token cap, so no downscaling applies and the figure is exact rather than approximate (calculated). Anthropic's own size table lists the same 2,691 for a 1920x1080 image on the high-resolution tier, which is what makes this number independently checkable rather than a vendor claim.
Why did a measured screenshot match the formula exactly?#
Because on the high-resolution tier the patch formula is not an estimate — it is the price. A real 1288x811 test screenshot captured on an M-series Mac was billed 1,334 tokens (measured), and the formula predicts 1,334 to the token.
ceil(1288 / 28) = 46 patches wide
ceil(811 / 28) = 29 patches tall
46 × 29 = 1,334 visual tokens
That is the whole verification. Two independently sized images, one calculated from the published rule and one measured off a real bill, agreeing exactly. Rounding matters and rounding goes up: 811 ÷ 28 is 28.96, which becomes 29 whole patches, because a partial patch is still a patch. That is why a 1288x811 image costs slightly more than its raw area would suggest, not slightly less.
The standard tier checks out the same way. Anthropic downscales a 1920x1080 image to 1456x819 for standard-tier models, and ceil(1456/28) × ceil(819/28) = 52 × 30 = 1,560 — the exact figure the documentation publishes for that row.
Does image resolution change how many tokens an image costs?#
Yes, in direct proportion to patch count, until the tier's cap intervenes. Because the formula multiplies width patches by height patches, resolution is the only input that matters: a sharper capture of the same window costs more even though nothing about the content changed.
This is easy to miss because resolution feels like a display setting rather than a cost lever. Two screenshots of the same window at different Retina scaling carry meaningfully different bills. What does not happen is unbounded scaling — see the cap below, which is why a 4K capture costs less than naive area maths predicts.
The other consequence is that content is free and geometry is not. A screenshot of a blank settings page and a screenshot of a dense dashboard cost identically at the same dimensions, because the model is charged for patches, not for what is in them.
Is there a cap on how many tokens one image can cost?#
Yes, and both tiers are published. The high-resolution tier used by Claude 4.7 and later accepts a long edge up to 2576 px and caps at 4,784 visual tokens. Standard-tier models accept a 1568 px long edge and cap at 1,568 visual tokens. Images past either limit are downscaled before they are priced.
| Resolution tier | Models | Max long edge | Max visual tokens | A 1920x1080 screenshot |
|---|---|---|---|---|
| High-resolution | Claude 4.7 and later | 2576 px | 4,784 | 2,691 tokens, not resized |
| Standard | All other models | 1568 px | 1,568 | 1,560 tokens, resized to 1456x819 |
Both rows come from Anthropic's vision documentation, and the exact resize behaviour — scaling to the largest size that fits the tier while preserving aspect ratio — is spelled out in Anthropic's image resizing and coordinates guide.
This is where an old rule of thumb goes wrong. A 3840x2160 Retina capture is four times the pixel area of 1080p, but it does not cost four times 2,691 tokens. On the high-resolution tier it is downsized to 2576x1449 and lands on the 4,784-token cap; on the standard tier it is downsized to 1456x819 and costs 1,560, exactly the same as the 1080p image. Capturing at 1x instead of full Retina saves about 44% on Claude 4.7 and later (2,691 against 4,784), not 75%.
What does the formula cost at common screen sizes?#
Every row below is the same three steps — divide by 28, round up, multiply — with the cap applied at the end. The two middle rows are load-bearing: one calculated from the published rule, one measured against a real bill.
| Image | Dimensions | Patch grid | Visual tokens | Type |
|---|---|---|---|---|
| Cropped detail view | 800x600 | 29 x 22 | 638 | calculated |
| 1-megapixel square | 1000x1000 | 36 x 36 | 1,296 | calculated, matches Anthropic's table |
| Real test screenshot | 1288x811 | 46 x 29 | 1,334 | measured, M-series Mac |
| 1080p full screen | 1920x1080 | 69 x 39 | 2,691 | calculated, exact |
| Phone screenshot | 1170x2532 | 42 x 91 | 3,822 | calculated |
| 4K capture, high-resolution tier | 3840x2160 | downsized to 2576x1449 | 4,784 | capped |
A 1080p screenshot at 2,691 tokens costs roughly $0.027 at $10 per million input tokens (2,691 ÷ 1,000,000 × $10). Anthropic's pricing page publishes input rates from $1 per million tokens on Claude Haiku 4.5 up to $10 per million on Claude Fable 5, so quote the rate rather than a model — the rate is checkable arithmetic and the model attribution goes stale. You can run any dimensions through the Walkie token calculator instead of doing it by hand.
How do you estimate an image's token cost before sending it?#
Three steps, no tools:
- Divide each dimension by 28 and round up. 1920 becomes 69 patches, 1080 becomes 39.
- Multiply the two patch counts. 69 × 39 = 2,691 visual tokens.
- Cap the result at 4,784 for Claude 4.7 and later, or 1,568 on standard-tier models — and remember that anything past the tier's long-edge limit is downscaled first.
For a mental shortcut with no arithmetic: a full-screen 1080p capture costs about 2,700 tokens on the high-resolution tier, an 800x600 crop about 640, and nothing ever exceeds 4,784. When you want certainty rather than a shortcut, Anthropic's token-counting endpoint accepts the same payload as a real request, images included, returns the exact input-token count, and is free to call.
How can you shrink a screenshot's token cost?#
Cut patches, not bytes. Cost is a patch count, so the only moves that work are the ones that reduce width or height.
- Crop before pasting. An 800x600 crop costs 638 tokens against 2,691 for a full 1080p screen on Claude 4.7 and later, a 76% reduction for about ten seconds of effort.
- Capture at 1x, not full Retina. A 4K capture hits the 4,784-token cap; the same content at 1080p costs 2,691, about 44% less.
- Send text where text will do. A stack trace, an error string, or a line number is cheaper and more precise as plain text than as a screenshot of a terminal.
- Compress nothing for token reasons. JPEG quality changes file size, not patch count. A smaller file saves disk space and upload time, never tokens.
- Narrate instead of capturing. A typical spoken review transcribes to about 150 tokens (measured) — cheaper than nearly any cropped image, and it carries what you meant rather than only what was visible.
That last point is where the arithmetic gets brutal for video: the same per-frame formula, multiplied across every extracted frame. Why your coding agent cannot watch a screen recording walks through that multiplication, and why UI work burns more Claude Code tokens than backend work covers the other multiplier, which is turns rather than frames.
The bottom line#
The formula is simple — patches of 28x28 pixels, one token each, capped at 4,784 on Claude 4.7 and later — and its consequences are not. A 1080p screenshot costs 2,691 tokens whether it shows a single button or an entire dashboard, because the rule only counts geometry. Cropping, capturing at 1x, and narrating instead of capturing are the three real levers, and all three are free. The measured baseline behind the narration figures, including the full bundle comparison, is published in the visual-context token cost research.