Ten of these eleven tactics cost nothing and take under a minute each: stop frame-dumping recordings, describe a bug in words before reaching for a screenshot, crop what you do capture, and clear context between unrelated tasks. The eleventh costs money — a narrated-review tool such as Walkie — and it only earns its place once visual review is a daily habit rather than a one-off.
Updated August 2026. Every figure below is either measured directly or calculated from Anthropic's published vision formula, and the tactics are ranked by tokens saved per minute of effort.
Why do costs run away?#
Because a Claude Code session re-sends its running context on every turn. Anything expensive added early — a raw screen recording, four screenshots instead of one — gets paid for again on every later message, not once when you pasted it. Anthropic's context-window documentation states that everything in a request counts, images and tool results included.
The clearest example is the ordinary habit of covering a page from a few angles instead of one:
- 10,764 tokens — four full-page screenshots, no narration
- 2,691 tokens — one cropped screenshot of the same bug, on Claude 4.7 and later
- 8,073 tokens saved — by sending one instead of four (10,764 − 2,691)
That ratio, tokens saved against minutes spent, is the sorting principle for the eleven tactics below. The arithmetic behind the per-screenshot number is worked step by step in how many tokens a screenshot costs an AI agent; this article is about what to do with it. Treat the figures as close estimates from a published formula and real measurements — the ratios between them are what should drive where you spend effort.
The eleven tactics#
Ranked by tokens saved per minute of effort: stop frame-dumping recordings, narrate or describe bugs before screenshotting, crop and batch what you capture, clear context between unrelated tasks, and reach for a paid bundler only once visual review is a daily habit. Ten of the eleven cost nothing.
Stop frame-dumping raw recordings into the agent. This is the single biggest sinkhole on the list. Saves: a 30-second review frame-dumped at one frame per second runs 80,730 tokens (30 × 2,691 on Claude 4.7 and later, under the 4,784-token per-image cap). Describing the same thing in words removes almost all of it.
Say it in words before reaching for a screenshot. If the bug is nameable — "the CTA button is red instead of teal" — text says it for a fraction of an image's cost. Saves: zero visual tokens instead of 2,691 for a full 1080p screenshot on Claude 4.7 and later.
Narrate instead of screenshotting. A spoken description of what is wrong, even dictated with no special tool, carries what and where in far fewer tokens than an image. Saves: a typical spoken review transcribes to roughly 100–200 tokens against 2,691 for one screenshot on Claude 4.7 and later — about 2,500 tokens per instance (2,691 − 150).
Send one screenshot, not four. Crop to the region that shows the problem instead of capturing the page from several angles. Saves: 8,073 tokens (10,764 − 2,691) for roughly 30 seconds of extra care.
Crop before you paste. Anthropic's vision documentation prices an image by pixel dimensions, tiled into 28x28-pixel patches at one visual token each, capped at 4,784 on the high-resolution tier. A tighter crop is a smaller grid of patches: an 800x600 crop is 29 × 22 = 638 tokens. Saves: scales directly with the dead space you remove.
Point at a file path instead of pasting the file. Claude Code can read a file itself; pasting it inline duplicates what it can already fetch, and the duplicate rides along on every later turn. Saves: the file's token cost, paid once instead of on every subsequent message.
Grep for the function instead of dumping the directory. A targeted search returns the ten lines that matter. Saves: the gap between ten relevant lines and an entire tree, repeated for as long as the context stays open.
Batch your feedback into one message. Five small follow-ups each re-send the context accumulated so far; one message listing all five issues sends it once. Saves: four re-sends' worth of everything already in the conversation.
Clear context between unrelated tasks. Claude Code's cost documentation recommends
/clearwhen switching to unrelated work, and/compactwith custom instructions when you need continuity. Saves: compounding, not one-time. An unmanaged thread is how a single 80,730-token frame dump becomes roughly 40% of a 200,000-token context window gone before the fix starts.Keep the bug report to one precise sentence. Text is already the cheapest input there is; a tight sentence against a hedging paragraph is a small, real saving with no downside. Saves: the smallest number on this list, and the easiest.
Reach for a bundler once visual review is a daily habit. A tool that narrates and selects frames automatically, such as Walkie, packages a review into one bundle instead of a raw recording. Saves: a measured bundle runs about 3,500 tokens against 80,730 for the frame-dumped equivalent — 23x fewer, roughly $0.04 instead of $0.81 at $10 per million input tokens. The free tier, 7 recordings and 10 screenshots with no card, covers casual use; it is $39 once, not a subscription, if it becomes how you work.
Which give the biggest return?#
The top three by a wide margin: never frame-dumping a recording, worth tens of thousands of tokens per review; narrating instead of screenshotting, worth about 2,500 tokens each time; and sending one cropped screenshot instead of four, worth 8,073 tokens for about 30 seconds of care. Everything past that is real but smaller.
| Tactic | Effort | What it saves |
|---|---|---|
| Do not frame-dump a recording | None — just do not | 80,730 tokens avoided per 30-second review |
| Narrate instead of screenshot | ~10 seconds | ~2,500 tokens per instance: 150 against 2,691 on Claude 4.7 and later |
| One screenshot, not four | ~30 seconds | 8,073 tokens per page reviewed |
| Crop before pasting | ~10 seconds | 2,053 tokens on an 800x600 crop: 638 against 2,691, high-resolution tier |
| Clear context between tasks | One command | Prevents compounding toward the ~40% of a 200k window one frame dump costs |
| Point at a file path, do not paste | Habit change | A file's full token cost, avoided on every later turn |
At $10 per million input tokens, the top row alone is $0.81 avoided per review; Anthropic's pricing page publishes input rates from $1 per million tokens on Claude Haiku 4.5 to $10 per million on Claude Fable 5, so quote the rate rather than the model. For the three-way format breakdown, see video, screenshots or words, and for why images specifically get re-billed turn after turn, why UI work burns more Claude Code tokens than backend work.
Which are not worth it?#
Three common instincts move the needle less than people expect, and one of them is a trade rather than a saving.
- Recompressing images for file size. Token cost comes from the patch grid, not JPEG quality or megabytes. A smaller file saves disk space, not tokens.
- Dropping your whole display to a lower resolution. This does cut cost — a 4K capture hits the 4,784-token cap while the same content at 1080p costs 2,691, about 44% less — but living at a worse resolution all day is a real cost. Cropping gets the same saving without the trade.
- Switching models specifically for image-heavy turns. This is a trade, not a free win. Standard-tier models downscale a 1080p screenshot to 1456x819 and charge 1,560 visual tokens instead of the 2,691 a Claude 4.7 or later model charges, and per-million rates differ by model — but you are also changing the quality of the reasoning you get back.
- Deleting messages mid-conversation. It can shrink context, but it is easy to delete something the agent still needed. The safer version is tactic 9: clear between tasks, not in the middle of one.
What to do first this week#
Start with the three free tactics that take under a minute: stop frame-dumping recordings, describe or narrate a bug before screenshotting it, and crop to one image instead of sending four. Add a fresh session per unrelated task after that.
None of this requires installing anything. Try those four for a week before touching anything else — they account for most of the savings available, and none of them change how you work with the agent, only what you hand it. The measured baseline behind the 3,500-token bundle figure is published in the visual-context token cost research, and you can price your own captures in the Walkie token calculator before you paste them.