A spoken description of your screen costs an AI coding agent roughly 150 tokens. One 1080p screenshot costs 2,691 on Claude 4.7 and later, where any single image is capped at 4,784. Frame-dumping 30 seconds of video into stills costs 80,730 tokens — about 40% of a 200,000-token context window gone before the agent reads your bug report. Words are cheapest by a wide margin, but only when they name an exact location.

Updated August 2026. This breakdown uses Anthropic's published patch formula plus figures measured on a real test bundle. It is part of the wider picture of Claude Code token costs: once you see where tokens go, the three formats turn out to have very different price tags for describing the same bug, and the cheapest is not automatically the best.

  • 80,730 tokens — to frame-dump a 30-second review at 1080p, 1 fps
  • 2,691 tokens — for one 1080p screenshot on Claude 4.7 and later
  • ~150 tokens — for a typical spoken review, transcribed

Can an AI agent watch a video?#

No. Claude's context accepts text and images, not video files. Anthropic's vision documentation lists JPEG, PNG, GIF and WebP as the supported image formats and states that animations are unsupported, with only the first frame used. The only route from a recording to something an agent can see is a sequence of stills, each billed as its own image.

Nobody actually pastes a raw .mp4 into an agent's context; there is no code path for it. What happens in practice is frame extraction: a tool samples the recording at some interval, usually once per second, and hands the agent that stack of stills. So "how much does video cost" collapses into "how many frames, at what resolution" — a question the next section answers exactly. The full mechanics are in why your coding agent cannot watch a screen recording.

What does a frame dump cost?#

A 30-second review at one frame per second is 30 still images. At 2,691 tokens per 1080p frame on Claude 4.7 and later, under the 4,784-token per-image cap, that is 80,730 tokens, or about $0.81 at $10 per million input tokens. Nothing in that total tells the agent which of the 30 frames matters.

The per-frame number comes straight from the published patch rule: each 28x28-pixel block is one visual token, so an image costs ceil(width / 28) × ceil(height / 28), capped at 4,784 on the high-resolution tier. For a full 1920x1080 frame that is 69 × 39 = 2,691 tokens, exact, no rounding. Multiply by 30 frames and the total is 80,730 — close to 40% of a 200,000-token context window spent before the agent has done anything but look. Anthropic now ships 1M-token windows by default on Claude Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 and Sonnet 4.6 (context-window documentation); that raises the ceiling without changing what a frame costs.

This article uses one baseline throughout: 2,691 tokens per 1080p frame on Claude 4.7 and later. A different baseline prices a full-resolution Retina recording at the flat 4,784-token cap per frame, because those frames exceed the cap outright. That is a different comparison for a different question, and the two are never mixed here.

What does a screenshot cost?#

One 1080p screenshot costs 2,691 tokens on Claude 4.7 and later, capped at 4,784 per image; a smaller real screenshot at 1288x811 measured 1,334. A realistic bug report often needs several views, and four screenshots covering one page cost 10,764 tokens with zero narration, so the agent still has to guess why each was sent.

This is the comparison that matters day to day. Almost nobody sends a coding agent a video; plenty of people paste screenshots. A single one costs the same as a single frame of the dump above, or less if it is a crop rather than a full-screen capture. The cost climbs fast once a bug needs context, and every one of those four images arrives with no explanation attached. For the full breakdown across common resolutions, see how many tokens a screenshot costs an AI agent, or price your own captures in the Walkie token calculator.

What do words cost?#

A spoken review of the same screen, transcribed, costs far less than any image. A measured 15-second test transcript ran 65 tokens; a typical spoken walkthrough runs about 150, in a 100–200 range. That is roughly 18x fewer tokens than one screenshot (2,691 ÷ 150 = 17.9, Claude 4.7 and later). But cheap words only help if they say exactly where the problem is.

Text is cheap because language is a compact way to encode intent; an image spends tokens encoding pixels whether or not those pixels are relevant. The catch shows up immediately in practice: "the spacing looks off" and "the spacing on the checkout card is off" cost almost the same number of tokens, and only one of them is useful. Most written feedback defaults to the first kind.

Which carries the most meaning per token?#

Words win on raw efficiency, but tokens are not the only unit that matters. "The button is misaligned" costs almost nothing and says almost nothing. "The submit button on the pricing page sits right of the card edge, not centred like the others" costs barely more and pinpoints the fix. Specificity, not brevity, makes words useful.

That cuts against the instinct to pick the cheapest format and move on. A vague sentence and a precise one are nearly the same price; the precise one just requires the writer to name what and where. Most people skip that work, which is why "words are cheapest" does not automatically mean "words are best" in practice.

  • Words are the cheapest format by a wide margin, but only when they name an exact location.
  • A vague sentence and a precise one cost almost the same, so precision is nearly free once you are already writing.
  • What reliably works is words plus one or two selected frames — not words alone, and not a full frame dump.

What is the practical answer?#

For a single obvious bug, a specific sentence is enough on its own. For anything visual — spacing, alignment, colour, motion — pair that sentence with one or two frames showing exactly where. That combination measures about 3,500 tokens typical, roughly 23x fewer than frame-dumping the same review.

Approach Tokens, one 30-second review Cost at $10 per million input tokens What the agent actually gets
Video, frame-dumped at 1 fps, 1080p 80,730 — 30 × 2,691, Claude 4.7 and later $0.81 Every frame, no signal for which one matters
Screenshots, four covering one page 10,764, zero narration $0.11 A few fixed views, no why, no order
Words only, typical spoken review ~150 under $0.01 Cheap, but only useful if it names an exact location
Words plus a couple of selected frames ~3,500, measured typical $0.04 Location, intent and a visual anchor together

Quote the rate rather than the model: Anthropic's pricing page publishes input rates from $1 per million tokens on Claude Haiku 4.5 to $10 per million on Claude Fable 5, and those move while the arithmetic does not. On a Pro or Max subscription the same tokens are drawn from a plan allowance rather than billed in dollars, and Claude Code's cost documentation covers where to see that. The measurement method behind the 3,500-token bundle figure is published in the visual-context token cost research.

No single format wins every case. A one-line, unambiguous bug — "the footer copyright year is wrong" — needs nothing but a precise sentence. Anything where "where" or "what it looks like" is the whole point needs a frame attached, because prose is a poor tool for describing geometry. What consistently loses is the two extremes: a full frame dump nobody asked for, and a vague sentence with no image to anchor it.

A short checklist for keeping feedback cheap and specific:

  1. Name the exact element and where it lives on the page, not just "the layout".
  2. State what is wrong and what you expected, in the same sentence.
  3. If the bug is spatial — spacing, alignment, colour, overlap — attach one frame instead of describing geometry in prose.
  4. Leave everything else out. A short, precise note at the right location beats a long, rambling one.
  5. Skip the full recording. If a frame-by-frame dump feels necessary, the real problem is usually that the location has not been named yet.

Several tools now generate the words-plus-frames pairing automatically, narration transcribed alongside the frames it points at, and they differ mainly in how frames get selected and how they are priced — how Walkie compares to Clipy and Loom lines up the current field. The same result is achievable by hand: record nothing, write one precise sentence, and attach a single screenshot only when the bug is genuinely visual.