Updated August 2026. Screen recording feels like the obvious move when something breaks: hit record, narrate the bug, hand the file to your coding agent. It does not work, and not because the idea is bad. Claude's API was never built to read video files. The only route from a recording to something an agent can look at is turning it into a stack of stills, and that route costs far more than it looks like it should.
Can Claude watch a video file?#
No. Anthropic's vision documentation lists JPEG, PNG, GIF and WebP as the supported formats, states that animations are unsupported and only the first frame is used, and describes images as content blocks supplied as base64 data, a URL, or a Files API reference. There is no video content block and no "watch this recording" call.
Anything shaped like a video has to become a sequence of still frames before an agent sees any part of it at all. That conversion is the whole story of this article, because it is where the cost appears.
Why doesn't my coding agent support .mp4 or .webm?#
Coding agents such as Claude Code and Cursor inherit whatever the underlying model API supports, and the vision input is images only. Native video would mean reasoning over motion and audio streams directly — a materially different capability from the per-image patch pricing the API actually implements.
That is not a client-side limitation someone forgot to build. It is the shape of the API. Every tool that claims to let an agent "watch" a recording is doing the same thing underneath: sampling frames, discarding what a still image cannot capture, and sending what is left as ordinary image tokens.
This explains some confusing behaviour. If you have ever pasted a video link or a file path into a chat and had the agent respond as though it understood the recording, it did not watch anything. It was reasoning from a filename, a transcript you supplied separately, or a tool result — not from the pixels.
What happens if you frame-dump a video into images?#
Frame-dumping means extracting stills from a recording at a fixed rate, commonly one frame per second, and sending each as a separate image. A 30-second recording becomes 30 full-resolution screenshots, each priced by the patch formula, stacked one after another into the same conversation.
Nothing about that process is smart. It does not know which second shows the bug and which second is you scrolling past it. Every frame is billed the same, whether it is the exact moment the button breaks or three seconds of an unchanged loading spinner.
Does a lower frame rate fix the token problem?#
Not really. Dropping the sample rate lowers the bill without fixing the underlying issue: a lower rate is more likely to skip the exact moment the bug shows up. You trade cost for coverage — cheaper, and blinder to the second that mattered.
No frame rate reliably solves this, because a fixed-interval sampler has no idea which second matters. Only the person who saw the bug happen does. Narrating the moment you see something break already carries that judgment; a timer sampling every second, or every five, is guessing. That is the deeper problem a slider cannot fix: it treats every second of screen time as equally important, when almost none of it is.
How many tokens does a 30-second recording cost as frames?#
About 80,730 tokens (calculated): 30 frames at 1 fps, each a 1080p image priced at 2,691 tokens on Claude 4.7 and later, where a single image is capped at 4,784. That is close to 40% of a 200,000-token context window gone before your agent has read a word of what you meant.
The per-frame figure comes from the published patch rule — one visual token per 28x28-pixel block, ceil(width / 28) × ceil(height / 28), so 69 × 39 = 2,691 for 1920x1080. On standard-tier models each frame is downscaled to 1456x819 and costs 1,560, which still totals 46,800 tokens for the same 30 seconds. Window sizes have grown — Anthropic's context-window documentation now lists 1M tokens by default on Claude Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 and Sonnet 4.6 — but a bigger window does not make a frame cheaper. For the full arithmetic at other resolutions, see how Claude calculates image tokens, or run your own dimensions through the Walkie token calculator. If you want certainty rather than an estimate, Anthropic's token-counting endpoint returns the exact input-token count for a payload, images included, and is free to call.
How does a frame-dumped recording compare to sending screenshots?#
Frame-dumping is the most expensive way to get a screen recording in front of an agent: more expensive than a handful of hand-picked screenshots, and dramatically more expensive than a purpose-built review bundle. Four approaches to the same review, each labelled measured or calculated:
| Approach | Tokens | Cost at $10 per million input tokens | Type |
|---|---|---|---|
| 30s recording, frame-dumped at 1 fps | 80,730 | $0.81 | calculated |
| 4 screenshots covering the same page | 10,764 | $0.11 | calculated |
| Full narrated bundle: transcript, pointed frames, REVIEW.md | ~3,500 | $0.04 | measured |
| Spoken transcript alone, no images | ~150 | under $0.01 | measured, range 100–200 |
Quote the rate rather than a model — Anthropic's pricing page lists input rates from $1 per million tokens on Claude Haiku 4.5 to $10 per million on Claude Fable 5, and those change while the arithmetic does not.
The gap is not subtle. Frame-dumping loses to plain screenshots because most of a 30-second recording is visually redundant: you pay full image price for near-duplicate frames of a screen that has not changed. It loses harder to anything carrying narration and intent instead of raw pixels, because words are nearly free next to images. Even the screenshot column costs more than it needs to if the four overlap or include screen area unrelated to the bug. The measured basis for the bundle row is published in the visual-context token cost research.
What should I send my agent instead of a screen recording?#
Skip the recording. Narrate the bug in your own words, take one or two targeted screenshots of exactly what is wrong, and state what you expected against what happened. A spoken walkthrough runs about 150 tokens (measured), and two well-chosen screenshots beat thirty redundant frames of a mostly-static screen.
- Narrate first, in plain language. Describe the bug before touching a screenshot tool; it forces you to name what is wrong instead of hoping pixels explain it.
- Point at what matters. Crop or annotate so the agent's attention lands on the broken element, not the whole desktop.
- Take one or two targeted screenshots, not a reel of near-duplicate frames.
- Write the expected-against-actual sentence yourself. Do not make the agent infer intent from a pile of images.
For the three-way cost comparison behind that advice, see video, screenshots or words, and for the category of tools built around this constraint, what visual feedback for AI coding agents actually is.
What do screen-recording-for-AI tools actually do under the hood?#
Every tool in this category, Walkie included, ultimately hands the agent images plus text, because that is the only visual input the API accepts. The difference is not whether video gets sent raw. It is which frames get chosen, and how much surrounding context — voice, pointer position, timing — rides along with them.
A tool that samples frames blindly at a fixed interval is doing the same expensive, low-signal thing as manual frame-dumping, just automated. A tool that lets your voice and pointer decide which moments matter is solving a different problem: not how to get pixels to the model, but how to get the right pixels there with enough context that the agent does not have to guess. None of this makes Walkie the only option, and it is not — the constraint comes from the model API, not from any vendor's design choices.
The bottom line#
Video-in, agent-reads-it-out was never on the table. Claude reads images, and a screen recording has to be broken into a stack of them first. Frame-dumping does that the expensive way: 80,730 tokens for 30 seconds at 1 fps on Claude 4.7 and later, most of it redundant. Screenshots do it more cheaply at 10,764 for four. Narration plus a couple of targeted frames does it cheapest of all, at about 3,500 tokens, because it replaces guesswork with a sentence. The full per-image breakdown is in how many tokens a screenshot costs an AI agent.