Updated August 2026. If you have never opened a REVIEW.md file, here is what is actually in one, why it is the format Walkie hands your agent instead of the recording itself, and how it compares to what you are probably doing now: pasting screenshots, or describing the bug in a sentence and hoping. The category this belongs to is covered in the pillar on visual feedback for AI coding agents.

What is inside a REVIEW.md bundle?#

Six things, in one folder. REVIEW.md is the entry point; transcript.md holds the narration on its own; frames/ holds the stills; manifest.json holds the machine-readable token counts and file map; SAVINGS.md appears when there is a saving to report; and the raw recording — recording.mp4 or .webm — sits alongside for humans.

File What it holds Who reads it
REVIEW.md Title, a how-to-read block for Claude, the transcript with inline frame references, a frame index table, a duration line The agent, first and fully
transcript.md The narration on its own, without the surrounding structure Either
frames/ The captured stills, as JPEGs The agent, selectively
manifest.json Token counts for transcript, frames and total, plus the file map and transcription engine Tooling
SAVINGS.md Tokens saved against frame-dumping the same recording, written only when that saving is positive The agent, at the end
recording.mp4 or .webm The raw video Humans only

The entry point does more than carry data. Right after the title, REVIEW.md contains a section headed "How to read this (for Claude)" — plain rules telling the agent that the transcript is the primary signal, that frame references sit inline exactly where they happened, that it should open only the frames relevant to the question, and that it must never open the recording video. The file tells the agent how to spend its budget reading the file.

Why markdown instead of video?#

Because a language model cannot open a video and can open a markdown file. Anthropic's vision documentation lists the supported image formats as JPEG, PNG, GIF and WebP, and states that animations are unsupported and only the first frame is used. There is no API call that reads an .mp4.

Video optimises for a human watching in real time; markdown optimises for an agent that needs to locate one fact fast. A transcript can be searched, quoted and reasoned over. That is why the template says, in plain language, never to open the recording video file — it is for humans only. The video still gets saved in case a person wants to watch it back, but the agent's job is done entirely inside the markdown. The recording is scaffolding; the roughly 3,500-token bundle is the product.

Can any AI agent read a Walkie bundle?#

Yes. A bundle is standard markdown, standard JSON and standard JPEGs in a normal folder — formats every current coding agent already parses as part of ordinary file reading. No Walkie-specific plugin is needed to open the transcript or the instructions block.

That claim is checkable rather than a slogan. If an agent can read a README, it can read a REVIEW.md. Nothing depends on Walkie's app running, on a proprietary viewer, or even on an MCP server being connected — although Walkie ships one, and Claude Code's MCP documentation covers how a server like it gets registered. The bundle works the moment it lands in a project folder, whether the agent reaches it through a file read, a pasted path or a tool call. Walkie's own docs page on reading the bundle describes the same structure publicly.

How many tokens does a bundle cost?#

About 3,500, measured across real recordings — transcript, the frames the agent opens, and the frame index table combined. For scale, one full 1080p screenshot alone costs 2,691 tokens on the high-resolution tier used by Claude 4.7 and later, which caps any single image at 4,784 visual tokens.

A bundle that covers an entire review — narration, context and several pointed frames — for roughly the cost of one raw screenshot is the whole argument for the format. Zoom out and the gap widens. Sampling the same 30-second review at one frame per second, the naive way an agent might try to watch a recording, comes to 80,730 tokens: 30 screenshots at 2,691 tokens apiece on that same high-resolution tier, against its 4,784 cap, calculated from the formula ceil(width / 28) x ceil(height / 28) in Anthropic's documentation. A bundle covering that same review measures about 3,500 tokens — a 23x reduction, or 96% fewer tokens, for a review an agent can act on. Run your own dimensions through the token calculator.

What actually makes it into the bundle?#

Not everything on screen. The format is built around selectivity, and REVIEW.md states its own constraints so the agent does not have to guess:

  1. The full transcript — everything you said, in order, forming the primary narrative the agent reads first.
  2. Frame references inline — placed in the transcript exactly at the moment a frame was captured, not batched at the end.
  3. A frame index table — a compact lookup so the agent can see which frames exist before deciding whether to open any.
  4. A token-budget line and duration line — metadata telling the agent roughly what it is spending before it spends it.
  5. An explicit instruction to open five frames or fewer — a ceiling written into the file rather than left to the agent's judgement.

What is deliberately absent is a dump of every frame from the recording. A static screen held for ten seconds would bill for ten separate images in a frame dump; in a bundle it produces at most one frame, because nothing about it earned a second.

How does Walkie decide which frames to include?#

By what you pointed at and how long you dwelled there, or what you circled while talking — not by a timer sampling the screen at a fixed interval. Your voice and your cursor are the selection algorithm, and a frame that never got that attention never enters the file.

This is where the honest version of the pitch matters. Clipy already ships a similar recording-to-markdown format — transcript, extracted frames, click coordinates, and an MCP server on npm — free for your first 15 recordings or 2 hours and $9 a month after. Walkie's difference is not that nobody else does this. It is that pointing decides which frames matter by intent rather than by every click logged, and that the format costs $39 once. The full field is in the priced comparison of agent feedback tools.

Is the bundle format open?#

Yes, in the way that matters to a sceptical reader: every file in a bundle is a format you already have a tool for. Markdown, JSON and JPEG, in a plain folder, with no encryption and no proprietary container. Walkie's public docs describe the structure, so you can check it before installing anything.

That matters more in this category than most, because the audience considering a tool like this is the audience that audits vendor claims for a living. A closed format asks for trust; a readable one asks for nothing. It is also a deliberate position on where the moat is: the markdown structure is copyable by anyone in an afternoon. What is not copyable is what decides which frames go in it.

How does a bundle compare with screenshots or raw video?#

A bundle costs about 3,500 tokens and the agent can act on it; four bare screenshots cost 10,764 tokens and carry no context; a raw video costs nothing because the agent cannot open it at all. That last row is not a strawman — it is what actually happens when you hand an agent a recording.

Format Token cost Can an agent open it?
Walkie REVIEW.md bundle about 3,500 tokens (measured) Yes — plain markdown with inline frame references
Raw screen recording (video) Not applicable No — the API reads images, not video
A pile of 4 screenshots, no narration 10,764 tokens (calculated) Yes, but with no context for what to look at

The screenshot row is the comparison that matters, because that is what most people paste today: four screenshots to cover one page cost more tokens than an entire narrated bundle, and still say nothing about which part you care about or why.

Why does the format matter more than the file shape?#

Because the file shape is trivial and the selection is not. A REVIEW.md is not a compression trick applied to a video — it is a different object, built from the transcript outward, with frames earning their place by what you pointed at rather than by existing.

That is why it reads at roughly a third the cost of four bare screenshots, for coverage a screenshot pile cannot match, and why publishing the structure openly costs nothing: the value is the pointer and the voice that decide what goes in. For the raw token math, see what a screenshot costs your coding agent to read; for why video was ruled out as the fix, see why AI can't watch your screen recordings.