# Why your AI agent can't see what it built

> An AI coding agent works in text, so it sees a render only when a tool hands it one. Claude Code can now screenshot a page or your Mac screen, but both are off by default and each look costs 2,691 input tokens at 1080p, capped at 4,784.

Author: Roberto Ercole · Published: 2026-08-05 · Updated: 2026-08-16 · Canonical URL: https://usewalkie.com/blog/why-your-agent-cant-see-what-it-built/

---

An AI coding agent can't see what it built because its working world is text: the code it writes, the files it reads, the output of the commands it runs. Rendering happens somewhere it has no access to. It sees a render only when a tool captures one and hands it back as an image — and in 2026 those tools exist, are off by default, and cost real tokens every time they fire.

**Updated August 2026.** The gap is narrower than it was and it has not closed. What changed is perception: Claude Code can now drive a browser and, on macOS, look at your actual screen. What has not changed is initiative and intent — nothing in an agent's loop decides, unprompted, that something looks wrong.

## What does an agent actually perceive by default?

By default, an agent perceives text only: source code, file contents, and terminal output. Anthropic's own description of Claude Code is "an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools" — every verb in that sentence operates on text, and none of them describe looking at a screen ([Claude Code overview](https://code.claude.com/docs/en/overview)).

A browser turns HTML, CSS and JavaScript into pixels through a process with no text representation the agent automatically receives. The agent can describe what CSS *implies* — `padding: 24px` reads as generous spacing — because that is inference from source, not observation of a result. In a transcript those two operations look identical. They are not.

This is the mechanism underneath the wider question of [what visual feedback for AI coding agents is](/blog/what-is-visual-feedback-for-ai-coding-agents/) and why it needs its own tooling at all.

## Can Claude Code take its own screenshot?

Yes, through two documented paths, both of which you have to switch on. The Chrome integration reads the DOM, console and network panel and can screenshot a page. Computer use, on macOS, lets Claude open apps, click, and see the screen. Neither is on out of the box.

**Chrome.** Claude Code connects to the Claude in Chrome extension for "live debugging: read console errors and DOM state directly" and "design verification: build a UI from a Figma mock, then open it in the browser to verify it matches." It requires extension version 1.0.36 or higher, a direct Anthropic plan, and the `--chrome` flag or `/chrome` set to enabled by default. Anthropic also warns that enabling it permanently "increases context usage since browser tools are always loaded" ([Claude Code Chrome docs](https://code.claude.com/docs/en/chrome)).

**Computer use.** On macOS, Claude Code exposes a built-in MCP server named `computer-use` that is disabled until you enable it in `/mcp` and grant Accessibility and Screen Recording permission. Anthropic ships it as "a research preview on macOS that requires a Pro or Max plan," not available on Team or Enterprise, and not available in non-interactive mode with the `-p` flag ([Claude Code computer-use docs](https://code.claude.com/docs/en/computer-use)).

So the honest 2026 statement is narrower than "it can't see." It is: an agent sees nothing until you wire a capture tool in, and even then it looks only when something tells it to.

## What does each look actually cost?

Every look is priced as an image. Anthropic's vision documentation states the rule directly: "Claude views images in patches instead of pixels. Each patch is a 28x28-pixel block of the image, referred to as a visual token. An image, therefore, costs `⌈width / 28⌉ × ⌈height / 28⌉` visual tokens" ([Anthropic vision docs](https://platform.claude.com/docs/en/build-with-claude/vision)).

On the high-resolution tier — Claude 4.7 and later models, long edge up to 2576 px — that count is capped at 4,784 visual tokens. On the standard tier, used by all other models, images are downscaled to a 1568 px long edge and capped at 1,568. A 1920×1080 screenshot lands at `⌈1920/28⌉ × ⌈1080/28⌉ = 69 × 39 = 2,691` tokens on the high-resolution tier, under the 4,784 cap, and at 1,560 on the standard tier after downscaling. Anthropic publishes both numbers in the same table, so the claim is checkable rather than asserted.

Computer use pays a smaller bill, because Claude Code downscales first: Anthropic documents that "a 16-inch MacBook Pro at native Retina resolution captures at 3456×2234 and downscales to roughly 1372×887." Run that through the same patch formula and one full-screen look costs `⌈1372/28⌉ × ⌈887/28⌉ = 49 × 32 = 1,568` tokens (calculated), still on the high-resolution tier with its 4,784 cap.

| Capture | Input tokens | Basis |
|---|---|---|
| One 1920×1080 screenshot, high-res tier | 2,691 | calculated, formula above, cap 4,784 |
| The same screenshot, standard tier | 1,560 | calculated, downscaled to 1456×819, cap 1,568 |
| One downscaled computer-use screenshot, 1372×887 | 1,568 | calculated, formula above, cap 4,784 |
| One real 1288×811 test screenshot | 1,334 | measured |
| A complete Walkie bundle | ~3,500 | measured, typical |

Run your own dimensions through the [screenshot token calculator](/calculator/), or read the cross-vendor version in the [visual context token cost benchmark](/research/visual-context-token-cost/).

## Why doesn't more resolution solve it?

Resolution is not the constraint; meaning is. Extra pixels give the vision system more patches to encode, which raises the token count until it hits the 4,784 cap on the high-resolution tier (Claude 4.7 and later). None of those patches tell the agent which detail matters or why.

The same math gets worse fast on video. Frame-dumped at 1 fps and priced as 1080p stills, 30 seconds of screen recording costs `30 × 2,691 = 80,730` input tokens (calculated) — about 40% of a 200k context window before a single line is fixed. At $10 per million input tokens that is $0.81. If the frames are full-resolution Retina rather than 1080p, every frame hits the cap instead, and the same 30 seconds costs `30 × 4,784 = 143,520` tokens (calculated) as an upper bound. Both numbers describe pixels, and neither carries a reason. The mechanics are worked through in [why your agent can't watch a screen recording](/blog/why-ai-cant-watch-screen-recordings/), and the per-token breakdown in [what a screenshot costs an AI coding agent](/blog/screenshot-token-cost-ai-coding-agent/).

## What closes the gap?

Two things together: a capture of the actual render, and a statement of what is wrong and where. A screenshot supplies the first alone. Pointing at the screen while narrating supplies the second. Neither one by itself gets an agent to a correct judgment about what it built.

| Method | Input tokens | Pixels included | Intent included |
|---|---|---|---|
| No image, code only | 0 | No | No |
| One pasted 1080p screenshot | 2,691, high-res tier, cap 4,784 (calculated) | One moment | No |
| 30 seconds frame-dumped at 1 fps | 80,730 (calculated) | Redundant | No |
| Narrated, pointed review bundle | ~3,500 (measured typical) | Selected | Yes |

That last row is 23× smaller than the frame dump, or 96% less (calculated), and costs $0.04 against $0.81 at $10 per million input tokens. The difference is not compression. It is that a spoken transcript — about 150 tokens for a typical review, measured, range 100–200 — states why a frame matters. That pairing is what a [REVIEW.md bundle](/blog/what-is-review-md-bundle/) is for.

In practice, closing the gap means three steps:

1. Capture the actual render — a screenshot, the Chrome integration, or computer use on macOS.
2. Select the moment that matters. A person doing this beats a script guessing.
3. Attach the reason: what is wrong, where, and what right looks like instead.

## Whose job is verification?

Yours, and Anthropic's documentation says so plainly. Its Claude Code best-practices guide opens the verification section with "Give Claude a check it can run: tests, a build, a screenshot to compare," and warns that "Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop" ([Claude Code best practices](https://code.claude.com/docs/en/best-practices)).

That guide's recommended prompt for UI work is explicit about who supplies the target: paste a screenshot, ask Claude to implement it, take a screenshot of the result, compare the two, and list the differences. The agent runs the comparison. You decide what it is comparing against.

Split the work accordingly. Let the agent verify what is checkable in text — types, tests, logic, error handling. Keep visual verification as a step a person owns, and make it fast with a structured pass like the [12-point checklist for verifying AI-built UI](/blog/verifying-ai-built-ui-checklist/) rather than eyeballing a page from scratch.

## What to do the next time an agent says done

Don't accept its word for the part it structurally cannot see unless you gave it a way to look. Open the app yourself, or hand it a screenshot with a specific caption naming the location and the expected state — the discipline covered in [how to describe a visual bug to an AI coding agent](/blog/describe-a-visual-bug-to-ai/).

If you do this often enough that typing captions is the slow part, narrating over the screen carries the same information for less typing and fewer tokens. How you capture matters less than making sure the agent receives both the render and the reason, not one without the other.

---

Read the HTML version: https://usewalkie.com/blog/why-your-agent-cant-see-what-it-built/
