Reviewing AI-generated code means checking two different things: whether the logic is correct, which the agent already verified with its own tests, types and linters, and whether the result works and looks the way you meant, which only a human eye confirms. The blind spot is structural, not a discipline problem — the agent never saw its own output rendered.

Updated August 2026. As more shipped code comes from Claude Code, Cursor and Codex, the review step has not shrunk. It has moved. The agent runs its test suite, resolves its type errors, and reports "done" with real confidence, because by every check it can run, it is done. What it has not done is look at the page.

  • 2,691 input tokens — what it costs an agent to read one 1920×1080 screenshot on Claude 4.7 and later, capped at 4,784 (calculated)
  • 0 — how many of those tokens tell it what is wrong, or what you meant instead

Why does AI-generated code need a different review?#

Because its author never experienced its own output. A human writes with a mental picture of the running app; an agent writes code, runs automated checks, and reports success without seeing a rendered pixel. Review has to cover exactly what the agent structurally cannot.

There is also a fluency problem. Generated code tends to be syntactically clean and confidently structured even when wrong, because the model optimises for plausibility as much as correctness. A junior developer's mistake often looks like a mistake — a weird variable name, an obviously missing case. An agent's mistake usually looks like correct code, because it was trained on correct code. That makes a fast skim less reliable on generated diffs than on human ones.

Memory is the other difference. You can ask a colleague why they made a decision; an agent's context resets with the session, so reasoning not captured during review is gone — including for the agent.

What does a plausible-but-wrong change look like?#

It looks like a finished feature with an unverified surface. An agent asked to add a loading spinner while search results fetch will generate the component, wire it to the fetch call, and report the task complete once its type checker and test suite are satisfied.

What it has no way to notice: that the spinner does not match the page's design language, that it flashes for 80 milliseconds on a fast connection and reads as a rendering bug, or that it never appears at all on a slow connection because the component unmounts before the network call resolves. None of that is a logic error. All of it is real, and all of it ships.

What can the agent check itself?#

Anything it can run and compare against a rule: does the code compile, do tests pass, are types correct, does lint pass, did the expected files change. These are closed-loop textual checks. It cannot verify anything needing the running result unless you connect a tool that captures one.

What gets reviewed Can the agent verify it? Why
Type errors Yes Static analysis reads source text
Unit test pass or fail Yes Tests are more code it wrote and ran
Lint and style rules Yes Pattern matching against the source
Whether the build compiles Yes Deterministic pass or fail from the toolchain
Whether the requested files changed Yes A diff is text
Whether the page renders as intended Only with a capture tool connected Requires actual rendered pixels
Whether the layout survives real content Only with a capture tool connected Requires real data and real screen widths
Whether the fix matches your intent No Requires knowing what you meant
Whether it feels slow, jarring or confusing No Requires a human experiencing the running app

Those self-checks are genuinely valuable: they eliminate a whole category of trivial regressions before you look. The mistake is treating them as a proxy for the review rather than the first half of it. Anthropic says as much in its own guide: "Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop" (Claude Code best practices).

Can the agent look at the page now?#

Two rows in that table have moved since 2025. Claude Code can screenshot a browser page through its Chrome integration, which also reads console errors and DOM state, and on macOS it can see your actual screen through the built-in computer-use MCP server.

Both are off until you switch them on, and neither decides on its own that something looks wrong — which is why the bottom three rows of that table have not moved at all. Perception improved; initiative and intent did not. The mechanics, and what each look costs, are in why your agent can't see what it built.

What can only you check?#

Whether the output matches your actual intent, whether it looks right with real content, whether messy data breaks it, and whether the experience feels right rather than merely functions. Four categories, none of which a toolchain reaches on its own.

  • Intent match — does the change solve the problem you had, not just the literal words of the request. An agent can build exactly what you asked and still miss what you needed.
  • Real-content behaviour — does a 40-character company name, an empty search result, or a price with four decimal places break a layout that looked fine with placeholder data.
  • Felt experience — does a transition feel slow, a form feel fussy, an error message read as hostile. None of that shows up in a passing test.
  • Cross-context rendering — dark mode, a narrow viewport, a focus ring after a keyboard tab, a slow connection.

A concrete version: an agent building an onboarding flow generates the form, validates the fields, and passes its tests using "Jane Doe" and "test@example.com." It has no occasion to try a 40-character company name, a name with an apostrophe, or a screen reader tabbing through fields in the wrong order — not because it is careless, but because nothing in its checklist asks it to.

How do you read a diff you didn't write?#

In a fixed order: data shape first, then boundaries — inputs, outputs, error paths — then the logic between them. Reading top to bottom the way you read your own code fails, because there is no narrative already in your head to follow.

Generated code fails in recognisable ways once you know to look: an error path that catches an exception and silently does nothing, a defensive null check on a value that cannot be null there, a test asserting against the function's own implementation instead of its behaviour so it stays green while the behaviour is wrong. None of those jump out on a linear read, because that style assumes you would notice if the general shape were off — and generated code's general shape is almost always fine.

Not everything needs equal scrutiny. Boilerplate and mechanical diffs can be skimmed. Anything touching money, authentication, a data mutation or a public API deserves the full pass. How to review AI-generated code you didn't write covers that order and the five specific failure patterns, with the strings to grep for.

How do you report a fix so it lands?#

Give the agent a precise location and a precise expectation. "The spacing looks off" gives it nothing to act on. "The card padding is 8px; it should be 16px" gives it a fact to check and a target to hit.

Length is rarely the difference. "The button looks wrong" is six words and useless. "The Save button on the settings page is filled teal — it should be outlined, like the Cancel button next to it" is barely longer, and an agent can act on it without a follow-up question.

Visual bugs are the hardest category to put into words, because "off" is a feeling and translating it into a coordinate takes effort people skip under time pressure. How to describe a visual bug to an AI coding agent has a ten-row translation table; how to give feedback to Claude Code so it lands breaks the anatomy down further.

Showing costs something too, and the spread is wide. One 1920×1080 screenshot costs 2,691 input tokens on the high-resolution tier — Claude 4.7 and later, capped at 4,784 — by the patch formula in Anthropic's vision documentation. Four to cover a page cost 10,764. A 30-second recording frame-dumped at 1 fps costs 80,730 tokens, or $0.81 at $10 per million input tokens, against about 3,500 tokens for a narrated review bundle — $0.04 at $10 per million input tokens, a 23× reduction (calculated). Price your own case in the screenshot token calculator, and see what visual feedback for AI coding agents is for how the four kinds of tool compare.

What does a good review loop look like?#

Plan, then execute, then verify — treating verify as its own step rather than something the agent's "done" message satisfies. Planning sets acceptance criteria before code is written; execution is the agent working; verification is a human comparing the running result to those criteria.

  1. Write the acceptance criteria before the agent starts, in plain language it can check itself against.
  2. Let the agent execute and run its own automated checks: types, tests, lint, build.
  3. Stop before merging and actually open what it built.
  4. Compare the running result against the criteria from step one, not against the agent's summary.
  5. Report back with a location and an expectation for anything that does not match.
  6. Re-verify the fix the same way. Do not take "fixed" as an answer either.

None of this has to be slow. Planning is a sentence or two; verification is the two or three minutes it takes to open the page and click through. The loop breaks down not because verification is expensive but because it is the step with no automatic prompt. The plan → execute → verify loop covers each stage, including the three documented ways to make Claude Code hold its own gate.

How do you know when it's actually done?#

When a human has checked the result against the original acceptance criteria on the real rendered output — not when the agent's automated checks pass and it says so. "Done" from an agent means its own checks are green. It says nothing about empty states, dark mode, or whether the change solved the problem you had.

Treat "done" as a two-part bar: automated checks pass, and a human confirmed the rendered result against intent. If either half is missing, what you have is "the agent believes it's done," a much weaker claim. The mismatch follows a pattern: the agent reports the feature works, the human skims the summary instead of opening the page, and the gap between "the tests are green" and "the thing does what I needed" ships unnoticed. The 12-point checklist for verifying AI-built UI lists the states agents most reliably miss.

What to do next#

The next time an agent tells you something is done, treat that as a claim rather than a fact. Open what it built. Run the real command, load the real page, click through with real data instead of the happy path it tested against. If something is off, name the exact location and the exact expectation instead of the general feeling.

None of it requires a specific tool. A notepad and two minutes of actually looking catches most of what matters. What it requires is treating verification as a distinct step you do on purpose.