Twelve states account for most of what an AI coding agent ships broken without knowing it: empty data, loading, errors, long text, dark mode, keyboard focus, mobile width, disabled controls, first run versus returning, slow network, list counts, and reduced motion. Running all twelve after the agent says "done" is the fastest way to catch what it could not see itself.
Updated August 2026. Below is the full list with what to look at, why each one is easy for an agent to miss, and the fastest way to force the state to appear — no new tooling, just the devtools panels and OS settings already on your machine. It is the practical half of how to review what your AI coding agent built.
Why does done not mean done?#
An agent's "done" means the code compiled, the tests passed, and the happy path worked when it checked. It says nothing about the states nobody prompted it to hit. Anthropic names this failure mode directly in its own guide: "the trust-then-verify gap. Claude produces a plausible-looking implementation that doesn't handle edge cases," with the fix stated as "always provide verification (tests, scripts, screenshots). If you can't verify it, don't ship it" (Claude Code best practices).
Checking its own work also costs the agent tokens. One 1920×1080 screenshot runs ⌈1920/28⌉ × ⌈1080/28⌉ = 69 × 39 = 2,691 input tokens on the high-resolution tier — Claude 4.7 and later, capped at 4,784 — per Anthropic's vision documentation, and 1,560 on standard-tier models after downscaling. So an agent that does look tends to look once, not twelve times. That is the architectural reason your agent can't see what it built across every state below.
The twelve checks#
These twelve cover most of what ships visually broken. Each item gives what to look at, why agents miss it, and the fastest trigger.
- Empty states. Look at the first screen a brand-new user or an empty list sees. Agents miss it because they build and test against a database that already has rows, so a genuinely empty state only appears if someone clears the data first. Trigger: wipe the table or local storage and reload on a fresh account.
- Loading states. Look at what renders between the request firing and the response landing — a spinner, a skeleton, or nothing. Agents miss it because local responses resolve in milliseconds, so the code path exists but is rarely seen. Trigger: throttle the network to Slow 3G and watch the first render.
- Error states. Look at what a user sees when the request fails: a timeout, a 500, a rejected form. Agents miss it because their check runs against a working backend, so nothing fails. Trigger: stop the API server, or force a 500 from the devtools network overrides.
- Long text and overflow. Look at names, titles and labels at real maximum length, not the short placeholder the agent typed. Agents miss it because a self-written fixture is almost always short. Trigger: paste a 50-character name into every text field on the screen.
- Dark mode. Look at text contrast, icon visibility, and any image or gradient with a background colour baked in. Agents miss it because most dev environments default to one theme, so a CSS variable can exist without ever being checked against a dark ground. Trigger: toggle the theme and step through every screen just built.
- Focus rings and keyboard navigation. Look at whether tabbing moves focus in a sane order and whether focus is ever visible. Agents miss it because they interact through the DOM or a click simulator, not a keyboard, so a missing
:focus-visiblestyle never registers as a failure. Trigger: put the mouse down and tab through the screen from the top. - Mobile width. Look at the layout at real phone width, not a desktop window dragged smaller. Agents miss it because a component can pass every check at desktop width while overflowing at phone width. Trigger: open the devtools device toolbar and load the page at 375×667.
- Disabled states. Look at buttons and inputs mid-action or pre-condition: submit before the form is valid, a button while a request is in flight. Agents miss it because a happy-path check clicks the button once, already enabled. Trigger: submit before filling the form, then double-click submit to see if it fires twice.
- First run versus returning. Look at whether onboarding, tooltips and empty-state prompts reappear for someone who has already seen them. Agents miss it because every test run starts from a clean slate, so both look identical. Trigger: complete the flow once, reload, and confirm the welcome tour does not replay.
- Slow network. Look at whether the interface stays usable, not just whether it eventually loads. Agents miss it because an instant response never exposes a second click or a mid-request navigation. Trigger: throttle to Slow 3G and interact while a request is still pending.
- Zero, one, many list counts. Look at a list with zero items, exactly one, and enough to force scrolling — three different layouts, not one. Agents miss it because seed data is usually three to five tidy rows, the one count that never breaks a grid, a plural label, or a "showing X of Y" string. Trigger: delete down to zero rows, then to one, then paste in fifty.
- Reduced motion. Look at whether animations respect a user's reduced-motion setting instead of running regardless. Agents miss it because the setting lives in OS accessibility preferences their environment does not have. Trigger: turn on Reduce Motion in system settings and retrigger the animation.
Same list, as a lookup table when you only need the trigger:
| # | Check | Fastest trigger |
|---|---|---|
| 1 | Empty states | Wipe the data, load on a fresh account |
| 2 | Loading states | Throttle the network to Slow 3G |
| 3 | Error states | Kill the API mid-request |
| 4 | Long text and overflow | Paste a 50-character name into every field |
| 5 | Dark mode | Toggle the theme switch |
| 6 | Focus rings and keyboard nav | Tab through with the mouse down |
| 7 | Mobile width | Load the page at 375×667 |
| 8 | Disabled states | Submit before the form is valid |
| 9 | First run versus returning | Reload after completing the flow once |
| 10 | Slow network | Throttle, then interact mid-request |
| 11 | Zero, one, many | Go to 0 rows, then 1, then 50 |
| 12 | Reduced motion | Turn on Reduce Motion, then retrigger |
Which checks fail most often?#
The checks needing no user interaction to surface fail most: empty states, dark mode, and long text. An agent's own test run never needs a genuinely empty database, a dark theme, or a name longer than the one it typed. Transient states — loading, error, slow network — fail almost as often, for the same root cause.
None of this means the agent did badly on the states it did check. The happy path, correctly implemented, is real work. The gap sits specifically in the states a normal build-and-verify loop has no reason to visit.
Can the agent run any of these for you?#
Some of them, if you connect a browser. Claude Code's Chrome integration lets it open a page, resize the window, read console errors and DOM state, and take a screenshot — Anthropic lists "web app testing: test form validation, check for visual regressions, or verify user flows" among its capabilities (Claude Code Chrome docs).
That covers the mechanical half of checks 2, 3, 7 and 10 nicely: throttling, resizing and forcing a failure are all things a browser can be told to do. What it does not cover is the judgment half. Nothing in the integration decides that your empty state reads as an error, or that the dark-mode contrast is technically passing and still unpleasant. Anthropic's framing holds: hand Claude a check it can run, and keep deciding what the check is comparing against.
How long should a pass take?#
A full pass is quick once you know the triggers, because most are a single devtools toggle, a pasted string, or a flipped OS setting rather than a rebuild. The first pass on a new feature takes longer while you locate each control; after that it becomes a habit inside how you already look at a screen.
Run the full twelve on a new screen or a substantial feature. On a smaller change, spot-check whichever checks the change actually touches — a copy edit does not need a dark-mode pass, but any new list view needs check 11 every time, because that is exactly where seed data lies to you.
How do you report what you find?#
Report each failing check by name, not a vague description: "the loading state has no spinner" beats "it looks broken while it loads." Screenshotting every failure adds up fast — four screenshots to cover one page cost 4 × 2,691 = 10,764 input tokens on Claude 4.7 and later, capped at 4,784 per image, with zero narration attached.
- Name the check that failed, using the numbered list above.
- State what you expected at that check.
- State what actually happened instead.
- Note whether it is every case or one edge case.
- Send it as one message, not a stream of commentary.
The same principle applies as in giving feedback that lands: name the location, state the expectation, state what happened. That is also most of what it takes to describe a visual bug precisely. Narrating the twelve checks as you trigger them, instead of screenshotting each one afterwards, is what visual feedback tooling exists to compress; how to run a visual review with Claude Code walks the setup, and the visual context token cost benchmark prices each option side by side.
Bookmark the list above and run it once against whatever an agent just told you was done. Twelve checks, a few minutes, no new tooling required.