# How to describe a visual bug to an AI coding agent

> Describe a visual bug by naming three things: the component and its state, the current value, and the target value. Never a feeling. A 1920x1080 screenshot costs 2,691 input tokens on Claude 4.7 and later, capped at 4,784, and still cannot say which four pixels you meant.

Author: Roberto Ercole · Published: 2026-08-05 · Updated: 2026-08-16 · Canonical URL: https://usewalkie.com/blog/describe-a-visual-bug-to-ai/

---

Describe a visual bug to an AI coding agent by naming the exact component, its current state, and the value it should have instead — never a feeling. "The spacing looks off" gives the agent nothing to act on. "The gap between the card title and body is 4px; it should be 12px, matching the rest of the stack" gives it an edit to make.

**Updated August 2026.** This is a translation problem, not a taste problem. Below: why visual bugs resist description, a ten-row table turning vague phrases into facts, and the three bug types that stay hard to put into words no matter how careful you are.

A full-resolution screenshot does not rescue a vague sentence. One 1920×1080 image costs an agent 2,691 input tokens to read on the high-resolution tier — Claude 4.7 and later, where the count is capped at 4,784 — calculated from the patch formula published in [Anthropic's vision documentation](https://platform.claude.com/docs/en/build-with-claude/vision), and it still does not say which four pixels you meant. Used precisely, words are usually cheaper than the picture, and clearer. This is the input-quality half of [reviewing what your AI coding agent built](/blog/review-ai-generated-code/).

## Why are visual bugs the hardest to report?

Visual bugs are hard to report because they are comparative and subjective by default: "off" describes a feeling relative to an expectation only you hold. A functional bug ships with an error message and a stack trace. A visual bug ships with nothing but your reaction, so the words you choose are the entire bug report.

An agent that gets `TypeError: cannot read property 'map' of undefined` knows exactly what broke and where. An agent that gets "the header looks weird" knows you are unhappy and nothing else. It has no default visual taste to fall back on, so it guesses — usually wrong — or asks a clarifying question that costs you a round trip. The remedy is the same one covered in [how to give feedback to Claude Code so the fix lands](/blog/give-feedback-to-claude-code/): specificity is the only lever you have.

## What does the agent actually receive?

It receives text: your sentence, the source already in its context, and, if you attach one, a screenshot converted into image tokens. It has no felt sense that something is off. Every ambiguity in what you write becomes ambiguity in what it changes.

That is the mechanism behind [why your AI agent can't see what it built](/blog/why-your-agent-cant-see-what-it-built/). Rendering happens where the agent has no access unless you wire in a capture tool, so whatever you hand over in words or pixels is the only evidence it works from. There is no implicit channel where it absorbs your intent. If a fact is not in the message, it is not in the model.

## How do you name a location precisely?

Name the component, not the page region: not "the header" but "the card title in the third row of the pricing table," including its state if relevant. Anthropic's own prompting guidance says the same in general terms — "reference specific files, mention constraints, and point to example patterns" ([Claude Code best practices](https://code.claude.com/docs/en/best-practices)).

A useful location has four parts, and most reports include only the first:

1. The component or section, named as specifically as you can.
2. The state it is in when the bug appears: default, hover, focus, after a click, loading.
3. The file or component name, if you have it open.
4. The viewport or window size, if the bug only happens at one size.

Skip the last two and the agent searches the codebase to find what you mean before it can start fixing anything.

## How do you describe what is wrong versus what you want?

State the current value and the target value as two separate, measurable facts. "The gap is 4px" is incomplete on its own; pair it with "it should be 12px, matching the spacing used elsewhere in the stack." A complaint gives the agent a direction with no destination.

The target does not have to be exact. "About twice the gap used on the card above it" is a real target, because the agent can go measure that gap and match it. What breaks reports is one-sidedness: all complaint and no destination, or a target with no stated problem to fix.

## When is a screenshot worth the tokens?

A screenshot earns its cost when the bug is a relationship a sentence cannot compress: overlapping elements, a layout broken in a way you cannot name, or a colour wrong only in context. For a single nameable property — a 4px gap, a wrong hex code — a precise sentence is cheaper and just as exact.

| What you send | Input tokens | Basis |
|---|---|---|
| A precise sentence | Tens of tokens | Text, priced as text |
| One 1288×811 cropped screenshot | 1,334 | measured |
| One 1920×1080 screenshot, standard tier | 1,560 | calculated, downscaled to 1456×819, cap 1,568 |
| One 1920×1080 screenshot, high-res tier | 2,691 | calculated, Claude 4.7 and later, cap 4,784 |
| Four screenshots to cover one page | 10,764 | calculated, 4 × 2,691 |
| A spoken walkthrough transcript, typical | ~150 | measured, range 100–200 |

Neither screenshot figure tells the agent which part of the image you meant. A screenshot without a location and a target value is still an incomplete report — just a more expensive one. Price your own dimensions in the [screenshot token calculator](/calculator/), or read the [full breakdown of what a screenshot costs](/blog/screenshot-token-cost-ai-coding-agent/).

## How do you translate vague phrases into facts?

Swap the feeling for a number, a token name, or a named state. Spacing and colour are usually just unmeasured. Alignment and typography get described relatively rather than numerically. Motion and state get described as a feeling instead of a before-and-after. Ten real examples, same fix each time:

| What people say | What the agent actually needs |
|---|---|
| "The spacing looks off." | "The gap between the card title and body is 4px; it should match the 12px used elsewhere in the stack." |
| "There's too much padding on that button." | "The button has 24px of horizontal padding; every other button in this file uses 16px." |
| "The blue is wrong." | "The link colour is #3B82F6; the rest of the UI uses #2563EB. Change this one to match." |
| "It doesn't feel dark enough." | "The panel background is #1F2937; drop it to #111827 to match the sidebar." |
| "Things aren't lined up." | "The icon and label are centred against different baselines. Align both to the row's vertical centre." |
| "The button's in the wrong place." | "The submit button is left-aligned; it should be right-aligned, flush with the input field above it." |
| "The text is too big." | "The heading is 32px; this level uses 24px everywhere else, and this one should match." |
| "The font looks different here." | "This heading renders in the fallback sans-serif, not the site's usual font. Check whether font-family is being overridden." |
| "The animation feels janky." | "The modal's enter transition drops a frame around 100ms in. Try a standard easing curve or a shorter transform distance." |
| "The button doesn't look right when I click it." | "The button's pressed state has no visual change at all. Add a background shift or a slight scale-down on the active state." |

Every right-hand entry carries what the left-hand one lacks: a location, a current value, and a target value.

## How do you report a motion bug?

Name where in the motion it goes wrong and which property is at fault — speed, easing, or a dropped frame — because a sentence can describe a start and an end but not the curve between them. "The animation feels janky" is a symptom, not a diagnosis.

You are detecting something real without being able to name it from memory: a dropped frame, a linear ease where a spring was expected, a duration that runs long. Say roughly when in the motion it happens and which of those three it feels like. A three-second clip settles in seconds what a paragraph will not.

## How do you report a state-transition bug?

Report the exact sequence of actions that triggers it, in order, and what you expected at the moment it did not happen. Some bugs exist only between two states, so describing the before and after separately misses the bug entirely — the transition is the bug.

The loading spinner that never clears, the tooltip that stays open after you click elsewhere, the form that shows stale data for one render before catching up: each of those looks correct in both end states. Write the trigger sequence as numbered steps and mark the step where the expectation broke.

## How do you report a viewport-specific bug?

Give the exact width, or the breakpoint name if the codebase has one, and describe what happens at that width that does not happen above or below it. A layout correct at 1440px and broken at 768px is invisible until someone resizes the window to that width.

"It breaks on mobile" is a hint, not a bug report. If Claude Code has its [Chrome integration](https://code.claude.com/docs/en/chrome) connected, it can open the page, resize, and read the DOM and console itself — Anthropic lists "web app testing: test form validation, check for visual regressions, or verify user flows" among its documented capabilities. It still needs you to name the width and what should happen there. For the wider set of states agents miss, see the [12-point checklist for verifying AI-built UI](/blog/verifying-ai-built-ui-checklist/).

## What should you do the next time something looks wrong?

Name the component, its state, the current value, and the value you want, in that order. If the bug is motion, state-based or viewport-specific, add a clip, a repro path, or an exact width instead of a longer sentence.

1. Name the component and its state.
2. Write the current value.
3. Write the target value, and where it comes from — another component, a design spec, a past version.
4. If it is motion, state or viewport-specific, attach a clip or the exact width instead of more prose.
5. Send it before you second-guess the wording. A rough report with a real number beats a polished one without.

If pointing at the screen is faster than writing the sentence, that is the job [visual feedback tooling](/blog/what-is-visual-feedback-for-ai-coding-agents/) does — you talk while you point, and the agent gets the transcript and the selected frame instead of a paragraph you had to compose. [How to run a visual review with Claude Code](/blog/how-to-run-a-visual-review-with-claude-code/) walks that setup end to end. Either path works, as long as the location, the current value and the target value all make the trip.

---

Read the HTML version: https://usewalkie.com/blog/describe-a-visual-bug-to-ai/
