The plan → execute → verify loop is the three-step cycle behind any AI coding session: the agent states what it is about to change, makes the change, and then someone confirms the result does what was intended. Most people run the first two steps and stop, which is backwards — verification is where the quality decision actually gets made.

Updated August 2026. The pattern predates AI coding; it is how careful engineers already work. What agents changed is how easy the last step is to skip, because execution finishes fast and states it is done with real confidence. That confidence is precisely the part nobody verified.

The asymmetry shows up in the numbers. The agent's own checks — tests, types, lint — run in seconds and cost almost nothing. Catching a visual miss afterwards costs tokens: one 1920×1080 screenshot handed back to the agent costs ⌈1920/28⌉ × ⌈1080/28⌉ = 69 × 39 = 2,691 input tokens on Claude 4.7 and later, under that tier's 4,784-token cap, per Anthropic's vision documentation. On older standard-tier models the same image is downscaled and costs 1,560. That is before you have said a word about what is wrong with it.

What is the loop?#

The loop is three named stages for any change an agent makes. Plan states what will change and how success will be checked. Execute makes the change and runs whatever checks produce text. Verify confirms the result matches intent — the stage only a human fully closes, because only you know what you meant.

Claude Code's documented workflow maps onto it almost exactly, with four phases instead of three: Explore, Plan, Implement, Commit. Anthropic's guidance is to "separate research and planning from implementation to avoid solving the wrong problem," entering plan mode by "pressing Shift+Tab until the status bar shows ⏸ plan mode on, or start the session with claude --permission-mode plan" (Claude Code best practices). Explore and Plan are the plan stage. Implement is execute. What that four-phase list does not contain is a stage for checking the result against your intent — which is the point of naming verify separately.

What does a good plan look like?#

A good plan names the files that will change, describes the behaviour difference in terms you would observe, and states up front how you will know it worked. That last part matters most: a plan with no stated success condition gives execution nothing to check itself against later.

Three things separate a plan you can verify against from one you cannot:

  • Scope — exactly which files or components change, not "the login flow."
  • Expected behaviour — what should happen after, described as something observable rather than as implementation detail.
  • A success condition — the specific thing that, if true, means it worked. A passing test is one. "The button turns green on save" is another.

Anthropic's own advice is stronger than "write it down." For larger features it recommends having Claude interview you, write the spec to a file, and then "start a fresh session to execute it," because "the new session has clean context focused entirely on implementation." The most useful specs, it adds, "end with an end-to-end verification step that proves the feature works."

What happens during execution?#

Execution is the agent editing files, running commands, and iterating against feedback — test failures, type errors, lint warnings — until its own checks pass. It is mechanical and largely self-correcting, because the agent reads its own tool output and catches its own syntax and logic mistakes without you.

This is the stage agents are best at, for one reason: everything execution checks is text. A failing test prints text. A type error prints text. A linter prints text. The agent reads it, adjusts, re-runs. That loop can run dozens of times inside a single turn, which is exactly why execution feels fast and confident compared to the review stages around it. It is optimising against a target it can measure.

Verification needs eyes on the rendered result, and an agent has none unless you wire a capture tool in. Whether a layout looks right, a colour feels off, or a fix matches what you meant are judgment calls that no volume of passing tests settles.

Tests, types and lint check a claim the agent can state precisely: this input produces that output. "Looks right" is not that kind of claim — it compares a rendered page against a mental picture only you hold. That is the architectural reason your agent can't see what it built, and it is why visual feedback needs its own tooling. A green test suite tells you the logic runs. It says nothing about whether the button is reachable, the copy reads naturally, or the layout survives a long name.

Who can verify what?#

An agent can verify anything with a text-based pass or fail signal. A human has to verify anything needing judgment. The split is clean enough to write down and worth keeping visible, because "all tests pass" reads like completion and is only half of it.

What needs checking Can the agent verify it? Why
Tests pass Yes It ran them and reads the output
Types check Yes Compiler output is text
Lint rules Yes Linter output is text
Diff matches its own stated plan Yes It compares its changes to what it said
Layout renders correctly Only if given a capture It has no view of the page by default
Colour, spacing, motion feel right No Judgment call, not a pass or fail signal
Copy reads naturally No Tone is not checkable against a rule
The fix matches what you meant No Only you know your own intent
Edge cases nobody thought to test No Untested paths pass silently either way

The top four rows are the agent's half of the job, and it does them well. The bottom five are yours. A practical run-through of that bottom half is the 12-point checklist for verifying AI-built UI.

How do you make the agent hold its own gate?#

Give it a check it can run, then decide how hard that check gates the stop. Anthropic documents three escalating levels, and they trade setup effort for how much of your attention the run needs.

Gate How it works What it costs you
In the prompt Ask Claude to run the check and iterate in the same message Nothing; works on any task today
/goal condition A separate evaluator re-checks the condition after every turn until it resolves One setup step per session
Stop hook A script runs the check and blocks the turn from ending until it passes A script; Claude Code overrides it after 8 consecutive blocks

Anthropic's framing is the useful part: "Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop." It also recommends an adversarial step — a subagent reviewing the diff in a fresh context, "so the agent doing the work isn't the one grading it" — with the honest caveat that "a reviewer prompted to find gaps will usually report some, even when the work is sound."

Gating also pays for itself in tokens. Anthropic lists "give verification targets" among its documented ways to reduce Claude Code token usage, on the grounds that "when Claude can verify its own work, it catches issues before you need to request fixes" (Claude Code cost docs). A check that catches a mistake in the same turn is cheaper than a correction three turns later.

None of that closes the visual half. A Stop hook can assert a test passes; it cannot assert a modal looks right.

How do you close the loop quickly?#

Look at the actual result yourself right after execution, then report back specifically: where the problem is, what you expected, and what is there instead. A vague "this looks off" restarts the whole loop. A precise report lets the agent verify against something concrete.

  1. State the success condition in the plan, before execution starts.
  2. Let the agent run its own checks first. Tests, types and lint are cheap; do not eyeball what a machine confirms faster.
  3. Look at the actual result. Not the agent's summary of the result — the rendered thing.
  4. Report the gap precisely: location, expectation, what is actually there. How to give feedback to Claude Code so it lands covers the anatomy of a report that works.
  5. Confirm the fix against the same view you flagged, not a fresh description of it.

Step 3 is where workflows quietly break down, because looking costs attention and re-describing what you are pointing at costs tokens. A 30-second screen recording frame-dumped at 1 fps and priced as 1080p stills runs to 80,730 input tokens (calculated) — $0.81 at $10 per million input tokens — while a narrated review bundle runs about 3,500 tokens (measured typical), or $0.04 at $10 per million input tokens: a 23× reduction for the same 30 seconds. Check any resolution yourself in the screenshot token calculator, or see the full token breakdown. If you want the loop wired end to end, how to run a visual review with Claude Code walks each step.

Start here#

You do not need new tooling to close this loop today. Write the success condition into the plan before execution starts. Let the agent run its own checks. Then actually look — at the real result, not the agent's description of it — before you say done. That third step is the whole difference between an agent that assists you and one you have to double-check every time. This is the mechanical skeleton under how to review what your AI coding agent built.