# How to review AI-generated code you didn't write

> Read a generated diff in a fixed order: data shape first, then boundaries, then logic last. Top-to-bottom reading misses what generated code gets wrong. USENIX Security 2025 measured hallucinated package names in at least 5.2 percent of commercial-model outputs and 21.7 percent of open-source ones.

Author: Roberto Ercole · Published: 2026-08-05 · Updated: 2026-08-16 · Canonical URL: https://usewalkie.com/blog/reviewing-code-you-didnt-write/

---

Review AI-generated code in a fixed order: data shape first, then boundaries — inputs, outputs, and error paths — then logic last. Reading top to bottom fails because generated code is usually locally plausible, meaning each line looks reasonable in isolation, while being globally questionable, meaning the pieces do not cohere. A linear read is built to catch the first kind of problem and is structurally blind to the second.

**Updated August 2026.** The order holds regardless of which agent wrote the code. Claude Code, Cursor and Copilot fail in similar shapes, because none of them can watch their own code run and notice it does not do what its docstring claims. That check is still yours, and it is one half of [how to review what your AI coding agent built](/blog/review-ai-generated-code/).

## Why is generated code harder to read?

Generated code is harder to read because it is optimised to look finished, not to be understood. A human writing new code builds structure around a problem held in their head; a model predicts the next plausible token. The result can be syntactically clean and locally correct while skipping the reasoning trail a reviewer normally relies on.

When a teammate hands you a pull request, you can reconstruct why they made a choice: ask them, check the commit history, recognise the pattern from three other files. A model's stated reasoning, where you can see it, is a plan written before touching a file, not a record of tradeoffs made mid-implementation. It also cannot watch its own code run — the same blind spot that keeps it from [seeing what it built](/blog/why-your-agent-cant-see-what-it-built/) visually applies here.

That is why "this reads fine" is weaker evidence than usual with generated code. Reading fine and running correctly are different claims, and generated code collapses the two in your head before you have checked either.

## Where should you start?

Start with data shape: the structures moving through the code, not what the code does to them. Read type definitions, function signatures, schemas and API contracts before any logic. If the shape is wrong — an array where the caller expects an object, a field assumed required that is optional — every function built on it is wrong too.

It is also the fastest part of the review, which is why it goes first:

1. **Data shape** (2–5 minutes) — types, schemas, function signatures, request and response contracts. You are checking that what the code assumes about its data matches reality.
2. **Boundaries** (10–15 minutes) — every place data enters or leaves: inputs, outputs, error paths, external calls. This is where most real bugs live.
3. **Logic** (whatever is left) — the algorithm itself, read last, once you already trust the shape and the edges it operates on.

Reading logic first is tempting because it is the most interesting part of any diff. But logic review only tells you the code does what it appears to do. It cannot tell you whether that matches what the data actually looks like at runtime.

## What deserves the most scrutiny?

Boundaries: anywhere data crosses a trust line, or an error can occur. Input validation, API responses, database writes and catch blocks are where generated code most often looks right and behaves wrong, because the model is guessing at failure behaviour it has no feedback loop for.

A function's happy path is the easiest thing for a model to get right, because it is also the most common pattern in its training data. What is scarce in that data is the unhappy path specific to your system: what your API returns on a 429, what your database does with a duplicate key, what your frontend shows when a fetch times out mid-render. Generated code fills those gaps with something plausible-sounding rather than something correct for your system.

## What can you safely skim?

Skim self-contained pure logic — a sort comparator, a formatting helper, a small math utility — anything with no side effects and a bounded input space you can eyeball in seconds. If a function only transforms values already validated at a boundary, and its tests pass, your attention is worth more elsewhere in the diff.

The test is not "is this function simple," it is "can this function's mistakes hide." A one-line date formatter can only be wrong in ways you notice immediately. A function reshaping an API response before it reaches five other files can be wrong in ways that surface three screens later, in code that looks unrelated to the cause.

## What are the specific AI failure patterns?

Five patterns recur across agents: invented APIs, errors caught and silently discarded, logic duplicated instead of reused, defensive null checks that hide a real bug, and confident comments describing behaviour the code does not have. Each leaves a greppable signature.

| Failure pattern | What it looks like | What to grep for |
|---|---|---|
| Invented APIs | A method, package or config flag absent from the library actually installed | The exact import or method name against `package.json` or `requirements.txt` and the library's real type defs |
| Silently swallowed errors | A `catch` that logs and moves on, or catches nothing | `catch (e) {}`, `catch (e) { console.log`, `except Exception:` with no re-raise, `.catch(() => {})` |
| Duplicated logic | The same validation, parsing or formatting rewritten instead of reused | A distinctive literal — a regex, an error string, a magic number — that should appear once and shows up twice |
| Over-defensive null checks | `?.`, `?? []`, `\|\| {}` stacked on a value that cannot be null if the code above it is correct | Optional-chaining density per file, and any `?.` two lines below a check that already confirmed the value exists |
| Confident wrong comments | A comment claims retries, caching or validation the code below does not implement | Comments containing "retr", "cache", "valid", "sanitiz" — then read the code under each one |

## How often are invented APIs actually a problem?

Often enough to check every import. A peer-reviewed study presented at USENIX Security 2025 tested 16 code-generating LLMs across 576,000 code samples in two languages and found "the average percentage of hallucinated packages is at least 5.2% for commercial models and 21.7% for open-source models, including a staggering 205,474 unique examples of hallucinated package names" ([Spracklen et al., USENIX Security 2025](https://www.usenix.org/conference/usenixsecurity25/presentation/spracklen)). The paper won a Distinguished Paper Award, and it frames the risk as a supply-chain one: an invented name is a name an attacker can register.

Newer models narrow the range without closing it. A 2026 preprint re-running the method on five frontier models over 199,845 paired Python and JavaScript prompts reports hallucination rates between 4.62% and 6.10% ([Churilov, arXiv preprint, revised August 2026](https://arxiv.org/abs/2605.17062)). Treat that as a preprint rather than a peer-reviewed result — but the direction is the useful part: the floor has not reached zero, so verifying that a package exists stays a review step, not a formality.

Duplicated logic is the pattern most likely to slip past a diff review, because a duplicate function compiles, passes its own tests, and looks locally fine. The cost only appears later, when someone fixes a bug in one copy and not the other.

## How deep is deep enough?

Deep enough means every boundary touching money, auth or data you cannot recreate has been traced to a real input or output, not just read. Reading generates a hypothesis about what the code does; running it against a real case confirms or kills that hypothesis. Everything lower-risk can stop at a confident read.

This is the verify step of the [plan → execute → verify loop](/blog/plan-execute-verify-loop/), and it is worth treating as genuinely separate from reading. Anthropic's Claude Code guidance makes a related point about who does the reviewing: it recommends "an adversarial review step" in which a subagent reviews the diff in a fresh context, "so the agent doing the work isn't the one grading it," and warns in the same breath that "a reviewer prompted to find gaps will usually report some, even when the work is sound" ([Claude Code best practices](https://code.claude.com/docs/en/best-practices)). Both halves apply to you reading your own agent's diff right after watching it write the thing.

Use risk, not diff size, to set depth. A 40-line change to a currency-conversion function deserves a full boundary trace even if it looks trivial; a 400-line change to an internal logging helper might not.

## What if the code is right and the output still looks wrong?

Then a code review will not find it, because nothing in the code is wrong. The types check, the boundaries hold, the tests pass, and the rendered result is still off. That is a different problem with a different fix.

Pair the read with the [12-point checklist for verifying AI-built UI](/blog/verifying-ai-built-ui-checklist/) whenever the change has a visual surface — code review and UI verification are different checks, and generated code regularly passes one while failing the other. Reporting the visual half well is covered in [how to describe a visual bug to an AI coding agent](/blog/describe-a-visual-bug-to-ai/) and [how to give feedback to Claude Code so it lands](/blog/give-feedback-to-claude-code/).

Sending an image back is not free either. One 1920×1080 screenshot costs 2,691 input tokens on the high-resolution tier — Claude 4.7 and later, capped at 4,784 — by the patch formula in [Anthropic's vision documentation](https://platform.claude.com/docs/en/build-with-claude/vision), and it carries no reason attached. Price your own before sending in the [screenshot token calculator](/calculator/), or read why pairing pixels with narration beats either alone in [what visual feedback for AI coding agents is](/blog/what-is-visual-feedback-for-ai-coding-agents/).

## What to do next

Open type definitions and function signatures first. Trace every boundary — inputs, outputs, catch blocks — against one real case before trusting it. Verify that every imported package actually exists. Read logic last, and treat that pass as confirmation rather than discovery.

---

Read the HTML version: https://usewalkie.com/blog/reviewing-code-you-didnt-write/
