# The vibe coding workflow: a working setup for solo builders

> Vibe coding is a four-step loop: describe what you want, let the agent generate it, look at what it built, and react. Describe and generate are heavily tooled; look and react are not. One pasted 1080p screenshot costs 2,691 tokens on the high-resolution tier before you explain anything.

Author: Roberto Ercole · Published: 2026-08-05 · Updated: 2026-08-16 · Canonical URL: https://usewalkie.com/blog/vibe-coding-workflow/

---

Vibe coding is a four-step loop: describe what you want, let the agent generate it, look at what it actually built, and react with feedback, repeated dozens of times a session. The setup that works isn't about a faster model. It's about closing the last two steps, look and react, which carry almost none of the tooling the first two enjoy.

Updated August 2026: that loop hasn't changed shape since vibe coding got its name, but the tooling gap inside it has become the whole story. Everyone building agents raced to make describe and generate better: bigger context windows, faster models, sharper planning. Almost nobody built for the moment you stare at a running app and try to explain, precisely, what's wrong with it — the subject of the pillar on [visual feedback for AI coding agents](/blog/what-is-visual-feedback-for-ai-coding-agents/).

## What is vibe coding, really?

Building software by describing outcomes in plain language and letting an AI agent write the code, rather than typing every line yourself. It is a workflow, not a tool: you stay in charge of intent and judgment, the agent handles syntax and boilerplate, and the loop repeats until the result matches what you pictured.

In practice it means running Claude Code, Cursor, Codex or a similar agent as the primary way you write software, and treating your own typing as the exception. You describe a fix in a sentence or two; the agent reads your codebase, writes the change, and often runs it; you look at what came back and either accept it or say what's still wrong.

That last part is the whole trick, and it's easy to miss because the first half of the loop is so fast. Speed doesn't remove the two judgment-heavy steps — it makes them happen more often, back to back. For a solo builder that matters more than on a team: you're both the person who asked for the change and the only one checking it landed, and it's easy to let the second role slide when the first one feels finished.

## What does the loop look like?

The loop has four steps that repeat every few minutes during a vibe-coding session: describe what you want, generate the code, look at what actually happened, and react with the next instruction. Each pass is short, but a two-hour session runs the loop dozens of times, so small friction in any one step compounds fast.

Broken into steps:

1. **Describe:** you write a sentence or two of intent, sometimes pointing at a file for context.
2. **Generate:** the agent plans, edits files, runs commands, and sometimes tests its own work.
3. **Look:** you check the actual result: a page in a browser, a screen in a simulator, output in a terminal.
4. **React:** you tell the agent what's right, what's wrong, and what to do next.

Only generate happens without you. Look is entirely yours, every pass, because nothing else in the loop has eyes.

## Where does the time actually go?

Into look and react. Describe and generate are the tooled half: editors, agents and context management all target them. There is no dominant tool for capturing what's on screen or turning a reaction into something an agent can act on, so builders do both by hand, every single pass, and it shows up as time.

Here's the asymmetry in numbers instead of feel. Describing a change costs whatever your sentence costs, typically a few dozen tokens. Generation's cost is real but tooled. Look is the outlier: showing an agent a single 1080p screenshot costs 2,691 tokens on the high-resolution tier — Claude 4.7 and later, which caps any single image at 4,784 visual tokens — before you've said a word about what's wrong with it. The formula, `ceil(width / 28) x ceil(height / 28)`, and both tier rows are published in [Anthropic's vision documentation](https://platform.claude.com/docs/en/build-with-claude/vision); standard-tier models downscale the same screenshot to 1456x819 and pay 1,560. Frame-dump the same 30 seconds at one frame per second and the bill is 80,730 tokens. A narrated walkthrough of the identical screen measures about 150 tokens.

| Metric | Tokens |
|---|---|
| One 1080p screenshot, high-resolution tier, 4,784 cap | 2,691 (exact) |
| The same screenshot on the standard tier | 1,560 (calculated) |
| 30 seconds frame-dumped at 1 fps | 80,730 (calculated) |
| A spoken walkthrough of the same 30 seconds | ~150 (measured, typical) |

Try your own screen dimensions in the [token calculator](/calculator/).

None of that is really about tokens. Tokens are the measurable proxy for the actual cost, which is attention. Somebody still has to look at the running app, notice what's wrong, and put it into words the agent can use — a human step no amount of context-window growth removes.

The gap isn't an oversight. Describe and generate are problems every agent vendor had to solve to ship at all, so they got years of investment by default; look and react were nobody's core problem until agentic coding got fast enough to make them the bottleneck. The tooling is thin because the need is new.

## What tools does each stage need?

Describe and generate are served by the coding agent itself and by MCP servers that give it more to work with. Look needs a cheap way to capture and show the screen. React needs a structured way to turn what you saw into something the agent can act on, not a video link, not a paragraph of prose.

**Describe and generate.** This half runs on the agent itself: Claude Code, Cursor, Codex, Copilot, plus whatever it can reach. The reach matters. [MCP](/blog/what-is-mcp/) — described on its own site as [an open-source standard for connecting AI applications to external systems](https://modelcontextprotocol.io/) — is the plumbing under most of what makes 2026's agents more capable than 2024's. If your agent can only read code and never touch anything else, describe and generate are still where you'd start improving it.

**Look.** The stage without a default. Most people paste a screenshot, which costs 2,691 tokens on the high-resolution tier against its 4,784 cap, and still needs a caption to mean anything. A smaller group tries [screen recording built specifically for an agent](/blog/screen-recording-for-claude-code/): raw video, frame extraction, or narrated bundles. Reaching for Loom is a reasonable instinct, since it already sits in most toolbars, but [Loom's AI reads the transcript, not the pixels](/blog/loom-video-to-ai-coding-agent/), so the visual half of the review never reaches the agent.

**React.** Once you've looked, you have to say something. A wall of prose works but costs tokens and attention on both sides. A growing number of tools compress a review into a fixed, [agent-readable format](/blog/what-is-review-md-bundle/) an LLM scans in one pass. Before you pick one, note that [feedback tools for Claude Code, Cursor and Codex range from free browser extensions to buy-once desktop apps](/blog/agent-feedback-tools-compared/), and the right one depends on whether you're reviewing a browser tab, a native app, or a whole desktop.

| Loop stage | Default today | Typical cost | Dedicated tooling |
|---|---|---|---|
| Describe | Typed prompt | Tens of tokens | Mature, built into every agent's UI |
| Generate | Agent + MCP tools | Varies, tooled | Mature: Claude Code, Cursor, Codex |
| Look | Pasted screenshot | 2,691 tokens per image, high-resolution tier, 4,784 cap | Thin: a handful of 2026 entrants |
| React | Typed follow-up message | Tens to hundreds of tokens | Thin: structured formats are new |

Every tool in the look and react rows is new and single-vendor; none has the adoption record Claude Code or Cursor already have. Judge each on what it costs and what it does, not on reputation, because there isn't one yet — Walkie included.

## How do you keep quality up as speed rises?

Quality holds up when the react step gets more precise as the loop speeds up, not less: name the exact location of a problem, say what you expected as well as what's wrong, and check the fix in the same place you found the bug rather than trusting the agent's description of what it changed.

Speed is the appeal of vibe coding, and also the thing that erodes review discipline first. When a fix takes eight seconds instead of eight minutes, it's tempting to skim the result and accept "done" at face value. Both habits compound by pass fifty.

- Be precise: location, expectation, and the gap between them, not a vague "this still looks off"
- Keep the look step honest on the fifth pass of the hour: render the page, run the flow, rather than reading the agent's summary

An agent's account of its own work is text about a claim it can't verify by looking. You're the only part of the loop that can.

## What breaks at scale?

Two things: context, because each new file adds to what the agent re-reads every turn, and review bandwidth, because more surface area means more to look at per pass while your session time stays fixed. Both show up as the same symptom — slower turns that feel like the agent got worse.

On a small project, one review pass covers most of what changed. Past a few dozen files, it doesn't. [Anthropic's guidance on Claude Code costs](https://code.claude.com/docs/en/costs) states that Claude Code sends your full conversation with every request, so as CLAUDE.md files and file counts grow, more of every turn's budget goes to re-reading state. Image-heavy review makes that worse fastest, since a handful of screenshots per pass can outweigh the code itself in tokens.

The fix isn't slowing down generation. It's narrower reviews scoped to what actually changed: a layout tweak in a shared component might touch five pages; a fix to one form's validation almost never does.

## What does a good setup look like end to end?

A capable coding agent for describe and generate, plus a deliberate choice for look and react: a fast capture method, a compact format the agent reads cheaply, and feedback specific enough to act on first try. No single tool covers all four steps yet; most builders assemble two or three.

Stage by stage:

1. **Describe and generate:** a coding agent you already trust, Claude Code, Cursor or Codex, with an MCP server or two connected for anything outside the codebase itself.
2. **Look:** a capture method sized to the bug. A single screenshot for something contained to one screen. A short narrated recording when the problem is a sequence, a state change, or something easier to point at than to type.
3. **React:** feedback in a format the agent can scan in one pass rather than a paragraph it has to parse: a location, what's wrong, what you expected.
4. **Verify:** the same look step, run again, against the actual output, not the agent's account of what it changed.

None of these choices is fixed, and the honest state of the market in mid-2026 is that no single vendor has solved all four steps. For the fuller list of what's available at each stage, with prices re-read on the vendor pages on 16 August 2026, see [the current stack of vibe coding tools](/blog/vibe-coding-tools-2026/). Walkie is [one entry in the look-and-react half](/), a screen-and-voice recorder that turns a review into a bundle a coding agent reads for about 3,500 tokens instead of a video it can't open. It's a recent, single-vendor answer, not the only one.

## Which step should you fix first?

The one that stalled. Pick a single loop pass from this week that took longer than it should have and work out where it actually stuck: describe, generate, look or react.

If it stalled on describe or generate, that's a prompting or configuration problem, and plenty is written about it. If it stalled on look or react — rewriting the same bug report three times, pasting five screenshots to cover one page — that's the half this article is about. Write a rough structured review by hand for a week (location, issue, expectation, one screenshot) before deciding whether a dedicated tool is worth paying for.

---

Read the HTML version: https://usewalkie.com/blog/vibe-coding-workflow/
