A context window is the maximum amount of text a model can hold and reference in one request — the system prompt, the conversation so far, tool definitions, tool results, images, and the reply it is about to generate. Filling it is not only a room problem: recall and precision measurably decline before the window is full, a pattern Anthropic calls context rot.
Updated August 2026. Visual review is one of the fastest ways to fill a window without noticing. A 30-second screen recording frame-dumped at one frame per second costs roughly 80,730 tokens — close to 40% of a 200,000-token window gone before the agent starts fixing what you showed it, calculated from Anthropic's published patch formula.
What is a context window?#
A context window is the working memory of one request. Anthropic's context-window documentation defines it as all the text a model can reference when generating a response, including the response itself, and describes it as working memory rather than the corpus the model was trained on.
The distinction matters. Training data shapes what a model knows in general; the context window is what it can see right now, in this conversation, and it starts empty every time a new one begins. This is one piece of the wider picture in Claude Code token costs: every dollar spent in a session traces back to what is sitting in that window at the time.
What goes into it?#
Everything in the request. Anthropic states that the system prompt, every message in the conversation including tool results, images and documents, and your tool definitions all count — as does the output the model generates for the turn, including its extended thinking. Nothing is free once it is in context.
| What is in the window | Example | Cost |
|---|---|---|
| System prompt and tool definitions | Instructions and available tools, loaded before you type | Loaded automatically, not controlled per turn |
| One 1080p screenshot | A single full-screen image pasted into chat | 2,691 tokens on Claude 4.7 and later, capped at 4,784, calculated |
| The same image on a standard-tier model | Downscaled to 1456x819 before pricing | 1,560 tokens, calculated |
| Four screenshots to cover one page | Static images, zero narration | 10,764 tokens, calculated |
| A 30-second screen recording, frame-dumped at 1 fps | Raw frames, no narration | 80,730 tokens, calculated |
| A narrated review of the same 30 seconds | Transcript plus selected frames | ~3,500 tokens, measured typical |
The image rows come from Anthropic's patch rule: one visual token per 28x28-pixel block, capped at 4,784 on the high-resolution tier (vision documentation). Every row draws from the same budget as your actual instructions. See how many tokens a screenshot costs an AI agent for the full worked breakdown, or price a capture in the Walkie token calculator.
Anthropic also confirms images count toward the budget the model itself tracks: on models with context awareness the API injects the total window and a running remaining-capacity figure into each request, and its documentation notes plainly that image tokens are included in those budgets.
How big is a Claude context window now?#
It depends on the model, and the numbers moved in 2026. Claude Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5 and Sonnet 4.6 carry a 1M-token context window by default on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. Other models, including Sonnet 4.5, hold 200,000 tokens.
Anthropic is explicit that 1M is the default on those models rather than a beta opt-in, and that long-context requests are billed at standard pricing. A single request can also include up to 600 images or PDF pages, or 100 on models with a 200,000-token window.
- 200,000 tokens — the window on Claude Sonnet 4.5 and other non-1M models
- 80,730 tokens — burned by one 30-second frame-dumped review
- 40% — of a 200,000-token window gone before the fix even starts
What happens when it fills?#
Two different things, depending on when you hit the ceiling. If the input alone exceeds the window, the API returns a 400 invalid_request_error reading "prompt is too long" on every model. If generation reaches the limit mid-response on Claude 4.5 models and newer, it stops with stop_reason: "model_context_window_exceeded" rather than truncating silently.
Chat interfaces such as claude.ai handle it differently again, managing the window on a rolling first-in-first-out basis. For long agentic sessions, Anthropic points to server-side compaction, which summarises earlier parts of the conversation on the server so the conversation can continue past the limit; it is in beta for Claude 4.6 and later models. Context editing offers narrower tools, including clearing old tool results.
One thing compaction does not change: cached prefixes still occupy the window. Prompt caching changes what you pay for those tokens, not whether they count.
Is a bigger window the answer?#
Not by itself. Anthropic's own documentation says more context is not automatically better, that accuracy and recall degrade as token count grows — the phenomenon it names context rot — and that this makes curating what is in context as important as how much space is available.
That headroom genuinely helps with large codebases and long documents. It does not repeal the underlying problem. Anthropic's guidance on context engineering treats the window as a budget to be curated rather than filled, which is the same conclusion the measurements below point at.
A wider window changes when you hit the wall. It does not change whether stuffing it with irrelevant tokens hurts you before then.
How does context relate to output quality?#
Two separate effects are documented. Position matters: the 2023 paper Lost in the Middle: How Language Models Use Long Contexts found that performance is often highest when relevant information occurs at the beginning or end of the input context and degrades significantly when a model must retrieve it from the middle, including on models built for long contexts.
Volume matters too, and this evidence is newer and less settled. Chroma's July 2025 report Context Rot: How Increasing Input Tokens Impacts LLM Performance evaluated 18 models across Anthropic, OpenAI, Google and Alibaba families and found that performance grows increasingly unreliable as input length grows, even on simple tasks, and that models do not use their context uniformly. It is an industry research report rather than a peer-reviewed study, so treat the exact numbers as directional — but the direction matches what Anthropic's own documentation already states.
What both point at, cautiously: a short context containing only what is relevant tends to outperform a long one padded with things the model has to sort through.
How do you keep it clean?#
Treat the window as a budget rather than a dumping ground. Trim history you do not need, avoid pasting raw video or frame-dumped screenshots when a description or a single image will do, and use compaction and context editing to summarise or clear old tool results before they accumulate.
- Trim what you do not need. Old tool results and resolved side-conversations do not have to stay in context forever.
- Prefer a short description or one well-chosen screenshot over a raw recording or a frame dump.
- Use compaction and context editing where your tooling supports it, instead of letting history grow unbounded.
- Check before you send. Anthropic's token-counting endpoint returns the exact input-token count for a payload, images included, for free.
For the compounding version of this problem, why a Claude Max plan runs out faster than expected walks through the daily arithmetic, and 11 ways to cut Claude Code token costs ranks the highest-return habits by effort.
Narrated review bundles are one way to keep a window clean for visual feedback specifically — a transcript plus a handful of selected frames instead of a raw recording, which is the approach Walkie packages for Claude Code, measured at about 3,500 tokens against 80,730 for the frame-dumped equivalent. It is not the only way, and it is irrelevant if your context problem is not visual. The measurement method is published in the visual-context token cost research.