Context window limit: 200K, 500K or 1M
The context limit sets how much conversation a thread may carry before the gateway moves its oldest parts to the archive. You set it once, in Settings, and it applies to Claude Code and Codex on every machine.
Key facts
- What it is
- How much conversation a thread may carry before the gateway moves its oldest parts to the archive: 200K, 500K or 1M tokens.
- Who it is for
- Anyone with long Claude Code or Codex sessions.
- Clients
- Claude Code and Codex, set once per account in Settings.
- Measured result
- Not a cost feature, and not measured on its own. For the gateway as a whole, see Benchmarks.
- Limits
- An image older than the kept tail is replaced by a short note and cannot be read back. The model's own window is a separate, hard limit.
Where to set it
Open Settings in the console and find the Context limit card. It has three choices:
The limit belongs to your account, not to one machine or one client. A new value reaches every request within 30 seconds.
An account that has never picked a limit runs at 800K, and the card shows "now 800K" with no box selected. We recommend 1M.
How it works near the limit
The gateway measures each request before it goes to the model. It counts size from bytes, at about 3 bytes per token. When a thread grows, it acts in two steps:
- At 89% of the limit, old tool output and long answers from the model (over 1,500 characters) move to the archive. A short pointer takes their place. The last 3 messages stay as they are.
- At the limit, the earliest turns move to the archive in fixed steps, until the thread fits. At a 1M limit, each step is 150,000 tokens. The last 8 messages always stay. Your first message stays too if it is under 64 KB, because that is where a rule for the whole session usually sits.
Three things never change:
- The text of your own messages is never shortened. Old turns may leave the thread as a whole, and their text goes to the archive.
- Images are the exception. An image older than the kept tail is replaced by a short note, and it cannot be read back.
- A message that carries the model's extended thinking is never edited, because the provider checks it on the next turn.
- Text is not thrown away. Every text part that leaves the thread is in the archive, and the agent can read it back with Search, byte for byte except secrets and e-mail addresses, which are masked before storage.
The steps are fixed sizes on purpose. Between two steps, the start of the request stays the same, so the provider keeps reading it from its cache.
The model's own window
The limit is your choice. The model's window is a hard wall set by the provider. The gateway keeps each thread inside that wall, whatever limit you pick, so that the client does not have to compact the conversation itself:
| Client and window | What the gateway does |
|---|---|
| Claude Code, on requests that ask for the 1M window | Once a thread passes 600,000 tokens, the gateway plans a cut and keeps it at or under 850,000 tokens. Claude Code's own compaction starts higher, near 967,000 tokens by our reading of Claude Code 2.1.x (Anthropic does not publish the number), so it does not start. |
| Codex | Codex compacts at 90% of the model's window: 244,800 tokens for most models (read from codex-rs 0.156.1). The gateway keeps the thread at or under 208,080 tokens, so Codex does not compact, unless you lower model_context_window in Codex's config. |
| Any other request, for example Claude Code on a 200K window | If the provider refuses a request as too long, the gateway cuts it to fit, sends it again, and keeps that cut for the next requests. We do not promise that the client never compacts here. |
A limit above the model's window does no harm: the window rule keeps the thread inside it. A limit below the window trims earlier.
Claude Code asks for the 1M window with Anthropic's context-1m beta header. The gateway reads that header to pick the first row.
How to choose
A lower limit does not lower your bill. It raises it. Each trim changes the start of the request, and the provider then bills the whole start again at the cache write price, not the far lower read price.
We measured this live on 2026-07-29, with one pair of runs: 14 turns each, the same model and the same prompts, with only the limit changed. At a 100K limit, an older setting that is not on today's menu, the thread used 282,089 effective input tokens. At 1M it used 145,364. The 100K thread carried only 1.2% less context. Effective input tokens weigh cache reads and writes by their price, so they follow the bill. The step logic has changed since that run, so read it as the direction, not the size.
| Limit | Trims | Pick it when |
|---|---|---|
| 1M | rarely | You want the full history and the fewest trims, so the fewest cache rewrites. This is the right choice for most teams. |
| 500K | sometimes | You want old turns out of the model's view sooner and accept a higher input cost. |
| 200K | often | You want the smallest thread the model reads. Expect the most trims, so the most cache rewrites. |
We measured no speed gain from a smaller limit.
What you get
- The full history in the thread. At 1M, the model keeps more of the session in view, and trims are rare.
- A warm cache. Fewer trims mean fewer full rewrites of the request's start.
- Text is kept. Text that leaves the thread stays in the archive and can be read back. Old images are not kept.
- Fewer refusals at the window. The gateway cuts a thread before the provider can refuse it. In a simulated 60-turn Claude Code thread in our test harness, near the window, provider refusals fell from 37 to 1. In a simulated 52-turn Codex thread, they fell from 28 to 2. Each refusal costs a full extra round trip.
Compared with other approaches
| Approach | Where it runs | What happens near the window | What happens to the old part | Clients |
|---|---|---|---|---|
| Claude Code auto-compaction | The client, on each machine | Claude Code summarizes the conversation | The local session file keeps the full transcript; the client gives the model no search over it | Claude Code |
| Codex automatic compaction | The client, on each machine | Codex compacts the history | Keeping dropped turns is not documented | Codex |
| Anthropic context management | The Claude API, set by the app that calls it | Server-side compaction, or clearing old tool results and thinking | Removed from the request. Anthropic's docs say clearing tool results costs cache writes. | Apps built on the Claude API |
| Context Mode context limit | The gateway, one setting per account | Old tool output moves to the archive first, then the earliest turns, in fixed steps | Text kept in the archive, searchable, and readable byte for byte except masked secrets and e-mail addresses; old images are dropped | Claude Code and Codex |
Where they are stronger. Built-in compaction is free, first-party and needs no third party. A summary keeps the gist of old turns in the model's view. With Context Mode, the agent has to search the archive to see old turns again. Anthropic's context management works in any app you build on the API.
Where Context Mode differs. One setting covers Claude Code and Codex, on every machine. Text that leaves the thread stays searchable by the agent. On Codex, and on Claude Code requests for the 1M window, the gateway cuts before the client's own compaction line, so a summary does not replace the conversation.
FAQ
Which limit should I choose?
We recommend 1M. It trims rarely, and the cache stays warm.
Is this the model's own context window?
No. The model's window is a hard wall set by the provider. The gateway keeps each thread inside it, whatever limit you pick.
Is text lost when a thread is cut?
No. Text that leaves the thread goes to the archive, and the agent can read it back with Search, except secrets and e-mail addresses, which are masked before storage. Images are the exception.
Compared with other tools: see the landscape.