Context window limit: 200K, 500K or 1M

The context limit sets how much conversation a thread may carry before the gateway moves its oldest parts to the archive. You set it once, in Settings, and it applies to Claude Code and Codex on every machine.

Key facts

What it is
How much conversation a thread may carry before the gateway moves its oldest parts to the archive: 200K, 500K or 1M tokens.
Who it is for
Anyone with long Claude Code or Codex sessions.
Clients
Claude Code and Codex, set once per account in Settings.
Measured result
Not a cost feature, and not measured on its own. For the gateway as a whole, see Benchmarks.
Limits
An image older than the kept tail is replaced by a short note and cannot be read back. The model's own window is a separate, hard limit.

Where to set it

Open Settings in the console and find the Context limit card. It has three choices:

200KTrims often. Costs more.
500KTrims sometimes.
1MRecommended. Trims rarely, and the cache stays warm.

The limit belongs to your account, not to one machine or one client. A new value reaches every request within 30 seconds.

An account that has never picked a limit runs at 800K, and the card shows "now 800K" with no box selected. We recommend 1M.

How it works near the limit

The gateway measures each request before it goes to the model. It counts size from bytes, at about 3 bytes per token. When a thread grows, it acts in two steps:

  1. At 89% of the limit, old tool output and long answers from the model (over 1,500 characters) move to the archive. A short pointer takes their place. The last 3 messages stay as they are.
  2. At the limit, the earliest turns move to the archive in fixed steps, until the thread fits. At a 1M limit, each step is 150,000 tokens. The last 8 messages always stay. Your first message stays too if it is under 64 KB, because that is where a rule for the whole session usually sits.

Three things never change:

The steps are fixed sizes on purpose. Between two steps, the start of the request stays the same, so the provider keeps reading it from its cache.

The model's own window

The limit is your choice. The model's window is a hard wall set by the provider. The gateway keeps each thread inside that wall, whatever limit you pick, so that the client does not have to compact the conversation itself:

Client and windowWhat the gateway does
Claude Code, on requests that ask for the 1M windowOnce a thread passes 600,000 tokens, the gateway plans a cut and keeps it at or under 850,000 tokens. Claude Code's own compaction starts higher, near 967,000 tokens by our reading of Claude Code 2.1.x (Anthropic does not publish the number), so it does not start.
CodexCodex compacts at 90% of the model's window: 244,800 tokens for most models (read from codex-rs 0.156.1). The gateway keeps the thread at or under 208,080 tokens, so Codex does not compact, unless you lower model_context_window in Codex's config.
Any other request, for example Claude Code on a 200K windowIf the provider refuses a request as too long, the gateway cuts it to fit, sends it again, and keeps that cut for the next requests. We do not promise that the client never compacts here.

A limit above the model's window does no harm: the window rule keeps the thread inside it. A limit below the window trims earlier.

Claude Code asks for the 1M window with Anthropic's context-1m beta header. The gateway reads that header to pick the first row.

How to choose

A lower limit does not lower your bill. It raises it. Each trim changes the start of the request, and the provider then bills the whole start again at the cache write price, not the far lower read price.

We measured this live on 2026-07-29, with one pair of runs: 14 turns each, the same model and the same prompts, with only the limit changed. At a 100K limit, an older setting that is not on today's menu, the thread used 282,089 effective input tokens. At 1M it used 145,364. The 100K thread carried only 1.2% less context. Effective input tokens weigh cache reads and writes by their price, so they follow the bill. The step logic has changed since that run, so read it as the direction, not the size.

LimitTrimsPick it when
1MrarelyYou want the full history and the fewest trims, so the fewest cache rewrites. This is the right choice for most teams.
500KsometimesYou want old turns out of the model's view sooner and accept a higher input cost.
200KoftenYou want the smallest thread the model reads. Expect the most trims, so the most cache rewrites.

We measured no speed gain from a smaller limit.

What you get

Compared with other approaches

ApproachWhere it runsWhat happens near the windowWhat happens to the old partClients
Claude Code auto-compactionThe client, on each machineClaude Code summarizes the conversationThe local session file keeps the full transcript; the client gives the model no search over itClaude Code
Codex automatic compactionThe client, on each machineCodex compacts the historyKeeping dropped turns is not documentedCodex
Anthropic context managementThe Claude API, set by the app that calls itServer-side compaction, or clearing old tool results and thinkingRemoved from the request. Anthropic's docs say clearing tool results costs cache writes.Apps built on the Claude API
Context Mode context limitThe gateway, one setting per accountOld tool output moves to the archive first, then the earliest turns, in fixed stepsText kept in the archive, searchable, and readable byte for byte except masked secrets and e-mail addresses; old images are droppedClaude Code and Codex

Where they are stronger. Built-in compaction is free, first-party and needs no third party. A summary keeps the gist of old turns in the model's view. With Context Mode, the agent has to search the archive to see old turns again. Anthropic's context management works in any app you build on the API.

Where Context Mode differs. One setting covers Claude Code and Codex, on every machine. Text that leaves the thread stays searchable by the agent. On Codex, and on Claude Code requests for the 1M window, the gateway cuts before the client's own compaction line, so a summary does not replace the conversation.

FAQ

Which limit should I choose?

We recommend 1M. It trims rarely, and the cache stays warm.

Is this the model's own context window?

No. The model's window is a hard wall set by the provider. The gateway keeps each thread inside it, whatever limit you pick.

Is text lost when a thread is cut?

No. Text that leaves the thread goes to the archive, and the agent can read it back with Search, except secrets and e-mail addresses, which are masked before storage. Images are the exception.

Compared with other tools: see the landscape.