Active Context
Choose how much the model sees.
Pick 200K, 500K or 1M once. Each thread stays under it, and older turns move to an archive the agent can search.
Below 89% of your limit, nothing moves.
Old tool output and long answers move out. A short pointer stays.
The earliest turns move out in fixed steps, so the cache stays warm.
The thread stays in bounds. Nothing is thrown away.
- 01Pick a limit
200K, 500K or 1M, in Settings. It reaches every request within 30 seconds.
- 02At 89%
Old tool output and long answers move to the archive. A pointer stays.
- 03At the limit
The earliest turns move out in fixed steps. The last 8 messages always stay.
A bigger limit costs less.
Effective input tokens, same 14 turns
fewer cache rewrites
Provider refusals near the window, 60-turn thread
cut before refusal
One live pair, 2026-07-29; refusals from a simulated Claude Code thread in our test harness. Read them as the direction, not the size. Method →
Long sessions, warm cache.
Codex, and Claude Code on the 1M window, stay under the point where they compact.
Your own messages are never shortened. Text that leaves is in the archive.
Per account, for both agents and every machine.
Keep the whole session in reach.
npx @context-mode/cli connects your agents. 1,000 requests free, no card.