Active Context

Choose how much the model sees.

Pick 200K, 500K or 1M once. Each thread stays under it, and older turns move to an archive the agent can search.

Near the limit
In the threadyour first messageearlier turnsold tool outputrecent messages
In the archivenothing yet

Below 89% of your limit, nothing moves.

In the threadyour first messageearlier turnspointer to old outputthe last 3 messages, unchanged
In the archiveold tool outputlong answers

Old tool output and long answers move out. A short pointer stays.

In the threadyour first messagethe last 8 messages
In the archiveold tool outputthe earliest turns, word for word

The earliest turns move out in fixed steps, so the cache stays warm.

How it works

The thread stays in bounds. Nothing is thrown away.

  1. 01Pick a limit

    200K, 500K or 1M, in Settings. It reaches every request within 30 seconds.

  2. 02At 89%

    Old tool output and long answers move to the archive. A pointer stays.

  3. 03At the limit

    The earliest turns move out in fixed steps. The last 8 messages always stay.

Proof

A bigger limit costs less.

Effective input tokens, same 14 turns

100K limit282,089
1M limit145,364

fewer cache rewrites

Provider refusals near the window, 60-turn thread

Direct37
Context Mode1

cut before refusal

One live pair, 2026-07-29; refusals from a simulated Claude Code thread in our test harness. Read them as the direction, not the size. Method →

Features

Long sessions, warm cache.

No surprise compaction

Codex, and Claude Code on the 1M window, stay under the point where they compact.

Your words kept

Your own messages are never shortened. Text that leaves is in the archive.

One setting

Per account, for both agents and every machine.

Keep the whole session in reach.

npx @context-mode/cli connects your agents. 1,000 requests free, no card.