<!-- Published copy of a run plan from the Context Mode gateway repository. -->

Pre-registration: Long session on a 200K window, run 1

- Source file: docs/verify/recall-ab-200k/PREREG.md
- Commit that added it: 1c07881f, 2026-10-03T22:34:16Z
- Evidence: /benchmarks/recall/200k/ab.json (its "prereg_check" names the same commit and the first pair's start time).
- This copy is the file as it stood at that commit.
- Edits: private fields only (tenant ids, local paths, the internal Worker name and internal endpoint names),
  each marked [edited: ...]. The workload, the pass rule and the analysis are as committed.

---

# Recall A/B on a real 200K window: pre-registration (written and pushed before any run)

Earlier runs: `../recall-ab-2026-10/amendment4/` (1M window, 0 compactions, ship rule FAIL on pair 4). Design source:
`docs/prd/recall-quality-2026-10-03.md` §9. This file is the rule for this run. Nothing in it changes after a result.

## 1. The question

When Claude Code runs on a 200,000-token window and compacts by itself, does the gateway with Recall cost less **and**
remember at least as well? We report three things for every arm and every pair: **savings**, **quality** and
**performance**.

## 2. Arms

- **plain**: Claude Code 2.1.283 (`[edited: the local claude binary]`) straight to Anthropic, the owner's login, model
  `claude-opus-5-5` with `CLAUDE_CODE_DISABLE_1M_CONTEXT=1`. Checked before this file was written: the client's own
  result says `contextWindow: 200000` for that model with that setting. No `CLAUDE_CODE_AUTO_COMPACT_WINDOW`. It
  compacts when it decides to.
- **gateway**: the same client, same setting, through the deployed gateway [edited: the gateway Worker] on a fresh disposable
  tenant per pair. Flags memory on, skills off. Levers (admin key): `x-cm-lever-relayout: on`,
  `x-cm-lever-recall: search,decisions,episodes,auto`. Runtime switch [edited: an internal switch]: each A/B tenant is added to
  `CM_RECALL_QUALITY` (pay rule, v2 index, cut pins, **window rule**), `CM_RECALL_AUTO` (auto-recall), and to the
  stages the owner tenants run Recall with (`CM_RECALL_THREAD`, `CM_RECALL_CONFIRM`, `CM_RECALL_CHAIN_MEMO`,
  `CM_RECALL_YIELD`, `CM_DO_HINT_FRA`). The window rule gives L = 0.8 × 200,000 = 160,000 tokens for a client with no
  context-1m beta. The tenant's default Context limit is not changed.
- No gateway-without-Recall arm (dropped, as the PRD allows, to stay in budget).

Tenant names (fixed now so [edited: an internal switch] can name them before the run): [edited: a disposable test tenant per pair] for pair k,
[edited: a disposable test tenant] for the dry run. None is an owner or customer tenant.

## 3. Workload

`naturalPlan` (`scripts/e2e/recall-ab/natural-lib.mjs`), 30 steps, unchanged, with a **new seed 616161** (424242 was
the old A/B; 424242 and 515151 were used in the offline replay of the quality work). The 3 constant questions are the
same as before (they are fixed in the code, not drawn from the seed); the 4 typed values and the 3 report values are
new.

Scaling: none needed. In amendment 4 the same steps sent 228,514 tokens at step r2 and 440,411 at the end (pair 1
plain), so the history passes 200,000 at r2 and reaches 2.2 × the window. Plain must compact at least once before the
first question; if it does not, the pair is invalid (§5).

Question turns: plain allows no tool (hook refuses all, as before). Gateway allows one tool, `context-mode-search`
(amendment 4 sentence and hook, `--search-probes`). The tool list stays the same on every turn of an arm.

## 4. What is measured

**Savings.** Every turn priced at list price (`src/audit/model-pricing.ts`, 5m and 1h writes apart) from the client's
own result usage; the gateway arm is the higher of that and the gateway ledger. Per pair: plain $, gateway $, saved $
and saved %. Also tokens sent per call (mean, max) and Recall's signed net from the ledger.

**Quality.** Exact answers to the 10 memory questions (`probeExact`, banner line set aside), the "contains" checker
(reported only), the 8 task checks, re-asks of the 2 decisions, compactions (all, and before the first question),
upstream 4xx, re-layouts by trigger, archive searches on the question turns.

**Performance.** Client side, from the client's own stream with `--include-partial-messages`, every line stamped on
arrival: per request, wall time (start to `message_stop`) and time to first token (start to the first
`content_block_delta`), p50 / p90 / p99. A request starts at the last stream line before its `message_start`. Also
total task time (sum of all turn times) and per-turn time. Gateway side, from its ops rows (one per client request):
`pre_ms`, `first_call_ms`, `total_ms`, Recall's thread read (`recall_rt_read_ms`, else `recall_read_ms`) and the
auto-recall lookup (`recall_auto_ms`) p50 / p90 / p99, with auto-recall fires and fails. Performance is reported, not
part of the ship rule.

## 5. A pair is valid when

Both arms ran all 30 steps (no cap stop, no crash, not void); the deployed gateway (main head, "Workers Builds:
[edited: the gateway Worker]" build id, Worker version) **and** the A/B tenants' [edited: an internal switch] stages and kill were the same before the
pair and after every step; each arm's tool list stayed the same; the seed is in every file; both clients ran with a
200,000-token window (`modelUsage.contextWindow`); plain compacted at least once before the first question; the
gateway arm re-laid out at least once. Code: `ab200kPairValid` in `scripts/e2e/recall-ab/ab200k-lib.mjs`, tests in
`ab200k-lib.test.mjs`.

## 6. Ship rule

For each of the 4 pairs: **gateway $ < plain $** at list price **and** **gateway exact answers >= plain exact
answers**. Ship = all 4 pairs present, valid and passing. Anything else is FAIL. Task checks, re-asks, 4xx and
performance are reported next to it. Code: `ab200kPairPasses`, `ab200kShip`.

## 7. Freeze, order, budget

- The deploy lock (`[edited: the deploy lock file]`) is taken before this file is pushed and held until the results are pushed.
  This file and the harness change go to main in one push; the run starts after its build and smoke pass.
- Dry run first (`--plan dry4`, arms `plain` and `recall-b4k` = the gateway levers with `;b=4000`, at most $3): it must
  show a 200,000 window in both arms, client timing rows, gateway timing from ops rows, and a real `recall_relayout`
  row. If it fails on the harness, the harness is fixed and the fix is recorded here as an amendment before pair 1.
- Then the 4 pairs **in parallel**, started about 60 s apart; pair 1 and 3 start plain, pair 2 and 4 start gateway;
  the second arm starts 5 s after the first.
- Budget ~$53. **Hard cap $60** for everything (dry run and pairs), summed in `spend.json` and enforced by
  `--total 60` (an arm stops before a step when the total is within $1 of it). Per arm: plain $8.50, gateway $6.00 (an
  arm stops before a step when within $0.60 of its cap). Amendment 4's arms were $7.35 and $3.17 at the most on a 1M
  window; a 200K window bounds plain's reads, so both caps leave room. A stopped arm makes its pair invalid.
- One attempt. One restart only if the harness itself breaks before any pair completes. No tuning after results.

## 8. Commands

```
CM_CLAUDE=[edited: the local claude binary] node scripts/e2e/recall-ab/natural.mjs --plan dry4 --search-probes --window200k --auto --timing \
  --stamp ab200k1004 --topo-tenants [edited: a disposable test tenant] --seed 616161 \
  --arms plain,recall-b4k --cap-plain 2 --cap-recall 1.5 --total 60 --out docs/verify/recall-ab-200k
CM_CLAUDE=[edited: the local claude binary] node scripts/e2e/recall-ab/natural.mjs --pair <k> --first <plain|recall> --seed 616161 \
  --search-probes --window200k --auto --timing --stamp ab200k1004 --topo-tenants [edited: a disposable test tenant per pair] \
  --cap-plain 8.5 --cap-recall 6 --total 60 --out docs/verify/recall-ab-200k
node scripts/e2e/recall-ab/ab200k-summarize.mjs --dir docs/verify/recall-ab-200k --pairs 4
```

## 9. Files written

`pair-<k>-plain.json`, `pair-<k>-recall.json` (the gateway arm), `pair-<k>.json` (freeze record), `dry4*.json`,
`spend.json`, `ab.json` (all three measures, per arm and per pair, the validity and ship rule applied), `README.md`.
Counts, dollars, times and booleans only: no prompt, answer, planted value or key.
