<!-- Published copy of a run plan from the Context Mode gateway repository. -->

Pre-registration: Long session on a 200K window, run 2

- Source file: docs/verify/recall-ab-200k-r2/PREREG.md
- Commit that added it: 34acbe8b, 2026-10-03T23:54:04Z
- Evidence: /benchmarks/recall/200k-r2/ab.json (its "prereg_check" names the same commit and the first pair's start time).
- This copy is the file as it stood at that commit.
- Edits: private fields only (tenant ids, local paths, the internal Worker name and internal endpoint names),
  each marked [edited: ...]. The workload, the pass rule and the analysis are as committed.

---

# Recall A/B on a real 200K window, run 2: pre-registration (written and pushed before any run)

Run 1 (`../recall-ab-200k/`) found: the gateway arm was 52.6 % cheaper and answered 10 of 10 against plain's 4 of 10,
but **Recall never re-laid out**: the gateway's largest forward was 134,291 tokens, under Recall's line
L = 160,000. So run 1 did not test Recall. Its rule said: to test Recall on a 200K client, the gateway's forward must
pass L. This run does that in the one way a user can: the gateway tenant's **Context limit** (the Settings card,
[edited: the Context limit setting]) is set to **100,000**. Recall's L is never above the tenant's limit
(`cmRecallLimitFor`, `src/gateway/recall-plan.ts`), so L = 100,000 here. Nothing else changes. Run 1's files and
result stay as they are.

## 1. The question

Same as run 1: on a 200,000-token Claude Code window, does the gateway **with Recall acting** cost less **and**
remember at least as well as plain? Reported for every arm and pair: **savings**, **quality** and **performance**.

## 2. Arms

- **plain**: unchanged from run 1 (`../recall-ab-200k/PREREG.md` §2): Claude Code 2.1.283, `claude-opus-5-5`,
  `CLAUDE_CODE_DISABLE_1M_CONTEXT=1`, straight to Anthropic.
- **gateway**: run 1's gateway arm, plus **Context limit 100,000** on its fresh disposable tenant, written through
  [edited: the Context limit setting] and read back from `/__wm-config` before any spend (`--context-limit 100000`; the harness
  refuses if the read-back differs). Same levers, same [edited: an internal switch] stages for the A/B tenants (`CM_RECALL_QUALITY`,
  `CM_RECALL_AUTO`, `CM_RECALL_THREAD`, `CM_RECALL_CONFIRM`, `CM_RECALL_CHAIN_MEMO`, `CM_RECALL_YIELD`,
  `CM_DO_HINT_FRA`).

Tenants: [edited: a disposable test tenant per pair] for pair k. None is an owner or customer tenant.

## 3. Workload

Run 1's workload, unchanged: `naturalPlan`, 30 steps, the 16 ref files read from a checkout of run 1's head
`a653f040` (`--ref-repo`; one of them, `src/memory/fact-routes.ts`, grew by 17 lines on main since), **seed 616161** (the same seed as run 1, so run 1's gateway
arm and this one differ only in the Context limit). Question turns: plain no tool; gateway `context-mode-search` only.

## 4. What is measured

As run 1 §4: savings (list price, client result usage, gateway arm the higher of that and the ledger; saved $ and
%; tokens per call; Recall's net from the ledger), quality (exact of 10, contains, 8 task checks, re-asks,
compactions, 4xx, re-layouts by trigger, archive searches), performance (client wall time and time to first token
p50/p90/p99, task time; gateway `pre_ms`, Recall thread read, auto-recall lookup).

## 5. A pair is valid when

Run 1 §5 unchanged (`ab200kPairValid`): both arms ran all 30 steps, deploy and the A/B tenants' [edited: an internal switch] stages the
same before the pair and after every step, tool lists constant, seed recorded, both windows 200,000, plain compacted
at least once before the first question, **the gateway arm re-laid out at least once**.

## 6. Ship rule

For each of the **3** pairs: gateway $ < plain $ (list price) **and** gateway exact >= plain exact. Ship = all 3
pairs present, valid and passing. Anything else is FAIL. (Three pairs, not four: the night's budget is $60 and run 1
used $31.33.)

## 7. Freeze, order, budget

- The deploy lock is taken before this file is pushed and held until the results are pushed. This file and the
  harness change go to main in one push; the run starts after its build and smoke pass.
- No dry run: the only changes are a product setting, checked by read-back before any spend, and the ref files'
  source, recorded as `ref_head` in every arm file.
- The 3 pairs run in parallel, about 60 s apart; pairs 1 and 3 start plain, pair 2 starts gateway; the second arm
  starts 5 s after the first.
- **Hard cap: $28** for this run (`--total 28`, in this folder's `spend.json`), so the night stays under $60 with
  run 1's $31.33. Per arm: plain $6.50, gateway $4.00. A stopped arm makes its pair invalid.
- One attempt. One restart only if the harness itself breaks before any pair completes. No tuning after results.

## 8. Commands

```
CM_CLAUDE=[edited: the local claude binary] node scripts/e2e/recall-ab/natural.mjs --pair <k> --first <plain|recall> --seed 616161 \
  --search-probes --window200k --auto --timing --stamp ab200kr2 --topo-tenants [edited: a disposable test tenant per pair] \
  --context-limit 100000 --ref-repo [edited: a local folder] --cap-plain 6.5 --cap-recall 4 --total 28 --out docs/verify/recall-ab-200k-r2
node scripts/e2e/recall-ab/ab200k-summarize.mjs --dir docs/verify/recall-ab-200k-r2 --pairs 3
```

## 9. Files written

`pair-<k>-plain.json`, `pair-<k>-recall.json`, `pair-<k>.json`, `spend.json`, `ab.json`, `README.md`. Counts,
dollars, times and booleans only.
