<!-- Published copy of a run plan from the Context Mode gateway repository. -->

Pre-registration: Long agentic coding session, 1M window

- Source file: docs/verify/recall-ab-2026-10/final/PREREG.md
- Commit that added it: f2b5c567, 2026-10-03T04:01:17Z
- Evidence: /benchmarks/long-session/ab.json (its "prereg_check" names the same commit and the first pair's start time).
- This copy is the file as it stood at that commit.
- Edits: private fields only (tenant ids, local paths, the internal Worker name and internal endpoint names),
  each marked [edited: ...]. The workload, the pass rule and the analysis are as committed.

---

# Final Recall A/B: pre-registration (written before the dry run and before any pair)

Earlier runs: `../natural/PREREG.md` (design, checkers, amendment 1) and `../natural/README.md` (no valid pair).
This file is amendment 2 of that design. It changes only what is listed here. Everything else (workload, task checks,
re-ask checker, models, client) stays as in `../natural/PREREG.md` §2-§3.

## 1. The eight lessons, and what the harness does about each

| # | Lesson from the failed attempts | What the harness does now |
|---|---|---|
| 1 | Forcing Claude Code's compaction at 200K made the run unrepresentative | No `CLAUDE_CODE_AUTO_COMPACT_WINDOW`. Both arms use the client's real window. Compactions are counted. |
| 2 | Other pushes changed the gateway mid-pair | Before each pair and after every step of every arm, the harness reads three things fresh: the head of `origin/main`, the [edited: the gateway Worker's build check] build id of that head, and the live Worker version (`wrangler deployments status`, read only). Any change stops the pair, and the pair is invalid. |
| 3 | Probe turns removed tools, so plain re-wrote its whole cache ($3.46 on one turn) | The tool list is the same on every turn. The probe prompt asks for no tool. A `PreToolUse` hook (`probe-gate-hook.py`, in the same settings on every turn) refuses any client tool call while a probe marker file exists. Each turn's tool list is recorded; a change makes the pair invalid. |
| 4 | Attempt 1 did not record the seed | The seed is in every arm file and every pair file. A missing seed makes the pair invalid. |
| 5 | The egress token expired after about 1 h | Before every step, the client's credential is copied from the login when the login's is newer. The recall arm's ledger is read before every step; two failed reads in a row stop the arm. |
| 6 | Client usage of some turns was not priced | Every API call is priced from the client's own stream JSON (one entry per message id). The transcript and the client's own figure are cross-checks. Recall also reads the gateway ledger. |
| 7 | An arm stopped at its cap made $ incomparable | Caps come from the dry run (§5), with 40 % headroom, and both arms of a pair get the same cap. |
| 8 | Recall never acted when the session stayed under its window | The workload passes the Recall window (B = 150,000) by real work. A recall arm with no `recall_relayout` ops row makes its pair invalid. |

## 2. What changed in the harness (amendment 2)

- Probe turns: the tool list is kept; the gate is the hook plus one sentence in the prompt (`PROBE_NO_TOOLS` in
  `natural-lib.mjs`). The old `--disallowedTools` is gone.
- The probe checker for the ship rule is now **exact match** (`probeExact`): the whole answer, with spaces, backticks,
  quotes, asterisks and periods removed from both ends, equals the planted value, case-sensitive. The old "contains"
  checker is reported beside it and decides nothing.
- Pricing per call from the client's stream JSON. Recall arm $ = max(ledger, client JSON). Plain arm $ = client JSON.
- Recall is used as it ships on the owner tenants today: lever `x-cm-lever-relayout: on` (no `;rt`, so the record is
  read from the tenant store, not RecallThread) and `x-cm-lever-recall: search,decisions,episodes`, B = 150,000.
  Failed record reads (`recall_fail:*`) are counted and stay in the recall arm's cost.
- Per-step cost curve per arm (client JSON, and for recall the ledger).

## 3. Pairs and order

4 pairs, run one after another, one process per pair. Pair 1 starts plain, pair 2 starts recall, pair 3 plain,
pair 4 recall; the second arm starts 5 s after the first and both run at the same time. Seed 424242 and the same 30
prompts in every arm. At most 1 attempt. If the harness itself breaks before any pair completes, one restart is
allowed.

## 4. Ship rule (unchanged in substance)

For **each** of the 4 pairs: recall exact probes >= plain exact probes **and** recall $ < plain $ **and** task checks
equal **and** 0 upstream 4xx in the recall arm. Ship = all 4 pairs present, valid and passing. Anything else is FAIL.

A pair is **valid** when both arms ran every step (none stopped or failed), neither is void, the deployed gateway was
the same before the pair and after every step, each arm's tool list was the same on every turn, the seed is recorded,
and the recall arm re-laid out at least once (`natural-lib.mjs` `finalPairValid`).

## 5. Dry run, caps and budget

- Hard cap **$90** for everything in this phase (dry run plus pairs), summed in `final/spend.json`.
- **Dry run** (3 steps: `s01`, `r1`, the first probe; at most $3): arms `plain`, `recall` (the real levers) and
  `recall-b4k` (the same levers with `;b=4000`, only to show that the re-layout check reads a real `recall_relayout`
  row; its numbers are not used for anything else). It must show: no compaction (1), the deploy check ran after every
  step and saw no change (2), the same tool list on all 3 turns and a probe turn that wrote far less than its prompt to
  cache (3), the seed in the files (4), the credential check ran each step and the ledger was read each step (5),
  every turn priced from the client JSON and close to the client's own figure (6), the projected arm cost below (7),
  a `recall_relayout` row seen by the check in `recall-b4k` (8).
- **Cap per arm** (both arms, same number): `cap = min(1.4 × projected plain arm, (90 − dry spend) / 8)`.
  `projected plain arm = 6.3781 × k + 10 × probe`, where 6.3781 is the measured plain cost of the 20 work steps in
  `../natural/pair-1-a2-plain.json`, `k` = (dry plain cost of `s01` + `r1`) / 0.9075 (the same two steps there), and
  `probe` = 436,436 tokens × $0.20 per million × 1.15 (one probe re-reads the whole history at the cache-read price,
  plus 15 % for output and small writes). If `1.4 × projected` exceeds `(90 − dry spend) / 8`, the workload drops
  Read batch 4 (`r4`) for all pairs and this is recorded as amendment 3 before pair 1.

## 6. What is reported

Per arm and per pair: $ at list price (client JSON; recall also the ledger), forwarded tokens mean and max,
compactions, exact probes (and contains), re-asks of the two typed decisions, task checks, upstream 4xx, re-layouts,
failed record reads, Recall's signed net from the ledger (`/__recall-stats`), the per-step cost curve. Over the 4
pairs: the mean with the lowest and highest value. Nothing is tuned after a result is seen.
