Pre-registration: Opus 5.5 multi-agent and data analysis (v4)

The plan for one benchmark run: the tasks, the number of pairs, the cost basis and the analysis. It was committed before the run's first pair.

Commit74b3f851ff0b39794e5c4146d6f39f361da5d554, the commit that added this file to the gateway repository
Committed2026-09-29 17:04:33 UTC
Run started2026-09-29 18:20:50 UTC, from the evidence file
Evidencev4-opus55-W1-W2-W3-W4.json
This copyThe file as it stood at that commit. Private fields and passages about workloads that are not in the published evidence are removed; each edit is marked in the text. Plain text
NoteThis run also planned a workload that is not in the published evidence. The owner stopped Part B after W2 and W3 finished, before any pair of it ran; that amendment was added to the file after the runs began.
# Run v4, Opus 5.5: [edited: three workloads removed] W2, W3 — pre-registration, 2026-09-29

Written and pushed before the first pair. Nothing below changes after the data.

## Why this run

[edited: a paragraph and a table on a workload that is not in the published evidence]

Two fixes, same code for Claude Code and Codex (`src/memory/in-conversation.ts` `cmConversationOpening`,
`src/gateway/proxy-memory-tiers.ts` `cmP2FanOut`): the fact tier (`59052465`), then the knowledge-base search and the
episodic tier (`3cba9400`), read only while the conversation holds no assistant message — the first human turn, a
subagent's task, the first request after a compaction. A later human turn carries no cross-session recall; the model
has `context-mode-recall` and `context-mode-search`. Dry run [edited: on one workload] after the deploy: the opening request
carries the tail, the five later turn starts carry none.

## Freeze

- Freeze commit: `3cba94005e5e5b2cceb308d5cc926887d9e01d3b`. [edited: the gateway Worker's build] succeeded on it and on
  `59052465`.
- Real-client smoke (`claude -p`, `claude-opus-5-5`, fresh `CLAUDE_CONFIG_DIR`, disposable tenant, ToolSearch-deferred
  MCP tools with a mid-session tool addition, two `--resume` turns): before `59052465` 8/8 requests ok; after it 8/8;
  after `3cba9400` 7/7. 0 upstream 4xx.
- This pre-registration touches `docs/` only.

## Which rows change, and which do not

The fix changes the forwarded request only on a later human turn. W2, W3 [edited: and three more workloads] have more than one human
turn, so all five run again here. The six short tasks, `tests` and `minitest` are one human turn each: their forwarded
requests are unchanged, and their latest rows stand with the commit they were measured on.

## Part A

[edited: Part A ran one workload that is not in the published evidence]

## Part B — [edited: two workloads removed] W2, W3 (after part A ends, same freeze)

| Model | Workload | Measured pairs | Warm-up |
|---|---|---:|---:|
| [edited: one row removed] | | | |
| claude-opus-5-5 | W2 (subagents) | 20 | 1 |
| claude-opus-5-5 | W3 (data analysis) | 20 | 1 |
| [edited: one row removed] | | | |

Cap $70 (plan: $51.28 at Opus 5.5 list price, priced from the earlier runs).

```
node scripts/e2e/unified-ab.mjs --only [edited: one],W2,W3,[edited: one] --model claude-opus-5-5 --pairs-map [edited: one],W2:20,W3:20,[edited: one] \
  --conc 2 --cap 70 --freeze $F \
  --prior-from [edited: an earlier evidence file],docs/verify/launch-ab-2026-09-29/final-opus55-tests-W3-minitest-W4.json \
  --out docs/verify/launch-ab-2026-09-29/v4-opus55-W1-W2-W3-W4.json
```

A part that pauses (the 60-minute egress token, a 429) goes on with the same command plus `--resume <that file>`.

## Rules

- Real client only: `claude -p` (Claude Code 2.1.283), fresh `CLAUDE_CONFIG_DIR` per run, one disposable tenant per
  workload. Arm order alternates per pair. No pairs are added after seeing data.
- The runner reads the deployed commit with `gh api` before and after every pair. A pair during which it changed is
  invalid and runs again. A deployed commit that differs from the freeze outside `docs/`, `scripts/e2e/`, `site/` and
  root `*.md` stops the runner.
- Invalid attempts stay in the file, marked superseded.

## Cost basis

cost-basis.md: plain = the client's `total_cost_usd`; gateway = the sum of [edited: the gateway's per-call cost record]
over the run's sessions, hidden calls included, each call at the list price of the model it ran on. A gateway run whose
ops tokens fall below the client's in any class is invalid.

## Analysis

rerun-analyze.py, every measured valid pair, no outlier dropped: paired % difference (gateway −
plain, as % of the plain mean) with the 95% t interval; "less" or "more" only when the interval excludes zero, else
"same"; median paired difference; wins; success per arm; tag recall per arm [edited: on two workloads].

What we expect, stated now: [edited: two workloads] W2, W3 same or less. What would count against the fix: [edited: one workload]; a lower success or recall rate in the gateway arm on any row; an upstream 4xx on any run. Whatever the result,
it is published as measured.

## Result

Written after the runs, below this line.