Pre-registration: Opus 5.5 multi-agent and data analysis (v4) - Source file: docs/verify/launch-ab-2026-09-29/v4-opus55-prereg.md - Commit that added it: 74b3f851ff0b39794e5c4146d6f39f361da5d554, 2026-09-29T17:04:33Z - The run it pre-registers started at 2026-09-29T18:20:50Z (the runner's start time in the evidence file). - Evidence: /benchmarks/data/v4-opus55-W1-W2-W3-W4.json - This copy is the file as it stood at that commit. Later commits added results and notes to it; the results are in the evidence file. - Edits: private fields (tenant and session ids, local paths, internal host and endpoint names) and passages about workloads that are not in the published evidence are removed, and each such edit is marked [edited: ...]. Links to files that are not published are shown as plain names. - Note: This run also planned a workload that is not in the published evidence. The owner stopped Part B after W2 and W3 finished, before any pair of it ran; that amendment was added to the file after the runs began. --- # Run v4, Opus 5.5: [edited: three workloads removed] W2, W3 — pre-registration, 2026-09-29 Written and pushed before the first pair. Nothing below changes after the data. ## Why this run [edited: a paragraph and a table on a workload that is not in the published evidence] Two fixes, same code for Claude Code and Codex (`src/memory/in-conversation.ts` `cmConversationOpening`, `src/gateway/proxy-memory-tiers.ts` `cmP2FanOut`): the fact tier (`59052465`), then the knowledge-base search and the episodic tier (`3cba9400`), read only while the conversation holds no assistant message — the first human turn, a subagent's task, the first request after a compaction. A later human turn carries no cross-session recall; the model has `context-mode-recall` and `context-mode-search`. Dry run [edited: on one workload] after the deploy: the opening request carries the tail, the five later turn starts carry none. ## Freeze - Freeze commit: `3cba94005e5e5b2cceb308d5cc926887d9e01d3b`. [edited: the gateway Worker's build] succeeded on it and on `59052465`. - Real-client smoke (`claude -p`, `claude-opus-5-5`, fresh `CLAUDE_CONFIG_DIR`, disposable tenant, ToolSearch-deferred MCP tools with a mid-session tool addition, two `--resume` turns): before `59052465` 8/8 requests ok; after it 8/8; after `3cba9400` 7/7. 0 upstream 4xx. - This pre-registration touches `docs/` only. ## Which rows change, and which do not The fix changes the forwarded request only on a later human turn. W2, W3 [edited: and three more workloads] have more than one human turn, so all five run again here. The six short tasks, `tests` and `minitest` are one human turn each: their forwarded requests are unchanged, and their latest rows stand with the commit they were measured on. ## Part A [edited: Part A ran one workload that is not in the published evidence] ## Part B — [edited: two workloads removed] W2, W3 (after part A ends, same freeze) | Model | Workload | Measured pairs | Warm-up | |---|---|---:|---:| | [edited: one row removed] | | | | | claude-opus-5-5 | W2 (subagents) | 20 | 1 | | claude-opus-5-5 | W3 (data analysis) | 20 | 1 | | [edited: one row removed] | | | | Cap $70 (plan: $51.28 at Opus 5.5 list price, priced from the earlier runs). ``` node scripts/e2e/unified-ab.mjs --only [edited: one],W2,W3,[edited: one] --model claude-opus-5-5 --pairs-map [edited: one],W2:20,W3:20,[edited: one] \ --conc 2 --cap 70 --freeze $F \ --prior-from [edited: an earlier evidence file],docs/verify/launch-ab-2026-09-29/final-opus55-tests-W3-minitest-W4.json \ --out docs/verify/launch-ab-2026-09-29/v4-opus55-W1-W2-W3-W4.json ``` A part that pauses (the 60-minute egress token, a 429) goes on with the same command plus `--resume `. ## Rules - Real client only: `claude -p` (Claude Code 2.1.283), fresh `CLAUDE_CONFIG_DIR` per run, one disposable tenant per workload. Arm order alternates per pair. No pairs are added after seeing data. - The runner reads the deployed commit with `gh api` before and after every pair. A pair during which it changed is invalid and runs again. A deployed commit that differs from the freeze outside `docs/`, `scripts/e2e/`, `site/` and root `*.md` stops the runner. - Invalid attempts stay in the file, marked superseded. ## Cost basis cost-basis.md: plain = the client's `total_cost_usd`; gateway = the sum of [edited: the gateway's per-call cost record] over the run's sessions, hidden calls included, each call at the list price of the model it ran on. A gateway run whose ops tokens fall below the client's in any class is invalid. ## Analysis rerun-analyze.py, every measured valid pair, no outlier dropped: paired % difference (gateway − plain, as % of the plain mean) with the 95% t interval; "less" or "more" only when the interval excludes zero, else "same"; median paired difference; wins; success per arm; tag recall per arm [edited: on two workloads]. What we expect, stated now: [edited: two workloads] W2, W3 same or less. What would count against the fix: [edited: one workload]; a lower success or recall rate in the gateway arm on any row; an upstream 4xx on any run. Whatever the result, it is published as measured. ## Result Written after the runs, below this line.