Pre-registration: Opus 5.5 test loops and data analysis (final)

The plan for one benchmark run: the tasks, the number of pairs, the cost basis and the analysis. It was committed before the run's first pair.

Commit83429d30dcbfff8c5b7708270875b05c812b3a80, the commit that added this file to the gateway repository
Committed2026-09-29 07:18:09 UTC
Run started2026-09-29 07:18:17 UTC, from the evidence file
Evidencefinal-opus55-tests-W3-minitest-W4.json
This copyThe file as it stood at that commit. Private fields and passages about workloads that are not in the published evidence are removed; each edit is marked in the text. Plain text
NoteThis run also measured a workload that is not in the published evidence; its rows in the plan are removed and marked.
# Final run, Opus 5.5: tests, W3, minitest, [edited: one workload removed] (pre-registration)

Written and pushed before the first pair. Nothing below changes after the data.

## What is measured

- Gateway code: freeze commit `3bedf1ba` (the safe fold, close5.md Amendment 1). Its
  [edited: the gateway Worker's build] check passed at 2026-09-29T07:09:43Z, and its smoke check passed.
- Real-client smoke before this run (the hard rule for a change to the Messages request shape): `claude -p` on
  claude-opus-5-5 through the gateway, fresh `CLAUDE_CONFIG_DIR`, disposable tenant, deferred MCP tools through
  ToolSearch, Bash and Read, then two `--resume` turns. After the deploy: 9 of 9 requests ok, 0 errors, 0 incidents.
- Client: Claude Code 2.1.283, model `claude-opus-5-5`, both arms.

| Workload | Use case | Measured pairs | Warm-up pairs (discarded) |
|---|---|---:|---:|
| tests | test-fix loop, JS | 12 | 1 |
| W3 | data analysis with a script | 20 | 1 |
| minitest | test-fix loop, Ruby | 8 | 1 |
| [edited: one row removed] | | | |

## Command

```
F=3bedf1ba
D=docs/verify/launch-ab-2026-09-29
node scripts/e2e/unified-ab.mjs --only tests,W3,minitest,[edited: one] --model claude-opus-5-5 \
  --pairs-map tests:12,W3:20,minitest:8,[edited: one] --conc 3 --cap 50 --freeze $F \
  --prior-from $D/rerun-opus55-tests-W3-minitest-W4.json,$D/rerun20-opus55-W3-W4.json \
  --out $D/final-opus55-tests-W3-minitest-W4.json
```

A part that stops at the runner's deadline goes on with `--resume $D/final-opus55-tests-W3-minitest-W4.json` and the
same options.

`--prior-from` is new in this commit. It changes only the cost plan: the old plan priced each pair from Haiku 4.5
runs scaled by list price. For Opus 5.5 that gave $62.10, above the $50 cap, so the runner refused to start. The
plan now uses the cost per pair that we measured on Opus 5.5 in the two named files: $31.24 in total. The in-run
cap check still stops a workload before a pair would pass $50 of total spend.

## Rules the runner enforces

- Pairs run back to back. The arm order alternates each pair (even pair: plain first; odd pair: gateway first).
- Before and after every pair, the runner reads the deployed commit with `gh api` (the newest `main` commit whose
  [edited: the gateway Worker's build] check passed), plus the Cloudflare deployment id. If it changed during a pair,
  that pair is invalid and runs again. If the deployed commit differs from `3bedf1ba` outside `docs/` and
  `scripts/e2e/`, the runner stops.
- A 429 stops the pair. The runner waits, then runs it again.
- The gateway arm uses one disposable tenant per workload. Owner and customer tenants are refused.

## Analysis (fixed now)

rerun-analyze.py, unchanged, on every measured valid pair. No outlier is dropped.

- Plain cost: the client's `total_cost_usd`. Gateway cost: the sum of [edited: the gateway's per-call cost record] over the run's
  sessions, hidden calls included. The client's own figure is kept as a cross-check.
- Paired % difference: mean(gateway − plain) / mean(plain), with the 95% t interval.
- Verdict: "less" when the whole interval is below zero, "more" when it is above zero, else "same".
- Also reported: median paired difference, wins (pairs where the gateway cost less), success per arm.

What would count against the gateway: a lower success rate on any row, or a 400 on any run.

## Result

(Written after the runs, in a separate file.)