Pre-registration: Opus 5.5 test loops and data analysis (final) - Source file: docs/verify/launch-ab-2026-09-29/final-opus55-tests-W3-minitest-W4-prereg.md - Commit that added it: 83429d30dcbfff8c5b7708270875b05c812b3a80, 2026-09-29T07:18:09Z - The run it pre-registers started at 2026-09-29T07:18:17Z (the runner's start time in the evidence file). - Evidence: /benchmarks/data/final-opus55-tests-W3-minitest-W4.json - This copy is the file as it stood at that commit. The file has not changed since. - Edits: private fields (tenant and session ids, local paths, internal host and endpoint names) and passages about workloads that are not in the published evidence are removed, and each such edit is marked [edited: ...]. Links to files that are not published are shown as plain names. - Note: This run also measured a workload that is not in the published evidence; its rows in the plan are removed and marked. --- # Final run, Opus 5.5: tests, W3, minitest, [edited: one workload removed] (pre-registration) Written and pushed before the first pair. Nothing below changes after the data. ## What is measured - Gateway code: freeze commit `3bedf1ba` (the safe fold, close5.md Amendment 1). Its [edited: the gateway Worker's build] check passed at 2026-09-29T07:09:43Z, and its smoke check passed. - Real-client smoke before this run (the hard rule for a change to the Messages request shape): `claude -p` on claude-opus-5-5 through the gateway, fresh `CLAUDE_CONFIG_DIR`, disposable tenant, deferred MCP tools through ToolSearch, Bash and Read, then two `--resume` turns. After the deploy: 9 of 9 requests ok, 0 errors, 0 incidents. - Client: Claude Code 2.1.283, model `claude-opus-5-5`, both arms. | Workload | Use case | Measured pairs | Warm-up pairs (discarded) | |---|---|---:|---:| | tests | test-fix loop, JS | 12 | 1 | | W3 | data analysis with a script | 20 | 1 | | minitest | test-fix loop, Ruby | 8 | 1 | | [edited: one row removed] | | | | ## Command ``` F=3bedf1ba D=docs/verify/launch-ab-2026-09-29 node scripts/e2e/unified-ab.mjs --only tests,W3,minitest,[edited: one] --model claude-opus-5-5 \ --pairs-map tests:12,W3:20,minitest:8,[edited: one] --conc 3 --cap 50 --freeze $F \ --prior-from $D/rerun-opus55-tests-W3-minitest-W4.json,$D/rerun20-opus55-W3-W4.json \ --out $D/final-opus55-tests-W3-minitest-W4.json ``` A part that stops at the runner's deadline goes on with `--resume $D/final-opus55-tests-W3-minitest-W4.json` and the same options. `--prior-from` is new in this commit. It changes only the cost plan: the old plan priced each pair from Haiku 4.5 runs scaled by list price. For Opus 5.5 that gave $62.10, above the $50 cap, so the runner refused to start. The plan now uses the cost per pair that we measured on Opus 5.5 in the two named files: $31.24 in total. The in-run cap check still stops a workload before a pair would pass $50 of total spend. ## Rules the runner enforces - Pairs run back to back. The arm order alternates each pair (even pair: plain first; odd pair: gateway first). - Before and after every pair, the runner reads the deployed commit with `gh api` (the newest `main` commit whose [edited: the gateway Worker's build] check passed), plus the Cloudflare deployment id. If it changed during a pair, that pair is invalid and runs again. If the deployed commit differs from `3bedf1ba` outside `docs/` and `scripts/e2e/`, the runner stops. - A 429 stops the pair. The runner waits, then runs it again. - The gateway arm uses one disposable tenant per workload. Owner and customer tenants are refused. ## Analysis (fixed now) rerun-analyze.py, unchanged, on every measured valid pair. No outlier is dropped. - Plain cost: the client's `total_cost_usd`. Gateway cost: the sum of [edited: the gateway's per-call cost record] over the run's sessions, hidden calls included. The client's own figure is kept as a cross-check. - Paired % difference: mean(gateway − plain) / mean(plain), with the 95% t interval. - Verdict: "less" when the whole interval is below zero, "more" when it is above zero, else "same". - Also reported: median paired difference, wins (pairs where the gateway cost less), success per arm. What would count against the gateway: a lower success rate on any row, or a 400 on any run. ## Result (Written after the runs, in a separate file.)