Pre-registration: Haiku 4.5 (final) - Source file: docs/verify/launch-ab-2026-09-29/final-haiku45-prereg.md - Commit that added it: af66f1bdf43cf992c9d90786f0d0a7895235dc73, 2026-09-29T07:16:38Z - The run it pre-registers started at 2026-09-29T07:16:50Z (the runner's start time in the evidence file). - Evidence: /benchmarks/data/final-haiku45.json - This copy is the file as it stood at that commit. The file has not changed since. - Edits: private fields (tenant and session ids, local paths, internal host and endpoint names) and passages about workloads that are not in the published evidence are removed, and each such edit is marked [edited: ...]. Links to files that are not published are shown as plain names. --- # Final benchmark, Haiku 4.5: pre-registration Written and pushed before the first pair. Docs only. No pair is added or dropped after seeing data. ## Freeze - Freeze commit: `3bedf1ba29958a8c4512cc53de2fe6aac4eb045b` (the safe fold, see close5.md, Amendment 1). - Deployed 07:09:43Z. Real-client smoke before the run: 9 of 9 requests ok, 0 upstream 4xx. - The runner reads the deployed commit with `gh api` (newest `main` commit whose [edited: the gateway Worker's build] check-run succeeded) before and after every pair. - A pair during which the deployed commit changed is invalid and runs again. - A deployed commit that differs from the freeze outside `docs/` and `scripts/e2e/` stops the runner. ## Pairs Model: Haiku 4.5 (`claude-haiku-4-5-20251001`, the runner's default). | Workload | Measured pairs | |---|---:| | tests | 12 | | json | 20 | | log | 20 | | git | 12 | | grep | 20 | | edit | 8 | | minitest | 8 | | W3 | 12 | - Warm-up: 1 pair per workload, discarded. - Arm order alternates per pair. - A 429 stops the pair; the runner waits, then runs it again. - Cap: $15. ## Command ``` F=3bedf1ba29958a8c4512cc53de2fe6aac4eb045b OUT=docs/verify/launch-ab-2026-09-29/final-haiku45.json node scripts/e2e/unified-ab.mjs --only tests,json,log,git,grep,edit,minitest,W3 \ --pairs-map tests:12,json:20,log:20,git:12,grep:20,edit:8,minitest:8,W3:12 \ --conc 4 --retries 3 --cap 15 --freeze $F --out $OUT # when the runner pauses at its deadline, the same run continues with: node scripts/e2e/unified-ab.mjs --only tests,json,log,git,grep,edit,minitest,W3 \ --pairs-map tests:12,json:20,log:20,git:12,grep:20,edit:8,minitest:8,W3:12 \ --conc 4 --retries 3 --cap 15 --freeze $F --resume $OUT ``` ## Analysis (fixed now) On every measured valid pair, no outlier dropped, with rerun-analyze.py: - Paired % difference (gateway vs plain) with the 95% t interval. - "less" or "more" only when the interval excludes zero; else "same". - Median paired difference. - Wins (pairs where the gateway cost less). - Success per arm. - Plain cost: the client's `total_cost_usd`. - Gateway cost: the sum of [edited: the gateway's per-call cost record] over the run's sessions, hidden calls included, checked against the client's figure per run. ## What would count against the gateway - A lower success rate than plain on any row. - A 400 on any run.