Pre-registration: Haiku 4.5 (final)
The plan for one benchmark run: the tasks, the number of pairs, the cost basis and the analysis. It was committed before the run's first pair.
Commit
af66f1bdf43cf992c9d90786f0d0a7895235dc73, the commit that added this file to the gateway repositoryCommitted2026-09-29 07:16:38 UTC
Run started2026-09-29 07:16:50 UTC, from the evidence file
Evidencefinal-haiku45.json
This copyThe file as it stood at that commit. Private fields and passages about workloads that are not in the published evidence are removed; each edit is marked in the text. Plain text
# Final benchmark, Haiku 4.5: pre-registration Written and pushed before the first pair. Docs only. No pair is added or dropped after seeing data. ## Freeze - Freeze commit: `3bedf1ba29958a8c4512cc53de2fe6aac4eb045b` (the safe fold, see close5.md, Amendment 1). - Deployed 07:09:43Z. Real-client smoke before the run: 9 of 9 requests ok, 0 upstream 4xx. - The runner reads the deployed commit with `gh api` (newest `main` commit whose [edited: the gateway Worker's build] check-run succeeded) before and after every pair. - A pair during which the deployed commit changed is invalid and runs again. - A deployed commit that differs from the freeze outside `docs/` and `scripts/e2e/` stops the runner. ## Pairs Model: Haiku 4.5 (`claude-haiku-4-5-20251001`, the runner's default). | Workload | Measured pairs | |---|---:| | tests | 12 | | json | 20 | | log | 20 | | git | 12 | | grep | 20 | | edit | 8 | | minitest | 8 | | W3 | 12 | - Warm-up: 1 pair per workload, discarded. - Arm order alternates per pair. - A 429 stops the pair; the runner waits, then runs it again. - Cap: $15. ## Command ``` F=3bedf1ba29958a8c4512cc53de2fe6aac4eb045b OUT=docs/verify/launch-ab-2026-09-29/final-haiku45.json node scripts/e2e/unified-ab.mjs --only tests,json,log,git,grep,edit,minitest,W3 \ --pairs-map tests:12,json:20,log:20,git:12,grep:20,edit:8,minitest:8,W3:12 \ --conc 4 --retries 3 --cap 15 --freeze $F --out $OUT # when the runner pauses at its deadline, the same run continues with: node scripts/e2e/unified-ab.mjs --only tests,json,log,git,grep,edit,minitest,W3 \ --pairs-map tests:12,json:20,log:20,git:12,grep:20,edit:8,minitest:8,W3:12 \ --conc 4 --retries 3 --cap 15 --freeze $F --resume $OUT ``` ## Analysis (fixed now) On every measured valid pair, no outlier dropped, with rerun-analyze.py: - Paired % difference (gateway vs plain) with the 95% t interval. - "less" or "more" only when the interval excludes zero; else "same". - Median paired difference. - Wins (pairs where the gateway cost less). - Success per arm. - Plain cost: the client's `total_cost_usd`. - Gateway cost: the sum of [edited: the gateway's per-call cost record] over the run's sessions, hidden calls included, checked against the client's figure per run. ## What would count against the gateway - A lower success rate than plain on any row. - A 400 on any run.