Pre-registration: Haiku 4.5 (final)

The plan for one benchmark run: the tasks, the number of pairs, the cost basis and the analysis. It was committed before the run's first pair.

Commitaf66f1bdf43cf992c9d90786f0d0a7895235dc73, the commit that added this file to the gateway repository
Committed2026-09-29 07:16:38 UTC
Run started2026-09-29 07:16:50 UTC, from the evidence file
Evidencefinal-haiku45.json
This copyThe file as it stood at that commit. Private fields and passages about workloads that are not in the published evidence are removed; each edit is marked in the text. Plain text
# Final benchmark, Haiku 4.5: pre-registration

Written and pushed before the first pair. Docs only. No pair is added or dropped after seeing data.

## Freeze

- Freeze commit: `3bedf1ba29958a8c4512cc53de2fe6aac4eb045b` (the safe fold, see close5.md, Amendment 1).
- Deployed 07:09:43Z. Real-client smoke before the run: 9 of 9 requests ok, 0 upstream 4xx.
- The runner reads the deployed commit with `gh api` (newest `main` commit whose [edited: the gateway Worker's build]
  check-run succeeded) before and after every pair.
  - A pair during which the deployed commit changed is invalid and runs again.
  - A deployed commit that differs from the freeze outside `docs/` and `scripts/e2e/` stops the runner.

## Pairs

Model: Haiku 4.5 (`claude-haiku-4-5-20251001`, the runner's default).

| Workload | Measured pairs |
|---|---:|
| tests | 12 |
| json | 20 |
| log | 20 |
| git | 12 |
| grep | 20 |
| edit | 8 |
| minitest | 8 |
| W3 | 12 |

- Warm-up: 1 pair per workload, discarded.
- Arm order alternates per pair.
- A 429 stops the pair; the runner waits, then runs it again.
- Cap: $15.

## Command

```
F=3bedf1ba29958a8c4512cc53de2fe6aac4eb045b
OUT=docs/verify/launch-ab-2026-09-29/final-haiku45.json
node scripts/e2e/unified-ab.mjs --only tests,json,log,git,grep,edit,minitest,W3 \
  --pairs-map tests:12,json:20,log:20,git:12,grep:20,edit:8,minitest:8,W3:12 \
  --conc 4 --retries 3 --cap 15 --freeze $F --out $OUT
# when the runner pauses at its deadline, the same run continues with:
node scripts/e2e/unified-ab.mjs --only tests,json,log,git,grep,edit,minitest,W3 \
  --pairs-map tests:12,json:20,log:20,git:12,grep:20,edit:8,minitest:8,W3:12 \
  --conc 4 --retries 3 --cap 15 --freeze $F --resume $OUT
```

## Analysis (fixed now)

On every measured valid pair, no outlier dropped, with rerun-analyze.py:

- Paired % difference (gateway vs plain) with the 95% t interval.
  - "less" or "more" only when the interval excludes zero; else "same".
- Median paired difference.
- Wins (pairs where the gateway cost less).
- Success per arm.
- Plain cost: the client's `total_cost_usd`.
- Gateway cost: the sum of [edited: the gateway's per-call cost record] over the run's sessions, hidden calls included, checked against
  the client's figure per run.

## What would count against the gateway

- A lower success rate than plain on any row.
- A 400 on any run.