Cost-aware routing across three model gardens
AI Gateway- proposed
- voting
- running
- measuring
- concluded
Preregistration · v1 · 4 May 2026 · immutable after start
Hypothesis
A routing policy we own, keyed on our per-task eval pass rate, cuts blended cost per task by at least 25% against a fixed frontier default with no drop in eval pass rate.
Kill condition
Kill if blended cost saving is under 10% by day 12, or if eval pass rate on any task family drops more than 2 points at any saving level.
Result
ValidatedValidated. Routed configuration cut blended cost per task by 34% with no family dropping more than 1 point of pass rate. The vendor router comparison failed our extraction eval on 38% of calls. Published as r-gateway-default; g-single-gateway-vendor superseded.
Running notes · fed from harness sessions and by the pair
- manual30 Jun 2026PR Priya Raman
Concluded. 34% saving, no family dropped more than 1 point. The per-call ledger is the reusable artefact; r-gateway-default written.
- harness18 Jun 2026PR Priya Raman
vendor gateway comparison arm: their router sent 38% of extraction to a model failing our eval. logged to g-single-gateway-vendor
- manual9 Jun 2026SK Sam Kowalczyk
Replay done. Blended cost -34%. Extraction family routed to Qwen 71% of the time at pass rate 0.93 vs 0.94 frontier. Agentic loops stayed on frontier as expected.
- harness21 May 2026PR Priya Raman
gateway up; 3 gardens; per-call ledger writing to postgres; replay harness 18k calls loaded
- manual4 May 2026PR Priya Raman
Predicted 30% saving, confidence 0.55. The risk is that the extraction family has no cheap model that clears the bar.
Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.