AI Gateway
over all model gardens
One gateway the client controls — routing, observability and cost attribution over Bedrock, Vertex, Azure, direct APIs and on-prem — is table stakes; standardising on a vendor's gateway trades a small build for a large lock-in and loses the on-prem and open-weight legs.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
81%human-committedExpiry
63duntil review · 5 Nov 2026Lead time
−4moopened after mainstream — recorded honestlyOwnership
PRPriya Ramanmonthly cadenceWhere it is
The field converged in Q2. Every delivery pattern now routes through a gateway, the open-source gateways reached parity with the commercial ones on provider coverage, and our routing experiment found that static task-class rules capture nearly all the cost saving that learned routers promise. What clients actually buy the gateway for is cost attribution per use case, not routing. The remaining work is operational: adapter drift on streaming and tool-call formats recurs every provider release, and that is a tooling problem, not a research one. We expect to dissolve this field into practice by Q4.
Why a Quantium decision hinges on it
Telco and banking clients ask for one view of spend across two or three model gardens before they ask anything about models. Without a gateway they control, a client cannot route to an open-weight or on-prem model when sovereignty or cost demands it, and cannot answer 'what does use case X cost' at all. A hyperscaler gateway answers that question for the hyperscaler's garden only. The choice is made once per client and is expensive to reverse.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Static task-class routing across Bedrock, Vertex and a direct API cut cost 34% at equal eval score on four of six task classes (x-gateway-routing). A learned router added 2 points on top and lost 9 outside its training distribution.
- 02Per-use-case cost attribution through the gateway answered the telco client's spend question in one dashboard; the hyperscaler console could not.
- 03One open-source gateway covered all five legs we needed, including on-prem vLLM. The commercial gateway we trialled covered three.
- 01Learned routers. Outside their training distribution they underperform a lookup table, and the table is auditable.
- 02'Semantic caching' as a headline feature. Hit rates on our workloads were under 6%; prompt caching at the provider does the real work.
- 03Analyst claims of a discrete 'AI gateway market'. It is a feature of a platform, and most of it will be free.
- 01Provider tool-call and streaming formats stabilising enough that adapter drift stops being a monthly incident.
- 02Hyperscaler gateways routing outside their own garden, which none of them does today and none has announced.
- 03Cost attribution surviving agentic loops, where one task fans out across models and the per-call attribution stops meaning anything.
- 01Keep r-gateway-default as the delivery default and hand the field to practice in Q4.
- 02Stop re-running routing benchmarks; the answer is stable. Track adapter drift as an instrument-health metric instead.
- 03Fold the on-prem and open-weight legs into the standard gateway build so those fields do not need their own plumbing.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Routing experiment: static task-class rules cut cost 34% at equal eval score; learned router adds 2 points and loses 9 out of distribution
Six task classes routed across Bedrock, Vertex and a direct API through an open-source gateway. Static rules keyed on task class captured a 34% saving on four classes; learned routing added marginally in-distribution and underperformed the rules outside it. Agentic loops gained nothing.
extracted claimStatic task-class routing captures the cost saving; learned routers do not generalise.

Logged from Claude Code: gateway upgrade broke tool-call streaming on one provider for four days
Product engineer logged from a session: a provider changed its streamed tool-call delta format; the gateway adapter silently dropped arguments. Found by a failing eval, not by monitoring. Tried tier, one harness.

RouteBench: Learned LLM Routers Under Distribution Shift
Evaluates eight learned routers against static rules across shifted task mixes. Learned routers beat rules in-distribution by 1–3 points and lose by 6–12 under shift. Matches our experiment almost exactly.
extracted claimLearned routers underperform static task-class rules outside their training distribution.

Three AU banks post 'LLM Gateway Engineer' roles in the same fortnight
All three job descriptions lead with cost attribution and showback; routing appears in one. Argus inference: banks are building their own gateways and treating them as finance tooling.

'Your gateway is your lock-in'
Argues the gateway is the most durable lock-in in the stack because every prompt, log and cost centre tag lives in its schema. Agrees with our position on owning it; disagrees on whether open source is cheaper. Kept for the second half.

Hyperscaler gateway adds 'cross-provider routing' — to models inside its own garden only
Marketed as multi-model routing. Reads the fine print: every route terminates inside the provider's own catalogue. No direct-API leg, no on-prem leg. Confirms the pattern from the commercial trial.

Analyst note: '70% of enterprises will route inference through a gateway by 2027'
Demand-band signal with a discrete-market framing we do not share. Useful as evidence the pattern is mainstream; not useful on build versus buy, where its three-year TCO figure assumes zero adapter maintenance for the commercial option.

Commercial gateway trial ended: no on-prem leg, opaque cost attribution
Six-week trial of a commercial gateway on the telco pattern. Three of five legs covered; cost attribution exported as a monthly CSV with no use-case dimension. Became g-single-gateway-vendor.

Open-source gateway passes 30k stars; adds Bedrock, Vertex, Azure and vLLM parity in one release
The release that closed the provider-coverage gap with the commercial gateways. Streaming and tool-call adapters are where the issue tracker lives; a third of open issues are format drift after a provider release.
extracted claimOpen-source gateways reached provider-coverage parity with commercial ones in early 2026.

'Can we see cost per use case across Bedrock and Vertex in one place?'
Asked by a telco CFO's office, not the engineering team. The question that reframed the field from routing to attribution. Logged unanswered; the gateway build that followed answered it in one dashboard.
Claims · 5 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
Static task-class routing rules capture most of the achievable cost saving (30–35%); learned routers add little and lose outside their training distribution.
Vendor gateways cover their own garden well and every other garden badly; the on-prem and open-weight legs are where they fail.
Cost attribution per use case is the feature clients buy the gateway for; routing is secondary.
Routing gains vanish on agentic loops because the loop needs one model's tool-call semantics end to end.
Adapter drift on streaming and tool-call formats is the recurring operational cost of a gateway; it re-appears with every provider release.
A commercial gateway is cheaper over three years than an owned open-source one once staffing is included.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 50% → 81%.
Routing experiment concluded: static rules capture the saving; learned routers do not generalise; agentic loops do not benefit. Field converged. Recommendation published; hand to practice in Q4.
- Static task-class routing rules capture most of the achievable cost saving (30–35%); learned routers add little and lose outside their training distribution.
- Cost attribution per use case is the feature clients buy the gateway for; routing is secondary.
- Adapter drift on streaming and tool-call formats is the recurring operational cost of a gateway; it re-appears with every provider release.
- Routing gains vanish on agentic loops because the loop needs one model's tool-call semantics end to end.
- c-ai-gateway-3 ↑ 0.7 → 0.77
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · PREvery inference call in every pattern passes through it. Made once per client.
Timeline
committed · AWAlready mainstream. We opened this field late.
TAM
agent-estimatedAgent-estimated from gateway vendor revenue and platform spend; most of it will be bundled. Uncommitted.
Cost
committed · PROne engineer, one week for the standard build; drift maintenance thereafter.
Demand
committed · MLAsked in every telco engagement this year, always as a cost question.
Cost of being wrong
committed · LFWrong gateway is a migration, not an incident.
Workforce readiness
agent-estimatedDelivery teams can stand it up; adapter drift still lands on the lab. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Two model gardens under separate contracts and a finance team that wants spend per use case. The gateway is the only place that number exists.
Mechanism · Gateway tags every call with use case and cost centre; showback runs off the gateway log.
Banks need an on-prem or sovereign leg for some workloads and a frontier leg for others; the gateway is where the policy lives.
Mechanism · Routing rules by data classification; sovereign leg for PII-bearing calls.
Health clients hit residency constraints first and cost second; both are gateway questions.
Mechanism · Residency-aware routing with the on-prem leg as default for clinical text.
Agencies procure gardens through panels one at a time; multi-garden routing is not yet a question they can ask.
Mechanism · Would apply once a second garden lands on a whole-of-government arrangement.
Red team · the strongest case against
The strongest case against: the hyperscalers will make the gateway free and native, and 'a gateway you control' is a maintenance burden a twelve-person lab is telling delivery teams to carry forever. Most clients have one cloud contract, and multi-garden routing solves a problem they do not have.
- —Adapter drift is a permanent tax. We measured it at roughly one incident per provider release; over five providers that is a part-time engineer per client, which the cost score does not include.
- —Cost attribution is a logging feature. Any provider console could ship it next quarter and remove the main reason to own the gateway.
- —Routing gains of 34% were measured on classification and extraction. On agentic loops — where spend is growing fastest — the gain was zero.
- —A converged field with a low-priority score is a field that should already be in practice. Keeping it open flatters the lab's coverage.
Source diversity
- Open-source infra35%
- ML research15%
- Vendor20%
- Analyst10%
- Internal / Engel20%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Open weights become a routable leg only through the gateway; task-class parity is what makes routing to them worth doing.
Routing and caching both live in the gateway; the ledger that measures them is the gateway log.
The delegation broker needs one choke point for tool calls and the gateway already is one for model calls.
An on-prem leg is only usable in delivery when the gateway can route to it by data classification.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- OGOllie Grant · Product engineer2 drops
- PRPriya Raman · Research engineer · inference2 drops
- MLMarcus Lee · Delivery lead · Telco1 drop
- MTMei Tanaka · Research lead · evals1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Route through a gateway you control; do not standardise on a vendor's
strength strong · 33 citations · review 11 Jan 2027
Cost-aware routing across three model gardens
A routing policy we own, keyed on our per-task eval pass rate, cuts blended cost per task by at least 25% against a fixed frontier default with no drop in eval pass rate.
One commercial gateway over every model garden
“Standardised on someone else's roadmap.” · lived 7 months
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Does per-call cost attribution mean anything on an agentic loop that fans out across three models?
- 02What is the true adapter-drift cost per client per year, measured rather than estimated?
- 03When a hyperscaler first routes outside its own garden, does the own-the-gateway recommendation survive?