When does on-prem inference make sense?
When the data cannot leave a boundary that no AU-region API sits inside, or when steady-state utilisation is above ~60% of a rack for a year. Not for cost at typical enterprise volumes; the 18-month payback did not hold.
Tier is not strength
Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.
Machine-readable target
{
"taskType": "inference-hosting",
"configKey": "hosting.decision.utilisation_threshold"
}Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.
Body
As of 29 July 2026: on-prem inference makes sense on two grounds and not on a third. It makes sense when a data-residency or classification requirement rules out every AU-region API endpoint — some government and defence-adjacent work, and a small amount of banking. It makes sense when sustained utilisation is high enough that owned hardware beats metered pricing, which on our model is above roughly 60% of a rack, sustained, for a year. It does not make sense on cost at the volumes most enterprise clients actually run (c-on-prem-inference-1).
Evidence: two validation runs and g-onprem-h100-cluster, the graveyard entry where an 18-month payback case was built for a client and fell apart on utilisation — actual load ran at 20–30% of the modelled figure, and API prices fell 40% during the build (c-on-prem-inference-2). TEE-based confidential inference (x-tee-inference, in flight) may take the residency ground away from on-prem for banking PII; see pos-sovereign-inference for the longer view.
Caveat: the hardware price is the moving part. A cheaper inference-class card in the AU channel changes the utilisation threshold, and Argus flagged one such move on 1 September; the graveyard entry is marked resurrectable on it and has been resurfaced this week. This answer is due for review inside a fortnight.
What it rests on
The cost crossover between owned H100-class hardware and AU-region hosted inference for a 70B open-weight model sits at roughly 55–65% sustained utilisation at 2026 prices.
Enterprise inference clusters run at 20–40% sustained utilisation; the crossover is not reached in practice.
Field
On-Prem InferenceGraveyard · superseded
On-prem H100 cluster pays back inside 18 months “Payback period outran the price list.”