cavendish
Standing answerAssessedstrength · moderate

When does on-prem inference make sense?

When the data cannot leave a boundary that no AU-region API sits inside, or when steady-state utilisation is above ~60% of a rack for a year. Not for cost at typical enterprise volumes; the 18-month payback did not hold.

Tier is not strength

Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.

Evidence tierAssessed
Strengthmoderate
OwnerPRPriya Raman
Last validated29 Jul 2026
Review by12 Sep 2026
Half-life45 days
Citations22
Asked41× this quarter
VerticalsBanking, Government, Defence, Health
decay9d until review

Machine-readable target

{
  "taskType": "inference-hosting",
  "configKey": "hosting.decision.utilisation_threshold"
}

Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.

Body

As of 29 July 2026: on-prem inference makes sense on two grounds and not on a third. It makes sense when a data-residency or classification requirement rules out every AU-region API endpoint — some government and defence-adjacent work, and a small amount of banking. It makes sense when sustained utilisation is high enough that owned hardware beats metered pricing, which on our model is above roughly 60% of a rack, sustained, for a year. It does not make sense on cost at the volumes most enterprise clients actually run (c-on-prem-inference-1).

Evidence: two validation runs and g-onprem-h100-cluster, the graveyard entry where an 18-month payback case was built for a client and fell apart on utilisation — actual load ran at 20–30% of the modelled figure, and API prices fell 40% during the build (c-on-prem-inference-2). TEE-based confidential inference (x-tee-inference, in flight) may take the residency ground away from on-prem for banking PII; see pos-sovereign-inference for the longer view.

Caveat: the hardware price is the moving part. A cheaper inference-class card in the AU channel changes the utilisation threshold, and Argus flagged one such move on 1 September; the graveyard entry is marked resurrectable on it and has been resurfaced this week. This answer is due for review inside a fortnight.

PRSigned Priya Raman · Research engineer · inference · 29 Jul 2026

What it rests on

The cost crossover between owned H100-class hardware and AU-region hosted inference for a 70B open-weight model sits at roughly 55–65% sustained utilisation at 2026 prices.

Assessed c-on-prem-inference-1
80%

Enterprise inference clusters run at 20–40% sustained utilisation; the crossover is not reached in practice.

Assessed c-on-prem-inference-2
76%

Field

On-Prem Inference

Graveyard · superseded

On-prem H100 cluster pays back inside 18 months Payback period outran the price list.