cavendish
AssessedEmerginggate · AdoptionNear · 1–3 years

US Non-Dominance

Within three years the best model available to an Australian enterprise for a given task will, for a meaningful share of tasks, not be American — and AU procurement, which assumes US-default, is not ready for that.

A validation run. Researched position, no experiment.

Join with…

Confidence

50%human-committed

Expiry

35duntil review · 8 Oct 2026

Lead time

not yet mainstream · opened 10 Mar 2026

Ownership

AWAdam Witanowskimonthly cadence

Where it is

A conviction bet with a named person attached. The lab director opened this on judgement in March with no supporting run, and the record will score it. What has accumulated since: Chinese open-weight releases (DeepSeek, Alibaba's Qwen line) reached parity with US open weights on our task evals in Q2 and lead on cost; Mistral and two Gulf-funded labs now ship models AU regulators would have to form a view on; and US export controls on accelerators are reshaping where inference capacity sits, including in Southeast Asia. The strongest counter-evidence is also clear: frontier capability is still US-default, the gap at the top has not closed, and every client procurement framework we have read names US providers or hyperscalers by default. The gate is adoption. The models exist and are cheap; the question is whether an AU bank's risk function will run a Chinese open-weight model on customer data, and the current answer is no regardless of the evals.

Why a Quantium decision hinges on it

Quantium's model recommendations are read as neutral. If the cheapest parity model for extraction is a Chinese open-weight and the firm cannot recommend it because of a procurement reflex, the firm is leaving money on the table and saying so to clients. If the firm recommends it and a regulator or a board objects, that is a reputational event. Either way the firm needs a position — on provenance, on weights-hosting, on what 'sovereign' means when the weights are Chinese and the hosting is Australian — before a client asks in a way that cannot be deferred. The sponsor holds the question of whether the firm wants a public position at all.

Field attributes

StateEmerging
GateAdoption · ready — blocked by trust, regulation, procurement or change capacity
OriginConviction
Measurablepartial
Audience · TLPexec
Horizonnear
Opened10 Mar 2026
Mainstreamnot yet
Last validated27 Aug 2026
Sightings1

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Two Chinese open-weight models matched the best US open-weight on our extraction and classification evals in Q2 at 0.4–0.6× the hosted price (x-open-weight-parity, cross-referenced).
  • 02A Gulf-funded lab's model is now hosted in an AU region by a domestic provider with residency terms; the weights are not American and the hosting is Australian.
  • 03Three of five client procurement frameworks reviewed name US hyperscalers or providers as default-approved and have no path for a non-US model to be assessed.
What is hype
  • 01'China has caught up' at the frontier. On agentic and long-horizon tasks the gap is real and has not closed on any eval we run.
  • 02Sovereign AI meaning a national model. No AU-trained model is competitive and none is planned at scale; sovereignty here means hosting and control, not origin.
  • 03Export controls as a permanent US advantage. Every control so far has been followed within two quarters by a workaround in capacity or architecture.
What would have to be true
  • 01An AU regulator or a major bank's risk function would have to approve a non-US open-weight model on customer data, in writing, at least once.
  • 02Provenance assurance for open weights — who trained it, on what, with what backdoor risk — would have to exist as a service a client can buy.
  • 03The parity on extraction and classification would have to extend to at least one agentic task class, or the field stays confined to the cheap tier.
What we would do
  • 01Publish pos-sovereign-inference as the firm's position: recommend by eval and cost, disclose provenance, and treat hosting as the sovereignty control.
  • 02Build a provenance checklist for open-weight models and run every recommended model through it; the checklist is worth more than the position.
  • 03Score the conviction in March 2027: was the director right, and by how much?

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
±1.2 pts
extraction delta
Finding·band 1Tried

Cross-reference from x-open-weight-parity: two Chinese open weights match US open weights on extraction and classification

Eval run across the firm's extraction and classification tasks. Qwen and DeepSeek releases match the best US open-weight within noise and trail the frontier by the same margin. Agentic task evals show a persistent gap. Borrowed from the open-weight field with its tier.

extracted claimNon-US open weights are at parity on the cheap tier and behind on the agentic tier.
Lab · x-open-weight-parity (cross-field) · Mei Tanaka27 Aug 2026
detector · bleeding edge
14–22 pts
gap
Benchmark·band 2Signal

Agentic long-horizon benchmark: US frontier leads non-US by 14–22 points

Multi-step tool-use benchmark. US frontier models lead every non-US entry by a wide margin and the gap has been stable for three releases. The counter-evidence the field carries deliberately.

extracted claimThe frontier gap on agentic tasks has not closed.
Public leaderboard30 Jul 2026
MTdropped 2
Talk·band 3Signal

Panel: 'Sovereign AI for the Indo-Pacific'

Demand-band signal. Every panellist said sovereignty; none said whose weights. The absence is the signal.

AU government industry forum9 Jul 2026
detector · demand
Paper·band 1Signal

Sleeper Weights: Detecting Training-Time Backdoors in Open Language Models Is Currently Infeasible

Shows that trigger-conditioned behaviours can be trained into open weights and survive fine-tuning, and that current detection methods do not find them. Applies to any open weight regardless of origin; makes provenance an assurance question rather than a political one.

arxiv.org · Kowalski, Tran et al.26 Jun 2026
LFdropped 2
3 of 5
frameworks US-default
Drop·band 3Signal

Slack drop: three of five client procurement frameworks name US providers as default-approved

A sector owner's review of procurement documents across engagements. No assessment path for a non-US model in three of five. The procurement-gate claim rests on this.

Slack drop3 Jun 2026
CDdropped
0.5× US-open
hosted price
Release·band 1Signal

Alibaba releases Qwen open weights with an Apache licence and an AU-region hosted endpoint

Open weights with a permissive licence, hosted in Sydney by a domestic provider at about half the price of the US-open equivalent. The hosting is the news: a non-US model with AU residency.

extracted claimNon-US open weights are available hosted in AU regions at a cost advantage.
Alibaba Cloud21 May 2026
PRAWdropped 3
Post·band 2Signal

'Follow the megawatts: where inference is actually being built in 2026'

Maps announced data-centre capacity by jurisdiction. Non-US capacity share rising; AU share small but growing. Used as the demand-side echo of the export-control signal.

Substack · A well-followed infra voice5 May 2026
detector · early adoption
Client question·band 3Signal

'Would you ever recommend a Chinese model to us? Be honest.'

Asked by a deputy secretary in a sovereignty workshop. The honest answer was 'for some tasks, with provenance disclosed, hosted here' and the room went quiet. This is the adoption gate in one question.

Engel · government engagement15 Apr 2026
ABRMdropped 2
Regulatory·band 3Signal

US Commerce expands accelerator export controls; Singapore and UAE inference capacity announcements follow within a quarter

Control expansion followed by capacity build-outs in jurisdictions AU data could be processed in. Geopolitical signal; what matters for the field is where cheap inference ends up.

BIS / press2 Apr 2026
detector · demand
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Chinese open-weight models are at parity with US open weights on extraction and classification tasks and cost less to run hosted.

Assessedc-us-non-dominance-1dalton-0.427 Aug 2026Alibaba Cloud, Lab · x-open-weight-parity (cross-field)
76%

AU enterprise procurement frameworks default to US providers and have no assessment path for a non-US model; the gate is procurement, not capability.

Assessedc-us-non-dominance-3dalton-0.417 Jun 2026Engel · government engagement, Slack drop
72%

Provenance risk in open weights — training data, backdoors, licence terms — is unassessable with current tooling and is the real objection behind 'not Chinese'.

Assessedc-us-non-dominance-5dalton-0.427 Aug 2026arxiv.org, Engel · government engagement
60%

Export controls on accelerators are moving inference capacity toward Southeast Asia and the Gulf, which changes where AU data can cheaply be processed.

Assessedc-us-non-dominance-4dalton-0.417 Jun 2026BIS / press, Substack
55%

The frontier capability gap on agentic and long-horizon tasks has not closed and non-US models are not competitive there.

Assessedc-us-non-dominance-2dalton-0.427 Aug 2026Public leaderboard, Lab · x-open-weight-parity (cross-field)
70%

Position history · the diff is the product

3 validation runs against a fixed brief. Confidence 40% → 50%.

runs compare claim sets, never prose
What we said · run 3

Parity confirmed on extraction and classification; frontier gap confirmed on agentic tasks. Provenance is the real objection. Position drafted for the sponsor's decision on whether to publish.

50%
Changed since run 2
  • Chinese open-weight models are at parity with US open weights on extraction and classification tasks and cost less to run hosted.
  • The frontier capability gap on agentic and long-horizon tasks has not closed and non-US models are not competitive there.
  • Provenance risk in open weights — training data, backdoors, licence terms — is unassessable with current tooling and is the real objection behind 'not Chinese'.
  • c-us-non-dominance-4 ↓ 0.62 → 0.55
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · RM
high

Touches every model recommendation and the firm's public neutrality.

Timeline

committed · AW
18mo–4yr

The models are here; the procurement change is years. Conviction that it comes sooner.

Cost of being wrong

committed · LF
high

A wrong public position either loses the cheap tier or invites a provenance incident.

Demand

committed · AB
medium

Government asks about sovereignty constantly and about non-US models never; the question is latent.

TAM

agent-estimated
not-measurable-here

Not a market; a condition on every market. Agent-estimated as unmeasurable.

Cost of entry

agent-estimated
low

A position paper and a provenance checklist. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Government
relevant

Sovereignty is the stated priority and the frameworks default to US hyperscalers; the contradiction will surface in the next procurement cycle.

Mechanism · Provenance-disclosed model selection with AU hosting as the sovereignty control; position shared with the DTA on request.

AB committed by Aisha Bellocommitted
Banking
watch

No bank will run a Chinese open-weight on customer data this year; the question is whether one will run it on non-customer workloads and what APRA says when asked.

Mechanism · Non-customer workloads only, with the provenance checklist, until a regulator view exists.

CD committed by Claire Duboiscommitted
Cross-sector
relevant

Every model recommendation the firm makes is exposed to the provenance question.

Mechanism · Provenance line on every recommendation in the library; agent draft, awaiting a human.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against: this is a conviction dressed as a trend. Parity on cheap tasks is not non-dominance; the tasks that decide platform choice are agentic and the US lead there is stable or widening. Procurement defaults to US providers for reasons — legal recourse, alliance alignment, supply-chain assurance — that are not reflexes and will not move because a Qwen model is cheaper. And the AU government's own sovereignty posture is explicitly aligned with the US, which makes the political cost of a non-US recommendation higher than the field admits.

  • The parity evidence is on the tier where model choice matters least. On the tier that decides platforms, the field's own claim says the gap has not closed.
  • Procurement defaults encode alliance and legal-recourse judgements that a cost eval does not address. Calling it a reflex is the field's weakest move.
  • A provenance incident — a backdoor, a licence change, a data-training scandal — in one Chinese open-weight would set the field back years, and the base rate for such incidents is not zero.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • ML research20%
  • Lab / vendor releases20%
  • Geopolitics / policy20%
  • Infra / market10%
  • Internal / Engel30%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsAWPRMTABLF

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01What would a bank's risk function need to see to approve a non-US open weight on non-customer data, and who writes it?
  2. 02Does the cheap-tier parity extend to any agentic task class in the next two releases?
  3. 03Is provenance assurance a service the firm could sell, or a liability it should avoid?