cavendish
TestedValidatinggate · ReliabilityNow · 0–12 months×6 sightings

Open Weight Models

Open-weight models have reached frontier parity on classification, extraction and short-form generation and have not on long agentic loops; in AU regulated sectors, sovereignty and licence terms decide where they run, and capability parity is merely the permission.

Experiment run, measured result. The only tier that becomes a recommendation.

Join with…

Confidence

74%human-committed

Expiry

6doverdue for review

Lead time

+4moahead of mainstream awareness

Ownership

MTMei Tanakafortnightly cadence

Where it is

The gap closed on some task classes and stayed open on others, and it is now measurable rather than argued. Our parity run put Llama, Qwen, Gemma, Mistral and DeepSeek weights against frontier models on six of our own evals: within two points on classification, extraction and short generation; 14–22 points behind on multi-step agentic tasks, and that gap re-opens with every frontier release because open-weight gains are mostly distilled from frontier outputs. Government and health clients are asking for on-shore inference before they ask which model; the model's country of origin is becoming a procurement question in the public sector regardless of where it is hosted. Licence terms sort the models faster than benchmarks do.

Why a Quantium decision hinges on it

Half of the classification and extraction work we deliver could run on open weights at a tenth of the cost, on-shore, tomorrow. The other half — agentic loops — cannot, and clients who hear 'open source has caught up' will assume it applies to both. Getting the task-class boundary right is the single most reusable answer the lab holds this quarter; getting the sovereignty and licence answer right is what lets government and health engagements proceed at all.

Field attributes

StateValidating
GateReliability · possible, not yet dependable enough
OriginSignal
Measurablefull
Audience · TLPexec
Horizonnow
Opened22 Sep 2025
Mainstream28 Jan 2026
Last validated5 Aug 2026
Sightings6

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Parity within 2 points on classification, structured extraction and short-form generation across six of our own evals (x-open-weight-parity). The best open-weight model on extraction was a Qwen family member at roughly one-ninth the per-task cost.
  • 02A 14–22 point gap on multi-step agentic tasks, consistent across all five open-weight families, widening on the two tasks that need tool-call recovery.
  • 03The extraction step of a live health pattern moved to an open-weight model on vLLM with identical F1, logged from a session and later confirmed on the parity eval.
What is hype
  • 01'Open source has caught up.' On aggregate leaderboards, nearly; on the agentic suites clients care about, no, and the gap re-opens each frontier release.
  • 02Open weights as automatically sovereign. Weights are portable; the hosting, the fine-tune data and the operator are what residency rules look at.
  • 03Licence-free. The Llama community licence has use restrictions and an acceptable-use policy that a bank's legal team will read; Apache and MIT weights are a different conversation.
What would have to be true
  • 01Open-weight agentic performance holding within 5 points of frontier for two consecutive frontier releases — nobody has shown this.
  • 02A public-sector procurement position on model origin, so that Chinese-lab weights hosted on-shore are either allowed or not, rather than argued per agency.
  • 03An on-prem or sovereign serving stack that a delivery team can operate, which is the on-prem-inference field's problem and not yet solved.
What we would do
  • 01Ship r-open-weight-tier as the default: open weights for classification and extraction, frontier for agentic loops, decided per task class by eval.
  • 02Re-run the parity eval within four weeks of every frontier release; the gap measurement is only useful while it is current.
  • 03Maintain sa-which-model-extraction as a standing answer with a 60-day half-life and the licence table attached.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
14–22 pts
agentic gap
Finding·band 1Tested

Parity run: open weights within 2 points on extraction and classification, 14–22 behind on agentic loops

Five open-weight families against two frontier models on six lab evals. Parity on classification, structured extraction and short generation; a consistent 14–22 point gap on multi-step agentic tasks, widest where tool-call recovery is needed. Best cost ratio on extraction: roughly 9× cheaper per task.

extracted claimOpen-weight parity is real on three task classes and absent on agentic loops.
Lab · x-open-weight-parity · Mei Tanaka5 Aug 2026
detector · bleeding edge
Paper·band 1Signal

Where Do Open-Weight Gains Come From? Attributing Capability to Distillation

Attributes most of the 2025–26 open-weight improvement on reasoning and agentic suites to distillation from frontier outputs, and shows the gap re-opening within eight weeks of each frontier release. Explains why the parity boundary moves but does not close.

arxiv.org · Ferreira, Lindqvist et al.21 Jul 2026
detector · bleeding edge 2
>15 pts
agentic gap
Benchmark·band 2Signal

Independent leaderboard: open-weight aggregate gap under 5 points; agentic-suite gap over 15

The public number behind both the 'caught up' story and its rebuttal, depending on which tab you read. The split by suite matches our own evals; we cite the agentic tab, everyone else cites the aggregate.

extracted claimAggregate parity hides a persistent agentic gap.
Independent eval leaderboard30 Jun 2026
MTdropped 3
Analyst·band 3Signal

Analyst note: Chinese-origin model weights face procurement scrutiny in AU and Five Eyes public sectors

Reports informal guidance in three jurisdictions discouraging Chinese-lab weights in government workloads regardless of hosting. No published policy in Australia. Carried at moderate authority because the evidence is interviews, not documents.

Analyst brief24 Jun 2026
detector · demand 2
÷9
cost per task
Finding·band 1Tried

Logged from Claude Code: extraction step moved to Qwen on vLLM, same F1, one-ninth the cost

Logged from a session on a health pattern: swapped the frontier extraction call for an open-weight model on a vLLM endpoint inside the client tenancy. F1 unchanged on the client's held-out set. Tried tier; later confirmed on the parity eval.

MCP · log_finding · Priya Raman17 Jun 2026
PRdropped
Regulatory·band 1Signal

DTA guidance note: hosting location and model provenance for government AI procurement

Guidance for agencies on documenting where a model is hosted and who trained it. Does not prohibit any origin; requires it to be recorded and risk-assessed. In practice agencies are reading it as a reason to avoid the question.

Digital Transformation Agency13 May 2026
ABdropped 2
2.4M
HF downloads, 30d
Release·band 1Signal

Alibaba ships a new Qwen open-weight family under Apache 2.0, 0.6B to 200B+

Full family with a permissive licence, including a mixture-of-experts flagship. The Apache 2.0 licence cleared one client's legal review in three days; a community-licensed competitor took five weeks on the same engagement.

extracted claimPermissive-licence open weights clear enterprise legal review an order of magnitude faster.
Alibaba / Hugging Face29 Apr 2026
PRdropped 3
Client question·band 3Signal

'Can we run this on-shore without sending anything to a US lab?'

Asked in the same form in a state health department and a federal agency in the same fortnight. Both wanted the residency answer before any capability discussion. Logged as one demand signal with two sightings; reframed the field from cost to sovereignty.

Engel · government and health engagements10 Mar 2026
HNABdropped 2
Release·band 1Signal

DeepSeek releases open-weight reasoning model under MIT licence

The release that made 'open source has caught up' a mainstream claim. Strong on reasoning benchmarks; on our agentic evals it sat with the rest of the open-weight pack. Origin question raised immediately in two government engagements.

DeepSeek / Hugging Face22 Jan 2026
detector · bleeding edge 5
Post·band 2Signal

'Llama's licence is not open source and your lawyers will notice'

Walks through the community licence's use restrictions and the acceptable-use policy clause by clause. Became the document we send legal teams first. Widely shared; contested by the vendor's developer advocates.

Personal blog · A well-followed open-source licensing voice9 Dec 2025
detector · early adoption 3
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 5 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Open-weight models are within 2 points of frontier on classification, structured extraction and short-form generation on our task evals.

Testedc-open-weight-models-1dalton-0.45 Aug 2026Lab · x-open-weight-parity, MCP · log_finding
85%

The open-weight gap on multi-step agentic tasks is 14–22 points and re-opens with each frontier release, because open-weight gains are mostly distilled from frontier outputs.

Testedc-open-weight-models-2dalton-0.45 Aug 2026Lab · x-open-weight-parity, arxiv.org, Independent eval leaderboard
78%

In government and health, on-shore residency is the buying reason for open weights and capability parity is the permission; the order matters for how the pitch is written.

Assessedc-open-weight-models-3dalton-0.321 May 2026Engel · government and health engagements, Digital Transformation Agency
72%

Licence terms, not benchmarks, are the first filter a client's legal team applies; Apache/MIT weights clear in days, community licences take weeks.

Assessedc-open-weight-models-6dalton-0.321 May 2026Personal blog, Alibaba / Hugging Face
66%

Model origin (a Chinese lab) is a procurement blocker in the AU public sector regardless of hosting location.

Assessedc-open-weight-models-5dalton-0.42 Jul 2026Analyst brief, Digital Transformation Agency
55%

Open-weight models will reach agentic parity within twelve months, making a frontier tier unnecessary.

Assessedc-open-weight-models-4dalton-0.312 Feb 2026DeepSeek / Hugging Face, Independent eval leaderboard
27%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 48% → 74%.

runs compare claim sets, never prose
What we said · run 4

Parity experiment concluded: within 2 points on three task classes, 14–22 behind on agentic, gap re-opens each frontier release. Origin-as-blocker carried at moderate confidence. Recommendation and standing answer published.

74%
Changed since run 3
  • Open-weight models are within 2 points of frontier on classification, structured extraction and short-form generation on our task evals.
  • The open-weight gap on multi-step agentic tasks is 14–22 points and re-opens with each frontier release, because open-weight gains are mostly distilled from frontier outputs.
  • Model origin (a Chinese lab) is a procurement blocker in the AU public sector regardless of hosting location.
  • c-open-weight-models-3 ↑ 0.62 → 0.72
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · MT
high

Changes the model choice on roughly half of delivered task volume.

Timeline

committed · MT
0–18mo

Parity on the relevant classes is measured, not forecast.

TAM

agent-estimated
>$10B

Agent-estimated from global inference spend that could shift to open weights. Uncommitted and probably too broad.

Cost

committed · PR
low

Re-running the parity eval is a day; the serving stack is another field's cost.

Demand

committed · AB
high

Every government engagement this year asked about on-shore inference; three named specific open-weight families.

Cost of being wrong

agent-estimated
medium

Wrong task-class boundary is a rework, not an incident; wrong licence read is a contract problem. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Government
relevant

Residency and provenance questions arrive before capability questions. Agencies want a defensible answer on where the weights came from and where they run.

Mechanism · On-shore serving of Apache/MIT-licensed weights for extraction and classification; frontier via a sovereign-region endpoint for agentic work.

AB committed by Aisha Bellocommitted
Health
relevant

Clinical and claims text cannot leave the jurisdiction in several of our engagements; open weights on-shore are the only path for those workloads.

Mechanism · Extraction on open weights inside the client's tenancy; the parity eval is the assurance artifact.

Agent draft · awaiting a sector owneragent-estimated
Banking
watch

APRA does not care about weight origin; cost does. Banks will move extraction to open weights when the gateway can route to them, not before.

Mechanism · Task-class routing through the gateway once the on-prem leg exists.

CD committed by Claire Duboiscommitted
Retail & FMCG
not-relevant

Retail clients run on frontier models through hyperscaler commitments they have already paid for. Sovereignty is not a buying reason and the cost delta is inside their committed spend.

Mechanism · None until committed-spend contracts roll over.

DS committed by Dev Sharmacommitted

Red team · the strongest case against

The strongest case against: the parity finding is a snapshot of a moving target measured with our own evals, and the frontier labs' own pricing collapse removes most of the cost argument for open weights outside sovereignty cases. We may be recommending a serving stack clients must operate, to save money the frontier labs are about to give away.

  • Frontier list prices fell roughly 4× in twelve months. At that rate the cost case for self-hosted open weights on extraction narrows to residency-driven workloads only, which is a smaller field than the one we have scored.
  • Six evals are ours. A 2-point parity on our extraction tasks says nothing about a client's messy PDFs; we have one tried-tier confirmation on a live pattern, not a tested one.
  • Origin-as-blocker is carried at 0.55 on analyst and anecdotal evidence. If it is wrong, half the licence-and-provenance advice is noise; if it is right, the best-performing open weights are off the table for the sector that most wants them.
  • Every open-weight family we tested is under twelve months old. The recommendation names models that will be superseded before the review date.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis holds

Source diversity

  • ML research30%
  • Model labs / releases25%
  • Regulatory / analyst20%
  • Open-source10%
  • Internal / Engel15%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsMTPRABHNLF

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Does the 2-point parity on our extraction evals hold on a client's scanned PDFs, or only on clean text?
  2. 02Will the DTA guidance harden into a published origin policy, and in which direction?
  3. 03At what frontier list price does the cost case for self-hosted open weights disappear outside residency-driven workloads?