Open Weight Models
Open-weight models have reached frontier parity on classification, extraction and short-form generation and have not on long agentic loops; in AU regulated sectors, sovereignty and licence terms decide where they run, and capability parity is merely the permission.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
74%human-committedExpiry
6doverdue for reviewLead time
+4moahead of mainstream awarenessOwnership
MTMei Tanakafortnightly cadenceWhere it is
The gap closed on some task classes and stayed open on others, and it is now measurable rather than argued. Our parity run put Llama, Qwen, Gemma, Mistral and DeepSeek weights against frontier models on six of our own evals: within two points on classification, extraction and short generation; 14–22 points behind on multi-step agentic tasks, and that gap re-opens with every frontier release because open-weight gains are mostly distilled from frontier outputs. Government and health clients are asking for on-shore inference before they ask which model; the model's country of origin is becoming a procurement question in the public sector regardless of where it is hosted. Licence terms sort the models faster than benchmarks do.
Why a Quantium decision hinges on it
Half of the classification and extraction work we deliver could run on open weights at a tenth of the cost, on-shore, tomorrow. The other half — agentic loops — cannot, and clients who hear 'open source has caught up' will assume it applies to both. Getting the task-class boundary right is the single most reusable answer the lab holds this quarter; getting the sovereignty and licence answer right is what lets government and health engagements proceed at all.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Parity within 2 points on classification, structured extraction and short-form generation across six of our own evals (x-open-weight-parity). The best open-weight model on extraction was a Qwen family member at roughly one-ninth the per-task cost.
- 02A 14–22 point gap on multi-step agentic tasks, consistent across all five open-weight families, widening on the two tasks that need tool-call recovery.
- 03The extraction step of a live health pattern moved to an open-weight model on vLLM with identical F1, logged from a session and later confirmed on the parity eval.
- 01'Open source has caught up.' On aggregate leaderboards, nearly; on the agentic suites clients care about, no, and the gap re-opens each frontier release.
- 02Open weights as automatically sovereign. Weights are portable; the hosting, the fine-tune data and the operator are what residency rules look at.
- 03Licence-free. The Llama community licence has use restrictions and an acceptable-use policy that a bank's legal team will read; Apache and MIT weights are a different conversation.
- 01Open-weight agentic performance holding within 5 points of frontier for two consecutive frontier releases — nobody has shown this.
- 02A public-sector procurement position on model origin, so that Chinese-lab weights hosted on-shore are either allowed or not, rather than argued per agency.
- 03An on-prem or sovereign serving stack that a delivery team can operate, which is the on-prem-inference field's problem and not yet solved.
- 01Ship r-open-weight-tier as the default: open weights for classification and extraction, frontier for agentic loops, decided per task class by eval.
- 02Re-run the parity eval within four weeks of every frontier release; the gap measurement is only useful while it is current.
- 03Maintain sa-which-model-extraction as a standing answer with a 60-day half-life and the licence table attached.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Parity run: open weights within 2 points on extraction and classification, 14–22 behind on agentic loops
Five open-weight families against two frontier models on six lab evals. Parity on classification, structured extraction and short generation; a consistent 14–22 point gap on multi-step agentic tasks, widest where tool-call recovery is needed. Best cost ratio on extraction: roughly 9× cheaper per task.
extracted claimOpen-weight parity is real on three task classes and absent on agentic loops.

Where Do Open-Weight Gains Come From? Attributing Capability to Distillation
Attributes most of the 2025–26 open-weight improvement on reasoning and agentic suites to distillation from frontier outputs, and shows the gap re-opening within eight weeks of each frontier release. Explains why the parity boundary moves but does not close.

Independent leaderboard: open-weight aggregate gap under 5 points; agentic-suite gap over 15
The public number behind both the 'caught up' story and its rebuttal, depending on which tab you read. The split by suite matches our own evals; we cite the agentic tab, everyone else cites the aggregate.
extracted claimAggregate parity hides a persistent agentic gap.

Analyst note: Chinese-origin model weights face procurement scrutiny in AU and Five Eyes public sectors
Reports informal guidance in three jurisdictions discouraging Chinese-lab weights in government workloads regardless of hosting. No published policy in Australia. Carried at moderate authority because the evidence is interviews, not documents.

Logged from Claude Code: extraction step moved to Qwen on vLLM, same F1, one-ninth the cost
Logged from a session on a health pattern: swapped the frontier extraction call for an open-weight model on a vLLM endpoint inside the client tenancy. F1 unchanged on the client's held-out set. Tried tier; later confirmed on the parity eval.

DTA guidance note: hosting location and model provenance for government AI procurement
Guidance for agencies on documenting where a model is hosted and who trained it. Does not prohibit any origin; requires it to be recorded and risk-assessed. In practice agencies are reading it as a reason to avoid the question.

Alibaba ships a new Qwen open-weight family under Apache 2.0, 0.6B to 200B+
Full family with a permissive licence, including a mixture-of-experts flagship. The Apache 2.0 licence cleared one client's legal review in three days; a community-licensed competitor took five weeks on the same engagement.
extracted claimPermissive-licence open weights clear enterprise legal review an order of magnitude faster.

'Can we run this on-shore without sending anything to a US lab?'
Asked in the same form in a state health department and a federal agency in the same fortnight. Both wanted the residency answer before any capability discussion. Logged as one demand signal with two sightings; reframed the field from cost to sovereignty.

DeepSeek releases open-weight reasoning model under MIT licence
The release that made 'open source has caught up' a mainstream claim. Strong on reasoning benchmarks; on our agentic evals it sat with the rest of the open-weight pack. Origin question raised immediately in two government engagements.

'Llama's licence is not open source and your lawyers will notice'
Walks through the community licence's use restrictions and the acceptable-use policy clause by clause. Became the document we send legal teams first. Widely shared; contested by the vendor's developer advocates.
Claims · 5 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
Open-weight models are within 2 points of frontier on classification, structured extraction and short-form generation on our task evals.
The open-weight gap on multi-step agentic tasks is 14–22 points and re-opens with each frontier release, because open-weight gains are mostly distilled from frontier outputs.
In government and health, on-shore residency is the buying reason for open weights and capability parity is the permission; the order matters for how the pitch is written.
Licence terms, not benchmarks, are the first filter a client's legal team applies; Apache/MIT weights clear in days, community licences take weeks.
Model origin (a Chinese lab) is a procurement blocker in the AU public sector regardless of hosting location.
Open-weight models will reach agentic parity within twelve months, making a frontier tier unnecessary.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 48% → 74%.
Parity experiment concluded: within 2 points on three task classes, 14–22 behind on agentic, gap re-opens each frontier release. Origin-as-blocker carried at moderate confidence. Recommendation and standing answer published.
- Open-weight models are within 2 points of frontier on classification, structured extraction and short-form generation on our task evals.
- The open-weight gap on multi-step agentic tasks is 14–22 points and re-opens with each frontier release, because open-weight gains are mostly distilled from frontier outputs.
- Model origin (a Chinese lab) is a procurement blocker in the AU public sector regardless of hosting location.
- c-open-weight-models-3 ↑ 0.62 → 0.72
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · MTChanges the model choice on roughly half of delivered task volume.
Timeline
committed · MTParity on the relevant classes is measured, not forecast.
TAM
agent-estimatedAgent-estimated from global inference spend that could shift to open weights. Uncommitted and probably too broad.
Cost
committed · PRRe-running the parity eval is a day; the serving stack is another field's cost.
Demand
committed · ABEvery government engagement this year asked about on-shore inference; three named specific open-weight families.
Cost of being wrong
agent-estimatedWrong task-class boundary is a rework, not an incident; wrong licence read is a contract problem. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Residency and provenance questions arrive before capability questions. Agencies want a defensible answer on where the weights came from and where they run.
Mechanism · On-shore serving of Apache/MIT-licensed weights for extraction and classification; frontier via a sovereign-region endpoint for agentic work.
Clinical and claims text cannot leave the jurisdiction in several of our engagements; open weights on-shore are the only path for those workloads.
Mechanism · Extraction on open weights inside the client's tenancy; the parity eval is the assurance artifact.
APRA does not care about weight origin; cost does. Banks will move extraction to open weights when the gateway can route to them, not before.
Mechanism · Task-class routing through the gateway once the on-prem leg exists.
Retail clients run on frontier models through hyperscaler commitments they have already paid for. Sovereignty is not a buying reason and the cost delta is inside their committed spend.
Mechanism · None until committed-spend contracts roll over.
Red team · the strongest case against
The strongest case against: the parity finding is a snapshot of a moving target measured with our own evals, and the frontier labs' own pricing collapse removes most of the cost argument for open weights outside sovereignty cases. We may be recommending a serving stack clients must operate, to save money the frontier labs are about to give away.
- —Frontier list prices fell roughly 4× in twelve months. At that rate the cost case for self-hosted open weights on extraction narrows to residency-driven workloads only, which is a smaller field than the one we have scored.
- —Six evals are ours. A 2-point parity on our extraction tasks says nothing about a client's messy PDFs; we have one tried-tier confirmation on a live pattern, not a tested one.
- —Origin-as-blocker is carried at 0.55 on analyst and anecdotal evidence. If it is wrong, half the licence-and-provenance advice is noise; if it is right, the best-performing open weights are off the table for the sector that most wants them.
- —Every open-weight family we tested is under twelve months old. The recommendation names models that will be superseded before the review date.
Source diversity
- ML research30%
- Model labs / releases25%
- Regulatory / analyst20%
- Open-source10%
- Internal / Engel15%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Open weights are only usable in delivery once the gateway can route to them by task class.
Open weights are the only thing on-prem inference can serve; without parity on some task class there is nothing to host.
Chinese-lab weights reaching parity is the concrete form of the sovereign-inference position.
Distilling to a small model starts from an open-weight base; the parity measurement is the baseline that field needs.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- PRPriya Raman · Research engineer · inference2 drops
- ABAisha Bello · Sector owner · Government2 drops
- MTMei Tanaka · Research lead · evals1 drop
- HNHana Novak · Sector owner · Health1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Open-weight models for classification and extraction; frontier for agentic loops
strength moderate · 47 citations · review 25 Nov 2026
Which model for structured extraction?
strength strong · 54 citations · review 24 Sep 2026
Open-weight parity on our task evals
At least two current open-weight models (Qwen, DeepSeek, Llama, Mistral) come within 3 points of the frontier default on the extraction and classification eval families.
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Does the 2-point parity on our extraction evals hold on a client's scanned PDFs, or only on clean text?
- 02Will the DTA guidance harden into a published origin policy, and in which direction?
- 03At what frontier list price does the cost case for self-hosted open weights disappear outside residency-driven workloads?