cavendish
Do thisTestedstrength · moderate

Open-weight models for classification and extraction; frontier for agentic loops

Default classification and structured extraction to an open-weight 30B-class model behind the gateway. Keep multi-step agentic loops on frontier until the open-weight tool-use gap closes.

Tier is not strength

Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.

Evidence tierTested
Strengthmoderate
OwnerPRPriya Raman
Last validated27 Aug 2026
Review by25 Nov 2026
Half-life90 days
Citations47
VerticalsCross-sector, Banking, Retail & FMCG
decay83d until review

Machine-readable target

{
  "model": "qwen3.5-32b",
  "taskType": "structured-extraction",
  "configKey": "gateway.routing.tiers.extraction"
}

Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.

Body

x-open-weight-parity ran the current open-weight tier (Qwen 3.5 32B, Llama 4 Maverick, DeepSeek V4, Gemma 4 27B) against Claude Sonnet 5 and GPT-5.5 on our own task evals, not public leaderboards. On classification and single-pass structured extraction the best open-weight model was within 1.5 points of frontier on five of six tasks at roughly one-sixth of the per-token cost (c-open-weight-models-1). On agentic loops with tool use over more than four steps the gap was 11–18 points and did not close with prompting (c-open-weight-models-2).

This supersedes r-slm-edge. That recommendation was written when the choice was frontier or a fine-tuned 7B, and it said: measure the frontier baseline first. The baseline is now measured, and the answer that fell out is that a stock 30B-class open-weight model does the classification job without a fine-tune at all. g-fine-tune-domain-slm records that the fine-tuned 7B never beat the stock 32B on banking classification once the eval was held constant.

The 'moderate' strength is honest about two things. The open-weight tier moves monthly, and parity on our evals in August is not parity in November. And the cost advantage narrows when the AU-region hosting premium is included; for low volume, the frontier mid-tier through the gateway is cheaper in engineering time than standing up an open-weight host (c-open-weight-models-3).

What to do: set the gateway's extraction and classification tiers to the open-weight default; pin the frontier for anything with a tool loop; re-run parity on the eval set when either side ships a new version.

PRSigned Priya Raman · Research engineer · inference · 27 Aug 2026

What it rests on

Open-weight models are within 2 points of frontier on classification, structured extraction and short-form generation on our task evals.

Tested c-open-weight-models-1
85%

The open-weight gap on multi-step agentic tasks is 14–22 points and re-opens with each frontier release, because open-weight gains are mostly distilled from frontier outputs.

Tested c-open-weight-models-2
78%

In government and health, on-shore residency is the buying reason for open weights and capability parity is the permission; the order matters for how the pitch is written.

Assessed c-open-weight-models-3
72%

Field

Open Weight Models

Experiment · superseded

Open-weight parity on our task evals

Supersedes

Distil to a small model only after the frontier baseline is measured on the same eval Prior version visible. Everyone who cited it was notified.