Open-weight models for classification and extraction; frontier for agentic loops
Default classification and structured extraction to an open-weight 30B-class model behind the gateway. Keep multi-step agentic loops on frontier until the open-weight tool-use gap closes.
Tier is not strength
Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.
Machine-readable target
{
"model": "qwen3.5-32b",
"taskType": "structured-extraction",
"configKey": "gateway.routing.tiers.extraction"
}Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.
Body
x-open-weight-parity ran the current open-weight tier (Qwen 3.5 32B, Llama 4 Maverick, DeepSeek V4, Gemma 4 27B) against Claude Sonnet 5 and GPT-5.5 on our own task evals, not public leaderboards. On classification and single-pass structured extraction the best open-weight model was within 1.5 points of frontier on five of six tasks at roughly one-sixth of the per-token cost (c-open-weight-models-1). On agentic loops with tool use over more than four steps the gap was 11–18 points and did not close with prompting (c-open-weight-models-2).
This supersedes r-slm-edge. That recommendation was written when the choice was frontier or a fine-tuned 7B, and it said: measure the frontier baseline first. The baseline is now measured, and the answer that fell out is that a stock 30B-class open-weight model does the classification job without a fine-tune at all. g-fine-tune-domain-slm records that the fine-tuned 7B never beat the stock 32B on banking classification once the eval was held constant.
The 'moderate' strength is honest about two things. The open-weight tier moves monthly, and parity on our evals in August is not parity in November. And the cost advantage narrows when the AU-region hosting premium is included; for low volume, the frontier mid-tier through the gateway is cheaper in engineering time than standing up an open-weight host (c-open-weight-models-3).
What to do: set the gateway's extraction and classification tiers to the open-weight default; pin the frontier for anything with a tool loop; re-run parity on the eval set when either side ships a new version.
What it rests on
Open-weight models are within 2 points of frontier on classification, structured extraction and short-form generation on our task evals.
The open-weight gap on multi-step agentic tasks is 14–22 points and re-opens with each frontier release, because open-weight gains are mostly distilled from frontier outputs.
In government and health, on-shore residency is the buying reason for open weights and capability parity is the permission; the order matters for how the pitch is written.
Field
Open Weight ModelsExperiment · superseded
Open-weight parity on our task evalsSupersedes
Distil to a small model only after the frontier baseline is measured on the same eval Prior version visible. Everyone who cited it was notified.