Which model for structured extraction?
Qwen 3.5 32B through the gateway's extraction tier for flat and moderately nested schemas; Claude Sonnet 5 for schemas over ~40 fields or with cross-field constraints. As of 27 August.
Tier is not strength
Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.
Machine-readable target
{
"model": "qwen3.5-32b",
"taskType": "structured-extraction",
"configKey": "gateway.routing.tiers.extraction"
}Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.
Body
As of 27 August 2026: use the gateway's extraction tier, which resolves to Qwen 3.5 32B in the AU region. For schemas over roughly 40 fields, deeply nested output, or cross-field validation rules, route to Claude Sonnet 5. Do not fine-tune for extraction without a measured gap on your own eval.
Evidence: x-open-weight-parity, re-run 27 August on the current open-weight tier. On five of six extraction tasks the open-weight default was within 1.5 points of frontier at about one-sixth the token cost (c-open-weight-models-1). The sixth task — a 60-field insurance claim schema with conditional fields — is where the frontier lead was 9 points and held across three runs (c-open-weight-models-3).
Caveat: this answer moves monthly. Qwen 3.5 displaced Llama 4 Maverick as the default in July; a new DeepSeek release is on the eval queue now. Check the gateway config rather than this page if you are reading it after the review date.
What it rests on
Open-weight models are within 2 points of frontier on classification, structured extraction and short-form generation on our task evals.
In government and health, on-shore residency is the buying reason for open weights and capability parity is the permission; the order matters for how the pitch is written.
Field
Open Weight ModelsExperiment · superseded
Open-weight parity on our task evals