Open-weight parity on our task evals
Open Weight Models- proposed
- voting
- running
- measuring
- concluded
Preregistration · v1 · 7 Apr 2026 · immutable after start
Hypothesis
At least two current open-weight models (Qwen, DeepSeek, Llama, Mistral) come within 3 points of the frontier default on the extraction and classification eval families.
Kill condition
Type 1: no kill condition required. Freshness date 2026-05-07.
Result
SupersededSuperseded by a vendor benchmark published 2026-04-14 that covered nine open-weight models with a public harness; our four-model numbers agreed with it within a point. The question was answered elsewhere and the lab adopted the public harness.
Running notes · fed from harness sessions and by the pair
- manual17 Apr 2026MT Mei Tanaka
Superseded. Adopted the vendor harness; re-ran on our evals in an afternoon to confirm. r-open-weight-tier and sa-which-model-extraction draw on that, not this run.
- manual14 Apr 2026MT Mei Tanaka
A vendor published a benchmark this morning covering nine open-weight models across the same two families plus agentic, with a public harness. Their numbers on our four match mine within a point.
- harness7 Apr 2026MT Mei Tanaka
4 models x 2 families run; qwen -2.1, deepseek -2.8, llama -6.4, mistral -5.0 on extraction; classification all within 3
Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.
Artifacts