cavendish
QueueExperimentsScorecard
Type 1SignalSuperseded

Open-weight parity on our task evals

Open Weight Models
  1. proposed
  2. voting
  3. running
  4. measuring
  5. concluded

Preregistration · v1 · 7 Apr 2026 · immutable after start

An extension creates a new version rather than editing the old.

Hypothesis

At least two current open-weight models (Qwen, DeepSeek, Llama, Mistral) come within 3 points of the frontier default on the extraction and classification eval families.

Kill condition

Type 1: no kill condition required. Freshness date 2026-05-07.

MethodHalf a day. Run four open-weight models through the extraction and classification families of the harness with the standard prompts; report the gap to frontier per family.
Expected cost$120 in tokens, half a day
Expected durationHalf a day

Result

Superseded

Superseded by a vendor benchmark published 2026-04-14 that covered nine open-weight models with a public harness; our four-model numbers agreed with it within a point. The question was answered elsewhere and the lab adopted the public harness.

Running notes · fed from harness sessions and by the pair

AW
manual · Adam Witanowski · 3 Sep 2026· ⌘↩ to post
  1. manual17 Apr 2026MT Mei Tanaka

    Superseded. Adopted the vendor harness; re-ran on our evals in an afternoon to confirm. r-open-weight-tier and sa-which-model-extraction draw on that, not this run.

  2. manual14 Apr 2026MT Mei Tanaka

    A vendor published a benchmark this morning covering nine open-weight models across the same two families plus agentic, with a public harness. Their numbers on our four match mine within a point.

  3. harness7 Apr 2026MT Mei Tanaka

    4 models x 2 families run; qwen -2.1, deepseek -2.8, llama -6.4, mistral -5.0 on extraction; classification all within 3

Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.

Artifacts

measFour-model gap-to-frontier, two families (half-day run)measurement