cavendish
QueueExperimentsScorecard
Type 3Pressure

Release canary suite v2

Eval Harnesses
  1. proposed
  2. voting
  3. running
  4. measuring · 7d
  5. concluded

Preregistration · v1 · 27 Jul 2026 · immutable after start

An extension creates a new version rather than editing the old.

Hypothesis

A 120-item canary suite with calibrated judges catches at least 80% of the behaviour regressions that the last six vendor model releases introduced on our patterns, at under 15 minutes wall-clock per run.

Kill condition

Kill if the suite catches fewer than 50% of the 22 known regressions on replay by day 10, or if wall-clock exceeds 30 minutes.

MethodAssemble 120 items from the calibration set and the pattern evals; replay six historical model releases against the suite and score recall of the 22 regressions logged at the time. Judges from x-judge-calibration with κ attached.
Expected cost$700 in replay tokens, two people for three weeks
Expected duration3 weeks

Running notes · fed from harness sessions and by the pair

AW
manual · Adam Witanowski · 3 Sep 2026· ⌘↩ to post
  1. manual27 Aug 2026JP Jun Park

    Measuring. Raw recall 18/22 (82%). Wall-clock 11 min. The four misses are all tone or formatting, as predicted. Checking whether any pass is judge-flattered.

  2. harness11 Aug 2026MT Mei Tanaka

    suite: 120 items; 6 releases replayed; 22 known regressions loaded

  3. manual27 Jul 2026MT Mei Tanaka

    Predicted 75% recall, confidence 0.55. The regressions that scare me are the tone shifts; those are the hardest to canary.

Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.

Artifacts

evalCanary suite v2, 120 items with judge κ per itemeval
measRegression recall on six historical releasesmeasurement