Release canary suite v2
Eval Harnesses- proposed
- voting
- running
- measuring · 7d
- concluded
Preregistration · v1 · 27 Jul 2026 · immutable after start
Hypothesis
A 120-item canary suite with calibrated judges catches at least 80% of the behaviour regressions that the last six vendor model releases introduced on our patterns, at under 15 minutes wall-clock per run.
Kill condition
Kill if the suite catches fewer than 50% of the 22 known regressions on replay by day 10, or if wall-clock exceeds 30 minutes.
Running notes · fed from harness sessions and by the pair
- manual27 Aug 2026JP Jun Park
Measuring. Raw recall 18/22 (82%). Wall-clock 11 min. The four misses are all tone or formatting, as predicted. Checking whether any pass is judge-flattered.
- harness11 Aug 2026MT Mei Tanaka
suite: 120 items; 6 releases replayed; 22 known regressions loaded
- manual27 Jul 2026MT Mei Tanaka
Predicted 75% recall, confidence 0.55. The regressions that scare me are the tone shifts; those are the hardest to canary.
Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.
Artifacts