Judge calibration against human panel
Eval Harnesses- proposed
- voting
- running
- measuring
- concluded
Preregistration · v1 · 16 Mar 2026 · immutable after start
Hypothesis
An uncalibrated frontier LLM judge, as used in the harness through Q1, agrees with a five-person human panel at Cohen's κ of at least 0.8 across extraction, summarisation and agentic-trace grading.
Kill condition
Kill if κ is below 0.6 on any task family after 200 paired items in that family.
Result
RefutedRefuted. The uncalibrated judge reached κ 0.58 / 0.49 / 0.41 against the panel, with 0.87 self-consistency masking it. A rubric-calibrated judge reached 0.79–0.84 on the same items. The refuted hypothesis is that an uncalibrated judge agrees with humans; the recommendation r-judge-calibration came out of the second arm. Graveyard: g-llm-judge-unchecked.
Running notes · fed from harness sessions and by the pair
- manual8 May 2026MT Mei Tanaka
Concluded, refuted. Two Q1 release decisions would have reversed under the panel. Recommendation r-judge-calibration: every judge ships with a κ or does not ship.
- harness30 Apr 2026MT Mei Tanaka
rubric-calibrated arm: κ 0.84 / 0.79 / 0.81 on same items
- manual22 Apr 2026LF Lena Fischer
κ: extraction 0.58, summarisation 0.49, trace 0.41. Kill condition fired on two families. Trace grading: judge rewards length and tool-call count; panel penalises both.
- harness8 Apr 2026MT Mei Tanaka
panel: 5 graders, 600 items, majority labels computed; judge run x3; self-consistency 0.87
- manual16 Mar 2026MT Mei Tanaka
Predicted κ 0.7 on extraction, lower elsewhere, confidence 0.5. If the kill fires on trace grading I will not be surprised.
Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.
Artifacts