cavendish
QueueExperimentsScorecard
Type 3ObservationRefuted

Judge calibration against human panel

Eval Harnesses
  1. proposed
  2. voting
  3. running
  4. measuring
  5. concluded

Preregistration · v1 · 16 Mar 2026 · immutable after start

An extension creates a new version rather than editing the old.

Hypothesis

An uncalibrated frontier LLM judge, as used in the harness through Q1, agrees with a five-person human panel at Cohen's κ of at least 0.8 across extraction, summarisation and agentic-trace grading.

Kill condition

Kill if κ is below 0.6 on any task family after 200 paired items in that family.

Method600 items, 200 per family, each graded blind by five lab and firm people and by the production judge prompt. Majority label per item. κ per family, judge self-consistency across three re-runs, and a length/tool-call bias check. Second arm: a rubric-calibrated judge on the same items.
Expected cost$900 in tokens, 40 person-hours of panel grading, two people for three weeks
Expected duration3 weeks plus two weeks of panel scheduling

Result

Refuted

Refuted. The uncalibrated judge reached κ 0.58 / 0.49 / 0.41 against the panel, with 0.87 self-consistency masking it. A rubric-calibrated judge reached 0.79–0.84 on the same items. The refuted hypothesis is that an uncalibrated judge agrees with humans; the recommendation r-judge-calibration came out of the second arm. Graveyard: g-llm-judge-unchecked.

Running notes · fed from harness sessions and by the pair

AW
manual · Adam Witanowski · 3 Sep 2026· ⌘↩ to post
  1. manual8 May 2026MT Mei Tanaka

    Concluded, refuted. Two Q1 release decisions would have reversed under the panel. Recommendation r-judge-calibration: every judge ships with a κ or does not ship.

  2. harness30 Apr 2026MT Mei Tanaka

    rubric-calibrated arm: κ 0.84 / 0.79 / 0.81 on same items

  3. manual22 Apr 2026LF Lena Fischer

    κ: extraction 0.58, summarisation 0.49, trace 0.41. Kill condition fired on two families. Trace grading: judge rewards length and tool-call count; panel penalises both.

  4. harness8 Apr 2026MT Mei Tanaka

    panel: 5 graders, 600 items, majority labels computed; judge run x3; self-consistency 0.87

  5. manual16 Mar 2026MT Mei Tanaka

    Predicted κ 0.7 on extraction, lower elsewhere, confidence 0.5. If the kill fires on trace grading I will not be surprised.

Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.

Artifacts

eval600-item panel-labelled calibration set (three families)eval
measκ by family, uncalibrated vs rubric-calibrated judgemeasurement
noteLength and tool-call bias analysisnotebook