Eval Harnesses
An eval harness is five instruments, not one — capability, task, release canary, judge calibration, cost and latency — and the one that pays first is the canary; a judge without a human agreement score is not an eval, it is an opinion with a decimal point.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
84%human-committedExpiry
42duntil review · 15 Oct 2026Lead time
+4moahead of mainstream awarenessOwnership
MTMei Tanakaweekly cadenceWhere it is
Turing was built first, on conviction, before the graph existed, and it is the field that has converged. The release canary has fired on three foundation-model releases and caught a nine-point regression on structured extraction within six hours of one of them; no client pipeline caught it first. Judge calibration against a human panel put the best judge at kappa 0.71 on rubric-scored task evals and 0.38 on open-ended quality, and showed an uncalibrated judge over-scoring the newer model by eleven points — the finding that sent 'judge without a human score' to the graveyard. Public benchmarks saturate inside nine months and stop discriminating; task evals on the firm's own data are the only durable instrument. The gate is skills: the tooling is open and adequate, foundation labs now ship eval products, and the constraint is that eval engineering is a job title the market is only now creating. Competitors are hiring for it.
Why a Quantium decision hinges on it
Every recommendation the lab publishes rests on an eval; every client asks 'how do we know the new model didn't break it'. The canary is the answer to the second and the calibration score is what keeps the first honest. Evals are also the handover artefact — the thing the firm keeps if the lab stops — and the table stake most likely to be a brief moat before competitors staff up. Health clients face a software-as-medical-device question where a calibrated eval is the difference between a regulated product and an unregulated one.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Release canary v2 fired on three foundation releases and caught a 9-point regression on structured extraction within six hours; no client pipeline caught it first (x-eval-canary).
- 02Judge calibration against a twelve-person human panel: best judge kappa 0.71 on rubric-scored task evals, 0.38 on open-ended quality; an uncalibrated judge over-scored the newer same-family model by 11 points (x-judge-calibration).
- 03Rubric-with-references judging raised agreement from 0.58 to 0.71 on the same tasks; free-form judging did not improve with a stronger judge model.
- 04Public benchmarks in our watchlist saturated (>95% top score) a median of nine months after release; task evals on our own data have not.
- 01'We use LLM-as-judge.' Without a human agreement score it measures the judge's preferences, including its preference for its own family.
- 02Leaderboard positions. A model's rank on a saturated public benchmark says nothing about its behaviour on a claims-triage task.
- 03Eval platforms as a product category. The platform is the easy part; the rubric, the references and the panel are the work.
- 01Judge agreement above 0.6 on open-ended quality, which no judge configuration we have tried reaches; until then open-ended evals are human-scored.
- 02A canary that fires on gateway routing changes as well as model releases; today it fires on release announcements only.
- 03Eval engineering as a delivery skill, not a lab skill: a rubric a sector team can write and calibrate without Turing's owner in the room.
- 01Keep r-judge-calibration as a hard rule: every judge ships with a human agreement score or does not ship.
- 02Wire the canary to the gateway so it fires on any model or routing change, not just announcements; this is the Argus teeth.
- 03Move the field toward dissolution by absorption: the tooling is practice now, and the lab's remaining job is calibration and the skills gap.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Judge calibration against a human panel: kappa 0.71 on rubric tasks, 0.38 open-ended; uncalibrated judge over-scored the newer model by 11 points
Twelve-person panel scored 600 items across four task evals and two open-ended quality sets. Four judge configurations compared. Rubric-with-references reached kappa 0.71; free-form judging stayed under 0.4 regardless of judge model, and the free-form judge preferred the newer same-family model by 11 points against the panel.
extracted claimJudge reliability is a property of the rubric and references, not the judge model, and uncalibrated judges favour their own family.

'AI Evaluation Engineer' openings in Australia double in six months
Argus hiring scan across consultancies, banks and two foundation-lab AU offices. Forty-one open roles; the title barely existed in January. Inference: the skills gate is being closed by the market, and the lab's advantage on evals has a shelf life.

Release canary v2 caught a 9-point extraction regression within six hours of a model release
Canary fires on foundation-lab release announcements and runs the task-eval suite against the new default. On the third release it flagged a nine-point drop on structured extraction; the client pipeline running the same model noticed four days later from support tickets.
extracted claimA release canary on own-data task evals catches material regressions before client pipelines do.

Logged from Claude Code: switching the judge to a rubric with reference answers lifted agreement on the merchandising eval from 0.55 to 0.7
Product engineer re-ran the retail pilot's eval with reference-anchored rubrics and a small human sample. Agreement moved from 0.55 to 0.70. One eval, one person; tried tier, and consistent with the calibration experiment.


Judges prefer their own: self- and family-preference bias in LLM evaluation
Measures judge preference for outputs from the judge's own model family across five families. Finds a consistent 6–12 point preference that survives prompt-level debiasing and disappears only with reference-anchored rubrics. Our calibration reproduced it.

Watchlist benchmark saturation: median nine months from release to a >95% top score
Argus tracking of fourteen public benchmarks in the capability matrix. Median time from release to saturation was nine months; three saturated within four. The reason the capability matrix is scored on our own task evals.

'Your eval is measuring your prompt'
Argues that task evals overfit to the prompt and rubric that produced them and that a good score is a tautology. Partly right — it is why the panel exists — and kept as the sharpest critique of the field's own instrument.

'How do we know the new model version didn't break our pipeline?'
Asked by a head of model risk after a provider deprecated a model mid-engagement. Logged as unanswered at the time; it is now the canary's job description and the most-asked question in the banking pipeline.

Foundation lab ships a hosted evals product with a frontier model as default grader
Managed eval runs with the vendor's own model as the default judge and no calibration step. Convenient, and the design choice our calibration result argues against. Retained as the source of the 'judge is good enough' claim.
Claims · 5 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.
Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.
Uncalibrated judges systematically favour newer and same-family models.
Public benchmarks saturate within about nine months of release and stop discriminating; task evals on own data are the only durable instrument.
Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.
A frontier model as judge is reliable enough without human calibration for most enterprise uses.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 50% → 84%.
Both experiments concluded. Canary caught a real regression before any client did; judge calibration gives 0.71 on rubric tasks and 0.38 on open-ended. Unchecked judge to the graveyard; recommendation and standing answer current. Field converged; skills is the gate.
- LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.
- Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.
- Uncalibrated judges systematically favour newer and same-family models.
- Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.
- c-eval-harnesses-5 ↓ 0.30 → 0.18
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · MTEvery recommendation and every canary rests on it; it is also the handover artefact.
Timeline
committed · MTIn production. The remaining timeline is the skills gap in delivery.
Cost
committed · MTThe human panel is the expense: twelve people, two days per calibration cycle.
Demand
committed · CD'Did the new model break it' is the most-asked question in banking steering committees.
TAM
agent-estimatedAgent-estimated from eval-tooling and AI-QA spend. Small market; the value is in avoided regressions. Uncommitted.
Workforce readiness
agent-estimatedDelivery can run the canary; writing and calibrating a rubric still needs the lab. Agent-estimated.
Cost of being wrong
committed · LFAn uncalibrated judge passing a regressed model into a claims process is an assurance failure with a name on it.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Model changes in a servicing or credit process need a documented regression check; the canary is that check.
Mechanism · Canary suite on the bank's task evals, fired on every model or routing change; calibration score in the model-risk file.
Claims triage evals are the closest thing to a conduct test for an agent; the judge calibration score is what makes them defensible.
Mechanism · Rubric-with-references judge, human panel drawn from claims assessors, agreement score published with the eval. Agent draft.
Merchandising and product-classification evals already run in the Woolworths pilots; the canary caught a regression there first.
Mechanism · Task evals on category data; canary on the pilot's model endpoint.
A calibrated eval is what a TGA software-as-medical-device conversation needs; an uncalibrated one is what it fears.
Mechanism · Clinician panel for calibration; eval record as part of the quality-management evidence.
Red team · the strongest case against
The strongest case against: the field has converged on measuring what is easy to measure. Rubric-scored tasks give good kappa because rubrics make humans agree with each other, not because the judge understands quality. The 0.38 on open-ended work is the honest number, and open-ended work is most of what clients want. Meanwhile the canary catches regressions on the tasks we thought to write evals for and is blind to the ones we did not. A converged field can be a field that stopped asking the hard question.
- —Kappa on rubric tasks measures rubric quality. The judge may be learning the rubric's surface features, which is overfitting with a human agreement score attached.
- —Three canary firings is a small sample; the base rate of material regressions per release is unknown and the false-negative rate is unmeasured.
- —Twelve panellists from the firm are not the client's assessors; calibration may not transfer across panels.
- —Public benchmark saturation is a claim about the benchmarks we watch; the labs' private evals may not saturate and we cannot see them.
Source diversity
- ML research25%
- Open-source tooling15%
- Foundation labs and vendors15%
- Competitors / Argus10%
- Internal / Engel35%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Evals are the table stake most likely to be a brief moat; this field's timing sets that field's.
Parity claims for open-weight models are only claims until they run on the same task evals.
The canary needs to fire on gateway routing changes, not just release announcements; the gateway is the trigger.
Agent-written tests are graded by evals; without calibrated judges the near field cannot be assessed.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- CDClaire Dubois · Sector owner · Banking1 drop
- MTMei Tanaka · Research lead · evals1 drop
- ?Anonymous · Anonymous drop1 drop
- OGOllie Grant · Product engineer1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Every LLM judge ships with a human agreement score or does not ship
strength strong · 29 citations · review 8 Dec 2026
What eval tooling do we use?
strength moderate · 27 citations · review 19 Sep 2026
Judge calibration against human panel
An uncalibrated frontier LLM judge, as used in the harness through Q1, agrees with a five-person human panel at Cohen's κ of at least 0.8 across extraction, summarisation and agentic-trace grading.
Release canary suite v2
A 120-item canary suite with calibrated judges catches at least 80% of the behaviour regressions that the last six vendor model releases introduced on our patterns, at under 15 minutes wall-clock per run.
LLM-as-judge without human calibration
“Agreed with itself, mostly.” · lived 4 months
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Is there any judge configuration that reaches 0.6 agreement on open-ended quality, or is that work permanently human-scored?
- 02What is the canary's false-negative rate — how many regressions has it missed that clients found?
- 03When does the field dissolve into practice, and what does the lab keep?