cavendish
TestedConvergedgate · SkillsNow · 0–12 months×6 sightings

Eval Harnesses

An eval harness is five instruments, not one — capability, task, release canary, judge calibration, cost and latency — and the one that pays first is the canary; a judge without a human agreement score is not an eval, it is an opinion with a decimal point.

Experiment run, measured result. The only tier that becomes a recommendation.

Join with…

Confidence

84%human-committed

Expiry

42duntil review · 15 Oct 2026

Lead time

+4moahead of mainstream awareness

Ownership

MTMei Tanakaweekly cadence

Where it is

Turing was built first, on conviction, before the graph existed, and it is the field that has converged. The release canary has fired on three foundation-model releases and caught a nine-point regression on structured extraction within six hours of one of them; no client pipeline caught it first. Judge calibration against a human panel put the best judge at kappa 0.71 on rubric-scored task evals and 0.38 on open-ended quality, and showed an uncalibrated judge over-scoring the newer model by eleven points — the finding that sent 'judge without a human score' to the graveyard. Public benchmarks saturate inside nine months and stop discriminating; task evals on the firm's own data are the only durable instrument. The gate is skills: the tooling is open and adequate, foundation labs now ship eval products, and the constraint is that eval engineering is a job title the market is only now creating. Competitors are hiring for it.

Why a Quantium decision hinges on it

Every recommendation the lab publishes rests on an eval; every client asks 'how do we know the new model didn't break it'. The canary is the answer to the second and the calibration score is what keeps the first honest. Evals are also the handover artefact — the thing the firm keeps if the lab stops — and the table stake most likely to be a brief moat before competitors staff up. Health clients face a software-as-medical-device question where a calibrated eval is the difference between a regulated product and an unregulated one.

Field attributes

StateConverged
GateSkills · operable, not yet staffed
OriginConviction
Measurablefull
Audience · TLPpractice
Horizonnow
Opened15 Sep 2025
Mainstream15 Jan 2026
Last validated31 Aug 2026
Sightings6

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Release canary v2 fired on three foundation releases and caught a 9-point regression on structured extraction within six hours; no client pipeline caught it first (x-eval-canary).
  • 02Judge calibration against a twelve-person human panel: best judge kappa 0.71 on rubric-scored task evals, 0.38 on open-ended quality; an uncalibrated judge over-scored the newer same-family model by 11 points (x-judge-calibration).
  • 03Rubric-with-references judging raised agreement from 0.58 to 0.71 on the same tasks; free-form judging did not improve with a stronger judge model.
  • 04Public benchmarks in our watchlist saturated (>95% top score) a median of nine months after release; task evals on our own data have not.
What is hype
  • 01'We use LLM-as-judge.' Without a human agreement score it measures the judge's preferences, including its preference for its own family.
  • 02Leaderboard positions. A model's rank on a saturated public benchmark says nothing about its behaviour on a claims-triage task.
  • 03Eval platforms as a product category. The platform is the easy part; the rubric, the references and the panel are the work.
What would have to be true
  • 01Judge agreement above 0.6 on open-ended quality, which no judge configuration we have tried reaches; until then open-ended evals are human-scored.
  • 02A canary that fires on gateway routing changes as well as model releases; today it fires on release announcements only.
  • 03Eval engineering as a delivery skill, not a lab skill: a rubric a sector team can write and calibrate without Turing's owner in the room.
What we would do
  • 01Keep r-judge-calibration as a hard rule: every judge ships with a human agreement score or does not ship.
  • 02Wire the canary to the gateway so it fires on any model or routing change, not just announcements; this is the Argus teeth.
  • 03Move the field toward dissolution by absorption: the tooling is practice now, and the lab's remaining job is calibration and the skills gap.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
0.71
judge–human kappa (rubric)
Finding·band 1Tested

Judge calibration against a human panel: kappa 0.71 on rubric tasks, 0.38 open-ended; uncalibrated judge over-scored the newer model by 11 points

Twelve-person panel scored 600 items across four task evals and two open-ended quality sets. Four judge configurations compared. Rubric-with-references reached kappa 0.71; free-form judging stayed under 0.4 regardless of judge model, and the free-form judge preferred the newer same-family model by 11 points against the panel.

extracted claimJudge reliability is a property of the rubric and references, not the judge model, and uncalibrated judges favour their own family.
Lab · x-judge-calibration · Mei Tanaka31 Aug 2026
detector · bleeding edge
41
openings
Job posting·band 3Signal

'AI Evaluation Engineer' openings in Australia double in six months

Argus hiring scan across consultancies, banks and two foundation-lab AU offices. Forty-one open roles; the title barely existed in January. Inference: the skills gate is being closed by the market, and the lab's advantage on evals has a shelf life.

SEEK / competitor careers pages11 Aug 2026
detector · demand
3 of 3 releases
regressions caught
Finding·band 1Tested

Release canary v2 caught a 9-point extraction regression within six hours of a model release

Canary fires on foundation-lab release announcements and runs the task-eval suite against the new default. On the third release it flagged a nine-point drop on structured extraction; the client pipeline running the same model noticed four days later from support tickets.

extracted claimA release canary on own-data task evals catches material regressions before client pipelines do.
Lab · x-eval-canary · Mei Tanaka23 Jul 2026
detector · bleeding edge 2
0.55 → 0.70
agreement
Finding·band 1Tried

Logged from Claude Code: switching the judge to a rubric with reference answers lifted agreement on the merchandising eval from 0.55 to 0.7

Product engineer re-ran the retail pilot's eval with reference-anchored rubrics and a small human sample. Agreement moved from 0.55 to 0.70. One eval, one person; tried tier, and consistent with the calibration experiment.

MCP · log_finding · Ollie Grant19 Jun 2026
OGdropped
28k
stars
Repository·band 2Tried

Open eval framework passes 28k stars; adds calibration and human-panel modules

The framework Turing is built on. The calibration module landed after we filed the issue; the tooling gate is closed, which is why the field's gate is skills.

github.com19 May 2026
MTdropped 2
6–12 pts
same-family preference
Paper·band 1Signal

Judges prefer their own: self- and family-preference bias in LLM evaluation

Measures judge preference for outputs from the judge's own model family across five families. Finds a consistent 6–12 point preference that survives prompt-level debiasing and disappears only with reference-anchored rubrics. Our calibration reproduced it.

arxiv.org · Iqbal, Sørensen et al.12 Mar 2026
detector · bleeding edge 3
9 months
median time to saturation
Benchmark·band 2Signal

Watchlist benchmark saturation: median nine months from release to a >95% top score

Argus tracking of fourteen public benchmarks in the capability matrix. Median time from release to saturation was nine months; three saturated within four. The reason the capability matrix is scored on our own task evals.

Public leaderboards14 Jan 2026
detector · early adoption 2
Post·band 2Signal

'Your eval is measuring your prompt'

Argues that task evals overfit to the prompt and rubric that produced them and that a good score is a tautology. Partly right — it is why the panel exists — and kept as the sharpest critique of the field's own instrument.

Substack · An applied-ML voice9 Jan 2026
?dropped 3
Client question·band 3Signal

'How do we know the new model version didn't break our pipeline?'

Asked by a head of model risk after a provider deprecated a model mid-engagement. Logged as unanswered at the time; it is now the canary's job description and the most-asked question in the banking pipeline.

Engel · banking engagement10 Dec 2025
CDdropped 5
Release·band 1Signal

Foundation lab ships a hosted evals product with a frontier model as default grader

Managed eval runs with the vendor's own model as the default judge and no calibration step. Convenient, and the design choice our calibration result argues against. Retained as the source of the 'judge is good enough' claim.

OpenAI2 Oct 2025
detector · bleeding edge 3
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 5 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.

Testedc-eval-harnesses-1dalton-0.431 Aug 2026Lab · x-judge-calibration, MCP · log_finding
88%

Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.

Testedc-eval-harnesses-2dalton-0.431 Aug 2026Lab · x-eval-canary, Engel · banking engagement
84%

Uncalibrated judges systematically favour newer and same-family models.

Testedc-eval-harnesses-3dalton-0.44 Jun 2026Lab · x-judge-calibration, arxiv.org
76%

Public benchmarks saturate within about nine months of release and stop discriminating; task evals on own data are the only durable instrument.

Assessedc-eval-harnesses-4dalton-0.321 Jan 2026Public leaderboards, Substack
70%

Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.

Assessedc-eval-harnesses-6dalton-0.431 Aug 2026SEEK / competitor careers pages, github.com
60%

A frontier model as judge is reliable enough without human calibration for most enterprise uses.

Testedc-eval-harnesses-5dalton-0.38 Oct 2025OpenAI, Lab · x-judge-calibration
18%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 50% → 84%.

runs compare claim sets, never prose
What we said · run 4

Both experiments concluded. Canary caught a real regression before any client did; judge calibration gives 0.71 on rubric tasks and 0.38 on open-ended. Unchecked judge to the graveyard; recommendation and standing answer current. Field converged; skills is the gate.

84%
Changed since run 3
  • LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.
  • Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.
  • Uncalibrated judges systematically favour newer and same-family models.
  • Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.
  • c-eval-harnesses-5 ↓ 0.30 → 0.18
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · MT
high

Every recommendation and every canary rests on it; it is also the handover artefact.

Timeline

committed · MT
0–18mo

In production. The remaining timeline is the skills gap in delivery.

Cost

committed · MT
medium

The human panel is the expense: twelve people, two days per calibration cycle.

Demand

committed · CD
high

'Did the new model break it' is the most-asked question in banking steering committees.

TAM

agent-estimated
$100M–1B

Agent-estimated from eval-tooling and AI-QA spend. Small market; the value is in avoided regressions. Uncommitted.

Workforce readiness

agent-estimated
medium

Delivery can run the canary; writing and calibrating a rubric still needs the lab. Agent-estimated.

Cost of being wrong

committed · LF
high

An uncalibrated judge passing a regressed model into a claims process is an assurance failure with a name on it.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Banking
relevant

Model changes in a servicing or credit process need a documented regression check; the canary is that check.

Mechanism · Canary suite on the bank's task evals, fired on every model or routing change; calibration score in the model-risk file.

CD committed by Claire Duboiscommitted
Insurance
relevant

Claims triage evals are the closest thing to a conduct test for an agent; the judge calibration score is what makes them defensible.

Mechanism · Rubric-with-references judge, human panel drawn from claims assessors, agreement score published with the eval. Agent draft.

Agent draft · awaiting a sector owneragent-estimated
Retail & FMCG
relevant

Merchandising and product-classification evals already run in the Woolworths pilots; the canary caught a regression there first.

Mechanism · Task evals on category data; canary on the pilot's model endpoint.

DS committed by Dev Sharmacommitted
Health
relevant

A calibrated eval is what a TGA software-as-medical-device conversation needs; an uncalibrated one is what it fears.

Mechanism · Clinician panel for calibration; eval record as part of the quality-management evidence.

HN committed by Hana Novakcommitted

Red team · the strongest case against

The strongest case against: the field has converged on measuring what is easy to measure. Rubric-scored tasks give good kappa because rubrics make humans agree with each other, not because the judge understands quality. The 0.38 on open-ended work is the honest number, and open-ended work is most of what clients want. Meanwhile the canary catches regressions on the tasks we thought to write evals for and is blind to the ones we did not. A converged field can be a field that stopped asking the hard question.

  • Kappa on rubric tasks measures rubric quality. The judge may be learning the rubric's surface features, which is overfitting with a human agreement score attached.
  • Three canary firings is a small sample; the base rate of material regressions per release is unknown and the false-negative rate is unmeasured.
  • Twelve panellists from the firm are not the client's assessors; calibration may not transfer across panels.
  • Public benchmark saturation is a claim about the benchmarks we watch; the labs' private evals may not saturate and we cannot see them.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis holds

Source diversity

  • ML research25%
  • Open-source tooling15%
  • Foundation labs and vendors15%
  • Competitors / Argus10%
  • Internal / Engel35%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsMTAWJPLFOGCDHN

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Is there any judge configuration that reaches 0.6 agreement on open-ended quality, or is that work permanently human-scored?
  2. 02What is the canary's false-negative rate — how many regressions has it missed that clients found?
  3. 03When does the field dissolve into practice, and what does the lab keep?