cavendish
QueueExperimentsScorecard

Lab app · Scorecard · Q3 2026

The track record.

Lead time measures whether the lab is distinctive. Answer rate measures whether it is useful. A lab needs both, and only one of them will be asked about in a budget review.

0mo
Median lead time
11 fields reached mainstream
0%
Answer rate
share of questions the library answers well
$0.0k
Cost per validated recommendation
attributed through the graph's cost edges
0%
Top-50 freshness
1 library items overdue for review
Fields
30
Candidates
10
Experiments
20
Validated
3
Killed
16
Recommendations
12
Corrections
1
Claims
208

Lead time · per field

When we opened it versus when it went mainstream.

Negative lead time is recorded honestly. An honest negative is the only thing that makes a positive one believable.

Calibration

What we said versus what happened.

Predictions are dated, confidence-scored and resolvable. A perfectly calibrated lab sits on the diagonal. We are over-confident at the top and under-confident at the bottom, which is the usual shape.

bucketpredictedactualn
90%+
92%
83%
6
70–90%
79%
71%
14
50–70%
60%
58%
19
30–50%
41%
47%
11
<30%
22%
30%
8

Brier 0.19 across 58 resolved predictions. Deferred to a spreadsheet for the first year, as the design says.

Instrument health · weekly

The system applies its own decay model to itself.

Cavendish tells the firm when its knowledge is stale. It also has to know when it is degrading — the observability nobody would accept omitting from a client system and everybody omits from their own.

Source pool diversity

38% ML research

largest single community share; drift +4pts this quarter

Source yield distribution

23 of 61

sources that have ever produced an elected or validated result

Claim extraction quality

0.87

sampled human agreement with dalton-0.4 output, n=120

Extraction cost per signal

$0.041

against a $0.06 per-source cap

Cluster separability

6%

share of clusters failing the separability test — the blob failure mode

Ranking quality

2.6×

acceptance of top-ranked candidates vs a random sample from the pool

Diff signal-to-noise

0.71

share of claim changes attributable to source changes; two fields carry a degraded marker

Dream journal acceptance

31%

rolling four weeks; a feed perceived as noise is abandoned

Cost drift

1 field

voice-and-vision 18% over cap after the latency bench

Human decision load

17 / wk

against the ~20 target

Human vs detector origin

67%

elected fields originating from people rather than detectors. If this runs heavily human, the automation is a filing system, not a discovery engine — still valuable, worth knowing.

Human drop share

60%

signals with at least one human drop attached. The human route stays first-class permanently.

Experiments requested from outside

9

the strongest measure available from the first cycle. It needs no instrumentation and measures whether the firm finds the lab useful.

16 graveyard entries This week’s dream journal Annual adversarial review scheduled November, under Asilomar.