Lab app · Scorecard · Q3 2026
The track record.
Lead time measures whether the lab is distinctive. Answer rate measures whether it is useful. A lab needs both, and only one of them will be asked about in a budget review.
Lead time · per field
When we opened it versus when it went mainstream.
Negative lead time is recorded honestly. An honest negative is the only thing that makes a positive one believable.
- AI-SDLC+6mo
- SLM / Edge / Tuning+4mo
- Open Weight Models+4mo
- Agentic Memory System+4mo
- Fully Agentic QA+4mo
- Eval Harnesses+4mo
- Cost redux on tokens−3mo
- ROI−4mo
- AI Gateway−4mo
- Voice and Vision−4mo
- Quantum Encryption−20mo
Calibration
What we said versus what happened.
Predictions are dated, confidence-scored and resolvable. A perfectly calibrated lab sits on the diagonal. We are over-confident at the top and under-confident at the bottom, which is the usual shape.
Brier 0.19 across 58 resolved predictions. Deferred to a spreadsheet for the first year, as the design says.
Instrument health · weekly
The system applies its own decay model to itself.
Cavendish tells the firm when its knowledge is stale. It also has to know when it is degrading — the observability nobody would accept omitting from a client system and everybody omits from their own.
Source pool diversity
largest single community share; drift +4pts this quarter
Source yield distribution
sources that have ever produced an elected or validated result
Claim extraction quality
sampled human agreement with dalton-0.4 output, n=120
Extraction cost per signal
against a $0.06 per-source cap
Cluster separability
share of clusters failing the separability test — the blob failure mode
Ranking quality
acceptance of top-ranked candidates vs a random sample from the pool
Diff signal-to-noise
share of claim changes attributable to source changes; two fields carry a degraded marker
Dream journal acceptance
rolling four weeks; a feed perceived as noise is abandoned
Cost drift
voice-and-vision 18% over cap after the latency bench
Human decision load
against the ~20 target
Human vs detector origin
elected fields originating from people rather than detectors. If this runs heavily human, the automation is a filing system, not a discovery engine — still valuable, worth knowing.
Human drop share
signals with at least one human drop attached. The human route stays first-class permanently.
Experiments requested from outside
the strongest measure available from the first cycle. It needs no instrumentation and measures whether the firm finds the lab useful.