cavendish
AssessedEmerginggate · AdoptionNow · 0–12 months×2 sightings

Lab Effectivity

A twelve-person research lab justifies itself on three numbers — lead time, calibration and cost per validated recommendation — or on nothing; labs that report output volume instead of decisions changed are restructured in their second budget cycle.

A validation run. Researched position, no experiment.

Join with…

Confidence

52%human-committed

Expiry

19doverdue for review

Lead time

not yet mainstream · opened 14 Apr 2026

Ownership

AWAdam Witanowskimonthly cadence

Where it is

A meta-field, opened from the question the sponsor will ask at the next budget review: what did the lab change that the firm would not have done anyway. The validation run looked at how eleven corporate and consultancy research labs measured themselves and what happened to them; the ones that survived reported decisions changed and lead time over the market, the ones that died reported papers, demos and headcount. Our own first measurement, computed from field open dates against mainstream dates, gives a median lead time of 4.2 months across sixteen now-fields with three honest negatives. Two quarters of data is not proof of anything. The field exists so that the measurement is captured from day one rather than reconstructed under pressure.

Why a Quantium decision hinges on it

The lab costs the firm roughly a delivery pod. The sponsor has to defend that line in FY27 planning and cannot do it with a list of things the lab looked at. If lead time, calibration and cost per validated recommendation are not being measured now, they cannot be produced in March, and the default outcome for research functions without a number is absorption into delivery. This is the field where the lab is the client.

Field attributes

StateEmerging
GateAdoption · ready — blocked by trust, regulation, procurement or change capacity
OriginQuestion
Measurablepartial
Audience · TLPlab
Horizonnow
Opened14 Apr 2026
Mainstreamnot yet
Last validated31 Jul 2026
Sightings2

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Lead time is computable from the graph today: median +4.2 months across sixteen now-fields, three negative, from openedAt against mainstreamAt (logged from a session, tried tier).
  • 02Of eleven labs in the validation run, the six that reported decisions changed and lead time were intact at 36 months; four of the five that reported output volume were folded into delivery or closed.
  • 03Calibration can be measured per axis: expert technology forecasters in the literature are calibrated on direction and badly miscalibrated on timing, which is the axis our position runs mostly commit to.
What is hype
  • 01Lab scorecards with a dozen metrics. The labs that survived tracked three; the ones with dashboards spent their last year building dashboards.
  • 02Citations and publications as proof of worth. They measure the field's interest in the lab, not the firm's.
  • 03'Innovation ROI' frameworks from the consultancies. They price the lab's outputs; nobody has shown one that priced the decisions the lab prevented.
What would have to be true
  • 01At least four quarters of position runs, so calibration has enough resolved predictions to score — currently most of our predictions have not resolved.
  • 02A cost-per-validated-recommendation figure the sponsor accepts as the denominator; the numerator is easy and the cost allocation is contested.
  • 03Engel first-mention dates, so lead time can be measured against client demand and not only against the public mainstream date.
What we would do
  • 01Ship a three-number scorecard — lead time, calibration, cost per validated recommendation — to the sponsor monthly, with the honest negatives shown.
  • 02Score every position run's confidence as a prediction with a resolution date, so calibration exists by FY27 planning.
  • 03Decide by December whether this stays a field or dissolves into the instrument-health work in Nightingale.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
6 of 11
labs surviving
Finding·band 1Assessed

Validation run: how eleven research labs measured themselves, and which ones survived

Desk research over eleven corporate and consultancy research labs with public records from 2019–2026. Six reported decisions changed and lead time over the market; all six intact at 36 months. Five reported output volume; four were folded into delivery or closed within two budget cycles. No experiment; a researched position.

extracted claimReporting decisions changed predicts lab survival; reporting output predicts absorption.
Lab · validation run 2 · Adam Witanowski24 Jun 2026
detector · bleeding edge
+4.2 mo
median lead
Finding·band 1Tried

Logged from Claude Code: lead time computed across sixteen now-fields — median +4.2 months, three negative

Measurement engineer logged from a session: a script over openedAt and mainstreamAt for every now-field. Median lead +4.2 months; three fields opened after mainstream, one by four months. Tried tier: mainstream dates are hand-assigned and the number is only as good as they are.

MCP · log_finding · Jun Park22 Jul 2026
JPdropped
Announcement·band 3Signal

A big-four firm folds its AI research lab into its delivery practice

Announced as 'bringing research closer to clients'. The lab had published forty pieces in eighteen months. Anonymous drop with the note 'this is what the second budget review looks like from outside'.

Press release8 Jul 2026
?dropped 2
4
openings
Job posting·band 3Signal

'Head of Research Operations' postings up at AI-native consultancies

Four postings in a quarter, all describing lab measurement, portfolio review and 'research-to-delivery conversion'. Argus inference: competitors are professionalising the lab-measurement role rather than leaving it to the lab head.

Consultancy careers pages16 Jun 2026
detector · demand
40%
restructured
Analyst·band 3Signal

Analyst note: 40% of consultancy AI research functions restructured within 24 months of founding

Survey of professional-services firms that stood up AI research functions in 2023–24. Restructuring reasons cited: unclear contribution to revenue, duplication with delivery. Carried as demand-band context, not as evidence about any single lab.

Analyst brief28 May 2026
detector · demand
Client question·band 3Signal

'What did the lab change this quarter that we would not have done anyway?'

Asked by the sponsor at the Q3 budget check-in. Not a client question in the Engel sense; logged as one because the lab is the client for this field. Answered with a list of activities, which was the wrong shape. This is the question that opened the field.

Exec sponsor · budget review9 Apr 2026
RMdropped
Post·band 2Signal

'Why corporate research labs die, and what the ones that don't have in common'

First-hand account across three labs. The pattern: the lab that could name the decisions it changed survived the second budget review; the ones with the best demos did not. The post that turned the sponsor's question into a field.

extracted claimLabs die at the second budget review unless they can name decisions changed.
Personal blog · A former corporate lab director3 Apr 2026
RMdropped 3
2.4 yr early
timing bias
Paper·band 1Signal

Calibration of Expert Technology Forecasters: Brier Scores Over a Decade of Resolved Predictions

Scores ten years of resolved technology predictions from analysts, academics and practitioners. Direction calls were well calibrated; timing calls were systematically early by a median of 2.4 years, and confidence did not predict accuracy on timing at all.

extracted claimExperts are calibrated on whether and not on when.
arxiv.org · Whitfield, Osei et al.18 Mar 2026
AWdropped 2
Talk·band 2Signal

'Measuring a research org without killing it'

A lab head's talk on three-metric scorecards and why dashboards with twelve metrics correlate with closure. The 'three numbers' framing came from here.

Industry research-leadership summit26 Feb 2026
detector · early adoption
~1,400
resolved predictions
Dataset·band 2Signal

Resolved AI-capability prediction set, 2019–2026

Roughly 1,400 resolved public predictions about AI capabilities with timestamps and forecaster confidence. The dataset behind the calibration paper and the one we would score the lab's own position runs against.

Public forecasting platform export15 Jan 2026
detector · early adoption
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Lead time is the only proof of worth a lab cannot reconstruct retrospectively; it has to be captured at field open or it does not exist.

Assessedc-lab-effectivity-1dalton-0.431 Jul 2026MCP · log_finding, Lab · validation run 2
78%

Expert technology forecasters are calibrated on direction and miscalibrated on timing, so calibration must be scored per axis or it flatters the lab.

Assessedc-lab-effectivity-3dalton-0.36 May 2026arxiv.org, Public forecasting platform export
70%

Labs that report output volume — papers, demos, headcount — rather than decisions changed are restructured in their second budget cycle.

Assessedc-lab-effectivity-2dalton-0.431 Jul 2026Personal blog, Analyst brief, Press release
64%

The lab's current median lead time is +4.2 months across sixteen now-fields with three negatives; too early to claim value, early enough to claim honesty.

Triedc-lab-effectivity-5dalton-0.422 Jul 2026MCP · log_finding
55%

Cost per validated recommendation is a sufficient single metric for the lab.

Assessedc-lab-effectivity-4dalton-0.424 Jun 2026Lab · validation run 2, Exec sponsor · budget review
30%

Position history · the diff is the product

3 validation runs against a fixed brief. Confidence 40% → 52%.

runs compare claim sets, never prose
What we said · run 3

First lead-time measurement from the graph: median +4.2 months, three negatives. Lead time confirmed as the number that cannot be reconstructed later. Field stays open until December, then dissolves or does not.

52%
Changed since run 2
  • Lead time is the only proof of worth a lab cannot reconstruct retrospectively; it has to be captured at field open or it does not exist.
  • The lab's current median lead time is +4.2 months across sixteen now-fields with three negatives; too early to claim value, early enough to claim honesty.
  • c-lab-effectivity-2 ↑ 0.58 → 0.64
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

Decides whether the lab exists in FY28.

Timeline

committed · RM
0–18mo

FY27 planning is the deadline. Nothing after that matters if this is wrong.

Cost

committed · AW
low

The measurement is a graph query and a monthly page. The cost is in honesty, not effort.

Cost of being wrong

agent-estimated
high

A flattering scorecard that fails at budget review is worse than none. Agent-estimated.

Demand

agent-estimated
not-measurable-here

No client asks this. The demand is one sponsor and one budget cycle; Engel cannot see it. Agent-estimated.

Workforce readiness

committed · JP
medium

Nightingale can compute all three numbers; the lab has to agree to be scored.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Cross-sector
relevant

The lab's own justification, and the framing the sponsor will use across every sector conversation about research spend.

Mechanism · Three-number scorecard in the monthly sponsor note; negatives shown.

RM committed by Rohan Mehtacommitted
Government
watch

Public-sector research and foresight units face the same measurement question and are periodically restructured for the same reason. A possible standing answer, not a field.

Mechanism · Would apply as advisory content if an agency asked how to measure its own foresight function.

Agent draft · awaiting a sector owneragent-estimated
Banking
not-relevant

No client mechanism. Banks do not buy the lab's self-measurement and should not see it.

Mechanism · None.

CD committed by Claire Duboiscommitted

Red team · the strongest case against

The strongest case against: a twelve-person lab measuring itself is navel-gazing on the firm's dime, with a built-in incentive to flatter. Two quarters of lead time over sixteen fields is noise, calibration cannot be scored until predictions resolve, and the real test of the lab is whether execs use it — which is the visibility field's job, not this one's.

  • Median lead time of 4.2 months on sixteen fields with self-assigned mainstream dates is a number the lab gave itself. An outside reviewer would call the mainstream dates the softest input in the graph.
  • Self-measurement has an incentive problem this field cannot solve from the inside. Every lab in the validation run that died also had a scorecard.
  • The three numbers reward opening fields early and vaguely. A lab optimising lead time will elect more candidates and commit to less, which is the opposite of what the sponsor wants.
  • The field's own cost-of-being-wrong score is high and its confidence is 0.52. That is an argument for folding it into instrument health, not for running it as a field.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • Research leadership / practice35%
  • Forecasting research25%
  • Analyst / press15%
  • Internal / sponsor25%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsAWJPLFRM

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Who assigns mainstream dates, and can they be assigned by someone outside the lab so lead time is not self-graded?
  2. 02How many resolved predictions does calibration need before it says anything, and do we reach that by FY27 planning?
  3. 03Does the sponsor accept a cost-per-validated-recommendation denominator that includes the lab's own time on this field?