Lab Effectivity
A twelve-person research lab justifies itself on three numbers — lead time, calibration and cost per validated recommendation — or on nothing; labs that report output volume instead of decisions changed are restructured in their second budget cycle.
A validation run. Researched position, no experiment.
Confidence
52%human-committedExpiry
19doverdue for reviewLead time
—not yet mainstream · opened 14 Apr 2026Ownership
AWAdam Witanowskimonthly cadenceWhere it is
A meta-field, opened from the question the sponsor will ask at the next budget review: what did the lab change that the firm would not have done anyway. The validation run looked at how eleven corporate and consultancy research labs measured themselves and what happened to them; the ones that survived reported decisions changed and lead time over the market, the ones that died reported papers, demos and headcount. Our own first measurement, computed from field open dates against mainstream dates, gives a median lead time of 4.2 months across sixteen now-fields with three honest negatives. Two quarters of data is not proof of anything. The field exists so that the measurement is captured from day one rather than reconstructed under pressure.
Why a Quantium decision hinges on it
The lab costs the firm roughly a delivery pod. The sponsor has to defend that line in FY27 planning and cannot do it with a list of things the lab looked at. If lead time, calibration and cost per validated recommendation are not being measured now, they cannot be produced in March, and the default outcome for research functions without a number is absorption into delivery. This is the field where the lab is the client.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Lead time is computable from the graph today: median +4.2 months across sixteen now-fields, three negative, from openedAt against mainstreamAt (logged from a session, tried tier).
- 02Of eleven labs in the validation run, the six that reported decisions changed and lead time were intact at 36 months; four of the five that reported output volume were folded into delivery or closed.
- 03Calibration can be measured per axis: expert technology forecasters in the literature are calibrated on direction and badly miscalibrated on timing, which is the axis our position runs mostly commit to.
- 01Lab scorecards with a dozen metrics. The labs that survived tracked three; the ones with dashboards spent their last year building dashboards.
- 02Citations and publications as proof of worth. They measure the field's interest in the lab, not the firm's.
- 03'Innovation ROI' frameworks from the consultancies. They price the lab's outputs; nobody has shown one that priced the decisions the lab prevented.
- 01At least four quarters of position runs, so calibration has enough resolved predictions to score — currently most of our predictions have not resolved.
- 02A cost-per-validated-recommendation figure the sponsor accepts as the denominator; the numerator is easy and the cost allocation is contested.
- 03Engel first-mention dates, so lead time can be measured against client demand and not only against the public mainstream date.
- 01Ship a three-number scorecard — lead time, calibration, cost per validated recommendation — to the sponsor monthly, with the honest negatives shown.
- 02Score every position run's confidence as a prediction with a resolution date, so calibration exists by FY27 planning.
- 03Decide by December whether this stays a field or dissolves into the instrument-health work in Nightingale.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Validation run: how eleven research labs measured themselves, and which ones survived
Desk research over eleven corporate and consultancy research labs with public records from 2019–2026. Six reported decisions changed and lead time over the market; all six intact at 36 months. Five reported output volume; four were folded into delivery or closed within two budget cycles. No experiment; a researched position.
extracted claimReporting decisions changed predicts lab survival; reporting output predicts absorption.

Logged from Claude Code: lead time computed across sixteen now-fields — median +4.2 months, three negative
Measurement engineer logged from a session: a script over openedAt and mainstreamAt for every now-field. Median lead +4.2 months; three fields opened after mainstream, one by four months. Tried tier: mainstream dates are hand-assigned and the number is only as good as they are.

A big-four firm folds its AI research lab into its delivery practice
Announced as 'bringing research closer to clients'. The lab had published forty pieces in eighteen months. Anonymous drop with the note 'this is what the second budget review looks like from outside'.

'Head of Research Operations' postings up at AI-native consultancies
Four postings in a quarter, all describing lab measurement, portfolio review and 'research-to-delivery conversion'. Argus inference: competitors are professionalising the lab-measurement role rather than leaving it to the lab head.

Analyst note: 40% of consultancy AI research functions restructured within 24 months of founding
Survey of professional-services firms that stood up AI research functions in 2023–24. Restructuring reasons cited: unclear contribution to revenue, duplication with delivery. Carried as demand-band context, not as evidence about any single lab.

'What did the lab change this quarter that we would not have done anyway?'
Asked by the sponsor at the Q3 budget check-in. Not a client question in the Engel sense; logged as one because the lab is the client for this field. Answered with a list of activities, which was the wrong shape. This is the question that opened the field.

'Why corporate research labs die, and what the ones that don't have in common'
First-hand account across three labs. The pattern: the lab that could name the decisions it changed survived the second budget review; the ones with the best demos did not. The post that turned the sponsor's question into a field.
extracted claimLabs die at the second budget review unless they can name decisions changed.

Calibration of Expert Technology Forecasters: Brier Scores Over a Decade of Resolved Predictions
Scores ten years of resolved technology predictions from analysts, academics and practitioners. Direction calls were well calibrated; timing calls were systematically early by a median of 2.4 years, and confidence did not predict accuracy on timing at all.
extracted claimExperts are calibrated on whether and not on when.

'Measuring a research org without killing it'
A lab head's talk on three-metric scorecards and why dashboards with twelve metrics correlate with closure. The 'three numbers' framing came from here.

Resolved AI-capability prediction set, 2019–2026
Roughly 1,400 resolved public predictions about AI capabilities with timestamps and forecaster confidence. The dataset behind the calibration paper and the one we would score the lab's own position runs against.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
Lead time is the only proof of worth a lab cannot reconstruct retrospectively; it has to be captured at field open or it does not exist.
Expert technology forecasters are calibrated on direction and miscalibrated on timing, so calibration must be scored per axis or it flatters the lab.
Labs that report output volume — papers, demos, headcount — rather than decisions changed are restructured in their second budget cycle.
The lab's current median lead time is +4.2 months across sixteen now-fields with three negatives; too early to claim value, early enough to claim honesty.
Cost per validated recommendation is a sufficient single metric for the lab.
Position history · the diff is the product
3 validation runs against a fixed brief. Confidence 40% → 52%.
First lead-time measurement from the graph: median +4.2 months, three negatives. Lead time confirmed as the number that cannot be reconstructed later. Field stays open until December, then dissolves or does not.
- Lead time is the only proof of worth a lab cannot reconstruct retrospectively; it has to be captured at field open or it does not exist.
- The lab's current median lead time is +4.2 months across sixteen now-fields with three negatives; too early to claim value, early enough to claim honesty.
- c-lab-effectivity-2 ↑ 0.58 → 0.64
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · AWDecides whether the lab exists in FY28.
Timeline
committed · RMFY27 planning is the deadline. Nothing after that matters if this is wrong.
Cost
committed · AWThe measurement is a graph query and a monthly page. The cost is in honesty, not effort.
Cost of being wrong
agent-estimatedA flattering scorecard that fails at budget review is worse than none. Agent-estimated.
Demand
agent-estimatedNo client asks this. The demand is one sponsor and one budget cycle; Engel cannot see it. Agent-estimated.
Workforce readiness
committed · JPNightingale can compute all three numbers; the lab has to agree to be scored.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
The lab's own justification, and the framing the sponsor will use across every sector conversation about research spend.
Mechanism · Three-number scorecard in the monthly sponsor note; negatives shown.
Public-sector research and foresight units face the same measurement question and are periodically restructured for the same reason. A possible standing answer, not a field.
Mechanism · Would apply as advisory content if an agency asked how to measure its own foresight function.
No client mechanism. Banks do not buy the lab's self-measurement and should not see it.
Mechanism · None.
Red team · the strongest case against
The strongest case against: a twelve-person lab measuring itself is navel-gazing on the firm's dime, with a built-in incentive to flatter. Two quarters of lead time over sixteen fields is noise, calibration cannot be scored until predictions resolve, and the real test of the lab is whether execs use it — which is the visibility field's job, not this one's.
- —Median lead time of 4.2 months on sixteen fields with self-assigned mainstream dates is a number the lab gave itself. An outside reviewer would call the mainstream dates the softest input in the graph.
- —Self-measurement has an incentive problem this field cannot solve from the inside. Every lab in the validation run that died also had a scorecard.
- —The three numbers reward opening fields early and vaguely. A lab optimising lead time will elect more candidates and commit to less, which is the opposite of what the sponsor wants.
- —The field's own cost-of-being-wrong score is high and its confidence is 0.52. That is an argument for folding it into instrument health, not for running it as a field.
Source diversity
- Research leadership / practice35%
- Forecasting research25%
- Analyst / press15%
- Internal / sponsor25%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Decisions changed cannot be counted unless consumption is instrumented; that field supplies the numerator.
Same measurement problem pointed at the lab instead of at a client's AI programme; the survey-versus-telemetry lesson transfers directly.
Scoring position-run confidence as a prediction borrows the judge-calibration method against a resolved panel.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- RMRohan Mehta · Exec sponsor2 drops
- JPJun Park · Measurement (Nightingale)1 drop
- AWAdam Witanowski · Lab Director (acting)1 drop
- ?Anonymous · Anonymous drop1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Who assigns mainstream dates, and can they be assigned by someone outside the lab so lead time is not self-graded?
- 02How many resolved predictions does calibration need before it says anything, and do we reach that by FY27 planning?
- 03Does the sponsor accept a cost-per-validated-recommendation denominator that includes the lab's own time on this field?