ROI
measure and quantify
AI return is only measurable from telemetry and cycle-time deltas on instrumented workflows; self-reported productivity overstates the delta by two to three times, and a firm that cannot show the telemetry number will lose the budget to one that can.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
74%human-committedExpiry
36duntil review · 9 Oct 2026Lead time
−4moopened after mainstream — recorded honestlyOwnership
JPJun Parkfortnightly cadenceWhere it is
The field was opened under pressure in January when the exec sponsor asked for a number the board would accept. Nine months on, the honest answer is narrower than the question. Our Nightingale phase-1 run on fourteen internal repos found coding assistants cut median PR cycle time by 23% while pushing review load up, for a net delivery gain in the 10–18% band; the same engineers surveyed at the same time reported 41%. Vendor telemetry (acceptance rate, retained lines) measures usage and does not correlate with the cycle-time change. Public data agrees with the shape: ABS series show most firms using AI cannot say what it returned, and the widely quoted 'most pilots show no P&L impact' analyst line is a statement about missing baselines, not missing value. The gate is adoption: the measurement method exists and is cheap, and clients are still choosing surveys because a survey gives them the number they want.
Why a Quantium decision hinges on it
Every AI engagement Quantium sells is now asked to defend its return, and the December board cycle will ask again. If the firm's own number comes from a survey it is indefensible the moment a client's CFO runs the telemetry; if it comes from telemetry the firm can price outcomes rather than hours. The measurement method is also the only thing that turns the lab's other fields into money — a memory layer or a backpressure pattern is worth nothing until its delivery effect is on a cycle-time chart. This field is the bridge between the lab and the P&L.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Telemetry-measured coding-assistant gain on fourteen internal repos: median PR cycle time −23%, review load +14%, net delivery gain 10–18% depending on repo age (x-roi-measurement).
- 02Self-reported gain from the same engineers in the same fortnight: 41%. The survey/telemetry ratio held between 2.3× and 2.8× across three cohorts.
- 03Vendor acceptance-rate dashboards do not predict cycle-time change: r = 0.11 across repos. Acceptance measures how often people press tab.
- 04The measurement run is cheap. One engineer, four days, from existing git and CI metadata plus assistant session logs. No client data touched.
- 01'10× developer' and '40% productivity' claims. Every one we traced came from a survey or a vendor-run study with no control repo.
- 02Vendor ROI calculators. They multiply seats by an assumed hours-saved figure that the vendor supplies.
- 03'AI-native' P&L transformation stories. The public cases with hard numbers are cost-out on headcount, not measured productivity.
- 01Clients agreeing to instrument delivery before the engagement starts, which today is rare — the baseline is the thing nobody has.
- 02A telemetry method for non-engineering work (analysts, servicing, claims) that is as clean as git metadata. We do not have one; cycle time on a claims queue is the nearest candidate.
- 03Phase-2 Nightingale access to Quantium's own engagement records so the method can be run on the firm's book, not just the lab's repos.
- 01Ship r-roi-measure as the default for any engagement that will be asked for a return figure. Instrument first, survey never.
- 02Run the Nightingale measurement again in November on the same repos with the same query, so the board gets a diff rather than a number.
- 03Take the survey method to the graveyard publicly (done: g-roi-survey-method) and use the tombstone in client conversations.
- 04Extend the method to one non-engineering workflow — claims triage cycle time — in Q4 as a Type 3.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Nightingale phase 1: coding-assistant ROI on fourteen internal repos
Versioned measurement run over git and CI metadata plus assistant session logs, twelve-month baseline. Median PR cycle time fell 23%, review load rose 14%, net delivery gain 10–18%. The same engineers surveyed the same fortnight reported 41%. Query stored; re-runnable in November.
extracted claimSelf-reported coding-assistant productivity gains overstate telemetry-measured gains by two to three times.

'What I want from an AI business case' — a listed-company CFO
A CFO at an ASX 100 company sets out what they will and will not accept as evidence: telemetry against a baseline yes, staff surveys no, vendor calculators no. Demand-band confirmation of the method from the buyer's side.

Big-four consultancy hiring 'AI Value Realisation Lead' ×6 in Sydney and Melbourne
Argus competitor watch. Six identical postings; the role exists to produce a return figure for clients after the fact. Inference: the competitor's ROI evidence is a service line, not a measurement.

'Surveys are fine if you ask a counterfactual question'
Argues that 'how long would this have taken without the assistant' recovers a defensible number cheaply. Our run asked exactly that question and still got 2.3× the telemetry figure. Kept as the strongest disconfirming voice.

ABS Business Characteristics Survey: AI use and measured benefit
Latest release adds a question on whether the business measured a benefit from AI use. Of businesses reporting AI use, a small minority report any measurement at all. Used as the public baseline in the Nightingale run.

Logged from Claude Code: hooks give per-task token cost and edit acceptance for free
Product engineer wired session hooks to emit task id, token spend and whether the edit survived to merge. Two hours of work; gave the measurement run its per-task cost column. Tried tier, one harness, one person.

Coding-assistant vendor ships org-level 'productivity' dashboard
Acceptance rate, retained lines, active seats. No cycle-time or defect measure. We joined it to our own run: acceptance rate did not predict cycle-time change across repos.

'The board wants a number for the AI spend by December. What number?'
Logged by the sector owner after a steering committee. The client's own figure came from a staff survey; their CFO did not believe it. This is the demand signal that turned a pressure-opened field into a measurement run.

Replicating the experienced-developer RCT on enterprise monorepos
Follows the 2025 randomised trial that found experienced developers were slower with assistants while believing they were faster. On enterprise monorepos the sign flips to a modest gain for juniors and near zero for seniors; the perception gap persists in every cohort.
extracted claimDevelopers' perceived assistant speed-up exceeds measured speed-up in every cohort studied, including the cohorts where the measured effect is negative.

'Most enterprise gen-AI pilots show no P&L impact'
Widely quoted headline that reached boardrooms in weeks. Reading the method: it counts pilots with no measured P&L line, which is not the same as pilots with no return. Retained as the mainstream framing our thesis answers.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
Self-reported productivity gains from coding assistants overstate telemetry-measured gains by a factor of two to three.
Coding assistants cut PR cycle time but raise review load; the net delivery gain is in the 10–20% band, not 40%.
Vendor acceptance-rate telemetry measures usage, not value; it does not correlate with cycle-time change.
Most firms cannot show a P&L effect from AI because they have no baseline, not because there is no effect.
A survey with a well-designed counterfactual question recovers the same figure as telemetry at a fraction of the cost.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 40% → 74%.
Run concluded. Survey overstates telemetry 2–3×; net delivery gain 10–18%. Survey method sent to the graveyard; recommendation published. Gate moved from tooling to adoption.
- Self-reported productivity gains from coding assistants overstate telemetry-measured gains by a factor of two to three.
- Coding assistants cut PR cycle time but raise review load; the net delivery gain is in the 10–20% band, not 40%.
- c-ai-roi-3 ↑ 0.55 → 0.68
- c-ai-roi-5 ↓ 0.35 → 0.24
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · RMDecides whether the AI budget survives December. Nothing else in the lab has that property.
Timeline
committed · JPMethod is live; the remaining work is getting clients to instrument before they start.
TAM
agent-estimatedAgent-estimated from global AI-services spend that is nominally conditional on demonstrated return. Not a meaningful band for this field; uncommitted.
Demand
committed · CDAsked in every banking steering committee this half. Two clients have made the next tranche conditional on a number.
Cost
committed · JPFour engineer-days per measurement run from existing metadata.
Cost of being wrong
agent-estimatedA survey number quoted to a board and later contradicted by the client's own telemetry is a credibility event. Agent-estimated.
Workforce readiness
committed · AWDelivery leads can run the git-metadata method; nobody outside the lab can yet run it on a non-engineering workflow.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Banks are the clients asking for a number, and the ones most likely to run their own telemetry to check it.
Mechanism · Instrument the engagement's delivery repos and servicing queues on day one; report cycle-time deltas against a pre-engagement baseline.
Woolworths engineering already has assistant seats at scale and a survey-based number nobody trusts.
Mechanism · Same git-metadata method; the group has the CI history to make a twelve-month baseline.
Agencies are under a whole-of-government direction to report AI benefits but the reporting template is a survey. Our method would contradict their own returns.
Mechanism · Depends on whether the Digital Transformation Agency's benefits framework accepts telemetry in place of self-report. Agent draft.
Contact-centre AI is sold on handle-time reduction; handle time is telemetry the client already has, so the method transfers directly.
Mechanism · Cycle time on a servicing queue rather than a PR; baseline from the client's own ACD data. Agent draft.
Red team · the strongest case against
The strongest case against: cycle time is a proxy that measures what is easy to measure. A 23% faster PR that ships worse code, or shifts effort into review that telemetry books as overhead, is not a 23% gain. We may have replaced an inflated number with a precise one that is wrong in a different direction, and the precision makes it more dangerous in a boardroom, not less.
- —Fourteen internal repos is a lab sample. Client repos have different review cultures and CI maturity; the survey/telemetry ratio may not transfer.
- —We measured throughput, not quality. Defect escape rate over six months is the number that would settle it and we do not have it yet.
- —The published RCT literature disagrees with itself by cohort: experienced engineers on familiar code have measured slower with assistants. A single ratio hides that.
- —Telemetry is politically harder to get than a survey. If clients will not instrument, the 'right' method is one nobody uses, and the survey wins by default.
Source diversity
- Internal / Nightingale30%
- ML and SE research20%
- Public statistics15%
- Vendor10%
- Firm / Engel / buyers25%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
The same telemetry that prices client work is what makes the lab's own outcomes legible.
Instrumented agentic delivery is where the cycle-time data comes from; without it there is nothing to measure.
A measured return on client data is the one claim competitors cannot copy from a practice page.
Lead time and yield are the lab's own ROI; the method is shared.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- JPJun Park · Measurement (Nightingale)1 drop
- SKSam Kowalczyk · Research engineer · SDLC1 drop
- CDClaire Dubois · Sector owner · Banking1 drop
- OGOllie Grant · Product engineer1 drop
- ?Anonymous · Anonymous drop1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Measure ROI from telemetry and cycle time, not surveys
strength moderate · 38 citations · review 17 Dec 2026
Nightingale measurement run: coding-assistant ROI
Telemetry-derived cycle-time and rework measures on 84 licensed developers show a return on the coding-assistant licence of at least 2× annual cost, and self-reported time saved does not predict which developers' measures moved.
Self-reported productivity surveys as ROI evidence
“Everyone said it was working.” · lived 9 months
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01What does the defect escape rate do over six months on assistant-heavy repos, and does it eat the cycle-time gain?
- 02Is there a telemetry method for claims and servicing work that is as clean as git metadata?
- 03Will a client let us instrument before the engagement, or is the baseline permanently missing outside the lab?