cavendish
SignalCandidategate · AdoptionNear · 1–3 years×2 sightings

Education

AI tutoring has crossed from promising to evidenced in narrow, well-instrumented settings; the field that matters to the firm is not the schools market but professional upskilling — including the firm's own — where the same efficacy evidence applies and the buyer already exists.

Clustered only. No lab work behind it. Cannot be cited.

Join with…

Confidence

42%unresearched

Expiry

14doverdue for review

Lead time

7moopened after mainstream — recorded honestly

Ownership

Unownedcandidate — a named human elects

Where it is

The evidence base changed in the last year. Two large randomised trials — one in higher education, one in corporate technical training — showed AI tutoring producing learning gains of a third to half a standard deviation over instructor-led baselines, with the effect concentrated in structured, assessable skills. Assessment integrity is the counterweight: universities are moving to in-person and oral assessment because take-home work no longer evidences anything. Quantium is not an education company and has no sector owner for it. What it has is a 400-person upskilling problem in agentic delivery, a government practice whose clients include education departments, and three clients who have asked whether the firm's own AI capability-building programme is something they could buy. Candidate; not elected; no owner.

Why a Quantium decision hinges on it

Two decisions hinge on it. The first is internal: the firm's ability to execute on half the fields on this board is gated by skills, and AI tutoring is the only upskilling approach with trial evidence at the scale the firm needs. The second is whether 'we upskilled ourselves and can do it for you' is a product. The schools and university market is a distraction the firm should explicitly decline.

Field attributes

StateCandidate
GateAdoption · ready — blocked by trust, regulation, procurement or change capacity
OriginSignal
Measurablepartial
Audience · TLPexec
Horizonnear
Opened28 May 2026
Mainstream4 Nov 2025
Last validated16 Jul 2026
Sightings2

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Randomised trials in higher education and corporate technical training show AI tutoring gains of 0.3–0.5 SD on structured, assessable skills over instructor-led baselines (band-1 papers; not lab work).
  • 02Take-home assessment no longer evidences learning; universities are moving to in-person, oral and process-based assessment at scale.
  • 03The firm's own agentic-delivery upskilling cohort (40 people, Q2) used an AI tutor over the lab's material; completion 88% against 51% for the previous cohort's video course (tried, no control).
What is hype
  • 01'Personalised learning for every child.' The trial evidence is on structured, assessable skills in motivated adult learners; the extrapolation to schools is unsupported.
  • 02AI tutors as a replacement for instructors. The trials with the largest effects kept the instructor and changed what they did.
  • 03Detection tools for AI-written assessment. Every one tested has false-positive rates that make it unusable for a decision about a student.
What would have to be true
  • 01A client paying for the firm's capability programme rather than treating it as a pre-sales conversation.
  • 02The internal cohort result holding under a control, which means running the next cohort as a Type 2 with a comparison arm.
  • 03A named owner. Nobody in the lab or the firm owns education, and a candidate without one dies quietly.
What we would do
  • 01Keep it a candidate. Run the next internal upskilling cohort as a Type 2 with a comparison arm so the firm has its own evidence.
  • 02Decline the schools and university market explicitly and record the rejection.
  • 03Ask the government sector owner to log any education-department demand in Engel so the field has a demand score in a quarter.

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
0.41 SD
effect size
Paper·band 1Signal

AI Tutoring at Scale: A Pre-Registered Randomised Trial Across 11,400 University Students

Pre-registered RCT across four universities and three disciplines. AI tutor plus instructor beat instructor alone by 0.41 SD on end-of-course assessment in quantitative subjects; no significant effect in essay-based subjects. Effect strongest for students in the bottom tercile at baseline.

extracted claimAI tutoring produces large learning gains on structured, assessable skills and none on open-ended ones.
arxiv.org · A multi-university learning-sciences consortium9 Mar 2026
AWMTdropped 3
88% vs 51%
completion
Finding·band 1Tried

Internal: Q2 agentic-delivery cohort with an AI tutor over lab material — 88% completion

Forty people, six weeks, AI tutor built on the lab's own agentic-delivery material. Completion 88% against 51% for the prior cohort's video course. No comparison arm, no post-test; tried tier, novelty effect likely.

extracted claimAn AI tutor over the firm's own material raises upskilling completion sharply (uncontrolled).
Lab · capability programme · Mei Tanaka10 Jul 2026
MTdropped
Client question·band 3Signal

'Could we buy the programme you ran for your own people?'

Asked by a retail client's chief data officer after a delivery review where the firm's upskilling cohort came up. Logged informally by the retail sector owner. The third such question this half; none has become a proposal.

Engel · retail engagement30 Jun 2026
DSdropped 3
Post·band 2Signal

'The tutor taught them SQL. It did not teach them what to ask.'

Argues the RCT effects are confined to skills with a right answer and that judgement-heavy skills show no gain because the tutor cannot assess them. Kept as the strongest disconfirming voice on transfer.

Substack · A learning-science researcher14 Jun 2026
AWdropped 2
80
respondents
Analyst·band 3Signal

Analyst: AU corporate L&D budgets shifting 20% from content licences to AI tutoring platforms by 2027

Survey of 80 AU L&D leaders. Directional; vendor-sponsored. Supports the buyer-exists half of the thesis and nothing else.

Industry analyst briefing26 May 2026
RMdropped
−34%
time to certification
Paper·band 1Signal

Does It Work at Work? AI Tutoring for Corporate Technical Upskilling: A Field Experiment

Field experiment in a 2,000-person technology firm's data-engineering upskilling programme. AI-tutored arm reached certification 34% faster with a 0.29 SD higher practical assessment score. The instructor's role shifted to review and unblocking.

extracted claimAI tutoring transfers to adult technical upskilling with a smaller but material effect.
arxiv.org · Vasquez, Ng et al.12 May 2026
MTdropped 2
8–24%
false positives
Benchmark·band 2Signal

AI-writing detector evaluation: false-positive rates on non-native English writers

Open evaluation of seven commercial detectors. False-positive rates of 8–24% on human-written text by non-native English speakers. No detector clears a threshold that would be defensible for an individual academic-misconduct finding.

github.com22 Apr 2026
detector · early adoption 2
Regulatory·band 2Signal

TEQSA guidance: assessment must evidence learning in a way that is 'secure against generative AI'

Australian higher-education regulator's request-for-action. Universities are responding with in-person, oral and process-based assessment. The spend has moved from detection to redesign.

TEQSA18 Feb 2026
ABdropped 2
Release·band 1Signal

Foundation lab ships a 'learning mode' that withholds answers and asks Socratic questions

Product mode designed for education use, released after the higher-ed RCT's pre-registration became public. Marks the point the tutoring use case became mainstream; the field was opened seven months later, which is a negative lead time and recorded as one.

Anthropic4 Nov 2025
detector · bleeding edge 2
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

AI-writing detection tools have false-positive rates that make them unusable for individual assessment decisions.

Signalc-education-5dalton-0.330 Apr 2026github.com
80%

Take-home assessment no longer evidences learning; the assessment redesign, not the tutor, is where institutions are spending.

Signalc-education-3dalton-0.422 Jun 2026TEQSA, Industry analyst briefing
75%

AI tutoring produces learning gains of 0.3–0.5 SD over instructor-led baselines on structured, assessable skills in adult learners.

Signalc-education-1dalton-0.416 Jul 2026arxiv.org, arxiv.org
72%

The firm's own upskilling is the tractable application; the external education market is not one Quantium should enter.

Signalc-education-4dalton-0.416 Jul 2026Lab · capability programme, Engel · retail engagement
60%

The tutoring effect transfers to open-ended, judgement-heavy skills at the same magnitude.

Signalc-education-2dalton-0.416 Jul 2026arxiv.org, Substack
18%

Position history · the diff is the product

2 validation runs against a fixed brief. Confidence 35% → 42%.

runs compare claim sets, never prose
What we said · run 2

Internal upskilling cohort result logged. The firm's own capability building is the tractable application; the external market is not. Remain a candidate; run the next cohort with a control.

42%
Changed since run 1
  • The tutoring effect transfers to open-ended, judgement-heavy skills at the same magnitude.
  • The firm's own upskilling is the tractable application; the external education market is not one Quantium should enter.
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

agent-estimated
medium

High for the firm's own capability; low as a market. Agent-estimated on the internal case.

Timeline

committed · AW
0–18mo

The internal upskilling need is this year; the trial evidence is already published.

TAM

agent-estimated
>$10B

Agent-estimated global education and corporate-training spend. Almost none addressable by the firm. Uncommitted.

Cost

committed · MT
low

A Type 2 on the next cohort with a comparison arm is a fortnight of measurement work.

Demand

committed · AB
low

One education-department mention in Engel this half, about assessment policy, not tutoring. Three clients asked about the firm's own programme informally.

Workforce readiness

agent-estimated
low

Nobody in the firm has learning-science training; the lab can read the trials but not design one. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Government
watch

State education departments are clients of the government practice for other things. Assessment-integrity policy is live; tutoring procurement is not.

Mechanism · Would need an education department asking for evidence synthesis or policy analytics, which the practice already does in other domains.

AB committed by Aisha Bellocommitted
Cross-sector
relevant

Every client's AI programme is gated by skills. The firm's own capability programme is the artefact, and the tutoring evidence says how to run it.

Mechanism · AI tutor over the lab's material for client teams as part of delivery; instructor role redesigned around it.

Agent draft · awaiting a sector owneragent-estimated
Banking
not-relevant

Banks run their own L&D at scale and buy content, not tutoring platforms. Nothing here changes a bank's decision.

Mechanism · None.

CD committed by Claire Duboiscommitted

Red team · the strongest case against

The strongest case against: this is a candidate because it is interesting, not because the firm can act on it. The internal cohort result has no control and a strong novelty effect; the trial evidence is on skills the firm does not primarily need to teach; and 'we can upskill you' is a sentence every consultancy says and none is paid for. Meanwhile the real education market is a policy and procurement domain the firm has no standing in.

  • Completion rate is not learning. The internal cohort measured who finished, not what they could do afterwards.
  • The 0.3–0.5 SD gains are on structured, assessable skills. Agentic delivery judgement — the skill the firm is actually short of — is the kind the trials show no effect on.
  • No owner, no sector, no demand score above low. Every mechanism in the system says this field should not be elected, and it is right.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis in doubt

Source diversity

  • Learning-sciences research40%
  • Regulators / institutions15%
  • Foundation lab / vendor / analyst20%
  • Internal / Engel25%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsAWMTAB

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Does the internal cohort's completion gain survive a comparison arm and a practical post-test?
  2. 02Is there a version of the firm's capability programme a client would pay for, or is it permanently pre-sales?
  3. 03What does the lab need to teach that the tutoring evidence says tutors cannot — and how is that taught instead?