Education
AI tutoring has crossed from promising to evidenced in narrow, well-instrumented settings; the field that matters to the firm is not the schools market but professional upskilling — including the firm's own — where the same efficacy evidence applies and the buyer already exists.
Clustered only. No lab work behind it. Cannot be cited.
Confidence
42%unresearchedExpiry
14doverdue for reviewLead time
−7moopened after mainstream — recorded honestlyOwnership
Unownedcandidate — a named human electsWhere it is
The evidence base changed in the last year. Two large randomised trials — one in higher education, one in corporate technical training — showed AI tutoring producing learning gains of a third to half a standard deviation over instructor-led baselines, with the effect concentrated in structured, assessable skills. Assessment integrity is the counterweight: universities are moving to in-person and oral assessment because take-home work no longer evidences anything. Quantium is not an education company and has no sector owner for it. What it has is a 400-person upskilling problem in agentic delivery, a government practice whose clients include education departments, and three clients who have asked whether the firm's own AI capability-building programme is something they could buy. Candidate; not elected; no owner.
Why a Quantium decision hinges on it
Two decisions hinge on it. The first is internal: the firm's ability to execute on half the fields on this board is gated by skills, and AI tutoring is the only upskilling approach with trial evidence at the scale the firm needs. The second is whether 'we upskilled ourselves and can do it for you' is a product. The schools and university market is a distraction the firm should explicitly decline.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Randomised trials in higher education and corporate technical training show AI tutoring gains of 0.3–0.5 SD on structured, assessable skills over instructor-led baselines (band-1 papers; not lab work).
- 02Take-home assessment no longer evidences learning; universities are moving to in-person, oral and process-based assessment at scale.
- 03The firm's own agentic-delivery upskilling cohort (40 people, Q2) used an AI tutor over the lab's material; completion 88% against 51% for the previous cohort's video course (tried, no control).
- 01'Personalised learning for every child.' The trial evidence is on structured, assessable skills in motivated adult learners; the extrapolation to schools is unsupported.
- 02AI tutors as a replacement for instructors. The trials with the largest effects kept the instructor and changed what they did.
- 03Detection tools for AI-written assessment. Every one tested has false-positive rates that make it unusable for a decision about a student.
- 01A client paying for the firm's capability programme rather than treating it as a pre-sales conversation.
- 02The internal cohort result holding under a control, which means running the next cohort as a Type 2 with a comparison arm.
- 03A named owner. Nobody in the lab or the firm owns education, and a candidate without one dies quietly.
- 01Keep it a candidate. Run the next internal upskilling cohort as a Type 2 with a comparison arm so the firm has its own evidence.
- 02Decline the schools and university market explicitly and record the rejection.
- 03Ask the government sector owner to log any education-department demand in Engel so the field has a demand score in a quarter.
Signals · 9 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

AI Tutoring at Scale: A Pre-Registered Randomised Trial Across 11,400 University Students
Pre-registered RCT across four universities and three disciplines. AI tutor plus instructor beat instructor alone by 0.41 SD on end-of-course assessment in quantitative subjects; no significant effect in essay-based subjects. Effect strongest for students in the bottom tercile at baseline.
extracted claimAI tutoring produces large learning gains on structured, assessable skills and none on open-ended ones.

Internal: Q2 agentic-delivery cohort with an AI tutor over lab material — 88% completion
Forty people, six weeks, AI tutor built on the lab's own agentic-delivery material. Completion 88% against 51% for the prior cohort's video course. No comparison arm, no post-test; tried tier, novelty effect likely.
extracted claimAn AI tutor over the firm's own material raises upskilling completion sharply (uncontrolled).

'Could we buy the programme you ran for your own people?'
Asked by a retail client's chief data officer after a delivery review where the firm's upskilling cohort came up. Logged informally by the retail sector owner. The third such question this half; none has become a proposal.

'The tutor taught them SQL. It did not teach them what to ask.'
Argues the RCT effects are confined to skills with a right answer and that judgement-heavy skills show no gain because the tutor cannot assess them. Kept as the strongest disconfirming voice on transfer.

Analyst: AU corporate L&D budgets shifting 20% from content licences to AI tutoring platforms by 2027
Survey of 80 AU L&D leaders. Directional; vendor-sponsored. Supports the buyer-exists half of the thesis and nothing else.

Does It Work at Work? AI Tutoring for Corporate Technical Upskilling: A Field Experiment
Field experiment in a 2,000-person technology firm's data-engineering upskilling programme. AI-tutored arm reached certification 34% faster with a 0.29 SD higher practical assessment score. The instructor's role shifted to review and unblocking.
extracted claimAI tutoring transfers to adult technical upskilling with a smaller but material effect.

AI-writing detector evaluation: false-positive rates on non-native English writers
Open evaluation of seven commercial detectors. False-positive rates of 8–24% on human-written text by non-native English speakers. No detector clears a threshold that would be defensible for an individual academic-misconduct finding.

TEQSA guidance: assessment must evidence learning in a way that is 'secure against generative AI'
Australian higher-education regulator's request-for-action. Universities are responding with in-person, oral and process-based assessment. The spend has moved from detection to redesign.

Foundation lab ships a 'learning mode' that withholds answers and asks Socratic questions
Product mode designed for education use, released after the higher-ed RCT's pre-registration became public. Marks the point the tutoring use case became mainstream; the field was opened seven months later, which is a negative lead time and recorded as one.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
AI-writing detection tools have false-positive rates that make them unusable for individual assessment decisions.
Take-home assessment no longer evidences learning; the assessment redesign, not the tutor, is where institutions are spending.
AI tutoring produces learning gains of 0.3–0.5 SD over instructor-led baselines on structured, assessable skills in adult learners.
The firm's own upskilling is the tractable application; the external education market is not one Quantium should enter.
The tutoring effect transfers to open-ended, judgement-heavy skills at the same magnitude.
Position history · the diff is the product
2 validation runs against a fixed brief. Confidence 35% → 42%.
Internal upskilling cohort result logged. The firm's own capability building is the tractable application; the external market is not. Remain a candidate; run the next cohort with a control.
- The tutoring effect transfers to open-ended, judgement-heavy skills at the same magnitude.
- The firm's own upskilling is the tractable application; the external education market is not one Quantium should enter.
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
agent-estimatedHigh for the firm's own capability; low as a market. Agent-estimated on the internal case.
Timeline
committed · AWThe internal upskilling need is this year; the trial evidence is already published.
TAM
agent-estimatedAgent-estimated global education and corporate-training spend. Almost none addressable by the firm. Uncommitted.
Cost
committed · MTA Type 2 on the next cohort with a comparison arm is a fortnight of measurement work.
Demand
committed · ABOne education-department mention in Engel this half, about assessment policy, not tutoring. Three clients asked about the firm's own programme informally.
Workforce readiness
agent-estimatedNobody in the firm has learning-science training; the lab can read the trials but not design one. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
State education departments are clients of the government practice for other things. Assessment-integrity policy is live; tutoring procurement is not.
Mechanism · Would need an education department asking for evidence synthesis or policy analytics, which the practice already does in other domains.
Every client's AI programme is gated by skills. The firm's own capability programme is the artefact, and the tutoring evidence says how to run it.
Mechanism · AI tutor over the lab's material for client teams as part of delivery; instructor role redesigned around it.
Banks run their own L&D at scale and buy content, not tutoring platforms. Nothing here changes a bank's decision.
Mechanism · None.
Red team · the strongest case against
The strongest case against: this is a candidate because it is interesting, not because the firm can act on it. The internal cohort result has no control and a strong novelty effect; the trial evidence is on skills the firm does not primarily need to teach; and 'we can upskill you' is a sentence every consultancy says and none is paid for. Meanwhile the real education market is a policy and procurement domain the firm has no standing in.
- —Completion rate is not learning. The internal cohort measured who finished, not what they could do afterwards.
- —The 0.3–0.5 SD gains are on structured, assessable skills. Agentic delivery judgement — the skill the firm is actually short of — is the kind the trials show no effect on.
- —No owner, no sector, no demand score above low. Every mechanism in the system says this field should not be elected, and it is right.
Source diversity
- Learning-sciences research40%
- Regulators / institutions15%
- Foundation lab / vendor / analyst20%
- Internal / Engel25%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
The lab's own capability programme is the first application; its effectiveness is measurable with the same instruments.
A tutoring assistant is an effective assistant with a curriculum; the interaction-design evidence is shared.
The upskilling programme is the paved road that field says citizen developers need.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- MTMei Tanaka · Research lead · evals3 drops
- AWAdam Witanowski · Lab Director (acting)2 drops
- ABAisha Bello · Sector owner · Government1 drop
- DSDev Sharma · Sector owner · Retail1 drop
- RMRohan Mehta · Exec sponsor1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Does the internal cohort's completion gain survive a comparison arm and a practical post-test?
- 02Is there a version of the firm's capability programme a client would pay for, or is it permanently pre-sales?
- 03What does the lab need to teach that the tutoring evidence says tutors cannot — and how is that taught instead?