cavendish
SignalCandidategate · BreakthroughDistant · 7+ years

AGI

still in the foothills

AGI is not a capability the lab can watch for directly; it is a set of specific capabilities — sustained learning in deployment, reliable long-horizon agency, transfer across unrelated professional domains — each of which the lab tracks elsewhere, and 'AGI' is what we call the state where all of them have fired.

Clustered only. No lab work behind it. Cannot be cited.

Join with…

Confidence

20%unresearched

Expiry

84duntil review · 26 Nov 2026

Lead time

36moopened after mainstream — recorded honestly

Ownership

RMRohan Mehtasponsor · prospective vertical

Where it is

The field exists because the exec sponsor asked the question and the lab needed a defensible answer that was not a lab's press release. The honest position: definitions vary so much that lab timelines cannot be compared, and the evidence available to us is our own canary suite, which shows frontier models clearing more of our professional task families each release and failing in the same three ways — they do not learn from the last task, they lose the thread over long horizons, and they cannot yet be trusted to be wrong in a predictable way. Lab timelines have moved earlier every year; our capability tracker has moved steadily but not exponentially. We hold the field at Distant because the specific things we would need to see have not been seen, and we record what they are so the field is watchable rather than a mood.

Why a Quantium decision hinges on it

Every board conversation the firm's executives have with clients now contains the question, and the answer shapes ten-year decisions: what to build, what to hire, what to defer. The value of holding the field is not predicting the date; it is being able to say what evidence would move us and to show that the evidence has not arrived. That is a more useful answer than any date.

Field attributes

StateCandidate
GateBreakthrough · not yet technically possible
OriginQuestion
Measurablepartial
Audience · TLPexec
Horizondistant
Opened2 Mar 2026
Mainstream14 Mar 2023
Last validated26 Aug 2026
Sightings1

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Frontier models clear an increasing share of the lab's professional task families each release: 41% of families at expert level in 2025-09, 58% in 2026-08 (canary suite).
  • 02The three consistent failure modes across every release: no learning from the previous task, thread loss over multi-day horizons, unpredictable error distribution.
  • 03Lab-published timelines have moved earlier each year since 2023; our own capability tracker shows a steady, not accelerating, slope.
What is hype
  • 01Timelines announced by people whose valuation depends on them.
  • 02Benchmark saturation presented as generality; every saturated benchmark was written by humans who could not imagine the failure mode.
  • 03'AGI achieved internally' as a genre of post.
What would have to be true
  • 01A model that improves on the lab's held-out task families across weeks of deployment without a release — the Continuous Learning trigger.
  • 02Reliable agency over a multi-day task with a measured error distribution the lab can bound — the Self-Organising Agents and Learning Agents triggers.
  • 03Expert-level performance on three professional task families the model was not trained toward and the lab did not publish, with the failures being the kind an expert makes.
What we would do
  • 01Answer the sponsor's question with the three triggers and the canary trend, quarterly, in one page.
  • 02Never issue a date. Issue what would move us.
  • 03If two of three triggers fire inside a year: elect the field, assign an owner, and treat every Now-horizon default as up for review.

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
58%
families at expert level
Finding·band 1Assessed

Canary suite, 2026-08 release: 58% of professional task families at expert level, up from 41% a year ago

The lab's release canary across 34 professional task families. Coverage at expert level rose from 41% to 58% over twelve months on a steady slope. The three failure modes recur in every family that has not cleared. The only AGI evidence the lab generates itself.

extracted claimFrontier models clear more professional task families each release on a steady, not accelerating, slope.
Lab · Turing canary suite · Mei Tanaka22 Aug 2026
detector · bleeding edge
Regulatory·band 3Signal

APRA letter to ADIs on AI concentration and dependency risk

Asks banks to assess dependency on a small number of model providers. Not about AGI, but the regulatory shadow of the same question: what happens to the system if the capability jumps. Logged for the banking brief.

APRA5 Aug 2026
CDdropped
Post·band 2Signal

'We have achieved AGI internally' — a thread

The genre. No evidence, very widely shared, arrived at the exec sponsor within a day. Logged so the lab's brief can name it.

Social · A frontier-lab employee30 Jul 2026
RM?dropped 5
Finding·band 1Tried

Logged from Claude Code: sealed task family (APRA reporting reconciliation) — expert on the parts it had seen, novice on the part it had not

Red team ran a task family written after the model's cutoff and never published. Expert-level on sub-tasks resembling public material; failed the novel reconciliation step in a way no trained analyst would. The transfer test in miniature. Tried tier.

MCP · log_finding · Lena Fischer6 Jul 2026
LFdropped
Paper·band 1Signal

Task Horizon and Failure: How Long Can an Agent Hold the Thread?

Measures the task duration at which agents' success rate halves; it has doubled roughly every seven months for two years and currently sits around one working day. An exponential in a metric that matters, which the red team notes and the position does not yet reflect.

~7 monthshorizon doubling
arxiv.org · An evals organisation19 Jun 2026
TOAWdropped 4
Benchmark·band 2Signal

Public expert-level benchmark saturated; successor announced within a month

The third 'final exam' style benchmark to saturate in eighteen months, followed by a harder one. The pattern the hype list refers to. The lab's canary uses sealed families to avoid it.

Benchmark leaderboard28 May 2026
MTdropped 2
Paper·band 2Signal

Forty Definitions of AGI and Why Their Timelines Cannot Be Compared

Catalogues published definitions from labs, academics and forecasters. Finds no two labs use a testable definition in common. The paper behind the lab's refusal to compare timelines.

extracted claimNo two frontier labs share a testable definition of AGI.
40definitions catalogued
arxiv.org · Philosophers and an evals researcher16 Apr 2026
MTLFdropped 2
Client question·band 3Signal

'What do we do if it happens in 2028?'

Asked by a bank director after the investor letter. The question that opened the field. Answered with the three triggers; the director asked for the quarterly page.

Engel · board briefing, banking20 Mar 2026
RMdropped
Announcement·band 1Signal

Frontier lab CEO: 'AGI by 2028' in an investor letter

The latest in a series of earlier-every-year timelines. The definition used is 'can do most economically valuable cognitive work', which is not testable. Argus positioning delta: the same lab said 2030 in 2024 and 2029 in 2025.

extracted claimAGI on the lab's own definition arrives by 2028.
Foundation lab10 Feb 2026
RM?MLdropped 6
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Published lab definitions of AGI differ enough that their timelines cannot be compared to each other or to our tracker.

Assessedc-agi-4dalton-0.330 Apr 2026arxiv.org, Foundation lab
82%

Frontier models clear more of the lab's professional task families each release on a steady, not accelerating, slope.

Assessedc-agi-1dalton-0.426 Aug 2026Lab · Turing canary suite, Benchmark leaderboard
78%

Three failure modes persist across every release: no learning from the last task, long-horizon thread loss, unpredictable error distribution.

Assessedc-agi-2dalton-0.426 Aug 2026Lab · Turing canary suite, MCP · log_finding, arxiv.org
74%

Transfer to an unpublished professional task family is the most discriminating test available to the lab, and no model has passed it at expert level.

Triedc-agi-5dalton-0.48 Jul 2026MCP · log_finding, Lab · Turing canary suite
60%

AGI on any published lab definition arrives before 2030.

Signalc-agi-3dalton-0.426 Aug 2026Foundation lab, Social
20%

Position history · the diff is the product

2 validation runs against a fixed brief. Confidence 18% → 20%.

runs compare claim sets, never prose
What we said · run 2

Canary coverage up to 58% of families at expert level, steady slope. Three failure modes unchanged in kind. Red team argues the tracker is capped by design; unanswered. Still Distant.

20%
Changed since run 1
  • Frontier models clear more of the lab's professional task families each release on a steady, not accelerating, slope.
  • Three failure modes persist across every release: no learning from the last task, long-horizon thread loss, unpredictable error distribution.
  • Transfer to an unpublished professional task family is the most discriminating test available to the lab, and no model has passed it at expert level.
  • c-agi-3 ↑ 0.15 → 0.2
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

Nothing else in the graph makes every other field moot.

Timeline

committed · AW
4yr+

On our triggers, none of which has fired. Not a date.

Cost

committed · MT
low

The canary suite already runs on every release; the field costs one page a quarter.

Cost of being wrong

agent-estimated
high

Agent-estimated in both directions: early means misallocated decade; late means the firm's defaults are wrong all at once.

Demand

agent-estimated
not-measurable-here

The question arrives from boards, not engagements; Engel does not see it. Agent-drafted.

Workforce readiness

agent-estimated
low

Not a readiness question on this horizon. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Cross-sector
watch

The field is a board question in every vertical and an engagement question in none.

Mechanism · Quarterly one-page brief from the triggers; no delivery mechanism.

RM committed by Rohan Mehtacommitted
Banking
watch

Bank boards ask it most; APRA has started asking about AI concentration risk, which is the regulatory shadow of the same question.

Mechanism · Brief reused for board-level conversations; nothing in delivery.

CD committed by Claire Duboiscommitted
Health
not-relevant

Health clients ask about specific clinical capabilities, which are tracked in their own fields; the general question does not arise.

Mechanism · None.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against our position: the lab's canary suite is a lagging indicator by construction — it tests what we already know how to test — and the three 'persistent' failure modes are each being worked on directly at frontier labs with results the lab has already logged (the continual-learning paper, the long-horizon agent results). A steady slope on our tracker with an accelerating slope on theirs is the signature of a tracker that measures the wrong thing. Also: our refusal to issue a date is itself a position, and it is the position of every institution that was late.

  • The canary suite saturates at expert level per family and cannot see beyond it; the slope is capped by our test design, not the models.
  • Each of the three failure modes has a credible lab result against it this year; three separate research programmes converging is exactly what 'before 2030' would look like from here.
  • Holding at Distant with no owner is the cheapest possible position and cheap positions are usually wrong on the things that matter most.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • Evals research35%
  • Frontier labs20%
  • Commentary / social15%
  • Internal / Engel30%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Trigger · A deployed frontier model improves on the lab's held-out task families across at least four weeks without a release — weight-level or otherwise. Detected by re-running the canary suite on a fixed model version monthly and seeing the score move.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

depends onEval Harnesses

Trigger · A frontier release scores at expert level on three professional task families the lab wrote after the model's training cutoff and never published, with failures a human expert would make; the canary suite is extended with sealed families for exactly this.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

Trigger · An agent or collective completes a multi-day professional task from the lab's suite with an error distribution the lab can bound in advance — the long-horizon reliability trigger.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

enablingNon-Weight-Bound Continuous Learning

Learning in deployment is the first of the three triggers and is tracked there.

enablingSelf-Organising Agents

Reliable long-horizon agency is the second trigger; the collective form is tracked there.

compoundingEval Harnesses

The canary suite is the only evidence the lab generates itself on this question.

compoundingPost Economy

The macro half of that field is what this one produces if it fires.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsAWMTLFRM

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Is the canary suite capped by design, and would a sealed-family extension change the slope?
  2. 02If the task-horizon doubling continues, at what date does it cross the multi-day trigger, and is that inside the Next horizon?
  3. 03What does the firm actually do differently on the day two of three triggers fire?