AGI
still in the foothills
AGI is not a capability the lab can watch for directly; it is a set of specific capabilities — sustained learning in deployment, reliable long-horizon agency, transfer across unrelated professional domains — each of which the lab tracks elsewhere, and 'AGI' is what we call the state where all of them have fired.
Clustered only. No lab work behind it. Cannot be cited.
Confidence
20%unresearchedExpiry
84duntil review · 26 Nov 2026Lead time
−36moopened after mainstream — recorded honestlyOwnership
RMRohan Mehtasponsor · prospective verticalWhere it is
The field exists because the exec sponsor asked the question and the lab needed a defensible answer that was not a lab's press release. The honest position: definitions vary so much that lab timelines cannot be compared, and the evidence available to us is our own canary suite, which shows frontier models clearing more of our professional task families each release and failing in the same three ways — they do not learn from the last task, they lose the thread over long horizons, and they cannot yet be trusted to be wrong in a predictable way. Lab timelines have moved earlier every year; our capability tracker has moved steadily but not exponentially. We hold the field at Distant because the specific things we would need to see have not been seen, and we record what they are so the field is watchable rather than a mood.
Why a Quantium decision hinges on it
Every board conversation the firm's executives have with clients now contains the question, and the answer shapes ten-year decisions: what to build, what to hire, what to defer. The value of holding the field is not predicting the date; it is being able to say what evidence would move us and to show that the evidence has not arrived. That is a more useful answer than any date.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Frontier models clear an increasing share of the lab's professional task families each release: 41% of families at expert level in 2025-09, 58% in 2026-08 (canary suite).
- 02The three consistent failure modes across every release: no learning from the previous task, thread loss over multi-day horizons, unpredictable error distribution.
- 03Lab-published timelines have moved earlier each year since 2023; our own capability tracker shows a steady, not accelerating, slope.
- 01Timelines announced by people whose valuation depends on them.
- 02Benchmark saturation presented as generality; every saturated benchmark was written by humans who could not imagine the failure mode.
- 03'AGI achieved internally' as a genre of post.
- 01A model that improves on the lab's held-out task families across weeks of deployment without a release — the Continuous Learning trigger.
- 02Reliable agency over a multi-day task with a measured error distribution the lab can bound — the Self-Organising Agents and Learning Agents triggers.
- 03Expert-level performance on three professional task families the model was not trained toward and the lab did not publish, with the failures being the kind an expert makes.
- 01Answer the sponsor's question with the three triggers and the canary trend, quarterly, in one page.
- 02Never issue a date. Issue what would move us.
- 03If two of three triggers fire inside a year: elect the field, assign an owner, and treat every Now-horizon default as up for review.
Signals · 9 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.
Canary suite, 2026-08 release: 58% of professional task families at expert level, up from 41% a year ago
The lab's release canary across 34 professional task families. Coverage at expert level rose from 41% to 58% over twelve months on a steady slope. The three failure modes recur in every family that has not cleared. The only AGI evidence the lab generates itself.
extracted claimFrontier models clear more professional task families each release on a steady, not accelerating, slope.
APRA letter to ADIs on AI concentration and dependency risk
Asks banks to assess dependency on a small number of model providers. Not about AGI, but the regulatory shadow of the same question: what happens to the system if the capability jumps. Logged for the banking brief.
'We have achieved AGI internally' — a thread
The genre. No evidence, very widely shared, arrived at the exec sponsor within a day. Logged so the lab's brief can name it.
Logged from Claude Code: sealed task family (APRA reporting reconciliation) — expert on the parts it had seen, novice on the part it had not
Red team ran a task family written after the model's cutoff and never published. Expert-level on sub-tasks resembling public material; failed the novel reconciliation step in a way no trained analyst would. The transfer test in miniature. Tried tier.
Task Horizon and Failure: How Long Can an Agent Hold the Thread?
Measures the task duration at which agents' success rate halves; it has doubled roughly every seven months for two years and currently sits around one working day. An exponential in a metric that matters, which the red team notes and the position does not yet reflect.
Public expert-level benchmark saturated; successor announced within a month
The third 'final exam' style benchmark to saturate in eighteen months, followed by a harder one. The pattern the hype list refers to. The lab's canary uses sealed families to avoid it.
Forty Definitions of AGI and Why Their Timelines Cannot Be Compared
Catalogues published definitions from labs, academics and forecasters. Finds no two labs use a testable definition in common. The paper behind the lab's refusal to compare timelines.
extracted claimNo two frontier labs share a testable definition of AGI.
'What do we do if it happens in 2028?'
Asked by a bank director after the investor letter. The question that opened the field. Answered with the three triggers; the director asked for the quarterly page.
Frontier lab CEO: 'AGI by 2028' in an investor letter
The latest in a series of earlier-every-year timelines. The definition used is 'can do most economically valuable cognitive work', which is not testable. Argus positioning delta: the same lab said 2030 in 2024 and 2029 in 2025.
extracted claimAGI on the lab's own definition arrives by 2028.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
Published lab definitions of AGI differ enough that their timelines cannot be compared to each other or to our tracker.
Frontier models clear more of the lab's professional task families each release on a steady, not accelerating, slope.
Three failure modes persist across every release: no learning from the last task, long-horizon thread loss, unpredictable error distribution.
Transfer to an unpublished professional task family is the most discriminating test available to the lab, and no model has passed it at expert level.
AGI on any published lab definition arrives before 2030.
Position history · the diff is the product
2 validation runs against a fixed brief. Confidence 18% → 20%.
Canary coverage up to 58% of families at expert level, steady slope. Three failure modes unchanged in kind. Red team argues the tracker is capped by design; unanswered. Still Distant.
- Frontier models clear more of the lab's professional task families each release on a steady, not accelerating, slope.
- Three failure modes persist across every release: no learning from the last task, long-horizon thread loss, unpredictable error distribution.
- Transfer to an unpublished professional task family is the most discriminating test available to the lab, and no model has passed it at expert level.
- c-agi-3 ↑ 0.15 → 0.2
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · AWNothing else in the graph makes every other field moot.
Timeline
committed · AWOn our triggers, none of which has fired. Not a date.
Cost
committed · MTThe canary suite already runs on every release; the field costs one page a quarter.
Cost of being wrong
agent-estimatedAgent-estimated in both directions: early means misallocated decade; late means the firm's defaults are wrong all at once.
Demand
agent-estimatedThe question arrives from boards, not engagements; Engel does not see it. Agent-drafted.
Workforce readiness
agent-estimatedNot a readiness question on this horizon. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
The field is a board question in every vertical and an engagement question in none.
Mechanism · Quarterly one-page brief from the triggers; no delivery mechanism.
Bank boards ask it most; APRA has started asking about AI concentration risk, which is the regulatory shadow of the same question.
Mechanism · Brief reused for board-level conversations; nothing in delivery.
Health clients ask about specific clinical capabilities, which are tracked in their own fields; the general question does not arise.
Mechanism · None.
Red team · the strongest case against
The strongest case against our position: the lab's canary suite is a lagging indicator by construction — it tests what we already know how to test — and the three 'persistent' failure modes are each being worked on directly at frontier labs with results the lab has already logged (the continual-learning paper, the long-horizon agent results). A steady slope on our tracker with an accelerating slope on theirs is the signature of a tracker that measures the wrong thing. Also: our refusal to issue a date is itself a position, and it is the position of every institution that was late.
- —The canary suite saturates at expert level per family and cannot see beyond it; the slope is capped by our test design, not the models.
- —Each of the three failure modes has a credible lab result against it this year; three separate research programmes converging is exactly what 'before 2030' would look like from here.
- —Holding at Distant with no owner is the cheapest possible position and cheap positions are usually wrong on the things that matter most.
Source diversity
- Evals research35%
- Frontier labs20%
- Commentary / social15%
- Internal / Engel30%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Trigger · A deployed frontier model improves on the lab's held-out task families across at least four weeks without a release — weight-level or otherwise. Detected by re-running the canary suite on a fixed model version monthly and seeing the score move.
When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.
Trigger · A frontier release scores at expert level on three professional task families the lab wrote after the model's training cutoff and never published, with failures a human expert would make; the canary suite is extended with sealed families for exactly this.
When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.
Trigger · An agent or collective completes a multi-day professional task from the lab's suite with an error distribution the lab can bound in advance — the long-horizon reliability trigger.
When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.
Learning in deployment is the first of the three triggers and is tracked there.
Reliable long-horizon agency is the second trigger; the collective form is tracked there.
The canary suite is the only evidence the lab generates itself on this question.
The macro half of that field is what this one produces if it fires.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- RMRohan Mehta · Exec sponsor3 drops
- ?Anonymous · Anonymous drop2 drops
- MTMei Tanaka · Research lead · evals2 drops
- LFLena Fischer · Red team & assurance2 drops
- MLMarcus Lee · Delivery lead · Telco1 drop
- TOTom Okafor · Research engineer · agents1 drop
- AWAdam Witanowski · Lab Director (acting)1 drop
- CDClaire Dubois · Sector owner · Banking1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Is the canary suite capped by design, and would a sealed-family extension change the slope?
- 02If the task-horizon doubling continues, at what date does it cross the multi-day trigger, and is that inside the Next horizon?
- 03What does the firm actually do differently on the day two of three triggers fire?