cavendish
SignalCandidategate · BreakthroughNext · 3–7 years

World Models

Learned simulators will replace hand-built ones for physical environments first and commercial environments last; a world model that plans over demand or pricing is not credible until it is calibrated on held-out real trajectories at a horizon longer than a quarter.

Clustered only. No lab work behind it. Cannot be cited.

Join with…

Confidence

27%unresearched

Expiry

19doverdue for review

Lead time

not yet mainstream · opened 18 Nov 2025

Ownership

Unownedcandidate — a named human elects

Where it is

Video-generation world models from Google DeepMind and others are the most visible thing in AI this year and almost none of it transfers to what Quantium does. The physical branch (robot policy training, game engines) is real and advancing. The commercial branch — a learned simulator of a market, a supply chain or a customer base you can plan against — is where the interest is and where the evidence is not. Our own attempt to use a world-model framing for retail demand planning lost to gradient-boosted trees on every horizon and is in the graveyard. The field stays open because the physical results keep improving and the gap between them and a commercial simulator is a research question, not a category error.

Why a Quantium decision hinges on it

Quantium's core product is a model of a commercial environment that a client plans against. If learned simulators ever reach calibration on real commercial trajectories, they are a direct substitute for a large part of the analytics business, not an adjacent tool. That is the reason to watch it closely and the reason to be honest that nothing we have seen does it yet.

Field attributes

StateCandidate
GateBreakthrough · not yet technically possible
OriginSignal
Measurablepartial
Audience · TLPlab
Horizonnext
Opened18 Nov 2025
Mainstreamnot yet
Last validated3 Jul 2026
Sightings1

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Interactive video world models sustain coherent physical environments for minutes and are being used to train robot policies (physical branch).
  • 02A learned simulator of a warehouse layout matched a hand-built discrete-event sim on throughput within 6% at a fraction of the build time.
  • 03Retail demand planning: a world-model framing lost to a gradient-boosted baseline on 4-, 13- and 26-week horizons (g-world-model-sim, refuted).
What is hype
  • 01'Simulate your business' pitches that mean a video model of a shop floor. Visual coherence is not causal calibration.
  • 02Counting frames of physically plausible video as evidence of planning capability.
  • 03Any commercial world model demo that does not report held-out trajectory error against a boring baseline.
What would have to be true
  • 01A learned simulator calibrated on held-out real commercial trajectories at a horizon of one quarter or longer, beating a tuned gradient-boosted baseline with error bars.
  • 02An intervention interface: the model has to answer 'what if we change price' not just 'what happens next'.
  • 03A data-residency story; the training corpus for a commercial simulator is the client's most sensitive asset.
What we would do
  • 01Nothing on the commercial branch until the calibration trigger fires; the graveyard entry is the standing answer.
  • 02Track the physical branch through Robots; a warehouse or store-layout simulator is the nearest adjacency to a Woolworths use case.
  • 03If a calibrated commercial result appears: Type 3 reproduction on a Nightingale phase-1 public series (ABS retail trade) before touching client data.

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
+3.1 pts worse
MAPE vs baseline
Finding·band 1Tested

Retail demand planning: world-model framing loses to gradient-boosted trees at every horizon

Learned latent-dynamics model of category demand versus a tuned LightGBM baseline on a public retail series, 4/13/26-week horizons. The world model was worse at every horizon and 40× more expensive to train. The experiment is g-world-model-sim.

extracted claimA world-model framing does not beat gradient-boosted trees for retail demand planning at any tested horizon.
Lab · graveyard experiment · Jun Park28 Mar 2026
detector · bleeding edge
Post·band 2Signal

'Simulate your entire business before you run it'

A simulation startup's launch post: a video world model of a retail floor, framed as a business simulator. No calibration, no baseline, no intervention interface. Kept as the canonical hype instance.

Vendor blog16 Jun 2026
?dropped 3
Client question·band 3Signal

'Could we simulate a new DC layout before we build it?'

Asked by a supply-chain lead. The answer today is a hand-built DES, which the client already has. Logged because it is the first client question that maps to the layout-simulation adjacency rather than the demand branch.

Engel · retail engagement11 Jun 2026
DSdropped
6%
throughput error
Paper·band 1Signal

Learned Layout Simulators Match Discrete-Event Models on Warehouse Throughput

Learned simulator of a fulfilment centre reaches within 6% of a hand-built DES on throughput and picks up a congestion effect the DES missed. Build time: days versus months. The nearest thing to a commercial result we have seen.

extracted claimLearned layout simulators reach DES parity on throughput at a fraction of build time.
arxiv.org · Osei, Brandt et al.9 Jun 2026
DSdropped 2
<10%
papers reporting held-out error
Paper·band 1Signal

World Models for Decision-Making: A Survey of Evaluation Practice

Surveys 140 world-model papers. Fewer than one in ten report held-out trajectory error; none in a commercial domain do so against a gradient-boosted baseline. Confirms that our calibration trigger has no published candidate.

arxiv.org · Petrova, Ng et al.27 May 2026
JPdropped
minutes
coherence
Release·band 1Signal

Google DeepMind ships a minutes-long interactive world model with agent training API

Coherent interactive environments sustained for several minutes, with an API for training embodied agents inside them. The physical branch's headline release this year. Nothing in it touches commercial planning.

extracted claimPhysical world models are now good enough to serve as a training environment for embodied agents.
Google DeepMind19 May 2026
MTOG?dropped 5
Analyst·band 3Signal

Analyst note: 'World models will be the next platform layer for enterprise planning'

Demand-band signal. Predicts enterprise planning platforms adopt world models within three years. Cites physical-branch results as evidence for the commercial branch. This conflation is the thing to watch for in client conversations.

Industry analyst8 Apr 2026
RMdropped 2
Dataset·band 2Signal

ABS Retail Trade + public scanner panel: the series behind the graveyard run

Public series assembled for the demand-planning experiment. Reusable for any future reproduction without touching client data, which is the point.

ABS / Nightingale phase 112 Feb 2026
JPdropped
Talk·band 2Signal

'Sim-to-real is solved for manipulation; now do it for markets'

A robotics researcher's keynote arguing the commercial branch is the same problem with worse data. Interesting because it names the gap honestly: markets have no ground-truth physics to bootstrap from.

CoRL6 Nov 2025
detector · early adoption
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

A world-model framing for retail demand planning does not beat a tuned gradient-boosted baseline at any horizon we tested.

Testedc-world-models-1dalton-0.330 Mar 2026Lab · graveyard experiment, ABS / Nightingale phase 1
83%

No published commercial world model reports held-out trajectory error against a gradient-boosted baseline at a horizon of a quarter or longer.

Assessedc-world-models-4dalton-0.43 Jul 2026arxiv.org, arxiv.org
79%

Physical world models are advancing on their own evidence and are already useful for robot policy training and layout simulation.

Assessedc-world-models-2dalton-0.43 Jul 2026Google DeepMind, arxiv.org, CoRL
76%

Learned layout simulators reach parity with discrete-event simulation on warehouse throughput at a fraction of build time.

Signalc-world-models-5dalton-0.412 Jun 2026arxiv.org, Engel · retail engagement
55%

Visual coherence in generated environments transfers to causal calibration in commercial planning.

Signalc-world-models-3dalton-0.420 Jun 2026Vendor blog, Industry analyst
14%

Position history · the diff is the product

2 validation runs against a fixed brief. Confidence 25% → 27%.

runs compare claim sets, never prose
What we said · run 2

Layout simulation is the nearest adjacency and the physical branch keeps advancing. Commercial calibration still has no published result. Red team argues our trigger is too strict; not yet changed.

27%
Changed since run 1
  • Physical world models are advancing on their own evidence and are already useful for robot policy training and layout simulation.
  • No published commercial world model reports held-out trajectory error against a gradient-boosted baseline at a horizon of a quarter or longer.
  • Learned layout simulators reach parity with discrete-event simulation on warehouse throughput at a fraction of build time.
  • c-world-models-3 ↓ 0.2 → 0.14
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

A calibrated commercial simulator substitutes for a large part of the analytics business. Conditional, but the condition is the whole point.

Timeline

committed · MT
4yr+

Physical branch is now; commercial calibration has no published result and our own attempt failed.

TAM

agent-estimated
>$10B

Agent-estimated on the substitution case for planning analytics globally. Uncommitted and probably meaningless at this distance.

Cost

committed · JP
medium

Reproducing a commercial result needs a real series and a Nightingale run; not a Type 2.

Cost of being wrong

agent-estimated
high

Agent-estimated: missing a real commercial simulator is an existential miss for a planning-analytics firm.

Demand

committed · DS
low

One warehouse-layout question in retail; nobody asks for a 'world model'.

Workforce readiness

agent-estimated
low

Nobody in the practice has trained a simulator; the graveyard experiment was the lab's first. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Retail & FMCG
watch

Demand planning is the substitution case and the graveyard case. Layout and logistics simulation is the nearest live adjacency for Woolworths.

Mechanism · Learned store/warehouse simulator replaces hand-built DES for layout decisions; demand branch stays parked.

DS committed by Dev Sharmacommitted
Energy & Utilities
watch

Grid and demand-response simulation is a natural target and AEMO already runs hand-built models; a learned one would have a customer.

Mechanism · Learned simulator of feeder-level demand as a planning tool for a distribution network.

Agent draft · awaiting a sector owneragent-estimated
Banking
not-relevant

Credit and market simulation are heavily regulated and model-risk functions will not accept a learned simulator without decades of validation.

Mechanism · None on any horizon we track.

CD committed by Claire Duboiscommitted
Cross-sector
watch

The physical branch enables Robots; the lab tracks it here rather than there.

Mechanism · World model as the training environment for foundation-model robot policies.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against our position: the graveyard experiment tested one architecture on one series in one quarter, and the field moved a great deal since. We may be treating a negative result as a category verdict. Also, the substitution framing is wrong — a learned simulator that is 80% as good and costs 5% as much does not need to beat gradient-boosted trees to change the business.

  • g-world-model-sim used a 2025-era architecture; the 2026 releases are qualitatively different and we have not re-run.
  • Parity at lower cost is a different question from beating the baseline, and it is the question a client would ask.
  • Our calibration trigger is strict enough that it may fire only after the market has already moved; a weaker leading indicator is needed.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • ML research40%
  • Robotics15%
  • Vendor / analyst20%
  • Internal / Engel25%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

depends onSensing Agents

Trigger · x-sensing-agents or a successor produces a continuous store-telemetry trajectory dataset of at least two quarters — the minimum a commercial simulator could be calibrated and held-out tested against.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

Trigger · A world-model checkpoint with an intervention interface ('what if price changes') is released with open weights and a published held-out calibration on a commercial or logistics series; Argus flags the release and the lab reproduces on an ABS retail-trade series.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

enablingRobots

Physical world models are the training environment for foundation-model robot policies.

compoundingSensing Agents

Store telemetry from a sensing agent is the trajectory dataset a commercial simulator would be calibrated on.

compoundingLearning Agents

An agent that plans against a simulator needs a memory of which simulated decisions held in reality.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsJPDSMT

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Is parity-at-lower-cost the right trigger for the commercial branch, rather than beating the baseline?
  2. 02Does the warehouse-layout result hold for a supermarket floor, where the agents are customers rather than pickers?
  3. 03What would a commercial simulator's training corpus look like under APP 6 and client data residency?