World Models
Learned simulators will replace hand-built ones for physical environments first and commercial environments last; a world model that plans over demand or pricing is not credible until it is calibrated on held-out real trajectories at a horizon longer than a quarter.
Clustered only. No lab work behind it. Cannot be cited.
Confidence
27%unresearchedExpiry
19doverdue for reviewLead time
—not yet mainstream · opened 18 Nov 2025Ownership
Unownedcandidate — a named human electsWhere it is
Video-generation world models from Google DeepMind and others are the most visible thing in AI this year and almost none of it transfers to what Quantium does. The physical branch (robot policy training, game engines) is real and advancing. The commercial branch — a learned simulator of a market, a supply chain or a customer base you can plan against — is where the interest is and where the evidence is not. Our own attempt to use a world-model framing for retail demand planning lost to gradient-boosted trees on every horizon and is in the graveyard. The field stays open because the physical results keep improving and the gap between them and a commercial simulator is a research question, not a category error.
Why a Quantium decision hinges on it
Quantium's core product is a model of a commercial environment that a client plans against. If learned simulators ever reach calibration on real commercial trajectories, they are a direct substitute for a large part of the analytics business, not an adjacent tool. That is the reason to watch it closely and the reason to be honest that nothing we have seen does it yet.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Interactive video world models sustain coherent physical environments for minutes and are being used to train robot policies (physical branch).
- 02A learned simulator of a warehouse layout matched a hand-built discrete-event sim on throughput within 6% at a fraction of the build time.
- 03Retail demand planning: a world-model framing lost to a gradient-boosted baseline on 4-, 13- and 26-week horizons (g-world-model-sim, refuted).
- 01'Simulate your business' pitches that mean a video model of a shop floor. Visual coherence is not causal calibration.
- 02Counting frames of physically plausible video as evidence of planning capability.
- 03Any commercial world model demo that does not report held-out trajectory error against a boring baseline.
- 01A learned simulator calibrated on held-out real commercial trajectories at a horizon of one quarter or longer, beating a tuned gradient-boosted baseline with error bars.
- 02An intervention interface: the model has to answer 'what if we change price' not just 'what happens next'.
- 03A data-residency story; the training corpus for a commercial simulator is the client's most sensitive asset.
- 01Nothing on the commercial branch until the calibration trigger fires; the graveyard entry is the standing answer.
- 02Track the physical branch through Robots; a warehouse or store-layout simulator is the nearest adjacency to a Woolworths use case.
- 03If a calibrated commercial result appears: Type 3 reproduction on a Nightingale phase-1 public series (ABS retail trade) before touching client data.
Signals · 9 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Retail demand planning: world-model framing loses to gradient-boosted trees at every horizon
Learned latent-dynamics model of category demand versus a tuned LightGBM baseline on a public retail series, 4/13/26-week horizons. The world model was worse at every horizon and 40× more expensive to train. The experiment is g-world-model-sim.
extracted claimA world-model framing does not beat gradient-boosted trees for retail demand planning at any tested horizon.

'Simulate your entire business before you run it'
A simulation startup's launch post: a video world model of a retail floor, framed as a business simulator. No calibration, no baseline, no intervention interface. Kept as the canonical hype instance.

'Could we simulate a new DC layout before we build it?'
Asked by a supply-chain lead. The answer today is a hand-built DES, which the client already has. Logged because it is the first client question that maps to the layout-simulation adjacency rather than the demand branch.

Learned Layout Simulators Match Discrete-Event Models on Warehouse Throughput
Learned simulator of a fulfilment centre reaches within 6% of a hand-built DES on throughput and picks up a congestion effect the DES missed. Build time: days versus months. The nearest thing to a commercial result we have seen.
extracted claimLearned layout simulators reach DES parity on throughput at a fraction of build time.

World Models for Decision-Making: A Survey of Evaluation Practice
Surveys 140 world-model papers. Fewer than one in ten report held-out trajectory error; none in a commercial domain do so against a gradient-boosted baseline. Confirms that our calibration trigger has no published candidate.

Google DeepMind ships a minutes-long interactive world model with agent training API
Coherent interactive environments sustained for several minutes, with an API for training embodied agents inside them. The physical branch's headline release this year. Nothing in it touches commercial planning.
extracted claimPhysical world models are now good enough to serve as a training environment for embodied agents.

Analyst note: 'World models will be the next platform layer for enterprise planning'
Demand-band signal. Predicts enterprise planning platforms adopt world models within three years. Cites physical-branch results as evidence for the commercial branch. This conflation is the thing to watch for in client conversations.

ABS Retail Trade + public scanner panel: the series behind the graveyard run
Public series assembled for the demand-planning experiment. Reusable for any future reproduction without touching client data, which is the point.

'Sim-to-real is solved for manipulation; now do it for markets'
A robotics researcher's keynote arguing the commercial branch is the same problem with worse data. Interesting because it names the gap honestly: markets have no ground-truth physics to bootstrap from.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
A world-model framing for retail demand planning does not beat a tuned gradient-boosted baseline at any horizon we tested.
No published commercial world model reports held-out trajectory error against a gradient-boosted baseline at a horizon of a quarter or longer.
Physical world models are advancing on their own evidence and are already useful for robot policy training and layout simulation.
Learned layout simulators reach parity with discrete-event simulation on warehouse throughput at a fraction of build time.
Visual coherence in generated environments transfers to causal calibration in commercial planning.
Position history · the diff is the product
2 validation runs against a fixed brief. Confidence 25% → 27%.
Layout simulation is the nearest adjacency and the physical branch keeps advancing. Commercial calibration still has no published result. Red team argues our trigger is too strict; not yet changed.
- Physical world models are advancing on their own evidence and are already useful for robot policy training and layout simulation.
- No published commercial world model reports held-out trajectory error against a gradient-boosted baseline at a horizon of a quarter or longer.
- Learned layout simulators reach parity with discrete-event simulation on warehouse throughput at a fraction of build time.
- c-world-models-3 ↓ 0.2 → 0.14
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · AWA calibrated commercial simulator substitutes for a large part of the analytics business. Conditional, but the condition is the whole point.
Timeline
committed · MTPhysical branch is now; commercial calibration has no published result and our own attempt failed.
TAM
agent-estimatedAgent-estimated on the substitution case for planning analytics globally. Uncommitted and probably meaningless at this distance.
Cost
committed · JPReproducing a commercial result needs a real series and a Nightingale run; not a Type 2.
Cost of being wrong
agent-estimatedAgent-estimated: missing a real commercial simulator is an existential miss for a planning-analytics firm.
Demand
committed · DSOne warehouse-layout question in retail; nobody asks for a 'world model'.
Workforce readiness
agent-estimatedNobody in the practice has trained a simulator; the graveyard experiment was the lab's first. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Demand planning is the substitution case and the graveyard case. Layout and logistics simulation is the nearest live adjacency for Woolworths.
Mechanism · Learned store/warehouse simulator replaces hand-built DES for layout decisions; demand branch stays parked.
Grid and demand-response simulation is a natural target and AEMO already runs hand-built models; a learned one would have a customer.
Mechanism · Learned simulator of feeder-level demand as a planning tool for a distribution network.
Credit and market simulation are heavily regulated and model-risk functions will not accept a learned simulator without decades of validation.
Mechanism · None on any horizon we track.
The physical branch enables Robots; the lab tracks it here rather than there.
Mechanism · World model as the training environment for foundation-model robot policies.
Red team · the strongest case against
The strongest case against our position: the graveyard experiment tested one architecture on one series in one quarter, and the field moved a great deal since. We may be treating a negative result as a category verdict. Also, the substitution framing is wrong — a learned simulator that is 80% as good and costs 5% as much does not need to beat gradient-boosted trees to change the business.
- —g-world-model-sim used a 2025-era architecture; the 2026 releases are qualitatively different and we have not re-run.
- —Parity at lower cost is a different question from beating the baseline, and it is the question a client would ask.
- —Our calibration trigger is strict enough that it may fire only after the market has already moved; a weaker leading indicator is needed.
Source diversity
- ML research40%
- Robotics15%
- Vendor / analyst20%
- Internal / Engel25%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Trigger · x-sensing-agents or a successor produces a continuous store-telemetry trajectory dataset of at least two quarters — the minimum a commercial simulator could be calibrated and held-out tested against.
When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.
Trigger · A world-model checkpoint with an intervention interface ('what if price changes') is released with open weights and a published held-out calibration on a commercial or logistics series; Argus flags the release and the lab reproduces on an ABS retail-trade series.
When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.
Physical world models are the training environment for foundation-model robot policies.
Store telemetry from a sensing agent is the trajectory dataset a commercial simulator would be calibrated on.
An agent that plans against a simulator needs a memory of which simulated decisions held in reality.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- ?Anonymous · Anonymous drop2 drops
- DSDev Sharma · Sector owner · Retail2 drops
- JPJun Park · Measurement (Nightingale)2 drops
- MTMei Tanaka · Research lead · evals1 drop
- OGOllie Grant · Product engineer1 drop
- RMRohan Mehta · Exec sponsor1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Is parity-at-lower-cost the right trigger for the commercial branch, rather than beating the baseline?
- 02Does the warehouse-layout result hold for a supermarket floor, where the agents are customers rather than pickers?
- 03What would a commercial simulator's training corpus look like under APP 6 and client data residency?