World-model simulation for retail demand planning
“Simulated a world that had no shelves in it.”
Cause
Not selected: no published world model operates on tabular demand series, the proposed run would have been a conventional forecasting comparison with a new name, and the field is 'next' with a stated breakthrough gate that has not moved.
What was tried, what was hit, what it means
The proposal followed two large video world-model releases from Google DeepMind and a well-followed open-weight lab in May and June, and asked whether the same class of model could simulate store-level demand under promotional scenarios better than the gradient-boosted baselines the retail practice runs. It came in with strong sightings; three people placed it independently after the releases.
The scorers separated the field from the run. The field is real and stays at 'next'. The run was not: none of the released models take tabular series as input, the proposer's method was to fine-tune a sequence model on demand history and call the result a world model, and the comparison it would produce is the same one the practice runs every quarter against its own baselines. Blind scores averaged 4.1; the shadow panel scored it 3. Rejected with a reason: this run would have measured a forecasting model, not a world model, and would have told the field nothing about its gate.
The trigger that would change this is named. When a world model that ingests structured, non-visual state and is trained at scale exists, the demand-planning question becomes a Type 3 worth running, and it will run against the practice's baselines on the practice's data.
Returned notes · what the next proposer sees
Rejected 2026-07-10. The field is live at 'next'; the run was a forecasting comparison relabelled. Re-propose when a model exists that takes the input the problem actually has. The retail practice's baselines and promo-scenario dataset are the comparison to use when it does.
The difference between “someone tried this before” and “someone tried this in June, here is what they hit, here is what would need to be true now.”
Lessons
- 01Sightings measure attention, not readiness; three placements after a release is the release talking.
- 02Renaming a known method after a new field does not make it a test of the field.
- 03Reject the run, keep the field, and write down the trigger; that is what makes the rejection useful later.
Record
Try again?
Proposals that touch this entry get its returned notes attached automatically.