cavendish
SignalCandidategate · ReliabilityNext · 3–7 years

Robots

Foundation-model robotics will reach AU logistics, retail and mining as a data and integration problem before it reaches them as a hardware problem; the firm's relevance is the telemetry and decision layer, not the robot.

Clustered only. No lab work behind it. Cannot be cited.

Join with…

Confidence

38%unresearched

Expiry

34doverdue for review

Lead time

6moopened after mainstream — recorded honestly

Ownership

Unownedcandidate — a named human elects

Where it is

Vision-language-action models made general-purpose manipulation demos routine in 2026 and made none of them dependable. Published task success rates on novel objects sit around 70–85% in the lab; a warehouse needs four nines on the tasks it already automates. The economic reality in AU is that Woolworths' automated DCs, the Pilbara's autonomous haul fleets and Amazon's fulfilment robots all run on narrow, non-foundation systems that work. Where foundation models enter is the long tail: exception handling, novel-object picking, natural-language tasking. That is a software layer sitting on telemetry and decision data, which is closer to what Quantium does than the hardware is.

Why a Quantium decision hinges on it

Woolworths is the owner and the largest AU operator of automated retail logistics. If foundation-model robotics changes the economics of the long tail — the 15% of picks that still go to humans — the decision data around it is a Quantium-shaped problem. The firm does not build robots and should not; it should know when the layer above the robot becomes a modelling problem.

Field attributes

StateCandidate
GateReliability · possible, not yet dependable enough
OriginObservation
Measurablepartial
Audience · TLPexec
Horizonnext
Opened7 Apr 2026
Mainstream14 Oct 2025
Last validated25 Jun 2026
Sightings1

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Open-weight VLA models perform novel-object pick-and-place at 70–85% success in controlled settings across several robot bodies.
  • 02Narrow automation in AU logistics and mining is mature, profitable and not foundation-model based.
  • 03Natural-language tasking of a warehouse robot works in demos; no published deployment reports exception rates.
What is hype
  • 01Humanoid form factor. The body is the story and the story is not the economics.
  • 02Demo success rates on curated object sets presented as deployment readiness.
  • 03'General-purpose' as a claim about a model that has been trained on one lab's robots.
What would have to be true
  • 01A published deployment where a foundation-model policy handles the long-tail exceptions in a live DC at a success rate the operator accepts, with the exception rate reported.
  • 02Sim-to-real transfer from a learned world model to a real robot on a task we could specify, at a published success rate.
  • 03A telemetry and decision layer that the operator will let a third party model — the Quantium entry point.
What we would do
  • 01Nothing hardware-side, ever. The field is watched for the software layer.
  • 02Ask Woolworths' DC automation team what their long-tail exception rate is and whether they log it; that number is the demand signal we lack.
  • 03If the deployment trigger fires: Type 3 survey run into the decision-data layer around a live deployment, with the retail sector owner.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
81%
novel-object success
Benchmark·band 1Signal

Open cross-embodiment manipulation benchmark: best VLA at 81% on novel objects

Standardised novel-object pick-and-place across six robot bodies. Best open-weight policy 81%, best closed 84%; humans 99%+. The gap to deployment is this number, and it is public.

extracted claimFoundation-model manipulation sits at 70–85% success on novel objects.
Benchmark leaderboard3 Jun 2026
TOMTdropped 3
Client question·band 3Signal

'Can you model which exceptions we should automate next?'

Asked by a supply-chain analytics lead on the owner's side. Not a robotics question; a prioritisation question with robot options. Logged as the field's one real demand signal.

Engel · retail (owner) engagement19 Jun 2026
DSdropped
Paper·band 1Signal

Training Manipulation Policies Inside Learned World Models: Sim-to-Real at Scale

Policies trained mostly in a learned simulator, fine-tuned on 200 real demos, reach 76% on the cross-embodiment benchmark. Not yet parity with real-data training but a 10× cut in real-data collection. The bridge from World Models to this field.

200real demos
arxiv.org · A frontier-lab robotics team17 Jun 2026
MTdropped 2
~15%
manual picks
Talk·band 3Signal

'The last 15%: what still goes to humans in an automated DC'

An automation lead at a large AU grocer's DC describes the exception tail: damaged packaging, novel SKUs, mixed pallets. Says the number is logged but never modelled. This is the closest thing to a demand signal the field has.

extracted claimThe long tail of DC exceptions is logged, unmodelled, and is where foundation models would enter.
AU supply-chain conference21 May 2026
DSdropped
Post·band 2Signal

'The humanoid is the wrong shape for every job except the demo'

Practitioner argument that wheeled, fixed and gantry systems beat humanoids on cost per pick and always will inside a building designed for them. Widely shared among operators. The counter-signal to the humanoid press cycle.

Personal blog · A warehouse-automation engineer30 Apr 2026
DSdropped 2
<1k
demos to adapt
Release·band 1Signal

Open-weight VLA model released with fine-tuning recipe for new robot bodies

First open-weight policy with a documented recipe for adapting to a new body in under a thousand demonstrations. Makes the field reproducible outside the labs that own robots.

extracted claimAdapting a foundation policy to a new robot body now needs hundreds, not tens of thousands, of demonstrations.
Foundation lab22 Apr 2026
TOdropped 2
Job posting·band 2Signal

Pilbara operator hiring 'Autonomy Decision Analytics Lead' ×2

Two roles for analytics over autonomous fleet data. Argus inference: the operator agrees the next decade is decisions, not driving, and is staffing it in-house rather than through a consultancy.

2openings
Corporate careers page15 Apr 2026
detector · early adoption
Announcement·band 2Signal

Humanoid vendor announces retail-store pilot 'shelf restocking' with a US grocer

Pilot announcement with no success rate, no throughput and no cost. Widely covered. Kept as the humanoid claim; nothing in the release supports store economics.

Vendor press release12 Mar 2026
RM?dropped 4
Talk·band 3Signal

'Autonomous haulage is done; the next decade is maintenance'

Pilbara operator's autonomy lead: driving is solved with narrow systems; the open problems are maintenance scheduling, inspection and exception decisions. Foundation models enter as analytics, not as drivers.

AU mining technology forum19 Feb 2026
detector · demand
Repository·band 2Signal

openfleet — fleet exception analytics for mixed robot estates

Open-source exception-logging and analytics layer for heterogeneous robot fleets. Low stars, but the existence of an unbundled layer is evidence against the red team's bundling case.

900stars
github.com28 Jan 2026
detector · early adoption
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Foundation-model manipulation policies sit at 70–85% task success on novel objects; deployment needs above 99% on the tasks currently automated.

Assessedc-robots-1dalton-0.425 Jun 2026Benchmark leaderboard, Foundation lab
80%

Mining autonomy in the Pilbara is mature and narrow; foundation models enter through maintenance and exception decisions, not driving.

Assessedc-robots-4dalton-0.425 Jun 2026AU mining technology forum, Corporate careers page
70%

The economic entry point in AU logistics is the long tail of exceptions, which is a decision-data problem before it is a hardware one.

Assessedc-robots-2dalton-0.425 Jun 2026Engel · retail (owner) engagement, AU supply-chain conference
61%

Sim-to-real from learned world models will shorten the data-collection bottleneck for new tasks by an order of magnitude.

Signalc-robots-5dalton-0.425 Jun 2026arxiv.org, github.com
45%

Humanoid robots will be economically deployed in retail stores within the Next horizon.

Signalc-robots-3dalton-0.35 May 2026Vendor press release, Personal blog
18%

Position history · the diff is the product

2 validation runs against a fixed brief. Confidence 35% → 38%.

runs compare claim sets, never prose
What we said · run 2

Long-tail exceptions named as the entry point. Humanoid claims down-weighted. Red team's bundling argument is unanswered and the review date has now lapsed.

38%
Changed since run 1
  • The economic entry point in AU logistics is the long tail of exceptions, which is a decision-data problem before it is a hardware one.
  • Mining autonomy in the Pilbara is mature and narrow; foundation models enter through maintenance and exception decisions, not driving.
  • Sim-to-real from learned world models will shorten the data-collection bottleneck for new tasks by an order of magnitude.
  • c-robots-3 ↓ 0.26 → 0.18
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · DS
medium

Real for the owner's logistics arm; the firm's slice is the decision layer, which is medium at best.

Timeline

agent-estimated
18mo–4yr

Agent-estimated: long-tail deployment in a DC within four years is plausible on the reliability trajectory. Uncommitted.

TAM

agent-estimated
>$10B

Agent-estimated on global logistics robotics; the addressable decision-layer slice is unknown.

Cost

committed · TO
low

Watching costs nothing; a survey run into a live deployment is a Type 3.

Cost of being wrong

committed · LF
low

The firm is not building hardware; a late start on the software layer costs an engagement, not a business.

Demand

committed · DS
low

One question from the owner's supply-chain side about exception handling; nothing from other retail clients.

Cost of entry

agent-estimated
medium

Domain credibility with DC operators has to be earned; the firm has none in robotics. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Retail & FMCG
relevant

The owner runs AU's largest automated retail logistics estate. The long-tail exception layer is where a foundation-model policy would enter and where the decision data lives.

Mechanism · Exception-rate telemetry from automated DCs modelled as a decision problem; a foundation-model policy as one option in the decision.

DS committed by Dev Sharmacommitted
Energy & Utilities
watch

Mining and utilities run mature narrow autonomy. Foundation models enter through maintenance and inspection decisions, which are analytics problems.

Mechanism · Inspection-robot telemetry as input to asset-health models.

Agent draft · awaiting a sector owneragent-estimated
Health
not-relevant

Surgical and care robotics are a regulatory and hardware domain the firm has no path into on this horizon.

Mechanism · None.

HN committed by Hana Novakcommitted
Defenceprospective
watch

Autonomy is the fastest-moving defence procurement category and the firm has no presence; a sponsor question, not an owner one.

Mechanism · Would require a sponsor to hold the 'should we be here' question.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against: our 'software layer' framing is a consultancy telling itself the hardware transition is someone else's problem. The operators who deploy foundation-model robots will get the decision layer from the robot vendor, not from an analytics firm. Our entry point may not exist.

  • Robot vendors ship the fleet-management and exception-analytics layer bundled; there is no unbundled decision problem for a third party to model.
  • 70–85% on novel objects was 40% two years ago; the reliability trajectory is steeper than our timeline band assumes.
  • The owner's DC automation is run by an integrator under a long contract; the firm's ownership link is not an access link.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • Robotics research35%
  • AU operators / practitioners30%
  • Vendor / analyst15%
  • Internal / Engel20%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

depends onWorld Models

Trigger · A published sim-to-real result where a policy trained in a learned world model reaches ≥90% success on a real manipulation task the lab can specify, on a robot body available in AU; Argus tracks the three VLA labs' release notes for it.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

depends onSensing Agents

Trigger · A DC operator (the owner's or a competitor's) publishes or shares a long-tail exception rate from a live foundation-model deployment; until an exception rate exists there is no decision problem to model.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

enablingWorld Models

Learned simulators are the training environment and the sim-to-real path for new tasks.

compoundingSensing Agents

Robot fleet telemetry is the richest sensing source in a DC; the sensing agent pattern is the entry point.

compoundingVoice and Vision

Natural-language tasking of robots is a voice-and-vision interface problem.

blockingDeciding Table Stakes

The firm has no robotics credibility; the right to play in this room has not been bought.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsDSTOJP

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Does an unbundled decision layer exist in a live DC, or do robot vendors ship it bundled?
  2. 02What is the owner's actual long-tail exception rate, and who owns the data?
  3. 03At what novel-object success rate does a DC operator let a foundation-model policy touch a live pick line?