cavendish
QueueExperimentsScorecard
Type 3Signal

Sensing agent over store telemetry

Sensing Agents
  1. proposed · 13d
  2. voting
  3. running
  4. measuring
  5. concluded

Preregistration · v1 · 21 Aug 2026 · immutable after start

An extension creates a new version rather than editing the old.

Hypothesis

An agent that reads a store's shelf-event and footfall telemetry every fifteen minutes and raises an alert only when it can name a cause detects out-of-stock and planogram-drift events at least four hours earlier than the current end-of-day report, with a false-positive rate under 20%.

Kill condition

Kill if lead time over the end-of-day report is under one hour on the replay set by day 10, or if false-positive rate exceeds 40% at any threshold that keeps recall above 0.7.

MethodReplay of six weeks of telemetry from two stores (de-identified, from the retail practice) with the labelled event log as ground truth. Agent built on x-slm-edge-classifier's shelf classifier plus a frontier reasoner for cause-naming. Measure lead time, precision, recall, cost per store-day.
Expected cost$1,300 in tokens, two people for three weeks
Expected duration3 weeks

Running notes · fed from harness sessions and by the pair

AW
manual · Adam Witanowski · 3 Sep 2026· ⌘↩ to post
  1. manual29 Aug 2026TO Tom Okafor

    Dev (retail) has confirmed the six-week replay set can be released de-identified. Not client data under the governance doc; stays Type 3.

  2. manual21 Aug 2026PR Priya Raman

    Depends on the edge classifier concluding; proposed now so it can enter the next voting round if x-slm-edge-classifier clears. Predicted lead time 5h, FP 25%, confidence 0.45.

Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.