Edge SLM for in-store classification
SLM / Edge / Tuning- proposed
- voting
- running
- measuring · 29d
- concluded
Preregistration · v2 · 29 Jun 2026 · immutable after start
Hypothesis
A 3B model distilled from a measured frontier baseline reaches within 3 points macro-F1 of that baseline on a 22-class shelf-event classification task, running under 150ms on a retail edge box with no network.
Kill condition
Kill if the gap to the frontier baseline is over 8 points after distillation on the held-out split by day 12, or if p95 latency on the target box exceeds 400ms.
Running notes · fed from harness sessions and by the pair
- harness1 Sep 2026PR Priya Raman
older box: p95 212ms; still under kill line. waiting on practice sign-off to conclude
- manual19 Aug 2026MT Mei Tanaka
Measuring is overrunning. The retail practice wants a second box with the older CPU before they will read the latency number; box arrives next week.
- harness5 Aug 2026PR Priya Raman
3B quantised: 0.87 F1 (gap 4), p95 138ms on the box. Into measuring.
- harness16 Jul 2026PR Priya Raman
frontier baseline 0.91 macro-F1 with class defs + 12 examples; distillation set 46k built
- manual29 Jun 2026PR Priya Raman
Prereg v2 after vote: added the frontier-baseline-first step. Predicted gap 4 points, latency passes, confidence 0.5.
Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.