cavendish
QueueExperimentsScorecard
Type 3Signal

Temporal decision memory for a claims agent

Learning Agents
  1. proposed
  2. voting · 10d
  3. running
  4. measuring
  5. concluded

Preregistration · v1 · 24 Aug 2026 · immutable after start

An extension creates a new version rather than editing the old.

Hypothesis

A claims-triage agent that records each decision with its reasoning and timestamp in the episodic store from r-memory-layer, and retrieves prior decisions on similar claims, improves triage consistency on a 500-claim synthetic corpus by at least 12 points without a drop in accuracy.

Kill condition

Kill if consistency improves under 5 points by day 12 on the held-out split, or if accuracy drops more than 2 points at any point.

Method500 synthetic claims with injected near-duplicates. Two agents: with and without temporal decision memory. Measure consistency (same decision on near-duplicates), accuracy against the labelled outcome, and latency.
Expected cost$1,000 in tokens, two people for three weeks
Expected duration3 weeks

Running notes · fed from harness sessions and by the pair

AW
manual · Adam Witanowski · 3 Sep 2026· ⌘↩ to post
  1. manual28 Aug 2026JP Jun Park

    Consistency metric defined: agreement on near-duplicate pairs, 180 pairs in the held-out split.

  2. manual24 Aug 2026TO Tom Okafor

    First dependent run on r-memory-layer. Depends on the memory store; does not need the forgetting mechanism because the corpus is synthetic. Predicted +15, confidence 0.55.

Harness notes are auto-captured from Claude Code sessions: model, date, commit, session reference. Never the transcript, code or paths.