cavendish
QueueExperimentsScorecard

Kekulé · the dream cycle · 1 Sep 2026

A dream journal, not a dashboard.

Four passes run as one scheduled job: replay and consolidation, remote association, pruning, speculation. The pass that justifies it crosses the graveyard with capability tracking — when a constraint moves, every dead experiment depending on it is re-checked.

5 of 7 decidedacceptance 40%published rolling rate 31%Maximum seven items. The volume cap and the published rejection rate are the only defences.
  1. 1
    Remote associationSignal

    Expiring delegated credentials are a backpressure signal the orchestrator is not using

    auth-broker and backpressure sit close in embedding space and have never shared a claim, a citation or a person. The broker's short-lived, task-scoped credentials already encode a bound on how long and how wide an agent may act; the bounded-fanout queue enforces a similar bound by a separate mechanism, and the two could share one budget. Accepted as a join hypothesis because the shared constraint is structural, not lexical — Tom and Lena to draft a thought experiment.

  2. 2
    Remote associationSignal

    A compiled skill wiki as the retention policy for citizen-developer artefacts

    personal-wiki and citizen-developers-org-slop are near in embedding space — both are about accumulating agent-built artefacts — and far in the graph. The proposal was that the wiki's consolidation step is a retention policy in disguise. Rejected: the proximity is the words 'wiki' and 'artefact' doing the work; the wiki compiles one person's experience under one owner, and the slop problem is precisely that nobody owns anything. This is the spurious-connection shape the volume cap exists for.

  3. 3
    ReplaySignal

    Summarise-before-store beat store-then-retrieve in three unconnected experiments

    x-memory-bench, x-personal-wiki and x-learning-agent-memory each found, separately, that consolidating experience into a summary before storing it outperformed storing raw and retrieving later — on precision, on latency and on decision consistency. The three experiments share no brief and only one person. Proposed as a generalised claim on continuous-learning: 'consolidation at write time beats retrieval-time synthesis for agent experience.' Accepted for promotion; Mei to confirm the claim text and its supporting signal set.

  4. 4
    PruningSignal

    Three claims in cyber-cold-war have had no supporting signal added since April

    The field's supporting claims were sourced from a cluster of band-2 posts in March and nothing has been added since; two of the three have decayed below 0.4 and the third rests on a single analyst note. Pruning proposes demoting the field from validating to contested and re-running the brief. Deferred rather than actioned — the field has a fortnightly cadence but the owner has been on the eval work; the proposal is correct in mechanism, and the risk of a fluent case for demotion that is really just neglect is why a human reads it first.

  5. 5
    SpeculationSignal

    Calibrate sensing-agent anomaly alerts against a human panel the way judges are calibrated

    x-judge-calibration's method — a small human panel, an agreement score, and a refusal to publish without one — has never been tried on x-sensing-agents, whose store-telemetry alerts are currently unscored against any human ground truth. The problem classes look different (alerting versus grading) but the failure mode is the same: a confident signal with no measured agreement. Lands in the proposal queue as a Type 1 run. Undecided; Jun to say whether the telemetry has enough labelled incidents to make a panel.

  6. 6
    SpeculationSignal

    Run the token cost ledger method on realtime voice sessions

    The ledger from x-token-cost-ledger attributes every token to a cache state and a pattern; the proposal was to run it on x-voice-latency's sessions to get a per-minute cost curve for realtime voice. Rejected: realtime speech-to-speech is priced per second of session and per audio minute, not per token, and the cache-state attribution that makes the ledger useful has no equivalent. The analogy is fluent and the transfer is empty. A voice cost model is worth having, but it is a new instrument, not this one.

  7. 7
    Graveyard × capabilitySignal

    Argus: AU-channel inference-card pricing moved; g-onprem-h100-cluster depends on that constraint

    Argus detected a ~40% drop in $/TFLOP for inference-class cards in the AU channel across two distributors in the last three weeks, plus one colocation provider's Sydney rack pricing falling. g-onprem-h100-cluster died on utilisation and on API prices falling during the build; its resurrection trigger names hardware cost as one of two conditions. The constraint has moved, but only one of the two — API prices have kept falling too, so the utilisation threshold in sa-onprem-when may not shift. Resurfaced for Priya; undecided until the threshold is recomputed.

Note

LLM-generated connections are fluent, confident and mostly spurious. Accepted items land in the proposal queue, so the cycle produces work rather than a report.