cavendish

Share graph · provenance running forward

MT

Mei Tanaka

Research lead · evals · lab

What they dropped, which clusters it landed in, what got elected, and which experiments and findings descended from it. Everyone’s is visible to everyone, symmetrically. Show outcomes, not counts — no totals, no rankings, no rollups.

Dropped

23

signals with this person attached

Landed in

16

distinct clusters

Elected

11

of those, now fields

Descended

24

experiments and published items

Drops

What descended

recommendationTested

Use a structured episodic store with summarised recall, not a raw vector memory

recommendationTested

Route through a gateway you control; do not standardise on a vendor's

recommendationTested

Prompt caching: use for stable prefixes over 2k tokens; expect 30–45%, not 60%

recommendationTested

Every LLM judge ships with a human agreement score or does not ship

recommendationTested

Open-weight models for classification and extraction; frontier for agentic loops

recommendationTested

Distil to a small model only after the frontier baseline is measured on the same eval

standing answerTested

Which model for structured extraction?

standing answerTested

What does inference actually cost right now?

standing answerTested

Retrieval or fine-tuning for this?

standing answerTested

Which memory layer should a new agent use?

standing answerAssessed

What eval tooling do we use?

positionAssessed

Sovereign inference and the end of US default

positionAssessed

Learning without weights: where continual learning actually lands

experiment · concludedValidated

Memory layer bake-off on a 40-session support corpus

experiment · concludedValidated

Cost-aware routing across three model gardens

experiment · concludedValidated

Token cost ledger across six client patterns

experiment · concludedRefuted

Judge calibration against human panel

experiment · concludedSuperseded

Open-weight parity on our task evals

experiment · measuring

Release canary suite v2

experiment · measuring

Edge SLM for in-store classification

experiment · measuring

Agentic QA on a regression-heavy codebase

experiment · running

Compiling agent experience into a persistent skill wiki

experiment · voting

Temporal decision memory for a claims agent

experiment · proposed

Non-weight-bound learning via retrieval-updated skills

Follow this person’s finds

Following someone whose drops are consistently good is the internal version of the external voice watchlist, and often a better source than any detector.