Share graph · provenance running forward
Mei Tanaka
Research lead · evals · lab
What they dropped, which clusters it landed in, what got elected, and which experiments and findings descended from it. Everyone’s is visible to everyone, symmetrically. Show outcomes, not counts — no totals, no rankings, no rollups.
Dropped
signals with this person attached
Landed in
distinct clusters
Elected
of those, now fields
Descended
experiments and published items
Drops
What descended
Use a structured episodic store with summarised recall, not a raw vector memory
Route through a gateway you control; do not standardise on a vendor's
Prompt caching: use for stable prefixes over 2k tokens; expect 30–45%, not 60%
Every LLM judge ships with a human agreement score or does not ship
Open-weight models for classification and extraction; frontier for agentic loops
Distil to a small model only after the frontier baseline is measured on the same eval
Which model for structured extraction?
What does inference actually cost right now?
Retrieval or fine-tuning for this?
Which memory layer should a new agent use?
What eval tooling do we use?
Sovereign inference and the end of US default
Learning without weights: where continual learning actually lands
Memory layer bake-off on a 40-session support corpus
Cost-aware routing across three model gardens
Token cost ledger across six client patterns
Judge calibration against human panel
Open-weight parity on our task evals
Release canary suite v2
Edge SLM for in-store classification
Agentic QA on a regression-heavy codebase
Compiling agent experience into a persistent skill wiki
Temporal decision memory for a claims agent
Non-weight-bound learning via retrieval-updated skills
Follow this person’s finds
Following someone whose drops are consistently good is the internal version of the external voice watchlist, and often a better source than any detector.