Library · highest traffic in the system
Three vehicles, split by decay rate.
Not by topic. What a thing is worth depends on how fast it rots. Fast-decaying facts are never a document. Recommendations carry strength and evidence tier as separate fields. Positions are long-form, argued and signed.
- Recommendations
- 12
- Standing answers
- 8
- Positions
- 4
fast-decaying · refreshed on schedule, human re-committed
- Answerstanding answerTested
Which model for structured extraction?
Qwen 3.5 32B through the gateway's extraction tier for flat and moderately nested schemas; Claude Sonnet 5 for schemas over ~40 fields or with cross-field constraints. As of 27 August.
StrengthstrongCitations54Asked210×validated 27 Aug 202621d leftMTMei TanakaOpen Weight Models - Answerstanding answerTested
What does inference actually cost right now?
Blended across our six instrumented patterns, 1 September: $0.9–1.6 per thousand agent turns on the open-weight tier, $4–11 on frontier mid-tier. Auto-refreshed from the ledger; not yet human re-committed.
StrengthmoderateCitations48Asked184×auto-refreshed · not re-committed19d leftPRPriya RamanCost redux on tokens - Answerstanding answerTested
Retrieval or fine-tuning for this?
Retrieval, almost always. Fine-tune only for format or tone at scale, or for offline edge hardware, and only after the stock model is measured on the same eval.
StrengthstrongCitations36Asked132×validated 12 Aug 202623d leftPRPriya RamanSLM / Edge / Tuning - Answerstanding answerTested
Which memory layer should a new agent use?
Under ten sessions of history: none, just the context window. Over ten: the lab's episodic store with summarised recall via the memory adapter. Not a vendor memory product yet.
StrengthstrongCitations31Asked96×validated 21 Aug 202632d leftTOTom OkaforAgentic Memory System - Answerstanding answerAssessed
What eval tooling do we use?
The lab's canary suite for release gating, a calibrated LLM judge with a published kappa for scoring, and a per-engagement eval set in the client's repo. No commercial eval platform as of August.
StrengthmoderateCitations27Asked77×validated 5 Aug 202616d leftMTMei TanakaEval Harnesses - Answerstanding answerAssessed
When does on-prem inference make sense?
When the data cannot leave a boundary that no AU-region API sits inside, or when steady-state utilisation is above ~60% of a rack for a year. Not for cost at typical enterprise volumes; the 18-month payback did not hold.
StrengthmoderateCitations22Asked41×validated 29 Jul 20269d leftPRPriya RamanOn-Prem Inference - Answerstanding answerAssessed
Is the vendor's 'agentic' claim real?
Usually not in the sense the deck implies. Ask four questions: does it plan across more than one tool call, can it recover from a failed call, what happens on ambiguity, and where is the eval. As of 14 August, most vendors fail two of four.
StrengthmoderateCitations19Asked58×validated 14 Aug 202610d leftLFLena FischerDeciding Table Stakes - Answerstanding answerAssessed
When do we need to move to post-quantum crypto?
Inventory now, migrate long-lived secrets by 2028, everything else on the ASD timeline (2030). Do not sell a 'quantum readiness audit' as a product; it is a crypto inventory and clients can do it themselves.
StrengthmoderateCitations11Asked14×validated 1 Aug 202612d leftLFLena FischerQuantum Encryption