Library · highest traffic in the system
Three vehicles, split by decay rate.
Not by topic. What a thing is worth depends on how fast it rots. Fast-decaying facts are never a document. Recommendations carry strength and evidence tier as separate fields. Positions are long-form, argued and signed.
- Recommendations
- 12
- Standing answers
- 8
- Positions
- 4
- DorecommendationTested
Prompt caching: use for stable prefixes over 2k tokens; expect 30–45%, not 60%
Cache when the shared prefix is over ~2k tokens and reused inside the provider's TTL. Budget on 30–45% input-cost reduction across a real pattern; the 60% figure is a best-case single call.
StrengthmoderateCitations61Half-life90dvalidated 28 Aug 202684d leftPRPriya RamanCost redux on tokens - Answerstanding answerTested
Which model for structured extraction?
Qwen 3.5 32B through the gateway's extraction tier for flat and moderately nested schemas; Claude Sonnet 5 for schemas over ~40 fields or with cross-field constraints. As of 27 August.
StrengthstrongCitations54Asked210×validated 27 Aug 202621d leftMTMei TanakaOpen Weight Models - PositionpositionAssessed
Table stakes, not moats: what a right to play costs in 2026
Most of what consultancies sold as AI differentiation in 2024 is now table stakes: a gateway, an eval harness, a calibrated judge, a memory default, delegated auth. The moat is not in having them. It is in being measurably right about which ones to use, faster than the client's own team.
StrengthstrongCitations51Half-life365dvalidated 25 Aug 2026356d leftAWAdam WitanowskiDeciding Table Stakes - Answerstanding answerTested
What does inference actually cost right now?
Blended across our six instrumented patterns, 1 September: $0.9–1.6 per thousand agent turns on the open-weight tier, $4–11 on frontier mid-tier. Auto-refreshed from the ledger; not yet human re-committed.
StrengthmoderateCitations48Asked184×auto-refreshed · not re-committed19d leftPRPriya RamanCost redux on tokens - DorecommendationTested
Open-weight models for classification and extraction; frontier for agentic loops
Default classification and structured extraction to an open-weight 30B-class model behind the gateway. Keep multi-step agentic loops on frontier until the open-weight tool-use gap closes.
StrengthmoderateCitations47Half-life90dvalidated 27 Aug 202683d leftPRPriya RamanOpen Weight Models - PositionpositionAssessed
The direction of agentic development
Agentic development is real, the throughput gain is real, and the bottleneck has moved from writing code to specifying and verifying it. The firms that win will be the ones that industrialise the spec, not the ones that let the agent merge.
StrengthstrongCitations44Half-life365dvalidated 18 Aug 2026349d leftSKSam KowalczykAI-SDLC - DorecommendationTested
Use a structured episodic store with summarised recall, not a raw vector memory
For any agent with more than ten sessions of history, store episodes with summaries and recall the summary first. Raw vector recall over transcripts loses on precision and latency above ~50k tokens.
StrengthstrongCitations41Half-life120dvalidated 21 Aug 2026107d leftTOTom OkaforAgentic Memory System - Don'trecommendationTested
Measure ROI from telemetry and cycle time, not surveys
Do not report AI productivity gains from self-reported surveys. Instrument the workflow and measure cycle time, throughput and rework before and after.
StrengthmoderateCitations38Half-life120dvalidated 19 Aug 2026105d leftJPJun ParkROI - PositionpositionAssessed
Sovereign inference and the end of US default
US frontier models remain the capability ceiling, but the default that every serious workload runs on a US model through a US cloud is ending for Australian regulated sectors. The lab's position is to design for a model garden with an AU-resident open-weight tier from the start, and to treat sovereignty as a routing decision rather than a hardware one.
StrengthmoderateCitations37Half-life365dvalidated 10 Jul 2026310d leftPRPriya RamanUS Non-Dominance - Answerstanding answerTested
Retrieval or fine-tuning for this?
Retrieval, almost always. Fine-tune only for format or tone at scale, or for offline edge hardware, and only after the stock model is measured on the same eval.
StrengthstrongCitations36Asked132×validated 12 Aug 202623d leftPRPriya RamanSLM / Edge / Tuning - DorecommendationTested
Route through a gateway you control; do not standardise on a vendor's
Every client pattern goes through a gateway the delivery team owns the config for. Vendor gateways are fine as a backend, never as the routing policy.
StrengthstrongCitations33Half-life150dvalidated 14 Aug 2026130d leftPRPriya RamanAI Gateway - Answerstanding answerTested
Which memory layer should a new agent use?
Under ten sessions of history: none, just the context window. Over ten: the lab's episodic store with summarised recall via the memory adapter. Not a vendor memory product yet.
StrengthstrongCitations31Asked96×validated 21 Aug 202632d leftTOTom OkaforAgentic Memory System - Don'trecommendationTested
Every LLM judge ships with a human agreement score or does not ship
Do not report an eval number produced by an LLM judge unless the judge has a published agreement rate against a human panel on the same task. Below 0.8 Cohen's kappa the judge is not a measurement.
StrengthstrongCitations29Half-life120dvalidated 10 Aug 202696d leftMTMei TanakaEval Harnesses - Answerstanding answerAssessed
What eval tooling do we use?
The lab's canary suite for release gating, a calibrated LLM judge with a published kappa for scoring, and a per-engagement eval set in the client's repo. No commercial eval platform as of August.
StrengthmoderateCitations27Asked77×validated 5 Aug 202616d leftMTMei TanakaEval Harnesses - Don'trecommendationTested
Agents act under delegated, scoped, expiring authority — never a service account
Do not give an agent a long-lived service account. Every tool call runs under authority delegated from a named person, scoped to the task, and expiring with it.
StrengthstrongCitations26Half-life180dvalidated 7 Aug 2026153d leftLFLena FischerAuth Broker - PositionpositionAssessed
Learning without weights: where continual learning actually lands
Continual learning in the weights is a 4yr+ research problem. Continual learning outside the weights — retrieval-updated skills, episodic memory, compiled experience — is here, works, and is where the firm's investment should go. The distinction is the position.
StrengthmoderateCitations23Half-life365dvalidated 30 Jun 2026300d leftMTMei TanakaNon-Weight-Bound Continuous Learning - DorecommendationTested
Agentic delivery works when the spec is the artifact; do not start with the code
The reviewable, versioned artefact in an agentic delivery is the specification and its acceptance tests. Code is generated output. Teams that review code first lose the throughput gain to review time.
StrengthmoderateCitations22Half-life120dvalidated 18 Aug 2026104d leftSKSam KowalczykAI-SDLC - Answerstanding answerAssessed
When does on-prem inference make sense?
When the data cannot leave a boundary that no AU-region API sits inside, or when steady-state utilisation is above ~60% of a rack for a year. Not for cost at typical enterprise volumes; the 18-month payback did not hold.
StrengthmoderateCitations22Asked41×validated 29 Jul 20269d leftPRPriya RamanOn-Prem Inference - Answerstanding answerAssessed
Is the vendor's 'agentic' claim real?
Usually not in the sense the deck implies. Ask four questions: does it plan across more than one tool call, can it recover from a failed call, what happens on ambiguity, and where is the eval. As of 14 August, most vendors fail two of four.
StrengthmoderateCitations19Asked58×validated 14 Aug 202610d leftLFLena FischerDeciding Table Stakes - Not yetrecommendationTested
Not yet: realtime voice for AU contact centres above tier-1 triage
Do not commit a client to realtime voice agents beyond tier-1 triage and routing. The latency floor from Sydney is not there, and the cost of a public failure in a regulated contact centre is high.
StrengthstrongCitations18Half-life90dvalidated 31 Jul 202656d leftTOTom OkaforVoice and Vision - DorecommendationTested
Put backpressure on agent fan-out before you put it on the model
Bound the number of in-flight sub-agents and tool calls at the orchestrator with a queue that applies backpressure. Model-side rate limits are the wrong place to discover you have a fan-out problem.
StrengthstrongCitations15Half-life150dvalidated 25 Aug 2026141d leftTOTom OkaforBackpressure - Answerstanding answerAssessed
When do we need to move to post-quantum crypto?
Inventory now, migrate long-lived secrets by 2028, everything else on the ASD timeline (2030). Do not sell a 'quantum readiness audit' as a product; it is a crypto inventory and clients can do it themselves.
StrengthmoderateCitations11Asked14×validated 1 Aug 202612d leftLFLena FischerQuantum Encryption - DorecommendationAssessed
Give citizen developers a paved road and a retention policy, not a review board
Review boards do not stop org slop; they slow the people who would have done it well. A paved road with a default retention window keeps the volume survivable.
StrengthweakCitations9Half-life120dvalidated 16 Apr 202620d overdueSKSam KowalczykCitizen Developers and Org Slop - DorecommendationTested
Distil to a small model only after the frontier baseline is measured on the same eval
Superseded by r-open-weight-tier. Before fine-tuning or distilling any small model, run the frontier model on the same eval and record the gap you are trying to close.
StrengthmoderateCitations12Half-life120dsuperseded16d leftPRPriya RamanSLM / Edge / Tuning