Crystal ball · queryable
Not “we think X”. Here is what we said in March, what changed, and why.
Validation runs are versioned, never overwritten. The brief is fixed; only the world varies. A run emits claims, the report is rendered from the claim set, and the diff is a set operation — stable, attributable, defensible in front of a board.
- Validation runs
- 124
- Fields with a position
- 40
- What is our current position on X? · default
- Mid-horizon, banking, above medium confidence · faceted
- What changed since June? · the diff
- What did we say twelve months ago? · time travel
Fields by horizon and confidence
What changed
The diff — the thing nobody else can produce.
- Tested
Both experiments concluded. Canary caught a real regression before any client did; judge calibration gives 0.71 on rubric tasks and 0.38 on open-ended. Unchecked judge to the graveyard; recommendation and standing answer current. Field converged; skills is the gate.
- LLM-judge agreement with humans is task-dependent: around 0.7 kappa on…
- Release canaries catch material regressions on structured tasks within…
- Uncalibrated judges systematically favour newer and same-family models…
- Eval engineering is becoming a job title; the skills gate is closing o…
- c-eval-harnesses-5 ↓ 0.30 → 0.18
- Tried
Experiment interim: repeat-task time down 44%; 6 of 52 compiled pages wrong and acted on. The page-error rate is now the primary open question and the gate is tooling for the review loop.
- The compilation step produces confidently wrong pages at a material ra…
- c-personal-wiki-1 ↑ 0.6 → 0.7
- Tested
Run concluded. Survey overstates telemetry 2–3×; net delivery gain 10–18%. Survey method sent to the graveyard; recommendation published. Gate moved from tooling to adoption.
- Self-reported productivity gains from coding assistants overstate tele…
- Coding assistants cut PR cycle time but raise review load; the net del…
- c-ai-roi-3 ↑ 0.55 → 0.68
- c-ai-roi-5 ↓ 0.35 → 0.24
- Signal
Lab's own analytics production is the reference case, running dark since June. Headcount claims overstated by uncounted control roles. Red team's marginal-value argument is the open question.
- The lab's own analytics production can run unattended for a quarter wi…
- Task-level automation rates overstate the headcount effect because the…
- c-dark-factory-3 ↓ 0.22 → 0.15
- Tested
Pilot concluded. Effort saving is real, escape rate is flat, coverage is a vanity metric. The role is re-shaped around the risk model, not removed. Full-replacement thesis sent to the graveyard.
- Agent-run test maintenance and flaky-triage cut QA effort by more than…
- Line coverage produced by agents is uncorrelated with escape rate; cov…
- Agent-designed suites systematically under-test business-risk paths be…
- Agent-written and human-written suites catch different defect classes;…
- c-fully-agentic-qa-4 ↓ 0.38 → 0.24
- Assessed
Parity confirmed on extraction and classification; frontier gap confirmed on agentic tasks. Provenance is the real objection. Position drafted for the sponsor's decision on whether to publish.
- Chinese open-weight models are at parity with US open weights on extra…
- The frontier capability gap on agentic and long-horizon tasks has not …
- Provenance risk in open weights — training data, backdoors, licence te…
- c-us-non-dominance-4 ↓ 0.62 → 0.55
- Tested
Routing experiment concluded: static rules capture the saving; learned routers do not generalise; agentic loops do not benefit. Field converged. Recommendation published; hand to practice in Q4.
- Static task-class routing rules capture most of the achievable cost sa…
- Cost attribution per use case is the feature clients buy the gateway f…
- Adapter drift on streaming and tool-call formats is the recurring oper…
- Routing gains vanish on agentic loops because the loop needs one model…
- c-ai-gateway-3 ↑ 0.7 → 0.77
- Signal
Canary coverage up to 58% of families at expert level, steady slope. Three failure modes unchanged in kind. Red team argues the tracker is capped by design; unanswered. Still Distant.
- Frontier models clear more of the lab's professional task families eac…
- Three failure modes persist across every release: no learning from the…
- Transfer to an unpublished professional task family is the most discri…
- c-agi-3 ↑ 0.15 → 0.2
- Tried
Calibration claim strengthened by a paper and a second engagement. Proactivity has a ceiling; above it the assistant gets muted. Standing answer drafted; experiment proposed.
- Visible uncertainty and honest refusal increase use rather than reduce…
- Proactivity above a low threshold becomes interruption and drives the …
- c-effective-assistants-3 ↑ 0.58 → 0.66
- Assessed
Position drafted: parity on capability, differentiation on measurement and the graveyard. Relationship counter-case rated higher than expected; verdict in doubt until Engel can show measured pitches win. Position and standing answer published with the doubt stated.
- The only capability a client cannot obtain from any competitor is a me…
- Capability is irrelevant to consulting moats; relationships and distri…
- c-table-stakes-3 ↑ 0.52 → 0.60
- Assessed
Interim tax measured at about 12% at production batch shape; attestation verified client-side. Overhead-negligible claim retired. Regulator briefing drafted.
- GPU enclave inference for a 70B model carries an 8–35% performance tax…
- Remote attestation can be verified by the client independently of the …
- c-tee-secure-inference-4 ↓ 0.35 → 0.15
- Tried
Drift signal prototype caught half the failures. Acceptance cost did not fall until the fleet argued for its own PRs. Labour-saving claim downgraded; Type 3 proposed; client repos ruled out in writing.
- The dominant failure mode of dark runs is plausible code that satisfie…
- A second-agent spec audit before PR catches most drift at a cost that …
- c-dark-harness-3 ↓ 0.4 → 0.23
- Tested
Bench concluded: rate limits fail first, retries amplify 3.2×, admission control bounds spend at 1.1×. Recommendation published; gate is tooling because no harness ships the control.
- Under load spikes, unbounded agent fan-out fails on provider rate limi…
- Admission control at the orchestrator with a per-task cost cap bounds …
- c-backpressure-4 ↑ 0.6 → 0.7
- Tested
Ledger concluded. List-price falls dominate the levers over twelve months; distillation and pruning are the next factor with a maintenance bill. Recommendation and standing answer published; field converged.
- List-price falls of roughly 4× a year dominate every engineering lever…
- Task-specific distillation gives the next cost factor on narrow tasks …
- c-token-cost-reduction-1 ↑ 0.78 → 0.86
- Tested
Bench confirms the default. Vendors converging on three-store design; consolidation policy is the differentiator and remains hand-rolled. Recommendation published.
- Vendors are converging on a working/episodic/semantic three-store desi…
- c-agentic-memory-1 ↑ 0.74 → 0.84
- Assessed
Hosted AU-region prices down about 40% year on year; crossover moved further from typical utilisation. Position held on a third client measurement. Standing answer refreshed.
- The cost crossover between owned H100-class hardware and AU-region hos…
- Enterprise inference clusters run at 20–40% sustained utilisation; the…
- Hosted AU-region inference prices fell about 40% in twelve months, mov…
- c-on-prem-inference-4 ↓ 0.28 → 0.20
- Signal
Logical-qubit counts up, error rates down, both on schedule. The client-facing date is regulatory, not physical — a regulator will set a PQC deadline first. Overhead reductions are the risk to our estimate.
- Quantum chemistry is the first application to cross classical parity, …
- AU regulators will set a PQC deadline before the compute threshold is …
- c-quantum-compute-2 ↓ 0.72 → 0.66
- Assessed
Second validation run plus a Type 2 on the lab's own infrastructure. Inventory dominates the migration; agent-assisted scanning makes it a days-not-months exercise. Entry window closes around 2029 as big-four practices consolidate. Standing answer published.
- The inventory phase dominates PQC migration effort and timeline; the a…
- Agent-assisted scanning of code and configuration repositories can pro…
- The big-four consultancies' PQC practices will have consolidated the A…
- c-quantum-encryption-1 ↑ 0.78 → 0.86
- Assessed
Late, noisy feedback is the open problem and the experiment is aimed at it. Structured policy stores do not beat text rules yet. Field stays emerging until x-learning-agent-memory concludes.
- The gain flattens within a few hundred decisions; there is no evidence…
- Late, noisy outcome feedback destroys the gain; nobody has demonstrate…
- c-learning-agents-4 ↓ 0.45 → 0.30
- Tested
Latency bench concluded. US-hosted floor is unusable; regional endpoint halves it; barge-in is the abandonment driver. Recommendation published; IVR-replacement thesis to the graveyard. Field stays contested because the floor is moving.
- From Sydney, p95 turn latency on US-hosted realtime models exceeds the…
- Barge-in failure, not recognition accuracy, is the dominant cause of c…
- c-voice-and-vision-3 ↑ 0.62 → 0.70
- Tested
Experiment concluded: 38% lead-time cut under spec-first, review now half of lead time. Harness vendors shipping persisted plans. Recommendation and position published; skills is the gate.
- Spec-first agentic delivery cuts change lead time by a third or more; …
- Harness vendors are moving the plan or spec to a persisted first-class…
- c-ai-sdlc-2 ↑ 0.6 → 0.76
- Assessed
Skill library shows an 11-point gain at six weeks; contradictions are the named risk. Weight-level continual learning stays Distant. Position drafted for publication.
- A retrieval-updated skill library improves a deployed agent's task acc…
- Test-time training on the current task context improves long-document …
- Skill libraries accumulate contradictions over time; without a consoli…
- Nightly fine-tuning as a substitute for continual learning regresses h…
- c-continuous-learning-3 ↓ 0.36 → 0.28
- Tried
Humans edit the pages in practice, which is the property that distinguishes this from opaque summaries. Vendor skills primitives announced; compile-and-review identified as the differentiator.
- Human-readable wiki pages are edited by humans in practice; the readab…
- Foundation labs are converging on a first-party 'skills' primitive, wh…
- Tried
Risk function accepted the agent-drafted model document. Claims handling identified as the highest-value target and the least specified. Type 3 on internal analytics delivery proposed; regulated lifecycles held.
- Model-risk documentation can be drafted by an agent from the model rep…
- Claims handling is the highest-value non-software lifecycle for the fi…
- c-ai-dlcs-4 ↓ 0.35 → 0.2
- Tried
Second independent placement on the board. Draft-only retention confirmed in a second pod. The lab still disagrees on whether the field is distinct from Effective Assistants; decision due Q4.
- Draft-only ambient behaviour retains users where act-unprompted does n…
- c-ambient-agents-1 ↑ 0.66 → 0.74
- Assessed
SOCI guidance makes AI model supply a dependency; competitors are staffing. The sceptic's case (flat incident data, defence automating equally) logged and rated above what we expected. Field stays assessed.
- SOCI-regulated entities will be required to treat AI model supply as a…
- Incident statistics show no AI effect; the threat is priced-in vendor …
- c-cyber-cold-war-1 ↓ 0.76 → 0.70
- Tested
PoC concluded: the pattern works at negligible latency and the cost is policy authoring. Two vendor previews shipped mid-run. Recommendation published; policy tooling is the next gate.
- A policy engine at the tool-call boundary adds under 15 ms p50; the co…
- Identity vendors will ship agent-identity primitives before a standard…
- c-auth-broker-1 ↑ 0.7 → 0.82
- Signal
Vendor 'team' primitives change the urgency, not the evidence. Contamination by a wrong agent is now the named reliability gate. Trigger unchanged.
- Harness vendors shipping 'team' primitives means the pattern will appe…
- One confidently wrong agent contaminates the collective; no published …
- c-self-organising-agents-3 ↓ 0.31 → 0.24
- Signal
Turn-taking failures are frequent enough to need a policy. Shared memory, not presence, is where the value is. Coordination-cost evidence cuts against the thesis. Experiment proposed; election pending a second pod.
- When more than one human can address the same agent, contradictory-ins…
- A shared memory store is what makes a multiplayer surface useful; pres…
- c-multiplayer-ai-surfaces-4 ↓ 0.40 → 0.33
- Tested
Edge classifier concluded at 94% of frontier F1. Distillation is the tuning path; retrieval wins knowledge tasks; the cost case erodes with frontier pricing but residency and latency do not. Recommendation and standing answer published.
- Distillation from frontier-generated labels reaches more than 90% of f…
- Retrieval plus a frontier model beats fine-tuning for knowledge-bearin…
- Frontier price declines erode the small-model cost case on roughly a t…
- Assessed
Ninety-day measurement: 412 opens, 38 distinct readers, four people at 61%. Opens rejected as the KPI figure; distinct readers and citations-in-pitch proposed instead. Citation capture depends on Engel.
- Library opens are concentrated: four people account for 61% of reads, …
- Opens are not consumption; citation in a pitch is the measure that ref…
- c-lab-outcomes-visibility-4 ↓ 0.3 → 0.2
- Tried
Early experiment result supports the two-stage design at roughly 4% of naive cost with actionable events the rules engine missed. Additive to rules, not a replacement. Explanation of sources checked is what keeps operators reading.
- A cheap-gate two-stage design cuts the cost of watching a stream by 20…
- Sensing agents find a class of event that threshold rules miss — multi…
- Operators will accept sensed events alongside rules-engine alerts only…
- c-sensing-agents-4 ↓ 0.38 → 0.24
- Signal
Own-delivery substitution measured at 20–30% of junior hours. Outcome pricing appearing at competitors on the substituted share only. Jevons counter-case unanswered; horizon still contested.
- Frontier-class inference price at constant capability has fallen rough…
- In the lab's own delivery, tokens now substitute for 20–30% of junior …
- Consultancies that move to outcome pricing do so on the parts of their…
- c-post-economy-3 ↓ 0.38 → 0.3
- Tested
Parity experiment concluded: within 2 points on three task classes, 14–22 behind on agentic, gap re-opens each frontier release. Origin-as-blocker carried at moderate confidence. Recommendation and standing answer published.
- Open-weight models are within 2 points of frontier on classification, …
- The open-weight gap on multi-step agentic tasks is 14–22 points and re…
- Model origin (a Chinese lab) is a procurement blocker in the AU public…
- c-open-weight-models-3 ↑ 0.62 → 0.72
- Assessed
Offensive capability is rising faster than SOC adoption; export controls arrive as tenant obligations rather than denial. The near-term client question is regulatory, not technical.
- Model capability on offensive-security evals is rising faster than def…
- Export controls will reach Australian clients as cloud-tenant KYC and …
- Tried
Four properties named. Proactivity and integration depth are supported by telemetry; calibration and refusal by one tried result and thin literature. Gate set to adoption: the properties are buildable and not asked for.
- Assistants with a proactive surface retain roughly three times the wee…
- Organisations do not procure for the four properties because the procu…
- c-effective-assistants-4 ↓ 0.30 → 0.20
- Assessed
Guardrail thesis to the graveyard. Retention pattern worked in our tenant; the board case-study shows the value cost. Recommendation published at assessed tier with the second-tenant caveat.
- Template guardrails reduce obvious failures but not sprawl, because sp…
- Platform-enforced retention (owner, last-run TTL, auto-archive) remove…
- Review boards reduce agent count and reduce value roughly in proportio…
- c-citizen-developers-org-slop-5 ↓ 0.42 → 0.30
- Assessed
First lead-time measurement from the graph: median +4.2 months, three negatives. Lead time confirmed as the number that cannot be reconstructed later. Field stays open until December, then dissolves or does not.
- Lead time is the only proof of worth a lab cannot reconstruct retrospe…
- The lab's current median lead time is +4.2 months across sixteen now-f…
- c-lab-effectivity-2 ↑ 0.58 → 0.64
- Tried
Two tried results logged: analytics delivery and model-risk documentation. Both had a versioned artifact beforehand. The scaffold, not the agent, looks like the precondition.
- The spec-first agentic pattern transfers to non-software lifecycles on…
- Analytics delivery cycle time falls by more than half under an agentic…
- Tested
Structured episodic + summarised recall is the working default. Graph memory is a special case, not a default. Long-context is not a substitute.
- Summarised episodic recall beats raw chunk retrieval on decision-consi…
- Graph-structured memory improves entity-heavy tasks and degrades gener…
- c-agentic-memory-4 ↓ 0.41 → 0.22
- Assessed
Competitor 'proprietary' claims map to open tooling; evals are a job description in six months. The moat candidate is measured outcomes on client data. Red team's relationship case logged.
- Evals are claimed by half the competitor set and demonstrated publicly…
- 'Proprietary orchestration' claims in competitor pitches map to open-s…
- Tried
Migration backlogs identified as the production-ready class. Two engineers ran an overnight fleet on an internal repo; the split between precise and imprecise tickets is stark. Field marked contested — the lab disagrees on whether the gap is spec or signal.
- Unattended agent fleets clear well-specified, testable tickets at high…
- Migration and dependency-upgrade backlogs are the one production-ready…
- Tried
Reversibility predicts survival better than accuracy. Draft-only is deliverable; act-unprompted is not. Delegated authority is a hard requirement in regulated sectors.
- Reversibility of the action, not accuracy of the decision, predicts wh…
- Ambient agents acting under an integration token rather than delegated…
- c-ambient-agents-4 ↓ 0.30 → 0.18
- Signal
Internal upskilling cohort result logged. The firm's own capability building is the tractable application; the external market is not. Remain a candidate; run the next cohort with a control.
- The tutoring effect transfers to open-ended, judgement-heavy skills at…
- The firm's own upskilling is the tractable application; the external e…
- Assessed
Field experiment at a consultancy gives a push-over-pull multiplier with decay. Staff-privacy question raised by the product engineer who built the telemetry; aggregation threshold adopted before the KPI reports.
- Push digests roughly triple reach against pull for about eight weeks, …
- Per-person read telemetry on internal documents is a staff-privacy que…
- Tested
Vendor claims of unattended QA do not survive contact with a real repo. Maintenance and triage are the tractable half; design and sign-off are not. Pilot preregistered with escape rate as primary.
- A fully unattended QA function — design through sign-off — is deployab…
- The QA role is being re-titled, not removed: openings for 'test engine…
- Assessed
Regulator survey: no AU acceptance of attestation as a control. Field's gate confirmed as economics plus adoption; economics first because the number is ours to produce.
- No AU regulator has accepted attestation as a control for PII; until o…
- Tested
Providers moving to token-priority tiers; the 'buy throughput' counter-case logged and rated low. Bench running; interim results consistent with hypothesis.
- Provider rate limits are moving to token-based priority tiers, which m…
- Provisioned throughput removes the need for client-side backpressure.…
- Signal
Layout simulation is the nearest adjacency and the physical branch keeps advancing. Commercial calibration still has no published result. Red team argues our trigger is too strict; not yet changed.
- Physical world models are advancing on their own evidence and are alre…
- No published commercial world model reports held-out trajectory error …
- Learned layout simulators reach parity with discrete-event simulation …
- c-world-models-3 ↓ 0.2 → 0.14
- Tested
Hallucinated confirmations observed in test calls; vendor containment claim rated low for AU queues. Regional endpoint announced but not yet measured. Position: not yet above triage.
- Voice agents hallucinate confirmations at a rate (5–8% of calls) that …
- Vendor-reported 70% containment generalises to Australian queues once …
- Assessed
First validation run. The audit-as-product sold once and went nowhere; the buyer is the data-governance owner, not the CISO. QKD is irrelevant. Reframed as an inventory module inside existing engagements.
- The PQC migration case is independent of when a cryptographically rele…
- A standalone 'quantum readiness audit' does not sell repeatably; the i…
- Quantum key distribution hardware is a material part of the enterprise…
- Assessed
Gains are real on clean-feedback tasks and bounded. The regulatory constraint — inspectable, revocable rules — is firm and shapes the design more than the model does.
- A temporal decision log plus outcome feedback improves agent task succ…
- Learned behaviour that cannot be inspected and revoked by a human will…
- Tested
Nightingale run in progress. Vendor telemetry shown not to correlate with delivery outcomes. Counterfactual-survey defence logged as the strongest disconfirming voice and rated low.
- Vendor acceptance-rate telemetry measures usage, not value; it does no…
- A survey with a well-designed counterfactual question recovers the sam…
- Signal
Long-tail exceptions named as the entry point. Humanoid claims down-weighted. Red team's bundling argument is unanswered and the review date has now lapsed.
- The economic entry point in AU logistics is the long tail of exception…
- Mining autonomy in the Pilbara is mature and narrow; foundation models…
- Sim-to-real from learned world models will shorten the data-collection…
- c-robots-3 ↓ 0.26 → 0.18
- Tested
APRA treats agents as non-human identities under CPS 234; audit trail is the binding requirement. Broker PoC scoped as a Type 3 with a policy engine at the tool-call boundary.
- Regulators will treat agents as non-human identities under existing co…
- c-auth-broker-4 ↓ 0.31 → 0.18
- Assessed
Validation run over eleven labs: survivors reported decisions changed and lead time; the rest reported output. A single cost-per-recommendation metric rejected as sufficient. Three numbers proposed.
- Labs that report output volume — papers, demos, headcount — rather tha…
- Cost per validated recommendation is a sufficient single metric for th…
- Tested
Mid-run: spec-first arm ahead on lead time; review time climbing as a share. Industry throughput-up, stability-down report matches. Client question logged from retail confirms the demand.
- Review is the binding constraint on agentic delivery; throughput gains…
- c-ai-sdlc-4 ↓ 0.28 → 0.15
- Assessed
Procurement frameworks are the gate, not capability. Export controls are moving capacity. Cheap-tier parity is likely but not yet shown on our evals.
- AU enterprise procurement frameworks default to US providers and have …
- Export controls on accelerators are moving inference capacity toward S…
- Assessed
Inventory done: most flows idle, many duplicated, many unowned. Guardrails helped quality and not sprawl. The regulatory pressure is real; the board reflex is the risk.
- In the enterprise tenants inventoried, most citizen-built agents are i…
- A human review board is the only control APRA will accept for user-bui…
- Tested
Six patterns measured. One saving inverted after a prompt reorder. Routing adds 20–35% on narrow classes only. Field converging; caching claim strengthened, universal claim near dead.
- Prompt caching yields 30–45% per-task savings on stable prefixes over …
- Cache hit rate is fragile to prompt ordering; a reorder that moves a d…
- Batching halves cost on any workload that tolerates hours of latency, …
- c-token-cost-reduction-4 ↓ 0.3 → 0.12