cavendish

Crystal ball · queryable

Not “we think X”. Here is what we said in March, what changed, and why.

Validation runs are versioned, never overwritten. The brief is fixed; only the world varies. A run emits claims, the report is rendered from the claim set, and the diff is a set operation — stable, attributable, defensible in front of a board.

Validation runs
124
Fields with a position
40
  • What is our current position on X? · default
  • Mid-horizon, banking, above medium confidence · faceted
  • What changed since June? · the diff
  • What did we say twelve months ago? · time travel
Time traveltoday

Fields by horizon and confidence

90%70%50%30%10%
Now · 0–12 months
Near · 1–3 years
Next · 3–7 years
Distant · 7+ years

What changed

The diff — the thing nobody else can produce.

  • Eval Harnesses
    run 4 · 31 Aug 2026
    70%84%

    Both experiments concluded. Canary caught a real regression before any client did; judge calibration gives 0.71 on rubric tasks and 0.38 on open-ended. Unchecked judge to the graveyard; recommendation and standing answer current. Field converged; skills is the gate.

    • LLM-judge agreement with humans is task-dependent: around 0.7 kappa on
    • Release canaries catch material regressions on structured tasks within
    • Uncalibrated judges systematically favour newer and same-family models
    • Eval engineering is becoming a job title; the skills gate is closing o
    • c-eval-harnesses-5 ↓ 0.30 → 0.18
    Tested
  • Personal Wiki
    run 3 · 28 Aug 2026
    50%53%

    Experiment interim: repeat-task time down 44%; 6 of 52 compiled pages wrong and acted on. The page-error rate is now the primary open question and the gate is tooling for the review loop.

    • The compilation step produces confidently wrong pages at a material ra
    • c-personal-wiki-1 ↑ 0.6 → 0.7
    Tried
  • ROI
    run 4 · 28 Aug 2026
    61%74%

    Run concluded. Survey overstates telemetry 2–3×; net delivery gain 10–18%. Survey method sent to the graveyard; recommendation published. Gate moved from tooling to adoption.

    • Self-reported productivity gains from coding assistants overstate tele
    • Coding assistants cut PR cycle time but raise review load; the net del
    • c-ai-roi-3 ↑ 0.55 → 0.68
    • c-ai-roi-5 ↓ 0.35 → 0.24
    Tested
  • Dark Factory
    run 2 · 27 Aug 2026
    36%40%

    Lab's own analytics production is the reference case, running dark since June. Headcount claims overstated by uncounted control roles. Red team's marginal-value argument is the open question.

    • The lab's own analytics production can run unattended for a quarter wi
    • Task-level automation rates overstate the headcount effect because the
    • c-dark-factory-3 ↓ 0.22 → 0.15
    Signal
  • Fully Agentic QA
    run 4 · 27 Aug 2026
    55%61%

    Pilot concluded. Effort saving is real, escape rate is flat, coverage is a vanity metric. The role is re-shaped around the risk model, not removed. Full-replacement thesis sent to the graveyard.

    • Agent-run test maintenance and flaky-triage cut QA effort by more than
    • Line coverage produced by agents is uncorrelated with escape rate; cov
    • Agent-designed suites systematically under-test business-risk paths be
    • Agent-written and human-written suites catch different defect classes;
    • c-fully-agentic-qa-4 ↓ 0.38 → 0.24
    Tested
  • US Non-Dominance
    run 3 · 27 Aug 2026
    46%50%

    Parity confirmed on extraction and classification; frontier gap confirmed on agentic tasks. Provenance is the real objection. Position drafted for the sponsor's decision on whether to publish.

    • Chinese open-weight models are at parity with US open weights on extra
    • The frontier capability gap on agentic and long-horizon tasks has not
    • Provenance risk in open weights — training data, backdoors, licence te
    • c-us-non-dominance-4 ↓ 0.62 → 0.55
    Assessed
  • AI Gateway
    run 4 · 27 Aug 2026
    66%81%

    Routing experiment concluded: static rules capture the saving; learned routers do not generalise; agentic loops do not benefit. Field converged. Recommendation published; hand to practice in Q4.

    • Static task-class routing rules capture most of the achievable cost sa
    • Cost attribution per use case is the feature clients buy the gateway f
    • Adapter drift on streaming and tool-call formats is the recurring oper
    • Routing gains vanish on agentic loops because the loop needs one model
    • c-ai-gateway-3 ↑ 0.7 → 0.77
    Tested
  • AGI
    run 2 · 26 Aug 2026
    18%20%

    Canary coverage up to 58% of families at expert level, steady slope. Three failure modes unchanged in kind. Red team argues the tracker is capped by design; unanswered. Still Distant.

    • Frontier models clear more of the lab's professional task families eac
    • Three failure modes persist across every release: no learning from the
    • Transfer to an unpublished professional task family is the most discri
    • c-agi-3 ↑ 0.15 → 0.2
    Signal
  • Effective Assistants
    run 4 · 26 Aug 2026
    56%60%

    Calibration claim strengthened by a paper and a second engagement. Proactivity has a ceiling; above it the assistant gets muted. Standing answer drafted; experiment proposed.

    • Visible uncertainty and honest refusal increase use rather than reduce
    • Proactivity above a low threshold becomes interruption and drives the
    • c-effective-assistants-3 ↑ 0.58 → 0.66
    Tried
  • Deciding Table Stakes
    run 3 · 26 Aug 2026
    52%56%

    Position drafted: parity on capability, differentiation on measurement and the graveyard. Relationship counter-case rated higher than expected; verdict in doubt until Engel can show measured pitches win. Position and standing answer published with the doubt stated.

    • The only capability a client cannot obtain from any competitor is a me
    • Capability is irrelevant to consulting moats; relationships and distri
    • c-table-stakes-3 ↑ 0.52 → 0.60
    Assessed
  • TEE and Secure Inference
    run 4 · 25 Aug 2026
    52%57%

    Interim tax measured at about 12% at production batch shape; attestation verified client-side. Overhead-negligible claim retired. Regulator briefing drafted.

    • GPU enclave inference for a 70B model carries an 8–35% performance tax
    • Remote attestation can be verified by the client independently of the
    • c-tee-secure-inference-4 ↓ 0.35 → 0.15
    Assessed
  • Dark Harness
    run 3 · 24 Aug 2026
    45%47%

    Drift signal prototype caught half the failures. Acceptance cost did not fall until the fleet argued for its own PRs. Labour-saving claim downgraded; Type 3 proposed; client repos ruled out in writing.

    • The dominant failure mode of dark runs is plausible code that satisfie
    • A second-agent spec audit before PR catches most drift at a cost that
    • c-dark-harness-3 ↓ 0.4 → 0.23
    Tried
  • Backpressure
    run 4 · 24 Aug 2026
    66%76%

    Bench concluded: rate limits fail first, retries amplify 3.2×, admission control bounds spend at 1.1×. Recommendation published; gate is tooling because no harness ships the control.

    • Under load spikes, unbounded agent fan-out fails on provider rate limi
    • Admission control at the orchestrator with a per-task cost cap bounds
    • c-backpressure-4 ↑ 0.6 → 0.7
    Tested
  • Cost redux on tokens
    run 4 · 24 Aug 2026
    76%83%

    Ledger concluded. List-price falls dominate the levers over twelve months; distillation and pruning are the next factor with a maintenance bill. Recommendation and standing answer published; field converged.

    • List-price falls of roughly 4× a year dominate every engineering lever
    • Task-specific distillation gives the next cost factor on narrow tasks
    • c-token-cost-reduction-1 ↑ 0.78 → 0.86
    Tested
  • Agentic Memory System
    run 4 · 21 Aug 2026
    68%72%

    Bench confirms the default. Vendors converging on three-store design; consolidation policy is the differentiator and remains hand-rolled. Recommendation published.

    • Vendors are converging on a working/episodic/semantic three-store desi
    • c-agentic-memory-1 ↑ 0.74 → 0.84
    Tested
  • On-Prem Inference
    run 4 · 20 Aug 2026
    70%74%

    Hosted AU-region prices down about 40% year on year; crossover moved further from typical utilisation. Position held on a third client measurement. Standing answer refreshed.

    • The cost crossover between owned H100-class hardware and AU-region hos
    • Enterprise inference clusters run at 20–40% sustained utilisation; the
    • Hosted AU-region inference prices fell about 40% in twelve months, mov
    • c-on-prem-inference-4 ↓ 0.28 → 0.20
    Assessed
  • Quantum Compute General
    run 2 · 19 Aug 2026
    50%55%

    Logical-qubit counts up, error rates down, both on schedule. The client-facing date is regulatory, not physical — a regulator will set a PQC deadline first. Overhead reductions are the risk to our estimate.

    • Quantum chemistry is the first application to cross classical parity,
    • AU regulators will set a PQC deadline before the compute threshold is
    • c-quantum-compute-2 ↓ 0.72 → 0.66
    Signal
  • Quantum Encryption
    run 3 · 19 Aug 2026
    56%63%

    Second validation run plus a Type 2 on the lab's own infrastructure. Inventory dominates the migration; agent-assisted scanning makes it a days-not-months exercise. Entry window closes around 2029 as big-four practices consolidate. Standing answer published.

    • The inventory phase dominates PQC migration effort and timeline; the a
    • Agent-assisted scanning of code and configuration repositories can pro
    • The big-four consultancies' PQC practices will have consolidated the A
    • c-quantum-encryption-1 ↑ 0.78 → 0.86
    Assessed
  • Learning Agents
    run 3 · 19 Aug 2026
    48%54%

    Late, noisy feedback is the open problem and the experiment is aimed at it. Structured policy stores do not beat text rules yet. Field stays emerging until x-learning-agent-memory concludes.

    • The gain flattens within a few hundred decisions; there is no evidence
    • Late, noisy outcome feedback destroys the gain; nobody has demonstrate
    • c-learning-agents-4 ↓ 0.45 → 0.30
    Assessed
  • Voice and Vision
    run 4 · 19 Aug 2026
    58%63%

    Latency bench concluded. US-hosted floor is unusable; regional endpoint halves it; barge-in is the abandonment driver. Recommendation published; IVR-replacement thesis to the graveyard. Field stays contested because the floor is moving.

    • From Sydney, p95 turn latency on US-hosted realtime models exceeds the
    • Barge-in failure, not recognition accuracy, is the dominant cause of c
    • c-voice-and-vision-3 ↑ 0.62 → 0.70
    Tested
  • AI-SDLC
    run 4 · 19 Aug 2026
    63%71%

    Experiment concluded: 38% lead-time cut under spec-first, review now half of lead time. Harness vendors shipping persisted plans. Recommendation and position published; skills is the gate.

    • Spec-first agentic delivery cuts change lead time by a third or more;
    • Harness vendors are moving the plan or spec to a persisted first-class
    • c-ai-sdlc-2 ↑ 0.6 → 0.76
    Tested
  • Non-Weight-Bound Continuous Learning
    run 4 · 14 Aug 2026
    47%58%

    Skill library shows an 11-point gain at six weeks; contradictions are the named risk. Weight-level continual learning stays Distant. Position drafted for publication.

    • A retrieval-updated skill library improves a deployed agent's task acc
    • Test-time training on the current task context improves long-document
    • Skill libraries accumulate contradictions over time; without a consoli
    • Nightly fine-tuning as a substitute for continual learning regresses h
    • c-continuous-learning-3 ↓ 0.36 → 0.28
    Assessed
  • Personal Wiki
    run 2 · 14 Aug 2026
    45%50%

    Humans edit the pages in practice, which is the property that distinguishes this from opaque summaries. Vendor skills primitives announced; compile-and-review identified as the differentiator.

    • Human-readable wiki pages are edited by humans in practice; the readab
    • Foundation labs are converging on a first-party 'skills' primitive, wh
    Tried
  • AI-DLCs
    run 3 · 14 Aug 2026
    50%55%

    Risk function accepted the agent-drafted model document. Claims handling identified as the highest-value target and the least specified. Type 3 on internal analytics delivery proposed; regulated lifecycles held.

    • Model-risk documentation can be drafted by an agent from the model rep
    • Claims handling is the highest-value non-software lifecycle for the fi
    • c-ai-dlcs-4 ↓ 0.35 → 0.2
    Tried
  • Ambient Agents
    run 4 · 14 Aug 2026
    42%44%

    Second independent placement on the board. Draft-only retention confirmed in a second pod. The lab still disagrees on whether the field is distinct from Effective Assistants; decision due Q4.

    • Draft-only ambient behaviour retains users where act-unprompted does n
    • c-ambient-agents-1 ↑ 0.66 → 0.74
    Tried
  • Cyber Cold War
    run 3 · 14 Aug 2026
    45%48%

    SOCI guidance makes AI model supply a dependency; competitors are staffing. The sceptic's case (flat incident data, defence automating equally) logged and rated above what we expected. Field stays assessed.

    • SOCI-regulated entities will be required to treat AI model supply as a
    • Incident statistics show no AI effect; the threat is priced-in vendor
    • c-cyber-cold-war-1 ↓ 0.76 → 0.70
    Assessed
  • Auth Broker
    run 4 · 14 Aug 2026
    60%68%

    PoC concluded: the pattern works at negligible latency and the cost is policy authoring. Two vendor previews shipped mid-run. Recommendation published; policy tooling is the next gate.

    • A policy engine at the tool-call boundary adds under 15 ms p50; the co
    • Identity vendors will ship agent-identity primitives before a standard
    • c-auth-broker-1 ↑ 0.7 → 0.82
    Tested
  • Self-Organising Agents
    run 2 · 12 Aug 2026
    28%31%

    Vendor 'team' primitives change the urgency, not the evidence. Contamination by a wrong agent is now the named reliability gate. Trigger unchanged.

    • Harness vendors shipping 'team' primitives means the pattern will appe
    • One confidently wrong agent contaminates the collective; no published
    • c-self-organising-agents-3 ↓ 0.31 → 0.24
    Signal
  • Multiplayer AI Surfaces
    run 2 · 12 Aug 2026
    32%36%

    Turn-taking failures are frequent enough to need a policy. Shared memory, not presence, is where the value is. Coordination-cost evidence cuts against the thesis. Experiment proposed; election pending a second pod.

    • When more than one human can address the same agent, contradictory-ins
    • A shared memory store is what makes a multiplayer surface useful; pres
    • c-multiplayer-ai-surfaces-4 ↓ 0.40 → 0.33
    Signal
  • SLM / Edge / Tuning
    run 4 · 12 Aug 2026
    60%70%

    Edge classifier concluded at 94% of frontier F1. Distillation is the tuning path; retrieval wins knowledge tasks; the cost case erodes with frontier pricing but residency and latency do not. Recommendation and standing answer published.

    • Distillation from frontier-generated labels reaches more than 90% of f
    • Retrieval plus a frontier model beats fine-tuning for knowledge-bearin
    • Frontier price declines erode the small-model cost case on roughly a t
    Tested
  • Lab outcomes visibility
    run 3 · 12 Aug 2026
    50%56%

    Ninety-day measurement: 412 opens, 38 distinct readers, four people at 61%. Opens rejected as the KPI figure; distinct readers and citations-in-pitch proposed instead. Citation capture depends on Engel.

    • Library opens are concentrated: four people account for 61% of reads,
    • Opens are not consumption; citation in a pitch is the measure that ref
    • c-lab-outcomes-visibility-4 ↓ 0.3 → 0.2
    Assessed
  • Sensing Agents
    run 3 · 7 Aug 2026
    50%58%

    Early experiment result supports the two-stage design at roughly 4% of naive cost with actionable events the rules engine missed. Additive to rules, not a replacement. Explanation of sources checked is what keeps operators reading.

    • A cheap-gate two-stage design cuts the cost of watching a stream by 20
    • Sensing agents find a class of event that threshold rules miss — multi
    • Operators will accept sensed events alongside rules-engine alerts only
    • c-sensing-agents-4 ↓ 0.38 → 0.24
    Tried
  • Post Economy
    run 2 · 5 Aug 2026
    30%35%

    Own-delivery substitution measured at 20–30% of junior hours. Outcome pricing appearing at competitors on the substituted share only. Jevons counter-case unanswered; horizon still contested.

    • Frontier-class inference price at constant capability has fallen rough
    • In the lab's own delivery, tokens now substitute for 20–30% of junior
    • Consultancies that move to outcome pricing do so on the parts of their
    • c-post-economy-3 ↓ 0.38 → 0.3
    Signal
  • Open Weight Models
    run 4 · 5 Aug 2026
    64%74%

    Parity experiment concluded: within 2 points on three task classes, 14–22 behind on agentic, gap re-opens each frontier release. Origin-as-blocker carried at moderate confidence. Recommendation and standing answer published.

    • Open-weight models are within 2 points of frontier on classification,
    • The open-weight gap on multi-step agentic tasks is 14–22 points and re
    • Model origin (a Chinese lab) is a procurement blocker in the AU public
    • c-open-weight-models-3 ↑ 0.62 → 0.72
    Tested
  • Cyber Cold War
    run 2 · 1 Aug 2026
    38%45%

    Offensive capability is rising faster than SOC adoption; export controls arrive as tenant obligations rather than denial. The near-term client question is regulatory, not technical.

    • Model capability on offensive-security evals is rising faster than def
    • Export controls will reach Australian clients as cloud-tenant KYC and
    Assessed
  • Effective Assistants
    run 3 · 31 Jul 2026
    48%56%

    Four properties named. Proactivity and integration depth are supported by telemetry; calibration and refusal by one tried result and thin literature. Gate set to adoption: the properties are buildable and not asked for.

    • Assistants with a proactive surface retain roughly three times the wee
    • Organisations do not procure for the four properties because the procu
    • c-effective-assistants-4 ↓ 0.30 → 0.20
    Tried
  • Citizen Developers and Org Slop
    run 3 · 31 Jul 2026
    52%60%

    Guardrail thesis to the graveyard. Retention pattern worked in our tenant; the board case-study shows the value cost. Recommendation published at assessed tier with the second-tenant caveat.

    • Template guardrails reduce obvious failures but not sprawl, because sp
    • Platform-enforced retention (owner, last-run TTL, auto-archive) remove
    • Review boards reduce agent count and reduce value roughly in proportio
    • c-citizen-developers-org-slop-5 ↓ 0.42 → 0.30
    Assessed
  • Lab Effectivity
    run 3 · 31 Jul 2026
    47%52%

    First lead-time measurement from the graph: median +4.2 months, three negatives. Lead time confirmed as the number that cannot be reconstructed later. Field stays open until December, then dissolves or does not.

    • Lead time is the only proof of worth a lab cannot reconstruct retrospe
    • The lab's current median lead time is +4.2 months across sixteen now-f
    • c-lab-effectivity-2 ↑ 0.58 → 0.64
    Assessed
  • AI-DLCs
    run 2 · 30 Jul 2026
    45%50%

    Two tried results logged: analytics delivery and model-risk documentation. Both had a versioned artifact beforehand. The scaffold, not the agent, looks like the precondition.

    • The spec-first agentic pattern transfers to non-software lifecycles on
    • Analytics delivery cycle time falls by more than half under an agentic
    Tried
  • Agentic Memory System
    run 3 · 30 Jul 2026
    58%68%

    Structured episodic + summarised recall is the working default. Graph memory is a special case, not a default. Long-context is not a substitute.

    • Summarised episodic recall beats raw chunk retrieval on decision-consi
    • Graph-structured memory improves entity-heavy tasks and degrades gener
    • c-agentic-memory-4 ↓ 0.41 → 0.22
    Tested
  • Deciding Table Stakes
    run 2 · 21 Jul 2026
    45%52%

    Competitor 'proprietary' claims map to open tooling; evals are a job description in six months. The moat candidate is measured outcomes on client data. Red team's relationship case logged.

    • Evals are claimed by half the competitor set and demonstrated publicly
    • 'Proprietary orchestration' claims in competitor pitches map to open-s
    Assessed
  • Dark Harness
    run 2 · 20 Jul 2026
    40%45%

    Migration backlogs identified as the production-ready class. Two engineers ran an overnight fleet on an internal repo; the split between precise and imprecise tickets is stark. Field marked contested — the lab disagrees on whether the gap is spec or signal.

    • Unattended agent fleets clear well-specified, testable tickets at high
    • Migration and dependency-upgrade backlogs are the one production-ready
    Tried
  • Ambient Agents
    run 3 · 17 Jul 2026
    40%42%

    Reversibility predicts survival better than accuracy. Draft-only is deliverable; act-unprompted is not. Delegated authority is a hard requirement in regulated sectors.

    • Reversibility of the action, not accuracy of the decision, predicts wh
    • Ambient agents acting under an integration token rather than delegated
    • c-ambient-agents-4 ↓ 0.30 → 0.18
    Tried
  • Education
    run 2 · 16 Jul 2026
    35%42%

    Internal upskilling cohort result logged. The firm's own capability building is the tractable application; the external market is not. Remain a candidate; run the next cohort with a control.

    • The tutoring effect transfers to open-ended, judgement-heavy skills at
    • The firm's own upskilling is the tractable application; the external e
    Signal
  • Lab outcomes visibility
    run 2 · 15 Jul 2026
    42%50%

    Field experiment at a consultancy gives a push-over-pull multiplier with decay. Staff-privacy question raised by the product engineer who built the telemetry; aggregation threshold adopted before the KPI reports.

    • Push digests roughly triple reach against pull for about eight weeks,
    • Per-person read telemetry on internal documents is a staff-privacy que
    Assessed
  • Fully Agentic QA
    run 3 · 14 Jul 2026
    50%55%

    Vendor claims of unattended QA do not survive contact with a real repo. Maintenance and triage are the tractable half; design and sign-off are not. Pilot preregistered with escape rate as primary.

    • A fully unattended QA function — design through sign-off — is deployab
    • The QA role is being re-titled, not removed: openings for 'test engine
    Tested
  • TEE and Secure Inference
    run 3 · 14 Jul 2026
    48%52%

    Regulator survey: no AU acceptance of attestation as a control. Field's gate confirmed as economics plus adoption; economics first because the number is ours to produce.

    • No AU regulator has accepted attestation as a control for PII; until o
    Assessed
  • Backpressure
    run 3 · 9 Jul 2026
    58%66%

    Providers moving to token-priority tiers; the 'buy throughput' counter-case logged and rated low. Bench running; interim results consistent with hypothesis.

    • Provider rate limits are moving to token-based priority tiers, which m
    • Provisioned throughput removes the need for client-side backpressure.
    Tested
  • World Models
    run 2 · 3 Jul 2026
    25%27%

    Layout simulation is the nearest adjacency and the physical branch keeps advancing. Commercial calibration still has no published result. Red team argues our trigger is too strict; not yet changed.

    • Physical world models are advancing on their own evidence and are alre
    • No published commercial world model reports held-out trajectory error
    • Learned layout simulators reach parity with discrete-event simulation
    • c-world-models-3 ↓ 0.2 → 0.14
    Signal
  • Voice and Vision
    run 3 · 2 Jul 2026
    52%58%

    Hallucinated confirmations observed in test calls; vendor containment claim rated low for AU queues. Regional endpoint announced but not yet measured. Position: not yet above triage.

    • Voice agents hallucinate confirmations at a rate (5–8% of calls) that
    • Vendor-reported 70% containment generalises to Australian queues once
    Tested
  • Quantum Encryption
    run 2 · 30 Jun 2026
    50%56%

    First validation run. The audit-as-product sold once and went nowhere; the buyer is the data-governance owner, not the CISO. QKD is irrelevant. Reframed as an inventory module inside existing engagements.

    • The PQC migration case is independent of when a cryptographically rele
    • A standalone 'quantum readiness audit' does not sell repeatably; the i
    • Quantum key distribution hardware is a material part of the enterprise
    Assessed
  • Learning Agents
    run 2 · 30 Jun 2026
    40%48%

    Gains are real on clean-feedback tasks and bounded. The regulatory constraint — inspectable, revocable rules — is firm and shapes the design more than the model does.

    • A temporal decision log plus outcome feedback improves agent task succ
    • Learned behaviour that cannot be inspected and revoked by a human will
    Assessed
  • ROI
    run 3 · 30 Jun 2026
    52%61%

    Nightingale run in progress. Vendor telemetry shown not to correlate with delivery outcomes. Counterfactual-survey defence logged as the strongest disconfirming voice and rated low.

    • Vendor acceptance-rate telemetry measures usage, not value; it does no
    • A survey with a well-designed counterfactual question recovers the sam
    Tested
  • Robots
    run 2 · 25 Jun 2026
    35%38%

    Long-tail exceptions named as the entry point. Humanoid claims down-weighted. Red team's bundling argument is unanswered and the review date has now lapsed.

    • The economic entry point in AU logistics is the long tail of exception
    • Mining autonomy in the Pilbara is mature and narrow; foundation models
    • Sim-to-real from learned world models will shorten the data-collection
    • c-robots-3 ↓ 0.26 → 0.18
    Signal
  • Auth Broker
    run 3 · 25 Jun 2026
    52%60%

    APRA treats agents as non-human identities under CPS 234; audit trail is the binding requirement. Broker PoC scoped as a Type 3 with a policy engine at the tool-call boundary.

    • Regulators will treat agents as non-human identities under existing co
    • c-auth-broker-4 ↓ 0.31 → 0.18
    Tested
  • Lab Effectivity
    run 2 · 24 Jun 2026
    40%47%

    Validation run over eleven labs: survivors reported decisions changed and lead time; the rest reported output. A single cost-per-recommendation metric rejected as sufficient. Three numbers proposed.

    • Labs that report output volume — papers, demos, headcount — rather tha
    • Cost per validated recommendation is a sufficient single metric for th
    Assessed
  • AI-SDLC
    run 3 · 18 Jun 2026
    55%63%

    Mid-run: spec-first arm ahead on lead time; review time climbing as a share. Industry throughput-up, stability-down report matches. Client question logged from retail confirms the demand.

    • Review is the binding constraint on agentic delivery; throughput gains
    • c-ai-sdlc-4 ↓ 0.28 → 0.15
    Tested
  • US Non-Dominance
    run 2 · 17 Jun 2026
    40%46%

    Procurement frameworks are the gate, not capability. Export controls are moving capacity. Cheap-tier parity is likely but not yet shown on our evals.

    • AU enterprise procurement frameworks default to US providers and have
    • Export controls on accelerators are moving inference capacity toward S
    Assessed
  • Citizen Developers and Org Slop
    run 2 · 9 Jun 2026
    40%52%

    Inventory done: most flows idle, many duplicated, many unowned. Guardrails helped quality and not sprawl. The regulatory pressure is real; the board reflex is the risk.

    • In the enterprise tenants inventoried, most citizen-built agents are i
    • A human review board is the only control APRA will accept for user-bui
    Assessed
  • Cost redux on tokens
    run 3 · 3 Jun 2026
    66%76%

    Six patterns measured. One saving inverted after a prompt reorder. Routing adds 20–35% on narrow classes only. Field converging; caching claim strengthened, universal claim near dead.

    • Prompt caching yields 30–45% per-task savings on stable prefixes over
    • Cache hit rate is fragile to prompt ordering; a reorder that moves a d
    • Batching halves cost on any workload that tolerates hours of latency,
    • c-token-cost-reduction-4 ↓ 0.3 → 0.12
    Tested