cavendish
AssessedEmerginggate · AdoptionNow · 0–12 months×4 sightings

Cyber Cold War

AI is tilting the offence/defence balance toward whoever automates vulnerability discovery first, but for Australian critical-infrastructure clients the near-term consequence is not a novel attack — it is a compliance and procurement squeeze from model and compute export controls plus SOCI obligations, and that lands before any measurable change in incident rates.

A validation run. Researched position, no experiment.

Join with…

Confidence

48%human-committed

Expiry

6doverdue for review

Lead time

not yet mainstream · opened 8 Jul 2026

Ownership

LFLena Fischerfortnightly cadence

Where it is

Opened in July from a band-1 signal: a foundation lab's public attribution of a state-run intrusion campaign in which an agent did most of the tactical work. Since then the picture has split in two. On the offensive side, model performance on capture-the-flag evals roughly doubled inside nine months and an open-source pentest agent found real misconfigurations in our own staging harness in under an hour. On the defensive side, AI-assisted patching is closing the disclosure-to-patch window and the incident statistics show no shift yet. What has moved for clients is the rulebook: revised US controls on advanced compute and model weights arrive as know-your-customer obligations on cloud tenants, and CISC guidance now treats AI model supply as a critical-asset dependency under SOCI. Competitors have noticed — a big-four firm is hiring an 'AI cyber resilience' practice of twelve. The field is emerging and assessed only; nothing has been tested, and the red team's case that this is priced-in marketing is not weak.

Why a Quantium decision hinges on it

Quantium serves energy, telco, banking and government, all of which hold SOCI-designated assets, and Woolworths itself is a food-and-grocery critical asset. The question those clients will ask is not 'will an AI attack us' but 'does the US model in our SOC vendor's stack count as a supply-chain risk in our risk management program, and does our sovereign-cloud tenancy survive the export rule'. Answering that requires a position, not an experiment, and it has to exist before the first RFP asks. It is also the field where the firm's data-and-AI credibility can be tested against a cyber practice it does not have.

Field attributes

StateEmerging
GateAdoption · ready — blocked by trust, regulation, procurement or change capacity
OriginSignal
Measurablepartial
Audience · TLPexec
Horizonnow
Opened8 Jul 2026
Mainstreamnot yet
Last validated14 Aug 2026
Sightings4

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01A foundation lab attributed a state-run intrusion campaign in which an agent performed the majority of tactical work with limited human direction. Public, first-party, band 1.
  • 02Frontier-model solve rates on published offensive-security evals roughly doubled between mid-2025 and early 2026.
  • 03An open-source pentest agent found two authorisation misconfigurations in our own staging harness in forty minutes with no bespoke prompting (tried tier, one run).
  • 04CISC guidance now names AI model supply as a dependency SOCI risk management programs must cover.
What is hype
  • 01'AI superworm' and autonomous-cyberweapon headlines. Nothing in the incident record shows autonomous propagation; the attributed campaign was agent-assisted, human-directed.
  • 02Vendor 'AI-native SOC' rebrands. Most are anomaly detection with a chat front-end, sold into a fear the vendor helped create.
  • 03Export controls as a ban. For Australia they are a paperwork and residency obligation, not a denial of access.
What would have to be true
  • 01A measurable divergence between attacker and defender automation rates — currently both sides are accelerating and the net is unknown.
  • 02SOCI enforcement action against an entity for an AI supply-chain gap, which would convert guidance into a procurement requirement.
  • 03A named client engagement where the risk-management question is asked in an RFP, not a corridor. Engel has one corridor question so far.
What we would do
  • 01Hold the field at assessed and write a one-page position on 'AI model supply as a SOCI dependency' for sector owners to use, tiered as assessed and marked as such.
  • 02Run the open pentest agent against every lab harness as a Type 2 and log what it finds; cheap and it produces our own evidence.
  • 03Do not build a cyber practice. Partner for the SOC work; own the data-governance and model-supply half of the question.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
80–90%
agent share of tactical work
Announcement·band 1Signal

Foundation lab attributes agent-orchestrated intrusion campaign to a state actor

First-party disclosure that a state-sponsored group used an agentic model to run reconnaissance, exploitation and exfiltration against around thirty targets, with the agent performing most tactical steps and humans intervening at a handful of decision points. The signal that opened the field, eight months later.

extracted claimAn agentic model performed the majority of tactical intrusion work in a real state-run campaign.
Anthropic13 Nov 2025
LFAWdropped 5
12
openings
Job posting·band 3Signal

Big-four firm hiring 'AI Cyber Resilience' practice ×12 in Canberra

Argus competitor watch. Twelve roles across SOCI advisory, AI supply-chain assurance and model-risk. Inference: the competitor reads the guidance the same way we do and is staffing ahead of audit demand.

Competitor careers page5 Aug 2026
detector · demand
2
findings in 40 min
Finding·band 1Tried

Logged from Claude Code: open pentest agent found two auth misconfigs in our staging harness in forty minutes

Red-team lead pointed the open-source agent at the lab's staging gateway with a stock prompt. It found a missing scope check on a tool endpoint and an over-broad CORS policy. Both real, both fixed. One run, one harness, tried tier.

MCP · log_finding · Lena Fischer30 Jul 2026
LFdropped
Post·band 2Signal

'The AI cyber threat is a marketing threat'

Points at flat breach counts and unchanged initial-access vectors (phishing, credentials, unpatched edge devices) and argues nothing an agent does changes the chain's bottleneck. Well argued; kept as the strongest disconfirming voice.

Personal blog · A veteran incident-response practitioner25 Jul 2026
MLdropped 3
Client question·band 3Signal

'If our SOC vendor runs a US model, is that a SOCI supply-chain risk?'

Asked in a corridor after a data-platform review, not in the engagement scope. Unanswered at the time; the sector lead logged it. First demand signal for the field and the one that set its compliance framing.

Engel · energy engagement24 Jul 2026
?dropped 2
−54%
median disclosure→patch
Paper·band 1Signal

Closing the window: AI-assisted patch generation and disclosure-to-patch latency

Measures time from CVE disclosure to merged patch across open-source projects using AI-assisted patching. Median window fell by more than half. Defence-side evidence; supports the sceptic's case that the net balance is unknown.

arxiv.org · Haddad, Lindqvist et al.18 Jun 2026
detector · bleeding edge 2
9.8k
stars
Repository·band 2Tried

Open-source autonomous pentest agent crosses 10k stars

Agent that plans and executes a web-app penetration test end to end using a frontier model. Star velocity tripled after a conference demo. We ran it against our own harness (see finding below).

github.com5 Jun 2026
detector · early adoption 2
Regulatory·band 2Signal

CISC guidance: AI model supply as a dependency in SOCI risk management programs

Updated guidance for responsible entities under the SOCI Act names third-party AI models and inference services as supply-chain hazards a risk management program must identify and treat. Guidance, not a rule; the audit questions will follow it.

extracted claimSOCI responsible entities are expected to treat AI model supply as a critical-asset dependency.
Cyber and Infrastructure Security Centre20 May 2026
detector · early adoption 2
34% → 66%
CTF solve rate
Benchmark·band 1Signal

Offensive-security eval: frontier CTF solve rate doubles in nine months

Published capture-the-flag and vulnerability-discovery suite versioned across model releases. Top solve rate went from the low thirties to the mid sixties between mid-2025 and this release. We carry it as capability signal, not threat signal.

Public eval leaderboard2 Apr 2026
detector · bleeding edge 3
Regulatory·band 2Signal

Revised US rule on advanced-compute and model-weight exports

Replaces the rescinded 2025 diffusion framework. Australia is in the least-restricted tier, but the rule attaches know-your-customer and reporting obligations to cloud providers serving foreign tenants, which flow down to AU customers as contract terms and residency attestations.

US Bureau of Industry and Security15 Jan 2026
detector · early adoption 3
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

State actors are already using agentic models for the majority of tactical intrusion work, with humans directing rather than executing.

Assessedc-cyber-cold-war-1dalton-0.415 Jul 2026Anthropic
70%

Export controls will reach Australian clients as cloud-tenant KYC and weight-residency obligations, not as denial of access.

Assessedc-cyber-cold-war-3dalton-0.41 Aug 2026US Bureau of Industry and Security, Engel · energy engagement
62%

Model capability on offensive-security evals is rising faster than defensive-tooling adoption in Australian SOCs.

Assessedc-cyber-cold-war-2dalton-0.41 Aug 2026Public eval leaderboard, github.com, MCP · log_finding
58%

SOCI-regulated entities will be required to treat AI model supply as a critical-asset dependency within eighteen months.

Assessedc-cyber-cold-war-4dalton-0.414 Aug 2026Cyber and Infrastructure Security Centre, Competitor careers page
55%

Incident statistics show no AI effect; the threat is priced-in vendor marketing and defensive automation cancels the offensive gain.

Assessedc-cyber-cold-war-5dalton-0.414 Aug 2026Personal blog, arxiv.org
35%

Position history · the diff is the product

3 validation runs against a fixed brief. Confidence 38% → 48%.

runs compare claim sets, never prose
What we said · run 3

SOCI guidance makes AI model supply a dependency; competitors are staffing. The sceptic's case (flat incident data, defence automating equally) logged and rated above what we expected. Field stays assessed.

48%
Changed since run 2
  • SOCI-regulated entities will be required to treat AI model supply as a critical-asset dependency within eighteen months.
  • Incident statistics show no AI effect; the threat is priced-in vendor marketing and defensive automation cancels the offensive gain.
  • c-cyber-cold-war-1 ↓ 0.76 → 0.70
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · LF
high

If the compliance reading is right it touches every SOCI client's risk program; if the offence reading is right it touches everything.

Timeline

committed · LF
0–18mo

The regulatory half is already in guidance; the capability half is already in the incident record.

TAM

agent-estimated
$1B–10B

Agent-estimated from AU cyber-services spend attributable to AI-specific obligations. Wide error bars; uncommitted.

Cost of being wrong

committed · AW
high

Wrong in either direction is expensive: a false alarm burns credibility with CISOs, a miss lands in a SOCI audit.

Demand

committed · AB
medium

One corridor question in government, one in energy. No RFP language yet.

Cost of entry

agent-estimated
high

The firm has no cyber practice. Entry means partnering. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Energy & Utilities
relevant

Energy assets are the first SOCI sector where AI model supply appears in a risk management program review.

Mechanism · Map every AI component in the client's OT-adjacent stack to its model provider and jurisdiction; treat as a material service provider. Agent draft.

Agent draft · awaiting a sector owneragent-estimated
Government
relevant

Agencies are both SOCI regulators and SOCI entities, and the sovereign-cloud question is theirs to answer first.

Mechanism · Position paper on weight residency and cloud-tenant KYC; align to the hosting certification framework.

AB committed by Aisha Bellocommitted
Retail & FMCG
watch

Woolworths is a food-and-grocery critical asset, so the obligations exist, but the state-actor target profile does not.

Mechanism · Compliance mapping only; no threat-driven work unless a supply-chain incident names a retailer.

DS committed by Dev Sharmacommitted
Insurance
not-relevant

Insurers' AI exposure runs through the cyber-insurance product line, which is an underwriting question, not a SOCI or export-control one.

Mechanism · None from this field; the underwriting question belongs in a different candidate if it appears.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against: this field is a headline with a compliance tail. Attackers have automated for a decade; agentic models change the speed of one step in a chain that is limited elsewhere. Defenders are automating just as fast, the incident record is flat, and the regulatory guidance would have arrived without AI. We opened a field on a single lab's attribution, four months after the story was mainstream, and dressed it in SOCI language to make it ours.

  • One attributed campaign from one lab with a commercial interest in describing the threat. No independent confirmation of the agent's share of work.
  • Offensive-eval solve rates measure a benchmark, not a breach. CTF performance has never been shown to predict intrusion rates.
  • The disclosure-to-patch window is shrinking faster than the discovery window; the net could favour defence and we have not measured it.
  • The firm has no cyber practice. Advising on this from a data-science position invites the question of why anyone should listen.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • Foundation labs20%
  • Security research and practitioners30%
  • Regulators25%
  • Competitors / Argus10%
  • Internal / Engel15%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsLFAWABMTTO

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Is the net offence/defence balance measurable from any public series, or only from incident data nobody publishes?
  2. 02When does CISC guidance on AI model supply become an audit finding, and against which sector first?
  3. 03Can the firm credibly advise on the model-supply half of the question without a cyber practice, and who is the partner for the other half?