
Leading platform reports 48 logical qubits at 1e-4 logical error, sustained for hours
Tens of logical qubits at 1e-4 error are now demonstrated, not projected.
The Wire · continuously watched, ranked, cited
External signal and the lab's own output in one stream, ranked by velocity rather than volume, with a synthesis composed over whatever you have in scope. Configure it once and it is remembered — including whose drops you want to follow.
Composed for you · 439 items in scope
Rendered from the graph. Nothing here is written by hand.36 items landed in the last fortnight, led by Quantum Compute General and Eval Harnesses.12
3 sources have now been seen independently by more than two people or detectors, which is a strength signal rather than a duplicate.134
Demand is loudest around Deciding Table Stakes, Ambient Agents and Personal Wiki — 53 signals in the last quarter came from clients asking rather than from the field moving, and that gap is where advisory value sits.56
From the lab: 39 items are now tested and citable and 16 lines were killed and published with the cause.27
15 people have dropped into this scope directly; where several people's drops meet, an informal working group usually already exists.13
Sources

Tens of logical qubits at 1e-4 error are now demonstrated, not projected.

Judge reliability is a property of the rubric and references, not the judge model, and uncalibrated judges favour their own family.
3Skill storage and retrieval are becoming a first-party platform primitive.

Dark runs succeed on precise tickets and fail on implicit requirements, and human hours move rather than fall.

Self-reported coding-assistant productivity gains overstate telemetry-measured gains by two to three times.

Agentic QA cuts maintenance effort by more than half without moving the escape rate.

Static task-class routing captures the cost saving; learned routers do not generalise.

Compiled, page-granular skill knowledge outperforms trajectory retrieval on long-horizon agent tasks.
Analytics report production can run unattended for a quarter at an exception rate under 5%.

Agent fan-out fails on rate limits before latency, and orchestrator-level admission control bounds both spend and tail latency.

The levers are real, bounded and fragile, and none of them is 60%.
Cache when the shared prefix is over ~2k tokens and reused inside the provider's TTL. Budget on 30–45% input-cost reduction across a real pattern; the 60% figure is a best-case single call.
Default classification and structured extraction to an open-weight 30B-class model behind the gateway. Keep multi-step agentic loops on frontier until the open-weight tool-use gap closes.
5Capability parity is irrelevant because consulting is won on relationships and distribution.

Above ~50k tokens of history, retrieval latency — not storage — is the dominant cost of an agent memory layer.

Calibrated uncertainty sustains use; miscalibrated uncertainty destroys it faster than overconfidence does.

From Sydney, US-hosted realtime voice exceeds the talk-over threshold at p95; barge-in failure, not recognition, drives abandonment.

The spec, not the code, is where agentic delivery gains come from; review is where they stop.

Humans edit compiled pages in practice, and the edit propagates to future sessions.

Agents can draft model-risk documentation to a risk function's standard from the repo alone.
Blended across our six instrumented patterns, 1 September: $0.9–1.6 per thousand agent turns on the open-weight tier, $4–11 on frontier mid-tier. Auto-refreshed from the ledger; not yet…
Energy's share of inference marginal cost is rising and on track to dominate before 2030.

A pre-PR spec audit catches roughly half of silent drift at low cost.
For any agent with more than ten sessions of history, store episodes with summaries and recall the summary first. Raw vector recall over transcripts loses on precision and latency above…
Do not report AI productivity gains from self-reported surveys. Instrument the workflow and measure cycle time, throughput and rework before and after.
Frontier models clear more professional task families each release on a steady, not accelerating, slope.

A tool-call policy engine is cheap to run and expensive to author.
The reviewable, versioned artefact in an agentic delivery is the specification and its acceptance tests. Code is generated output. Teams that review code first lose the throughput gain to…
Bound the number of in-flight sub-agents and tool calls at the orchestrator with a queue that applies backpressure. Model-side rate limits are the wrong place to discover you have a…
A retrieval-updated skill library improves a deployed agent's accuracy over weeks without weight updates.

Distillation from a frontier teacher, not fine-tuning on domain labels, is the tuning path that reaches frontier-adjacent quality on narrow tasks.
Qwen 3.5 32B through the gateway's extraction tier for flat and moderately nested schemas; Claude Sonnet 5 for schemas over ~40 fields or with cross-field constraints. As of 27 August.

Owned H100-class inference beats AU-region hosted only above roughly 60% sustained utilisation.

AU-region hosted inference prices are falling faster than owned-hardware costs.

The wiki cuts repeat-task time and produces wrong pages agents trust; both effects are large.
Most of what consultancies sold as AI differentiation in 2024 is now table stakes: a gateway, an eval harness, a calibrated judge, a memory default, delegated auth. The moat is not in…

Non-US open weights are at parity on the cheap tier and behind on the agentic tier.
Every client pattern goes through a gateway the delivery team owns the config for. Vendor gateways are fine as a backend, never as the routing policy.

Proactivity and integration depth are the strongest correlates of week-eight assistant retention in the telemetry we hold.

The enclave performance tax at production batch shape is about 12% for a 70B model on AU-region hardware.

Off-the-shelf offensive agents find real misconfigurations in a competent team's staging environment without bespoke prompting.

A regional realtime endpoint removes roughly half the turn latency from Sydney.

Without a human-designed verifier role, single-agent errors reach most of a collective within three rounds.

Enterprise inference clusters run well below the cost crossover.

Open-weight parity is real on three task classes and absent on agentic loops.
Do not report an eval number produced by an LLM judge unless the judge has a published agreement rate against a human panel on the same task. Below 0.8 Cohen's kappa the judge is not a…
Under ten sessions of history: none, just the context window. Over ten: the lab's episodic store with summarised recall via the memory adapter. Not a vendor memory product yet.

The lab director placed the cluster on the board without knowing a product engineer had placed it in January. Two independent placements is the operating model's convergence signal and the…

Most dark-run failures pass tests and violate unstated constraints; a spec-audit agent catches about half.

A release canary on own-data task evals catches material regressions before client pipelines do.

De-duplicated reach is an order of magnitude below opens.
Agentic development is real, the throughput gain is real, and the bottleneck has moved from writing code to specifying and verifying it. The firms that win will be the ones that…
Tokens substitute for 20–30% of junior analyst hours; senior hours are unchanged.

An agent-assisted scan can produce a usable crypto-agility inventory in days on a mid-sized estate.
Do not give an agent a long-lived service account. Every tool call runs under authority delegated from a named person, scoped to the task, and expiring with it.

Agent collectives develop useful division of labour without human-designed roles.
Energy becomes the dominant component of inference cost between 2028 and 2030.
Weight-level continual learning without forgetting is demonstrated at small scale in a narrow domain.

A spec-first agentic loop cuts analytics delivery time by more than half when the artifact is versioned.

The lab's lead time is positive and small, and the negatives are real.
Showing the top 60 of 439. Narrow the configuration rather than scrolling.