cavendish
TriedContestedgate · ReliabilityNear · 1–3 years×2 sightings

Ambient Agents

Agents that act unprompted inside the tools people already use — Slack, email, calendar, the CRM — will reach more users than any chat surface, and will be switched off faster than any chat surface unless their false-action rate is near zero.

Someone ran it in their own harness. Artifact, no protocol. Decays fast.

Join with…

Confidence

44%human-committed

Expiry

22duntil review · 25 Sep 2026

Lead time

not yet mainstream · opened 27 Jan 2026

Ownership

TOTom Okaforfortnightly cadence

Where it is

Contested, and the disagreement is inside the lab. The pull is obvious: an agent that drafts the follow-up in the CRM after a call, moves a meeting when a flight changes, or replies to the routine email is the productivity story executives actually want. The push is our own graveyard. The Slack agent we built to act unprompted was switched off in nineteen days, not because it was wrong often but because when it was wrong it was wrong in public. Two people placed this cluster on the board independently — the lab director and a product engineer — and the convergence is why it stays open despite the tombstone. The field's evidence is one killed experiment, three vendor releases that ship 'ambient' as a feature, a paper on unprompted-action trust, and a steady drip of Engel questions. Reliability is the gate: the interesting number is not accuracy, it is the cost of a wrong unprompted action relative to the value of a right one, and that ratio is different in every tool.

Why a Quantium decision hinges on it

Every productivity engagement the firm sells eventually arrives at 'can it just do it without being asked'. The answer decides the engagement's shape and its risk. If ambient agents work in some tools and not others, a rule for which is a delivery asset; if they do not work anywhere yet, saying so saves a client from a vendor's demo. The share graph matters too: two independent placements is the strongest sightings signal in the near horizon and the operating model says that is election-grade evidence in itself.

Field attributes

StateContested
GateReliability · possible, not yet dependable enough
OriginSignal
Measurablepartial
Audience · TLPpractice
Horizonnear
Opened27 Jan 2026
Mainstreamnot yet
Last validated14 Aug 2026
Sightings2

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01An ambient Slack agent acting on an internal channel achieved 91% acceptable actions over nineteen days and was switched off anyway; the 9% were public and one was embarrassing (g-ambient-slack-agent).
  • 02Draft-only ambient behaviour — the agent writes, a human sends — retained 100% of its users over the same period in the same pod; the difference is the action, not the intelligence.
  • 03A product engineer's calendar agent that only proposes reschedules and never commits them has run for four months without being disabled (tried, one user).
What is hype
  • 01'Set it and forget it' agents in vendor decks. Every shipped ambient feature we inspected defaults to draft mode, which is an admission.
  • 02Accuracy as the metric. A 95%-accurate agent that acts unprompted in a customer-facing channel is a 5% public-failure machine.
  • 03Ambient agents as a replacement for assistants. The retention evidence from Effective Assistants says proactivity has a ceiling and unprompted action is above it in most tools.
What would have to be true
  • 01A per-tool false-action budget that a client would sign — a number, not a principle — and an agent that stays under it for a quarter.
  • 02Reversibility: every unprompted action needs an undo that is as cheap as the action, which email send and CRM writes do not have today.
  • 03A delegation model so that when the agent acts, it acts as a named person with scoped authority, not as an integration token.
What we would do
  • 01Hold the position that draft-only ambient behaviour is deliverable now and act-unprompted is not, and say so in every productivity engagement.
  • 02Propose a Type 3 that measures false-action cost per tool rather than accuracy; the graveyard entry's returned notes describe the design.
  • 03Resolve the contest inside the lab by the end of Q4: either the field narrows to reversible actions or it dissolves into Effective Assistants.

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
19
days to switch-off
Finding·band 1Tried

Ambient Slack agent killed at day nineteen: 91% acceptable, switched off for the 9%

Internal agent acted unprompted in an operations channel — answering, tagging, closing threads. 91% of actions rated acceptable in review; the pod switched it off after a wrong public action in week three. The graveyard entry carries the full account and the returned notes for the next proposer.

extracted claimIn a shared channel, a low rate of public wrong actions outweighs a high rate of right ones.
Lab · g-ambient-slack-agent · Tom Okafor16 Apr 2026
detector · bleeding edge
2
independent placements
Drop·band 3Signal

Whiteboard placement: 'ambient agents' — second independent placement

The lab director placed the cluster on the board without knowing a product engineer had placed it in January. Two independent placements is the operating model's convergence signal and the reason the field stays open despite the graveyard entry.

Whiteboard11 Aug 2026
AWOGdropped 2
4
months running
Finding·band 1Tried

Logged from Claude Code: calendar agent that only proposes has run four months without being disabled

Product engineer's personal ambient agent watches flight and meeting changes and proposes reschedules; it never commits. Four months, still on. One user, self-reported. The contrast with the Slack agent is the field's clearest datum.

MCP · log_finding · Ollie Grant8 Jul 2026
OGdropped
Post·band 2Signal

'My AI sent the email. I spent a week apologising.'

First-person account of an auto-send email agent replying to a customer with a draft meant for an internal thread. Anecdote; the failure mode is exactly the irreversibility problem.

Personal blog · A founder at a small SaaS company29 Jun 2026
?dropped 2
Release·band 1Signal

Three productivity suites ship 'ambient' agent features in one quarter — all default to draft

Email, calendar and CRM features that watch context and prepare actions. Every one defaults to draft-only with an opt-in for auto-send, and all run under an integration token. Naming event: 'ambient' in all three.

extracted claimVendors ship ambient agents as draft-only by default; act-unprompted is opt-in and unadvertised.
Vendor changelogs16 Jun 2026
SKdropped 2
Client question·band 3Signal

'If it sends the email, whose name is on it, and who gets fired?'

Asked by a head of customer operations. Nobody in the room had an answer. It became the delegation requirement in the position and the reason auth-broker is a blocking join.

Engel · telco engagement27 May 2026
MLdropped 2
Paper·band 1Signal

Asymmetric Trust: Why Users Abandon Proactive Agents After Rare Visible Errors

Lab study of unprompted-action agents. Trust recovers from private errors and does not recover from errors visible to colleagues; the effect is stronger than accuracy across the tested range.

extracted claimVisible errors by a proactive agent cause abandonment independent of accuracy.
arxiv.org · Petrova, Mensah et al.19 May 2026
AWLFdropped 3
Talk·band 3Signal

Keynote demo: agent closes a support ticket, refunds the customer, updates the CRM — unprompted

Demand-band signal. Flawless on stage. The feature shipped two months later in draft-only mode. The gap between the demo and the default is the field in miniature.

Vendor conference8 Apr 2026
detector · demand
Regulatory·band 3Signal

ACCC guidance on automated customer communications: the business is the sender

Clarifies that automated messages to customers are the business's representations regardless of how they were generated. An ambient agent's wrong email is a misleading-conduct question, not a software bug.

ACCC5 Mar 2026
detector · demand
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Unprompted action in a shared channel is switched off at a false-action rate far above zero even when overall accuracy exceeds 90%; public failures dominate the decision.

Triedc-ambient-agents-1dalton-0.414 Aug 2026Lab · g-ambient-slack-agent, arxiv.org
74%

Ambient agents acting under an integration token rather than delegated personal authority will fail audit in banking and telco regardless of accuracy.

Assessedc-ambient-agents-5dalton-0.417 Jul 2026Engel · telco engagement, ACCC
72%

Draft-only ambient behaviour retains users where act-unprompted does not, with the same underlying model and prompt.

Triedc-ambient-agents-2dalton-0.414 Aug 2026Lab · g-ambient-slack-agent, MCP · log_finding, Vendor changelogs
68%

Reversibility of the action, not accuracy of the decision, predicts whether an ambient agent survives; calendar proposals survive, email sends do not.

Triedc-ambient-agents-3dalton-0.417 Jul 2026MCP · log_finding, Personal blog
60%

Frontier-model accuracy is now high enough that unprompted action in customer-facing tools is safe.

Assessedc-ambient-agents-4dalton-0.320 Apr 2026Vendor changelogs, Vendor conference
18%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 50% → 44%.

runs compare claim sets, never prose
What we said · run 4

Second independent placement on the board. Draft-only retention confirmed in a second pod. The lab still disagrees on whether the field is distinct from Effective Assistants; decision due Q4.

44%
Changed since run 3
  • Draft-only ambient behaviour retains users where act-unprompted does not, with the same underlying model and prompt.
  • c-ambient-agents-1 ↑ 0.66 → 0.74
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

Reaches every user of every tool; that is the whole attraction.

Timeline

committed · TO
18mo–4yr

Draft-only is now; act-unprompted in consequential tools is years off on the reversibility evidence.

Cost of being wrong

committed · LF
high

A wrong unprompted email to a customer is a conduct event. We killed our own for less.

Demand

agent-estimated
high

Agent-estimated from Engel: 'can it just do it' appears in most productivity engagements. Uncommitted.

TAM

agent-estimated
>$10B

Agent-estimated from productivity-software spend. Uncommitted.

Cost

committed · TO
low

The graveyard experiment cost a pair for two weeks; a false-action-cost study costs the same.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Telco
relevant

Field-service scheduling is the reversible, high-volume case where ambient rescheduling proposals could pay immediately.

Mechanism · Agent proposes reschedules in the dispatch tool; a dispatcher commits. Never touches the customer channel directly.

ML committed by Marcus Leecommitted
Retail & FMCG
relevant

Supplier follow-ups and range-review actions in the merchandising CRM are drafted today by hand and mostly routine.

Mechanism · Draft-only in the CRM; send remains with the category manager.

Agent draft · awaiting a sector owneragent-estimated
Banking
watch

Any unprompted customer contact is a conduct and CPS 230 question; draft-only for internal channels is the ceiling until the delegation model exists.

Mechanism · Internal Slack and case-note drafting only; no customer-facing action.

CD committed by Claire Duboiscommitted

Red team · the strongest case against

The strongest case against: the field is defined by the failure it is trying to avoid. Draft-only ambient behaviour is not an ambient agent, it is a well-integrated assistant, and Effective Assistants already covers it. The remaining territory — acting unprompted in consequential tools — has one experiment in the graveyard and no evidence anywhere that a false-action rate near zero is achievable with current models. Two people placing it on the board reflects how appealing the idea is, not how likely it is.

  • Remove draft-only from the field and nothing survives except a killed experiment and vendor features that default to draft-only. The field may be empty.
  • The false-action budget clients would need is one nobody has written down, and the one client asked said 'zero'.
  • Every vendor 'ambient' release we opened ships the agent under an integration token, which fails the delegation requirement before accuracy is even measured.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis in doubt

Source diversity

  • ML / HCI research20%
  • Vendor20%
  • Practitioner15%
  • Regulator10%
  • Internal / Engel35%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsAWOGTOLFSK

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01What false-action rate would a client actually sign, per tool, and is it achievable?
  2. 02Is draft-only ambient behaviour a field, or is it Effective Assistants with a different name?
  3. 03Can an undo be made as cheap as the action for email and CRM writes, or is irreversibility structural?