cavendish
SignalCandidategate · ToolingNear · 1–3 years

Multiplayer AI Surfaces

The single-user chat window is a transitional form. The durable surface is a shared workspace where several people and one agent hold the same context, and the problems that decide whether it works — presence, turn-taking, attribution — are collaboration problems the AI industry has not yet had to solve.

Clustered only. No lab work behind it. Cannot be cited.

Join with…

Confidence

36%unresearched

Expiry

73duntil review · 15 Nov 2026

Lead time

not yet mainstream · opened 24 Jun 2026

Ownership

Unownedcandidate — a named human elects

Where it is

Candidate. Every AI product a delivery pod uses today is one-person-one-agent, and the pod works around it by pasting transcripts into Slack. The signals that put this on the board are three: two collaboration-tool vendors shipping an agent as a co-present participant in a document rather than a sidebar, a well-cited paper on turn-taking failures when more than one human addresses the same agent, and a product engineer's note that his pod had built exactly this by hand with a shared memory store and a bot. The unsolved parts are not model parts. Who the agent is answering, who is credited for what it produced, and what happens when two people give it contradictory instructions in the same minute are questions from twenty years of groupware research that have not been re-asked with an agent in the room. No lab work has been done. An experiment is proposed and awaiting election.

Why a Quantium decision hinges on it

Quantium delivers in pods, not as individuals. If the effective unit of AI use is the pod and the tooling assumes the individual, the firm is paying for a productivity gain that does not reach the unit that ships work. This is also where citizen-developer slop originates: five people each with their own agent and their own context produce five slightly different versions of the same artifact. A shared surface either fixes that or makes it worse faster, and the lab should know which before a vendor sells one to a client.

Field attributes

StateCandidate
GateTooling · possible and affordable, not yet operable
OriginSignal
Measurablepartial
Audience · TLPlab
Horizonnear
Opened24 Jun 2026
Mainstreamnot yet
Last validated12 Aug 2026
Sightings1

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01A product pod inside the firm ran a shared-context bot with a common memory store for eight weeks and stopped pasting transcripts into Slack; attributed, tried, one pod (s-multiplayer-ai-surfaces-finding-ollie).
  • 02Two collaboration vendors ship an agent as a named participant with a cursor in a shared document; presence is solved at the UI level in both.
What is hype
  • 01'AI teammate' marketing that means a bot with a profile picture. Presence in the UI is not shared context.
  • 02Claims that a shared surface removes the need for a human coordinator. In the one paper that measured it, coordination cost went up when the agent could be addressed by anyone.
What would have to be true
  • 01A turn-taking policy that survives two humans issuing contradictory instructions inside a minute, with the agent asking rather than picking one.
  • 02Attribution that a client and a partner would both accept: whose work is it when three people and an agent produced the deck?
  • 03A shared memory store with per-person visibility, or the surface leaks one person's context into another's the first time it is used on client data.
What we would do
  • 01Elect if a second pod reproduces the tried result; otherwise hold as a candidate and keep the vendor releases under watch.
  • 02If elected, run x-multiplayer-surface as a Type 3 with turn-taking failures as the primary measure, not satisfaction.
  • 03Decline to take a position for clients until then; the honest answer today is 'we have one pod's experience'.

Signals · 8 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
1 in 12
contradiction rate
Paper·band 1Signal

Who Is the Agent Talking To? Turn-Taking Failures in Multi-Party Human–Agent Sessions

Logs 400 multi-party sessions with a shared agent and codes the failures. Contradictory instructions inside a minute occur in about one session in twelve; the agent picked one silently in most cases. Coordination cost, measured as human messages about the agent, rose 18%.

extracted claimMulti-party sessions need an explicit turn-taking policy, and shared agents raise coordination cost rather than lowering it.
arxiv.org · Haddad, Brennan et al.16 Jul 2026
SKdropped 2
Drop·band 3Signal

Slack drop: 'we have three people prompting the same agent on the retail deck and it is a mess'

A delivery engineer's unprompted complaint about the current workaround. Not a sighting of the cluster — a sighting of the problem. Counted as demand.

Slack drop5 Aug 2026
SKdropped
310
stars
Repository·band 2Signal

roomctx — shared agent context with per-participant scopes

Early open-source attempt at per-participant visibility over a shared memory store. Two contributors, no releases. The only project we found trying to solve the right problem.

github.com1 Aug 2026
detector · early adoption
1
openings
Job posting·band 1Signal

Frontier lab hiring 'Multi-user Session Researcher'

One opening describing 'sessions with multiple human principals' and 'attribution of agent output'. Argus inference: a first-party multiplayer surface is under design. One opening; weak.

Foundation lab careers page28 Jul 2026
detector · bleeding edge
Post·band 2Signal

'The shared AI workspace told my client what my colleague thought of their budget'

First-person account of a shared agent surfacing an internal note during a client session. Anecdote, but the precise failure mode we would expect without per-person visibility.

Personal blog · A consultant at a mid-size advisory firm22 Jul 2026
LFdropped 2
Finding·band 1Tried

Logged from Claude Code: pod ran a shared-context bot for eight weeks and stopped pasting transcripts

A product pod wired a bot to a common memory store so that any member's session could see what the others had established. Transcript pasting into Slack stopped within a week. One pod, self-built, self-reported. This is the drop that put the cluster on the board.

extracted claimA shared memory store behind a bot removes the transcript-pasting workaround inside a pod.
MCP · log_finding · Ollie Grant19 Jun 2026
OGdropped
Release·band 1Signal

Two collaboration vendors ship an agent as a named participant with a cursor

Both add an agent that appears in the participant list and edits a shared document in real time. Neither gives the agent per-person context; it sees the whole document and everyone's messages. Naming event: 'AI participant' in both.

extracted claimShipped multiplayer surfaces solve presence and ignore per-person context.
Vendor changelogs10 Jun 2026
AWdropped 2
Talk·band 3Signal

'The team is the user' — product talk on multiplayer AI

Demand-band signal. Argues the individual chat window is a stopgap and shows a pod-level surface. Applause line, no data; the speaker conceded turn-taking was unsolved when asked.

Figma Config14 May 2026
detector · demand
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 3 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Per-person context visibility is a hard requirement before any shared surface touches client data; no shipped product has it.

Signalc-multiplayer-ai-surfaces-3dalton-0.430 Jul 2026Vendor changelogs, Personal blog
60%

A shared memory store is what makes a multiplayer surface useful; presence and cursors without shared context are cosmetic.

Triedc-multiplayer-ai-surfaces-2dalton-0.412 Aug 2026MCP · log_finding, Vendor changelogs
55%

When more than one human can address the same agent, contradictory-instruction failures occur often enough (roughly one in twelve multi-party sessions in the paper we rate) to need an explicit turn-taking policy.

Signalc-multiplayer-ai-surfaces-1dalton-0.412 Aug 2026arxiv.org
52%

Shared surfaces reduce coordination cost inside a pod.

Signalc-multiplayer-ai-surfaces-4dalton-0.430 Jul 2026arxiv.org, Figma Config
33%

Position history · the diff is the product

2 validation runs against a fixed brief. Confidence 32% → 36%.

runs compare claim sets, never prose
What we said · run 2

Turn-taking failures are frequent enough to need a policy. Shared memory, not presence, is where the value is. Coordination-cost evidence cuts against the thesis. Experiment proposed; election pending a second pod.

36%
Changed since run 1
  • When more than one human can address the same agent, contradictory-instruction failures occur often enough (roughly one in twelve multi-party sessions in the paper we rate) to need an explicit turn-taking policy.
  • A shared memory store is what makes a multiplayer surface useful; presence and cursors without shared context are cosmetic.
  • c-multiplayer-ai-surfaces-4 ↓ 0.40 → 0.33
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

agent-estimated
medium

Agent-estimated. Affects how the firm uses AI internally more than what it sells.

Timeline

agent-estimated
18mo–4yr

Agent-estimated from vendor release cadence; the groupware problems are slow.

TAM

agent-estimated
$1B–10B

Agent-estimated from collaboration-software spend. Uncommitted.

Cost of entry

committed · AW
low

One pod already built it by hand; a Type 3 costs a pair and three weeks.

Demand

committed · AW
low

No client has asked. Demand is internal and from one pod.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Cross-sector
relevant

The delivery pod is the unit of work in every sector; whatever changes the pod's tooling changes every engagement.

Mechanism · Shared-context surface for the pod, with the client invited into scoped sessions once per-person visibility exists.

Agent draft · awaiting a sector owneragent-estimated
Government
watch

Multi-party sessions with an agent produce a record whose ownership and retention status under the Archives Act is unclear.

Mechanism · Would need a records ruling before any public-sector pod used it.

AB committed by Aisha Bellocommitted
Banking
not-relevant

Client-facing use is blocked by per-person context leakage; internal use is a productivity question, not a banking one.

Mechanism · None proposed.

CD committed by Claire Duboiscommitted

Red team · the strongest case against

The strongest case against: groupware has been trying to make shared context work since the 1990s and the surviving products are the ones that kept context per person and made sharing explicit. Adding an agent does not change the human dynamics that killed shared workspaces before; it adds a participant that cannot read the room. The one pod that liked it was the pod that built it.

  • The tried result is from the pod that built the tool. Builder enthusiasm is the least transferable evidence there is.
  • The paper's coordination-cost finding points the other way: when anyone can address the agent, someone has to referee, and that is a new job, not a removed one.
  • Vendors will ship presence and cursors because they are easy and demo well; the hard problems (turn-taking, attribution, per-person context) may simply never be solved by the products people actually buy.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis in doubt

Source diversity

  • ML research25%
  • Product / design25%
  • Vendor15%
  • Internal35%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsOGSKAW

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Does a second pod, not the one that built it, get the same result?
  2. 02What turn-taking policy does a pod actually want — ask, queue, or last-writer-wins?
  3. 03Whose work is the artifact when three people and an agent made it, and would a client accept that answer?