cavendish
TestedContestedgate · ReliabilityNow · 0–12 months×4 sightings

Voice and Vision

Realtime voice agents in Australian contact centres are gated by reliability, not capability: a round-trip floor from Sydney to US-hosted realtime models, accent-driven recognition errors, barge-in failures and hallucinated confirmations mean they hold for tier-1 triage and fail above it; vision in field operations and retail is further along because latency is not the product.

Experiment run, measured result. The only tier that becomes a recommendation.

Join with…

Confidence

63%human-committed

Expiry

28duntil review · 1 Oct 2026

Lead time

4moopened after mainstream — recorded honestly

Ownership

TOTom Okaforfortnightly cadence

Where it is

The field is contested because the two halves are moving at different speeds. Vision on structured imagery — meter photos, shelf gaps, claim forms — is at production accuracy on frontier models and on a distilled small model, and the retail sector owner is already using it. Voice is not. Our latency run for an AU telco measured a p95 turn latency of 2.6 seconds via US-hosted realtime endpoints, well past the point at which callers talk over the agent; a Sydney-region endpoint launched in July halves that but does not fix barge-in, which failed on 18% of interruptions, or hallucinated confirmations, which occurred in 4 of 50 test calls. Vendors claim 70% containment; the analyst line says 40% of AU volume by 2027. Both may be true for tier-1 queues and neither is true for anything transactional. The recommendation is 'not yet' above triage, and the graveyard holds the IVR-replacement thesis.

Why a Quantium decision hinges on it

Telco and insurance clients are being pitched voice agents by every contact-centre vendor, and the pitch numbers are containment rates measured on US English in US data centres. Quantium's edge is a measured AU floor: the latency, the accent error rate and the confirmation-hallucination rate on real queues. That number is the difference between a triage deployment that works and a payments deployment that generates complaints to the TIO. Vision is the opposite story — a quiet production capability in field ops and retail that nobody is asking the lab about because it already works.

Field attributes

StateContested
GateReliability · possible, not yet dependable enough
OriginQuestion
Measurablepartial
Audience · TLPpractice
Horizonnow
Opened12 Mar 2026
Mainstream1 Nov 2025
Last validated19 Aug 2026
Sightings4

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01AU voice latency floor: p50 1.4 s and p95 2.6 s turn latency via US-hosted realtime endpoints; 0.78 s p50 via the Sydney-region endpoint launched in July (x-voice-latency).
  • 02Barge-in failed on 18% of caller interruptions across three stacks; abandonment correlated with barge-in failure more strongly than with recognition errors.
  • 03Hallucinated confirmations — the agent reading back a reference number, amount or date that did not exist — in 4 of 50 scripted test calls. Rules out unsupervised transactional use.
  • 04Vision on structured field imagery (meter dials, shelf gaps, handwritten forms) at 96% task accuracy on frontier models; a distilled 3B model reached 94% on shelf gaps on in-store hardware.
What is hype
  • 01'70% containment' from contact-centre vendors. Measured on tier-1 US-English queues with generous definitions of contained; nobody publishes the AU number.
  • 02'Voice AI handles 40% of AU volume by 2027.' Possible for triage; the transactional share of volume does not move on current reliability.
  • 03'Human-level' ASR. True for a US newsreader; error rates on Australian-accented and non-native English were 2.1× the US baseline across the stacks we tested.
What would have to be true
  • 01A regional realtime endpoint with barge-in handled at the edge, holding p95 under 1.2 s from Sydney on a real queue rather than a test rig.
  • 02A confirmation-grounding pattern — the agent may only read back values it retrieved, never generated — that survives a full call without a latency cost.
  • 03AU-accent recognition error within 1.2× of the US baseline; currently 2.1×.
What we would do
  • 01Keep r-voice-not-yet as the position for anything above tier-1 triage; revisit in Q4 with the regional endpoint on a live queue.
  • 02Ship vision on structured imagery as a default pattern for field ops and store audit; it does not need the lab any more.
  • 03Run the latency bench again against the regional endpoint with barge-in at the edge, and add the confirmation-grounding pattern as the intervention.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
2.6 s
p95 turn latency (US-hosted)
Finding·band 1Tested

Voice latency floor for AU telco: p95 2.6 s via US endpoints, barge-in fails 18%

Three realtime stacks driven from a Sydney telephony rig with scripted callers and timed interruptions. US-hosted endpoints gave p50 1.4 s and p95 2.6 s turn latency; the regional endpoint gave p50 0.78 s. Barge-in failed on 18% of interruptions and predicted abandonment better than recognition errors did.

extracted claimFrom Sydney, US-hosted realtime voice exceeds the talk-over threshold at p95; barge-in failure, not recognition, drives abandonment.
Lab · x-voice-latency · Tom Okafor19 Aug 2026
detector · bleeding edge
0.78 s
p50 turn latency (regional)
Release·band 1Tested

Frontier lab launches Sydney-region realtime voice endpoint

Realtime speech-to-speech API available in an AU region with data residency. Arrived mid-experiment; we added it as a fourth arm. Halves the p50 and brings p95 under 1.5 s on the rig. Does not change barge-in behaviour.

extracted claimA regional realtime endpoint removes roughly half the turn latency from Sydney.
OpenAI15 Jul 2026
detector · bleeding edge 3
4 / 50
hallucinated confirmations
Finding·band 1Tried

Logged from Claude Code: voice agent read back a booking reference that did not exist in 4 of 50 calls

During harness setup the agent confidently confirmed reference numbers and amounts it had not retrieved. Grounding the read-back to tool output removed it in a follow-up run of twenty calls. Logged as tried; became a bench arm.

MCP · log_finding · Tom Okafor24 Jun 2026
TOdropped
Post·band 2Signal

'The hardest part of voice is knowing when the human stopped talking'

Practitioner post on endpointing and barge-in: the model is rarely the problem; the turn-detection layer is. Matches our bench, where barge-in failure predicted abandonment. Widely shared in the infra community.

Engineering blog · A voice-infrastructure engineer2 Jun 2026
OGdropped 3
70%
claimed containment
Announcement·band 3Signal

Contact-centre platform vendor claims 70% containment with voice agents

Headline containment figure from US-English tier-1 deployments; 'contained' includes calls that ended after the agent read the account balance. No AU customer named. Carried as the claim our bench contradicts.

Vendor press release6 May 2026
detector · demand 4
Client question·band 3Signal

'Can the voice bot take a payment over the phone without a human on the line?'

Asked by a contact-centre GM after a vendor demo. The delivery lead did not have an answer and logged it. Became the transactional-intent arm of the bench and the reason the recommendation draws its line at triage.

Engel · telco engagement30 Apr 2026
MLdropped 3
96%
meter-read accuracy
Benchmark·band 2Signal

Field-imagery benchmark: meter dials, forms and asset photos

Public multimodal benchmark of structured field imagery. Frontier models above 96% task accuracy on meter reads and handwritten forms; the energy sector's use case is solved at the model level and blocked on integration only.

huggingface.co9 Apr 2026
detector · early adoption 2
2.1×
WER vs US baseline
Paper·band 1Signal

Accent-conditioned error rates in realtime speech models: an Australian and non-native English study

Measures word error rate across three realtime stacks on Australian-accented, Indian-English and Mandarin-accented English against a US baseline. AU-accented error rate was 2.1× baseline; the gap narrowed but did not close with vendor accent options.

arxiv.org · Whitmore, Nguyen et al.20 Mar 2026
detector · bleeding edge 2
6.1k
stars
Repository·band 2Tried

shelfsight — open vision-language model for shelf-gap and planogram detection

Small open VLM fine-tuned on retail shelf imagery. The retail sector owner ran it on store photos during a Woolworths pilot and it matched the frontier model on gap detection. Basis for the edge classifier work in the SLM field.

github.com11 Feb 2026
DSdropped 2
40%
forecast share by 2027
Analyst·band 3Signal

'Voice AI will handle 40% of Australian contact-centre volume by 2027'

Demand-band forecast that appears in every vendor deck since. Plausible for triage volume; no reliability caveat. Included as the mainstream framing.

Analyst research note28 Jan 2026
detector · demand 5
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.

Testedc-voice-and-vision-1dalton-0.419 Aug 2026Lab · x-voice-latency, OpenAI
86%

Vision on structured field imagery is production-ready now; the reliability gate is specific to voice.

Assessedc-voice-and-vision-4dalton-0.314 May 2026huggingface.co, github.com
79%

Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.

Testedc-voice-and-vision-2dalton-0.419 Aug 2026Lab · x-voice-latency, Engineering blog
74%

Voice agents hallucinate confirmations at a rate (5–8% of calls) that rules out unsupervised transactional use.

Testedc-voice-and-vision-3dalton-0.42 Jul 2026MCP · log_finding, Lab · x-voice-latency
70%

Vendor-reported 70% containment generalises to Australian queues once the accent model is tuned.

Assessedc-voice-and-vision-5dalton-0.42 Jul 2026Vendor press release, arxiv.org
25%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 42% → 63%.

runs compare claim sets, never prose
What we said · run 4

Latency bench concluded. US-hosted floor is unusable; regional endpoint halves it; barge-in is the abandonment driver. Recommendation published; IVR-replacement thesis to the graveyard. Field stays contested because the floor is moving.

63%
Changed since run 3
  • From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.
  • Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.
  • c-voice-and-vision-3 ↑ 0.62 → 0.70
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · ML
high

Contact-centre cost is the largest AI line item in every telco and insurance pipeline we have.

Timeline

committed · TO
0–18mo

Triage is deployable now; transactional depends on grounding and regional endpoints, both inside the window.

TAM

agent-estimated
$1B–10B

Agent-estimated from AU contact-centre labour spend addressable by tier-1 automation. Uncommitted.

Demand

committed · ML
high

Every telco engagement this half has asked; two have live vendor pilots we are being asked to assess.

Cost

committed · TO
medium

The bench needs a telephony rig and scripted callers; two engineers, three weeks per cycle.

Cost of being wrong

committed · LF
high

A hallucinated payment confirmation on a regulated queue is a complaint to the TIO or AFCA, not a bug.

Workforce readiness

agent-estimated
low

Nobody in delivery has built a barge-in-safe voice loop; the vendor stacks hide it until it fails. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Telco
relevant

The largest queues in the country and the most vendor pressure. The AU latency and accent numbers are the assessment they are asking us for.

Mechanism · Tier-1 triage with a regional endpoint and human handoff on any transactional intent; measure barge-in and confirmation errors on the live queue.

ML committed by Marcus Leecommitted
Retail & FMCG
relevant

Vision, not voice. Shelf-gap and planogram compliance from store photos is already in a Woolworths pilot.

Mechanism · Distilled vision model on in-store hardware; frontier model for exception review.

DS committed by Dev Sharmacommitted
Energy & Utilities
relevant

Meter-reading and asset-inspection photos are structured imagery where frontier accuracy is already sufficient.

Mechanism · Batch vision over field-app uploads; no latency constraint, so no gate. Agent draft.

Agent draft · awaiting a sector owneragent-estimated
Insurance
watch

Claims intake by voice is attractive and the confirmation-hallucination rate makes it dangerous; vision on claim photos is fine.

Mechanism · Vision now; voice only after the grounding pattern is measured. Agent draft.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against: we measured a floor that is already moving under us. The regional endpoint arrived mid-experiment, barge-in is a stack-engineering problem the vendors will fix within a release or two, and confirmation grounding is a prompt-and-tool pattern, not a research problem. A 'not yet' that expires in a quarter is a 'yes' with extra steps, and clients who wait on our advice will be a quarter behind the ones who did not.

  • The 2.6 s p95 was measured on US endpoints that no serious AU deployment would use after July. Our headline number is already historical.
  • Fifty scripted test calls is a small sample for a 4-in-50 hallucination rate; the confidence interval includes rates that are acceptable with a human check.
  • Contained tier-1 calls are most of the volume in most queues. 'Holds for triage' may be 80% of the business case, which makes 'not yet' the wrong headline.
  • The vision half is not contested by anyone and does not need a field; keeping it here inflates the field's confidence.
Stored permanently alongside the thesis. Sources are correlated; without an adversary, synthesis converges on consensus and calls it insight.thesis weakened

Source diversity

  • Speech and ML research25%
  • Voice infra practitioners20%
  • Vendor and analyst25%
  • Internal / Engel30%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsTOPRMLDSOGLF

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01What is the p95 on a live queue, not a rig, with the regional endpoint and barge-in at the edge?
  2. 02Does confirmation grounding survive a full call without adding a turn of latency?
  3. 03Should vision leave this field and go to practice now, given nobody contests it?