Voice and Vision
Realtime voice agents in Australian contact centres are gated by reliability, not capability: a round-trip floor from Sydney to US-hosted realtime models, accent-driven recognition errors, barge-in failures and hallucinated confirmations mean they hold for tier-1 triage and fail above it; vision in field operations and retail is further along because latency is not the product.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
63%human-committedExpiry
28duntil review · 1 Oct 2026Lead time
−4moopened after mainstream — recorded honestlyOwnership
TOTom Okaforfortnightly cadenceWhere it is
The field is contested because the two halves are moving at different speeds. Vision on structured imagery — meter photos, shelf gaps, claim forms — is at production accuracy on frontier models and on a distilled small model, and the retail sector owner is already using it. Voice is not. Our latency run for an AU telco measured a p95 turn latency of 2.6 seconds via US-hosted realtime endpoints, well past the point at which callers talk over the agent; a Sydney-region endpoint launched in July halves that but does not fix barge-in, which failed on 18% of interruptions, or hallucinated confirmations, which occurred in 4 of 50 test calls. Vendors claim 70% containment; the analyst line says 40% of AU volume by 2027. Both may be true for tier-1 queues and neither is true for anything transactional. The recommendation is 'not yet' above triage, and the graveyard holds the IVR-replacement thesis.
Why a Quantium decision hinges on it
Telco and insurance clients are being pitched voice agents by every contact-centre vendor, and the pitch numbers are containment rates measured on US English in US data centres. Quantium's edge is a measured AU floor: the latency, the accent error rate and the confirmation-hallucination rate on real queues. That number is the difference between a triage deployment that works and a payments deployment that generates complaints to the TIO. Vision is the opposite story — a quiet production capability in field ops and retail that nobody is asking the lab about because it already works.
Field attributes
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01AU voice latency floor: p50 1.4 s and p95 2.6 s turn latency via US-hosted realtime endpoints; 0.78 s p50 via the Sydney-region endpoint launched in July (x-voice-latency).
- 02Barge-in failed on 18% of caller interruptions across three stacks; abandonment correlated with barge-in failure more strongly than with recognition errors.
- 03Hallucinated confirmations — the agent reading back a reference number, amount or date that did not exist — in 4 of 50 scripted test calls. Rules out unsupervised transactional use.
- 04Vision on structured field imagery (meter dials, shelf gaps, handwritten forms) at 96% task accuracy on frontier models; a distilled 3B model reached 94% on shelf gaps on in-store hardware.
- 01'70% containment' from contact-centre vendors. Measured on tier-1 US-English queues with generous definitions of contained; nobody publishes the AU number.
- 02'Voice AI handles 40% of AU volume by 2027.' Possible for triage; the transactional share of volume does not move on current reliability.
- 03'Human-level' ASR. True for a US newsreader; error rates on Australian-accented and non-native English were 2.1× the US baseline across the stacks we tested.
- 01A regional realtime endpoint with barge-in handled at the edge, holding p95 under 1.2 s from Sydney on a real queue rather than a test rig.
- 02A confirmation-grounding pattern — the agent may only read back values it retrieved, never generated — that survives a full call without a latency cost.
- 03AU-accent recognition error within 1.2× of the US baseline; currently 2.1×.
- 01Keep r-voice-not-yet as the position for anything above tier-1 triage; revisit in Q4 with the regional endpoint on a live queue.
- 02Ship vision on structured imagery as a default pattern for field ops and store audit; it does not need the lab any more.
- 03Run the latency bench again against the regional endpoint with barge-in at the edge, and add the confirmation-grounding pattern as the intervention.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Voice latency floor for AU telco: p95 2.6 s via US endpoints, barge-in fails 18%
Three realtime stacks driven from a Sydney telephony rig with scripted callers and timed interruptions. US-hosted endpoints gave p50 1.4 s and p95 2.6 s turn latency; the regional endpoint gave p50 0.78 s. Barge-in failed on 18% of interruptions and predicted abandonment better than recognition errors did.
extracted claimFrom Sydney, US-hosted realtime voice exceeds the talk-over threshold at p95; barge-in failure, not recognition, drives abandonment.

Frontier lab launches Sydney-region realtime voice endpoint
Realtime speech-to-speech API available in an AU region with data residency. Arrived mid-experiment; we added it as a fourth arm. Halves the p50 and brings p95 under 1.5 s on the rig. Does not change barge-in behaviour.
extracted claimA regional realtime endpoint removes roughly half the turn latency from Sydney.

Logged from Claude Code: voice agent read back a booking reference that did not exist in 4 of 50 calls
During harness setup the agent confidently confirmed reference numbers and amounts it had not retrieved. Grounding the read-back to tool output removed it in a follow-up run of twenty calls. Logged as tried; became a bench arm.

'The hardest part of voice is knowing when the human stopped talking'
Practitioner post on endpointing and barge-in: the model is rarely the problem; the turn-detection layer is. Matches our bench, where barge-in failure predicted abandonment. Widely shared in the infra community.

Contact-centre platform vendor claims 70% containment with voice agents
Headline containment figure from US-English tier-1 deployments; 'contained' includes calls that ended after the agent read the account balance. No AU customer named. Carried as the claim our bench contradicts.

'Can the voice bot take a payment over the phone without a human on the line?'
Asked by a contact-centre GM after a vendor demo. The delivery lead did not have an answer and logged it. Became the transactional-intent arm of the bench and the reason the recommendation draws its line at triage.

Field-imagery benchmark: meter dials, forms and asset photos
Public multimodal benchmark of structured field imagery. Frontier models above 96% task accuracy on meter reads and handwritten forms; the energy sector's use case is solved at the model level and blocked on integration only.

Accent-conditioned error rates in realtime speech models: an Australian and non-native English study
Measures word error rate across three realtime stacks on Australian-accented, Indian-English and Mandarin-accented English against a US baseline. AU-accented error rate was 2.1× baseline; the gap narrowed but did not close with vendor accent options.

shelfsight — open vision-language model for shelf-gap and planogram detection
Small open VLM fine-tuned on retail shelf imagery. The retail sector owner ran it on store photos during a Woolworths pilot and it matched the frontier model on gap detection. Basis for the edge classifier work in the SLM field.

'Voice AI will handle 40% of Australian contact-centre volume by 2027'
Demand-band forecast that appears in every vendor deck since. Plausible for triage volume; no reliability caveat. Included as the mainstream framing.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.
Vision on structured field imagery is production-ready now; the reliability gate is specific to voice.
Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.
Voice agents hallucinate confirmations at a rate (5–8% of calls) that rules out unsupervised transactional use.
Vendor-reported 70% containment generalises to Australian queues once the accent model is tuned.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 42% → 63%.
Latency bench concluded. US-hosted floor is unusable; regional endpoint halves it; barge-in is the abandonment driver. Recommendation published; IVR-replacement thesis to the graveyard. Field stays contested because the floor is moving.
- From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.
- Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.
- c-voice-and-vision-3 ↑ 0.62 → 0.70
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · MLContact-centre cost is the largest AI line item in every telco and insurance pipeline we have.
Timeline
committed · TOTriage is deployable now; transactional depends on grounding and regional endpoints, both inside the window.
TAM
agent-estimatedAgent-estimated from AU contact-centre labour spend addressable by tier-1 automation. Uncommitted.
Demand
committed · MLEvery telco engagement this half has asked; two have live vendor pilots we are being asked to assess.
Cost
committed · TOThe bench needs a telephony rig and scripted callers; two engineers, three weeks per cycle.
Cost of being wrong
committed · LFA hallucinated payment confirmation on a regulated queue is a complaint to the TIO or AFCA, not a bug.
Workforce readiness
agent-estimatedNobody in delivery has built a barge-in-safe voice loop; the vendor stacks hide it until it fails. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
The largest queues in the country and the most vendor pressure. The AU latency and accent numbers are the assessment they are asking us for.
Mechanism · Tier-1 triage with a regional endpoint and human handoff on any transactional intent; measure barge-in and confirmation errors on the live queue.
Vision, not voice. Shelf-gap and planogram compliance from store photos is already in a Woolworths pilot.
Mechanism · Distilled vision model on in-store hardware; frontier model for exception review.
Meter-reading and asset-inspection photos are structured imagery where frontier accuracy is already sufficient.
Mechanism · Batch vision over field-app uploads; no latency constraint, so no gate. Agent draft.
Claims intake by voice is attractive and the confirmation-hallucination rate makes it dangerous; vision on claim photos is fine.
Mechanism · Vision now; voice only after the grounding pattern is measured. Agent draft.
Red team · the strongest case against
The strongest case against: we measured a floor that is already moving under us. The regional endpoint arrived mid-experiment, barge-in is a stack-engineering problem the vendors will fix within a release or two, and confirmation grounding is a prompt-and-tool pattern, not a research problem. A 'not yet' that expires in a quarter is a 'yes' with extra steps, and clients who wait on our advice will be a quarter behind the ones who did not.
- —The 2.6 s p95 was measured on US endpoints that no serious AU deployment would use after July. Our headline number is already historical.
- —Fifty scripted test calls is a small sample for a 4-in-50 hallucination rate; the confidence interval includes rates that are acceptable with a human check.
- —Contained tier-1 calls are most of the volume in most queues. 'Holds for triage' may be 80% of the business case, which makes 'not yet' the wrong headline.
- —The vision half is not contested by anyone and does not need a field; keeping it here inflates the field's confidence.
Source diversity
- Speech and ML research25%
- Voice infra practitioners20%
- Vendor and analyst25%
- Internal / Engel30%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
On-device recognition and vision are where the latency floor and the residency question both disappear.
Regional or on-prem realtime inference is the only path under the 1.2 s p95 target from Sydney.
Audio tokens are priced an order of magnitude above text; the cost ledger needs a voice column.
Voice is the interface most assistants will eventually need; the reliability lessons transfer.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- MLMarcus Lee · Delivery lead · Telco1 drop
- TOTom Okafor · Research engineer · agents1 drop
- DSDev Sharma · Sector owner · Retail1 drop
- OGOllie Grant · Product engineer1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Not yet: realtime voice for AU contact centres above tier-1 triage
strength strong · 18 citations · review 29 Oct 2026
Voice agent latency floor for AU telco
At least one realtime voice endpoint reachable from Sydney holds median turn latency under 800ms across 40 scripted tier-1 telco triage calls over a partner SIP trunk.
Realtime voice agents replace tier-1 IVR
“Hung up before the answer.” · lived 3 months
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01What is the p95 on a live queue, not a rig, with the regional endpoint and barge-in at the edge?
- 02Does confirmation grounding survive a full call without adding a turn of latency?
- 03Should vision leave this field and go to practice now, given nobody contests it?