Not yet: realtime voice for AU contact centres above tier-1 triage
Do not commit a client to realtime voice agents beyond tier-1 triage and routing. The latency floor from Sydney is not there, and the cost of a public failure in a regulated contact centre is high.
Tier is not strength
Tier says how much we know. Strength says how hard we are telling you to act. Scored independently.
Machine-readable target
{
"taskType": "realtime-voice-agent",
"configKey": "voice.deploy.scope"
}Nothing consumes it yet. Day two: findings ship as defaults into the gateway, routing config and skill library.
Body
x-voice-latency measured end-to-end latency for three speech-to-speech stacks from an AU telco contact centre through Sydney-region endpoints. Median turn latency was 640–910 ms; the p95 on the best stack was 1.4 s. Human conversational tolerance in our own panel broke at around 700 ms, and callers began talking over the agent above that (c-voice-and-vision-1). At tier-1 — 'which department do you need' — that is fine. For anything involving a decision it is not (c-voice-and-vision-2).
The strength is strong on evidence that is only moderate, and that is deliberate. One telco, one measurement campaign, one panel. But this is a 'not yet', and the asymmetry runs one way: being wrong by waiting costs a quarter of lead time; being wrong by shipping costs a client a public failure in an APRA-regulated channel. When the cost of being wrong is this lopsided, the recommendation should be firm even when the evidence is not.
g-voice-ivr-replacement is the failure this guards against. A realtime agent replacing tier-1 IVR was piloted and pulled after complaint volume rose 22% in the first fortnight; the drivers were latency and interruption handling, not comprehension. Comprehension was fine, which is why the pitch keeps coming back (c-voice-and-vision-4).
What would change this: a Sydney-region speech-to-speech endpoint with sub-400 ms p95, or an on-prem stack that hits it. The graveyard entry is marked resurrectable on exactly that trigger, and Argus watches for it. Re-review in ninety days regardless.
What it rests on
From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.
Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.
Vision on structured field imagery is production-ready now; the reliability gate is specific to voice.
Field
Voice and VisionExperiment · abandoned
Voice agent latency floor for AU telcoGraveyard · abandoned
Realtime voice agents replace tier-1 IVR “Hung up before the answer.”