Machine Learning Street Talk · Thursday, October 1, 2026
Shawn Wen highlights that latency in voice agents is crucial for user experience, but it's not just about raw speed. By collapsing ASR and LLM, models can better predict natural turn-taking based on audio frames and semantic information, reducing perceived latency. Additionally, hosting models and pre-caching processes are key to optimizing response times, allowing for more complex reasoning within tight budgets.
“Latency, first of all, the major thing is you need, to be honest, all these model processings, if you are using the right size of model, the model processing itself, latency is not really a problem.”
“So that is the major bit by collapsing the ASR and LM together. So then you have this more natural turn-taking built into the model. So the model is now making that prediction by itself.”
“A lot of the pre-cached processing, as we talked about earlier, is another thing to actually how to actually optimize it.”