← Front page

Machine Learning Street Talk · Thursday, October 1, 2026

End-to-End Models Necessary for Real Conversation Turn-Taking

Shawn Wen argues that for business use cases requiring natural conversation and turn-taking, companies must train their own models, as frontier labs aren't optimizing for these specific needs. He highlights that traditional cascaded systems, with separate speech recognizers, language models, and turn-taking modules, struggle to work together effectively, necessitating an end-to-end approach.

The tape

3 quotes
“we probably have to really train our own model for these use cases, because I think the frontier labs are not optimizing for what we care about. It's a real conversation turn-taking for the business use cases.”
Shawn Wen
“Because the traditional Cascade system, you have a speech recognizer, you have a larger language model, and you also have the turn-taking on top. And that is actually making the end-of-the-sentence detection. These three models don't really work with each other very well.”
Shawn Wen
“And you have the tween-lapid parameters. can have another detector to say oh this is an old lady so i have to tweet a parameter here and there but it's all very crunky and then reality is that it doesn't really work like that and it doesn't really work very well as well so therefore, it's almost necessary to collapse all these pipelines together so you have a large language model directly perceive the old audio and the”
Shawn Wen
Heard on Machine Learning Street Talk — “How a Voice Agent Learns the Rhythm of Conversation — Shawn Wen”, published Thursday, October 1, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via deepinfra · $0.01
End-to-End Models Necessary for Real Conversation Turn-Taking — Heardvine