Machine Learning Street Talk · Thursday, October 1, 2026
Shawn Wen discusses the 'cocktail party problem' in voice AI, where agents need to focus on a specific speaker amidst background noise and crosstalk. While traditional methods involved speaker diarization, end-to-end models with LLM integration can natively perceive and potentially adapt to or ignore background conversations based on prompts, making them more natural for handling complex audio environments.
“I remember for many, many years that there was this phenomenon called the cocktail party problem, which is that, you know, like when we were at a cocktail party and there's many people talking at the same time, we can focus our attention on one person and our brain can just filter out the other people.”
“But with a dialogue system, what you really need to do is that the agent focuses on the right kind of subject that it is talking to.”
“Because now you combine the LM with the speech recognition together, so the LM can natively perceive the speech recognition. there's a background like you know cross talk somehow i mean to be honest this lm is also promptable right so you can prompt it to say hey don't respond to the background cross talk, apart from the major speaker you are speaking to.”