Machine Learning Street Talk · Thursday, October 1, 2026
Shawn Wen details their company's approach to voice agents, using an audio-native large language model that directly processes audio input and outputs text. This architecture allows for native understanding of audio streams and integrates turn-taking signals as the first output token, streamlining the response generation process.
“Yeah, so the high level architecture is that this is basically So, audio native model. So the base model itself is a large language model, so it actually outputs text.”
“So we are not training models to do end-to-end speech yet. The reason is because enterprise is one level of control, so text is a lot easier to actually apply guardrail on top, while speech is a lot fiddly to deal with. And so we train the model to output text,”
“The way that we are structuring the model is that the model is doing multiple things at once. So first of all, the model would be based on the streaming in audio. predicting a first token. The first token is basically the turn-taking signal.”