← Front page

Machine Learning Street Talk · Thursday, October 1, 2026

Audio-Native LLMs for Voice Agents

Shawn Wen details their company's approach to voice agents, using an audio-native large language model that directly processes audio input and outputs text. This architecture allows for native understanding of audio streams and integrates turn-taking signals as the first output token, streamlining the response generation process.

The tape

3 quotes
“Yeah, so the high level architecture is that this is basically So, audio native model. So the base model itself is a large language model, so it actually outputs text.”
Shawn Wen
“So we are not training models to do end-to-end speech yet. The reason is because enterprise is one level of control, so text is a lot easier to actually apply guardrail on top, while speech is a lot fiddly to deal with. And so we train the model to output text,”
Shawn Wen
“The way that we are structuring the model is that the model is doing multiple things at once. So first of all, the model would be based on the streaming in audio. predicting a first token. The first token is basically the turn-taking signal.”
Shawn Wen
Heard on Machine Learning Street Talk — “How a Voice Agent Learns the Rhythm of Conversation — Shawn Wen”, published Thursday, October 1, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via deepinfra · $0.01