Machine Learning Street Talk · Monday, September 14, 2026
Building on its audio research, Mistral AI is working towards creating foundational building blocks for audio agents, including real-time transcription and text-to-speech (TTS) models. The long-term goal is to develop an end-to-end speech-to-speech model.
“And then since then, our focus was to build foundational building blocks for audio agents. Uh, and we released a transcription model, uh, real time variant of that, and then a TTS model, uh, earlier this year, and we are continuing to work on improving those models.”
“Uh, that's one layer and the hope eventually is to build an end to end, uh, speech to speech model in this space.”