Machine Learning Street Talk · Monday, September 14, 2026
Mistral AI has developed Voxler Chat, an audio-input, text-to-text LLM designed for general audio understanding. The model can perform tasks such as transcription, speaker segmentation, and summarization, and can answer questions about audio content, making it useful for analyzing meetings or podcasts.
“So we, uh, we started working on audio, uh, last year, and the very first model we released was Voxler Chat. It's an audio input, text to text, uh, LLM model.”
“Uh, the idea behind that is to have a general interface for audio understanding.”
“Uh, so the model can do transcription, uh, speaker segmentation, uh, summarization, or, uh, any questions like you can have a big audio document, uh, which is like an earnings call or a meeting recording.”