Machine Learning Street Talk · Monday, September 14, 2026
Pavan Muddireddy observes a convergence in AI model architectures across different modalities, including audio, text, and vision. This trend is driven by the need for unified user interfaces where models can accept various inputs (text, vision, speech) and produce diverse outputs (text, speech, images).
“Yeah, uh, I mean, it's fascinating that architectures are so, uh, similar across modalities.”
“Uh, even in audio, uh, which is a modality I started working on, uh, since, uh, two years ago, at this point, the techniques are converging more and more to a unified approach.”
“Uh, it's a model that you would like to give input through text, uh, communication through vision, input, and, uh, talk to it, and hope it writes back or speaks back and in some cases generates images.”