Machine Learning Street Talk · Thursday, October 1, 2026
Evaluating voice agents is complex due to multiple factors like latency, accuracy, and information retrieval, which standard ASR or LM benchmarks don't capture. Shawn Wen notes that releasing public voice datasets is also difficult due to privacy concerns. His company builds benchmarks on real conversation data, including response time and understanding, and plans to release it to the community.
“The voice is really difficult to evaluate because it's a voice agent has the latency considerations. You have to reply fast enough.”
“And also, another problem is that a lot of these voice agent benchmark on the market are synthetic to a degree. Because voice is a lot more private, releasing those voice that are set for evaluation and then make it public, it's non-trivial as well.”
“Our benchmark is built on the real data set we have in the conversations with our consumers. So we build it in a way so it measures a lot of different important access.”