← Front page

Machine Learning Street Talk · Thursday, October 1, 2026

Leveraging Enterprise Data for Voice Model Training

Poly AI trains its voice models using data from enterprise contact centers, including banking, logistics, and retail clients. Shawn Wen explains they use 'lost data' – real customer interaction data – and redact personally identifiable information (PII) to train the model on conversational mechanics rather than specific user data, ensuring privacy while improving model robustness.

The tape

3 quotes
“Well, us, we don't have to collect them because our platform generates a lot of voice data. So because we have been deployed to enterprise contact centers, we have banking clients, logistic clients, restaurants, hotels, a lot of these retail, outbound sale use cases.”
Shawn Wen
“And lost data is exactly... the kind of data we wanted to solve the problem for, because we serve predominantly enterprise at the moment. So therefore, we kind of use leveraging a lot of the listed data.”
Shawn Wen
“But really, the data is actually just to let the model learn the mechanics of the conversation rather than the actual user data.”
Shawn Wen
Heard on Machine Learning Street Talk — “How a Voice Agent Learns the Rhythm of Conversation — Shawn Wen”, published Thursday, October 1, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via deepinfra · $0.01
Leveraging Enterprise Data for Voice Model Training — Heardvine