Machine Learning Street Talk · Wednesday, September 23, 2026
Frank Hutter highlights that tabular data is ubiquitous but notoriously difficult to work with, presenting challenges like missing values, categorical features, and feature engineering. He contrasts this with image data, where general models can be trained effectively.
“I think one thing that people might not appreciate is that tabular data is absolutely everywhere.”
“And speaking from bitter experience, it's a nightmare to work with tabular data.”
“So anyone who's built a machine learning model with tabular data, you've, I mean, obviously like Pandas and, you know, scikit-learn has got some stuff in there, but you have to deal with missing values, you know, what do you do with the categorical features, how do you do some feature engineering, how do you do transformations that reflect the semantics of the problem.”