Training Data · Wednesday, July 29, 2026
Rohan Anil reflects on his past belief that scaling reinforcement learning (RL) would be the direct path to AGI. Despite significant advancements and scaling of RL models, he notes that real-world tasks remained largely unsolved. Anil attributes this to a disconnect between benchmark performance and real-world applicability, suggesting that current models do not adequately learn from real-world distributions.
“I basically believed that scaling up reinforcement learning is a necessary stepping stone on a path to AGI”
“I saw us training model after model that model was getting better and better all the benchmarks scores were going up And what did we also solve all the real world tasks at that moment Unfortunately unfortunately not”
“Our training data didn't really replicate the real world use cases and despite us basically maximizing all the tasks if you see ask anyone training models Hey what is one of your main issues I don't have hard enough tasks I don't have what to train our model on yet”