How I AI · Tuesday, June 30, 2026
In a surprising turn, Gemini 3 Pro and the new Claude Sonnet 5 tied for the highest score in the blind benchmark evaluation, with GPT 4.5 also performing strongly. This outcome was unexpected, as the host noted Opus 48 and Sonnet 46 scored lower than anticipated.
“The model that I forgot we were testing scored the best. So Gemini 3 Pro up here at the top of the leaderboard tied with the brand new drop Sonnet 5. GPT 4.5, my personal favorite also in this three-horse race at the top of the leaderboard.”
“And then poor Opus to vibes are off at the bottom of the list, as well as Sonnet 46. With lots of red flags on Sonnet 46.”