← Front page

How I AI · Tuesday, June 30, 2026

Gemini 3 Pro and Sonnet 5 Tie for Top Spot in Blind Test

In a surprising turn, Gemini 3 Pro and the new Claude Sonnet 5 tied for the highest score in the blind benchmark evaluation, with GPT 4.5 also performing strongly. This outcome was unexpected, as the host noted Opus 48 and Sonnet 46 scored lower than anticipated.

companyGooglecompanyAnthropic

The tape

2 quotes
The model that I forgot we were testing scored the best. So Gemini 3 Pro up here at the top of the leaderboard tied with the brand new drop Sonnet 5. GPT 4.5, my personal favorite also in this three-horse race at the top of the leaderboard.
And then poor Opus to vibes are off at the bottom of the list, as well as Sonnet 46. With lots of red flags on Sonnet 46.
Heard on How I AI — “Sonnet 5 review: I ran 64 generations to find out if it's worth it, published Tuesday, June 30, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.02
Gemini 3 Pro and Sonnet 5 Tie for Top Spot in Blind Test — Heardvine