How I AI · Tuesday, June 30, 2026
The 'How I AI' benchmark incorporates a multi-stage evaluation process, including a human-scored 'vibe check' via an HTML page and LLM-based scoring for objective metrics. The process involved blind testing five models, including Sonnet 5, Opus 48, and potentially GPT 5.5, Gemini 5.2, and GLM.
“What's really fun is the evals are not quite done running. So they are running in a sub agent right now for the final scores. So I will actually be surprised at the end of the episode about what I think of Sonnet 5 amongst all these other models.”
“So you can see here, I have a blind set of models A through E. I believe we tested Opus 48, 5, Sonnet 46, Sonnet 5, and maybe GLM. I'm not actually sure what the fifth one was. We'll see when we get the scores.”