← Front page

How I AI · Tuesday, June 30, 2026

The 'How I AI' Benchmark Includes Human Vibe Check and LLM Scoring

The 'How I AI' benchmark incorporates a multi-stage evaluation process, including a human-scored 'vibe check' via an HTML page and LLM-based scoring for objective metrics. The process involved blind testing five models, including Sonnet 5, Opus 48, and potentially GPT 5.5, Gemini 5.2, and GLM.

The tape

2 quotes
What's really fun is the evals are not quite done running. So they are running in a sub agent right now for the final scores. So I will actually be surprised at the end of the episode about what I think of Sonnet 5 amongst all these other models.
So you can see here, I have a blind set of models A through E. I believe we tested Opus 48, 5, Sonnet 46, Sonnet 5, and maybe GLM. I'm not actually sure what the fifth one was. We'll see when we get the scores.
Heard on How I AI — “Sonnet 5 review: I ran 64 generations to find out if it's worth it, published Tuesday, June 30, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.02
The 'How I AI' Benchmark Includes Human Vibe Check and LLM Scoring — Heardvine