← Front page

How I AI · Tuesday, June 30, 2026

Host Develops 'How I AI' Benchmark Using Claude Code

The podcast host details the creation of the 'How I AI' benchmark, built using Claude code, to provide a more repeatable and less subjective evaluation of AI models. This process involved using Claude to brainstorm benchmark design principles and tasks, focusing on builder-relevant use cases.

companyClaude

The tape

2 quotes
And I asked just a very simple question. Based on our work together, can you help me brainstorm a How I AI benchmark and Eval set we can test every time a new model comes out to consistently score different tasks that would be relevant to our podcast audience?
So, I'm going to show you how I built and will build the How I AI benchmark. And on a blind test, how these models did across a couple of use cases.
Heard on How I AI — “Sonnet 5 review: I ran 64 generations to find out if it's worth it, published Tuesday, June 30, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.02
Host Develops 'How I AI' Benchmark Using Claude Code — Heardvine