How I AI · Tuesday, June 30, 2026
Tired of subjective 'vibe checks,' the podcast host is introducing the 'How I AI bench,' a new set of AI and Claude-graded benchmarks designed for repeatable testing of AI models. The benchmark will assess models on tasks like writing PRDs, solving bugs, and one-shotting designs.
“What I want to start developing is a set of benchmarks we can regularly test these new models against that you'll care about.”
“So today, I'm going to be introducing the how I AI bench, a set of AI and Claude graded benchmarks that are going to tell us if this model and any model is good at writing PRDs, solving bugs, and one-shotting designs.”