How I AI · Tuesday, June 30, 2026
The podcast host details the creation of the 'How I AI' benchmark, built using Claude code, to provide a more repeatable and less subjective evaluation of AI models. This process involved using Claude to brainstorm benchmark design principles and tasks, focusing on builder-relevant use cases.
“And I asked just a very simple question. Based on our work together, can you help me brainstorm a How I AI benchmark and Eval set we can test every time a new model comes out to consistently score different tasks that would be relevant to our podcast audience?”
“So, I'm going to show you how I built and will build the How I AI benchmark. And on a blind test, how these models did across a couple of use cases.”