The a16z Show · Wednesday, September 9, 2026
Ben Horowitz and Rayan Krishnan discussed the difficulty in evaluating AI, comparing it to the challenge of evaluating human intelligence without agreed-upon frameworks. They noted that AI models are adept at 'hacking' benchmarks, underscoring the need to make evaluation criteria more explicit and robust.
“How do you deal with the kind of issue that it's a little bit of an AI complete problem in that we still aren't really good at evaluating humans? Or we haven't agreed on it.”
“And then of course, models are really good at hacking the benchmarks. So the fourth proved.”
“What is really the distinction between an associate and a partner at a law firm? And there is no clear test or e-val for that.”