The a16z Show · Wednesday, September 9, 2026
Rayan Krishnan of Val's pointed out a significant difference in Meta's Llama 4 performance, noting that while it excelled on public benchmarks, Val's private benchmarks showed underperformance. This suggests a potential issue with how AI model capabilities are publicly measured.
“So there's a huge disconnect between what was self-reported based on these open benchmarks, and then what we were actually finding with our higher quality, higher signal benchmarks.”
“When Meta released Llama 4, on our held out private benchmarks, the model is actually underperforming. But on all of the major public benchmarks, it was showing incredible capabilities.”