← Front page

The a16z Show · Wednesday, September 9, 2026

Meta's Llama 4 Performance Discrepancy Highlighted by Val's Benchmarks

Rayan Krishnan of Val's pointed out a significant difference in Meta's Llama 4 performance, noting that while it excelled on public benchmarks, Val's private benchmarks showed underperformance. This suggests a potential issue with how AI model capabilities are publicly measured.

companyMeta

The tape

2 quotes
So there's a huge disconnect between what was self-reported based on these open benchmarks, and then what we were actually finding with our higher quality, higher signal benchmarks.
When Meta released Llama 4, on our held out private benchmarks, the model is actually underperforming. But on all of the major public benchmarks, it was showing incredible capabilities.
Heard on The a16z Show — “Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan, published Wednesday, September 9, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.05