The Twenty Minute VC (20VC) · Saturday, September 26, 2026
Pagg Agrawal explains that making AI models more efficient (smaller) while maintaining performance can lead to a loss of parametric memory. He contrasts this with the recall of 'head facts' that models can memorize, noting that finding patterns rather than memorizing is key for more complex information, especially for individuals not in the public eye.
“So today by and large, I would say models have good recall from parametric memory on like, they call them head facts. It's like for someone famous, everything about them, Wikipedia, it can memorize, so it'll tell you who the president was in a certain year, right? Because the model can memorize those things.”
“The model couldn't tell you what year I graduated from college. Maybe if it can, but it can't tell you for somebody who works at parallel. Because again, both of them actually trying to find patterns rather than memorize them.”
“And then further, as you make models efficient, which is you make them smaller and smaller while keeping the performance, you lose more of the parametric memory. While trying to keep the reasoning.”