The Twenty Minute VC (20VC) · Saturday, September 26, 2026
Pagg Agrawal details how web search can be optimized to save compute for AI models. The process involves narrowing down trillions of documents to a few thousand tokens for the model's context window, with the amount of compute allocated to web search depending on the cost and efficiency of the AI model itself.
“So the problem is going from a trillion URLs with let's call it a few thousand tokens each, down to a thousand total tokens.”
“So you're essentially allocating compute to web search in order to save compute on the model. That's roughly what's going on here.”
“So if your Luna model is really cheap, you don't want to do too much compute in web search because it's okay to leaf a little bit more information into Luna's context because it's cheap.”