← Front page

The Twenty Minute VC (20VC) · Saturday, September 26, 2026

Optimizing Compute for AI Models: The Role of Web Search

Pagg Agrawal details how web search can be optimized to save compute for AI models. The process involves narrowing down trillions of documents to a few thousand tokens for the model's context window, with the amount of compute allocated to web search depending on the cost and efficiency of the AI model itself.

The tape

3 quotes
“So the problem is going from a trillion URLs with let's call it a few thousand tokens each, down to a thousand total tokens.”
“So you're essentially allocating compute to web search in order to save compute on the model. That's roughly what's going on here.”
“So if your Luna model is really cheap, you don't want to do too much compute in web search because it's okay to leaf a little bit more information into Luna's context because it's cheap.”
Heard on The Twenty Minute VC (20VC) — “20VC: Five Predictions for a World of Agents | The Ads Business Model Will Die | Biggest Lessons from Working with Elon Musk at Twitter with Parag Agrawal, Parallel”, published Saturday, September 26, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.06
Optimizing Compute for AI Models: The Role of Web Search — Heardvine