The a16z Show · Thursday, August 6, 2026
Matt Bornstein notes that the VLLM project, an open-source inference engine, is now running on an impressive half a million GPUs. Simon Mo further elaborates that VLLM acts as the inference engine, akin to databases or operating systems, turning GPUs into endpoints for AI.
“So VLLM actually has its origins kind of back in 2022, pre chat GPT. And your team set out to make a slow open source demo faster and instead just found this pile of unsolved problems.”
“VLM is an inference engine. That means its job is to turn available GPUs into a running endpoint for intelligence. So that means it is kind of like databases and operating system and other critical software to power this economy or power of AI that everybody really uses today to ensure they can have cost effectiveness, efficiency, reliability, and also always stay on the frontier.”