The Twenty Minute VC (20VC) · Saturday, September 19, 2026
Thomas Sohmers explained the fundamental differences between training and inference in AI hardware. Training is primarily compute-bound, requiring more FLOPS, while inference is heavily memory-bound due to the need to read parameters for each token generated.
“So training, I would say, from the underlying compute level, fundamentally, you know, is a compute-bound problem.”
“The big difference with inference, and so the deployment of those models, is the fact that for the actual math and the, the steps that you're doing, is about half of what you're doing during training in terms of that the steps, not shouldn't be thought of as like the actual compute involved.”
“But what it turns out to be is that that forward pass, that inference portion of it is heavily, heavily memory-bound due to the fact that basically for every single token that is generated, every little bit of output, that requires going through the weights, the parameters, you know, from a biological sense, think of the neurons, you have to read the values of that for every single token.”
“But when you're inferring, because that's actually generative, you don't know what the token is five words down the line.”