Machine Learning Street Talk · Monday, July 13, 2026
Alistair Pullen discusses the importance of both total and active parameters in LLM architecture, citing examples like Mistral's 675B model and Anthropic's Sonnet and Opus. He suggests that architectural decisions are heavily influenced by inference capabilities and practical deployment constraints.
“I believe the largest model that Mistral have made to date is the 675B Mistral 3 Large, right?”
“um That that model exists in the way that it does because it fits a use case that they have seen. um It probably fits a GPU deployment profile that they have seen in the enterprises that they're trying to sell to in France or Europe.”
“um That something like Sonnet is in the, I believe like 1.3 to 1.5 trillion total params, um MOE, and has probably 100 plus billion active, right?”
“um And an Opus is probably in the 1.5 to 1.8 range and probably has 150 to 180 billion active, depending on the detail you're talking about, whether it's FPA or FP4, whatever.”