Invest Like the Best with Patrick O'Shaughnessy · Tuesday, June 30, 2026
For the decode stage of AI inference, Etched emphasizes the critical role of memory bandwidth across the entire cluster, not just on a single chip. They have developed interconnects that provide higher bandwidth and lower latency, enabling more effective use of memory and improving time per token.
“For decode, it is all a memory game. More memory bandwidth, you can load the model faster, load the KV cache faster, and serve more tokens per second per user.”
“People often ask, how much memory bandwidth are you on your chip? You should be asking how much memory bandwidth is on your full-scale up cluster.”
“What we were able to do is add way, way more bandwidth and a much lower latency for chip-to-chip. To our interconnects, that allows us to then use the memory of other chips, much more effectively.”