← Front page

Invest Like the Best with Patrick O'Shaughnessy · Tuesday, June 30, 2026

Memory Bandwidth Focus for AI Decode Performance

For the decode stage of AI inference, Etched emphasizes the critical role of memory bandwidth across the entire cluster, not just on a single chip. They have developed interconnects that provide higher bandwidth and lower latency, enabling more effective use of memory and improving time per token.

personRob Locken

The tape

3 quotes
For decode, it is all a memory game. More memory bandwidth, you can load the model faster, load the KV cache faster, and serve more tokens per second per user.
Rob Locken
People often ask, how much memory bandwidth are you on your chip? You should be asking how much memory bandwidth is on your full-scale up cluster.
Rob Locken
What we were able to do is add way, way more bandwidth and a much lower latency for chip-to-chip. To our interconnects, that allows us to then use the memory of other chips, much more effectively.
Rob Locken
Heard on Invest Like the Best with Patrick O'Shaughnessy — “Etched - Building AI Hardware to Make Inference Faster and Cheaper - [Invest Like the Best, EP.480], published Tuesday, June 30, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.08