
LLM Inference Hardware Crisis Worse Than Thought
Google DeepMind engineers' paper exposes how GPUs/TPUs, optimized for training, fail at memory-intensive LLM inference decoding, driving up costs. Trends like MoE and long contexts amplify the gap as memory bandwidth lags compute. They propose HBF, PNM, and 3D stacking to address core pain points.







