Search

Tag: #datacenter12 results

NVIDIA Rubin+Groq Hits $1T GPU Projection

NVIDIA Rubin+Groq Hits $1T GPU Projection

At GTC, NVIDIA unveiled the Vera Rubin system integrating Groq LPU, a 7-chip AI infrastructure optimized for inference with disaggregated prefill on GPU and decode on LPU. This addresses GPU limits in high-speed token generation, boosting 1GW datacenter token output 350x. NVIDIA projects $1T GPU market by 2027 driven by agentic AI inference surge.

虎嗅MediaMar 16#gtc#lpu#inference
Cross-DC KV Cache for LLMs

Cross-DC KV Cache for LLMs

Moonshot's Kimi enables prefill/decode disaggregation across datacenters using hybrid Kimi Linear model to shrink KV cache. Delivers 1.54x throughput and 64% lower P90 TTFT on 20x scaled model. Lowers token costs via PaaS.

Reddit r/LocalLLaMACommunityApr 18#kv-cache#prefill-decode#datacenter
Page 1 of 2