
DeepSeek's DualPath Boosts Agent LLM Inference 1.9x
DeepSeek collaborated with Tsinghua and Peking University on a new paper introducing DualPath, an inference system optimizing KV-Cache loading for agentic LLMs. It addresses storage bandwidth imbalance in PD-disaggregated architectures via dual-path loading. Achieves 1.87x offline throughput and 1.96x online service throughput.





