Baidu Cloud Upgrades to Full-Stack AI Cloud for Agents
๐กLearn how Baidu is optimizing 10k-GPU clusters and KV caching to power the next generation of AI agents.
โก 30-Second TL;DR
What Changed
Introduced Token Factory for agent-first inference optimization
Why It Matters
This upgrade signals a shift toward agent-native infrastructure, potentially lowering the cost and complexity of deploying autonomous AI agents at scale.
What To Do Next
Evaluate your current inference pipeline against the Token Factory architecture to see if KV Cache optimizations can reduce your agent latency.
Key Points
- โขIntroduced Token Factory for agent-first inference optimization
- โขAchieved KV Cache hit rates exceeding 90% for improved latency
- โขKunlun P800 hardware reached 97% training efficiency in 10,000-GPU clusters
๐ง Deep Insight
Web-grounded analysis with 13 cited sources.
๐ Enhanced Key Takeaways
- โขBaidu's Token Factory, an upgrade from its MaaS Model Service, features an Agent-first product architecture designed to minimize token recalculation, resulting in approximately 25% faster inference generation compared to market benchmarks and supporting major domestic models like Ernie, DeepSeek, GLM, and MiniMax.
- โขBaidu Cloud introduced "Harness Engineering," which provides capabilities for long-context management, persistent memory, tool calling, and sub-agent scheduling, achieving a 95% task success rate in browser and Office scenarios with 23% less token consumption than OpenAI's offerings.
- โขThe Kunlun P800, Baidu's third-generation AI accelerator, boasts 345 FP16 TFLOPS, comparable to NVIDIA A100 and Huawei Ascend 910B, and features a unique architectural separation of communication and matrix multiplication units to enable simultaneous data transfer and computation, enhancing scalability.
- โขBaidu activated the Tianchi 256-card supernode, based on Kunlun chips, in April 2026, with an official launch set for June 2026, promising 25% higher throughput and a 50% improvement in inference efficiency for models such as Ernie 5.1, DeepSeek, GLM, and MiniMax.
- โขBaidu's CEO Robin Li proposed "Daily Active Agents" (DAA) as a new core industry metric for the AI era, shifting focus from token consumption or daily active users to the actual task completion loops and value generated by AI agents.
๐ Competitor Analysisโธ Show
| Feature/Category | Baidu AI Cloud | Alibaba Cloud | Tencent Cloud | Huawei Cloud |
|---|---|---|---|---|
| Primary AI Focus | Agent-centric applications, LLM training, autonomous driving (AV), intelligent search. | Broad cloud infrastructure, e-commerce, retail, general business applications. | Social media, video, gaming ecosystems, real-time communication, content delivery. | Hardware-heavy enterprise AI infrastructure, industrial projects, edge AI, 5G integration. |
| Proprietary AI Chips | Kunlun (P800 specialized for LLM training, AV, low latency). | Yitian / Hanguang (specialized for cloud databases and general apps). | Canghai (specialized for video transcoding), focus on real-time experience. | Ascend (focused on heavy AI computing, comparable to Nvidia Blackwell). |
| Market Position (China AI Cloud) | #1 by most measures, leading in smart segments. | Largest cloud provider overall, strong in infrastructure. | Media and Connectivity Titan, strong in social/gaming. | Focus on state-owned enterprises and government-backed projects. |
๐ ๏ธ Technical Deep Dive
- Token Factory: Rebuilt with an Agent-first product architecture to minimize token recalculation, achieving approximately 25% faster inference generation than market benchmarks. It supports major domestic models including Ernie, DeepSeek, GLM, and MiniMax.
- Harness Engineering: Covers long-context management, persistent memory, tool calling, sub-agent scheduling, and Runtime capabilities. It achieves 95% task success rates in browser and Office task scenarios, with 23% less token consumption compared to OpenAI's offerings.
- Kunlun P800 Hardware: A third-generation AI accelerator with a computing power of 345 FP16 TFLOPS. Its architecture features a physical separation of communication units from matrix multiplication units, enabling simultaneous data transfer and computation to improve scalability and reduce latency. It has completed scale validation, delivering multiple 10,000-GPU clusters with a 97% effective training rate and 85%+ linear scaling.
- KV Cache Optimization: Utilizes a layered pooling architecture for GPU memory, DRAM, and SSD, resulting in KV Cache hit rates exceeding 90%.
- Unified Multimodal Training Framework: Delivers 2x training efficiency compared to community standards.
- Tianchi 256-card Supernode: Based on Kunlun chips, this supernode offers 25% higher throughput and a 50% improvement in inference efficiency for adapted models like Ernie 5.1, DeepSeek, GLM, and MiniMax. Its network architecture has been upgraded to HPN5.0, optimizing end-to-end latency by 50%.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (13)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily โ

