๐ŸผStalecollected in 28m

Baidu Cloud Upgrades to Full-Stack AI Cloud for Agents

PostLinkedIn
๐ŸผRead original on Pandaily

๐Ÿ’กLearn how Baidu is optimizing 10k-GPU clusters and KV caching to power the next generation of AI agents.

โšก 30-Second TL;DR

What Changed

Introduced Token Factory for agent-first inference optimization

Why It Matters

This upgrade signals a shift toward agent-native infrastructure, potentially lowering the cost and complexity of deploying autonomous AI agents at scale.

What To Do Next

Evaluate your current inference pipeline against the Token Factory architecture to see if KV Cache optimizations can reduce your agent latency.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขIntroduced Token Factory for agent-first inference optimization
  • โ€ขAchieved KV Cache hit rates exceeding 90% for improved latency
  • โ€ขKunlun P800 hardware reached 97% training efficiency in 10,000-GPU clusters

๐Ÿง  Deep Insight

Web-grounded analysis with 13 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขBaidu's Token Factory, an upgrade from its MaaS Model Service, features an Agent-first product architecture designed to minimize token recalculation, resulting in approximately 25% faster inference generation compared to market benchmarks and supporting major domestic models like Ernie, DeepSeek, GLM, and MiniMax.
  • โ€ขBaidu Cloud introduced "Harness Engineering," which provides capabilities for long-context management, persistent memory, tool calling, and sub-agent scheduling, achieving a 95% task success rate in browser and Office scenarios with 23% less token consumption than OpenAI's offerings.
  • โ€ขThe Kunlun P800, Baidu's third-generation AI accelerator, boasts 345 FP16 TFLOPS, comparable to NVIDIA A100 and Huawei Ascend 910B, and features a unique architectural separation of communication and matrix multiplication units to enable simultaneous data transfer and computation, enhancing scalability.
  • โ€ขBaidu activated the Tianchi 256-card supernode, based on Kunlun chips, in April 2026, with an official launch set for June 2026, promising 25% higher throughput and a 50% improvement in inference efficiency for models such as Ernie 5.1, DeepSeek, GLM, and MiniMax.
  • โ€ขBaidu's CEO Robin Li proposed "Daily Active Agents" (DAA) as a new core industry metric for the AI era, shifting focus from token consumption or daily active users to the actual task completion loops and value generated by AI agents.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature/CategoryBaidu AI CloudAlibaba CloudTencent CloudHuawei Cloud
Primary AI FocusAgent-centric applications, LLM training, autonomous driving (AV), intelligent search.Broad cloud infrastructure, e-commerce, retail, general business applications.Social media, video, gaming ecosystems, real-time communication, content delivery.Hardware-heavy enterprise AI infrastructure, industrial projects, edge AI, 5G integration.
Proprietary AI ChipsKunlun (P800 specialized for LLM training, AV, low latency).Yitian / Hanguang (specialized for cloud databases and general apps).Canghai (specialized for video transcoding), focus on real-time experience.Ascend (focused on heavy AI computing, comparable to Nvidia Blackwell).
Market Position (China AI Cloud)#1 by most measures, leading in smart segments.Largest cloud provider overall, strong in infrastructure.Media and Connectivity Titan, strong in social/gaming.Focus on state-owned enterprises and government-backed projects.

๐Ÿ› ๏ธ Technical Deep Dive

  • Token Factory: Rebuilt with an Agent-first product architecture to minimize token recalculation, achieving approximately 25% faster inference generation than market benchmarks. It supports major domestic models including Ernie, DeepSeek, GLM, and MiniMax.
  • Harness Engineering: Covers long-context management, persistent memory, tool calling, sub-agent scheduling, and Runtime capabilities. It achieves 95% task success rates in browser and Office task scenarios, with 23% less token consumption compared to OpenAI's offerings.
  • Kunlun P800 Hardware: A third-generation AI accelerator with a computing power of 345 FP16 TFLOPS. Its architecture features a physical separation of communication units from matrix multiplication units, enabling simultaneous data transfer and computation to improve scalability and reduce latency. It has completed scale validation, delivering multiple 10,000-GPU clusters with a 97% effective training rate and 85%+ linear scaling.
  • KV Cache Optimization: Utilizes a layered pooling architecture for GPU memory, DRAM, and SSD, resulting in KV Cache hit rates exceeding 90%.
  • Unified Multimodal Training Framework: Delivers 2x training efficiency compared to community standards.
  • Tianchi 256-card Supernode: Based on Kunlun chips, this supernode offers 25% higher throughput and a 50% improvement in inference efficiency for adapted models like Ernie 5.1, DeepSeek, GLM, and MiniMax. Its network architecture has been upgraded to HPN5.0, optimizing end-to-end latency by 50%.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Baidu's emphasis on "Daily Active Agents" (DAA) could become a new industry standard for measuring AI platform success.
By proposing DAA as a core metric, Baidu is advocating for a shift from raw token consumption or user numbers to the actual value and productivity generated by AI agents, potentially influencing how AI solutions are evaluated across the industry.
The upgraded full-stack AI cloud, with innovations like "Harness Engineering" and "Token Factory," will significantly accelerate the adoption of complex AI agents by enterprises.
These optimizations, including faster inference, reduced token consumption, and robust agent management, make large-scale, efficient, and cost-effective deployment of AI agent applications more accessible for businesses.
Baidu's continued vertical integration with proprietary Kunlun chips and an optimized AI stack will solidify its leadership in China's domestic AI infrastructure market.
The performance and scalability of the Kunlun P800 and Tianchi supernode demonstrate Baidu's capability to provide competitive homegrown AI compute, reducing reliance on foreign hardware amidst geopolitical considerations.

โณ Timeline

2000-01
Baidu founded by Robin Li and Eric Xu.
2017
Baidu launched its Apollo autonomous driving platform, marking a significant AI initiative.
2023
Baidu unveiled Ernie Bot, its large language model.
2024-04
Baidu introduced AgentBuilder, AppBuilder, and ModelBuilder toolkits at Create 2024, laying groundwork for agent development.
2025-01
Baidu's Qianfan platform was upgraded to an agent-centric architecture.
2026-05
Baidu Cloud upgraded to a full-stack AI cloud for agents at Create 2026, introducing Token Factory and Harness Engineering.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily โ†—