SourceStalecollected in 28m

Baidu Cloud Upgrades to Full-Stack AI Cloud for Agents

Read original on Pandaily
#cloud-computing#agentic-ai#gpu-clusters

Learn how Baidu is optimizing 10k-GPU clusters and KV caching to power the next generation of AI agents.

30-Second TL;DR

What Changed

Introduced Token Factory for agent-first inference optimization

Why It Matters

This upgrade signals a shift toward agent-native infrastructure, potentially lowering the cost and complexity of deploying autonomous AI agents at scale.

What To Do Next

Evaluate your current inference pipeline against the Token Factory architecture to see if KV Cache optimizations can reduce your agent latency.

Who should care:Developers & AI Engineers

Key Points

  • Introduced Token Factory for agent-first inference optimization
  • Achieved KV Cache hit rates exceeding 90% for improved latency
  • Kunlun P800 hardware reached 97% training efficiency in 10,000-GPU clusters
Key numbers25%95%23%50%

Deep Insight

Background and context from public sources — not the original article. 13 sources cited.

Enhanced Key Takeaways

  • Baidu's Token Factory, an upgrade from its MaaS Model Service, features an Agent-first product architecture designed to minimize token recalculation, resulting in approximately 25% faster inference generation compared to market benchmarks and supporting major domestic models like Ernie, DeepSeek, GLM, and MiniMax.
  • Baidu Cloud introduced "Harness Engineering," which provides capabilities for long-context management, persistent memory, tool calling, and sub-agent scheduling, achieving a 95% task success rate in browser and Office scenarios with 23% less token consumption than OpenAI's offerings.
  • The Kunlun P800, Baidu's third-generation AI accelerator, boasts 345 FP16 TFLOPS, comparable to NVIDIA A100 and Huawei Ascend 910B, and features a unique architectural separation of communication and matrix multiplication units to enable simultaneous data transfer and computation, enhancing scalability.
  • Baidu activated the Tianchi 256-card supernode, based on Kunlun chips, in April 2026, with an official launch set for June 2026, promising 25% higher throughput and a 50% improvement in inference efficiency for models such as Ernie 5.1, DeepSeek, GLM, and MiniMax.
  • Baidu's CEO Robin Li proposed "Daily Active Agents" (DAA) as a new core industry metric for the AI era, shifting focus from token consumption or daily active users to the actual task completion loops and value generated by AI agents.

Competitor Analysis

Primary AI Focus
Baidu AI Cloud
Agent-centric applications, LLM training, autonomous driving (AV), intelligent search.
Alibaba Cloud
Broad cloud infrastructure, e-commerce, retail, general business applications.
Tencent Cloud
Social media, video, gaming ecosystems, real-time communication, content delivery.
Huawei Cloud
Hardware-heavy enterprise AI infrastructure, industrial projects, edge AI, 5G integration.
Proprietary AI Chips
Baidu AI Cloud
Kunlun (P800 specialized for LLM training, AV, low latency).
Alibaba Cloud
Yitian / Hanguang (specialized for cloud databases and general apps).
Tencent Cloud
Canghai (specialized for video transcoding), focus on real-time experience.
Huawei Cloud
Ascend (focused on heavy AI computing, comparable to Nvidia Blackwell).
Market Position (China AI Cloud)
Baidu AI Cloud
#1 by most measures, leading in smart segments.
Alibaba Cloud
Largest cloud provider overall, strong in infrastructure.
Tencent Cloud
Media and Connectivity Titan, strong in social/gaming.
Huawei Cloud
Focus on state-owned enterprises and government-backed projects.

Technical Deep Dive

  • Token Factory: Rebuilt with an Agent-first product architecture to minimize token recalculation, achieving approximately 25% faster inference generation than market benchmarks. It supports major domestic models including Ernie, DeepSeek, GLM, and MiniMax.
  • Harness Engineering: Covers long-context management, persistent memory, tool calling, sub-agent scheduling, and Runtime capabilities. It achieves 95% task success rates in browser and Office task scenarios, with 23% less token consumption compared to OpenAI's offerings.
  • Kunlun P800 Hardware: A third-generation AI accelerator with a computing power of 345 FP16 TFLOPS. Its architecture features a physical separation of communication units from matrix multiplication units, enabling simultaneous data transfer and computation to improve scalability and reduce latency. It has completed scale validation, delivering multiple 10,000-GPU clusters with a 97% effective training rate and 85%+ linear scaling.
  • KV Cache Optimization: Utilizes a layered pooling architecture for GPU memory, DRAM, and SSD, resulting in KV Cache hit rates exceeding 90%.
  • Unified Multimodal Training Framework: Delivers 2x training efficiency compared to community standards.
  • Tianchi 256-card Supernode: Based on Kunlun chips, this supernode offers 25% higher throughput and a 50% improvement in inference efficiency for adapted models like Ernie 5.1, DeepSeek, GLM, and MiniMax. Its network architecture has been upgraded to HPN5.0, optimizing end-to-end latency by 50%.

Future ImplicationsAI analysis grounded in cited sources

Baidu's emphasis on "Daily Active Agents" (DAA) could become a new industry standard for measuring AI platform success.
By proposing DAA as a core metric, Baidu is advocating for a shift from raw token consumption or user numbers to the actual value and productivity generated by AI agents, potentially influencing how AI solutions are evaluated across the industry.
The upgraded full-stack AI cloud, with innovations like "Harness Engineering" and "Token Factory," will significantly accelerate the adoption of complex AI agents by enterprises.
These optimizations, including faster inference, reduced token consumption, and robust agent management, make large-scale, efficient, and cost-effective deployment of AI agent applications more accessible for businesses.
Baidu's continued vertical integration with proprietary Kunlun chips and an optimized AI stack will solidify its leadership in China's domestic AI infrastructure market.
The performance and scalability of the Kunlun P800 and Tianchi supernode demonstrate Baidu's capability to provide competitive homegrown AI compute, reducing reliance on foreign hardware amidst geopolitical considerations.

Timeline

2000-01
Baidu founded by Robin Li and Eric Xu.
2017
Baidu launched its Apollo autonomous driving platform, marking a significant AI initiative.
2023
Baidu unveiled Ernie Bot, its large language model.
2024-04
Baidu introduced AgentBuilder, AppBuilder, and ModelBuilder toolkits at Create 2024, laying groundwork for agent development.
2025-01
Baidu's Qianfan platform was upgraded to an agent-centric architecture.
2026-05
Baidu Cloud upgraded to a full-stack AI cloud for agents at Create 2026, introducing Token Factory and Harness Engineering.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.