T-Head Launches Panmai 920 AI Smart NIC

💡Fixes AI cluster net bottlenecks, boosts efficiency 14% with 400G smart NIC.
⚡ 30-Second TL;DR
What Changed
Built-in PCIe Switch enables direct low-latency GPU/CPU connections, uniform paths.
Why It Matters
Panmai 920 tackles 'net force' shortages, potentially lifting GPU utilization above 60% in wan-card clusters. Enhances cost-efficiency for AI training/inference at scale. Bolsters Alibaba's full-stack AI infra with compute, storage, and now networking.
What To Do Next
Benchmark Panmai 920 in Aliyun trials for your next GPU cluster deployment.
Key Points
- •Built-in PCIe Switch enables direct low-latency GPU/CPU connections, uniform paths.
- •Multi-path RDMA achieves single QP 400G bandwidth, reduces switch buffer by 90%.
- •Programmable congestion control and fine-grained perception for proactive scheduling.
- •Cuts large model training/inference time by 14% in cluster tests.
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •The Panmai 920 utilizes a custom-designed ASIC architecture that integrates T-Head's proprietary XuanTie processor cores to handle offloaded network protocol processing, distinguishing it from FPGA-based smart NICs.
- •The integration of the PCIe Switch directly on the NIC silicon facilitates a 'disaggregated' server architecture, allowing for more flexible GPU-to-NIC mapping and reducing the physical footprint of high-density AI server racks.
- •The device supports advanced telemetry features that provide real-time, microsecond-level visibility into network congestion states, which are exported directly to Aliyun's centralized AI cluster orchestration software for dynamic traffic shaping.
📊 Competitor Analysis▸ Show
| Feature | T-Head Panmai 920 | NVIDIA BlueField-3 | Broadcom Thor 2 |
|---|---|---|---|
| Throughput | 400G | 400G | 400G |
| Primary Architecture | Custom ASIC w/ XuanTie | DPU (Arm-based) | ASIC (Ethernet-focused) |
| PCIe Switch Integration | Yes (On-die) | No (External) | No (External) |
| Primary Ecosystem | Aliyun / Proprietary | NVIDIA CUDA / Spectrum | Open Ethernet / Broadcom |
🛠️ Technical Deep Dive
- Architecture: Custom ASIC design incorporating multi-core XuanTie RISC-V processors for control plane offloading.
- PCIe Interface: Supports PCIe Gen5 x16, with internal switch fabric enabling direct peer-to-peer (P2P) communication between GPUs without traversing the host CPU root complex.
- RDMA Implementation: Proprietary multi-path RoCE (RDMA over Converged Ethernet) engine designed to handle out-of-order packet delivery and minimize tail latency in large-scale GPU clusters.
- Congestion Control: Hardware-accelerated congestion notification (ECN) combined with a programmable credit-based flow control mechanism to prevent buffer overflow in leaf-spine network topologies.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 极客公园 ↗


