Alibaba Unveils New Qwen Models and Custom AI Chips

💡Alibaba is building a full-stack 'AI factory' to challenge global leaders in model training and inference infrastructure
⚡ 30-Second TL;DR
What Changed
Introduction of new Qwen model iterations for enhanced training and inference.
Why It Matters
Alibaba's move to vertically integrate hardware and software signals a major push to dominate the Chinese AI infrastructure market. This could significantly lower costs for developers relying on Alibaba Cloud for large-scale model deployment.
What To Do Next
Review the Alibaba Cloud documentation for the latest Qwen model APIs to evaluate their performance against current open-source alternatives.
Key Points
- •Introduction of new Qwen model iterations for enhanced training and inference.
- •Development of custom AI chips to support large-scale model workloads.
- •Strategic shift toward 'AI factory' model, focusing on revenue generation through compute-intensive services.
- •Integration of cloud infrastructure with autonomous agent capabilities.
🧠 Deep Insight
Web-grounded analysis with 15 cited sources.
🔑 Enhanced Key Takeaways
- •Alibaba's AI-related product revenue has reached an annualized run rate of CNY 35.8 billion (approximately USD 5.2 billion) and currently accounts for 30% of its Cloud Intelligence Group's external revenue, with expectations to exceed 50% within a year.
- •The newly introduced Zhenwu M890 AI chip, developed by Alibaba's T-Head subsidiary, delivers three times the performance of its predecessor, the Zhenwu 810E, and is specifically engineered to handle the high memory and communication demands of autonomous AI agent workloads.
- •Alibaba has committed to an investment exceeding RMB 380 billion (approximately USD 53 billion) over three years in cloud and AI infrastructure, indicating a significant capital expenditure to support its 'AI factory' ambitions.
- •The Qwen model family includes both open-source (e.g., Qwen 3.5-7B, MIT-licensed) and proprietary variants, with some open-weight models demonstrating superior performance in benchmarks like MMLU compared to competitors such as GPT-4o-mini, at significantly lower inference costs.
- •Alibaba's 'AI factory' strategy is driven by the rapid expansion of AI agent workloads, encompassing training, inference, and orchestration, with current demand for compute capacity reportedly outstripping available supply.
📊 Competitor Analysis▸ Show
| Feature/Metric | Alibaba Qwen 3.5-7B (Alibaba Cloud API) | OpenAI GPT-4o-mini | Anthropic Claude 3.5 Haiku | Google Gemma 3-9B |
|---|---|---|---|---|
| MMLU Score | 74.2% | 72.9% | N/A | N/A |
| Input Price (per 1M tokens) | $0.008 | $0.15 | $0.08 | $0.03 |
| Output Price (per 1M tokens) | $0.01 (Together AI) | N/A | N/A | N/A |
| License | MIT-licensed (for open-weight models) | Proprietary | Proprietary | Proprietary |
| On-device Inference | Yes (0.5B model on iPhone 15 Pro at 40 tokens/sec) | No (API-only) | No (API-only) | N/A |
Note: Qwen2.5-Max, an MoE model, is also positioned to compete with DeepSeek V3, OpenAI's GPT-4o, Anthropic's Claude-3.5-Sonnet, Meta's Llama-3.1–405B, and Google's Gemini 2.0 Flash, with claims of outperforming them in certain aspects.
🛠️ Technical Deep Dive
- Qwen Model Architecture: The Qwen family includes both dense transformer models (like Qwen 3.5, optimized for edge hardware) and Mixture-of-Expert (MoE) architectures (like Qwen 2.5-Max).
- Qwen Training Data: Qwen 2.5-Max was pretrained on over 20 trillion tokens, encompassing multi-lingual textual data and domain-specific corpora.
- Qwen Multimodality: Models like Qwen-Omni are end-to-end multimodal, processing text, images, audio, and video, and delivering real-time streaming responses.
- Qwen Thinking Modes: Qwen3 models feature hybrid 'Thinking' and 'Non-Thinking' modes, allowing flexible control over reasoning performance, speed, and costs.
- Zhenwu M890 AI Chip: This custom chip features 144 gigabytes of GPU memory, an upgrade from its predecessor's 96 gigabytes, enabling it to process significantly larger data for complex AI agent workloads.
- Hanguang 800 AI Chip (Predecessor): Alibaba's first AI inference chip, launched in 2019, was built on a 12-nm process with 17 billion transistors. It achieved a peak performance of 78,563 images per second (IPS) on ResNet-50 inference tests.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (15)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: SCMP Technology ↗
