Alibaba Unveils New Qwen Models and Custom AI Chips

Alibaba is building a full-stack 'AI factory' to challenge global leaders in model training and inference infrastructure
30-Second TL;DR
What Changed
Introduction of new Qwen model iterations for enhanced training and inference.
Why It Matters
Alibaba's move to vertically integrate hardware and software signals a major push to dominate the Chinese AI infrastructure market. This could significantly lower costs for developers relying on Alibaba Cloud for large-scale model deployment.
What To Do Next
Review the Alibaba Cloud documentation for the latest Qwen model APIs to evaluate their performance against current open-source alternatives.
Key Points
- •Introduction of new Qwen model iterations for enhanced training and inference.
- •Development of custom AI chips to support large-scale model workloads.
- •Strategic shift toward 'AI factory' model, focusing on revenue generation through compute-intensive services.
- •Integration of cloud infrastructure with autonomous agent capabilities.
Deep Insight
Background and context from public sources — not the original article. 15 sources cited.
Enhanced Key Takeaways
- •Alibaba's AI-related product revenue has reached an annualized run rate of CNY 35.8 billion (approximately USD 5.2 billion) and currently accounts for 30% of its Cloud Intelligence Group's external revenue, with expectations to exceed 50% within a year.
- •The newly introduced Zhenwu M890 AI chip, developed by Alibaba's T-Head subsidiary, delivers three times the performance of its predecessor, the Zhenwu 810E, and is specifically engineered to handle the high memory and communication demands of autonomous AI agent workloads.
- •Alibaba has committed to an investment exceeding RMB 380 billion (approximately USD 53 billion) over three years in cloud and AI infrastructure, indicating a significant capital expenditure to support its 'AI factory' ambitions.
- •The Qwen model family includes both open-source (e.g., Qwen 3.5-7B, MIT-licensed) and proprietary variants, with some open-weight models demonstrating superior performance in benchmarks like MMLU compared to competitors such as GPT-4o-mini, at significantly lower inference costs.
- •Alibaba's 'AI factory' strategy is driven by the rapid expansion of AI agent workloads, encompassing training, inference, and orchestration, with current demand for compute capacity reportedly outstripping available supply.
Competitor Analysis
- Alibaba Qwen 3.5-7B (Alibaba Cloud API)
- 74.2%
- OpenAI GPT-4o-mini
- 72.9%
- Anthropic Claude 3.5 Haiku
- N/A
- Google Gemma 3-9B
- N/A
- Alibaba Qwen 3.5-7B (Alibaba Cloud API)
- $0.008
- OpenAI GPT-4o-mini
- $0.15
- Anthropic Claude 3.5 Haiku
- $0.08
- Google Gemma 3-9B
- $0.03
- Alibaba Qwen 3.5-7B (Alibaba Cloud API)
- $0.01 (Together AI)
- OpenAI GPT-4o-mini
- N/A
- Anthropic Claude 3.5 Haiku
- N/A
- Google Gemma 3-9B
- N/A
- Alibaba Qwen 3.5-7B (Alibaba Cloud API)
- MIT-licensed (for open-weight models)
- OpenAI GPT-4o-mini
- Proprietary
- Anthropic Claude 3.5 Haiku
- Proprietary
- Google Gemma 3-9B
- Proprietary
- Alibaba Qwen 3.5-7B (Alibaba Cloud API)
- Yes (0.5B model on iPhone 15 Pro at 40 tokens/sec)
- OpenAI GPT-4o-mini
- No (API-only)
- Anthropic Claude 3.5 Haiku
- No (API-only)
- Google Gemma 3-9B
- N/A
| Feature/Metric | Alibaba Qwen 3.5-7B (Alibaba Cloud API) | OpenAI GPT-4o-mini | Anthropic Claude 3.5 Haiku | Google Gemma 3-9B |
|---|---|---|---|---|
| MMLU Score | 74.2% | 72.9% | N/A | N/A |
| Input Price (per 1M tokens) | $0.008 | $0.15 | $0.08 | $0.03 |
| Output Price (per 1M tokens) | $0.01 (Together AI) | N/A | N/A | N/A |
| License | MIT-licensed (for open-weight models) | Proprietary | Proprietary | Proprietary |
| On-device Inference | Yes (0.5B model on iPhone 15 Pro at 40 tokens/sec) | No (API-only) | No (API-only) | N/A |
Technical Deep Dive
- Qwen Model Architecture: The Qwen family includes both dense transformer models (like Qwen 3.5, optimized for edge hardware) and Mixture-of-Expert (MoE) architectures (like Qwen 2.5-Max).
- Qwen Training Data: Qwen 2.5-Max was pretrained on over 20 trillion tokens, encompassing multi-lingual textual data and domain-specific corpora.
- Qwen Multimodality: Models like Qwen-Omni are end-to-end multimodal, processing text, images, audio, and video, and delivering real-time streaming responses.
- Qwen Thinking Modes: Qwen3 models feature hybrid 'Thinking' and 'Non-Thinking' modes, allowing flexible control over reasoning performance, speed, and costs.
- Zhenwu M890 AI Chip: This custom chip features 144 gigabytes of GPU memory, an upgrade from its predecessor's 96 gigabytes, enabling it to process significantly larger data for complex AI agent workloads.
- Hanguang 800 AI Chip (Predecessor): Alibaba's first AI inference chip, launched in 2019, was built on a 12-nm process with 17 billion transistors. It achieved a peak performance of 78,563 images per second (IPS) on ResNet-50 inference tests.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2009-09Alibaba Cloud officially established.
- 2017-LateDAMO Academy, Alibaba's research institute, launched.
- 2018-09T-Head, Alibaba's semiconductor division, spun out of DAMO Academy.
- 2019-09Alibaba unveils Hanguang 800, its first AI inference chip.
- 2023-04Alibaba launches beta of Tongyi Qianwen (Qwen) large language model.
- 2026-05Alibaba unveils Zhenwu M890 AI chip and Qwen 3.7-Max model.
Sources (15)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: SCMP Technology ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.



