Qwen 3.7 Max Preview Released: Top-Tier Performance

💡New Qwen 3.7 Max preview claims top-tier performance in text and vision, challenging current industry benchmarks.
⚡ 30-Second TL;DR
What Changed
Qwen 3.7 Max preview version is now available.
Why It Matters
This release reinforces Alibaba's position as a leader in the domestic Chinese AI market, providing developers with a high-performance alternative to global models.
What To Do Next
Integrate the Qwen 3.7 Max API into your current workflow to benchmark its reasoning capabilities against your existing LLM stack.
Key Points
- •Qwen 3.7 Max preview version is now available.
- •Achieved top-tier performance rankings in both text and vision benchmarks.
- •Alibaba continues rapid iteration cycles despite key team changes.
🧠 Deep Insight
Web-grounded analysis with 24 cited sources.
🔑 Enhanced Key Takeaways
- •Qwen 3.7 Max is a proprietary model, distinguishing it from many earlier Qwen versions that were open-source.
- •The model boasts over 1 trillion parameters, positioning it among an elite group of AI models globally.
- •It features an ultra-long context window of up to 262,144 tokens, significantly enhancing its ability to process extensive documents and complex conversations.
- •Qwen 3.7 Max demonstrates strong performance in complex reasoning, coding, and handling structured data formats like JSON.
- •Despite recent departures of key personnel, including a core leader of the Qwen model, Alibaba is reorganizing leadership and increasing investment in AI research and development, reaffirming its long-term strategy for foundational large models.
📊 Competitor Analysis▸ Show
| Feature/Metric | Qwen 3.7 Max (Preview) | Claude 3.7 Sonnet |
|---|---|---|
| Parameters | >1 Trillion | Estimated hundreds of billions (for Claude Opus 4) |
| Context Window | 256,000 - 262,144 tokens | 200,000 tokens |
| Input Pricing (per 1M tokens) | Starts at $1.20 (for 0-32K tokens) | $3.00 |
| Output Pricing (per 1M tokens) | Starts at $6.00 (for 0-32K tokens) | $15.00 |
| GPQA Benchmark | 76.4% | 65.6% (Non-reasoning) / 77.2% (Reasoning) |
| MMLU Pro Benchmark | 84.1% | 80.3% |
| LiveCodeBench | 76.7% | 39.4% |
| AIME 2025 Benchmark | 80.7% | 21.0% |
| Key Capabilities | Complex reasoning, coding, structured data (JSON), agentic behaviors, multilingual (100+ languages), multimodal (text & vision) | Text and images, external tools/APIs, advanced reasoning |
🛠️ Technical Deep Dive
- Parameters: Qwen3-Max features over 1 trillion parameters.
- Training Data: Qwen3-Max was pretrained on 36 trillion tokens. Earlier Qwen series models were trained on diverse multilingual and multimodal datasets, with Qwen-7B on up to 3 trillion tokens and Qwen2 series on 7 trillion tokens.
- Architecture: Built on a transformer-based architecture. The Qwen3 series incorporates a Mixture-of-Experts (MoE) architecture, with Qwen3-Next utilizing a highly sparse MoE design (80B total parameters, ~3B activated per inference step) with 512 total experts. It also features a hybrid architecture combining Gated DeltaNet with Gated Attention in a 3:1 ratio.
- Attention Mechanisms: Includes innovations in attention mechanisms, Rotary Positional Embeddings (RoPE), and SwiGLU activation functions.
- Context Window: Supports an ultra-long context window of 256,000 to 262,144 tokens, with some models like Qwen3.6-35B-A3B extensible up to 1,010,000 tokens.
- Training Efficiency: Optimized by PAI-FlashMoE's efficient multi-level pipeline parallelism strategy, leading to a 30% relative increase in MFU (Model FLOPs Utilization) for Qwen3-Max-Base and over 300% improvement in MoE training acceleration for the Qwen series.
- Operational Modes: Supports both a 'Thinking Mode' for step-by-step reasoning on complex problems and a 'Non-Thinking Mode' for quick, near-instant responses to simpler queries.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (24)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗