Qwen 3.7 Max Preview Released: Top-Tier Performance

New Qwen 3.7 Max preview claims top-tier performance in text and vision, challenging current industry benchmarks.
30-Second TL;DR
What Changed
Qwen 3.7 Max preview version is now available.
Why It Matters
This release reinforces Alibaba's position as a leader in the domestic Chinese AI market, providing developers with a high-performance alternative to global models.
What To Do Next
Integrate the Qwen 3.7 Max API into your current workflow to benchmark its reasoning capabilities against your existing LLM stack.
Key Points
- •Qwen 3.7 Max preview version is now available.
- •Achieved top-tier performance rankings in both text and vision benchmarks.
- •Alibaba continues rapid iteration cycles despite key team changes.
Deep Insight
Background and context from public sources — not the original article. 24 sources cited.
Enhanced Key Takeaways
- •Qwen 3.7 Max is a proprietary model, distinguishing it from many earlier Qwen versions that were open-source.
- •The model boasts over 1 trillion parameters, positioning it among an elite group of AI models globally.
- •It features an ultra-long context window of up to 262,144 tokens, significantly enhancing its ability to process extensive documents and complex conversations.
- •Qwen 3.7 Max demonstrates strong performance in complex reasoning, coding, and handling structured data formats like JSON.
- •Despite recent departures of key personnel, including a core leader of the Qwen model, Alibaba is reorganizing leadership and increasing investment in AI research and development, reaffirming its long-term strategy for foundational large models.
Competitor Analysis
- Qwen 3.7 Max (Preview)
1 Trillion
- Claude 3.7 Sonnet
- Estimated hundreds of billions (for Claude Opus 4)
- Qwen 3.7 Max (Preview)
- 256,000 - 262,144 tokens
- Claude 3.7 Sonnet
- 200,000 tokens
- Qwen 3.7 Max (Preview)
- Starts at $1.20 (for 0-32K tokens)
- Claude 3.7 Sonnet
- $3.00
- Qwen 3.7 Max (Preview)
- Starts at $6.00 (for 0-32K tokens)
- Claude 3.7 Sonnet
- $15.00
- Qwen 3.7 Max (Preview)
- 76.4%
- Claude 3.7 Sonnet
- 65.6% (Non-reasoning) / 77.2% (Reasoning)
- Qwen 3.7 Max (Preview)
- 84.1%
- Claude 3.7 Sonnet
- 80.3%
- Qwen 3.7 Max (Preview)
- 76.7%
- Claude 3.7 Sonnet
- 39.4%
- Qwen 3.7 Max (Preview)
- 80.7%
- Claude 3.7 Sonnet
- 21.0%
- Qwen 3.7 Max (Preview)
- Complex reasoning, coding, structured data (JSON), agentic behaviors, multilingual (100+ languages), multimodal (text & vision)
- Claude 3.7 Sonnet
- Text and images, external tools/APIs, advanced reasoning
| Feature/Metric | Qwen 3.7 Max (Preview) | Claude 3.7 Sonnet |
|---|---|---|
| Parameters | 1 Trillion | Estimated hundreds of billions (for Claude Opus 4) |
| Context Window | 256,000 - 262,144 tokens | 200,000 tokens |
| Input Pricing (per 1M tokens) | Starts at $1.20 (for 0-32K tokens) | $3.00 |
| Output Pricing (per 1M tokens) | Starts at $6.00 (for 0-32K tokens) | $15.00 |
| GPQA Benchmark | 76.4% | 65.6% (Non-reasoning) / 77.2% (Reasoning) |
| MMLU Pro Benchmark | 84.1% | 80.3% |
| LiveCodeBench | 76.7% | 39.4% |
| AIME 2025 Benchmark | 80.7% | 21.0% |
| Key Capabilities | Complex reasoning, coding, structured data (JSON), agentic behaviors, multilingual (100+ languages), multimodal (text & vision) | Text and images, external tools/APIs, advanced reasoning |
Technical Deep Dive
- Parameters: Qwen3-Max features over 1 trillion parameters.
- Training Data: Qwen3-Max was pretrained on 36 trillion tokens. Earlier Qwen series models were trained on diverse multilingual and multimodal datasets, with Qwen-7B on up to 3 trillion tokens and Qwen2 series on 7 trillion tokens.
- Architecture: Built on a transformer-based architecture. The Qwen3 series incorporates a Mixture-of-Experts (MoE) architecture, with Qwen3-Next utilizing a highly sparse MoE design (80B total parameters, ~3B activated per inference step) with 512 total experts. It also features a hybrid architecture combining Gated DeltaNet with Gated Attention in a 3:1 ratio.
- Attention Mechanisms: Includes innovations in attention mechanisms, Rotary Positional Embeddings (RoPE), and SwiGLU activation functions.
- Context Window: Supports an ultra-long context window of 256,000 to 262,144 tokens, with some models like Qwen3.6-35B-A3B extensible up to 1,010,000 tokens.
- Training Efficiency: Optimized by PAI-FlashMoE's efficient multi-level pipeline parallelism strategy, leading to a 30% relative increase in MFU (Model FLOPs Utilization) for Qwen3-Max-Base and over 300% improvement in MoE training acceleration for the Qwen series.
- Operational Modes: Supports both a 'Thinking Mode' for step-by-step reasoning on complex problems and a 'Non-Thinking Mode' for quick, near-instant responses to simpler queries.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2023-04Alibaba launched a beta of Qwen (Tongyi Qianwen).
- 2023-08Alibaba released its first open-model Qwen-7B.
- 2024-06Alibaba released the open-model Qwen2 series, encompassing dense and sparse models.
- 2025-01Alibaba released Qwen2.5-VL, a visual-language open model with remarkable multimodal capabilities.
- 2025-09Alibaba officially released Qwen3-Max, its largest LLM model with over 1 trillion parameters, introducing the Qwen3 family.
- 2026-05Alibaba launched the Qwen 3.7 Max preview, featuring significant agentic coding improvements and enhanced world knowledge.
Sources (24)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.