๐Ÿ‡ญ๐Ÿ‡ฐStalecollected in 30m

Alibaba's Qwen3.7-Max Outperforms OpenAI and Google in Coding

Alibaba's Qwen3.7-Max Outperforms OpenAI and Google in Coding
PostLinkedIn
๐Ÿ‡ญ๐Ÿ‡ฐRead original on SCMP Technology

๐Ÿ’กFirst non-US model to break the top 5 in global coding benchmarks, challenging OpenAI and Google's dominance.

โšก 30-Second TL;DR

What Changed

Qwen3.7-Max achieved a score of 1,541 on the Code Arena coding leaderboard.

Why It Matters

This ranking challenges the dominance of US-based AI labs in specialized coding tasks. It suggests that Alibaba's Qwen series is becoming a viable enterprise alternative for high-stakes software development workflows.

What To Do Next

Evaluate Qwen3.7-Max via API to benchmark its performance against your current coding assistant or automated code generation pipeline.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขQwen3.7-Max achieved a score of 1,541 on the Code Arena coding leaderboard.
  • โ€ขAlibaba is now the only developer other than Anthropic to hold a top-five position.
  • โ€ขThe model outperformed competing iterations from industry leaders OpenAI and Google.

๐Ÿง  Deep Insight

Web-grounded analysis with 16 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขQwen3.7-Max is positioned as an "Agent Foundation Model" specifically designed for long-term autonomous task execution, demonstrated by internal tests where it ran continuously for 35 hours, executing 1,158 tool calls and achieving a 10x geometric mean speedup in kernel optimization.
  • โ€ขThe model features a substantial 1-million-token context window, allowing it to process extensive codebases or lengthy technical documents in a single request, a significant increase from its predecessor's 256K limit.
  • โ€ขIt supports "cross-harness generalization," enabling native integration with diverse agent frameworks, including the Anthropic API protocol, for use with tools like Claude Code or OpenClaw.
  • โ€ขQwen3.7-Max is a proprietary and closed-weight model, marking a strategic shift from Alibaba's previous open-source approach for many models within the Qwen series.
  • โ€ขBeyond its Code Arena performance, Qwen3.7-Max also achieved strong scores on other benchmarks, including 44.5 on Apex Math Reasoning, 41.4 on Humanity's Last Exam, and 76.4 on MCP-Atlas, outperforming some competing models from Claude and DeepSeek in these areas.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature/BenchmarkAlibaba Qwen3.7-MaxAnthropic Claude Opus 4.7OpenAI GPT-5.5Google Gemini 3.1 Pro Preview
Code Arena Score1,541 (Rank 4)Higher than 1,541 (Rank 1 & 2)OutperformedOutperformed
Artificial Analysis Intelligence Index56.6 (Rank 5)57.3 (Rank 3)60.2 (Rank 1)57.2 (Rank 4)
Apex Math Reasoning44.534.5 (Opus 4.6 Max)N/AN/A
Context Window1M tokens1M tokens1M tokens1M tokens
Input Cost (per 1M tokens)$2.50$5.00$5.00$2.00-$4.00
Output Cost (per 1M tokens)$7.50$25.00$30.00$12.00-$18.00
Open-source WeightsNo (Proprietary)NoNoNo

๐Ÿ› ๏ธ Technical Deep Dive

  • Qwen3.7-Max is built on a transformer-based architecture, incorporating advanced attention mechanisms.
  • It belongs to the Qwen 3 series, which includes both dense and Mixture-of-Experts (MoE) variants, though Qwen3.7-Max is a proprietary, closed-weight model.
  • The model utilizes hybrid reasoning modes, referred to as "Thinking" and "Non-Thinking," allowing for flexible control over reasoning performance, speed, and costs.
  • It features a 1-million-token context window, designed to handle extensive inputs for complex tasks.
  • Qwen3.7-Max was trained using "environment scaling," involving a vast array of dynamic agentic environments to enhance its autonomous capabilities.
  • It incorporates built-in reward-hacking self-monitoring, enabling it to autonomously detect and correct its own behavior when attempting to exploit training environments.
  • The model employs explicit chain-of-thought reasoning, which generally improves performance on complex reasoning and mathematical tasks.
  • Designed for agent-centric workloads, it supports over 1,000 tool calls per session, facilitating multi-step task execution.
  • The Qwen Code architecture, which Qwen3.7-Max integrates with, is composed of a CLI (user-facing) package and a Core (backend) package, with the Core handling API client communication, prompt construction, tool registration and execution, and state management.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Alibaba's Qwen series will significantly intensify competition in the global AI agent and coding LLM market.
Qwen3.7-Max's strong performance on Code Arena and other benchmarks, coupled with its agentic capabilities, positions Alibaba as a major challenger to established Western AI leaders.
The trend of AI models focusing on long-horizon autonomous task execution will accelerate.
Qwen3.7-Max's demonstrated ability to run autonomously for extended periods and execute numerous tool calls highlights a critical capability for future AI applications beyond simple conversational agents.
The distinction between open-source and proprietary flagship models will become more pronounced in the competitive landscape.
While earlier Qwen models were open-source, Qwen3.7-Max is proprietary and API-only, indicating a strategy to monetize advanced capabilities while potentially maintaining an open-source ecosystem for other variants.

โณ Timeline

2023-04
Alibaba launched the initial Qwen models
2024
Qwen2 series launched, bringing leaps in reasoning, coding, and multilingual understanding
2025-01-29
Alibaba launched Qwen2.5-Max
2025-11-17
Code Arena, a new evaluation platform for AI coding performance, launched by LMArena
2026-05-18
Qwen3.7 Max stable release
2026-05-20
Alibaba officially unveiled Qwen3.7-Max at the Alibaba Cloud Summit
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: SCMP Technology โ†—