Qwen Code Posts 100% SWE-bench Smoke-Test Score
💡See how Qwen Code performed on SWE-bench—and why its Terminal-Bench result was excluded.
⚡ 30-Second TL;DR
What Changed
SWE-bench Verified resolved 1 of 1 cases with a 100.00% score.
Why It Matters
The SWE-bench result is an encouraging signal for Qwen Code’s coding-agent workflow, but the one-case sample is too small to establish broad performance. The Terminal-Bench infrastructure failure also highlights the need to separate model capability from evaluation reliability.
What To Do Next
Reproduce the SWE-bench Verified run with Qwen Code v0.21.13 and qwen3.7-plus across a larger case set, while logging infrastructure failures separately.
Key Points
- •SWE-bench Verified resolved 1 of 1 cases with a 100.00% score.
- •Terminal-Bench 2.0 completed 1 case but recorded 1 infrastructure failure and no valid grader results.
- •The run used Qwen Code v0.21.13 with the qwen3.7-plus model.
- •The benchmark was an isolated DSW EAS network and watchdog smoke test, not a broad benchmark campaign.
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •The qwen3.7-plus model represents a significant iteration in the Qwen series, specifically optimized for long-context reasoning and multi-step software engineering tasks.
- •DSW EAS (Data Science Workshop Elastic AI Service) is an Alibaba Cloud infrastructure environment designed to provide isolated, containerized execution for AI agent testing.
- •The 'smoke test' methodology indicates a shift toward rapid, targeted validation of agentic workflows rather than relying solely on massive, time-intensive benchmark suites.
- •Terminal-Bench 2.0 is an emerging evaluation framework focused on assessing an AI's ability to navigate complex, stateful command-line interfaces and file system operations.
- •The infrastructure failure during Terminal-Bench 2.0 highlights the ongoing challenge of maintaining stable, reproducible environments for autonomous agent evaluation at scale.
📊 Competitor Analysis▸ Show
| Feature | Qwen Code (qwen3.7-plus) | Claude 3.5 Sonnet | GPT-4o (SWE-bench) |
|---|---|---|---|
| SWE-bench Verified | 100% (Smoke Test) | ~60-70% (Full) | ~50-60% (Full) |
| Primary Focus | Agentic Coding/CLI | General Reasoning | General Reasoning |
| Infrastructure | DSW EAS | Anthropic API | OpenAI API |
🛠️ Technical Deep Dive
- Model Architecture: qwen3.7-plus utilizes a Mixture-of-Experts (MoE) architecture optimized for low-latency inference in coding tasks.
- Context Window: Enhanced support for repository-level context, allowing the model to maintain state across multiple file modifications.
- Execution Environment: The DSW EAS watchdog monitors for process hangs and infinite loops, automatically terminating non-responsive agentic sessions.
- Evaluation Protocol: The smoke test utilized a sandboxed Linux environment with restricted network access to ensure deterministic results.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Qwen (GitHub Releases: qwen-code) ↗
