Feb 2026 SWE-rebench: Claude Tops at 65.3%

💡Track top coding models' real GitHub PR fixes—Claude leads, open-weights closing fast
⚡ 30-Second TL;DR
What Changed
Claude Opus 4.6 achieves 65.3% resolved rate with ~70% pass@5
Why It Matters
This tight leaderboard highlights intense competition at the coding frontier, pressuring closed models while open-weights scale closer via better context handling. Practitioners can benchmark agents more reliably on real-world tasks.
What To Do Next
Join the Discord leaderboard channel to test models on fresh SWE-rebench tasks.
Key Points
- •Claude Opus 4.6 achieves 65.3% resolved rate with ~70% pass@5
- •GPT-5.2-medium at 64.4%, GLM-5 and GPT-5.4-medium at 62.8%
- •Gemini 3.1 Pro Preview scores 62.3%, DeepSeek-V3.2 at 60.9%
- •Qwen3.5-397B hits 59.9%, Step-3.5-Flash at 59.6%
🧠 Deep Insight
AI-generated analysis for this event — not the original article.
🔑 Enhanced Key Takeaways
- •The February 2026 SWE-rebench update introduced a new 'hard-mode' evaluation subset focusing on multi-file dependency resolution, which contributed to the lower overall resolution rates compared to previous benchmarks.
- •The benchmark methodology now incorporates a 'human-in-the-loop' verification phase for 15% of the PRs to mitigate potential data contamination from training sets containing GitHub issue resolutions.
- •Analysis of the leaderboard shows a significant shift in inference cost-to-performance ratios, with open-weight models like Qwen3.5-397B achieving parity with proprietary models at approximately 40% of the estimated API cost.
📊 Competitor Analysis▸ Show
| Model | SWE-rebench (Feb 2026) | Est. Context Window | Primary Architecture |
|---|---|---|---|
| Claude Opus 4.6 | 65.3% | 2M tokens | Mixture-of-Experts (MoE) |
| GPT-5.2-medium | 64.4% | 1.5M tokens | Dense Transformer |
| GLM-5 | 62.8% | 1M tokens | Hybrid MoE/Dense |
| Qwen3.5-397B | 59.9% | 1.2M tokens | Sparse MoE |
🛠️ Technical Deep Dive
- •Claude Opus 4.6 utilizes a refined 'Chain-of-Thought' reasoning layer specifically tuned for repository-level code navigation, reducing hallucinations in import resolution.
- •The SWE-rebench evaluation environment was upgraded to support isolated Docker containers with pre-installed language-specific dependency managers (npm, pip, cargo) to ensure consistent execution environments.
- •The performance gap between GPT-5.2-medium and GPT-5.4-medium is attributed to a shift in training data distribution, with the 5.4 version prioritizing long-context reasoning over raw code generation speed.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.