🦙Stalecollected in 6h

Monthly Open-Weight Models SOTA Rankings Update

Monthly Open-Weight Models SOTA Rankings Update
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#sota-rankings#llm-benchmarks#open-modelsopen-weight-sota-rankingslocalllamaopen-weight-models

💡Track open-weight LLMs closing gap to SOTA—key for model choices.

⚡ 30-Second TL;DR

What Changed

Monthly update to SOTA rankings for open-weight models

Why It Matters

Helps practitioners gauge open-source viability versus closed models, informing model selection for cost-effective deployments.

What To Do Next

Review the latest rankings on the r/LocalLLaMA post to benchmark your preferred open models.

Who should care:Researchers & Academics

Key Points

  • Monthly update to SOTA rankings for open-weight models
  • Assesses standing against proprietary leaders
  • Shared on r/LocalLLaMA with discussion thread
  • Tracks progress in LLM performance comparisons

🧠 Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

🔑 Enhanced Key Takeaways

  • As of January 2026, GLM-4.7 (Thinking) leads open-weight rankings with a quality score of 41.7, excelling in reasoning (89%), coding (95%), and math (86%) benchmarks[1][5].
  • DeepSeek V3.2, released building on the 2025 'DeepSeek moment' with R1, tops reasoning and agentic workflows among open models[2][4].
  • OpenAI's gpt-oss-120b (117B MoE) is the first open-weight release since GPT-2, matching o4-mini on AIME, MMLU, and surpassing GPT-4o in some areas[2][4].

🛠️ Technical Deep Dive

  • GLM-4.7 achieves ~42.8% on Humanity’s Last Exam (HLE) with tool use, ~84.9% on LiveCodeBench v6, and ~73.8% on SWE-Bench Verified[5].
  • Qwen3-235B-A22B (hybrid MoE + dual-mode, ~235B total / ~22B active) scores 89.2% on AIME 2025, 91.5% on HumanEval, with 262K native context[5].
  • Llama 4 Scout (109B total, 17B active) supports 10M token context and fits on single H100 with INT4 quantization; Maverick (400B total, 17B active) excels in image and coding[2].
  • Mistral Large 3 (MoE, 41B active / 675B total) trained on 3000 H200 GPUs, Apache 2.0 license, optimized for agentic systems[4].

🔮 Future ImplicationsAI analysis grounded in cited sources

Open-weight models will surpass 50% on HLE benchmark by mid-2026
GLM-4.7 already at ~42.8% HLE with tool use, showing rapid progress in reasoning benchmarks competitive with proprietary leaders[5].
MoE architectures dominate top open-weight rankings
Leading models like gpt-oss-120b (117B MoE), Qwen3-235B-A22B (hybrid MoE), and Mistral Large 3 (675B total MoE) leverage efficiency for frontier performance[2][4][5].

Timeline

2024-04
Meta releases Llama 3 (70B), sparking open-source LLM ecosystem growth[3]
2025-01
DeepSeek R1 launches 'DeepSeek moment' with low-cost frontier reasoning[2]
2025-12
OpenAI releases gpt-oss-120b, first open-weight since GPT-2[2][4]
2026-01
GLM-4.7 and DeepSeek V3.2 top January open-source rankings[1][5]
2026-02
r/LocalLLaMA posts monthly open-weight SOTA rankings update[article]
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.