Search

Tag: #leaderboard25 results

Feb 2026 SWE-rebench: Claude Tops at 65.3%

Feb 2026 SWE-rebench: Claude Tops at 65.3%

The February 2026 SWE-rebench leaderboard update features 57 fresh GitHub PR tasks, with Claude Opus 4.6 leading at 65.3% resolved rate. Top models like GPT-5.2-medium (64.4%), GLM-5 and GPT-5.4-medium (both 62.8%) form a tight race. Open-weight models such as Qwen3.5-397B (59.9%) are rapidly closing the gap.

Reddit r/LocalLLaMACommunityMar 23#leaderboard#swe-bench#coding-benchmark
🔬

IDP Leaderboard Benchmarks 16 VLMs

New open IDP Leaderboard evaluates 16 VLMs on 9,000+ documents across three benchmarks: OlmOCR, OmniDoc, and IDP Core. Gemini 3.1 Pro leads narrowly; cheaper variants match flagships except on reasoning tasks. Features a Results Explorer for predictions vs. ground truth.

Reddit r/MachineLearningCommunityMar 11#vlm-benchmark#document-ai#leaderboard
Page 2 of 3