
HappyHorse-1.0 Tops AI Video Arena at 1383 Elo
Video generation model HappyHorse-1.0 ranks No.1 on Artificial Analysis’ AI Video Arena. It achieved an Elo score of 1383. Developer and technical details are undisclosed.
Tag: #leaderboard25 results

Video generation model HappyHorse-1.0 ranks No.1 on Artificial Analysis’ AI Video Arena. It achieved an Elo score of 1383. Developer and technical details are undisclosed.

Alibaba's Qwen 3.6 Plus has secured the top position in the global large model invocation weekly leaderboard. A more powerful flagship model, Qwen 3.6 Max, is slated for upcoming release.

Alibaba's Qwen 3.6 Plus has shattered records with over 1.4 trillion daily token invocations. It now leads the global model usage leaderboard. This highlights its massive adoption and performance at scale.

The February 2026 SWE-rebench leaderboard update features 57 fresh GitHub PR tasks, with Claude Opus 4.6 leading at 65.3% resolved rate. Top models like GPT-5.2-medium (64.4%), GLM-5 and GPT-5.4-medium (both 62.8%) form a tight race. Open-weight models such as Qwen3.5-397B (59.9%) are rapidly closing the gap.

Arena, formerly LM Arena, has become the de facto public leaderboard for frontier LLMs, influencing funding, launches, and PR cycles. Funded by the companies it ranks, the startup evolved from a UC Berkeley PhD research project in just seven months.

Arena, formerly LM Arena, has become the de facto public leaderboard for frontier LLMs, founded by UC Berkeley PhD students. In just seven months, it evolved from a research project into a startup influencing AI funding, launches, and PR cycles amid fierce model competition.
ColQwen3.5-4.5B-v3 leads MTEB ViDoRe leaderboard at 75.67 mean, using half params and 13x fewer embedding dims than prior #1. Full eval trail public; supported by colpali-engine and vLLM on ROCm/CUDA. Apache 2.0, with 9B variant training.
New open IDP Leaderboard evaluates 16 VLMs on 9,000+ documents across three benchmarks: OlmOCR, OmniDoc, and IDP Core. Gemini 3.1 Pro leads narrowly; cheaper variants match flagships except on reasoning tasks. Features a Results Explorer for predictions vs. ground truth.
階躍星辰的Step3.5 Flash模型根據OpenRouter數據,已連續三天在OpenClaw上保持全球調用量第一。2026年3月以來,Kimi K2.5、Step3.5 Flash與MiniMax M2.5的調用量分居全球前三。

An independent developer known as yuxinlu1 has achieved a top ranking on the Hugging Face model leaderboard. This feat highlights the growing competitiveness of individual contributors against major tech corporations.