Search

Tag: #llm90 results

BotzoneBench: Scalable LLM Game Eval Benchmark

BotzoneBench: Scalable LLM Game Eval Benchmark

BotzoneBench introduces a scalable framework for evaluating LLMs' strategic reasoning in interactive games using fixed hierarchies of skill-calibrated game AIs. It assesses five flagship models across eight diverse games via 177,047 state-action pairs, revealing performance gaps and behaviors comparable to mid-tier game AIs. This enables linear-time absolute measurements with stable interpretability, unlike volatile LLM-vs-LLM rankings.

ArXiv AIResearchFeb 17#research#botzonebench#llm
397B Qwen 3.5 Beats Gemini 3

397B Qwen 3.5 Beats Gemini 3

Alibaba launches Qwen 3.5, strongest open-source model with 397B parameters on除夕. It surpasses Gemini 3 in benchmarks. Inference costs just 0.8 yuan per million tokens.

量子位MediaFeb 16#launch#alibaba#llm
Qwen3.5-Plus Breaks Cost-Performance Ceiling

Qwen3.5-Plus Breaks Cost-Performance Ceiling

Alibaba's Qwen3.5-Plus launches as top open-source model in multimodal, reasoning, coding, and agents. Priced at 0.8 yuan per million tokens, it's 18x cheaper than Gemini 3 Pro. Achieves superior performance with efficient architecture.

机器之心MediaFeb 16#launch#qwen#35-plus
Ant Open-Sources 1T Ling-2.5 Instant Model

Ant Open-Sources 1T Ling-2.5 Instant Model

Ant Group open-sourced Ling-2.5-1T, a trillion-parameter instant model with 63B active parameters, trained on 29T tokens. It upgrades architecture for 1M token context, boosts token efficiency, and enhances preference alignment and agent interactions. The model outperforms predecessors and rivals large instant models in reasoning and instruction following.

IT之家MediaFeb 16#launch#ant-group#ling-25-1t
Apple's Async Verified Semantic Caching for LLMs

Apple's Async Verified Semantic Caching for LLMs

Apple introduces asynchronous verified semantic caching to optimize tiered LLM architectures. It addresses tradeoffs in static and dynamic caches using embedding similarity thresholds. This reduces inference cost and latency in production workflows like search and agents.

Apple Machine LearningOfficialFeb 16#research#apple-ml#llm
JD Open-Sources 48B MoE JoyAI-Flash Model

JD Open-Sources 48B MoE JoyAI-Flash Model

JD.com open-sourced JoyAI-LLM-Flash, a 48B total parameter MoE model with 3B active params, pre-trained on 20T tokens. It excels in knowledge, reasoning, coding, and agents using FiberPO framework and Muon optimizer. Features 1.3x-1.7x throughput gains via dense MTP.

IT之家MediaFeb 15#launch#jd-com#joyai-llm-flash
Page 7 of 9