Qwen 3.5 MoE 35B 指令模式詢問
Reddit 用戶因辦公室頻寬限制,尋求 Qwen 3.5 MoE 35B 不含推理鏈的指令模式測試結果。對 Qwen 在 2507 發布後重返混合推理模型感到驚訝。社群討論持續中。
Tag: #benchmark195 results
Reddit 用戶因辦公室頻寬限制,尋求 Qwen 3.5 MoE 35B 不含推理鏈的指令模式測試結果。對 Qwen 在 2507 發布後重返混合推理模型感到驚訝。社群討論持續中。

ByteDance's Seed 2.0 debuts at #6 text, #3 vision on LMArena, leading domestic models. Excels in math, vision perception, reasoning, and agents, matching Gemini 3 Pro. Native multimodal upgrades drive benchmark dominance.
BrowseComp-V³ is a new benchmark with 300 challenging questions for evaluating multimodal browsing agents on deep multi-hop reasoning across text and visuals. It features subgoal-driven process evaluation and publicly searchable evidence for reproducibility. Experiments reveal state-of-the-art models achieve only 36% accuracy, highlighting integration bottlenecks.

Ohio State and Amazon release MMDR-Bench, a verifiable benchmark for multimodal Deep Research Agents. Focuses on process traceability, evidence alignment, and claim verification beyond superficial reports. Open resources include paper, GitHub, and Hugging Face datasets.
AgentLeak introduces the first full-stack benchmark for privacy leakage in multi-agent LLM systems, covering internal channels like inter-agent messages. It spans 1,000 scenarios across healthcare, finance, legal, and corporate domains. Tests on top models show internal channels cause 68.9% total leakage, missed by output audits.