Search

Tag: #benchmark195 results

Seed 2.0 Tops Arena for Chinese Models

Seed 2.0 Tops Arena for Chinese Models

ByteDance's Seed 2.0 debuts at #6 text, #3 vision on LMArena, leading domestic models. Excels in math, vision perception, reasoning, and agents, matching Gemini 3 Pro. Native multimodal upgrades drive benchmark dominance.

机器之心MediaFeb 16#update#seed#multimodal
BrowseComp-V³ Benchmark for Multimodal Agents

BrowseComp-V³ Benchmark for Multimodal Agents

BrowseComp-V³ is a new benchmark with 300 challenging questions for evaluating multimodal browsing agents on deep multi-hop reasoning across text and visuals. It features subgoal-driven process evaluation and publicly searchable evidence for reproducibility. Experiments reveal state-of-the-art models achieve only 36% accuracy, highlighting integration bottlenecks.

ArXiv AIResearchFeb 16#research#browsecomp#multimodal-ai
MMDR-Bench Verifies Multimodal Research

MMDR-Bench Verifies Multimodal Research

Ohio State and Amazon release MMDR-Bench, a verifiable benchmark for multimodal Deep Research Agents. Focuses on process traceability, evidence alignment, and claim verification beyond superficial reports. Open resources include paper, GitHub, and Hugging Face datasets.

机器之心MediaFeb 14#research#mmdr-bench#benchmark
AgentLeak: Multi-Agent Privacy Leak Benchmark

AgentLeak: Multi-Agent Privacy Leak Benchmark

AgentLeak introduces the first full-stack benchmark for privacy leakage in multi-agent LLM systems, covering internal channels like inter-agent messages. It spans 1,000 scenarios across healthcare, finance, legal, and corporate domains. Tests on top models show internal channels cause 68.9% total leakage, missed by output audits.

ArXiv AIResearchFeb 13#research#agentleak#multi-agent
Page 20 of 20