🧠机器之心•較早收集於 1m
多模態DeepResearch達到SOTA

💡SOTA open multimodal research agent crushes benchmarks with tiny params vs closed-source
⚡ 30-Second TL;DR
有什麼變化
建構多模態代理用於文字+圖像深度研究真實世界搜尋
為什麼重要
此進展多模態代理超越純文字,能可靠研究視覺證據如照片與圖表。降低複雜查詢幻覺風險,接近人類驗證方式。開放方法可民主化高性能研究工具。
下一步行動
Check Hugging Face daily papers for the multimodal DeepResearch model and replicate its 6 benchmarks.
誰應關注:Researchers & Academics
關鍵要點
- •建構多模態代理用於文字+圖像深度研究真實世界搜尋
- •以多尺度裁剪及多實體檢索修復引擎命中率
- •合成VQA資料與軌跡,經SFT+RL訓練內化能力
- •提升推理至數十輪、互動至數百次
- •以較小參數規模在6基準達SOTA
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 7 個來源。
🔑 增強重點摘要
- •MMDR-Bench is the first end-to-end benchmark for multimodal deep research agents, featuring 140 expert-crafted tasks across 21 domains in Daily and Research regimes to test report generation with image-text bundles.[1]
- •The model was evaluated alongside 25 state-of-the-art LLMs and DRAs on MMDR-Bench, revealing trade-offs in writing quality, citation faithfulness, and multimodal grounding.[1]
- •MMDR-Bench includes a unified evaluation pipeline assessing report quality (FLAE), citation-grounded faithfulness (TRACE), and text-visual evidence consistency (MOSAIC).[1]
🔮 前景展望AI analysis grounded in cited sources
Multimodal DRAs will close the 44.4% human-model gap on MMDR-Bench by 2027
Current top models like Gemini3-Pro-Preview score 49.7% versus human 94.1%, but SOTA advancements in benchmarks like MMDR-Bench drive rapid progress in multimodal reasoning.[1]
Compact multimodal models under 10B parameters will lead academic research benchmarks by 2027
Recent 10B models achieve 94.43% on AIME2025 and top STEM/OCR tasks, indicating efficiency gains outpace larger competitors.[3]
⏳ 時間線
2026-01
MMDR-Bench introduced as first multimodal deep research benchmark with 140 tasks across 21 domains.[1]
📎 來源 (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心 ↗
每週 AI 簡報
每週一封,可隨時退訂。