🧠較早收集於 1m

多模態DeepResearch達到SOTA

多模態DeepResearch達到SOTA
PostLinkedIn
🧠閱讀原文: 机器之心

💡SOTA open multimodal research agent crushes benchmarks with tiny params vs closed-source

⚡ 30-Second TL;DR

有什麼變化

建構多模態代理用於文字+圖像深度研究真實世界搜尋

為什麼重要

此進展多模態代理超越純文字,能可靠研究視覺證據如照片與圖表。降低複雜查詢幻覺風險,接近人類驗證方式。開放方法可民主化高性能研究工具。

下一步行動

Check Hugging Face daily papers for the multimodal DeepResearch model and replicate its 6 benchmarks.

誰應關注:Researchers & Academics

關鍵要點

  • 建構多模態代理用於文字+圖像深度研究真實世界搜尋
  • 以多尺度裁剪及多實體檢索修復引擎命中率
  • 合成VQA資料與軌跡,經SFT+RL訓練內化能力
  • 提升推理至數十輪、互動至數百次
  • 以較小參數規模在6基準達SOTA

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 7 個來源。

🔑 增強重點摘要

  • MMDR-Bench is the first end-to-end benchmark for multimodal deep research agents, featuring 140 expert-crafted tasks across 21 domains in Daily and Research regimes to test report generation with image-text bundles.[1]
  • The model was evaluated alongside 25 state-of-the-art LLMs and DRAs on MMDR-Bench, revealing trade-offs in writing quality, citation faithfulness, and multimodal grounding.[1]
  • MMDR-Bench includes a unified evaluation pipeline assessing report quality (FLAE), citation-grounded faithfulness (TRACE), and text-visual evidence consistency (MOSAIC).[1]

🔮 前景展望AI analysis grounded in cited sources

Multimodal DRAs will close the 44.4% human-model gap on MMDR-Bench by 2027
Current top models like Gemini3-Pro-Preview score 49.7% versus human 94.1%, but SOTA advancements in benchmarks like MMDR-Bench drive rapid progress in multimodal reasoning.[1]
Compact multimodal models under 10B parameters will lead academic research benchmarks by 2027
Recent 10B models achieve 94.43% on AIME2025 and top STEM/OCR tasks, indicating efficiency gains outpace larger competitors.[3]

時間線

2026-01
MMDR-Bench introduced as first multimodal deep research benchmark with 140 tasks across 21 domains.[1]
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。