來源較早收集於 11h

ReactBench:MLLMs 拓撲推理基準

ReactBench:MLLMs 拓撲推理基準
PostLinkedIn
📄閱讀原文: ArXiv AI
#benchmark#chemical-diagrams#structural-reasoningreactbenchreactbenchmllms

💡新基準揭露 MLLMs 圖形拓撲推理 30% 差距—多模態研究關鍵。(38字)

⚡ 30 秒速覽

有什麼變化

全新基準包含 1,618 個化學圖 QA 對

為什麼重要

突顯 MLLMs 在複雜圖形結構推理的根本限制,促使針對性改進。為科學應用視覺拓撲理解進展建立標準。

下一步行動

從 arXiv 下載 ReactBench 資料集,並基準測試你的 MLLM 於其 QA 任務。

誰應關注:Researchers & Academics

關鍵要點

  • 全新基準包含 1,618 個化學圖 QA 對
  • 測試多樣拓撲:線性鏈到循環圖
  • 17 個 MLLMs 在整體任務落後錨點任務 >30%
  • 消融證實推理缺陷而非感知問題

🧠 深度解析

本篇為 AI 生成分析,非原文內容。

🔑 增強重點摘要

  • ReactBench utilizes a specialized 'Topological Reasoning' framework that specifically targets the model's ability to map visual graph connectivity to chemical nomenclature, distinguishing it from standard OCR-based chemical benchmarks.
  • The benchmark incorporates a 'distractor-robustness' evaluation, where models are tested against diagrams containing extraneous chemical noise to isolate topological reasoning from general visual attention.
  • The 30% performance gap identified is specifically attributed to 'graph-traversal failure' in MLLMs, where models struggle to maintain state consistency when navigating complex, multi-ring fused structures.
📊 競品分析▸ Show
BenchmarkFocus AreaPrimary MetricData Modality
ReactBenchTopological/Chemical ReasoningStructural AccuracyChemical Diagrams
ChemBenchGeneral Chemical KnowledgeMultiple Choice AccuracyText/SMILES
SciBenchScientific Problem SolvingReasoning Chain AccuracyText/Diagrams

🛠️ 技術深入

  • Dataset Construction: 1,618 expert-annotated pairs derived from curated chemical reaction databases, ensuring ground-truth topological validity.
  • Task Hierarchy: Four levels of complexity: (1) Node identification, (2) Edge connectivity, (3) Sub-structure recognition, (4) Holistic reaction pathway reasoning.
  • Evaluation Protocol: Employs a multi-stage prompting strategy to decouple visual perception (object detection) from logical reasoning (graph traversal).
  • Ablation Methodology: Uses 'Perception-Masked' inputs to prove that even with perfect object detection, models fail to correctly infer the global topological structure.

🔮 前景展望基於引用來源的 AI 分析

Future MLLM architectures will prioritize graph-aware attention mechanisms.
The identified reasoning bottleneck suggests that standard transformer architectures lack the inductive bias necessary for complex topological graph traversal.
Chemical reasoning benchmarks will shift from text-based SMILES strings to pure visual-topological tasks.
The performance gap highlighted by ReactBench demonstrates that current models rely too heavily on text-based training data rather than true visual understanding of chemical structures.

時間線

2025-11
Initial curation of chemical reaction diagrams for ReactBench dataset.
2026-02
Completion of expert-annotation phase for 1,618 QA pairs.
2026-04
Public release of ReactBench on ArXiv and associated evaluation framework.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。