來源ArXiv AI•較早收集於 11h
ReactBench:MLLMs 拓撲推理基準

#benchmark#chemical-diagrams#structural-reasoningreactbenchreactbenchmllms
💡新基準揭露 MLLMs 圖形拓撲推理 30% 差距—多模態研究關鍵。(38字)
⚡ 30 秒速覽
有什麼變化
全新基準包含 1,618 個化學圖 QA 對
為什麼重要
突顯 MLLMs 在複雜圖形結構推理的根本限制,促使針對性改進。為科學應用視覺拓撲理解進展建立標準。
下一步行動
從 arXiv 下載 ReactBench 資料集,並基準測試你的 MLLM 於其 QA 任務。
誰應關注:Researchers & Academics
關鍵要點
- •全新基準包含 1,618 個化學圖 QA 對
- •測試多樣拓撲:線性鏈到循環圖
- •17 個 MLLMs 在整體任務落後錨點任務 >30%
- •消融證實推理缺陷而非感知問題
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •ReactBench utilizes a specialized 'Topological Reasoning' framework that specifically targets the model's ability to map visual graph connectivity to chemical nomenclature, distinguishing it from standard OCR-based chemical benchmarks.
- •The benchmark incorporates a 'distractor-robustness' evaluation, where models are tested against diagrams containing extraneous chemical noise to isolate topological reasoning from general visual attention.
- •The 30% performance gap identified is specifically attributed to 'graph-traversal failure' in MLLMs, where models struggle to maintain state consistency when navigating complex, multi-ring fused structures.
📊 競品分析▸ Show
| Benchmark | Focus Area | Primary Metric | Data Modality |
|---|---|---|---|
| ReactBench | Topological/Chemical Reasoning | Structural Accuracy | Chemical Diagrams |
| ChemBench | General Chemical Knowledge | Multiple Choice Accuracy | Text/SMILES |
| SciBench | Scientific Problem Solving | Reasoning Chain Accuracy | Text/Diagrams |
🛠️ 技術深入
- •Dataset Construction: 1,618 expert-annotated pairs derived from curated chemical reaction databases, ensuring ground-truth topological validity.
- •Task Hierarchy: Four levels of complexity: (1) Node identification, (2) Edge connectivity, (3) Sub-structure recognition, (4) Holistic reaction pathway reasoning.
- •Evaluation Protocol: Employs a multi-stage prompting strategy to decouple visual perception (object detection) from logical reasoning (graph traversal).
- •Ablation Methodology: Uses 'Perception-Masked' inputs to prove that even with perfect object detection, models fail to correctly infer the global topological structure.
🔮 前景展望基於引用來源的 AI 分析
Future MLLM architectures will prioritize graph-aware attention mechanisms.
The identified reasoning bottleneck suggests that standard transformer architectures lack the inductive bias necessary for complex topological graph traversal.
Chemical reasoning benchmarks will shift from text-based SMILES strings to pure visual-topological tasks.
The performance gap highlighted by ReactBench demonstrates that current models rely too heavily on text-based training data rather than true visual understanding of chemical structures.
⏳ 時間線
2025-11
Initial curation of chemical reaction diagrams for ReactBench dataset.
2026-02
Completion of expert-annotation phase for 1,618 QA pairs.
2026-04
Public release of ReactBench on ArXiv and associated evaluation framework.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週電子報
每週一封,可隨時退訂。