VAMPS: New Benchmark for Visual-Assisted Mathematical Problem Solving

๐กDiscover why top LLMs fail to use visual tools effectively for math and how to benchmark your model's reasoning.
โก 30-Second TL;DR
What Changed
Introduces 1,168 multimodal, bilingual QA pairs based on Iranian University Entrance Exam problems.
Why It Matters
This benchmark highlights a critical bottleneck in agentic AI workflows where models fail to bridge the gap between tool usage and reasoning. It provides researchers with a standardized way to measure and improve the 'visual grounding' capabilities of future multimodal systems.
What To Do Next
Evaluate your current multimodal agent's performance on the VAMPS benchmark to identify if it struggles with visual tool integration.
Key Points
- โขIntroduces 1,168 multimodal, bilingual QA pairs based on Iranian University Entrance Exam problems.
- โขTests model capability in constructing and grounding answers in visual plots (intersections, asymptotes, etc.).
- โขFinds that direct analytical solving currently outperforms tool-enabled visual reasoning across most models.
- โขProvides a diagnostic framework for evaluating how models handle external tool outputs.
๐ง Deep Insight
Web-grounded analysis with 13 cited sources.
๐ Enhanced Key Takeaways
- โขVAMPS uniquely focuses on evaluating LLMs' capability to actively construct and then reason over tool-generated visual plots, distinguishing it from prior multimodal benchmarks that primarily assess reasoning over pre-existing, fixed visual inputs.
- โขThe benchmark's problems are sourced from the Iranian University Entrance Exams, with 218 real problems translated into English and expanded with human-reviewed, LLM-generated synthetic variants, ensuring a diverse and challenging dataset.
- โขA significant finding from VAMPS and related research (e.g., VisAidMath) is that current models frequently exhibit 'hallucination regarding the implicit visual reasoning process,' struggling to accurately interpret and integrate visual information, even when provided with explicit visual aids.
- โขVAMPS serves as both a benchmarking tool and a diagnostic framework, designed to not only measure performance but also to identify and analyze the specific reasons why models fail to effectively utilize external tool outputs for visual reasoning.
๐ Competitor Analysisโธ Show
| Benchmark | Focus / Key Feature | Dataset Size | Problem Source/Level | Multilinguality | Key Findings (if applicable) |
|---|---|---|---|---|---|
| VAMPS | Tool-enabled visual reasoning (constructing & grounding plots) | 1,168 QA pairs | Iranian University Entrance Exam (Algebra & Calculus) | Bilingual (Persian/English) | Direct analytical solving outperforms tool-enabled visual solving. |
| MathVista | Comprehensive math reasoning in diverse visual contexts | 6,141 examples | 28 existing + 3 new datasets (IQTest, FunctionQA, PaperQA) | Not specified | Systematically studies math reasoning in visual contexts. |
| U-MATH | University-level mathematical thinking, includes a meta-benchmark (ฮผ-MATH) for judging solutions | 1,100 problems (20% visual) | Curriculum from top US universities | Not specified | Challenging for current LLMs; Gemini 2.0 Flash Thinking achieved ~73.6% accuracy. |
| VisAidMath | Evaluating visual-aided mathematical reasoning, explicit and implicit visual contexts | 1,200 problems | Textbooks, exams, Olympiads | Not specified | GPT-4V achieved 45.33% accuracy, highlighting hallucination in visual reasoning. |
| VC-Bench | Explicit visual dependency in multimodal mathematical reasoning, multi-image tasks | 1,720 problems (6,697 images) | Six cognitive domains | Not specified | Top models unable to exceed 50% accuracy, highlighting challenges in visual-mathematical integration. |
| MATHNET | Olympiad-level math reasoning and retrieval | 30,676 problems | 47 countries, 2 decades of competitions | Multilingual (17 languages) | SOTA models (GPT-5, Gemini 2.5 Pro) challenged (72%, 66% accuracy respectively). |
๐ ๏ธ Technical Deep Dive
- Dataset Composition: VAMPS comprises 1,168 multimodal, bilingual (Persian and English) multiple-choice question-answer pairs.
- Problem Sourcing: The core of the benchmark consists of 218 real problems from Iranian University Entrance Exams, each provided in Persian and with a manually checked English translation, resulting in 436 original question instances.
- Data Augmentation: This core is extended with synthetic multimodal variants, generated with LLM assistance from the real question seeds and subsequently reviewed by humans to ensure quality and relevance.
- Problem Characteristics: Problems are specifically chosen where plotting offers a natural solution strategy, involving concepts like intersections, extrema, and asymptotes, to effectively test visual reasoning.
- Evaluation Focus: The benchmark is designed to assess whether models can successfully construct a useful graph from problem descriptions and then accurately ground their answers in the resulting visualization.
- Diagnostic Framework: VAMPS includes a diagnostic framework aimed at evaluating how well models process and integrate outputs from external visualization tools.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (13)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI โ
