CrowdMath: New Dataset for Collaborative Mathematical Reasoning

๐กNew benchmark reveals LLMs struggle with collaborative reasoning, despite high accuracy in simple next-token prediction.
โก 30-Second TL;DR
What Changed
Contains 164 expert-annotated progress chains from the MIT PRIMES-AoPS program.
Why It Matters
This dataset exposes a significant limitation in current LLMs regarding their ability to understand the functional significance of collaborative contributions, providing a new benchmark for future reasoning research.
What To Do Next
Download the CrowdMath dataset from the arXiv link to test your model's ability to identify the functional role of individual contributions in a multi-step proof process.
Key Points
- โขContains 164 expert-annotated progress chains from the MIT PRIMES-AoPS program.
- โขModels show 83-88% accuracy in next-post prediction but struggle with functional role classification.
- โขHighlights the difficulty of modeling collaborative, multi-participant mathematical reasoning.
- โขProvides a benchmark for evaluating how models identify errors and synthesize proofs.
๐ง Deep Insight
Web-grounded analysis with 11 cited sources.
๐ Enhanced Key Takeaways
- โขCrowdMath is an initiative by MIT PRIMES and Art of Problem Solving (AoPS), inspired by Terry Tao's Polymath projects, designed to provide high school and college students with collaborative research experience in mathematics.
- โขThe dataset captures the iterative and often non-linear nature of mathematical discovery, where participants contribute ideas, build upon others' work, and collectively identify errors, reflecting a more realistic research process than traditional benchmarks.
- โขPrevious CrowdMath projects have led to published research papers, such as 'The Broken Stick Project' (2018) and 'Results on Pattern Avoidance Games' (2017), demonstrating the program's success in fostering genuine mathematical contributions from students.
- โขThe online, asynchronous nature of CrowdMath, with mentors guiding dozens of students without fixed meetings, significantly lowers logistical barriers and costs compared to traditional research programs, making advanced mathematical research more accessible.
๐ Competitor Analysisโธ Show
| Feature/Dataset | CrowdMath (2026) | MathNet (2026) | MATH (2020) | IMProofBench (2025) | ProofBench (2026) |
|---|---|---|---|---|---|
| Primary Focus | Collaborative mathematical reasoning discussions | Olympiad-level problem-solving & retrieval | Competition problem-solving | Research-level proof writing | Formal proof generation & verification |
| Data Source | Expert-annotated student/mentor discussions (MIT PRIMES-AoPS) | Expert-authored competition problems/solutions (58 countries) | Competition problems with step-by-step solutions | Expert-reviewed research problems | Lean 4 formal mathematics |
| Scale | 164 discussion chains | 30,000+ problems/solutions | 12,500 problems | 39 peer-reviewed problems | Not specified, Lean 4 tasks |
| Modality | Text-only (discussions) | Multilingual, multimodal (text + images) | Text-only | Text-only (with tool use) | Text-only (Lean 4 code) |
| Key Challenge | Modeling collaborative dynamics, functional role classification | Olympic-level reasoning, semantic retrieval | Complex problem-solving, step-by-step reasoning | Generating detailed, research-level proofs | Translating informal to machine-checkable proofs |
| Benchmarks | Next-post prediction (83-88% accuracy), functional role classification (struggles) | Solving math problems, mathematical semantic retrieval, RAG | Problem-solving accuracy | Proof generation, final-answer subproblems (Grok-4: 52%, GPT-5: 22% for full proof) | Lean 4 proof compilation/correctness (Aristotle: 71%, GPT 5.4: 56%) |
๐ ๏ธ Technical Deep Dive
- Dataset Structure: Comprises 164 expert-annotated chains of collaborative mathematical research discussions.
- Data Source: Derived from the MIT PRIMES-AoPS program, which involves high school and college students collaborating online with mentors on open mathematical problems.
- Annotation: The discussions are 'expert-annotated,' implying human experts have labeled elements within the conversational turns.
- Evaluation Tasks: The dataset is used to benchmark models on tasks such as next-post prediction and functional role classification within the collaborative discussion.
- Model Performance: Models achieve 83-88% accuracy in next-post prediction but show difficulty with functional role classification, indicating a challenge in understanding the nuanced contributions of participants in a collaborative setting.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (11)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI โ
