📄ArXiv AI•Stalecollected in 40m
CreativityBench: LLM Creative Tool Benchmark

💡New benchmark shows LLMs fail creative tool repurposing—benchmark yours now!
⚡ 30-Second TL;DR
What Changed
Introduces CreativityBench for affordance-based creativity in LLMs
Why It Matters
Reveals critical gaps in LLM creative tool use, essential for agent development. Offers a testbed to advance planning and reasoning in future AI systems.
What To Do Next
Download CreativityBench dataset from arXiv:2605.02910v2 and evaluate your LLM agent.
Who should care:Researchers & Academics
Key Points
- •Introduces CreativityBench for affordance-based creativity in LLMs
- •Builds 4K-entity, 150K+ annotation affordance knowledge base
- •Generates 14K grounded tasks requiring non-obvious solutions
- •10 SOTA LLMs fail on parts, affordances, and physical mechanisms
- •Scaling saturates; CoT yields limited gains
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •CreativityBench utilizes a hierarchical taxonomy of physical objects, mapping specific sub-components to their functional affordances to test cross-domain reasoning.
- •The benchmark specifically targets the 'functional fixedness' bias in LLMs, where models struggle to decouple an object's primary utility from its potential secondary uses.
- •Evaluation metrics include a novel 'repurposing success rate' that penalizes models for hallucinating non-existent physical mechanisms or impossible material interactions.
🛠️ Technical Deep Dive
- •Knowledge Base Structure: Employs a graph-based representation where nodes are physical entities and edges represent affordance relationships (e.g., 'can-be-used-as', 'has-part').
- •Task Generation: Utilizes a constrained generation pipeline that forces models to propose solutions within a defined set of physical laws and material properties.
- •Evaluation Framework: Implements a multi-stage verification process involving both automated semantic similarity checks against ground-truth affordances and a secondary LLM-based 'physical feasibility' judge.
🔮 Future ImplicationsAI analysis grounded in cited sources
Future LLM training will shift toward embodied physical simulation data.
The failure of scaling and CoT suggests that abstract linguistic reasoning is insufficient for physical creativity without grounded world-model training.
CreativityBench will become a standard metric for evaluating agentic reasoning.
As LLMs transition into autonomous agents, the ability to repurpose tools in novel environments will be a critical safety and performance requirement.
⏳ Timeline
2025-11
Initial development of the affordance knowledge base and taxonomy.
2026-02
Completion of the 14K grounded task generation pipeline.
2026-04
Final evaluation of 10 SOTA LLMs and submission to ArXiv.
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗