Search

10 results on this page

New Benchmark Tests Whether AI Can Draw Geometry

New Benchmark Tests Whether AI Can Draw Geometry

Researchers introduce an open-source benchmark separating olympiad geometry solving from accurate diagram construction. Across 954 problems, current foundation models achieved only a 36.14% average diagram compilation success rate, revealing a substantial gap between mathematical reasoning and faithful visual construction.

ArXiv AIResearch22h ago#geometry-reasoning#benchmark#asymptote
A New Complexity Scorecard for Game World Models

A New Complexity Scorecard for Game World Models

The paper proposes Transition Complexity Profile (TCP), a reproducible framework for measuring how difficult game-world transition prediction is at a specified interface. It evaluates branching, interaction-driven uncertainty, opponent influence, and temporal or spatial dependencies to improve comparisons across game-modeling and reinforcement-learning benchmarks.

ArXiv AIResearch22h ago#game-world-modeling#benchmarking
Transfer More Knowledge with Less Multilingual Data

Transfer More Knowledge with Less Multilingual Data

Apple Machine Learning presents a lexical-intervention approach for improving cross-lingual knowledge transfer when target-language data is scarce. The work targets downstream capabilities such as scientific reasoning, commonsense inference, and world knowledge without relying heavily on parallel data, translation systems, or auxiliary models.

Apple Machine LearningOfficial1d ago#multilingual-models
Page 1