Search

10 results on this page

New Benchmark Tests Whether AI Can Draw Geometry

New Benchmark Tests Whether AI Can Draw Geometry

Researchers introduce an open-source benchmark separating olympiad geometry solving from accurate diagram construction. Across 954 problems, current foundation models achieved only a 36.14% average diagram compilation success rate, revealing a substantial gap between mathematical reasoning and faithful visual construction.

ArXiv AIResearch20h ago#geometry-reasoning#benchmark#asymptote
Claude Enters Protein Design

Claude Enters Protein Design

Claude is expanding into protein design, adding a scientific research capability to its existing AI portfolio. The update suggests potential use in computational biology and biotech workflows, though the article excerpt does not provide implementation or benchmark details.

The NeuronMedia14h ago#protein-design#scientific-ai
A New Complexity Scorecard for Game World Models

A New Complexity Scorecard for Game World Models

The paper proposes Transition Complexity Profile (TCP), a reproducible framework for measuring how difficult game-world transition prediction is at a specified interface. It evaluates branching, interaction-driven uncertainty, opponent influence, and temporal or spatial dependencies to improve comparisons across game-modeling and reinforcement-learning benchmarks.

ArXiv AIResearch20h ago#game-world-modeling#benchmarking
Page 1