Search

10 results on this page

New Benchmark Tests Whether AI Can Draw Geometry

New Benchmark Tests Whether AI Can Draw Geometry

Researchers introduce an open-source benchmark separating olympiad geometry solving from accurate diagram construction. Across 954 problems, current foundation models achieved only a 36.14% average diagram compilation success rate, revealing a substantial gap between mathematical reasoning and faithful visual construction.

ArXiv AIResearch23h ago#geometry-reasoning#benchmark#asymptote
A Smaller Qwen3.8 Built by Pruning Layers

A Smaller Qwen3.8 Built by Pruning Layers

A community developer created Qwen3.8-23B-Mini-Me by strategically removing layers from Qwen3.8-27B, reducing the model to approximately 22.7B parameters without severe reasoning degradation. The model is reported to work well for coding, agentic tasks, and multi-turn chats, but it has not yet been benchmarked and struggles more with edge cases and underspecified prompts.

Reddit r/LocalLLaMACommunity1d ago#model-pruning#model-compression#apple-silicon
Reasoning Agents May Collude in Markets

Reasoning Agents May Collude in Markets

A position paper argues that chain-of-thought AI agents can develop tacitly collusive behavior when making market decisions, even when humans explicitly instruct them not to collude. Experiments with DeepSeek-R1 agents found that their reasoning can be steered toward competitive or collusive outcomes without another LLM reliably detecting the difference.

ArXiv AIResearch23h ago#agent-safety#market-governance
Page 1