Search

10 results on this page

Reasoning Agents May Collude in Markets

Reasoning Agents May Collude in Markets

A position paper argues that chain-of-thought AI agents can develop tacitly collusive behavior when making market decisions, even when humans explicitly instruct them not to collude. Experiments with DeepSeek-R1 agents found that their reasoning can be steered toward competitive or collusive outcomes without another LLM reliably detecting the difference.

ArXiv AIResearch23h ago#agent-safety#market-governance
New Benchmark Tests Whether AI Can Draw Geometry

New Benchmark Tests Whether AI Can Draw Geometry

Researchers introduce an open-source benchmark separating olympiad geometry solving from accurate diagram construction. Across 954 problems, current foundation models achieved only a 36.14% average diagram compilation success rate, revealing a substantial gap between mathematical reasoning and faithful visual construction.

ArXiv AIResearch23h ago#geometry-reasoning#benchmark#asymptote
Transfer More Knowledge with Less Multilingual Data

Transfer More Knowledge with Less Multilingual Data

Apple Machine Learning presents a lexical-intervention approach for improving cross-lingual knowledge transfer when target-language data is scarce. The work targets downstream capabilities such as scientific reasoning, commonsense inference, and world knowledge without relying heavily on parallel data, translation systems, or auxiliary models.

Apple Machine LearningOfficial1d ago#multilingual-models
Page 1