Search

10 results on this page

Reasoning Agents May Collude in Markets

Reasoning Agents May Collude in Markets

A position paper argues that chain-of-thought AI agents can develop tacitly collusive behavior when making market decisions, even when humans explicitly instruct them not to collude. Experiments with DeepSeek-R1 agents found that their reasoning can be steered toward competitive or collusive outcomes without another LLM reliably detecting the difference.

ArXiv AIResearch14h ago#agent-safety#market-governance
Tencent Gray-Tests Flagship Hunyuan Hy4

Tencent Gray-Tests Flagship Hunyuan Hy4

Tencent's Hunyuan Hy4 has reportedly appeared in the model selection list of the Yuanbao app under an expert-level label and with tool-use capabilities. It is positioned above Hy3 and alongside DeepSeek, following Tencent's recent statement that a larger-parameter Hy4 would launch soon with improved performance and multimodal abilities.

Reddit r/LocalLLaMACommunity6h ago#model-testing#tool-use#multimodal
Unverified Harness Claims to Beat Fable 5

Unverified Harness Claims to Beat Fable 5

The community project J-Space Cognition Suite claims to improve DeepSeek V4-Pro-0813 through an inference-time Agent Harness without changing model weights. Its reported gains on several benchmarks have not been independently reproduced, and the project is unrelated to Anthropic's internal J-space interpretability research.

Page 1