Search

10 results on this page

Tencent Gray-Tests Flagship Hunyuan Hy4

Tencent Gray-Tests Flagship Hunyuan Hy4

Tencent's Hunyuan Hy4 has reportedly appeared in the model selection list of the Yuanbao app under an expert-level label and with tool-use capabilities. It is positioned above Hy3 and alongside DeepSeek, following Tencent's recent statement that a larger-parameter Hy4 would launch soon with improved performance and multimodal abilities.

Reddit r/LocalLLaMACommunity8h ago#model-testing#tool-use#multimodal
A New Complexity Scorecard for Game World Models

A New Complexity Scorecard for Game World Models

The paper proposes Transition Complexity Profile (TCP), a reproducible framework for measuring how difficult game-world transition prediction is at a specified interface. It evaluates branching, interaction-driven uncertainty, opponent influence, and temporal or spatial dependencies to improve comparisons across game-modeling and reinforcement-learning benchmarks.

ArXiv AIResearch16h ago#game-world-modeling#benchmarking
Z.ai Reframes Scaling Beyond Parameter Counts

Z.ai Reframes Scaling Beyond Parameter Counts

Z.ai argues that model scaling should account for data, compute allocation, inference cost, sparsity, effective depth, and post-training—not parameters alone. The post presents GLM-5.3 as a controlled experiment using the same total and activated parameters as GLM-5.2 while scaling long-horizon environments and reinforcement learning for one month.

Reddit r/LocalLLaMACommunity1d ago#scaling-laws#mixture-of-experts#post-training
Page 1