Search

10 results on this page

Z.ai Reframes Scaling Beyond Parameter Counts

Z.ai Reframes Scaling Beyond Parameter Counts

Z.ai argues that model scaling should account for data, compute allocation, inference cost, sparsity, effective depth, and post-training—not parameters alone. The post presents GLM-5.3 as a controlled experiment using the same total and activated parameters as GLM-5.2 while scaling long-horizon environments and reinforcement learning for one month.

Reddit r/LocalLLaMACommunity1d ago#scaling-laws#mixture-of-experts#post-training
Unverified Harness Claims to Beat Fable 5

Unverified Harness Claims to Beat Fable 5

The community project J-Space Cognition Suite claims to improve DeepSeek V4-Pro-0813 through an inference-time Agent Harness without changing model weights. Its reported gains on several benchmarks have not been independently reproduced, and the project is unrelated to Anthropic's internal J-space interpretability research.

🤖

Qwen KV Precision Shows Real Quality Gaps

A community test on an AMD R9700 with ROCm reports noticeable quality and long-context retention differences between FP16 and q8_0 KV cache for Qwen3.8-27B. FP16 reportedly produces more careful structured output and maintains performance beyond 120k tokens, challenging the assumption that both formats are equivalent.

Reddit r/LocalLLaMACommunity17h ago#kv-cache#long-context#quantization
Qwen3.8-27B Shows Remarkable Local Agency

Qwen3.8-27B Shows Remarkable Local Agency

A Reddit user reports that Qwen3.8-27B autonomously retrieved a university class schedule through 80 tool calls using only credentials and a university name. In another test, it downloaded and analyzed a social-media video, installed Whisper for transcription, and enhanced video frames without human intervention.

Reddit r/LocalLLaMACommunity21h ago#local-inference#tool-use#computer-use
A Smaller Qwen3.8 Built by Pruning Layers

A Smaller Qwen3.8 Built by Pruning Layers

A community developer created Qwen3.8-23B-Mini-Me by strategically removing layers from Qwen3.8-27B, reducing the model to approximately 22.7B parameters without severe reasoning degradation. The model is reported to work well for coding, agentic tasks, and multi-turn chats, but it has not yet been benchmarked and struggles more with edge cases and underspecified prompts.

Reddit r/LocalLLaMACommunity1d ago#model-pruning#model-compression#apple-silicon
Page 3