Search

10 results on this page

A Smaller Qwen3.8 Built by Pruning Layers

A Smaller Qwen3.8 Built by Pruning Layers

A community developer created Qwen3.8-23B-Mini-Me by strategically removing layers from Qwen3.8-27B, reducing the model to approximately 22.7B parameters without severe reasoning degradation. The model is reported to work well for coding, agentic tasks, and multi-turn chats, but it has not yet been benchmarked and struggles more with edge cases and underspecified prompts.

Reddit r/LocalLLaMACommunity1d ago#model-pruning#model-compression#apple-silicon
💼

Duan Yongping Tests Alibaba Again

Investor Duan Yongping bought back 301,400 Alibaba shares worth about $28.93 million, representing only 0.15% of his portfolio. The small position suggests renewed observation rather than strong conviction, especially as Alibaba explores AI cooperation with Apple while PDD Holdings receives a much larger allocation.

🤖

Qwen KV Precision Shows Real Quality Gaps

A community test on an AMD R9700 with ROCm reports noticeable quality and long-context retention differences between FP16 and q8_0 KV cache for Qwen3.8-27B. FP16 reportedly produces more careful structured output and maintains performance beyond 120k tokens, challenging the assumption that both formats are equivalent.

Reddit r/LocalLLaMACommunity20h ago#kv-cache#long-context#quantization
Page 1