Search

10 results on this page

🤖

Qwen KV Precision Shows Real Quality Gaps

A community test on an AMD R9700 with ROCm reports noticeable quality and long-context retention differences between FP16 and q8_0 KV cache for Qwen3.8-27B. FP16 reportedly produces more careful structured output and maintains performance beyond 120k tokens, challenging the assumption that both formats are equivalent.

Reddit r/LocalLLaMACommunity12h ago#kv-cache#long-context#quantization
Qwen3.8-27B Shows Remarkable Local Agency

Qwen3.8-27B Shows Remarkable Local Agency

A Reddit user reports that Qwen3.8-27B autonomously retrieved a university class schedule through 80 tool calls using only credentials and a university name. In another test, it downloaded and analyzed a social-media video, installed Whisper for transcription, and enhanced video frames without human intervention.

Reddit r/LocalLLaMACommunity16h ago#local-inference#tool-use#computer-use
A Smaller Qwen3.8 Built by Pruning Layers

A Smaller Qwen3.8 Built by Pruning Layers

A community developer created Qwen3.8-23B-Mini-Me by strategically removing layers from Qwen3.8-27B, reducing the model to approximately 22.7B parameters without severe reasoning degradation. The model is reported to work well for coding, agentic tasks, and multi-turn chats, but it has not yet been benchmarked and struggles more with edge cases and underspecified prompts.

Reddit r/LocalLLaMACommunity20h ago#model-pruning#model-compression#apple-silicon
Page 1