Search

Tag: #gguf28 results

Qwen3.5-27B Q4 Quant Rankings

Qwen3.5-27B Q4 Quant Rankings

Comprehensive Q4 quantization comparison of Qwen3.5-27B GGUF files using KLD against BF16 baseline on custom chat and Wikitext2 datasets. unsloth_Qwen3.5-27B-UD-Q4_K_XL leads with lowest KLD of 0.005087. IQ4_XS variants top efficiency for VRAM-size balance.

Reddit r/LocalLLaMACommunityMar 3#quantization#gguf#kld-benchmark
🤖

Uncensored Qwen3.5-4B Aggressive GGUF Drops

An uncensored version of the new Qwen3.5-4B model has been released in GGUF format with zero refusals out of 465 tests. It retains full capabilities, supports multimodal inputs, and offers various quants from 2.6GB to 7.9GB. Upcoming uncensored variants for larger Qwen3.5 sizes are in progress.

Reddit r/LocalLLaMACommunityMar 3#uncensored#gguf#multimodal
Qwen3.5-9B Launches on Hugging Face

Qwen3.5-9B Launches on Hugging Face

Unsloth released Qwen3.5-9B-GGUF on Hugging Face, a 9B causal LM with vision encoder. It features advanced architecture like Gated DeltaNet and supports up to 1M token context. Trained with multi-step MTP for pre- and post-training.

Reddit r/LocalLLaMACommunityMar 2#gguf#long-context#9b
LFM2.5 Runs at 17 Tok/s on OnePlus 13

LFM2.5 Runs at 17 Tok/s on OnePlus 13

A developer demonstrated LFM2.5-2.6B running at approximately 17 tokens per second on a OnePlus 13 using only the phone’s CPU. The setup uses a Q4_K_M GGUF model, a custom 450 KB inference engine, and an ADB-based device probe suite, with a target of around 30 tokens per second.

Reddit r/LocalLLaMACommunityAug 5#on-device-inference#android#gguf
🤖

Savant Commander 48B: 12-Distill MOE

Savant Commander 48B is a custom Qwen3-based 4x12B MOE merging distills from Claude, Gemini, OpenAI, DeepSeek, and more with hand-coded routing. Users control activation via prompts and test differences easily. GGUF versions include regular and Heretic uncensored, available on Hugging Face.

Reddit r/LocalLLaMACommunityMar 24#moe#distillation#uncensored
Page 2 of 3