๐Ÿฆ™Stalecollected in 8h

Qwen3.5-35B Dynamic GGUFs Hit SOTA Benchmarks

Qwen3.5-35B Dynamic GGUFs Hit SOTA Benchmarks
PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA
#quantization#benchmarks#kl-divergenceqwen3.5-35b-a3b-unsloth-dynamic-ggufsqwen3.5-35bunslothgguf

๐Ÿ’กSOTA quantized Qwen3.5 GGUFs + 9TB benchmarks for local LLMs

โšก 30-Second TL;DR

What Changed

SOTA 99.9% KL Divergence on UD-Q4_K_XL, IQ3_XXS

Why It Matters

Improves efficiency for local inference of large models like Qwen3.5, enabling better quantized performance without quality loss. Community gains extensive benchmarks for future quant work.

What To Do Next

Download updated Qwen3.5-35B-A3B GGUFs and re-quantize with Imatrix for SOTA perplexity.

Who should care:Researchers & Academics

Key Points

  • โ€ขSOTA 99.9% KL Divergence on UD-Q4_K_XL, IQ3_XXS
  • โ€ขRetiring MXFP4 from most Q2_K_XL, Q3_K_XL, Q4_K_XL quants
  • โ€ขImatrix reduces KLD & PPL; I-quants 5-10% slower
  • โ€ข9TB artifacts available; sensitive tensors like ssm_out avoided

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 6 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขQwen3.5-35B-A3B features 35B total parameters with only 3B activated, using a hybrid architecture of Gated Delta Networks and sparse Mixture-of-Experts (256 experts, 8 routed + 1 shared active) for efficient inference[2].
  • โ€ขModel scores 37 on Artificial Analysis Intelligence Index, significantly above the median of 15 for similar open-weight models, and generates output at 166 tokens per second via Alibaba's API[1].
  • โ€ขSupports native 262,144 token context length, vision-language capabilities with early fusion training, and expanded coverage of 201 languages and dialects[2].

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ขHybrid architecture: Gated Delta Networks + sparse MoE with 256 total experts (8 routed + 1 shared active) for high-throughput inference and low latency[2].
  • โ€ขReasoning model using extended chain-of-thought: Enable via 'Enable Thinking' boolean parameter; supports step-by-step reasoning display in APIs like OpenRouter[2][5].
  • โ€ขMultimodal: Native vision-language with early fusion training on multimodal tokens, outperforming prior Qwen3-VL on reasoning, coding, agents, and visual benchmarks[2].
  • โ€ขPerformance: 166.1 t/s output speed (above median 90.9 t/s); generated 100M tokens on Intelligence Index eval (high vs. median 12M)[1].
  • โ€ขRL training: Scaled across million-agent environments with complex task distributions for real-world adaptability[2].

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Qwen3.5-35B-A3B enables cost-effective local deployment of advanced reasoning on consumer hardware
Its 3B activated parameters from 35B total allow fitting into 24GB VRAM at high quantizations like AQ8 while outperforming larger dense models[4].
Drives adoption of dynamic quants in local inference ecosystems
SOTA KL Divergence on UD-Q4_K_XL and IQ3_XXS plus 9TB artifacts provide benchmarks and configs to standardize high-fidelity low-bit quantization[article].

โณ Timeline

2026-02
Qwen3.5 series released with 35B-A3B as native multimodal reasoning model
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.