Qwen3.5-35B Dynamic GGUFs Hit SOTA Benchmarks

๐กSOTA quantized Qwen3.5 GGUFs + 9TB benchmarks for local LLMs
โก 30-Second TL;DR
What Changed
SOTA 99.9% KL Divergence on UD-Q4_K_XL, IQ3_XXS
Why It Matters
Improves efficiency for local inference of large models like Qwen3.5, enabling better quantized performance without quality loss. Community gains extensive benchmarks for future quant work.
What To Do Next
Download updated Qwen3.5-35B-A3B GGUFs and re-quantize with Imatrix for SOTA perplexity.
Key Points
- โขSOTA 99.9% KL Divergence on UD-Q4_K_XL, IQ3_XXS
- โขRetiring MXFP4 from most Q2_K_XL, Q3_K_XL, Q4_K_XL quants
- โขImatrix reduces KLD & PPL; I-quants 5-10% slower
- โข9TB artifacts available; sensitive tensors like ssm_out avoided
๐ง Deep Insight
Background and context from public sources โ not the original article. 6 sources cited.
๐ Enhanced Key Takeaways
- โขQwen3.5-35B-A3B features 35B total parameters with only 3B activated, using a hybrid architecture of Gated Delta Networks and sparse Mixture-of-Experts (256 experts, 8 routed + 1 shared active) for efficient inference[2].
- โขModel scores 37 on Artificial Analysis Intelligence Index, significantly above the median of 15 for similar open-weight models, and generates output at 166 tokens per second via Alibaba's API[1].
- โขSupports native 262,144 token context length, vision-language capabilities with early fusion training, and expanded coverage of 201 languages and dialects[2].
๐ ๏ธ Technical Deep Dive
- โขHybrid architecture: Gated Delta Networks + sparse MoE with 256 total experts (8 routed + 1 shared active) for high-throughput inference and low latency[2].
- โขReasoning model using extended chain-of-thought: Enable via 'Enable Thinking' boolean parameter; supports step-by-step reasoning display in APIs like OpenRouter[2][5].
- โขMultimodal: Native vision-language with early fusion training on multimodal tokens, outperforming prior Qwen3-VL on reasoning, coding, agents, and visual benchmarks[2].
- โขPerformance: 166.1 t/s output speed (above median 90.9 t/s); generated 100M tokens on Intelligence Index eval (high vs. median 12M)[1].
- โขRL training: Scaled across million-agent environments with complex task distributions for real-world adaptability[2].
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.