🦙Stalecollected in 2h

Qwen3.5-122B Heretic GGUF on Hugging Face

Qwen3.5-122B Heretic GGUF on Hugging Face
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#quantization#moe-model#local-llmqwen3.5-122b-a10b-heretic-ggufqwen3.5-122bggufhugging-facesabomako

💡Huge 122B Qwen3.5 heretic in GGUF—local inference ready

⚡ 30-Second TL;DR

What Changed

122B parameter Qwen3.5 model in GGUF format

Why It Matters

Provides accessible quantized large model for local deployment, expanding options for high-performance inference without cloud.

What To Do Next

Download Qwen3.5-122B-A10B-heretic-GGUF from Hugging Face for LM Studio testing.

Who should care:Developers & AI Engineers

Key Points

  • 122B parameter Qwen3.5 model in GGUF format
  • A10B-heretic variant optimized for local runs
  • Repo by Sabomako on Hugging Face for download

🧠 Deep Insight

Background and context from public sources — not the original article. 5 sources cited.

🔑 Enhanced Key Takeaways

  • Qwen3.5-122B-A10B is a Mixture-of-Experts (MoE) model with 10B active parameters out of 122B total, designed for higher intelligence per compute than dense scaling[3][4].
  • Unsloth provides state-of-the-art Dynamic GGUF quantizations for Qwen3.5 models, achieving 99.9% KL Divergence on benchmarks with over 150 tests and 9TB of artifacts[2][3].
  • The model series, including 122B-A10B, supports up to 1M context length in Flash variants and built-in tools, with strong practitioner feedback on local performance[4].

🛠️ Technical Deep Dive

  • MoE architecture: Qwen3.5-122B-A10B activates 10B parameters, balancing efficiency and performance[3][4].
  • Quantization insights: Unsloth benchmarks show ffn_up_exps, ffn_gate_exps optimal at 3-bit (iq3_xxs); avoid MXFP4 on attn_gate, attn_q, ssm_beta/alpha—prefer Q4_K[2].
  • GGUF updates: Fixed tool calling chat template bug across all quant types; 122B GGUFs pending re-upload as of Mar 2, 2026[2][3].
  • Hardware fit: Full 397B checkpoint ~807GB, but 3-bit GGUF fits 192GB RAM, 4-bit (MXFP4) on 256GB[3].

🔮 Future ImplicationsAI analysis grounded in cited sources

Qwen3.5 MoE GGUFs will dominate local inference by Q2 2026
Unsloth's SOTA quantizations and community enthusiasm position them ahead of dense models for efficiency on consumer hardware[2][4].
122B-A10B GGUF re-uploads complete by Mar 4, 2026
Unsloth announced 122B conversion ongoing with updates expected 'today' on Mar 2[3].

Timeline

2026-03
Alibaba releases Qwen3.5 series including 122B-A10B MoE model
2026-03-02
Unsloth updates 35B-A3B GGUFs, announces fixes and benchmarks for 122B/27B
2026-03-03
Sabomako uploads Qwen3.5-122B-A10B-heretic GGUF to Hugging Face, shared on r/LocalLLaMA

📎 Sources (5)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. news.ycombinator.com — Item
  2. unsloth.ai — Gguf Benchmarks
  3. unsloth.ai — Qwen3
  4. latent.space — Ainews the Unreasonable Effectiveness
  5. forums.developer.nvidia.com — 361639
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.