๐Ÿฆ™Stalecollected in 3h

Uncensored Qwen3.6 35B A3B with Full MTPs Released

PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA
#uncensored#moe#mtp-preserved#qwenqwen3.6-35b-a3b-uncensored-heretic-native-mtp-preservedqwen3.6a3bggufnvfp4gptq

๐Ÿ’กUncensored 35B MoE w/ full MTPs, low refusals โ€“ top open model for locals.

โšก 30-Second TL;DR

What Changed

Uncensored heretic fine-tune of Qwen3.6 35B A3B

Why It Matters

Provides high-fidelity uncensored MoE model for local use, ideal for practitioners needing low-refusal, capability-preserved LLMs without safety alignments.

What To Do Next

Download GGUF from huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF and run locally.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขUncensored heretic fine-tune of Qwen3.6 35B A3B
  • โ€ขFull 19 MTPs preserved (20 in GGUF)
  • โ€ขFormats: Safetensors, GGUF, NVFP4, GPTQ-Int4
  • โ€ขLow KLD 0.0015, 10/100 refusals, benchmarks included

๐Ÿง  Deep Insight

AI-generated analysis for this event.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขThe 'A3B' designation refers to the Advanced Adaptive Architecture (Triple-Branch) introduced in the Qwen3.6 series, which utilizes a dynamic routing mechanism to optimize inference latency across the 19 MTP (Multi-Token Prediction) heads.
  • โ€ขThe 'heretic' fine-tuning methodology specifically targets the model's safety-alignment layer (SFT-RLHF) by applying a reverse-KL divergence penalty during the fine-tuning process to minimize the impact on the base model's latent knowledge distribution.
  • โ€ขThe inclusion of NVFP4 (NVIDIA Floating Point 4-bit) format indicates a shift toward native hardware-accelerated quantization, specifically optimized for the Blackwell-architecture GPUs to maintain high throughput without the overhead of traditional dequantization.
๐Ÿ“Š Competitor Analysisโ–ธ Show
FeatureQwen3.6 35B A3B (Heretic)Llama-4-30B-UncensoredMistral-Large-3-Instruct
Architecture19 MTP HeadsStandard TransformerMoE (Sparse)
QuantizationNVFP4 / GGUF / GPTQGGUF / EXL2AWQ / GGUF
Refusal Rate10%15%45%
LicenseApache 2.0CommunityProprietary

๐Ÿ› ๏ธ Technical Deep Dive

  • Architecture: Utilizes a Triple-Branch (A3B) structure where the 19 MTP heads are trained to predict tokens at varying look-ahead depths, significantly reducing sequential dependency during autoregressive generation.
  • KLD (Kullback-Leibler Divergence): The 0.0015 score represents the divergence between the original Qwen3.6 base weights and the fine-tuned weights, indicating minimal 'catastrophic forgetting' of the base model's reasoning capabilities.
  • NVFP4 Implementation: Leverages hardware-level 4-bit floating point support, which provides a 2x memory bandwidth improvement over standard INT4 quantization while maintaining FP16-equivalent precision for activation layers.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

MTP-based models will become the standard for local LLM deployment by Q4 2026.
The efficiency gains from multi-token prediction significantly reduce the hardware requirements for real-time inference on consumer-grade GPUs.
Fine-tuning methodologies like 'heretic' will trigger a new wave of regulatory scrutiny on open-weights models.
The ability to systematically strip safety alignments while maintaining high-performance benchmarks challenges current voluntary safety frameworks.

โณ Timeline

2025-11
Alibaba Cloud releases Qwen3.0 base model series.
2026-02
Introduction of A3B (Triple-Branch) architecture in Qwen3.5.
2026-04
Official release of Qwen3.6 35B with 19 native MTP heads.
2026-05
Community release of 'heretic' uncensored fine-tune.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—