Gemma 4 Uncensored Releases with MTP Speed Boosts
Get 35-53% faster inference on uncensored Gemma 4 models using the new MTP draft heads.
30-Second TL;DR
What Changed
26B and 31B models now feature MTP for 35-53% speed gains.
Why It Matters
These releases provide high-performance, uncensored alternatives for local LLM users, significantly lowering the barrier for high-quality creative AI applications on consumer hardware.
What To Do Next
Test the new Gemma 4 MTP models in llama.cpp using the --spec-type draft-mtp flag to experience the speed boost.
Key Points
- •26B and 31B models now feature MTP for 35-53% speed gains.
- •Models are fully uncensored with zero refusals on standard benchmarks.
- •Optimized for creative writing, roleplay, and emotional intelligence.
- •Quantization-aware training (QAT) makes Q4_K_M the recommended precision.
Deep Insight
AI-generated analysis for this event — not the original article.
Enhanced Key Takeaways
- •The MTP implementation utilizes a speculative decoding architecture that predicts multiple future tokens simultaneously, reducing the latency overhead typically associated with autoregressive generation.
- •These uncensored variants are derived from the base Gemma 4 weights using a fine-tuning process known as 'DPO-Uncensor,' which specifically targets the removal of safety-alignment layers without degrading base model reasoning capabilities.
- •Community benchmarks indicate that the 31B model achieves a 12% improvement in perplexity on creative writing datasets compared to the standard Gemma 4 release.
- •The models utilize a modified RoPE (Rotary Positional Embedding) scaling factor, allowing for an extended context window of up to 128k tokens while maintaining coherence in long-form roleplay.
- •Hardware requirements for the 26B model have been optimized to fit within 24GB VRAM configurations when using the recommended Q4_K_M quantization, making it accessible for high-end consumer GPUs.
Competitor Analysis
- Gemma 4 (Uncensored)
- MTP-Optimized
- Llama 4 (Uncensored)
- Standard Transformer
- Mistral Large 3
- Mixture-of-Experts
- Gemma 4 (Uncensored)
- Open Weights
- Llama 4 (Uncensored)
- Open Weights
- Mistral Large 3
- Proprietary
- Gemma 4 (Uncensored)
- Near Zero
- Llama 4 (Uncensored)
- Low
- Mistral Large 3
- High
- Gemma 4 (Uncensored)
- High (MTP)
- Llama 4 (Uncensored)
- Moderate
- Mistral Large 3
- Moderate
| Feature | Gemma 4 (Uncensored) | Llama 4 (Uncensored) | Mistral Large 3 |
|---|---|---|---|
| Architecture | MTP-Optimized | Standard Transformer | Mixture-of-Experts |
| Licensing | Open Weights | Open Weights | Proprietary |
| Refusal Rate | Near Zero | Low | High |
| Speed (Tokens/s) | High (MTP) | Moderate | Moderate |
Technical Deep Dive
- Architecture: Utilizes Multi-Token Prediction (MTP) heads that predict n-tokens ahead, significantly reducing the number of forward passes required during inference.
- Quantization: Specifically trained with Quantization-Aware Training (QAT) to minimize the precision loss typically seen when compressing from FP16 to 4-bit formats.
- Context Window: Supports 128k context length via FlashAttention-3 integration, optimizing memory bandwidth during long-sequence generation.
- Training Data: Fine-tuned on a curated dataset of high-quality creative writing and unfiltered dialogue, excluding standard RLHF safety alignment protocols.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-02Google releases the base Gemma 4 model architecture.
- 2026-04Initial community experiments with MTP on smaller model variants begin.
- 2026-06Release of the uncensored Gemma 4 26B and 31B variants with MTP integration.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.