🦙Stalecollected in 2h

Qwen3.5-9B Abliterated: 0% Refusals + Vision

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA

💡Best uncensored 9B LLM yet: 0% refusals + vision on Ollama—game-changer for local apps

⚡ 30-Second TL;DR

What Changed

0% refusal rate vs heretic's 46% using orthogonal projection + LoRA

Why It Matters

Provides top uncensored local 9B model with vision for builders avoiding refusals. Boosts accessible uncensored inference on consumer hardware. Raises bar for ablation techniques in open LLMs.

What To Do Next

Run 'ollama run lukey03/qwen3.5-9b-abliterated-vision' to test zero-refusal vision inference.

Who should care:Developers & AI Engineers

Key Points

  • 0% refusal rate vs heretic's 46% using orthogonal projection + LoRA
  • Full vision/multimodal support added
  • Ollama commands: lukey03/qwen3.5-9b-abliterated-vision (vision) or text-only
  • Append /no_think tag for faster inference
  • Full methodology in Hugging Face model card

🧠 Deep Insight

Background and context from public sources — not the original article. 6 sources cited.

🔑 Enhanced Key Takeaways

  • The abliterated version uses a two-stage orthogonal projection plus LoRA method to achieve 0% refusals, as detailed in the Hugging Face model card methodology[5].
  • Model size is 6.6GB with a 256K context window, supporting both text and image inputs for local deployment[5].
  • Qwen3.5 series, the base for this model, features hybrid attention architecture like Gated DeltaNet (3 layers linear attention : 1 layer full attention) for efficient edge performance[3].
  • Abliteration technique originates from the 'remove-refusals-with-transformers' method to uncensor models while warning of risks for sensitive outputs[4].

🛠️ Technical Deep Dive

  • Base model: Qwen3.5-9B-Instruct from Alibaba, a compact multimodal model with native vision support and scaled RL training for edge devices[3].
  • Abliteration process: Applies orthogonal projection in two stages combined with LoRA adapters to eliminate refusal behaviors without retraining[1][4].
  • Quantization and size: Deployed at 6.6GB (likely Q4 or similar GGUF quantization), enabling 256K context on consumer hardware like LM Studio or Ollama[4][5].
  • Vision enhancements inherited from Qwen3.5: Supports image understanding with hybrid attention for flat memory usage during long-context multimodal inference[2][3].

🔮 Future ImplicationsAI analysis grounded in cited sources

Abliterated Qwen3.5-9B will accelerate uncensored local AI adoption on edge devices
Its 6.6GB size, 0% refusals, and full vision make it ideal for on-device agents, as shown in Ollama and iPhone demos of similar Qwen3.5 models[3][5].
Hybrid attention in Qwen3.5 enables broader lightweight multimodal deployment
Gated DeltaNet pattern maintains quality with flat memory, outperforming larger models in benchmarks for edge use cases[3].

Timeline

2026-02
Alibaba releases Qwen3.5 small model series (0.8B to 9B) with native multimodal and hybrid attention for edge devices[3]
2026-02
huihui_ai publishes initial qwen3.5-abliterated uncensored versions across sizes on Ollama[4]
2026-03
lukey03 releases Qwen3.5-9B-abliterated-vision with 0% refusals and full vision support on Ollama[5]
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.