Qwen3.5-9B Abliterated: 0% Refusals + Vision
💡Best uncensored 9B LLM yet: 0% refusals + vision on Ollama—game-changer for local apps
⚡ 30-Second TL;DR
What Changed
0% refusal rate vs heretic's 46% using orthogonal projection + LoRA
Why It Matters
Provides top uncensored local 9B model with vision for builders avoiding refusals. Boosts accessible uncensored inference on consumer hardware. Raises bar for ablation techniques in open LLMs.
What To Do Next
Run 'ollama run lukey03/qwen3.5-9b-abliterated-vision' to test zero-refusal vision inference.
Key Points
- •0% refusal rate vs heretic's 46% using orthogonal projection + LoRA
- •Full vision/multimodal support added
- •Ollama commands: lukey03/qwen3.5-9b-abliterated-vision (vision) or text-only
- •Append /no_think tag for faster inference
- •Full methodology in Hugging Face model card
🧠 Deep Insight
Background and context from public sources — not the original article. 6 sources cited.
🔑 Enhanced Key Takeaways
- •The abliterated version uses a two-stage orthogonal projection plus LoRA method to achieve 0% refusals, as detailed in the Hugging Face model card methodology[5].
- •Model size is 6.6GB with a 256K context window, supporting both text and image inputs for local deployment[5].
- •Qwen3.5 series, the base for this model, features hybrid attention architecture like Gated DeltaNet (3 layers linear attention : 1 layer full attention) for efficient edge performance[3].
- •Abliteration technique originates from the 'remove-refusals-with-transformers' method to uncensor models while warning of risks for sensitive outputs[4].
🛠️ Technical Deep Dive
- •Base model: Qwen3.5-9B-Instruct from Alibaba, a compact multimodal model with native vision support and scaled RL training for edge devices[3].
- •Abliteration process: Applies orthogonal projection in two stages combined with LoRA adapters to eliminate refusal behaviors without retraining[1][4].
- •Quantization and size: Deployed at 6.6GB (likely Q4 or similar GGUF quantization), enabling 256K context on consumer hardware like LM Studio or Ollama[4][5].
- •Vision enhancements inherited from Qwen3.5: Supports image understanding with hybrid attention for flat memory usage during long-context multimodal inference[2][3].
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
