
Qwen3.5-397B Uncensored NVFP4 Released
A new uncensored quantization of Qwen3.5-397B in NVFP4 format has been shared. Posted on Reddit r/LocalLLaMA with a link. Likely enables efficient local inference of the massive model.
Tag: #qwen24 results

A new uncensored quantization of Qwen3.5-397B in NVFP4 format has been shared. Posted on Reddit r/LocalLLaMA with a link. Likely enables efficient local inference of the massive model.

PhyDrawGen is a neuro-symbolic pipeline that generates physically accurate diagrams by combining LLM-based scene understanding with deterministic geometric solvers. It addresses common generative model failures like force vector hallucinations and geometric constraint violations.

Shanghai Jiao Tong University and the Qwen team introduced CodePercept, a paradigm that uses executable code as a 'second language' to improve visual perception in STEM tasks. By training models to generate code that reproduces images, the system overcomes the ambiguity of natural language descriptions.
New uncensored 'heretic' version of Qwen3.6 35B A3B preserves all 19 Native MTPs with KLD 0.0015 and only 10/100 refusals. Available in Safetensors, GGUF, NVFP4, NVFP4 GGUF, and GPTQ-Int4 formats on HuggingFace. Benchmarks confirm full MTP retention across formats.

ByteShape released quantized versions of Qwen 3.5 9B with benchmarks across GPUs like 5090/4080 and CPUs. Key GPU picks: 5.10 bpw baseline, 4.43 bpw balanced, 3.60 bpw fast. Blog offers interactive graphs for hardware-specific selection; first of more Qwen drops.
ik_llama.cpp fork delivers 26x faster prompt evaluation (43 to 1,122 tok/s) and 3.5x generation speed on Qwen 3.5 27B Q4_K_M using RTX PRO 4000. Fused GDN kernels reduce graph splits from 34 to 2 for full GPU utilization. Pre-built Windows binaries available as drop-in replacement.
Developer releases abliterated Qwen3.5-9B with vision, achieving 0% refusal rate via two-stage orthogonal projection + LoRA method. Outperforms heretic version's 46% refusals. Available on Ollama and Hugging Face model card.
An uncensored version of the new Qwen3.5-4B model has been released in GGUF format with zero refusals out of 465 tests. It retains full capabilities, supports multimodal inputs, and offers various quants from 2.6GB to 7.9GB. Upcoming uncensored variants for larger Qwen3.5 sizes are in progress.

Qwen 3.5 Plus is now available on Vercel's AI Gateway, featuring a 1M context window and built-in adaptive tool use. It excels in agentic workflows, coding, web development, and multimodal tasks, outperforming Qwen 3 VL in scientific problem-solving and visual reasoning. Developers can access it via AI SDK by setting the model to alibaba/qwen3.5-plus.

Horus Hiero is a new open-source multimodal model built on Qwen 3.5, specifically designed for translating Ancient Egyptian hieroglyphs. It supports 150 languages and features a massive 512K context window.