NVIDIA Releases Quantized Qwen3.6-35B-A3B Model

๐กDeploy large MoE models with 3x less VRAM using NVIDIA's new NVFP4 quantized Qwen3.6 model.
โก 30-Second TL;DR
What Changed
Quantized to NVFP4 format using NVIDIA Model Optimizer
Why It Matters
This release provides a high-efficiency deployment path for large MoE models on consumer or enterprise hardware with limited VRAM. It demonstrates the viability of 4-bit quantization for maintaining high-level reasoning capabilities.
What To Do Next
Download the NVFP4 weights from Hugging Face and test them in your vLLM pipeline to reduce your inference hardware footprint.
Key Points
- โขQuantized to NVFP4 format using NVIDIA Model Optimizer
- โขAchieves ~3.06x reduction in disk size and GPU memory usage
- โขMaintains benchmark accuracy comparable to BF16 precision
- โขOptimized specifically for inference with vLLM
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
