
ExLlamav3 Adds CPU Offload and Flash Model Support
ExLlamav3 has received a major update with CPU offloading for MoE experts, disk-based ngram offload for Qwen-3.8-Flash-Next, and support for GLM-5.3-Flash. It also introduces a self-calibrated optimization technique and additional performance improvements for NVIDIA users.
Reddit r/LocalLLaMA · 17d ago























