TQ3_1S Matches Q4_0 on 27B GPUs

💡Run 27B at Q4 quality on 16GB GPUs – local AI breakthrough for mid-range hardware.
⚡ 30-Second TL;DR
What Changed
TQ3_1S PPL 7.2570 vs Q4_0 7.2431 on wiki.test.raw
Why It Matters
This enables larger models like 27B on consumer 16GB GPUs, reducing API reliance for local inference. It democratizes high-quality local AI for hobbyists and devs with mid-range hardware.
What To Do Next
Fork llama.cpp and quantize Qwen3.5-27B to TQ3_1S to test on your 16GB GPU.
Key Points
- •TQ3_1S PPL 7.2570 vs Q4_0 7.2431 on wiki.test.raw
- •12.9GB size vs 14.4GB for Q4_0, 10% smaller
- •Fits fully on 16GB RTX 5060 Ti for 27B models
- •Prompt speed 130.87 tok/s, gen 15.55 tok/s
- •Inspired by TurboQuant and RaBitQ
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.