SourceStalecollected in 2h

TQ3_1S Matches Q4_0 on 27B GPUs

TQ3_1S Matches Q4_0 on 27B GPUs
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#quantization#local-inference#cuda-supporttq3_1stq3_1sqwen3.5-27bllama.cppturboquantrtx-5060-ti

💡Run 27B at Q4 quality on 16GB GPUs – local AI breakthrough for mid-range hardware.

⚡ 30-Second TL;DR

What Changed

TQ3_1S PPL 7.2570 vs Q4_0 7.2431 on wiki.test.raw

Why It Matters

This enables larger models like 27B on consumer 16GB GPUs, reducing API reliance for local inference. It democratizes high-quality local AI for hobbyists and devs with mid-range hardware.

What To Do Next

Fork llama.cpp and quantize Qwen3.5-27B to TQ3_1S to test on your 16GB GPU.

Who should care:Developers & AI Engineers

Key Points

  • TQ3_1S PPL 7.2570 vs Q4_0 7.2431 on wiki.test.raw
  • 12.9GB size vs 14.4GB for Q4_0, 10% smaller
  • Fits fully on 16GB RTX 5060 Ti for 27B models
  • Prompt speed 130.87 tok/s, gen 15.55 tok/s
  • Inspired by TurboQuant and RaBitQ
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.