TurboQuant: 4-bit LLM Weights, 3.2x Savings
TurboQuant algorithm adapted from KV-cache to model weights, offering drop-in nn.Linear replacement with near-optimal 4-bit quantization and lossless 8-bit residual. Benchmarks on Qwen2.5-0.5B show zero PPL increase at 50% size; 4B tests confirm minimal degradation. GitHub includes Triton kernels and full docs.
