Qwen3.8-27B Hits 2,000 Prefill Tokens per Second
An independent developer reports pushing Qwen3.8-27B inference on an RTX 3090 to nearly 2,000 prefill tokens per second and 132 decode tokens per second. The main improvement comes from a custom int8 kernel claimed to achieve 0.99997 similarity to fp32 output.
Reddit r/LocalLLaMA · 15d ago























