
Qwen3.8-27B Hits 218 Tok/s on Dual RTX 3090s
A community benchmark reports up to 218.3 tokens per second for Qwen3.8-27B on two RTX 3090 GPUs using vLLM, INT4 quantization, and DFlash2 speculative decoding. The setup achieved a 131K context ceiling, 168–178 ms time to first token, and peak VRAM usage of 22.3 GB per card.






