DFlash2 Hits 138 TPS on RTX 3090
A hyper-optimized Qwen3.8-27B inference stack using DFlash2 reaches about 138 tokens per second on a power-limited RTX 3090, while 64-request throughput reaches 942 TPS. Prefix caching also reduces long-chat follow-up latency from roughly 23 seconds to under 1.4 seconds in reported tests.





