vLLM-Radiance Hits 17.6K Prefill on Four GPUs

π‘A real-world report shows four AMD GPUs reaching 17.6K prefill with an alternative vLLM stack.
β‘ 30-Second TL;DR
What Changed
The reported prefill performance reached 17,636 tokens per second, with generation throughput recorded at 36.6 tokens per second during the spike.
Why It Matters
The report suggests that an alternative vLLM implementation can unlock strong multi-GPU throughput on hardware that performs poorly with the official stack. If reproducible, this could make high-throughput local inference more attractive for developers operating AMD-based systems.
What To Do Next
Benchmark vLLM-Radiance against official vLLM on your multi-GPU AMD host using identical Qwen model, context, batch, and MTP settings.
Key Points
- β’The reported prefill performance reached 17,636 tokens per second, with generation throughput recorded at 36.6 tokens per second during the spike.
- β’The best reported generation speed was 106 tokens per second with an 80% MTP acceptance rate.
- β’Four GPUs were consistently shown at 100% utilization, even though official support reportedly covers only dual-GPU setups.
- β’The workload used Hermes with Qwen 3.8 27B FP8 and a 262K context window.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
