πŸ¦™Freshcollected in 15m

vLLM-Radiance Hits 17.6K Prefill on Four GPUs

vLLM-Radiance Hits 17.6K Prefill on Four GPUs
PostLinkedIn
πŸ¦™Read original on Reddit r/LocalLLaMA
#inference#multi-gpu#amd-gpu#throughputvllm-radiancevllm-radiancevllmqwen 3.8 27bamd r9700 ai prohermes

πŸ’‘A real-world report shows four AMD GPUs reaching 17.6K prefill with an alternative vLLM stack.

⚑ 30-Second TL;DR

What Changed

The reported prefill performance reached 17,636 tokens per second, with generation throughput recorded at 36.6 tokens per second during the spike.

Why It Matters

The report suggests that an alternative vLLM implementation can unlock strong multi-GPU throughput on hardware that performs poorly with the official stack. If reproducible, this could make high-throughput local inference more attractive for developers operating AMD-based systems.

What To Do Next

Benchmark vLLM-Radiance against official vLLM on your multi-GPU AMD host using identical Qwen model, context, batch, and MTP settings.

Who should care:Developers & AI Engineers

Key Points

  • β€’The reported prefill performance reached 17,636 tokens per second, with generation throughput recorded at 36.6 tokens per second during the spike.
  • β€’The best reported generation speed was 106 tokens per second with an 80% MTP acceptance rate.
  • β€’Four GPUs were consistently shown at 100% utilization, even though official support reportedly covers only dual-GPU setups.
  • β€’The workload used Hermes with Qwen 3.8 27B FP8 and a 262K context window.
πŸ“°

Weekly AI Recap

Read this week's curated digest of top AI events β†’

πŸ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.

vLLM-Radiance Hits 17.6K Prefill on Four GPUs | Reddit r/LocalLLaMA | SetupAI | SetupAI