πŸ¦™Freshcollected in 2h

Dual R9700 Rig Delivers 111 Tokens per Second

PostLinkedIn
πŸ¦™Read original on Reddit r/LocalLLaMA
#local-inference#multi-gpu#quantization#amd-gpuamd-radeon-ai-pro-r9700amdradeon ai pro r9700qwen3.8vllm radiancebetterbench

πŸ’‘See how a €4,000 dual-AMD rig compares with a far more expensive RTX 5090 for local LLM inference.

⚑ 30-Second TL;DR

What Changed

Two 32 GB Radeon AI PRO R9700 GPUs operate through PCIe 5.0 x8 with 64 GB of system memory.

Why It Matters

The benchmark suggests that multi-GPU AMD workstations can be a cost-effective alternative to a single high-end NVIDIA card for local LLM experimentation. It also highlights the practical trade-offs between quantization, expert offloading, thermals, and sustained generation speed.

What To Do Next

Benchmark your target Qwen3.8 quantization on dual GPUs with BetterBench before choosing between FP8, AWQ MXFP4, and expert offloading.

Who should care:Researchers & Academics

Key Points

  • β€’Two 32 GB Radeon AI PRO R9700 GPUs operate through PCIe 5.0 x8 with 64 GB of system memory.
  • β€’Qwen3.8-27B in Quark AWQ MXFP4 reaches 111.4 tokens per second, compared with 87.6 tokens per second in native FP8.
  • β€’Qwen3.8-Flash-Next in UD-IQ4_XS GGUF reaches 35.4 tokens per second using tiered expert offload.
  • β€’The system records 73–81 ms TTFT for Qwen3.8-27B and prefill rates above 4,200 tokens per second.
  • β€’One GPU runs 10–15Β°C hotter, prompting planned power limiting to 210 W and undervolting.
πŸ“°

Weekly AI Recap

Read this week's curated digest of top AI events β†’

πŸ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.

Dual R9700 Rig Delivers 111 Tokens per Second | Reddit r/LocalLLaMA | SetupAI | SetupAI