Dual R9700 Rig Delivers 111 Tokens per Second
π‘See how a β¬4,000 dual-AMD rig compares with a far more expensive RTX 5090 for local LLM inference.
β‘ 30-Second TL;DR
What Changed
Two 32 GB Radeon AI PRO R9700 GPUs operate through PCIe 5.0 x8 with 64 GB of system memory.
Why It Matters
The benchmark suggests that multi-GPU AMD workstations can be a cost-effective alternative to a single high-end NVIDIA card for local LLM experimentation. It also highlights the practical trade-offs between quantization, expert offloading, thermals, and sustained generation speed.
What To Do Next
Benchmark your target Qwen3.8 quantization on dual GPUs with BetterBench before choosing between FP8, AWQ MXFP4, and expert offloading.
Key Points
- β’Two 32 GB Radeon AI PRO R9700 GPUs operate through PCIe 5.0 x8 with 64 GB of system memory.
- β’Qwen3.8-27B in Quark AWQ MXFP4 reaches 111.4 tokens per second, compared with 87.6 tokens per second in native FP8.
- β’Qwen3.8-Flash-Next in UD-IQ4_XS GGUF reaches 35.4 tokens per second using tiered expert offload.
- β’The system records 73β81 ms TTFT for Qwen3.8-27B and prefill rates above 4,200 tokens per second.
- β’One GPU runs 10β15Β°C hotter, prompting planned power limiting to 210 W and undervolting.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
Same topic
Explore #local-inference
Same product
More on amd-radeon-ai-pro-r9700
Same source
Latest from Reddit r/LocalLLaMA

Spark-X2.5 Brings 1M Context to Compact Models

Teaching Qwen Next 3D Sculpting with GPT Astra
Eight Qwen 3.8 27B Uncensored Variants Compared
New Benchmarks Probe Deep Software Engineering Ability
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.