Dual 7900 XTX Hits 123 tok/s on Qwen3.5-35B
💡123 tok/s on dual AMD GPUs beats NVIDIA—key for local 35B MoE runs
⚡ 30-Second TL;DR
What Changed
Dual 7900 XTX Vulkan: 123 tok/s TG128, 2,647 tok/s PP512 on Qwen3.5-35B-A3B Q4_K_M
Why It Matters
Demonstrates Vulkan backend maturity on AMD GPUs for high-speed MoE inference, making dual consumer cards viable for 35B models. Highlights llama.cpp superiority over vLLM on ROCm for now.
What To Do Next
Benchmark your dual 7900 XTX with llama.cpp Vulkan build b8516 on Qwen3.5-35B.
Key Points
- •Dual 7900 XTX Vulkan: 123 tok/s TG128, 2,647 tok/s PP512 on Qwen3.5-35B-A3B Q4_K_M
- •Beats single 7900 XTX HIP (76-78 tok/s) and RTX 3090 CUDA (111 tok/s)
- •vLLM ROCm broken: garbage output at 5 tok/s, OOM issues
- •llama.cpp HIP graphs: 86 tok/s TG, solid alternative
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.