SourceStalecollected in 62m

Dual 7900 XTX Hits 123 tok/s on Qwen3.5-35B

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#amd-gpu#moe-inference#llama-benchllama.cpp-vulkan-on-dual-rx-7900-xtxrx-7900-xtxqwen3.5-35bllama.cppvulkan

💡123 tok/s on dual AMD GPUs beats NVIDIA—key for local 35B MoE runs

⚡ 30-Second TL;DR

What Changed

Dual 7900 XTX Vulkan: 123 tok/s TG128, 2,647 tok/s PP512 on Qwen3.5-35B-A3B Q4_K_M

Why It Matters

Demonstrates Vulkan backend maturity on AMD GPUs for high-speed MoE inference, making dual consumer cards viable for 35B models. Highlights llama.cpp superiority over vLLM on ROCm for now.

What To Do Next

Benchmark your dual 7900 XTX with llama.cpp Vulkan build b8516 on Qwen3.5-35B.

Who should care:Developers & AI Engineers

Key Points

  • Dual 7900 XTX Vulkan: 123 tok/s TG128, 2,647 tok/s PP512 on Qwen3.5-35B-A3B Q4_K_M
  • Beats single 7900 XTX HIP (76-78 tok/s) and RTX 3090 CUDA (111 tok/s)
  • vLLM ROCm broken: garbage output at 5 tok/s, OOM issues
  • llama.cpp HIP graphs: 86 tok/s TG, solid alternative
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.