Search

Tag: #gpu-benchmark8 results

⚙️

Arc B70 hits 135 tps on Qwen3.5-27B

Benchmarks show Intel Arc Pro B70 32GB achieving 12 tps on single Qwen3.5-27B@Q4 queries via llama.cpp/vllm, scaling to 135 tps at 32 concurrency – 20% behind RTX PRO 4500 but with 50% higher power draw. Requires Ubuntu 26.04 beta and beta vllm fork. Detailed Docker command provided for setup.

Reddit r/LocalLLaMACommunityApr 11#gpu-benchmark#inference#intel-xpu