SourceReddit r/LocalLLaMA•Stalecollected in 5h
Arc B70 hits 135 tps on Qwen3.5-27B
#gpu-benchmark#inference#intel-xpu#high-concurrencyintel-arc-pro-b70intel-arc-pro-b70qwen3.5-27bvllmllama.cpp
💡Intel GPU nears Nvidia LLM speeds at 1/2 price? Benchmarks + setup guide
⚡ 30-Second TL;DR
What Changed
12 tps single query, 135 tps at 32 concurrency
Why It Matters
Validates Intel Arc for cost-effective LLM inference at scale, though power efficiency lags Nvidia; appeals to budget-conscious practitioners avoiding CUDA lock-in.
What To Do Next
Deploy vllm on Arc B70 using the post's Docker command on Ubuntu 26.04 beta.
Who should care:Developers & AI Engineers
Key Points
- •12 tps single query, 135 tps at 32 concurrency
- •20% slower TG than RTX PRO 4500 at high load
- •50% higher power consumption noted
- •Needs Ubuntu 26.04 and vllm Intel beta fork
- •Docker command for easy vllm setup shared
📰
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.