SourceStalecollected in 5h

Arc B70 hits 135 tps on Qwen3.5-27B

PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#gpu-benchmark#inference#intel-xpu#high-concurrencyintel-arc-pro-b70intel-arc-pro-b70qwen3.5-27bvllmllama.cpp

💡Intel GPU nears Nvidia LLM speeds at 1/2 price? Benchmarks + setup guide

⚡ 30-Second TL;DR

What Changed

12 tps single query, 135 tps at 32 concurrency

Why It Matters

Validates Intel Arc for cost-effective LLM inference at scale, though power efficiency lags Nvidia; appeals to budget-conscious practitioners avoiding CUDA lock-in.

What To Do Next

Deploy vllm on Arc B70 using the post's Docker command on Ubuntu 26.04 beta.

Who should care:Developers & AI Engineers

Key Points

  • 12 tps single query, 135 tps at 32 concurrency
  • 20% slower TG than RTX PRO 4500 at high load
  • 50% higher power consumption noted
  • Needs Ubuntu 26.04 and vllm Intel beta fork
  • Docker command for easy vllm setup shared
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.