Qwen3.5 27B at 100+ t/s on 2x3090s

๐กUnlock 100+ t/s Qwen3.5 on 2x3090s for local multi-user inference (585 t/s throughput)
โก 30-Second TL;DR
What Changed
vLLM with tensor parallelism and NVLink for multi-GPU efficiency
Why It Matters
Enables high-throughput local inference on consumer hardware, making powerful LLMs accessible for multi-user setups without enterprise GPUs.
What To Do Next
Compile vLLM from source with tensor parallelism and test MTP=5 on your 3090 setup.
Key Points
- โขvLLM with tensor parallelism and NVLink for multi-GPU efficiency
- โขMTP enabled at 5 tokens for optimal acceptance length
- โขint4 quantization on full attention layers leveraging 3090 hardware
- โขCustom vLLM compilation and engine tweaks boost performance
- โข585 t/s throughput across 8 simultaneous requests
๐ง Deep Insight
Background and context from public sources โ not the original article. 8 sources cited.
๐ Enhanced Key Takeaways
- โขQwen3.5 27B is a dense model with 27.8 billion parameters and native vision-language capabilities using a linear attention mechanism for fast responses[1][5].
- โขIt scores 42 on the Artificial Analysis Intelligence Index, excelling in reasoning, knowledge, mathematics, and coding benchmarks like 72.4 on SWE-bench Verified, matching GPT-5 mini[1][4].
- โขThe model supports a 256K-262K context window, full tool use, and multimodal input under Apache 2.0 license allowing commercial use[1][5].
- โขAPI benchmarks show 100.2 t/s output speed and 1.38s TTFT, above average for similar open-weight models[1].
๐ Competitor Analysisโธ Show
| Feature | Qwen3.5 27B (Dense) | Qwen3.5 35B-A3B (MoE) |
|---|---|---|
| Total Parameters | 27B | 35B |
| Active Parameters | 27B | ~3B |
| Intelligence/Reasoning | High (top-tier coding, logic) | Medium (faster but shallower) |
| Speed (local tests) | ~7-7.5 t/s at Q8 | ~46 t/s at Q8 |
| Benchmarks | 72.4 SWE-bench; 42 Intelligence Index | Surpasses prior 235B-A22B in some tasks |
๐ ๏ธ Technical Deep Dive
- โข27.8 billion parameters, all active per token in dense architecture, providing high reasoning density[1][4].
- โขIncorporates linear attention mechanism in native vision-language setup for balanced inference speed and performance[5].
- โขContext window of 256K-262K tokens with multimodal (vision) input and full tool use support[5].
- โขReleased under Apache 2.0; available via Ollama library[1][7].
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- artificialanalysis.ai โ Qwen3 5 27b
- vertu.com โ Qwen 3 5 27b vs Qwen 3 5 35b A3b Which Local LLM Reigns Supreme
- sonusahani.com โ Qwen 27b vs Qwen 35b
- digitalapplied.com โ Qwen 3 5 Medium Model Series Benchmarks Pricing Guide
- designforonline.com โ Qwen Qwen3 5 27b
- youtube.com โ Watch
- ollama.com โ Qwen3.5:27b
- openrouter.ai โ Qwen3.5 Plus 02 15
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.