๐Ÿฆ™Stalecollected in 10h

Qwen3.5 27B at 100+ t/s on 2x3090s

Qwen3.5 27B at 100+ t/s on 2x3090s
PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA

๐Ÿ’กUnlock 100+ t/s Qwen3.5 on 2x3090s for local multi-user inference (585 t/s throughput)

โšก 30-Second TL;DR

What Changed

vLLM with tensor parallelism and NVLink for multi-GPU efficiency

Why It Matters

Enables high-throughput local inference on consumer hardware, making powerful LLMs accessible for multi-user setups without enterprise GPUs.

What To Do Next

Compile vLLM from source with tensor parallelism and test MTP=5 on your 3090 setup.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขvLLM with tensor parallelism and NVLink for multi-GPU efficiency
  • โ€ขMTP enabled at 5 tokens for optimal acceptance length
  • โ€ขint4 quantization on full attention layers leveraging 3090 hardware
  • โ€ขCustom vLLM compilation and engine tweaks boost performance
  • โ€ข585 t/s throughput across 8 simultaneous requests

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 8 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขQwen3.5 27B is a dense model with 27.8 billion parameters and native vision-language capabilities using a linear attention mechanism for fast responses[1][5].
  • โ€ขIt scores 42 on the Artificial Analysis Intelligence Index, excelling in reasoning, knowledge, mathematics, and coding benchmarks like 72.4 on SWE-bench Verified, matching GPT-5 mini[1][4].
  • โ€ขThe model supports a 256K-262K context window, full tool use, and multimodal input under Apache 2.0 license allowing commercial use[1][5].
  • โ€ขAPI benchmarks show 100.2 t/s output speed and 1.38s TTFT, above average for similar open-weight models[1].
๐Ÿ“Š Competitor Analysisโ–ธ Show
FeatureQwen3.5 27B (Dense)Qwen3.5 35B-A3B (MoE)
Total Parameters27B35B
Active Parameters27B~3B
Intelligence/ReasoningHigh (top-tier coding, logic)Medium (faster but shallower)
Speed (local tests)~7-7.5 t/s at Q8~46 t/s at Q8
Benchmarks72.4 SWE-bench; 42 Intelligence IndexSurpasses prior 235B-A22B in some tasks

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ข27.8 billion parameters, all active per token in dense architecture, providing high reasoning density[1][4].
  • โ€ขIncorporates linear attention mechanism in native vision-language setup for balanced inference speed and performance[5].
  • โ€ขContext window of 256K-262K tokens with multimodal (vision) input and full tool use support[5].
  • โ€ขReleased under Apache 2.0; available via Ollama library[1][7].

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Qwen3.5 27B enables consumer-grade hardware for frontier reasoning
Its optimization to 100+ t/s on 2x3090s combined with 72.4 SWE-bench score matching GPT-5 mini democratizes high-intelligence local inference[1][4].
Dense 27B outperforms larger MoE in complex tasks
Superior active parameters lead to better coding reliability and roleplay consistency over Qwen3.5 35B-A3B's speed advantage[2][3].

โณ Timeline

2026-02
Qwen3.5 series release including 27B dense model with vision and 256K context
2026-02-15
Qwen3.5 Plus variant compared to 27B
2026-02-27
Qwen3.5-27B assessment highlighting agentic capabilities and benchmarks
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.