Qwen3 9B runs 6+ t/s on Android phones

💡9B LLM hits 6t/s on phones—unlock mobile AI now
⚡ 30-Second TL;DR
What Changed
Runs at q4_0 on S25 Ultra with 12GB RAM
Why It Matters
Shows large LLMs like 9B models are feasible on high-end Android phones, enabling edge AI apps without cloud dependency.
What To Do Next
Quantize Qwen3 9B to q4_0 and benchmark on your Android device with llama.cpp.
Key Points
- •Runs at q4_0 on S25 Ultra with 12GB RAM
- •>6 tokens/s generation speed achieved
- •Utilizes Hexagon NPU for acceleration
- •Tested on Snapdragon 8 Elite chip
🧠 Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
🔑 Enhanced Key Takeaways
- •Qwen3.5-9B employs a hybrid architecture combining Gated DeltaNet with Sparse MoE, using a 3:1 ratio of linear to softmax attention for reduced memory and compute costs[1][3][7].
- •The model supports a 262K token context window and native multimodality, processing text and visual data in a unified latent space[2][6][7].
- •Qwen3.5-9B outperforms prior Qwen3-30B on MMLU and math benchmarks like GSM8K/MATH due to Scaled RL training[1][3].
- •At Q4 quantization, it runs on ~5GB RAM with CPU-only inference at 20-30 t/s, and higher speeds like 115-167 t/s in optimized desktop tests[3][5].
🛠️ Technical Deep Dive
- •Hybrid architecture: Gated DeltaNet + Sparse Mixture-of-Experts (MoE), with 3:1 linear attention to softmax attention ratio, reducing computational cost for long contexts[3][7].
- •Parameter count: 9.65B, natively multimodal vision-language model[6][7].
- •Training: Scaled Reinforcement Learning (RL) optimizes logical reasoning, closing gap with 30B+ models[1][2][3].
- •Context window: 262K tokens[6].
- •Quantized (Q4 GGUF): ~5GB RAM footprint, supports CUDA/NVIDIA GPU, Metal/Apple Silicon, or CPU inference[3].
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- marktechpost.com — Alibaba Just Released Qwen 3 5 Small Models a Family of 0 8b to 9b Parameters Built for on Device Applications
- mlq.ai — Alibaba Releases Open Source Qwen35 Small Models for Edge Devices
- oflight.co.jp — Qwen35 9b Complete Guide
- dev.to — Qwen 3 Benchmarks Comparisons Model Specifications and More 4hoa
- youtube.com — Watch
- rdworldonline.com — We Installed Alibabas 9 Billion Parameter Qwen3 5 9b AI on a Usb Hard Drive It Said It Was Made by Google
- llm-stats.com — Qwen3.5 9b
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.