🦙Stalecollected in 62m

Qwen 0.8B Runs on Old S10E at 12 t/s

Qwen 0.8B Runs on Old S10E at 12 t/s
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA

💡Tiny Qwen 0.8B hits 12 t/s on 7yo phone—edge AI now feasible on old hardware.

⚡ 30-Second TL;DR

What Changed

Qwen launches 0.8B model

Why It Matters

Proves tiny LLMs viable for mobile/edge AI, lowering hardware barriers for on-device inference and privacy-focused apps.

What To Do Next

Compile Qwen3.5-0.8B with llama.cpp in Termux on your Android phone.

Who should care:Developers & AI Engineers

Key Points

  • Qwen launches 0.8B model
  • Runs at 12 tokens/sec on Samsung S10E
  • Uses llama.cpp and Termux with C library fixes
  • Capable of conversations and serious tasks

🧠 Deep Insight

Background and context from public sources — not the original article. 5 sources cited.

🔑 Enhanced Key Takeaways

  • Qwen3.5-0.8B is part of Alibaba's newly released compact model series (0.8B, 2B, 4B, 9B) announced on March 2, 2026, all in dense format with open-weight licensing under Apache 2.0[1][2]
  • The model natively supports multimodal processing (text, images, and video) with 262K context window and 201 language coverage, achieving MathVista 62.2 and OCRBench 74.5 benchmarks despite sub-1B parameter count[1]
  • Memory efficiency enables deployment across diverse edge devices: ~1.6GB VRAM at full precision (BF16), ~0.8GB at 8-bit quantization, and ~0.5GB at 4-bit quantization for phone and Raspberry Pi deployment[1]

🛠️ Technical Deep Dive

  • Architecture: Gated DeltaNet + Gated Attention hybrid (3:1 ratio) with 0.8B total parameters, all dense (no mixture-of-experts)[1]
  • Multimodal capabilities: Processes screenshots, documents, and basic video understanding natively within a single model[1]
  • Deployment flexibility: Available as base model (Qwen3.5-0.8B-Base) and seven quantized variants across HuggingFace and ModelScope[1]
  • Context window: 262K tokens, enabling long-form document and conversation processing[1]
  • Language support: Covers 201 languages for multilingual inference[1]

🔮 Future ImplicationsAI analysis grounded in cited sources

Sub-1B models will become the standard for on-device AI applications
Qwen0.8B's multimodal capabilities at phone-deployable sizes (0.5GB 4-bit) suggest the performance-efficiency frontier has shifted, making larger edge models obsolete for many use cases.
Open-weight compact models will accelerate edge AI adoption in resource-constrained markets
Apache 2.0 licensing combined with Raspberry Pi and IoT compatibility removes licensing and hardware barriers for developers in emerging markets.

Timeline

2026-03
Alibaba releases Qwen 3.5 compact series (0.8B, 2B, 4B, 9B) with open-weight Apache 2.0 licensing
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.