🏠Stalecollected in 8h

Alibaba open-sources Qwen3.5 tiny models

Alibaba open-sources Qwen3.5 tiny models
PostLinkedIn
🏠Read original on IT之家
#open-source#small-models#multimodal#edge-deploymentqwen3.5-0.8b/2b/4b/9balibabaqwen3.5hugging-face

💡Tiny open models from 0.8B-9B rival giants for edge/server AI needs.

⚡ 30-Second TL;DR

What Changed

Models: Qwen3.5-0.8B/2B for ultra-light edge/IoT deployment

Why It Matters

Democratizes high-perf AI for resource-limited apps, accelerating edge AI and agent development across devices.

What To Do Next

Download Qwen3.5-0.8B from Hugging Face and benchmark on edge hardware.

Who should care:Developers & AI Engineers

Key Points

  • Models: Qwen3.5-0.8B/2B for ultra-light edge/IoT deployment
  • Qwen3.5-4B as strong base for lightweight AI agents
  • Qwen3.5-9B delivers GPT-4o-level perf in compact size
  • Native multimodal training and latest architecture

🧠 Deep Insight

Background and context from public sources — not the original article. 6 sources cited.

🔑 Enhanced Key Takeaways

  • Qwen3.5 series flagship is a 397B parameter model using hybrid MoE and Gated Delta Networks architecture, activating only 17B parameters per forward pass for optimized efficiency[1][3].
  • Qwen family has achieved over 700 million downloads on Hugging Face with more than 180,000 derivative models across 119 languages[2].
  • Qwen3.5-Plus hosted version offers 1M context window and built-in tools via Alibaba Cloud Model Studio[5].
  • Medium series released on February 24, 2026, includes Qwen3.5-35B-A3B which surpasses prior 235B model despite using only 3B active parameters[4].
📊 Competitor Analysis▸ Show
ModelActive ParamsKey Benchmark WinsCost/Perf Improvement
Qwen3.5 (flagship)17B (of 397B)Beats GPT-5.2, Claude Opus 4.5, Gemini 3 Pro60% cheaper, 8x better large workloads vs predecessor[1]
Qwen3.5-35B-A3B3B (of 35B)Surpasses Qwen3-235B-A22BFaster generation (1/6th time of Claude Sonnet 4.6)[4]

🛠️ Technical Deep Dive

  • Hybrid architecture combines Mixture of Experts (MoE) and Gated Delta Networks for native vision-language model (VLM) with UI navigation and reasoning[3].
  • Qwen3.5 flagship: ~400B total parameters, ~17B active per pass, supports coding, visual reasoning, chat, complex search; NVIDIA NIM optimized for deployment[3].
  • Medium models like Qwen3.5-35B-A3B (3B active of 35B total) and Qwen3.5-122B-A10B use Gated DeltaNet + MoE hybrid, enabling linear attention at scale and fitting on 8GB+ VRAM GPUs[4].
  • Supports OpenAI-compatible tool calling; 1M context window in hosted Qwen3.5-Plus[3][5].

🔮 Future ImplicationsAI analysis grounded in cited sources

Qwen3.5 tiny models will exceed 100 million downloads within 6 months
Qwen family already has 700 million downloads and rapid adoption like Qwen App reaching 30 million MAU in 23 days demonstrates strong community momentum[2].
Hybrid MoE architectures will become standard for edge AI by end of 2026
Qwen3.5 proves 3B active params outperform prior 22B models, shifting focus from scale to efficient architectures usable on compact hardware[4].
Alibaba Cloud will capture 20% more enterprise AI market share
60% cost reduction and 8x workload efficiency position Qwen3.5 as leader in inference cost per capability for agentic deployments[1].

Timeline

2025-09
Unveiled Qwen3-Max at Apsara conference for coding and agentic tasks
2025-11
Launched Qwen App reaching 30 million MAU in 23 days
2026-02-16
Released Qwen3.5 flagship 397B model in open-weight and hosted versions
2026-02-24
Released Qwen3.5 medium series including 35B-A3B model
2026-03-02
Open-sourced four tiny Qwen3.5 models (0.8B, 2B, 4B, 9B) on ModelScope and Hugging Face
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: IT之家

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.