🦙Stalecollected in 52m

Qwen3.5 Small Few-Shot Benchmarks Surprise

Qwen3.5 Small Few-Shot Benchmarks Surprise
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA
#few-shot#benchmarks#small-modelsqwen-3.5-smallqwen3.5lm-studio

💡Few-shot hurts Qwen3.5 0.8B code perf; 4B shines—pick wisely

⚡ 30-Second TL;DR

What Changed

0.8B code fix collapses from 67% zero-shot to 33% with 1+ examples

Why It Matters

Reveals few-shot pitfalls for tiny models on code tasks, guiding size selection for efficient deployments.

What To Do Next

Benchmark your task with 0/1/2-shot on Qwen3.5 0.8B before adding examples.

Who should care:Researchers & Academics

Key Points

  • 0.8B code fix collapses from 67% zero-shot to 33% with 1+ examples
  • Classification: 0.8B improves to 100% at 8-shot; larger perfect zero-shot
  • 4B stable across code/classification/summarization; sweet spot for speed
  • 9B summarization low due to thinking artifacts
  • Tested on LM Studio with TF-IDF example selection

🧠 Deep Insight

Background and context from public sources — not the original article. 6 sources cited.

🔑 Enhanced Key Takeaways

  • Qwen3.5 series introduces Mixture-of-Experts (MoE) architecture with models like 397B-A17B (397B total/17B active parameters), enabling 19x faster text generation than Qwen3-Max while matching reasoning and coding performance.[1][2]
  • Qwen3.5 excels in multimodal agentic tasks, scoring 67.5 on ERQA embodied reasoning (vs. Qwen3-VL's 52.5) and 90.8% on OmniDocBench document recognition, surpassing GPT-5.2 and Claude Opus 4.5.[1]
  • Models support 1M token context window, 250k vocabulary for token efficiency (10-60% cost reduction across 201 languages), and accept text, images, video inputs under Apache 2.0 license on Hugging Face.[1][3][4]
📊 Competitor Analysis▸ Show
ModelParametersKey BenchmarksPricing (Input/Output per 1M tokens)
Qwen3.5-397B-A17B397B/17B activeIntelligence Index: 45; GDPval-AA ELO: 1221Open-source (free); Flash API: $0.10/$0.40
GLM-5744B/40BIntelligence Index: 50Not specified
Kimi K2.51T/32BIntelligence Index: >45Not specified
DeepSeek V3.2671B/37BCompetitive agenticNot specified

🛠️ Technical Deep Dive

  • Heterogeneous infrastructure: Vision and language components trained separately but simultaneously for ~100% training throughput vs. pure text models.[1]
  • Asynchronous reinforcement learning with FP8 compression and speculative decoding enables 3-5x faster agent skill acquisition (e.g., UI clicking, multi-step tasks).[1]
  • Multi-token prediction guesses several future words per step, paired with 250k vocabulary for 10-60% token cost reduction across 201 languages.[1]
  • MoE variants like Qwen3.5-397B-A17B (397B total, 17B active), Qwen3.5-122B-A10B, Qwen3.5-35B-A3B, and Qwen3.5-27B support 1M context and multimodal inputs.[2][3][4]

🔮 Future ImplicationsAI analysis grounded in cited sources

Qwen3.5 small models will dominate local inference on consumer hardware by mid-2026
4B model's stability across tasks combined with MoE efficiency matches larger models' performance at 19x speed, ideal for edge deployment.[1]
Open MoE agents will reduce proprietary API reliance by 30% in enterprise workflows
Apache 2.0 licensing, low-cost Flash API ($0.10/$0.40 per 1M), and superior agentic benchmarks enable commercial customization without vendor lock-in.[3][4]
Hallucination rates in Qwen3.5 remain above peers at 88%, limiting high-stakes use
AA-Omniscience Index of -32 shows accuracy gains but persistent high hallucination vs. competitors, despite agentic improvements.[2]

Timeline

2026-02
Qwen3.5 series launched with initial 397B-A17B model, focusing on efficient MoE for multimodal agents.[3]
2026-02
Expanded lineup released including Qwen3.5-Flash, 35B-A3B, 122B-A10B, and 27B variants outperforming predecessors.[3]
2026-03
Community benchmarks on LM Studio reveal few-shot behaviors in small Qwen3.5 models (0.8B-9B).[]
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.