Step 3.5 Flash Tops OpenClaw Post-OpenSource
💡Open-source Step 3.5 Flash hits OpenClaw #1—new top model to deploy now
⚡ 30-Second TL;DR
What Changed
Full open-sourcing of Step 3.5 Flash yesterday
Why It Matters
Boosts open-source AI momentum, pressuring closed models and enabling faster innovation for practitioners.
What To Do Next
Download Step 3.5 Flash from its repo and benchmark against your current LLMs.
Key Points
- •Full open-sourcing of Step 3.5 Flash yesterday
- •Quickly reached #1 invocation on OpenClaw globally
- •New generation base model gaining massive traction
🧠 Deep Insight
Background and context from public sources — not the original article. 9 sources cited.
🔑 Enhanced Key Takeaways
- •Step 3.5 Flash achieves 100-350 tokens/second throughput using 3-way Multi-Token Prediction (MTP-3), enabling real-time agentic reasoning that significantly outpaces competitors like DeepSeek V3.2 (33 tok/s on equivalent hardware)[2][4]
- •The model employs sparse Mixture-of-Experts architecture with 196B total parameters but only ~11B active per token, combining frontier reasoning capabilities with 11B-class inference efficiency and cost ($0.1/M input tokens)[2][3]
- •Step 3.5 Flash implements a 3:1 Sliding Window Attention ratio with 262K context window, reducing memory overhead while maintaining performance across both short queries and massive datasets—a key architectural advantage for long-context processing[1][2][3]
- •StepFun's open-source release strategy prioritizes architectural efficiency and practical utility over parameter count, positioning the model as purpose-built for coding assistants, agentic workflows, and GUI automation rather than general-purpose chat[3][5]
📊 Competitor Analysis▸ Show
| Feature | Step 3.5 Flash | DeepSeek V3.2 | GPT-5.2 High / Claude Opus 4.5 / Gemini 3 Pro |
|---|---|---|---|
| Total Parameters | 196B | 671B | Proprietary |
| Active Parameters | ~11B per token | Full (sparse MoE) | Proprietary |
| Throughput | 100-350 tok/s (peak 350 on coding) | 33 tok/s (Hopper) | Proprietary |
| Context Window | 262K | Proprietary | Proprietary |
| Architecture | Sparse MoE + MTP-3 + 3:1 SWA | Sparse MoE | Proprietary |
| Pricing (via SiliconFlow) | $0.1/M input, $0.3/M output | N/A | Proprietary (typically higher) |
| Open-Weight Status | Fully open-sourced | Proprietary | Proprietary |
| Primary Use Case | Agentic coding, deep reasoning, long-context | General-purpose | General-purpose |
| Reasoning Depth | Competitive with top-tier proprietary | Competitive | Frontier |
| Known Limitation | Longer generation trajectories than Gemini 3.0 Pro for comparable quality[5] | N/A | N/A |
🛠️ Technical Deep Dive
- Sparse MoE Architecture: 288 routed experts + 1 shared expert (always active), with Top-8 expert selection per token, enabling selective parameter activation[3]
- Multi-Token Prediction (MTP-3): Predicts 4 tokens simultaneously in a single forward pass during both training and inference (unusual—most models disable MTP at inference)[1][4]
- Hybrid Sliding Window Attention (SWA): 3:1 ratio integrating three SWA layers per full-attention layer with 512-token window size, significantly reducing VRAM overhead while maintaining 262K context consistency[1][3]
- Model Dimensions: 45 transformer layers, 4,096 hidden size, 128,896 vocabulary size, 196.81B total parameters (196B backbone + 0.81B MTP head)[3]
- Inference Optimization: Achieves 100-300 tok/s typical throughput, peaking at 350 tok/s for single-stream coding tasks on NVIDIA Hopper GPUs; optimized for DGX Spark deployment[2][3]
- Tool-Calling & Agentic Features: Built-in support for tool-calling and agentic applications with auto-tool-choice capability[3][5]
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (9)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- i-scoop.eu — Step 3 5 Flash
- siliconflow.com — Step 3.5 Flash Now on Siliconflow Deep Thinking at Flash Speed
- build.nvidia.com — Modelcard
- magazine.sebastianraschka.com — A Dream of Spring for Open Weight
- Hugging Face — Step 3.5 Flash
- youtube.com — Watch
- GitHub — Step 3.5 Flash
- pandaily.com — Step Fun Fully Open Sources Step 3 5 Flash
- openrouter.ai — Step 3.5 Flash:free
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.