🔥Stalecollected in 4m

Step 3.5 Flash Tops OpenClaw Post-OpenSource

Step 3.5 Flash Tops OpenClaw Post-OpenSource
PostLinkedIn
🔥Read original on 36氪
#open-source#benchmark-top#base-modelstep-3.5-flashstep-3.5-flashopenclaw

💡Open-source Step 3.5 Flash hits OpenClaw #1—new top model to deploy now

⚡ 30-Second TL;DR

What Changed

Full open-sourcing of Step 3.5 Flash yesterday

Why It Matters

Boosts open-source AI momentum, pressuring closed models and enabling faster innovation for practitioners.

What To Do Next

Download Step 3.5 Flash from its repo and benchmark against your current LLMs.

Who should care:Developers & AI Engineers

Key Points

  • Full open-sourcing of Step 3.5 Flash yesterday
  • Quickly reached #1 invocation on OpenClaw globally
  • New generation base model gaining massive traction

🧠 Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

🔑 Enhanced Key Takeaways

  • Step 3.5 Flash achieves 100-350 tokens/second throughput using 3-way Multi-Token Prediction (MTP-3), enabling real-time agentic reasoning that significantly outpaces competitors like DeepSeek V3.2 (33 tok/s on equivalent hardware)[2][4]
  • The model employs sparse Mixture-of-Experts architecture with 196B total parameters but only ~11B active per token, combining frontier reasoning capabilities with 11B-class inference efficiency and cost ($0.1/M input tokens)[2][3]
  • Step 3.5 Flash implements a 3:1 Sliding Window Attention ratio with 262K context window, reducing memory overhead while maintaining performance across both short queries and massive datasets—a key architectural advantage for long-context processing[1][2][3]
  • StepFun's open-source release strategy prioritizes architectural efficiency and practical utility over parameter count, positioning the model as purpose-built for coding assistants, agentic workflows, and GUI automation rather than general-purpose chat[3][5]
📊 Competitor Analysis▸ Show
FeatureStep 3.5 FlashDeepSeek V3.2GPT-5.2 High / Claude Opus 4.5 / Gemini 3 Pro
Total Parameters196B671BProprietary
Active Parameters~11B per tokenFull (sparse MoE)Proprietary
Throughput100-350 tok/s (peak 350 on coding)33 tok/s (Hopper)Proprietary
Context Window262KProprietaryProprietary
ArchitectureSparse MoE + MTP-3 + 3:1 SWASparse MoEProprietary
Pricing (via SiliconFlow)$0.1/M input, $0.3/M outputN/AProprietary (typically higher)
Open-Weight StatusFully open-sourcedProprietaryProprietary
Primary Use CaseAgentic coding, deep reasoning, long-contextGeneral-purposeGeneral-purpose
Reasoning DepthCompetitive with top-tier proprietaryCompetitiveFrontier
Known LimitationLonger generation trajectories than Gemini 3.0 Pro for comparable quality[5]N/AN/A

🛠️ Technical Deep Dive

  • Sparse MoE Architecture: 288 routed experts + 1 shared expert (always active), with Top-8 expert selection per token, enabling selective parameter activation[3]
  • Multi-Token Prediction (MTP-3): Predicts 4 tokens simultaneously in a single forward pass during both training and inference (unusual—most models disable MTP at inference)[1][4]
  • Hybrid Sliding Window Attention (SWA): 3:1 ratio integrating three SWA layers per full-attention layer with 512-token window size, significantly reducing VRAM overhead while maintaining 262K context consistency[1][3]
  • Model Dimensions: 45 transformer layers, 4,096 hidden size, 128,896 vocabulary size, 196.81B total parameters (196B backbone + 0.81B MTP head)[3]
  • Inference Optimization: Achieves 100-300 tok/s typical throughput, peaking at 350 tok/s for single-stream coding tasks on NVIDIA Hopper GPUs; optimized for DGX Spark deployment[2][3]
  • Tool-Calling & Agentic Features: Built-in support for tool-calling and agentic applications with auto-tool-choice capability[3][5]

🔮 Future ImplicationsAI analysis grounded in cited sources

Open-weight sparse MoE models will displace smaller proprietary models in cost-sensitive agentic applications
Step 3.5 Flash's combination of frontier reasoning, 11B-class inference cost, and open-source availability creates a compelling alternative to proprietary models for enterprises optimizing total-cost-of-ownership in agent deployment[2][3]
Multi-Token Prediction will become standard inference optimization for open-weight models
Step 3.5 Flash's successful deployment of MTP-3 during inference (contrary to industry practice) demonstrates that architectural innovations enabling 3-5x throughput gains will be rapidly adopted across competing open-source projects[1][4]
Specialized domain models (coding, research, agentic) will outcompete generalist models in their respective niches
StepFun's deliberate optimization for coding and agentic workflows, combined with acknowledged limitations in general-purpose chat, signals a market shift toward purpose-built models rather than one-size-fits-all foundation models[5]

Timeline

2026-03
Step 3.5 Flash fully open-sourced by StepFun; reaches #1 on OpenClaw global leaderboard within 24 hours of release
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.