🔥Stalecollected in 14m

Step3.5 Flash Tops OpenClaw for 3 Days

Step3.5 Flash Tops OpenClaw for 3 Days
PostLinkedIn
🔥Read original on 36氪

💡中國模型連3天霸榜全球API調用,開發者必試其效能優勢!

⚡ 30-Second TL;DR

What Changed

Step3.5 Flash連續三天蟬聯OpenClaw全球調用量榜首

Why It Matters

此成就凸顯中國LLM在全球API使用中的崛起,可能加速開發者轉向高效本土模型,影響OpenRouter生態競爭格局。

What To Do Next

透過OpenRouter API測試Step3.5 Flash模型的高速推理效能

Who should care:Developers & AI Engineers

Key Points

  • Step3.5 Flash連續三天蟬聯OpenClaw全球調用量榜首
  • OpenRouter數據顯示其為全球最大AI模型API聚合平台
  • 自2026年3月起穩居前三,與Kimi K2.5及MiniMax M2.5並列

🧠 Deep Insight

Background and context from public sources — not the original article. 9 sources cited.

🔑 Enhanced Key Takeaways

  • Step 3.5 Flash is a 196B parameter sparse Mixture-of-Experts (MoE) model with only ~11B active parameters per token, enabling high throughput of 100-300 tokens/sec and peaking at 350 tok/s for coding tasks on NVIDIA Hopper GPUs.[1][2][4]
  • It supports a 256K-262K context window via a hybrid 3:1 Sliding Window Attention (SWA) ratio, reducing compute overhead for long-context tasks like massive datasets and codebases.[2][4]
  • The model excels in agentic benchmarks, scoring 74.4% on SWE-bench Verified, 51.0% on Terminal-Bench 2.0, and 88.2 on τ²-Bench, matching or exceeding top closed-source systems.[2][3][4]
📊 Competitor Analysis▸ Show
Feature/BenchmarkStep 3.5 FlashDeepSeek V3.2Kimi K2.5
Total Parameters196B (11B active MoE)671B (37B active MoE)Not specified
Throughput (tok/s, 128k ctx)10033Not specified
SWE-bench Verified74.4%Not specified73.1%
Terminal-Bench 2.051.0%Not specified46.4%
τ²-Bench88.2Not specified80.3

🛠️ Technical Deep Dive

  • Architecture: Transformer-based sparse MoE with 196.81B total parameters (196B backbone + 0.81B MTP head), ~11B active per token; 45 layers, hidden size 4,096, vocabulary 128,896.[4]
  • Experts: 288 routed experts + 1 shared expert (always active), top-8 selection per token.[4]
  • Attention: 3:1 SWA ratio (three sliding-window layers of size 512 per full-attention layer) for efficient 256K context.[2][4]
  • Prediction: MTP-3 (3-way Multi-Token Prediction) used in both training and inference for accelerated generation.[1][2][4]

🔮 Future ImplicationsAI analysis grounded in cited sources

Step 3.5 Flash will capture >20% of OpenRouter agentic API market share by mid-2026
Its top ranking on OpenClaw for three days, combined with open-source availability and superior agent benchmarks, positions it to attract developers building long-horizon agents over slower competitors.[1][3][8]
MoE models like Step 3.5 Flash will reduce inference costs by 3x for 256K+ contexts
The 3:1 SWA and sparse activation enable high throughput at long contexts, outperforming dense models like DeepSeek V3.2 in tokens/sec on same hardware.[1][2][4]

Timeline

2026-01
StepFun releases Step 3.5 Flash as open-source MoE model with frontier agentic performance.[1]
2026-02
Model integrated into platforms like SiliconFlow, NVIDIA NIM, and OpenRouter for API access.[2][4][8]
2026-03
Step 3.5 Flash tops OpenClaw global calls for three consecutive days per OpenRouter data.[article]
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.