๐ŸผStalecollected in 80m

Stepfun Open-Sources Step 3.7 Flash LLM for Agents

Stepfun Open-Sources Step 3.7 Flash LLM for Agents
PostLinkedIn
๐ŸผRead original on Pandaily

๐Ÿ’กHigh-speed 196B MoE model with native tool-calling designed specifically for agentic AI workflows.

โšก 30-Second TL;DR

What Changed

196B-parameter sparse MoE architecture optimized for agent workflows

Why It Matters

This release provides developers with a high-performance, open-source alternative for building responsive agentic systems. Its speed and tool-calling focus could significantly reduce latency in real-time AI applications.

What To Do Next

Download the Step 3.7 Flash weights and benchmark its tool-calling latency against your current production model to evaluate potential performance gains.

Who should care:Developers & AI Engineers

Key Points

  • โ€ข196B-parameter sparse MoE architecture optimized for agent workflows
  • โ€ขHigh-speed inference performance reaching 400 tokens/s
  • โ€ขNative support for complex tool-calling and agentic tasks

๐Ÿง  Deep Insight

Web-grounded analysis with 18 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขStep 3.7 Flash is a multimodal vision-language model, integrating a 196B-parameter language backbone with a 1.8B-parameter vision encoder for native image and video understanding, enabling it to process diverse visual inputs like UI elements, charts, and documents.
  • โ€ขThe model operates with a sparse Mixture-of-Experts (MoE) architecture, activating approximately 11 billion parameters per token during inference, which contributes to its high efficiency and a reported throughput of up to 400 tokens per second.
  • โ€ขIt features selectable reasoning levels (low, medium, and high), allowing developers to dynamically adjust the trade-off between inference speed, operational cost, and the depth of cognitive processing required for specific agentic tasks.
  • โ€ขStep 3.7 Flash demonstrates strong performance in agentic reliability and multimodal perception, leading the ClawEval-1.1 benchmark with a score of 67.1 and achieving top-tier visual intelligence on SimpleVQA (Search) with 79.2.
  • โ€ขThe model is designed for broad compatibility with major agent frameworks, including Claude Code, KiloCode, RooCode, OpenCode, Hermes Agent, and OpenClaw, and supports MCP and Skills tool-calling protocols, significantly reducing integration complexity for developers.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature / ModelStep 3.7 FlashClaude Opus 4.6/4.8GPT-5.4/5.5 ProGrok 4.3DeepSeek V4 FlashGemini 3.1 Pro/3.5 Flash
Architecture196B sparse MoE (11B active), Vision-Language---284B (MoE implied)MoE implied
Parameters (Active)11B active-----
Inference Speed400 tokens/s-----
Context Window256K tokens1M tokens1M+ tokens1M tokens1M tokens1M tokens (Gemini 1.5 Pro)
MultimodalNative image/video understandingText, image, file inputsText, image inputsText, image inputs-Multimodal (video/voice/text)
Agentic FocusOptimized for agent workflows, tool-calling, coding, searchHighly autonomous agents, long-horizon work, multi-step reasoning, complex codingLong-horizon problem solving, agentic coding, multi-step workflowsAgentic workflows, instruction-followingAgentic Coding, Agentic Browser-UseConsumer agents, video/voice interaction, complex logic puzzles
ClawEval-1.1 Score67.170.8 (Opus 4.6)60.3 (GPT 5.4)-57.857.8 (Gemini 3.1 Pro)
SWE-Bench Verified Score73.7% (with Advisor Mode)78.7% (Opus 4.6)----
Pricing (API)Input $0.20/M, Output $1.15/M--Tiered pricing for >200k tokensInput $0.1/M, Output $0.4/M (Qwen3-Coder)Cost-efficient (Flash models)
Open-SourceYes (Apache 2.0)NoNoNoYesNo

๐Ÿ› ๏ธ Technical Deep Dive

  • Architecture: Step 3.7 Flash is a 198B-parameter sparse Mixture-of-Experts (MoE) vision-language model. It combines a 196B-parameter language backbone with a 1.8B-parameter vision encoder.
  • Active Parameters: During inference, approximately 11 billion parameters are activated per token, contributing to its efficiency.
  • Multimodal Capabilities: The integrated vision encoder enables native image and video understanding, allowing the model to process and convert complex visual information from UIs, charts, and documents into structured outputs or code.
  • Context Window: The model supports a 256K token context window.
  • Inference Optimization: It is engineered for high-frequency production workloads, achieving a throughput of up to 400 tokens per second. Optimizations include MTP-3 Acceleration (3-way Multi-Token Prediction) and a 3:1 Sliding Window Attention ratio for efficient long-context processing.
  • Reasoning Levels: Developers can choose from three selectable reasoning levels (low, medium, high) to balance speed, cost, and cognitive depth.
  • Tool Calling & Orchestration: Features reliable native tool calling, allowing it to stably invoke APIs, browsers, terminals, Office tools, and external systems in multi-turn agent workflows. It is compatible with major agent frameworks and protocols like MCP and Skills.
  • Advisor Mode: To enhance quality without sacrificing efficiency, Step 3.7 Flash supports an 'Advisor Mode.' In this mode, the model acts as an executor, escalating to a larger advisor model only at critical inflection points (e.g., complex planning or error recovery) to maintain high performance at reduced cost.
  • Deployment: The model is open-sourced and available on platforms like GitHub, Hugging Face, and ModelScope. It can be deployed via Stepfun's API or with open-source frameworks such as SGLang, NVIDIA TensorRT-LLM, and vLLM on NVIDIA-accelerated infrastructure (Blackwell, Hopper GPUs).

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Step 3.7 Flash will accelerate the adoption of AI agents in enterprise and production environments.
Its combination of high-speed inference, multimodal capabilities, robust tool-calling, and open-source availability lowers the barrier to entry for developing and deploying sophisticated AI agents in real-world applications.
The 'Advisor Mode' feature will become a standard optimization strategy for balancing performance and cost in agentic workflows.
By allowing a smaller, efficient model to handle most tasks and only escalating to a larger, more capable (and expensive) model for complex planning or error recovery, it offers a practical solution for cost-effective, high-performance agent deployment.
Stepfun's focus on multimodal and agentic models, coupled with its open-source strategy, will strengthen its position as a key player in the global AI market, particularly in China.
The release of advanced open-source models like Step 3.7 Flash, along with its existing multimodal portfolio and strategic partnerships, positions Stepfun to drive innovation and capture market share in the rapidly evolving AI landscape.

โณ Timeline

2023-04
StepFun founded by former Microsoft employees in Shanghai, China.
2024-03
Stepfun launched its 'Step series' of general-purpose large models.
2024-07
StepFun officially launched Step-2 (trillion-parameter LLM), Step-1.5V (multimodal), and Step-1X (image generation) at the World Artificial Intelligence Conference.
2024-11
Step-2 topped the LiveBench rankings, becoming the first Chinese-language model to break into the top 10.
2024-12
Stepfun secured a funding round of several hundred million dollars from investors including Tencent and Shanghai State-owned Capital Investment.
2025-02
Stepfun and Geely jointly announced the open-sourcing of two multimodal large models, Step-Video-T2V and Step-Audio.
2025-07
StepFun released Step 3, with optimization efforts for domestic chips announced through the 'Model-Chip Ecosystem Innovation Alliance'.
2026-02
Step-3.5-Flash, a 196B-parameter mixture-of-experts model with 11 billion active parameters, was released under the Apache 2.0 license.
2026-04
StepFun reportedly began dismantling its offshore red chip structure to pave the way for a planned Initial Public Offering (IPO) in Hong Kong.
2026-05-29
Stepfun open-sources Step 3.7 Flash LLM for agents.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily โ†—