๐Ÿฆ™Stalecollected in 71m

Intern-S2-Preview: Efficient 35B Scientific Multimodal Model

Intern-S2-Preview: Efficient 35B Scientific Multimodal Model
PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA

๐Ÿ’กA 35B model that rivals trillion-scale performance in scientific tasks using advanced CoT compression.

โšก 30-Second TL;DR

What Changed

35B parameter model continued from Qwen3.5

Why It Matters

This model bridges the gap between general reasoning and specialized scientific research, making high-level material science accessible to smaller-scale deployments.

What To Do Next

Evaluate Intern-S2-Preview for your next material science or scientific agent workflow to leverage its specialized domain performance.

Who should care:Researchers & Academics

Key Points

  • โ€ข35B parameter model continued from Qwen3.5
  • โ€ขSpecialized in scientific task scaling and material crystal structure generation
  • โ€ขUses shared-weight MTP and CoT compression for efficient RL reasoning
  • โ€ขAchieves performance comparable to trillion-scale models in scientific domains

๐Ÿง  Deep Insight

Web-grounded analysis with 12 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขIntern-S2-Preview is developed by Shanghai AI Laboratory, which is also responsible for other 'Intern' series models like InternLM, InternLM2, and Intern-S1.
  • โ€ขIt is the first open-source model to offer both material crystal structure generation capabilities and strong general capabilities, including enhanced spatial modeling for small-molecule structures and real-valued prediction modules.
  • โ€ขThe model achieves performance comparable to the trillion-scale Intern-S1-Pro on various core professional scientific tasks, despite utilizing only 35 billion parameters.
  • โ€ขThe base model, Qwen3.5, features a unified vision-language foundation achieved through 'Early Fusion' training on multimodal tokens, an efficient hybrid architecture combining Gated Delta Networks with sparse Mixture-of-Experts, and scalable reinforcement learning generalization.
  • โ€ขThe CoT compression techniques employed by Intern-S2-Preview aim to reduce response verbosity while maintaining strong reasoning, with research indicating such methods can cut response length by 20-40% without degrading accuracy.

๐Ÿ› ๏ธ Technical Deep Dive

  • Base Model: Intern-S2-Preview is a continuation of the Qwen3.5 model.
  • Architecture (inherited from Qwen3.5): Employs an efficient hybrid architecture that combines Gated Delta Networks with a sparse Mixture-of-Experts (MoE) design.
  • Multimodality: Features a Unified Vision-Language Foundation achieved through 'Early Fusion' training, where vision and language are interwoven at the foundational level.
  • Reinforcement Learning (RL): Utilizes efficient RL reasoning by adopting shared-weight Multi-Token Prediction (MTP) with KL loss. This approach aims to minimize the discrepancy between training and inference behavior, thereby enhancing the MTP acceptance rate and token generation speed.
  • CoT Compression: Incorporates Chain-of-Thought (CoT) compression techniques to shorten model responses while preserving robust reasoning capabilities. This is supported by methods like Fine-grained Group policy Optimization (FGO), which refines group responses by subdividing them and assigning weights based on length and entropy, and techniques that prune low-entropy intermediate steps.
  • MTP Mechanism: MTP involves a smaller 'drafter' model that rapidly generates candidate tokens, which are then verified by the larger target model in a single forward pass, accelerating inference without compromising output quality. The drafter reuses the target model's key-value cache and activations.
  • Scientific Specialization: The model scales hundreds of professional scientific tasks across its full-chain training pipeline (pre-training to RL). It specifically strengthens spatial modeling for small-molecule structures and integrates real-valued prediction modules.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Intern-S2-Preview will significantly accelerate scientific discovery in materials science and chemistry.
Its unique capability for material crystal structure generation and enhanced spatial modeling for small molecules could drastically speed up research and development cycles in these fields.
The model will increase the accessibility of advanced scientific AI to a broader range of researchers and institutions.
By achieving performance comparable to trillion-scale models with only 35 billion parameters, Intern-S2-Preview makes powerful scientific AI more deployable on less resource-intensive hardware.
The integrated efficient RL reasoning with MTP and CoT compression will drive further advancements in LLM efficiency.
These techniques could establish new benchmarks for efficient and accurate reasoning, influencing the design and optimization of future large language models, both specialized and general-purpose.

โณ Timeline

2023-06
InternLM, a 104B multilingual foundational language model, is presented by Shanghai AI Lab and SenseTime.
2024-01
InternLM2-1.8B, a smaller version of the InternLM series, is released.
2025-08
Intern-S1, a scientific multimodal Mixture-of-Experts (MoE) model with 241B total parameters (28B activated), is introduced.
2026-02
Alibaba's Qwen team launches the flagship Qwen3.5-397B-A17B model, part of the Qwen3.5 series.
2026-03
The Qwen3.5 Medium Model Series, including Qwen3.5-35B-A3B, is released, emphasizing efficiency and multimodal understanding.
2026-05-15
Intern-S2-Preview, an efficient 35B scientific multimodal foundation model continued from Qwen3.5, is introduced.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—