🐯Freshcollected in 11m

Marin Opens Live Training of a 535B MoE

Marin Opens Live Training of a 535B MoE
PostLinkedIn
🐯Read original on 虎嗅
#mixture-of-experts#foundation-models#open-trainingmarin-535b-a23bmarinstanford crfmnvidiagb200andrew ng

💡A 535B model is being trained in public, exposing the data, code, metrics, and failures usually kept secret.

⚡ 30-Second TL;DR

What Changed

Marin 535B-A23B has approximately 535 billion total parameters, with about 23 billion parameters activated per token.

Why It Matters

Marin turns large-scale model training into a publicly inspectable experiment rather than a closed corporate process. If successful, it could provide researchers and independent builders with a stronger template for reproducible, community-driven foundation-model development.

What To Do Next

Track Marin’s GitHub repository and W&B dashboards, then reproduce its 4K-context MoE routing and token-dropping measurements on a smaller expert model.

Who should care:Researchers & Academics

Key Points

  • Marin 535B-A23B has approximately 535 billion total parameters, with about 23 billion parameters activated per token.
  • The project plans to use 80% of its 18.75 trillion tokens for pretraining and 20% for mid-training, followed by post-training.
  • Training runs on 11 NVIDIA GB200 NVL72 systems and is expected to require roughly three months of computation.
  • Marin’s open-lab process publishes GitHub Issues, code, pull requests, W&B metrics, data, recipes, failures, and intermediate changes.
  • The team reduced MoE token dropping to about 3% at a 4K context length using a pooled/wave expert-parallel approach.

🧠 Deep Insight

Background and context from public sources — not the original article. 5 sources cited.

🔑 Enhanced Key Takeaways

  • The project is spearheaded by Percy Liang, a prominent Stanford professor and the founder of Simile AI, marking a high-profile academic-industry collaboration.
  • The infrastructure utilizes a cluster of 11 NVIDIA GB200 NVL72 systems, which provides a total of 792 GB200 GPUs for the training run.
  • The project aims to challenge the dominance of closed-source proprietary models by establishing a new standard for 'radical transparency' in large-scale model development.
  • Training metrics, including real-time throughput, gradient norms, and token balancing statistics, are being exposed publicly via Weights & Biases to allow for community-led diagnostic analysis.
  • The initiative is positioned as a direct response to the 'black box' nature of current frontier models, seeking to provide researchers with granular data on training failures and intermediate model states.
📊 Competitor Analysis▸ Show
FeatureMarin 535B-A23BLlama 3.1 405BMixtral 8x22B
TransparencyFull (Live logs/data)Weights/Inference onlyWeights/Inference only
Architecture535B MoEDense176B MoE
Compute792x GB200H100 ClusterH100 Cluster
AccessOpen-Lab (Live)Open WeightsOpen Weights

🛠️ Technical Deep Dive

  • Architecture: Mixture-of-Experts (MoE) with 535B total parameters and 23B active parameters per token.
  • Parallelism Strategy: Employs a pooled/wave expert-parallel approach to optimize communication overhead across the GB200 NVL72 interconnects.
  • Optimization: Specifically engineered to minimize MoE token dropping to 3% at 4K context lengths.
  • Infrastructure: 11 nodes of NVIDIA GB200 NVL72, leveraging NVLink Switch systems for high-bandwidth inter-GPU communication.
  • Monitoring: Integration with Weights & Biases for real-time telemetry of training loss, gradient stability, and expert utilization rates.

🔮 Future ImplicationsAI analysis grounded in cited sources

The project will force a shift in industry standards for open-source model releases.
By publishing intermediate failures and training recipes, Marin sets a precedent that makes 'weights-only' releases appear insufficient for scientific reproducibility.
MoE token-dropping mitigation will become a standard benchmark for large-scale training.
The project's focus on maintaining a 3% drop rate at 4K context highlights a critical bottleneck in MoE scaling that other developers will be forced to address.

Timeline

2026-05
Simile AI and Stanford announce the Marin initiative to build a transparent 535B MoE model.
2026-07
Finalization of the 18.75 trillion token dataset and procurement of the 11-node GB200 cluster.
2026-08
Official commencement of live training and public opening of the W&B dashboard and GitHub repository.

📎 Sources (5)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. 36kr.com
  2. htx.com
  3. latent.space
  4. wiredframe.xyz
  5. wandb.ai
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.