Marin Opens Live Training of a 535B MoE

💡A 535B model is being trained in public, exposing the data, code, metrics, and failures usually kept secret.
⚡ 30-Second TL;DR
What Changed
Marin 535B-A23B has approximately 535 billion total parameters, with about 23 billion parameters activated per token.
Why It Matters
Marin turns large-scale model training into a publicly inspectable experiment rather than a closed corporate process. If successful, it could provide researchers and independent builders with a stronger template for reproducible, community-driven foundation-model development.
What To Do Next
Track Marin’s GitHub repository and W&B dashboards, then reproduce its 4K-context MoE routing and token-dropping measurements on a smaller expert model.
Key Points
- •Marin 535B-A23B has approximately 535 billion total parameters, with about 23 billion parameters activated per token.
- •The project plans to use 80% of its 18.75 trillion tokens for pretraining and 20% for mid-training, followed by post-training.
- •Training runs on 11 NVIDIA GB200 NVL72 systems and is expected to require roughly three months of computation.
- •Marin’s open-lab process publishes GitHub Issues, code, pull requests, W&B metrics, data, recipes, failures, and intermediate changes.
- •The team reduced MoE token dropping to about 3% at a 4K context length using a pooled/wave expert-parallel approach.
🧠 Deep Insight
Background and context from public sources — not the original article. 5 sources cited.
🔑 Enhanced Key Takeaways
- •The project is spearheaded by Percy Liang, a prominent Stanford professor and the founder of Simile AI, marking a high-profile academic-industry collaboration.
- •The infrastructure utilizes a cluster of 11 NVIDIA GB200 NVL72 systems, which provides a total of 792 GB200 GPUs for the training run.
- •The project aims to challenge the dominance of closed-source proprietary models by establishing a new standard for 'radical transparency' in large-scale model development.
- •Training metrics, including real-time throughput, gradient norms, and token balancing statistics, are being exposed publicly via Weights & Biases to allow for community-led diagnostic analysis.
- •The initiative is positioned as a direct response to the 'black box' nature of current frontier models, seeking to provide researchers with granular data on training failures and intermediate model states.
📊 Competitor Analysis▸ Show
| Feature | Marin 535B-A23B | Llama 3.1 405B | Mixtral 8x22B |
|---|---|---|---|
| Transparency | Full (Live logs/data) | Weights/Inference only | Weights/Inference only |
| Architecture | 535B MoE | Dense | 176B MoE |
| Compute | 792x GB200 | H100 Cluster | H100 Cluster |
| Access | Open-Lab (Live) | Open Weights | Open Weights |
🛠️ Technical Deep Dive
- Architecture: Mixture-of-Experts (MoE) with 535B total parameters and 23B active parameters per token.
- Parallelism Strategy: Employs a pooled/wave expert-parallel approach to optimize communication overhead across the GB200 NVL72 interconnects.
- Optimization: Specifically engineered to minimize MoE token dropping to 3% at 4K context lengths.
- Infrastructure: 11 nodes of NVIDIA GB200 NVL72, leveraging NVLink Switch systems for high-bandwidth inter-GPU communication.
- Monitoring: Integration with Weights & Biases for real-time telemetry of training loss, gradient stability, and expert utilization rates.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (5)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
