535B Model Training Goes Fully Open

💡See how a 535B model documents months of training with code, data, and Loss curves in public.
⚡ 30-Second TL;DR
What Changed
The model has 535 billion parameters.
Why It Matters
Public access to training artifacts could improve reproducibility and help researchers study large-scale model training. It may also raise expectations for greater transparency in future foundation-model projects.
What To Do Next
Review the released training code, dataset documentation, and loss curves to compare the project’s scaling behavior with your own LLM experiments.
Key Points
- •The model has 535 billion parameters.
- •Training was livestreamed publicly for three months.
- •The project released its code, training data, and loss curves.
🧠 Deep Insight
Background and context from public sources — not the original article. 4 sources cited.
🔑 Enhanced Key Takeaways
- •The model is officially designated as Marin 535B-A23B and originated from the Stanford University Center for Research on Foundation Models (CRFM).
- •The training regimen involves processing a total of 18.75 trillion tokens, split between 80% pre-training and 20% mid-training phases.
- •Infrastructure deployment utilizes 11 sets of NVIDIA GB200 NVL72 systems to achieve the required computational throughput.
- •The project represents a total computational expenditure of approximately 2.7×10²⁴ FLOPs.
- •The initiative was formally announced in May 2025 under the leadership of Percy Liang and David Hall to promote radical transparency in large-scale model development.
📊 Competitor Analysis▸ Show
| Feature | Marin 535B-A23B | Proprietary LLMs (e.g., GPT-4/Claude) | Open Weights Models (e.g., Llama 3) |
|---|---|---|---|
| Transparency | Full (Data/Curves/Code) | Closed | Partial (Weights only) |
| Training Process | Publicly Livestreamed | Opaque | Opaque |
| Scale | 535B Parameters | Varies (Often >1T) | Varies (Up to 405B) |
| Pricing | Open Research | Commercial API | Open Weights/Commercial License |
🛠️ Technical Deep Dive
- Architecture: Large-scale transformer-based foundation model with 535 billion parameters.
- Compute Hardware: 11 clusters of NVIDIA GB200 NVL72 systems.
- Token Throughput: 18.75 trillion tokens total.
- Computational Load: 2.7×10²⁴ FLOPs.
- Training Methodology: Real-time disclosure of loss curves, data recipes, and model configurations.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (4)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: InfoQ中国 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.

