Olmo-core 3 Opens Large MoE Training

An open training stack could make large MoE experimentation more accessible to researchers and builders.
30-Second TL;DR
What Changed
Olmo-core 3 targets training infrastructure for large MoE models.
Why It Matters
Open training infrastructure can lower the barrier to reproducing and extending large-model research. Its practical value will depend on documentation, hardware efficiency and community adoption.
What To Do Next
Review the Olmo-core 3 repository and run its smallest MoE training example on your available accelerator stack.
Key Points
- •Olmo-core 3 targets training infrastructure for large MoE models.
- •The project is presented as open and scalable.
- •The supplied article does not specify supported hardware or training APIs.
Deep Insight
Background and context from public sources — not the original article. 6 sources cited.
Enhanced Key Takeaways
- •Developed by the Allen Institute for AI (Ai2) and University of Washington researchers, the release includes an open GitHub codebase, an interactive demo, and a companion technical report led by Tianhua Tao.
- •Olmo-core 3 shifts from FSDP weight-resharding to a DDP-based architecture where experts remain stationary in GPU memory while tokens are dynamically routed, cutting communication overhead.
- •Benchmark results on 8 NVIDIA B300 GPUs demonstrated a 2.7× throughput increase, sustaining 52,000 tokens/sec/GPU on a 47B MoE compared to 19,400 tokens/sec/GPU under the previous FSDP stack.
- •Throughput degraded by less than 5% when expanding the model capacity from 8 to 128 experts (total parameters scaling from 4.6B to 47B with active parameters held fixed at ~3.2B).
- •On a 512 NVIDIA B300 GPU cluster, the framework achieved 858 TFLOP/s per GPU while testing a 1.2-trillion-parameter MoE, with alternative backends like DeepEP v2 enabling tests up to 2.38 trillion total parameters.
Technical Deep Dive
- Parallelism Architecture: Replaces traditional Fully Sharded Data Parallelism (FSDP) weight gather/reshard loops with a Distributed Data Parallel (DDP) token-routing layout where experts remain pinned in local memory.
- Scaling Characteristics: Exhibits near-constant scaling overhead, maintaining consistent throughput (under 5% degradation) when scaling from 8 up to 128 experts with top-4 routing.
- Extreme-Scale Hardware Benchmarks: Tested across up to 512 NVIDIA Blackwell B300 GPUs, achieving 858 TFLOP/s per GPU on a 1.2-trillion-parameter MoE configuration (58.36B active parameters).
- High-Throughput Communication Backends: Incorporates support for alternative communication kernels including DeepEP v2, supporting validation scaling tests reaching up to 2.38 trillion parameters.
- Topology-Agnostic Checkpointing: Automatically serializes and reconstructs global FP32 states independent of cluster parallelism layouts, enabling training resumption across clusters with mismatched GPU counts or 3D-parallel schemes.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2024-02Ai2 releases original open OLMo architecture and dense pre-training stack
- 2024-07Ai2 introduces OLMoE-1B-7B sparse mixture-of-experts model
- 2026-10Ai2 launches Olmo-core 3 framework enabling trillion-parameter MoE training
Sources (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Hugging Face Blog ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.