SourceStalecollected in 7h

Olmo-core 3 Opens Large MoE Training

Read original on Hugging Face Blog
#mixture-of-experts#model-training#open-infrastructure

An open training stack could make large MoE experimentation more accessible to researchers and builders.

30-Second TL;DR

What Changed

Olmo-core 3 targets training infrastructure for large MoE models.

Why It Matters

Open training infrastructure can lower the barrier to reproducing and extending large-model research. Its practical value will depend on documentation, hardware efficiency and community adoption.

What To Do Next

Review the Olmo-core 3 repository and run its smallest MoE training example on your available accelerator stack.

Who should care:Researchers & Academics

Key Points

  • •Olmo-core 3 targets training infrastructure for large MoE models.
  • •The project is presented as open and scalable.
  • •The supplied article does not specify supported hardware or training APIs.
Key numbers52,000 tokens/s19,400 tokens/s5%

Deep Insight

Background and context from public sources — not the original article. 6 sources cited.

Enhanced Key Takeaways

  • •Developed by the Allen Institute for AI (Ai2) and University of Washington researchers, the release includes an open GitHub codebase, an interactive demo, and a companion technical report led by Tianhua Tao.
  • •Olmo-core 3 shifts from FSDP weight-resharding to a DDP-based architecture where experts remain stationary in GPU memory while tokens are dynamically routed, cutting communication overhead.
  • •Benchmark results on 8 NVIDIA B300 GPUs demonstrated a 2.7× throughput increase, sustaining 52,000 tokens/sec/GPU on a 47B MoE compared to 19,400 tokens/sec/GPU under the previous FSDP stack.
  • •Throughput degraded by less than 5% when expanding the model capacity from 8 to 128 experts (total parameters scaling from 4.6B to 47B with active parameters held fixed at ~3.2B).
  • •On a 512 NVIDIA B300 GPU cluster, the framework achieved 858 TFLOP/s per GPU while testing a 1.2-trillion-parameter MoE, with alternative backends like DeepEP v2 enabling tests up to 2.38 trillion total parameters.

Technical Deep Dive

  • Parallelism Architecture: Replaces traditional Fully Sharded Data Parallelism (FSDP) weight gather/reshard loops with a Distributed Data Parallel (DDP) token-routing layout where experts remain pinned in local memory.
  • Scaling Characteristics: Exhibits near-constant scaling overhead, maintaining consistent throughput (under 5% degradation) when scaling from 8 up to 128 experts with top-4 routing.
  • Extreme-Scale Hardware Benchmarks: Tested across up to 512 NVIDIA Blackwell B300 GPUs, achieving 858 TFLOP/s per GPU on a 1.2-trillion-parameter MoE configuration (58.36B active parameters).
  • High-Throughput Communication Backends: Incorporates support for alternative communication kernels including DeepEP v2, supporting validation scaling tests reaching up to 2.38 trillion parameters.
  • Topology-Agnostic Checkpointing: Automatically serializes and reconstructs global FP32 states independent of cluster parallelism layouts, enabling training resumption across clusters with mismatched GPU counts or 3D-parallel schemes.

Future ImplicationsAI analysis grounded in cited sources

Open-weight models will scale beyond 1 trillion parameters without proprietary training frameworks.
Olmo-core 3 provides an open-source, highly efficient MoE training pipeline tested up to 2.38 trillion parameters on Blackwell hardware, lowering the engineering barrier for frontier open research.
Token-routing architectures will largely displace FSDP for extreme sparse MoE workloads.
The 2.7× single-node throughput leap and sub-5% scaling penalty prove that pinning experts locally and routing activations avoids the severe network bottlenecks of weight-resharding.

Timeline

2024-02
Ai2 releases original open OLMo architecture and dense pre-training stack
2024-07
Ai2 introduces OLMoE-1B-7B sparse mixture-of-experts model
2026-10
Ai2 launches Olmo-core 3 framework enabling trillion-parameter MoE training

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Hugging Face Blog ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.