Covenant-72B: Largest Decentralized GPU Model

💡Largest LLM via decentralized GPUs—unlocks permissionless training tech.
⚡ 30-Second TL;DR
What Changed
Largest 72B model trained on decentralized permissionless GPUs
Why It Matters
Pioneers scalable decentralized training, lowering barriers for open AI development and challenging centralized cloud dominance.
What To Do Next
Download Covenant-72B from Hugging Face and test SparseLoco for your distributed training setups.
Key Points
- •Largest 72B model trained on decentralized permissionless GPUs
- •Introduces SparseLoco method built on DiLoCo for lower comms overhead
- •Uses local AdamW optimizer and reduced sync frequency
- •Applies aggressive top-K sparsification to fix bandwidth bottlenecks
🧠 Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
🔑 Enhanced Key Takeaways
- •Covenant-72B was trained on Bittensor subnet 3 (SN03 Templar) by over 70 independent GPU contributors using regular home and colocation hardware connected via commodity internet.[1][2]
- •The model underwent approximately 1.1 trillion tokens of pre-training, with a fine-tuned variant Covenant-72B-Chat developed via supervised fine-tuning (SFT).[3][4]
- •On the MMLU benchmark, Covenant-72B scored 67.1, surpassing LLaMA-2-70B (65.6) and LLM360 K2 (65.5).[1][5]
📊 Competitor Analysis▸ Show
| Feature | Covenant-72B | LLaMA-2-70B | LLM360 K2 |
|---|---|---|---|
| MMLU Score | 67.1 | 65.6 | 65.5 |
| Training Setup | Decentralized, permissionless on Bittensor SN03 | Centralized data centers | Centralized |
🛠️ Technical Deep Dive
- •72,747,327,488 total parameters; 80 layers; model dimension (width) of 8192.[4]
- •Architecture: dense decoder-only Transformer with 64 query heads, 8 KV heads, RoPE with θ=500,000, and Gemma 3 tokenizer (vocab size 262,208).[4]
- •SparseLoco achieved over 146× compression of pseudo-gradients compared to dense communication, enabling training despite 7.2× larger scale than prior decentralized model INTELLECT-1.[1][4]
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- tao.media — Templar Makes History with 72b Decentralized AI Training Run
- youtube.com — Watch
- emergentmind.com — 2603
- arXiv — 2603
- en.rattibha.com — 2031388295972929720
- simplytao.ai — Covenant 72b the Largest Decentralized LLM Training Run
- simplytao.ai — Covenant 72b Changes Everything Weekly Bittensor Update
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
