🦙Stalecollected in 25m

Covenant-72B: Largest Decentralized GPU Model

Covenant-72B: Largest Decentralized GPU Model
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA

💡Largest LLM via decentralized GPUs—unlocks permissionless training tech.

⚡ 30-Second TL;DR

What Changed

Largest 72B model trained on decentralized permissionless GPUs

Why It Matters

Pioneers scalable decentralized training, lowering barriers for open AI development and challenging centralized cloud dominance.

What To Do Next

Download Covenant-72B from Hugging Face and test SparseLoco for your distributed training setups.

Who should care:Researchers & Academics

Key Points

  • Largest 72B model trained on decentralized permissionless GPUs
  • Introduces SparseLoco method built on DiLoCo for lower comms overhead
  • Uses local AdamW optimizer and reduced sync frequency
  • Applies aggressive top-K sparsification to fix bandwidth bottlenecks

🧠 Deep Insight

Background and context from public sources — not the original article. 7 sources cited.

🔑 Enhanced Key Takeaways

  • Covenant-72B was trained on Bittensor subnet 3 (SN03 Templar) by over 70 independent GPU contributors using regular home and colocation hardware connected via commodity internet.[1][2]
  • The model underwent approximately 1.1 trillion tokens of pre-training, with a fine-tuned variant Covenant-72B-Chat developed via supervised fine-tuning (SFT).[3][4]
  • On the MMLU benchmark, Covenant-72B scored 67.1, surpassing LLaMA-2-70B (65.6) and LLM360 K2 (65.5).[1][5]
📊 Competitor Analysis▸ Show
FeatureCovenant-72BLLaMA-2-70BLLM360 K2
MMLU Score67.165.665.5
Training SetupDecentralized, permissionless on Bittensor SN03Centralized data centersCentralized

🛠️ Technical Deep Dive

  • 72,747,327,488 total parameters; 80 layers; model dimension (width) of 8192.[4]
  • Architecture: dense decoder-only Transformer with 64 query heads, 8 KV heads, RoPE with θ=500,000, and Gemma 3 tokenizer (vocab size 262,208).[4]
  • SparseLoco achieved over 146× compression of pseudo-gradients compared to dense communication, enabling training despite 7.2× larger scale than prior decentralized model INTELLECT-1.[1][4]

🔮 Future ImplicationsAI analysis grounded in cited sources

Decentralized GPU networks will train models rivaling centralized ones at 72B+ scale by 2027
Covenant-72B's competitive benchmarks on 1.1T tokens over permissionless internet prove viability, paving the way for broader adoption in Bittensor ecosystems like SN64 and SN27.[1][2]
Permissionless training reduces LLM development costs by 50%+ for SMEs via BYO-GPU marketplaces
The use of commodity GPUs and compressed gradients on Bittensor demonstrates scalable aggregation of small providers without data centers.[3]

Timeline

2026-03
Templar completes Covenant-72B pre-training on Bittensor SN03, largest decentralized 72B LLM run on 1.1T tokens.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.