來源量子位•較早收集於 51m
僅需1500美元訓練:高效能1B規模HRM模型引發關注
#rlhf#cost-efficiencyhrm-modelhuggingfaceyoshua bengio
💡了解這款僅需1500美元、獲頂尖AI研究團隊背書的1B模型,如何顛覆獎勵建模領域。
⚡ 30 秒速覽
有什麼變化
模型採用高效的1B參數架構。
為什麼重要
這凸顯了業界正轉向針對獎勵建模等特定任務進行高性價比的小型模型訓練。此趨勢挑戰了「頂尖效能必須依賴巨額運算預算」的傳統觀點。
下一步行動
評估您目前的RLHF流程,確認是否能以更小型的專用獎勵模型取代昂貴的大型替代方案。
誰應關注:Researchers & Academics
關鍵要點
- •模型採用高效的1B參數架構。
- •訓練總成本極低,僅需1500美元。
- •獲得HuggingFace執行長及Yoshua Bengio團隊等業界領袖的高度評價。
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 8 個來源。
🔑 增強重點摘要
- •The model, specifically named HRM-Text-1B, is a 1.15-billion-parameter language model developed by the Singapore-based AI startup Sapient Intelligence, founded in 2024.
- •HRM-Text-1B was trained on 16 GPUs over 1.9 days, utilizing approximately 40 billion structured instruction-response tokens, which is significantly less data (100 to 900 times fewer training tokens) compared to traditional Transformer models that often require trillions of tokens.
- •The model employs a Hierarchical Recurrent Model (HRM) architecture, which decouples computation into slow-evolving strategic and fast-evolving execution layers, enabling effectively unbounded compute depth with a bounded parameter count.
- •HRM-Text-1B demonstrates competitive performance against open models in the 2B-to-7B parameter range on several benchmarks, scoring 56.2 on MATH, 82.2 on DROP, 84.5% on GSM8K, and 81.9% on ARC-C. However, it exhibits weaker performance on MMLU (60.7%), suggesting a factual knowledge gap due to its specialized training data.
- •The model is fully open-sourced and available on platforms like GitHub and Hugging Face, making it accessible for inspection, modification, and deployment.
🛠️ 技術深入
- Architecture: Hierarchical Recurrent Model (HRM) architecture.
- Core Components: Features a dual-timescale recurrent design with two Transformer modules: a high-level (H) slow-updating module for abstract planning and a low-level (L) fast-updating module for detailed computations.
- Computational Depth: The H and L modules iterate over the same input embeddings for H_cycles × (L_cycles + 1) steps, incorporating additive state injection (z_L + z_H), which provides effectively unbounded computational depth at a fixed parameter count.
- Training Data: Pre-trained from scratch on approximately 40 billion structured instruction-response tokens, a stark contrast to the trillions of raw tokens typically used for conventional Transformer models.
- Training Objective: Utilizes a PrefixLM mask during pre-training, where prompt tokens attend bidirectionally to each other, and response tokens attend causally. The mask is controlled by
token_type_ids. - Hardware & Duration: Trained on a cluster of 16 GPUs over a period of 1.9 days.
- Limitations: The model is predominantly English-only due to its training corpus and shows expectedly weak performance on coding tasks as it was not trained on code datasets, though it demonstrates promising adaptation potential with additional code data.
🔮 前景展望基於引用來源的 AI 分析
Foundational LLM training will become more democratized.
The remarkably low training cost of HRM-Text-1B ($1,500) significantly lowers the barrier to entry for foundational model pre-training, making it accessible to a broader range of organizations and research labs beyond highly resourced institutions.
The AI industry will see a shift towards more sample-efficient training paradigms.
HRM-Text-1B's ability to achieve competitive performance with significantly fewer training tokens (40 billion versus trillions for traditional models) challenges the prevailing brute-force scaling approach, suggesting that specialized architectures and structured data can lead to useful general performance more efficiently.
There will be increased research and adoption of hierarchical and recurrent architectures in large language models.
The demonstrated success of the HRM architecture in achieving high performance with low computational cost and parameter count could inspire further development and integration of similar recurrent designs that offer unbounded effective depth in future LLMs.
⏳ 時間線
2024
Sapient Intelligence founded.
2025-01
Sapient Intelligence raised a $22 million seed round.
2025-06
The Hierarchical Reasoning Model (HRM) architecture debuted in a paper, demonstrating competitive performance with a 27 million parameter model.
2026-05-20
The `sapientinc/HRM-Text-1B` model checkpoint was released on Hugging Face.
2026-06-10
Sapient Intelligence publicly released HRM-Text, a 1.15-billion-parameter language model.
📎 來源 (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 量子位 ↗
每週電子報
每週一封,可隨時退訂。
