🍎較早收集於 21h

Apple 電路放大提升 LLM 數學推理

Apple 電路放大提升 LLM 數學推理
PostLinkedIn
🍎閱讀原文: Apple Machine Learning
#circuits#math-reasoning#subnetworks#interpretabilityconstructive-circuit-amplificationapplellms

💡Apple's targeted circuit method boosts LLM math reasoning efficiently—key for interpretability research.

⚡ 30-Second TL;DR

有什麼變化

識別 LLM 中負責特定任務的稀疏子網路(電路)

為什麼重要

實現無需完整再訓練的精準 LLM 改進,可能降低運算成本。推進機制解釋性,提升模型控制力。

下一步行動

Read Apple's full paper and test pivotal token identification on your LLM's math circuits.

誰應關注:Researchers & Academics

關鍵要點

  • 識別 LLM 中負責特定任務的稀疏子網路(電路)
  • 微調透過強化既有電路來提升效能
  • 提出建構性電路放大,利用關鍵權杖進行針對性更新

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 10 個來源。

🔑 增強重點摘要

  • CCA operates in three stages: generating reasoning traces to identify deviation points, pinpointing pivotal tokens and model components, and performing sparse targeted updates to amplify constructive signals.
  • On the GSM-Symbolic benchmark, CCA achieves up to +11.4% accuracy improvements across multiple model families while modifying only 1.59% of components like attention heads and MLP neurons.
  • CCA demonstrates minimal impact on unrelated abilities, with preserved performance on MMLU, TriviaQA, and TruthfulQA benchmarks.

🛠️ 技術深入

  • CCA's first stage generates reasoning traces to detect where the model deviates toward incorrect answers, building on prior circuit discovery work for single-pass tasks.
  • Updates target specific attention heads and MLP neurons responsible for correct reasoning, amplifying signals from components generating constructive responses.
  • Efficiency: Modifies as little as 1.59% of model components for significant gains, avoiding broad fine-tuning.
  • Tested on GSM-Symbolic (Mirzadeh et al., 2025), showing models possess latent math-solving capacity but deviate on certain tasks.

🔮 前景展望AI analysis grounded in cited sources

CCA enables 10%+ math accuracy gains with <2% parameter updates
Results on GSM-Symbolic show +11.4% improvements across models by modifying only 1.59% of components, preserving other capabilities.
Targeted circuit methods reduce compute needs for reasoning fine-tuning
Sparse updates focus on pivotal tokens and subnetworks, minimizing full-model retraining while enhancing specific task performance.

時間線

2025-12
arXiv publication of Constructive Circuit Amplification paper by Apple researchers
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Apple Machine Learning

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。