📄較早收集於 16h

溢出能量偵測大型語言模型幻覺

溢出能量偵測大型語言模型幻覺
PostLinkedIn
📄閱讀原文: ArXiv AI
#energy-based-models#training-freespilled-energyllamamistralgemmaqwen3

💡Training-free metrics detect LLM hallucinations on LLaMA/Mistral – integrate now for reliable inference.

⚡ 30-Second TL;DR

有什麼變化

將 LLM softmax 重新解釋為互動 EBM 以追蹤能量

為什麼重要

提供零成本、推理時幻覺偵測,可整合至任何 LLM 流程,提升可靠性而無需重新訓練。在 SOTA 模型和任務中泛化,助生產部署。

下一步行動

Compute spilled energy from your LLM logits during decoding to flag hallucinations in real-time.

誰應關注:Researchers & Academics

關鍵要點

  • 將 LLM softmax 重新解釋為互動 EBM 以追蹤能量
  • 從 logits 衍生無訓練的溢出能量和邊際能量指標
  • 能量溢出與幻覺、錯誤、偏見相關
  • 在 9 個基準上穩健,適用 LLaMA/Mistral/Gemma/Qwen3 預訓練/指令微調模型

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 8 個來源。

🔑 增強重點摘要

  • Spilled Energy method was submitted to ICLR 2026 conference on 01 Sept 2025, with revisions on 22 Nov 2025, highlighting its competition in top-tier venues.[4]
  • Unlike Semantic Energy, which requires multiple response samplings and semantic clustering on penultimate logits, Spilled Energy uses only output logits from subsequent generation steps without sampling.[1][4]
  • Spilled Energy improves on prior work like Orgad et al. (2025) by avoiding the need for trained classifiers or activation ablations, enabling zero-shot generalization across tasks and LLMs.[4]
📊 競品分析▸ Show
MethodTraining RequiredLogits UsedSampling NeededKey Benchmarks
Spilled EnergyNoFinal outputNo9 benchmarks (LLaMA, Mistral, Gemma, Qwen3) [4]
Semantic EnergyNoPenultimateYes (multiple responses)Multiple benchmarks, +13% AUROC over Semantic Entropy [1][3]
Semantic EntropyNoPost-softmaxYesHallucination detection [1]
DiffuTruthNoDiffusion reconstructionYes (noise corruption)FEVER (AUROC 0.70+), robust to shifts [5]

🛠️ 技術深入

  • Reinterprets LLM's final softmax layer as an Energy-Based Model (EBM), decomposing sequence probabilities into interacting EBMs during autoregressive decoding.[4]
  • Spilled energy: Measures discrepancy between energy values across two consecutive generation steps, which theoretically should be equal if no spill occurs.[4]
  • Marginalized energy: Computed from energy at a single generation step, providing a lightweight alternative for hallucination detection.[4]
  • Localizes errors to specific tokens without task-specific training, generalizing across pretrained and instruct-tuned LLMs like LLaMA, Mistral, Gemma, Qwen3.[4]

🔮 前景展望AI analysis grounded in cited sources

Spilled Energy enables plug-and-play hallucination detection in production LLMs
Its training-free nature using only output logits allows immediate integration into existing decoding pipelines without retraining or sampling overhead.[4]
Energy-based metrics outperform entropy in overconfident hallucination cases
Complements findings from Semantic Energy, which shows 13%+ AUROC gains over semantic entropy precisely where models are confidently wrong.[1][3]
Broadens EBM applications beyond detection to bias and error localization
Empirical correlation with biases and factual errors suggests utility in guiding interventions during generation.[4]

時間線

2025-09
Spilled Energy submitted to ICLR 2026 conference
2025-11
Paper revisions submitted (22 Nov 2025)
2026-02
Article published on ArXiv AI as 'Spilled Energy Detects LLM Hallucinations'

📎 來源 (8)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. arXiv — 2508
  2. openreview.net — Forum
  3. arXiv — 2508
  4. openreview.net — Forum
  5. arXiv — 2602
  6. arXiv — 2602
  7. arXiv — 2602
  8. ui.adsabs.harvard.edu — Abstract
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。