來源ArXiv AI•較早收集於 5h
針對人類與大型語言模型集合的對抗性社會知識論

#trust-and-safety#epistemology#llm-governanceadversarial-social-epistemology-(ase)arxiv
💡學習如何檢測並防止 AI 與人類溝通網路中對信任的策略性操縱。
⚡ 30 秒速覽
有什麼變化
引入對抗性社會知識論(ASE)以分析 LLM 輔助溝通中的信任剝削。
為什麼重要
該框架為開發 AI 整合資訊系統的開發者提供了一個關鍵視角,有助於減輕錯誤資訊並維持系統的可靠性。
下一步行動
在您的應用程式中為 LLM 生成的輸出加入自動化審計軌跡,以追蹤聲明的推論鏈。
誰應關注:Researchers & Academics
關鍵要點
- •引入對抗性社會知識論(ASE)以分析 LLM 輔助溝通中的信任剝削。
- •識別溝通主體如何扭曲或捏造資訊以破壞制度性認證。
- •提出審計推論鏈的機制,以確保公開斷言的完整性。
- •利用推論主義語義學來解釋並驗證 AI 輔助聲明的有效性。
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •The framework draws heavily on Robert Brandom’s inferentialism, treating LLM outputs as 'commitments' within a social game of giving and asking for reasons.
- •ASE specifically addresses the 'epistemic free-riding' problem, where agents use LLMs to generate high-volume, low-effort content that mimics institutional authority.
- •The research introduces a formal 'proof-of-provenance' protocol for inferential chains, requiring LLMs to cryptographically link claims to verifiable source datasets.
- •It identifies 'semantic drift' as a primary vulnerability, where LLMs subtly alter the inferential role of terms during multi-step reasoning to bypass safety filters.
- •The proposed auditing machinery utilizes 'adversarial verification,' where a secondary, specialized LLM acts as a dialectical opponent to stress-test the primary agent's inferential consistency.
🛠️ 技術深入
- Implementation utilizes a Directed Acyclic Graph (DAG) structure to map inferential dependencies across multi-agent interactions.
- Employs 'Inferential Traceability Tokens' (ITTs) to maintain a verifiable log of how specific premises lead to a final assertion.
- Integrates with existing Knowledge Graph (KG) architectures to cross-reference LLM-generated inferences against ground-truth ontologies.
- Utilizes a Bayesian belief-updating mechanism to quantify the 'trust score' of an agent based on its historical adherence to inferential validity.
🔮 前景展望基於引用來源的 AI 分析
Mandatory provenance metadata will become a standard for AI-generated public policy documents by 2027.
The increasing risk of institutional trust erosion necessitates verifiable inferential chains for high-stakes decision-making.
Adversarial verification will replace static safety fine-tuning as the primary method for LLM alignment.
Static fine-tuning fails to account for dynamic, context-dependent subversion, whereas adversarial auditing adapts to the agent's reasoning process.
⏳ 時間線
2025-03
Initial conceptualization of inferentialist semantics applied to LLM-human hybrid systems.
2025-11
Development of the first prototype for auditing inferential chains in closed-loop environments.
2026-06
Publication of the Adversarial Social Epistemology framework on ArXiv.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週電子報
每週一封,可隨時退訂。