Search

Tag: #llms11 results

LLM幻覺的幾何分類法

LLM幻覺的幾何分類法

研究人員提出LLM幻覺的幾何分類法,將其分為三類:不忠實(忽略上下文)、捏造(發明語義外內容)和事實錯誤(正確框架內錯誤)。基準測試幻覺具領域本地檢測力(AUROC 0.76-0.99),跨領域僅隨機水平;人類捏造則可用單一全球方向達0.96 AUROC。事實錯誤因嵌入僅編碼共現而無法檢測。

ArXiv AIResearchFeb 17#research#llms#hallucinations
SemaPop: Semantic Population Synthesis

SemaPop: Semantic Population Synthesis

SemaPop uses LLMs for semantic-conditioned population synthesis, deriving personas from surveys. Integrates with WGAN-GP for statistical alignment and behavioral realism. Achieves better marginal/joint distribution matches with diversity.

ArXiv AIResearchFeb 13#research#semapop#llms
ERM Fixes Causal Rung Collapse in LLMs

ERM Fixes Causal Rung Collapse in LLMs

New research identifies 'rung collapse' in LLMs, where models confuse associations with causal interventions, leading to flawed reasoning under distributional shifts. It proposes Epistemic Regret Minimization (ERM), a belief revision method that penalizes causal errors independently of task success. Experiments across six frontier LLMs show ERM recovers 53-59% of entrenched errors.

ArXiv AIResearchFeb 13#research#llms#v1
🤖

Metacognition Reduces LLM Slop, Aids Alignment

LLMs lack human-like metacognitive skills, causing errors, sycophancy, and 'slop' outputs. Enhancing metacognition could catch mistakes, stabilize alignment via reflective endorsement, and improve research utility. Benefits for alignment may outweigh capability risks, with work already underway.

AI Alignment ForumCommunityFeb 12#research#llms#ai-alignment
🤖

Metacognition Reduces LLM Slop

LLMs lack human-like metacognitive skills for error-catching and cognition management. Enhancing these could cut slop, sycophancy, and aid alignment research. Benefits for alignment may outweigh capability risks.

AI Alignment ForumCommunityFeb 12#research#llms#metacognition
FAC Synthesizes Diverse LLM Data

FAC Synthesizes Diverse LLM Data

Feature Activation Coverage (FAC) measures diversity in LLM feature space using sparse autoencoders. FAC Synthesis generates samples targeting missing features from seed data. Boosts diversity and performance on instruction, toxicity, reward, and steering tasks.

ArXiv AIResearchFeb 12#research#fac-synthesis#v1
Page 1 of 2