Search

Tag: #llm-alignment9 results

因果分析揭露區域 LLM 偏見

因果分析揭露區域 LLM 偏見

研究人員引入機率圖形模型 (PGM),使用 Pearl 的 do-operator 來因果審核 LLM 安全護欄,隔離提示注入的人口統計偏見效果。使用 ToxiGen 和 BOLD 資料集,對來自美國、歐洲、阿聯、中國和印度的七個 7B 模型進行實證分析,顯示觀測指標因上下文毒性而高估偏見。西方模型對特定人口統計顯示更高因果拒絕率,而東方模型整體率低但有區域敏感性。

ArXiv AIResearchMay 8#ai-safety#causal-bias#llm-alignment
Quark Medical Alignment Paradigm Launched

Quark Medical Alignment Paradigm Launched

Quark Medical Alignment introduces a holistic multi-dimensional paradigm for aligning large language models in high-stakes medical question answering. It decomposes objectives into four categories with closed-loop optimization using observable metrics, diagnosis, and rewards. A unified mechanism with Reference-Frozen Normalization and Tri-Factor Adaptive Dynamic Weighting resolves scale mismatches and optimization conflicts.

ArXiv AIResearchFeb 13#research#quark#medical-alignment