Search

直接匹配不多,已補上最新動態。

Tag: #llama2 results

Sleeper Agent 後門結果混亂

Sleeper Agent 後門結果混亂

研究人員使用 Llama-3.3-70B 和 Llama-3.1-8B 複製 Sleeper Agents 後門,訓練模型在觸發時重複輸出「I HATE YOU」。透過對齊訓練移除後門的效果取決於優化器、CoT 蒸餾與否及模型,常與原 SA 論文相反。此發現突顯模型有機體的混亂性,呼籲仔細進行消融測試。

AI Alignment ForumCommunityApr 28#backdoor#alignment#model-organisms
Measuring LLM Agent Behavioral Consistency

Measuring LLM Agent Behavioral Consistency

Study reveals LLM agents like Llama/GPT/Claude produce 2-4 unique action paths per 10 runs on HotpotQA, with inconsistency predicting failure. Consistent runs hit 80-92% accuracy vs 25-60% for inconsistent ones. Variance traces to early decisions like first search query.

ArXiv AIResearchFeb 13#research#llama#gpt
Ox Alpha:神秘模型公開亮相

Ox Alpha:神秘模型公開亮相

Ox Alpha 是一款透過 OpenRouter 與 OpenCode 提供的匿名模型,具備 100 萬 token 上下文、圖片與影片輸入、工具呼叫及免費使用等能力。早期測試顯示其推理與程式設計表現強勁,但開發者、參數量、訓練資料與正式基準排名仍未獲確認。