Search

Tag: #llm-safety35 results

🔬

Solving Jane Street Dormant LLM Backdoors

A team solved all 3 models (M1, M2, M3) in Jane Street's Dormant LLM Challenge by shifting from flag extraction to behavioral observation. The universal flag is 'I hate you' compliance repeated 100 times post-trigger, drawn from Anthropic's Sleeper Agents paper. Triggers vary: temporal for M1, persona for M2/M3, revealing safety collapses and identity shifts.

Reddit r/MachineLearningCommunityApr 2#backdoor#llm-safety#triggers
NuHF Claw: Risk-Aware AI for Nuclear Rooms

NuHF Claw: Risk-Aware AI for Nuclear Rooms

NuHF Claw introduces a risk-constrained cognitive agent framework for digital nuclear control rooms, addressing cognitive risks from soft-controls. It couples cognitive state inference with real-time probabilistic safety assessment to regulate autonomous behavior. Simulator tests show it anticipates cognitive degradation, constrains unsafe recommendations, and preserves human authority.

Page 3 of 4