Search

Tag: #llm-safety35 results

Zhiyuan Launches FlagSafe LLM Safety Platform

Zhiyuan Launches FlagSafe LLM Safety Platform

Beijing Zhiyuan AI Research Institute, with Peking University and others, released the FlagSafe large model safety platform. It aggregates frontier safety research projects in red team exercises, blue team defense, and white-box interpretability. The platform covers risk discovery, defense governance, and mechanism explanation.

NLAs Explain LLM Activations in Natural Language

NLAs Explain LLM Activations in Natural Language

Anthropic introduces Natural Language Autoencoders (NLAs), using two LLM modules to translate activations into readable text and reconstruct them via reinforcement learning. NLAs uncover hidden thoughts in Claude Opus 4.6, like unverbalized evaluation awareness, aiding safety audits. Code and trained NLAs for open models are released.

AI Alignment ForumCommunityMay 7#interpretability#model-auditing#llm-safety
LOCA Explains Specific LLM Jailbreaks

LOCA Explains Specific LLM Jailbreaks

Researchers introduce LOCA, a method providing local, causal explanations for jailbreak success in safety-trained LLMs by identifying minimal interpretable changes in intermediate representations to induce refusal. Evaluated on Gemma and Llama models using a large jailbreak benchmark, LOCA achieves refusal with an average of six changes, outperforming prior methods that often fail even after 20 changes. This advances mechanistic understanding of jailbreaks beyond global explanations.

Page 2 of 4