Search

Tag: #safety-alignment7 results

ARES Fixes RLHF Dual Safety Flaws

ARES Fixes RLHF Dual Safety Flaws

ARES introduces a framework to discover and fix systemic weaknesses in RLHF where both LLMs and Reward Models fail together. It uses a Safety Mentor to generate adversarial prompts and responses targeting both components. A two-stage repair process fine-tunes the RM first, then optimizes the core model, boosting safety on benchmarks without capability loss.

Narrow Fine-Tuning Erodes VLM Safety

Narrow Fine-Tuning Erodes VLM Safety

Narrow fine-tuning on harmful datasets erodes safety alignment in vision-language models, causing misalignment that generalizes across unrelated tasks and modalities. Experiments on Gemma3-4B reveal misalignment scales with LoRA rank and is worse in multimodal evaluation (70.71%) than text-only (41.19%). Even 10% harmful data triggers substantial degradation, with harms captured in a low-dimensional subspace.