Search

Tag: #robustness9 results

FaithSteer-BENCH: LLM Steering Stress-Test Benchmark

FaithSteer-BENCH: LLM Steering Stress-Test Benchmark

FaithSteer-BENCH introduces a deployment-aligned benchmark for stress-testing inference-time steering in LLMs, evaluating controllability, utility preservation, and robustness at fixed operating points. It reveals failure modes like illusory controllability, cognitive tax on unrelated tasks, and brittleness to perturbations in existing methods. Mechanism diagnostics show many approaches induce prompt-conditional alignment rather than stable latent shifts.

CORE: Robust OOD Detection via Orthogonal Scoring

CORE: Robust OOD Detection via Orthogonal Scoring

CORE disentangles confidence and membership signals by decomposing penultimate features into orthogonal subspaces: classifier-aligned confidence and residual. It scores each subspace independently and combines via normalized summation for robust OOD detection. Achieves SOTA performance across five architectures and benchmarks with negligible overhead.

ArXiv AIResearchMar 20#ood-detection#robustness
Robustness and CoT Consistency in RL-Finetuned VLMs

Robustness and CoT Consistency in RL-Finetuned VLMs

Apple researchers investigate the vulnerabilities of RL-finetuned vision-language models, specifically regarding visual grounding and hallucinations. The study reveals that textual perturbations significantly degrade model confidence and reasoning consistency.

Apple Machine LearningOfficialJul 2#robustness
Memvid Hires $800/Day 'AI Bully'

Memvid Hires $800/Day 'AI Bully'

California startup Memvid offers $800 for an 8-hour 'AI bully' role focused on testing leading chatbots' patience and memory. The job entails provoking AI to expose inconsistencies, forgetting, fudging, or hallucinations without meetings or emails. It highlights a unique approach to AI robustness evaluation.

The Guardian TechnologyMediaMar 19#ai-testing#hallucinations#robustness