Search

Tag: #llm-safety35 results

ARMOR 2025: Military LLM Safety Benchmark

ARMOR 2025: Military LLM Safety Benchmark

ARMOR 2025 introduces a military-aligned benchmark for evaluating LLM safety in defense applications, grounded in Law of War, Rules of Engagement, and Joint Ethics Regulation. It features a 12-category taxonomy based on the OODA framework with 519 doctrinally grounded prompts. Evaluations on 21 commercial LLMs reveal critical safety alignment gaps.

CourtGuard: Zero-Shot LLM Safety Framework

CourtGuard: Zero-Shot LLM Safety Framework

CourtGuard is a retrieval-augmented multi-agent framework that treats LLM safety as an evidentiary debate using external policy documents. It achieves state-of-the-art results on 7 safety benchmarks without fine-tuning, outperforming policy-following baselines. It excels in zero-shot adaptability (90% accuracy on Wikipedia Vandalism) and automated curation of 9 adversarial datasets.

ArXiv AIResearchFeb 28#llm-safety#zero-shot#multi-agent
Page 1 of 4