Search

Few direct matches — filled in with the latest updates.

Tag: #adversarial-debate1 results

CourtGuard: Zero-Shot LLM Safety Framework

CourtGuard: Zero-Shot LLM Safety Framework

CourtGuard is a retrieval-augmented multi-agent framework that treats LLM safety as an evidentiary debate using external policy documents. It achieves state-of-the-art results on 7 safety benchmarks without fine-tuning, outperforming policy-following baselines. It excels in zero-shot adaptability (90% accuracy on Wikipedia Vandalism) and automated curation of 9 adversarial datasets.

ArXiv AIResearchFeb 28#llm-safety#zero-shot#multi-agent