Search

Tag: #llm-security18 results

🔬

Best LLMs and Datasets for AI Red-Teaming

A developer seeks recommendations for high-performing LLMs to generate adversarial prompts and requests public 'golden' datasets for benchmarking AI agent security. The inquiry focuses on automating red-teaming for vulnerabilities like prompt injection, jailbreaks, and tool misuse.

Reddit r/MachineLearningCommunityJul 5#security#red-teaming#llm-security
315 Exposes Easy AI Poisoning

315 Exposes Easy AI Poisoning

China's 315 gala exposed GEO scams where 10+ fake articles generated by software fool LLMs into recommending fictional products like AstroTekk Apollo-9 smartband. Costs start at 299 RMB; AI treats repeated fakes as consensus. Highlights risks in product recommendations and competitor smears.

🔬

Why Roles Matter in Prompt Injection

This research-oriented post offers a mechanistic explanation of prompt injection and argues that understanding model roles is essential to studying the vulnerability. It encourages practitioners to examine how models distinguish system, developer, user, and other instruction sources.

Reddit r/MachineLearningCommunityAug 9#prompt-injection#llm-security
Page 2 of 2