Search

Tag: #llm-security18 results

🤖

New Multi-Turn Prompt Injection Patterns Discovered

Security researchers have identified a new class of prompt injection attacks where malicious payloads are delivered across multiple turns of conversation, bypassing standard single-message filters. The Bordair team has released an open-source dataset and a CLI tool to help developers evaluate their LLM endpoints against these sophisticated adversarial patterns.

Reddit r/MachineLearningCommunityJul 2#llm-security#prompt-injection#adversarial-testing
DynaTrust Defends MAS from Sleeper Agents

DynaTrust Defends MAS from Sleeper Agents

DynaTrust introduces dynamic trust graphs to protect LLM-based multi-agent systems from sleeper agents that build trust before attacking. It continuously updates agent trust based on behaviors and expert confidence, then restructures the graph to isolate threats while maintaining task connectivity. Evaluations on AdvBench and HumanEval show 41.7% better success rate than AgentShield, exceeding 86% with low false positives.

Page 1 of 2