Search

Few direct matches — filled in with the latest updates.

Tag: #tool-calls3 results

GAP: Text Safety Fails for LLM Agent Tools

GAP: Text Safety Fails for LLM Agent Tools

Researchers introduce the GAP benchmark to evaluate divergence between text-level and tool-call safety in LLM agents. Testing six frontier models across six domains reveals text refusals do not prevent harmful tool calls, with 219 persistent cases even under safety prompts. The study urges dedicated tool-call safety measures beyond text evaluations.