Search

Tag: #research297 results

BHI Framework Audits LLM Benchmarks

BHI Framework Audits LLM Benchmarks

Introduces Benchmark Health Index (BHI), a data-driven framework to audit LLM benchmarks amid reliability issues like score inflation. Evaluates along three axes: Capability Discrimination, Anti-Saturation, and Impact. Analyzes 106 benchmarks from 91 models in 2025.

ArXiv AIResearchFeb 13#research#bhi#v1
Benchmarking LLM Agents Under Noise

Benchmarking LLM Agents Under Noise

AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks. Reveals performance drops across models under perturbations.

ArXiv AIResearchFeb 13#research#arxiv#agentnoisebench
AT-RL Reinforces MLLM Anchors for Reasoning

AT-RL Reinforces MLLM Anchors for Reasoning

AT-RL selectively reinforces high-connectivity cross-modal anchor tokens (15% of total) in MLLM RLVR via attention graph clustering. 32B model hits 80.2% on MathVista, beating 72B baseline with 1.2% overhead. Low-connectivity training degrades performance.

ArXiv AIResearchFeb 13#research#mllm#at-rl
ARC Learns Dynamic Agent Configurations

ARC Learns Dynamic Agent Configurations

ARC introduces a reinforcement learning policy to dynamically configure LLM-based agent systems per query, selecting optimal workflows, tools, and prompts. It outperforms fixed templates on reasoning and tool-augmented QA benchmarks. The approach boosts accuracy by up to 25% while cutting token and runtime costs.

ArXiv AIResearchFeb 13#research#arc#llm-agents
AIR Boosts LLM Agent Safety

AIR Boosts LLM Agent Safety

AIR is the first incident response framework for LLM agents, focusing on detecting, containing, recovering from, and eradicating incidents post-occurrence. It integrates a domain-specific language into the agent's execution loop for autonomous management. Evaluations across agent types show over 90% success rates in all phases.

ArXiv AIResearchFeb 13#research#air#llm-agents
AgentLeak: Multi-Agent Privacy Leak Benchmark

AgentLeak: Multi-Agent Privacy Leak Benchmark

AgentLeak introduces the first full-stack benchmark for privacy leakage in multi-agent LLM systems, covering internal channels like inter-agent messages. It spans 1,000 scenarios across healthcare, finance, legal, and corporate domains. Tests on top models show internal channels cause 68.9% total leakage, missed by output audits.

ArXiv AIResearchFeb 13#research#agentleak#multi-agent
Page 12 of 30