Search

Tag: #llm-agents89 results

Benchmarking LLM Agents Under Noise

Benchmarking LLM Agents Under Noise

AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks. Reveals performance drops across models under perturbations.

ArXiv AIResearchFeb 13#research#arxiv#agentnoisebench
ARC Learns Dynamic Agent Configurations

ARC Learns Dynamic Agent Configurations

ARC introduces a reinforcement learning policy to dynamically configure LLM-based agent systems per query, selecting optimal workflows, tools, and prompts. It outperforms fixed templates on reasoning and tool-augmented QA benchmarks. The approach boosts accuracy by up to 25% while cutting token and runtime costs.

ArXiv AIResearchFeb 13#research#arc#llm-agents
AIR Boosts LLM Agent Safety

AIR Boosts LLM Agent Safety

AIR is the first incident response framework for LLM agents, focusing on detecting, containing, recovering from, and eradicating incidents post-occurrence. It integrates a domain-specific language into the agent's execution loop for autonomous management. Evaluations across agent types show over 90% success rates in all phases.

ArXiv AIResearchFeb 13#research#air#llm-agents
LLM Agents Auto-Optimize RecSys Models

LLM Agents Auto-Optimize RecSys Models

A self-evolving system uses Google's Gemini LLMs to autonomously generate, train, and deploy recommendation model improvements. It features an Offline Agent for hypothesis generation and an Online Agent for production validation. Deployed successfully at YouTube, surpassing manual workflows.

ArXiv AIResearchFeb 12#research#youtube#v1
Large-Scale AI Social Simulation Launched

Large-Scale AI Social Simulation Launched

AIvilization v0 deploys a resource-constrained artificial society with unified LLM agents. Features hierarchical planning, adaptive profiles, and human steering for long-horizon autonomy. Reproduces real market stylized facts like wealth stratification.

ArXiv AIResearchFeb 12#research#aivilization#v0
AgentTrace Enables AI Agent Observability

AgentTrace Enables AI Agent Observability

AgentTrace instruments LLM agents for structured logging across operational, cognitive, and contextual traces. Provides runtime transparency for security and monitoring in high-stakes settings. Minimal overhead supports accountability and risk analysis.

ArXiv AIResearchFeb 12#launch#agenttrace#v1
Apple Maps UX for LLM Computer Agents

Apple Maps UX for LLM Computer Agents

Apple's Machine Learning team conducted a two-phase study to explore user experience design for LLM-based computer use agents. Phase 1 reviewed existing systems and interviewed eight UX/AI practitioners to create a taxonomy covering user prompts, explainability, user control, and more. The work aims to understand optimal user interactions with these UI-interacting agents.

Apple Machine LearningOfficialFeb 12#research#apple#na
Page 9 of 9