Search

Few direct matches — filled in with the latest updates.

Tag: #noise-robustness1 results

Benchmarking LLM Agents Under Noise

Benchmarking LLM Agents Under Noise

AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks. Reveals performance drops across models under perturbations.

ArXiv AIResearchFeb 13#research#arxiv#agentnoisebench
Anduril Raises $5B on AI Defense Infrastructure

Anduril Raises $5B on AI Defense Infrastructure

Anduril Industries reportedly completed a $5 billion Series H financing at a $61 billion valuation, reinforcing its position as a major defense-technology company. The article highlights its Lattice AI platform, flexible Arsenal-1 manufacturing system, and Palmer Luckey’s view that China’s engineering and manufacturing efficiency could reshape future military competition.