Search

Tag: #arxiv10 results

Surveying Multi-Agent Communication Paradigms

Surveying Multi-Agent Communication Paradigms

This survey frames multi-agent communication via the Five Ws, tracing evolution from MARL's hand-designed protocols to emergent language and LLM-based systems. It highlights trade-offs in interpretability, scalability, and generalization across paradigms. Practical design patterns and open challenges are distilled for hybrid systems.

ArXiv AIResearchFeb 13#research#arxiv#multi-agent
NMIPS: Neuro-Symbolic PDE Solver

NMIPS: Neuro-Symbolic PDE Solver

NMIPS introduces a unified neuro-symbolic framework for solving PDE families with shared structures but varying parameters. It discovers interpretable analytical solutions via multifactorial optimization and affine transfer for efficiency. Experiments show up to 35.7% accuracy gains over baselines.

ArXiv AIResearchFeb 13#research#arxiv#nmips
INTENT: Budget Planning for Tool Agents

INTENT: Budget Planning for Tool Agents

INTENT is an inference-time planner for budget-constrained LLM agents using costly tools. Leverages hierarchical world model for intention-aware cost anticipation. Outperforms baselines on StableToolBench under budgets and price shifts.

ArXiv AIResearchFeb 13#research#arxiv#intent
Exposing Ground Truth Illusion in Annotations

Exposing Ground Truth Illusion in Annotations

Literature review critiques 'ground truth' in ML data annotation as a positivistic fallacy ignoring human subjectivity. Analyzes 346 papers from top venues revealing biases like anchoring and geographic hegemony. Proposes roadmap for pluralistic infrastructures embracing disagreement.

ArXiv AIResearchFeb 13#research#arxiv#none
DrIGM Enables Robust Multi-Agent RL

DrIGM Enables Robust Multi-Agent RL

DrIGM introduces distributionally robust IGM for MARL, ensuring decentralized actions align under uncertainties via robust value factorization. Compatible with VDN/QMIX/QTRAN without reward shaping. Boosts OOD performance in SustainGym and StarCraft.

ArXiv AIResearchFeb 13#research#arxiv#drigm
Benchmarking LLM Agents Under Noise

Benchmarking LLM Agents Under Noise

AgentNoiseBench evaluates tool-using LLM agents' robustness in noisy real-world environments. Categorizes noise into user-noise and tool-noise; injects controllable perturbations into benchmarks. Reveals performance drops across models under perturbations.

ArXiv AIResearchFeb 13#research#arxiv#agentnoisebench