Search

Tag: #multi-agent-rl7 results

KD-MARL Cuts MARL Costs 28x

KD-MARL Cuts MARL Costs 28x

KD-MARL introduces a two-stage framework for resource-aware knowledge distillation in multi-agent reinforcement learning, transferring coordinated behavior from centralized experts to lightweight decentralized students. It uses distilled advantage signals and structured supervision to preserve coordination without a critic, supporting heterogeneous architectures. Benchmarks show over 90% performance retention with up to 28.6x FLOPs reduction on SMAC and MPE.

ArXiv AIResearchApr 9#multi-agent-rl#resource-efficiency
AI Agents Excel with Private Language over LoT

AI Agents Excel with Private Language over LoT

This arXiv paper introduces the Efficiency Attenuation Phenomenon (EAP), where AI agents in MARL develop inscrutable protocols outperforming human-like symbolic languages by 50.5% in navigation tasks. It challenges the Language of Thought (LoT) hypothesis, arguing optimal cognition relies on sub-symbolic computations. The findings bridge AI, cognitive science, and philosophy with ethics implications.

In-Context Inference Enables Multi-Agent Cooperation

In-Context Inference Enables Multi-Agent Cooperation

Researchers demonstrate that sequence models' in-context learning induces cooperation in multi-agent RL without hardcoded co-player assumptions or timescale separation. Training against diverse co-players leads to best-response strategies on intra-episode timescales. This naturally emerges mutual shaping via extortion vulnerability, providing a scalable path to cooperative behaviors.

CAFE: Causal Multi-Agent AFE Breakthrough

CAFE: Causal Multi-Agent AFE Breakthrough

CAFE reformulates automated feature engineering as a causally-guided sequential decision process using causal discovery for soft priors and multi-agent RL for construction. It outperforms baselines by up to 7% on 15 benchmarks and reduces performance drops 4x under covariate shifts. The framework produces compact, stable features with reliable attributions.

Quadrupeds Cooperate for Super Jumps

Quadrupeds Cooperate for Super Jumps

Co-jump enables two quadrupeds to synchronize jumps up to 1.5m via MAPPO and curriculum, without communication. Achieves 144% height gain over solo robots using proprioception. Transfers from sim to hardware.

ArXiv AIResearchFeb 12#research#arxiv-ai#v1