
KD-MARL Cuts MARL Costs 28x
KD-MARL introduces a two-stage framework for resource-aware knowledge distillation in multi-agent reinforcement learning, transferring coordinated behavior from centralized experts to lightweight decentralized students. It uses distilled advantage signals and structured supervision to preserve coordination without a critic, supporting heterogeneous architectures. Benchmarks show over 90% performance retention with up to 28.6x FLOPs reduction on SMAC and MPE.