
PA2D-MORL Boosts Multi-Objective RL
PA2D-MORL proposes an efficient decomposition and policy improvement for multi-objective RL, achieving superior Pareto policy approximations in complex tasks. It uses Pareto ascent directions for scalarization weights and joint policy gradients, with evolutionary optimization of multiple policies and adaptive fine-tuning.
ArXiv AI · 190d ago












