Search

Few direct matches — filled in with the latest updates.

Tag: #multi-objective1 results

Blockwise Advantages for Multi-Objective RL

Blockwise Advantages for Multi-Objective RL

Introduces Blockwise Advantage Estimation for GRPO in structured generations, assigning per-objective advantages to avoid interference. Uses Outcome-Conditioned Baseline to estimate advantages without nested rollouts. Competitive on math tasks with uncertainty estimation.

ArXiv AIResearchFeb 12#research#grpo#v1