Search

Few direct matches — filled in with the latest updates.

Tag: #policy-gradient1 results

ARLArena: Stable Agentic RL Framework

ARLArena: Stable Agentic RL Framework

ARLArena offers a stable training recipe and analysis framework for agentic reinforcement learning (ARL) to combat training collapse. It decomposes policy gradients into four core dimensions and introduces SAMPO, a method mitigating key instability sources. SAMPO ensures consistent stability and superior performance across diverse agentic tasks.