Search

Few direct matches — filled in with the latest updates.

Tag: #vespo1 results

VESPO Stabilizes Off-Policy LLM Training

VESPO Stabilizes Off-Policy LLM Training

VESPO introduces variational sequence-level soft policy optimization to tackle training instability in RL for LLMs caused by policy staleness and async execution. It derives a closed-form reshaping kernel for importance weights without length normalization. Experiments demonstrate stable training up to 64x staleness on math benchmarks.

ArXiv AIResearchFeb 12#research#vespo#v1