Search

Tag: #llm-training33 results

VESPO Stabilizes Off-Policy LLM Training

VESPO Stabilizes Off-Policy LLM Training

VESPO introduces variational sequence-level soft policy optimization to tackle training instability in RL for LLMs caused by policy staleness and async execution. It derives a closed-form reshaping kernel for importance weights without length normalization. Experiments demonstrate stable training up to 64x staleness on math benchmarks.

ArXiv AIResearchFeb 12#research#vespo#v1
Page 4 of 4