Search

直接匹配不多,已補上最新動態。

Tag: #tsr1 results

TSR Boosts Multi-Turn RL for LLM Agents

TSR Boosts Multi-Turn RL for LLM Agents

TSR introduces trajectory-search rollouts to enhance multi-turn reinforcement learning for LLM agents. It uses lightweight tree-style search for high-quality trajectories, improving rollout generation and stabilizing training. Achieves up to 15% performance gains on tasks like Sokoban and WebShop.

ArXiv AIResearchFeb 13#research#tsr#llm-agents