Search

直接匹配不多,已補上最新動態。

Tag: #temporal-credit1 results

🔬

PPO 多時間尺度優勢解耦修復

研究人員發現 PPO 中動態路由多時間尺度優勢導致策略崩潰,原因為代理目標駭入與時間不確定性悖論。提出解耦 actor 與 critic,讓 actor 只用純長期優勢更新,而 critic 保留多尺度預測。GitHub 提供 PyTorch MRE,展示 LunarLander 中的崩潰、懸停與成功。

Reddit r/MachineLearningCommunityApr 16#temporal-credit#actor-critic