
SIE Breaks RL Env Scaling Bottleneck
Shanghai Jiao Tong Univ's SIE uses structured data like KGs for scalable, verifiable RL envs without expert labels. Models learn multi-hop reasoning, generalizing to math/logic puzzles. ICLR 2026 accepted; open-sources code.






