GRPO Scaling Produces Surprisingly Uneven Results
A researcher applied the same SFT and GRPO recipe to three from-scratch LLMs ranging from 316M to 672M parameters, but observed sharply different outcomes. GRPO barely affected the smallest model, severely degraded the middle model, and caused modest degradation in the largest, with no GSM8K transfer despite curriculum learning.

