WizardLM Releases Mix-GRM Paper
WizardLM released a new paper on improving Generative Reward Models by synergizing Breadth (B-CoT) and Depth (D-CoT) reasoning instead of just longer CoT. The Mix-GRM framework uses RLVR training to enable the model to autonomously select reasoning structures, achieving high performance with efficient token use. It addresses limitations in subjective vs. objective evaluation tasks.






