Search

Few direct matches — filled in with the latest updates.

Tag: #bnrm1 results

BNRM Prevents Reward Hacking in RLHF

BNRM Prevents Reward Hacking in RLHF

BNRM introduces Bayesian non-negative reward modeling to combat reward hacking in RLHF. It uses sparse latent factors for disentangled, debiased rewards. Scalable amortized VI enables end-to-end training on LLMs.

ArXiv AIResearchFeb 12#research#bnrm#v1