Ask a Question

Prefer a chat interface with context about you and your work?

MetaRM: Shifted Distributions Alignment via Meta-Learning

MetaRM: Shifted Distributions Alignment via Meta-Learning

The success of Reinforcement Learning from Human Feedback (RLHF) in language model alignment is critically dependent on the capability of the reward model (RM). However, as the training process progresses, the output distribution of the policy model shifts, leading to the RM's reduced ability to distinguish between responses. This issue …