PoEM: Predicting RL Outcomes from Existing Policies

Administrator 0 阅读

AI Digest - ArXiv AI

PoEM: Predicting RL Outcomes from Existing Policies

Foundation models are post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to be run from scratch every time the reward model changes or when we want to combine multiple rewards. We hence ask: given a new reward function, is it possible to predict the RL outcomes without actually running RL on it? We answer this in the affirma


Source: ArXiv AI