Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

Administrator 0 阅读

AI Digest - ArXiv AI

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work characterizes only broad trends (e.g., SFT dominates in low-data regimes), lacks a principled allocation framework, and does not examine whether the optimal ratio transfers across model sizes. We frame this problem in terms of near-optimality: rather than seeking a single optimal SFT-RL ratio, we characterize the near-optimal


Source: ArXiv AI