将近最优SFT-RL标注预算分配从小型大语言模型扩展至大型大语言模型
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
浏览论文内容
中文总结 AI 辅助
该研究提出近最优区域框架,发现SFT-RL的近最优区域可从小型代理模型迁移至大型LLM,据此提出无需大规模搜索的实用标注预算分配策略,且结果在多任务、多模型及不同RL方法中一致。
中文摘要 AI 辅助
在大语言模型(LLM)的后训练阶段,如何将固定的标注预算分配给监督微调(SFT)和强化学习(RL)仍是一个未解决的问题。现有研究仅描述了宽泛的趋势(例如,在低数据 regime 下SFT占主导),缺乏原则性的分配框架,也未探究最优比例是否能跨模型规模迁移。我们将该问题以近最优性的视角进行建模:不寻求单一最优的SFT-RL比例,而是刻画近最优区域,即处于峰值性能指定容差范围内的分配集合。实证结果表明,即使在较小的容差(2%-10%)下,该区域也较宽,且随模型规模增大而变宽,并能从小型代理模型可靠地迁移至大型目标模型。这产生了一种实用策略:小型代理模型的实验足以识别可迁移的近最优区域,无需进行 exhaustive 大规模搜索。我们的结果在不同任务、模型家族以及基于偏好的离线RL和基于奖励监督的在线RL方法中均保持一致。我们进一步分析了SFT和RL数据之间标注成本的不对称性如何改变近最优区域。
英文摘要
How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work characterizes only broad trends (e.g., SFT dominates in low-data regimes), lacks a principled allocation framework, and does not examine whether the optimal ratio transfers across model sizes. We frame this problem in terms of near-optimality: rather than seeking a single optimal SFT-RL ratio, we characterize the near-optimal region, the set of allocations within a specified tolerance of peak performance. Empirically, this region is wide even for small tolerances (2-10%), widens with model scale, and transfers reliably from small proxy models to large target models. This yields a practical strategy: small proxy-model experiments suffice to identify a transferable near-optimal region, eliminating the need for exhaustive large-scale search. Our results hold consistently across tasks, model families, and both preference-based off-policy and reward-supervision on-policy RL methods. We further analyze how the asymmetry in annotation costs between SFT and RL data shifts the near-optimal region.
发表机构
- National University of Singapore(新加坡国立大学)
- Agency for Science, Technology, and Research (A*STAR)(新加坡科技研究局)
- Singapore-MIT Alliance for Research and Technology Centre(新加坡-麻省理工研究与技术联盟中心)
机构由 AI 辅助整理,请以论文原文为准。