Coverage Improvement and Fast Convergence of On-policy Preference Learning
策略学习中的覆盖改进与快速收敛
机构 * UC Berkeley(伯克利大学) ; KRAFTON(KRAFTON公司) ; University of Arizona(亚利桑那大学)
AI总结 本文提出覆盖改进原理,通过分析在线策略学习中的覆盖变化,证明其在批量大小足够时能实现快速收敛,并设计混合采样器和奖励蒸馏方案,提升语言模型对齐性能。
Comments 46 pages, 2 figures, 2 tables