发表机构
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出轨迹式迁移的局部理论,探究替代更新改进决策的条件,推导相关梯度、校准等理论结果,通过网格世界和LLM后训练实验验证结论。
AI 中文摘要
广泛的模型面临着这样的不匹配:它们通过轨迹损失进行更新,但由下游任务奖励进行评估。这里,轨迹是一个训练实例,它会诱导出一个替代损失,而该损失的减少可能无法追踪模型的决策效用更新。从理论上讲,我们探究了轨迹训练的一步更新何时能同时降低总体替代损失和决策风险,以及迁移如何在重复更新中累积。为了将此形式化,我们首先固定一个检查点和一个受限更新空间,并将轨迹诱导的总体替代风险和决策风险的减少分别定义为其可学习性和决策效用。在此基础上,我们的理论得出四个主要结果:第一,一步迁移界将它们的差异分离为非负校准后的一阶梯度失配和二阶曲率;路径式扩展会在重复更新中累积相同的项。第二,当可访问的替代梯度非零时,对每个可访问方向的通用一阶迁移恰好成立,当且仅当可访问的替代梯度和决策梯度正共线。第三,校准差距界定了基于可学习性的轨迹选择的决策遗憾,而候选差异细化通过仅保留影响成对排名的方向来收紧这一保证。最后,我们在嵌套更新空间中建立了近似-校准权衡。受控网格世界和大语言模型(LLM)后训练实验得出的结果与我们的预测一致。
英文摘要
A broad range of models face the mismatch where they are updated through trajectory losses but are evaluated by downstream task reward. Here, a trajectory is a training instance that induces a surrogate loss whose reduction might not track the model's decision utility update. Theoretically, we ask when one step of trajectory training reduces both population surrogate loss and decision risk, and how transfer accumulates along repeated updates. To formalize this, we first fix a checkpoint and a restricted update space, and define the reductions in population surrogate risk and decision risk induced by a trajectory as its learnability and decision utility, respectively. On this basis, our theory yields four main results. First, a one-step transfer bound separates their discrepancy into first-order gradient misalignment after nonnegative calibration and second-order curvature; and a pathwise extension accumulates the same terms over repeated updates. Second, when the accessible surrogate gradient is nonzero, universal first-order transfer over every accessible direction holds exactly when the accessible surrogate and decision gradients are positively collinear. Third, the calibration gap bounds the decision regret of learnability-based trajectory selection, while a candidate-difference refinement tightens this guarantee by retaining only directions that affect pairwise rankings. Finally, we establish an approximation--calibration trade-off across nested update spaces. Controlled gridworld and LLM post-training experiments yield results consistent with our predictions.
Comments22 pages in total, including references and appendices; 3 composite figures (9 panels) and 4 tables. Code is available at https://github.com/Ethan-Shen-Individual-Lab/surrogate-to-decision-transfer