发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究分布转移下忠实生成问题,提出令牌级离策略标注(TOPL)方法,通过训练模型区分响应中好坏令牌引导生成,在文档摘要等任务实验中实现强大分布外泛化,能有效转移到机器翻译,且学习信号关键,更新具可解释性。
AI 中文摘要
我们提出了令牌级离策略标注(TOPL),这是一种离策略训练范式,将训练后重新构建为令牌级正确性预测任务。我们的关键直觉是,通过训练模型区分响应中的好令牌和坏令牌,自然地引导模型生成好令牌,同时避免直接训练模型生成离策略令牌带来的陷阱。在文档摘要任务上的实验表明,TOPL在11个数据集上相对于各种序列级和令牌级基线实现了强大的分布外泛化。我们进一步证明TOPL能有效转移到机器翻译,表明其好处能推广到不同的忠实生成任务。通过消融研究,我们确认令牌级学习信号对良好性能至关重要;序列级类似物没有类似好处。最后,我们表明TOPL能诱导可解释的模型更新:通过TOPL学习的LoRA适配器充当线性分类头和引导向量。
英文摘要
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-distribution generalization across 11 datasets against a diverse set of sequence-level and token-level baselines. We further demonstrate that TOPL transfers effectively to machine translation, suggesting that its benefits generalize across different faithful generation tasks. Through ablation studies, we confirm that our token-level learning signal is critical to good performance; sequence-level analogues do not confer similar benefits. Finally, we show that TOPL induces interpretable model updates: the LoRA adapters learned through TOPL function as linear classification heads and steering vectors.