发表机构
Waseda University(早稻田大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对持续预训练后知识检索不准的问题,提出风格去偏 DPO,通过反转事实正确但风格不同的偏好对并加权,高效提升知识引出与更新准确率。
AI 中文摘要
通过数据增强(如释义)进行的持续预训练(CPT)可以将小型源语料库的知识存储在大语言模型(LLM)中。然而,存储的知识并不总能被正确检索。我们研究的是知识的引出而非存储方面:我们使用偏好优化,它从成对的偏好(被选中的)和不偏好(被拒绝的)响应中学习,以使模型更准确地引出其存储的知识。一种提出的方法将模型自身的错误响应作为被拒绝的响应,将正确答案作为被选中的响应,从而抑制错误。然而,当目标知识部分已知时,这些被拒绝的响应中的大多数在事实上是正确的。使用直接偏好优化(DPO)会降低包含正确知识且仅在风格(如长度和措辞)上与被选中的答案不同的被拒绝响应的概率。我们提出了风格去偏 DPO(SD-DPO),它对每一对被拒绝的响应在事实上是否正确进行评分,反转此类对的偏好,并对其加权,使得由风格差异引起的学习信号在整体上相互抵消。我们首先测试了在 EntiGraph(一种代表性的存储侧方法,它对从语料库合成的文本运行 CPT)之上,我们的方法是否能高效地增加准确性。在 QuALITY(EntiGraph 评估所用的阅读理解问答基准)上,SD-DPO 超过了我们使用相同基础模型在 EntiGraph 的合成数据上进行 CPT 并采用相同程序评估的基线。实现相同增益所需的训练 token 比额外 CPT 所需的少几十倍。对于知识更新(这项工作的主要目标),我们使用 AToKE,一个针对随时间变化事实的知识编辑基准。在那里,SD-DPO 达到了 0.982 的总体准确率,并根据查询的时间段回答新事实或旧事实。
英文摘要
Continued pretraining (CPT) with data augmentation such as paraphrasing can store inside a large language model (LLM) the knowledge of a small source corpus. The stored knowledge, however, is not always retrieved correctly. We study the eliciting side rather than the storing side: we use preference optimization, which learns from pairs of a preferred (chosen) and a dispreferred (rejected) response, so that the model elicits its stored knowledge more accurately. One proposed approach takes the model's own erroneous response as rejected and the gold answer as chosen, so as to suppress the error. When the target knowledge is partially known, however, most of these rejected responses are factually correct. Using direct preference optimization (DPO) then pushes down rejected responses that contain correct knowledge and differ from the chosen answer only in style, such as length and wording. We propose style-debiased DPO (SD-DPO), which scores whether the rejected response of each pair is factually correct, inverts the preference of such pairs, and weights them so that the learning signal due to differences in style cancels out as a whole. We first test whether, on top of EntiGraph, a representative storing-side method that runs CPT on text synthesized from the corpus, our method adds accuracy efficiently. On QuALITY, the reading-comprehension QA benchmark on which EntiGraph was evaluated, SD-DPO exceeds a baseline we CPT on EntiGraph's synthetic data from the same base model and evaluate with the same procedure. The training tokens this requires are a few dozen times fewer than the additional CPT needed for the same gain. For knowledge updating, the main goal of this work, we use AToKE, a knowledge-editing benchmark for facts that change over time. There, SD-DPO reaches an overall accuracy of 0.982 and answers with the new or the old fact according to the queried period.
Comments23 pages, 3 figures, 13 tables