arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

坚持你所知:知识对齐的监督微调研究

Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning

Arthur Becker, Jakob Kemmler, David Thulke, Christine Schäfer, Christian Dugast, Hermann Ney

arXiv 2608.30987首次发表:更新:

发表机构

AppTek GmbH; RWTH Aachen University(AppTek有限公司; 亚琛工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出知识对齐的监督微调方法,引入证据重写、回忆重写两个变体,实验显示其可减少事实幻觉、保留通用能力,回忆重写效果最优且改善弃权行为,证实超知识SFT目标会引发幻觉。

AI 中文摘要

监督微调(SFT)用于训练基础语言模型以模仿目标响应,而这些目标可能需要基础模型尚未稳健内化的知识。我们将此视为幻觉的一个来源,并将一组缓解方法定义为「知识对齐的SFT」:将SFT训练目标约束在基础模型的参数化知识范围内。在统一设置下,我们比较了现有的基于生成和基于估计的知识对齐方法,并引入了两个新变体:证据重写(使用外部证据验证基础模型的生成结果)和回忆重写(仅保留基础模型能一致回忆的断言)。使用Qwen 3 4B和OLMo 3 7B开展的实验显示,知识对齐的SFT可减少WildHalu和Biography数据集上的事实幻觉,同时基本保留通用能力。回忆重写带来最强的事实性提升,并改善了UnknownBench上的弃权(不执行)行为。这证实了超出基础模型知识范围的SFT目标会驱动幻觉行为。

英文摘要

Supervised fine-tuning (SFT) trains a base language model to imitate target responses, and these targets may require knowledge the base model has not robustly internalized. We study this as a source of hallucinations and frame a group of mitigation methods as \emph{knowledge-aligned SFT}: constraining SFT training targets to the base model's parametric knowledge. Under a unified setup, we compare existing generation-based and estimation-based knowledge-alignment methods and introduce two new variants: Evidence Rewrite, which verifies base-model generations using external evidence, and Recall Rewrite, which retains claims only when they can be consistently recalled by the base model. Experiments with Qwen 3 4B and OLMo 3 7B show that knowledge-aligned SFT can reduce factual hallucinations on WildHalu and Biography while largely preserving general capabilities. Recall Rewrite yields the strongest factuality gains and improves refusal behavior on UnknownBench. It thereby confirms that SFT targets beyond the base model's knowledge drive hallucination behavior.

CommentsAccepted to EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑