发表机构
Simon Fraser University; Microsoft; University of Illinois at Chicago(西蒙 Fraser大学; 微软; 伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对高质量标注数据稀缺致识别LLM生成输出中幻觉困难及现有方法局限,提出幻觉自我博弈框架,含检测器和生成器,经多轮训练优化,能在无外部监督下提升小型LLM性能。
AI 中文摘要
由于高质量标注数据稀缺,识别大语言模型(LLM)生成输出中的忠实幻觉仍具有挑战性。近期工作依赖先进LLM合成训练数据,但将生成器视为静态组件限制了检测器的迭代改进。为此,我们引入幻觉自我博弈(HSP)框架,使检测器能与进化生成器协同引导。HSP包含由同一基础模型初始化的检测器和生成器,检测器先在人工标注数据上微调,再作为奖励模型通过基于人工智能反馈的强化学习(RLAIF)训练生成器,进化后的生成器合成幻觉数据通过基于规则的强化学习进一步优化检测器。在RAGTruth基准测试和两个模型系列上的实验表明,该框架能在无外部监督下逐步提升小型LLM以匹配甚至超越先进LLM。
英文摘要
Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data. Recent work relies on advanced LLMs to synthesize training data, including rationales, labels, and hallucinated claims. However, these methods treat the generator as a static component, limiting iterative improvement of the detector. To address this limitation, we introduce Hallucination Self-Play (HSP), a novel framework that enables the detector to bootstrap with an evolved generator. HSP involves two roles initialized from the same base model, a detector that assesses the faithfulness of model outputs, and a generator that produces increasingly hard-to-detect hallucinated responses. Specifically, the detector is first fine-tuned on human-labeled data and then employed as a reward model to train the generator via reinforcement learning from AI feedback (RLAIF). In turn, the evolved generator synthesizes hallucination data to further optimize the detector through rule-based reinforcement learning. Experiments on RAGTruth and LLM-AggreFact across three model families demonstrate that the proposed framework can progressively enhance a small LLM to match or even outperform advanced LLMs without external supervision. Our code is available at https://github.com/maybenotime/Hallucination_Self-Play.
CommentsCOLM 2026