幻觉-R1:面向事实一致性的鲁棒性导向改写生成
Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出HALLUCINATION-R1框架,通过两阶段优化生成语义忠实且具挑战性的改写,揭示并提升问答模型的事实一致性鲁棒性,实验验证其有效性与数据价值。
AI中文摘要:
事实幻觉通常被定义为不正确的事实输出。我们研究了一种由改写引发的幻觉场景,其中模型在原始形式下能正确回答事实性问题,但在语义等价的改写下却生成错误答案。这种不一致性揭示了语义不变性下潜在的事实不稳定性。然而,通用改写往往不足以作为面向鲁棒性的监督信号:近似复制的改写提供弱信号,而过度多样化的改写可能破坏语义等价性。在本文中,我们提出了HALLUCINATION-R1,一个面向鲁棒性的改写生成框架,学习生成语义忠实且具有鲁棒性挑战的改写,以用于事实一致性。通过两阶段优化,它首先稳定保持意义和多样化的改写,然后奖励能揭示下游问答模型中事实一致性退化的改写。在SimpleQuestions、PopQA和TruthfulQA上的实验表明,HALLUCINATION-R1实现了强一致性与多样性的权衡,并在多个模型家族和数据集上暴露了鲁棒性失败。进一步分析表明,这些失败不能归结为表面伪影或语义漂移,而是揭示了在保持意义的变化下非平凡的事实不稳定性。一项轻量级微调研究还表明,HALLUCINATION-R1生成的数据提高了在改写变化下的鲁棒准确性,表明其在面向鲁棒性训练中的实用性。我们的代码和模型在此https URL公开可用。
英文摘要:
Factual hallucination is commonly defined by incorrect factual outputs. We study a paraphrase-induced hallucination setting, where a model answers a factual question correctly in its original form but generates an incorrect answer under a semantically equivalent paraphrase. Such inconsistencies expose latent factual instability under semantic invariance. However, general-purpose paraphrases are often insufficient as robustness-oriented supervision: near-copy paraphrases provide weak signals, while overly diverse paraphrases may break semantic equivalence. In this paper, we propose HALLUCINATION-R1, a robustness-oriented paraphrase generation framework that learns to produce semantically faithful yet robustness-challenging paraphrases for factual consistency. Through two-stage optimization, it first stabilizes meaning-preserving and diverse paraphrasing, then rewards paraphrases that reveal factual consistency degradation in downstream QA models. Experiments on SimpleQuestions, PopQA, and TruthfulQA show that HALLUCINATION-R1 achieves a strong consistency--diversity trade-off and exposes robustness failures across multiple model families and datasets. Further analyses indicate that these failures are not reducible to surface-level artifacts or semantic drift, but reveal non-trivial factual instability under meaning-preserving variation. A lightweight fine-tuning study also shows that HALLUCINATION-R1-generated data improves robust accuracy under paraphrase variations, suggesting its utility for robustness-oriented training. Our code and models are publicly available at https://github.com/yuwenhan07/Hallucination-R1.