发表机构
Renmin University of China; Southeast University; Zhejiang University(中国人民大学; 东南大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM易受输入事实误导产生错误响应的问题,提出三阶段框架DEDUCE,结合新数据集与指标,在多基准上提升了Qwen、LLaMA等系列模型的鲁棒性与错误修正能力。
AI 中文摘要
大语言模型(LLM)在用户输入包含误导性前提时,常产生自信却事实错误的响应,我们将此现象归因于输入中的事实扰动。现有幻觉缓解方法通常假设用户输入可靠,忽略了此类事实错误会主动误导模型推理。为解决这一漏洞,我们提出DEDUCE,这是一个三阶段框架,可将LLM从被动响应者转变为主动错误修正者。DEDUCE分三个阶段运行:(1)通过细粒度事实提取与验证检测错误;(2)通过多视角审议制定修正策略;(3)纠正误解并提供可靠答案。我们还提出MisFactQA,一个包含不同程度事实错误的数据集,并提出用于评估模型鲁棒性的新指标。在TruthfulQA、FalseQA及我们的MisFactQA基准上的实验表明,DEDUCE显著提升了准确率与错误修正能力。在Qwen、LLaMA和Gemma系列上的一致增益,证实了其有效性与可扩展性。
英文摘要
Large language models (LLMs) frequently produce confident yet factually incorrect responses when user inputs contain misleading premises, a phenomenon we attribute to fact perturbations in the input. Existing approaches to hallucination mitigation typically assume reliable user inputs, overlooking how such factual errors can actively mislead model reasoning. To address this vulnerability, we propose DEDUCE, a three-stage framework that transforms LLMs from passive responders into proactive error correctors. DEDUCE operates in three stages: (1) detect errors through fine-grained fact extraction and verification; (2) devise correction strategies via multi perspective deliberation; and (3) correct misconceptions while delivering reliable answers. We also present MisFactQA, a dataset containing factual errors of varying degrees, and propose new metrics for evaluating model robustness. Experiments on TruthfulQA, FalseQA, and our MisFactQA benchmark demonstrate that DEDUCE significantly improves both accuracy and error correction capability. Consistent gains across Qwen, LLaMA, and Gemma families confirm its effectiveness and scalability.
CommentsAccepted to the Main Conference of EMNLP 2026