发表机构
National University of Singapore; Tsinghua University; Johns Hopkins University; Nanyang Technological University(新加坡国立大学; 清华大学; 约翰斯·霍普金斯大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出PHRBench基准,评估18个LLM在四个领域中的幻觉后推理行为,发现成功恢复罕见且与信念更新相关,提示属性可预测恢复,AUROC达0.847。
AI 中文摘要
幻觉信息可以在多阶段大型语言模型系统中传播,并成为后续推理上下文的一部分。现有的关于幻觉后推理(PHR)的研究主要刻画最终结果的变化和聚合推理动态,而忽略了模型在响应层面如何解决幻觉前提。在这项工作中,我们引入了PHRBench,一个用于行为结构化PHR的受控基准,涵盖四个领域和18个大型语言模型。PHRBench通过幻觉遵从、幻觉避免和启发式修正独立于最终答案正确性来刻画每个推理轨迹,并将有见地的轨迹定义为最终达到正确答案的成功修正。在4820个受控实例中,我们发现成功恢复仍然相对罕见,并且与推理轨迹上更频繁的信念更新相关。我们进一步发现,幻觉提示的属性包含对成功恢复的实质性预测信号,一个轻量级预测器达到了0.847的AUROC。这些发现提供了幻觉后推理的行为视角,刻画了大型语言模型如何解决错误上下文以及成功恢复何时可能发生。
英文摘要
Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.