arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

反思还是重新生成?为什么大型语言模型(LLM)的修正会失败,而人类的修正却能成功

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds

Yefan Tao, Gerald Friedland, Madhusudhanan Chandrasekaran, Luyang Kong

arXiv 2607.28908首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出人类-LLM反思框架(HRF)对比人类与LLM的修正,发现LLM反思在客观任务增益近零、主观任务增益为负,本质是条件重新生成,而非真正的错误驱动修正。

AI 中文摘要

反思,即重新审视并修正先前推理的能力,是人类改进答案的核心。人们越来越多地提示大型语言模型(LLM)进行“反思”,但这是否与人类的修正相似仍不清楚。我们引入人类-LLM反思框架(HRF),这是一个受控的两轮协议,在自我、同行和跨智能体设置下,在相同条件下比较人类和LLM的修正。使用基于每次迭代交叉熵减少的信息论分析,我们发现LLM反思的两种失败模式:在具有有限答案空间的客观任务上,反思产生的信息增益接近零(ΔI≈0),表现为与重新采样无法区分的中性重新生成;在主观任务上,它产生显著的负增益(ΔI<0),使预测偏离目标。相比之下,人类的修正会在两种设置下都产生正增益。跨智能体实验将失败定位在修正步骤而非输入质量:LLM即使对高质量的人类响应也会进行降级。诊断分析(基于第一次通过正确性的修正,以及针对随机重排基线的神谕引导修正)显示,哪个子步骤占主导因任务和模型而异,而非简化为单一机制:自我错误检测在客观多项选择任务上存在,但在主观任务上较弱;在神谕错误信号下的恢复,对某些模型超过基线,对另一些模型则低于基线。统一的解释是结构性的:没有外部信息时,基于自身条件的修正无法减少关于目标的不确定性,因此LLM的反思更应被理解为条件重新生成,而非真正的错误驱动修正。

英文摘要

Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to "reflect," yet whether this resembles human revision remains unclear. We introduce the Human-LLM Reflection Framework (HRF), a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings. Using an information-theoretic analysis based on per-iteration cross-entropy reduction, we find two failure modes of LLM reflection. On objective tasks with finite answer spaces, reflection yields near-zero information gain (Delta I approx 0), behaving as neutral re-generation indistinguishable from re-sampling. On subjective tasks, it yields significant negative gain (Delta I < 0), moving predictions away from the target. Human revision, by contrast, yields positive gain in both settings. Cross-agent experiments localize the failure to the revision step, not input quality: LLMs degrade even high-quality human responses. Diagnostic analyses (revision conditioned on first-pass correctness, and oracle-guided revision against a random-reshuffle baseline) show that which sub-step dominates varies by task and by model rather than reducing to a single mechanism: self-error detection is present on objective multiple-choice tasks but weak on subjective ones, and recovery under an oracle error signal exceeds the baseline for some models and falls below it for others. The unifying account is structural: without external information, self-conditioned revision cannot reduce uncertainty about the target, so LLM reflection is better understood as conditioned re-generation than as genuine error-driven revision.

Comments20 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑