arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中的现实监测:随对话记忆而转变的自我认知

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

Saurabh Ranjan, Konstantina Sokratous, Brian Odegaard

arXiv 2607.23927首次发表:更新:

发表机构

University of Florida; University of Missouri(佛罗里达大学; 密苏里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型中现实监测能力,通过两个实验和六个模型发现来源归因依赖对话记忆结构,反馈揭示模型存在内部外部判断互换、置信度与正确性脱钩等问题,表明评估模型知识时追踪来源很重要。

AI 中文摘要

无法区分自身输出与用户所说内容的对话式人工智能会将自身错误当作用户提供的事实。在人类中,这种能力被称为现实监测,其失败与幻觉、妄想和虚构有关,但大语言模型是否具备此能力仍未得到检验。通过两个实验和六个大语言模型,我们发现来源归因取决于对话记忆的结构:在最小记忆需求下,自我生成内容的最高准确率在情节性延迟消除捷径后会转变为对外部项目的微弱优势。反馈揭示了两个问题:在一些模型中,内部和外部判断会互换;在另一些模型中,准确率提高而置信度与正确性脱钩,现有基准无法察觉这些分离。跨模型来看,这种模式涉及活跃参数计数而非总参数计数。这表明随着人工智能系统承担自主、多轮角色,评估它们知道什么是不够的:追踪知识来源可能同样重要。

英文摘要

A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough: tracking where that knowledge came from may matter equally.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑