arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39473cs.AIcs.LG

超越柏拉图洞穴的阴影:通过反事实推理评估自主智能体中的虚假记忆

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning

  • Sydney AI Centre, The University of Sydney(悉尼大学悉尼人工智能中心)
  • Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所)
  • Australian Artificial Intelligence Institute, University of Technology Sydney(悉尼科技大学澳大利亚人工智能研究所)
  • Wuhan University(武汉大学)
  • School of Mathematics and Statistics, The University of Melbourne(墨尔本大学数学与统计学院)

机构由 AI 辅助整理,请以论文原文为准。

Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu

AI总结:

针对自主智能体在未见环境中的虚假记忆问题,提出无需训练的FAME框架,通过反事实推理测量概念漂移来评估记忆忠实性,在多个基准上显著优于基线。

AI中文摘要:

自主智能体日益依赖记忆来泛化到其训练环境之外。然而,智能体受限于其所见和所信,在未见环境中利用此类记忆可能在其内部信念中引入偏差。我们将这一现象形式化为“虚假记忆”,它可能源于虚假相关性、环境偏移和知识冲突。尽管其重要性,虚假记忆难以评估,因为它源于智能体的内部信念,且容易与普通泛化失败混淆。因此,我们提出FAME,一个无需训练的框架,通过反事实推理下智能体信念的演化来评估虚假记忆。具体而言,反事实场景揭示了在假设干预下,随着记忆的潜在概念偏移,信念如何变化;因此,测量由此产生的概念漂移提供了区分忠实记忆与虚假记忆的信号。此类概念可在答案生成前从智能体隐藏状态中估计,避免了对奖励设计或答案采样的需求。实证实验表明,仅监控答案往往无法检测虚假记忆,而FAME在虚假记忆设置中实现了76.2%至96.7%的AUROC,并在现实基准上优于最佳基线3.4%至23.3%,涵盖数学推理(GSM-Symbolic)、代码生成(GitChameleon)和复杂推理(BigBench-Hard)。我们进一步发布了相应的反事实模板,以促进未来对虚假记忆的研究。

英文摘要:

Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counterfactual scenarios reveal how beliefs change as the latent concept of memory shifts under hypothetical interventions; thus, measuring the resulting concept drift provides a signal for distinguishing faithful versus false memory. Such concepts can be estimated from agent hidden states before answer generation, avoiding the need for reward design or answer sampling. Empirical experiments reveal that simply monitoring answers often fails to detect false memory, while FAME achieves AUROCs of 76.2% - 96.7% across false-memory settings, and outperforms the best baseline by 3.4% - 23.3% across realistic benchmarks, spanning math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard). We further release corresponding counterfactual templates and facilitate future research on false memory.

↑