arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当记忆成为权威:在记忆整合边界处对权威崩溃进行基准测试

When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary

Qiuyang Zhan, Rui Zhang, Sheng Guo, Lepeng Zhao, Zhuotao Liu

arXiv 2608.01679首次发表:更新:

发表机构

Tsinghua University; East China Normal University(清华大学; 华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出AuthMem-Bench基准,发现智能体记忆整合存在权威崩溃问题,自动保留权威标签可大幅降低未授权动作率,证明记忆适配需同时保留权威信息。

AI 中文摘要

持久内存使(自演化)大语言模型(LLM)智能体能够通过将异构交互历史整合为可复用的事实、偏好、观察结果和规则,从而在不同任务间进行适配。然而,整合也设定了一个隐含的权威边界:它决定了存储的信息后续是否可作为用户事实、已证实的观察结果或长期指令被使用。我们识别出权威崩溃现象,即整合过程保留了某一主张,却抹去了对其授权使用的来源约束,导致存储的记忆所隐含的权威超出其来源允许的范围。我们推出AuthMem-Bench,这是一个受控配对基准,在固定核心主张和下游任务的同时,仅改变来源权威。它评估写入时崩溃、下游授权错误以及自动权威保留情况。在基于广泛使用的智能体记忆系统构建的7种整合器和7种LLM骨干模型中,我们在49种评估配置里的48种中观察到了权威崩溃。在受控的动作落地评估中,没有权威元数据的崩溃记忆会产生50.3%的平均未授权动作率。在端到端评估中,自动预测并持久化的权威标签将观察到的未授权动作率从16.9%降至0.0%,而良性任务成功率基本保持不变。这些发现表明,记忆驱动的适配不仅必须保留所学内容,还必须保留其可被复用的权威。

英文摘要

Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts, preferences, observations, and rules. Yet consolidation also imposes an implicit authorization boundary: it determines whether stored information may later be consumed as a user fact, an attested observation, or a standing instruction. We identify authority collapse, in which consolidation preserves a claim while erasing the source constraints governing its authorized use, causing the stored memory to imply greater authority than its source permits. We introduce AuthMem-Bench, a controlled paired benchmark that holds the focal claim and downstream task fixed while varying only source authority. It evaluates write-time collapse, downstream authorization errors, and automatic authority preservation. Across seven consolidators based on widely used agent-memory systems and seven LLM backbones, we observe authority collapse in 48 of 49 evaluated configurations. In a controlled action-grounded evaluation, collapsed memories without authority metadata yield a mean unauthorized-action rate of 50.3%. In an end-to-end evaluation, automatically predicted and persisted authority labels reduce the observed unauthorized-action rate from 16.9% to 0.0%, while benign task success remains essentially unchanged. These findings show that memory-driven adaptation must preserve not only what was learned, but also the authority under which it may be reused.

Comments38 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑