arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38275cs.DCcs.AI

当正确记忆出错时:对LLM智能体中持久化记忆使用的模糊测试

When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents

  • Binghamton University, State University of New York(纽约州立大学宾汉姆顿分校)
  • The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuqiao Meng, Luoxi Tang, Yingxue Zhang, Yuchen Yang, Zhaohan Xi

AI总结:

针对LLM智能体中持久化记忆使用错误,提出U-Fuzz模糊测试框架,通过变异查询和记忆状态,系统性地发现记忆使用失败,并在多种架构下验证其有效性。

AI中文摘要:

持久化记忆帮助LLM智能体在长时间交互中携带信息,但当查询发生变化或记忆状态演变时,即使正确的记忆也可能被错误地使用。现有工作主要研究记忆内容错误或评估固定测试用例,使得记忆使用失败难以被系统性地发现。我们将此问题形式化为一个模糊测试问题,并将此类失败归类为查询相关失败和记忆状态失败。随后,我们开发了U-Fuzz,它以记忆检查点作为测试种子,在明确的变异义务下对查询或记忆状态进行变异,验证每个变异体,并利用观察到的记忆行为来指导迭代测试,同时将失败标签排除在搜索之外。我们在多个记忆系统上针对多种模糊测试基线评估了U-Fuzz,并进一步测试了基于API的LLM的仅输出设置,其中记忆检索是隐藏的。在这些设置中,U-Fuzz始终发现更多已确认的记忆使用失败,表明其搜索在不同记忆架构下以及仅可观察最终响应时均保持有效。

英文摘要:

Persistent memory helps LLM agents carry information across long interactions, but correct memory can still be used incorrectly when queries change or memory states evolve. Existing work mainly studies memory content errors or evaluates fixed test cases, leaving memory-use failures hard to discover systematically. We formulate this issue as a fuzzing problem and categorize such failures into query-related and memory-state failures. We then develop U-Fuzz, which starts from memory checkpoints as test seeds, mutates queries or memory states under explicit mutation obligations, validates each mutant, and uses observed memory behavior to guide iterative testing while keeping failure labels outside the search. We evaluate U-Fuzz across several memory systems against diverse fuzzing baselines, and further test an output-only setting with API-based LLMs where memory retrieval is hidden. Across these settings, U-Fuzz consistently uncovers more confirmed memory-use failures, showing that its search remains effective across different memory architectures and even when only final responses are observable.

↑