arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

事后记忆-PRM:通过可审计的事后信用监督记忆管理

Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit

Haoxuan Jia, Yang Liu, Yingguang Yang, Yancheng Chen, Chongyang Zhang, Hao Zheng, Qian Li, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Hao Peng, Junyu Lu, Du Cheng, Philip S. Yu, Bin Chong

arXiv 2608.29605首次发表:更新:

发表机构

Peking University; University of Chinese Academy of Sciences; Fullive-AI; Nanyang Technological University; Beijing University of Posts and Telecommunications; Beihang University; Beijing Institute of Technology, Zhuhai; Northeastern University; University of Illinois Chicago(北京大学; 中国科学院大学; Fullive-AI; 南洋理工大学; 北京邮电大学; 北京航空航天大学; 北京理工大学珠海学院; 东北大学; 芝加哥大学伊利诺伊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Hindsight Memory-PRM,利用轨迹审计轨迹训练记忆效用评判器并确定代理奖励,在LoCoMo、LongMemEval数据集上的性能优于基线系统,实现高效的长视野LLM智能体记忆管理监督。

AI 中文摘要

长视野大语言模型(LLM)智能体的记忆操作难以监督:操作执行时其价值不可观测。但这类操作具有特殊性——它们会在轨迹中留下机器可读取的证据:检索命中数和回答时刻的引用。事后记忆-PRM(Hindsight Memory-PRM)两次利用该审计轨迹:离线训练操作条件下的记忆效用评判器,在线阶段则针对每个探测样本执行检索、引用操作及一次受控的删除后重新回答,以此确定经干预校准的条目级存在信用,该信用沿版本链传播作为动作级代理奖励,无需每操作的人工标签,也无需蒙特卡洛续演。在保留的LoCoMo数据集上,固定共享读取器下的本地8B策略达到77.5%的性能,超越其API教师(65.1%)及所有复现的外部系统,且上下文长度仅为Mem0官方运行点的八分之一;在LongMemEval上,性能达79.0%。 ablation实验表明,性能提升源于因果校准而非信号密度,且该策略收敛至多版本记忆组织,其增益无任何经测试的开环基线可复现。

英文摘要

Memory operations of long-horizon LLM agents are hard to supervise: an operation's value is unobservable when it is taken. But they are special -- they leave machine-readable evidence in the trajectory: retrieval hits and answer-time citations. Hindsight Memory-PRM exploits this audit trail twice: offline to train an operation-conditioned memory-utility critic, and online, where retrievals, citations, and one controlled deletion-and-reanswer per probe settle an intervention-calibrated entry-level presence credit, propagated along version chains as an action-level proxy reward -- no per-operation human labels, no Monte-Carlo replay of continuations. On held-out LoCoMo a local 8B policy reaches 77.5% under a fixed shared reader, surpassing its API teacher (65.1%) and all reproduced external systems, at one eighth the context of Mem0's official operating point; on LongMemEval, 79.0%. Ablations attribute the gain to causal calibration rather than signal density, and the policy converges to a multi-version memory organization whose gains no tested open-loop baseline reproduces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑