发表机构
Peking University; University of Chinese Academy of Sciences; Fullive-AI; Nanyang Technological University; Beijing University of Posts and Telecommunications; Beihang University; Beijing Institute of Technology, Zhuhai; Northeastern University; University of Illinois Chicago(北京大学; 中国科学院大学; Fullive-AI; 南洋理工大学; 北京邮电大学; 北京航空航天大学; 北京理工大学珠海学院; 东北大学; 芝加哥大学伊利诺伊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Hindsight Memory-PRM,利用轨迹审计轨迹训练记忆效用评判器并确定代理奖励,在LoCoMo、LongMemEval数据集上的性能优于基线系统,实现高效的长视野LLM智能体记忆管理监督。
AI 中文摘要
长视野大语言模型(LLM)智能体的记忆操作难以监督:操作执行时其价值不可观测。但这类操作具有特殊性——它们会在轨迹中留下机器可读取的证据:检索命中数和回答时刻的引用。事后记忆-PRM(Hindsight Memory-PRM)两次利用该审计轨迹:离线训练操作条件下的记忆效用评判器,在线阶段则针对每个探测样本执行检索、引用操作及一次受控的删除后重新回答,以此确定经干预校准的条目级存在信用,该信用沿版本链传播作为动作级代理奖励,无需每操作的人工标签,也无需蒙特卡洛续演。在保留的LoCoMo数据集上,固定共享读取器下的本地8B策略达到77.5%的性能,超越其API教师(65.1%)及所有复现的外部系统,且上下文长度仅为Mem0官方运行点的八分之一;在LongMemEval上,性能达79.0%。 ablation实验表明,性能提升源于因果校准而非信号密度,且该策略收敛至多版本记忆组织,其增益无任何经测试的开环基线可复现。
英文摘要
Memory operations of long-horizon LLM agents are hard to supervise: an operation's value is unobservable when it is taken. But they are special -- they leave machine-readable evidence in the trajectory: retrieval hits and answer-time citations. Hindsight Memory-PRM exploits this audit trail twice: offline to train an operation-conditioned memory-utility critic, and online, where retrievals, citations, and one controlled deletion-and-reanswer per probe settle an intervention-calibrated entry-level presence credit, propagated along version chains as an action-level proxy reward -- no per-operation human labels, no Monte-Carlo replay of continuations. On held-out LoCoMo a local 8B policy reaches 77.5% under a fixed shared reader, surpassing its API teacher (65.1%) and all reproduced external systems, at one eighth the context of Mem0's official operating point; on LongMemEval, 79.0%. Ablations attribute the gain to causal calibration rather than signal density, and the policy converges to a multi-version memory organization whose gains no tested open-loop baseline reproduces.