arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AttriMem:用于智能体记忆学习的归因引导过程反馈

AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction

Qinfeng Li, Yuntai Bao, Xinyan Yu, Hongze Chen, Yanming Liu, Huifeng Zhu, Yier Jin, Jintao Chen, Wenqi Zhang, Xuhong Zhang

arXiv 2607.21106首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对LLM智能体有效构建记忆难的问题,提出AttriMem框架,通过用token级对最终答案贡献的局部奖励增强全局结果奖励,以学习记忆构建策略,实验表明该框架性能优于多种基线,具有泛化性且稳定了RL优化。

AI 中文摘要

有效的记忆对语言模型智能体至关重要,但有效构建记忆仍具有挑战性。记忆构建策略决定随着交互积累提取、存储、更新、压缩或丢弃哪些信息。启发式记忆方法依赖主观、特定任务规则,可能与下游目标不一致并限制跨任务适应性。基于强化学习(RL)的方法从任务反馈中学习,但主要使用结果或模块级奖励,存在细粒度信用分配瓶颈。我们提出AttriMem,一种用于通过RL学习记忆构建策略的归因引导过程反馈框架。AttriMem用从token级对最终答案的贡献得出的局部奖励增强全局结果奖励。在长时对话问答实验表明,AttriMem优于基于检索、启发式和基于RL的基线,在基准和答案模型间具有泛化性,稳定了RL优化。

英文摘要

Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with downstream objectives and limit cross-task adaptability. RL-based methods, by contrast, learn from task feedback but mainly use outcome- or module-level rewards. These coarse signals indicate task success but cannot identify which intermediate memory contents support the final answer, creating a fine-grained credit-assignment bottleneck. However, constructing such process feedback is prohibitively difficult because intermediate memory decisions lack unique ground-truth targets, while the appropriate credit varies with the agent's uncertain reasoning trajectory and therefore cannot be specified in advance. We propose AttriMem, an attribution-guided process-feedback framework for learning memory-construction policies with RL. AttriMem augments the global outcome reward with local rewards derived from token-level contributions to the final answer. Experiments on long-horizon dialogue question answering show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑