发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLM智能体有效构建记忆难的问题,提出AttriMem框架,通过用token级对最终答案贡献的局部奖励增强全局结果奖励,以学习记忆构建策略,实验表明该框架性能优于多种基线,具有泛化性且稳定了RL优化。
AI 中文摘要
有效的记忆对语言模型智能体至关重要,但有效构建记忆仍具有挑战性。记忆构建策略决定随着交互积累提取、存储、更新、压缩或丢弃哪些信息。启发式记忆方法依赖主观、特定任务规则,可能与下游目标不一致并限制跨任务适应性。基于强化学习(RL)的方法从任务反馈中学习,但主要使用结果或模块级奖励,存在细粒度信用分配瓶颈。我们提出AttriMem,一种用于通过RL学习记忆构建策略的归因引导过程反馈框架。AttriMem用从token级对最终答案的贡献得出的局部奖励增强全局结果奖励。在长时对话问答实验表明,AttriMem优于基于检索、启发式和基于RL的基线,在基准和答案模型间具有泛化性,稳定了RL优化。
英文摘要
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with downstream objectives and limit cross-task adaptability. RL-based methods, by contrast, learn from task feedback but mainly use outcome- or module-level rewards. These coarse signals indicate task success but cannot identify which intermediate memory contents support the final answer, creating a fine-grained credit-assignment bottleneck. However, constructing such process feedback is prohibitively difficult because intermediate memory decisions lack unique ground-truth targets, while the appropriate credit varies with the agent's uncertain reasoning trajectory and therefore cannot be specified in advance. We propose AttriMem, an attribution-guided process-feedback framework for learning memory-construction policies with RL. AttriMem augments the global outcome reward with local rewards derived from token-level contributions to the final answer. Experiments on long-horizon dialogue question answering show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization.