发表机构
University of Science and Technology of China; Qwen Applications Business Group, Alibaba Group; Fudan University(中国科学技术大学; 阿里巴巴集团通义应用业务集团; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MemCalib基准评估LLM智能体记忆使用,发现模型存在过度使用或使用不足问题,并设计MemCalib-RL算法通过双向反事实信用分配优化,实现最佳整体性能并平衡两种偏差。
AI 中文摘要
智能体记忆的有效性最终取决于底层LLM是否给予上下文中的每条记忆对其响应适当程度的影响。然而,这一能力在很大程度上仍被忽视。为了评估这一能力,我们引入了MemCalib,一个基于现实记忆系统场景的基准,用于评估记忆使用并推进优化算法。在MemCalib测试集上的结果表明,前沿的开源和闭源模型难以适当地使用记忆。它们经常过度使用或使用不足记忆,而不是将每个命题的实际使用与其目标水平相匹配,导致有偏见的、低质量的响应。使用常见后训练算法(包括组相对策略优化和同策略自蒸馏)的实验进一步揭示了明确的方向性偏差:训练后的模型在一个方向上改进,而在另一个方向上恶化。因此,我们提出了MemCalib-RL,一种有序的双向反事实信用分配算法,该算法分离过度使用和使用不足信号,并通过精确原子消融将其信用定位到响应标记。跨模型家族和规模(Qwen3-8B、Ministral-3-8B-Instruct和Qwen3.5-35B-A3B)的结果表明,MemCalib-RL实现了最佳的整体性能,同时更好地平衡了过度使用和使用不足,其收益在外部基准评估中泛化到MemCalib之外。进一步的实验支持其设计选择和鲁棒性,并提供了对其训练动态的见解。
英文摘要
The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.