arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推荐反馈如何演化智能体记忆?

How Can Recommendation Feedback Evolve Agent Memory?

Shanwen Mao, Mingming Li, Hao Zhang, Zhiheng Li, Yige Wang, Penghua Yu, Junxiong Zhu

arXiv 2609.37544首次发表:更新:

发表机构

Harbin Institute of Technology; Alibaba Group; Institute of Automation, Chinese Academy of Sciences(哈尔滨工业大学; 阿里巴巴集团; 中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对推荐反馈延迟和噪声导致记忆演化困难的问题,提出TIDE框架,利用延迟反馈进行信用分配并演化记忆,在电商智能体上离线MEG提升7.75个百分点,在线UCTR和激活率显著提升,并在延迟标签基准上取得最优性能。

AI 中文摘要

内容生成智能体持续接收来自推荐系统的曝光、点击、转化和负面反馈,这些信号为记忆演化提供了真实世界的结果信号。然而,这些信号存在延迟和噪声,受到受众构成、展示位置和推荐策略的混淆,并且可能由多个记忆的共同影响所致,使得准确归因变得困难。现有方法主要依赖即时反馈或语义检索,因此难以可靠地将推荐结果转化为记忆适应性。为应对这一挑战,我们提出了TIDE(轨迹引导的定向记忆演化),一种由延迟推荐反馈驱动的外部记忆演化框架。我们进一步引入了记忆演化增益(MEG),用于衡量演化后的记忆相对于无记忆基线在严格未来任务上的效用提升。TIDE将记忆视为容量受限的经验群体:时间和语义信用分配估计上下文适应性,而责任信用根据生成过程中引用的记忆分配结果信号。这些信号随后用于强化、交叉、变异或驱逐记忆。在一个电子商务会员营销内容生成智能体上,TIDE在离线时间回放中实现了+7.75个百分点的MEG,并在在线A/B测试中显著提高了独特点击率(UCTR)和激活率。在一个延迟标签基准上,TIDE在比较方法中实现了最低的平均绝对误差(MAE)和均方根误差(RMSE)以及最高的MEG,证明了其有效性。

英文摘要

Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may result from the combined influence of multiple memories, making accurate attribution difficult. Existing methods rely primarily on immediate feedback or semantic retrieval and therefore struggle to reliably translate recommendation outcomes into memory fitness. To address this challenge, we propose TIDE (Trajectory-Informed Directed Memory Evolution), an external memory evolution framework driven by delayed recommendation feedback. We further introduce Memory Evolution Gain (MEG), which measures the utility improvement of evolved memory over a no memory baseline on strictly future tasks. TIDE treats memory as a capacity-constrained population of experiences: temporal and semantic credit assignment estimates contextual fitness, while responsibility credit distributes outcome signals according to the memories referenced during generation. These signals are then used to reinforce, crossover, mutate, or evict memories. On an e-commerce membership marketing content-generation agent, TIDE achieves a +7.75-percentage-point MEG in offline temporal replay and significantly improves both unique click-through rate (UCTR) and activation rate in an online A/B test. On a delayed-label benchmark, TIDE achieves the lowest mean absolute error (MAE) and root mean squared error (RMSE) and the highest MEG among the compared methods, demonstrating its effectiveness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑