arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37311cs.AIcs.IR

ReMem:重新思考长上下文推荐代理中的感知与记忆

ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents

  • The Hong Kong Polytechnic University(香港理工大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan, Dacheng Tao

AI总结:

针对推荐代理在条目感知和长上下文推理上的不足,提出ReMem框架,结合OCR多模态感知与分块动态记忆,实现线性复杂度推理,并在三个任务上平均提升5.16%。

AI中文摘要:

近期的推荐代理(RecAgents)提供了一种有前景的替代方案,将推荐转变为主动的、用户侧的范式,其中生成式代理自主感知外部平台、推理用户偏好并执行决策。然而,现有的推荐代理仍存在两个关键局限:基于嘈杂且异构的条目页面的脆弱条目感知,以及处理扩展用户历史和多步交互轨迹时的低效长上下文推理。为解决这些挑战,我们提出了一种新颖的推荐代理框架,称为ReMem,它结合了基于OCR的多模态感知与随时间演变的动态记忆。ReMem不解析原始HTML,而是通过截图观察条目页面,并借助OCR工具提取结构化的多模态信息,从而实现更类人且与平台无关的感知机制。为了支持长时程偏好建模,ReMem进一步引入了分块顺序记忆更新策略,其中代理在处理任意长上下文时,以线性推理复杂度和有界上下文长度,选择性地维护一个固定大小的、包含信息性历史交互的记忆。这种设计使代理能够保留演变的用户偏好,而无需依赖外部记忆模块或干扰标准的自回归生成过程。为了增强动态记忆指令,我们进一步开发了一种多记忆GRPO变体,该变体将最终答案的优势传播到所有对最终响应有贡献的中间对话中。在三个数据集上的广泛实验表明,ReMem持续优于最先进的基线,在三个推荐代理任务(即搜索、排序和评判)上平均提升了5.16%。

英文摘要:

Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16\% across three recommendation agent tasks, namely searching, ranking, and judging.

↑