arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36435cs.CL

MemFold:通过在线策略优化学习用于长上下文个性化的紧凑软记忆

MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

Jingxuan Wu, Yuzhe Yang, Yiqiao Huang, Chengzhi Liu, Qingni Wang, Chengxuan Qian, Shutong Wu, Jiawei Zhang, Xin Eric Wang

首次发表
浏览论文内容

中文总结 AI 辅助

MemFold通过在线策略优化将查询条件化文本记忆压缩为固定数量的连续向量,结合组相对奖励和置信度门控蒸馏训练阅读器,在长上下文个性化任务上取得最优性能并实现跨域迁移。

中文摘要 AI 辅助

一个在长时间跨度内服务同一用户的助手必须根据用户所透露的信息来回答:哪些偏好仍然有效,哪些已被修订,以及现在适用哪些约束。保留这些信息与根据这些信息行动并不相同,而这两者通常被当作同一件事来优化。将信息保留为文本会使阅读器的输入随保留的历史记录增长,而将其压缩为固定数量的潜在向量则限制了接口,但通常训练目标是重建文本或模仿参考答案,这两者都是在阅读器从未产生的序列上评分的。我们提出了MemFold,它通过所支持的行为来优化固定预算的软记忆。一个查询条件化的文本记忆被压缩成K个连续向量,形成阅读器的记忆接口,然后阅读器在其自身的轨迹上接受两种互补信号的训练:用于任务结果的组相对奖励,以及置信度门控的在线策略蒸馏,其中冻结的文本记忆教师模型在文本记忆下重新评分学生采样得到的令牌。教师模型从不被采样,因此监督保持在学生当前分布上,且不增加自回归解码;在推理时它被完全移除。在三个Qwen骨干网络上,MemFold在PersonaMem-32K和PersonaMem-128K上达到了我们测量的最高准确率,且差距在更长的历史长度下扩大,并无需目标领域训练即可迁移到PrefEval和LongMemEval。消融实验表明,大部分任务增益来自奖励项,教师信号带来较小的额外增益,记忆干预显示阅读器依赖于其软记忆的实例特定内容。

英文摘要

An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were. Keeping the information as text makes the reader's input grow with the retained history, while compressing it into a fixed number of latent vectors bounds the interface but is typically trained to reconstruct text or imitate reference answers, both of which are scored on sequences the reader never produced. We present MemFold, which optimizes a fixed-budget soft memory by the behavior it supports. A query-conditioned textual memory is compressed into K continuous vectors that form the reader's memory interface, and the reader is then trained on its own rollouts under two complementary signals: group-relative rewards for task outcomes, and confidence-gated on-policy distillation in which a frozen textual-memory teacher re-scores the student's sampled tokens under the textual memory. The teacher is never sampled from, so supervision stays on the student's current distribution and adds no autoregressive decoding; at inference it is removed entirely. Across three Qwen backbones, MemFold attains the highest accuracy we measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length, and transfers to PrefEval and LongMemEval without target-domain training. Ablations attribute most of the task gain to the reward term and a smaller additional gain to the teacher signal, and memory interventions show that the reader depends on the instance-specific content of its soft memory.

发表机构

  • University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • Harvard University(哈佛大学)
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

↑