arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DRIFTLENS: 测量个性化语言模型中记忆引发的推理漂移

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

Xi Fang, Weijie Xu, Yingqiang Ge, Yuhui Xu, Stephanie Eckman, Chandan K. Reddy

arXiv 2607.02374首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DRIFTLENS框架,通过将推理步骤映射到价值类别并比较有无用户属性记忆的轨迹,量化个性化语言模型中的推理漂移,发现记忆会引发中等至大的推理变化,且现有后训练方法仅部分缓解。

AI 中文摘要

个性化改变了模型对用户所说的话;我们表明它也可以改变用于证明响应的推理轨迹。现代LLM通过存储用户属性、偏好和先前上下文,然后将这些信息注入未来提示来个性化交互。我们研究这种记忆是否重塑了对开放性问题(不存在单一真实答案)的推理。为了量化这种效应,我们引入了DRIFTLENS,一个无真实答案的框架,它将每个表达的推理步骤映射到一个价值类别,并测量问题在无记忆轨迹与注入用户属性记忆下的轨迹之间的差异。我们首先验证DRIFTLENS能够区分无内容的语用噪声和实质性的推理变化。在四个LLM和10个用户属性类别(包括年龄、职业和残疾)中,用户属性记忆在每个模型的语用噪声基准之上引发了中等至大的推理漂移,即使最终答案保持流畅、相关且合理。然后,我们评估基于GRPO和DPO的后训练方法以减少漂移。两者都减少了漂移,但都没有统一占优;对下游能力、有用性和指令遵循的影响取决于模型和奖励。这些结果表明,记忆引发的推理漂移是个性化语言模型的一个可测量且仅部分缓解的失败模式。

英文摘要

Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question's no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model's pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑