arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PrefReward:学习用于个性化文本生成的用户偏好矩阵

PrefReward: Learning User Preference Matrix for Personalized Text Generation

Yue Wu, Chengbing Wang, Yimeng Bai, Xiaoyan Zhao, Yang Zhang, Fuli Feng

arXiv 2607.21067首次发表:更新:

发表机构

University of Science and Technology of China; The Chinese University of Hong Kong; National University of Singapore(中国科学技术大学; 香港中文大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型个性化生成中存在的问题,提出PrefReward框架,通过提取用户特定偏好矩阵并将其作为奖励信号指导生成,实验证明该框架在生成质量和个性化可解释性上优于基线。

AI 中文摘要

大语言模型在利用用户历史和上下文线索生成个性化内容方面展现出卓越能力。然而,多数现有个性化方法依赖模型参数内的隐式表示,难以解释用户特定偏好或有效处理长上下文依赖。为应对这些挑战,我们提出PrefReward,一种新颖的偏好感知生成框架,通过结构化偏好矩阵明确建模用户风格并将其作为奖励信号整合到解码过程中。PrefReward由两个阶段组成:提取总结个体风格倾向的用户特定偏好矩阵;使用该矩阵通过基于KL散度的奖励函数指导生成。在LongLaMP数据集上的实验表明,PrefReward在生成质量和个性化可解释性方面均优于非个性化和基于检索的基线。

英文摘要

Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. However, most existing personalization approaches rely on implicit representations within model parameters, making it difficult to interpret user-specific preferences or effectively handle long-context dependencies. To address these challenges, we propose PrefReward, a novel preference-aware generative framework that explicitly models user styles through a structured preference matrix and integrates it into the decoding process as a reward signal. PrefReward consists of two stages: (1) extracting a user-specific preference matrix that summarizes individual stylistic tendencies, and (2) using the matrix to guide generation via a KL-divergence-based reward function. Experiments on the LongLaMP dataset show that PrefReward outperforms non-personalized and retrieval-based baselines in both generation quality and personalization interpretability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑